Note Wisdom
Notes from Sam Altman's Stanford CS153 session, where he argues that scale itself is an underexplored design bet, walks through how ChatGPT and Codex were discovered rather than planned, and frames intelligence as a new utility. Includes candid listener skepticism on the points that went unproven.
Institution: Stanford
Original Course: Stanford CS153 Frontier Systems: Scale, AGI, and the Future of Everything
Instructor Bio: This session is co-taught by **Anjney Midha** and **Michael Abbott**, co-founders of AMP PBC and co-instructors of Stanford CS 153: Frontier Systems. Anjney Midha is a Stanford alumnus who previously served as a partner at Andreessen Horowitz (a16z) and held early leadership roles at Discord. He specializes in frontier AI ecosystem development and AI-native company building. Michael Abbott brings decades of engineering leadership experience from General Motors, Apple, Twitter, and Microsoft, where he oversaw global-scale cloud infrastructure and consumer platforms serving hundreds of millions of users. His expertise spans scalable system design, operational discipline, and infrastructure engineering.
Course Description: This lecture explores the exponential scaling dynamics of frontier AI and their transformative impact across every sector of the global economy. It examines the scaling laws of compute, data, and model capability, the technical and commercial trajectory toward artificial general intelligence, and the second-order effects of AI on work, innovation, and society. It frames both the unprecedented opportunities and systemic risks that come with building technology systems that scale at rates never seen before in human history.
Sam Altman came back to Stanford for a session of CS153, a systems class whose instructors built it partly out of nostalgia for his own 2014 course on starting startups. One of them took that class as an undergrad and says so up front, which sets the tone: this is less a lecture than a long, friendly interrogation. The interviewer pushes, reframes, and occasionally tries to make Sam spicier than he wants to be.
Nearly every thread in the conversation runs back to one word: scale. Not scale as in "bigger models," but scale as a general-purpose bet — the claim that pushing anything past the size anyone has tried tends to hand you capabilities you could not have predicted, and that almost nobody pushes hard enough. He says, twice and without embarrassment, that he cannot explain why this works. That admission is both the most interesting thing in the room and the weakest link in the argument, and I'll come back to it.
The warm-up is about whether the 2014 startup curriculum still holds. Sam's answer is that it mostly doesn't, and the reason is leverage: a modest token budget now buys you output that used to require something like a hundred-person engineering team. Ambition ceiling, speed, and the number of things you can run in parallel all moved at once (2:35).
When asked what problem he'd assign students at the end of a quarter, he refuses the premise. If an idea is obvious enough for him to hand out, it's obvious to a lot of other people. Back when OpenAI started, he says, they were generously one of maybe four serious AI efforts on the planet, and that kind of obscurity is what you want. He's fairly explicit that the students in the room are better positioned than he is to spot the next non-obvious, multi-trillion-dollar thing, because his head is full of OpenAI.
Then the actual thesis (5:01). He frames it as an observation he can't justify: the most interesting phenomena he's watched in his career all involved either properties that only appear at size, or returns that kept arriving long after consensus said they'd stop. Model scaling laws are the obvious case, but he lists others — more smart people on one problem, the economies a company gets from getting bigger.
The example he develops is Y Combinator. The conventional wisdom among smart people was that YC had gotten too big and should shrink back to something like ten companies per batch; the argument was that the best companies are obvious anyway, and funding the rest adds little. It was tempting, he says, partly because it would have been much less work. What that view missed is that a lot of the value came from network effects within a batch, which simply didn't exist at a batch size where nothing could compound. Nobody had tried funding startups at that volume in that way, so nobody had stumbled onto the effect.
The version he wants you to use is a heuristic: find something that already works in an interesting way at small size, push it to a scale nobody has attempted, and more often than not that turns out to have been the right call. He adds that most people don't do this nearly enough. He notes the same dynamic in deep learning, where most of the field's most decorated researchers thought continued scaling was barely a scientific result, and in founders who sense something interesting might happen if they scaled but are vaguely worried.
Here's where I got stuck. He says outright that he has no theory he finds satisfying, and that recommending it anyway makes him a little nervous — which I respect, genuinely. But the evidence offered is a recollection of cases that worked. We never get a denominator. There's no mention of the times someone pushed scale past what anyone had tried and it just broke, and given that he spends the next ten minutes explaining how much breaks when you scale, those cases must exist. "Most people don't do it enough" might be true, or it might be what survival looks like from the inside.
The interviewer asks what stops people, and the answer is refreshingly concrete. Systems don't degrade linearly as they grow; they break at an accelerating and unpredictable rate. Anything genuinely large is always a little bit broken. And there is a permanent supply of very smart people telling you not to be so ambitious, to try the smaller version first (8:27).
What I found useful is how he says to handle that: don't argue with "this is too ambitious" as a blob. Separate the objections and answer them one at a time. In the model-scaling case he names three. Technical — could anyone even orchestrate a single training run across tens of thousands of accelerators, which nobody had contemplated and which needed layers of engineering talent. Financial — what the capital requirement was, what the business would eventually look like, how to think about the risk. Cultural — researchers asking why all that compute should go into one project instead of being spread across a portfolio.
The human layer gets its own answer (10:42). Three things: a clear goal, a clear plan for reaching it, and a clear account of how decisions will get made along the way. He illustrates it with the decision to bet on scaling deep learning — stated plainly, including the part where if they were wrong they'd simply fail, and including the picture of the world they believed would exist if it worked.
Underneath that is a claim about cognition that he repeats later in a different form: people did not evolve to think in exponentials. Scaling curves, revenue, organizational complexity — all of it gets underestimated because linear intuition is what we have. His practical fix is unglamorous: it takes a lot of time, sitting with people, reasoning from first principles about what the continuation would actually imply.
This section would benefit from an example of something specific that broke. "It's always a little bit broken" is vivid but I wanted to know which part, and how they decided what to leave broken.
The best stretch of the session is the origin story, because it quietly demonstrates the earlier point about not assigning problems.
After GPT-3, they needed revenue to fund computers that cost a billion dollars and up, and they had a model that made a great demo but no product. Several attempts failed. So they did the honest thing and shipped an API in the summer of 2020, essentially outsourcing product discovery to developers. For about a month, nothing. Then it went viral on Twitter when several people independently got it to do something cool on the same day.
What follows is a useful corrective to hindsight. He insists the models were shockingly bad by present standards relative to the excitement they generated, and the only business that actually worked at any scale was copywriting, which nobody found thrilling. The signal was elsewhere: developers who couldn't make the API work for their product were using their API keys to chat with it.
So they built the thing users were already improvising. GPT-4 was finished; they had a 3.5 model staged in between, plus a post-training approach that made instruction-following good enough to hold a conversation. The chatbot was released as a research demo, meant to convince other people to build chat products and pay for the API.
Then the YC pattern-matching kicks in. His rule: when something is growing fast and still isn't very good, you have a guaranteed hit. He describes five days of traffic spiking and falling, with each spike written off as a hype cycle, each one higher than the last. By day four or five he called an all-hands and declared an emergency — the good kind — because they had to build a company and a product simultaneously. Two months of frantic scaling followed. Monetization was, by his telling, an afterthought: charge something so the compute bills get paid, figure out the business model later.
Codex is the mirror image. Going all-in on code was the plan before ChatGPT existed. The internal logic was about actuators — code as the way a model acts on computers, robots as the way it acts on the physical world — so a sufficiently capable model with those two handles could actually do things rather than just talk. It took longer than expected. He dates real quality to early this year, with the 5.5 release as the inflection.
On methodology, he's surprisingly blunt. The current pipeline — pre-training, then mid-training, then post-training, then reinforcement learning with a supervised feedback loop — is definitely the shape everyone uses, and he expects a major rewrite of it. He doesn't know when or in what form. The staged structure doesn't strike him as optimal; when asked what optimal would look like, he hands the question to the models. There's an internal target mentioned in passing — intern-level automated research by September of this year, a full end-to-end capable researcher by March 2028 — but the phrasing in the room was tangled and I wouldn't lean on those specifics.
The interviewer presses on metaphor, which sounds like a softball and turns into the most original part of the talk. Sam's claim is that AI isn't a product category, it's a new utility, and there are almost no precedents to study — electricity, the internet, water, and not much else.
So he went and read about electrification. The detail he found: the early power companies, at least the ones he could find records for, did not sell electricity. Nobody knew what it was, and it sounded like something that would come into your house and kill you. What they sold was "light at night." And when they tried to explain that the same current could eventually wash your clothes, people couldn't make the jump.
He doesn't have the equivalent for AI. He suspects that "we're selling intelligence" is the wrong frame, that it doesn't land, and that the industry needs to find some way of explaining what it means to have an intelligence pipe running into everything you do.
The interviewer points out that the utility framing has now come up repeatedly in this class attached to different things — Jensen Huang used it for compute, arguing a university should procure compute the way it procures power. Which is it? Sam's answer: from a user's seat, you'll think in tokens or one level above. Hardware gets abstracted the way the base station behind your phone bill does. What you'll care about is whether you can use it a lot, whether it's cheap, and whether it does a good job.
Asked what he'd build for the class's one-person frontier lab project, he picks the least glamorous option: inference. Great models are coming no matter what anyone in the room does, he says, but nobody has invested enough in delivering huge quantities of cheap intelligence, and every frontier lab is going to have to become an inference company to a meaningful degree. There's a quiet tension with his own earlier advice — the assigned problem is probably not the one you want — since he's now assigning one.
Q&A covers a lot of ground fast.
On the claim that LLMs are a dead end, he concedes the interesting half immediately: these systems beat humans at some things and are wildly worse at others, especially long-horizon work with sparse feedback and heavy judgment. Then he counters with a recent case of one of their models disproving one of the Erdős conjectures, and mathematicians publicly wondering what that means for their field. His diagnosis of the field's skepticism is generational — a cohort of scientists who were far too certain about what scaling would never produce, while a few people just looked at the curves and kept going. World models matter, he says, particularly for robotics; betting against scaling seems misguided to him.
The follow-up about being the "I told you so" guy produces the sharpest psychological observation of the hour. He's past enjoying it. What he thinks happens is that people fuse a belief to their identity, and when the evidence turns, they can't release it because releasing it means losing the identity. He's careful to say it cuts both ways.
On education, he's unusually self-critical. He expected roughly one bad year of cheating after ChatGPT launched, followed by a redesign of how students are taught and evaluated. Three and a half years on, he says he struggles to point to any significant systemic change, and calls it a prediction error. The risk he names is atrophy in the ability to think. But he's not a purist — he writes to think, and would keep teaching that even though machines now write.
His spiciest take, after being pushed for one, is that AI is simply going to keep going, and that this isn't widely believed. If it were, he argues, society would be shaking more than it is. Another three and a half years on the current curve and the set of things society can do looks completely different.
Two forks follow. First, whether the technology ends up broadly distributed or concentrated in a few firms, which he calls an attractor state and says would be both unstable and unfair. He puts roughly 80% on the democratic path while acknowledging his own company is pointed at that outcome — and the interviewer immediately notes the self-fulfilling nature of a forecast you have agency over. Second, how compute gets distributed as leverage shifts from labor to capital. He funded a basic income study years ago, but prefers ownership to a fixed monthly payment, sketching a citizen's wealth fund where everyone holds a slice. The interviewer counters with Norway's sovereign wealth fund, which owns about 1.5% of all publicly traded companies.
On supply, the interviewer cites roughly a 5x gap between reserved and spot pricing for current-generation accelerators. Sam thinks that specific number has eased but agrees the shortage is gigantic, that a wave of new hardware is coming, and that demand may outrun it anyway. His structural point: if models get capable enough and cheap enough, demand has no natural ceiling — you'd want ten personal agents running, or a hundred — so in a sense the shortage never ends.
My reservations cluster here. The 80% figure arrives with no stated basis, attached to the outcome he's actively steering toward, which is exactly the situation where you'd expect stated confidence to be inflated. His spiciest take is directional rather than concrete, and the electricity analogy is doing a lot of unexamined work — he admits he doesn't have his own "light at night," which means the one historical case he offers is being used to argue for a framing he hasn't found yet. The education complaint has no mechanism attached. And the internal capability dates are stated without visible error bars, from a stage where error bars would be the honest part.
None of that cancels the useful core. If you take one thing from the hour, it's the pairing: pick a scale nobody has attempted, and then treat every objection — technical, financial, cultural — as a separate system to be designed around rather than a single reason to stop. Whether scale is magic or just selection bias, the people who got somewhere seem to have done that part deliberately, and that's reproducible regardless of how the argument about scale itself turns out.
Content Disclaimer:
This article is for general reference only and does not constitute professional R&D guidance, production process advice or quality certification. All material performance data has specific test premises; readers should verify parameters against actual equipment and working conditions.
All contents below are exclusive to the paid Word file, NOT available on this web page

