Note Wisdom
Notes on Anjney Midha's CS153 lecture: why context — the verifiable environment an agent learns in — now matters more than model architecture, why GPU rental prices keep climbing, and what past infrastructure booms suggest about a coming standardization of compute.
Institution: Stanford
Original Course: Stanford CS153 Frontier Systems | Anjney Midha from AMP PBC on Frontier Systems
Instructor Bio: This closing session is led by **Anjney Midha**, Co-Founder of AMP PBC and co-instructor of Stanford CS 153: Frontier Systems. Anjney Midha is a Stanford alumnus who previously served as a partner at Andreessen Horowitz (a16z) and held early leadership roles at Discord. He specializes in frontier AI ecosystem development, AI-native company building, and infrastructure strategy. Through AMP PBC, he supports builders and founders working across the full frontier AI technology stack.
Course Description: As the concluding session of the course, this lecture synthesizes the core themes from CS 153 and provides a forward-looking framework for understanding and building frontier systems. Anjney Midha recaps the key insights across energy, silicon, infrastructure, models, applications, and business models, and outlines the most important challenges and opportunities on the road ahead. The session closes with guidance for students and builders on how to participate in the frontier AI era, maximize their individual impact, and build systems that create lasting value.
I went in expecting a talk about chips and came out with a talk about standards, taste, and why a two-year-old accelerator now rents for more per hour than it did when it was new. This is week two of Stanford's CS153 frontier systems sequence, and the speaker — credited in the file as Anjney Midha of Amp, though the auto-captions render his name as "Andre" for nearly the whole hour, which is worth knowing before you try to search the raw text — spends about sixty-five minutes doing three things at once: warming up a room of roughly five hundred students, laying out a systems-level account of where AI infrastructure is actually stuck, and quietly arguing that compute is headed for the same messy standardization that steel, electricity and bandwidth went through.
A note on reading conditions before anything else. Much of the hour is built around slides that aren't in the audio, so there are places where he says "what do you notice about the shape of the curve" and you have no curve to look at. There's also a stretch between roughly 46:04 and 47:25 where the caption track simply stops; that's where he was showing Claude Code commits on screen. I've flagged the slide-dependent moments below so you know what you're missing rather than assuming you misheard.
He opens on a joke, not a slide: a concert needs an opening act, and he's the opening act for a quarter full of headliners. Someone tweeted that students should be wary of classes that sound like "AI Coachella," and he reads the tweet out, agrees with it, and then cheerfully keeps the Coachella bit for another ten minutes. Under the joke is a real piece of advice he returns to at the very end. Students keep asking him how to plan a career, and his answer is that you can't forecast it, so you should optimize for something you can actually observe: spend your time with people you like. He met his wife here as a sophomore, he says, and both companies he's started were with former Stanford roommates. His one visible regret is that he never went to Coachella because he was too busy.
That's not filler, and it connects to the technical argument. When students ask how they could possibly build something novel given what the big labs are spending, his answer from the previous lecture was to use those labs, do things that don't scale, and be asymmetric — and one asymmetry available to a student that isn't available to a large organization is obsession. Love and taste, he argues, are the parts that large institutions can't scale up.
The biography matters because he's making an empirical argument all hour and wants you to know the sample it's drawn from. Born in India, high school in Singapore, undergrad at Stanford in economics plus a major he thinks no longer exists (mathematics and computational science), then graduate work in bioinformatics at the medical school, which he describes as machine learning applied to healthcare. Fifteen years of applied work. He's a visiting scientist in the physics department, running benchmarks on how good frontier models are at physics and science reasoning. Over the last decade or so he's been involved in the early days of more than ten AI labs, including Anthropic, Mistral and Black Forest Labs, variously as angel, co-founder, or the person who helps scientists turn a paper into a company. He puts up a list of these as an explicit disclosure slide, and makes the point directly: an observer of an empirical experiment is biased by the data they're fed, and he's telling you which data he was fed so you can discount accordingly.
Then the frame. There's a stack, and it runs roughly from capital (flexible, goes anywhere) down through land, power and shell, then chips, then the cloud software that makes chips usable, then models and agents, then applications and solutions, then governance — safety, security, trust, the frameworks that let the whole thing get deployed. Fifteen years of distributed systems and cloud had settled into something stable. AI has unsettled it, and everyone at every layer is now asking how to unblock their bottleneck and revisiting basic assumptions about where they sit in the value chain. He calls this the great transition, and he's candid that nobody, himself included, knows what the new world looks like. The course used to be called Security at Scale, started four years ago with fifty students when he was running platform at Discord and his co-instructor Mike was running infrastructure at Apple; it's now at five hundred with another fifty waitlisted, and they're not sure they'll get to teach it again.
The stated goal of the class is preparedness for the real world, explicitly not internships. He repeats that twice in the first ten minutes, which I took as a warning about what kind of notes to take.
Here's the part I'd most want if I'd missed the lecture. Four years ago, he says, building a frontier model was a craft process with a simple recipe — compute, data, algorithms; a transformer; some pre-training, some fine-tuning, plug it into an app. Model releases came once or twice a year.
Now it's industrial. Base model training happens at least twice a year on clusters on the order of a hundred thousand B300-class accelerators. Post-training runs on roughly a tenth of that compute, two to four times a year, through mid-training to add capabilities. On top of that sits continuous post-training — supervised fine-tuning plus reinforcement learning — running more or less constantly.
And then the number that reframes everything: the reinforcement learning stage at the end is now consuming nearly as much compute as every other stage of the pipeline put together. He says it lightly, as though it's a fun fact, but the rest of the lecture is downstream of it. If the expensive, growing part of the stack is the part that needs a live environment to learn in, then the environment becomes the strategic asset.
He pauses to make sure people can follow, since about a third of the room had done any RL problem set. The primer is short and uses a dog: you don't tell the system how to do the task, you tell it what outcome earns a reward, you hand out the reward when it gets there and withhold it when it doesn't, and you repeat. What changed about two years ago is that you now initialize that loop with a language model that has strong priors about the world. Old reinforcement learning — chess, Go — would clear human performance and then flatten out, and he links this to the bitter lesson from lecture one. His explanation for why it's different now is that the earlier models just weren't general enough to keep learning, and he flags this as genuinely unsettled rather than settled.
The flywheel story that follows is the concrete version. Roughly four years ago he got a call from two people then running research at OpenAI who wanted to leave and start a lab, and the business plan they sketched was: raise money, buy compute, add data, pre-train, ship a model good enough that programmers want it, run inference. Inference then feeds you two loops — revenue to buy the next round of compute, and observation of whether the model actually completed the task, which becomes training signal. He mentions pitching this up and down Sand Hill Road and collecting twenty-two introductions and twenty-one rejections, mostly on the grounds that there was no empirical proof yet. Four years on, he says, there is proof, and he cites revenue growth at Anthropic and OpenAI and adoption at Google.
This is the intellectual center of the talk, and it starts around (24:30). If the recipe is repeatable and everyone can run it, where does the value land? His answer is context, by which he means the environment the agent operates in. The dog-training version: if you're teaching a dog to fetch in a park, the park is the context — the kids running around, the grass, the rain. All of it shapes whether and how well the animal learns.
Two questions fall out of that, and he recommends students use them to pick projects. First: where is there context that can be reliably measured and verified? Code is verifiable because unit tests pass or fail. He says material science turns out to be verifiable too, and points to Periodic Labs, where people he's working with are using reinforcement learning against physical verification to hunt for superconductors, with a facility full of robots in Menlo Park that he floats as a possible field trip. Second: who captures the value? His answer is whoever has unique, defensible access to that context — teams that got there first, or that have an insight. Teams locked out of the context essential to improving on a domain, he says flatly, won't get a chance.
The illustration is the episode I found most interesting and also most in need of caveats. About a year ago, he says, news broke that OpenAI was trying to acquire an IDE called Windsurf, and a few days later Anthropic cut off model access to Windsurf's users. Cutting an API off without warning is unusual in this industry, he notes, and then explains why it made sense: if a competitor is watching how your model helps the customers you're trying to win, that's context leaking out. To people inside the industry this read as normal; to everyone else it was a surprise. His framing is that one comfortable assumption died that week — the assumption that if you're an application company, you can always count on your model provider to keep supplying intelligence.
The second illustration is longer and I think stronger. Mistral was founded by people behind Llama and Chinchilla, and their thesis was a distinction between ordinary context and what he calls sovereign context. A developer in Silicon Valley may not mind piping their coding context to a server somewhere. A government does mind, when the context is national records or defense work. That means you need models and weights running on infrastructure you control, locally. He spends a few minutes on why this is a break in the pattern: fifteen years of cloud history were driven by falling marginal costs, because Amazon and Google had piled up servers for their own needs and discovered they could rent the spare capacity out, and that's how you get AWS, GCP and Azure. It's very hard to beat that flywheel. For the first time in fifteen years, he says, it's changing — and the tell is a head of state and the head of Nvidia sharing a stage in Paris next to a thirty-three-year-old scientist who has never run a business. He brings in the Cloud Act here: if your workloads sit on servers operated by a US company, anywhere in the world, the US government can reach that data, and for some governments that's disqualifying. Expect to hear the phrase sovereign AI a lot, he says.
Then the playbook, stated compactly. Pick a frontier you want to move — material science, coding, whatever. Get enough research compute to run experiments. Ship something into a context you actually have access to. Run the feedback loop. Keep both flywheels turning, because they reinforce each other. Eventually, if a team executes well enough, the loops start propelling themselves, which is what people mean by recursive self-improvement. His own framing is deliberately at the level of the whole system rather than any single model, and he suggests students ask the guest speakers whether they see it that way.
He raises the biggest open question himself and I'm not sure he resolves it. Does reinforcement learning generalize outside the task distribution it was trained in? The philosophical view, which he describes without endorsing, is that given the right context and enough compute an agent should be able to learn anything — at which point you could tell your coding agent to go build itself a materials-science environment and run its own loop. The empirical view, which is his, is that life is messy and progress concentrates in easily verifiable domains, so you get relentless improvement in coding and much less in places where verification is hard. His examples for the hard case are aesthetics, beauty and love, and the concrete one is long-form writing: he asked a founder friend to sanity-check a blog post he'd outlined himself and had a model flesh out, and the friend replied in thirty seconds asking whether he'd used Claude. Amp now has an internal rule against circulating model-generated documents to each other. The counterexample he offers is Grant Sanderson of 3Blue1Brown, an old friend, whose value is taste — the ability to deconstruct a hard topic from first principles — and he says outright that RL is only one technique among many we'll invent.
Three things left me unsatisfied. The verifiability heuristic is genuinely useful as a filter, but he never says how you'd tell in advance whether a domain is verifiable enough; "code has unit tests, materials science has physical measurement" is clear in retrospect and much less clear when you're picking a thesis topic. The context framework also has a whiff of unfalsifiability — if a company wins, it had unique context, and if it loses, it was locked out, which makes the claim hard to test against the twenty-one rejections he mentioned earlier. And structurally, he promised four bottlenecks — context, compute, capital, culture — and only got through two, with a throwaway line about possibly needing an extra office hour. Capital and culture may well be abandoned rather than deferred, and I'd want to know before building a mental model around a four-part frame that only has two parts.
This is where he's most at home, and it's the most useful part of the hour if you care about infrastructure. Scaling, he says, works predictably: capabilities scale with compute. He puts up public estimates overlaying Anthropic's revenue against the compute the company brought online, and the pattern he points to is that each time new compute came up, roughly sixty to ninety days later there was a jump in capability and then a jump in revenue. His framing is deliberately financial. A dollar of hard assets — land, power, shell — normally trades at three to four times revenue in public markets. A dollar of software revenue trades at thirty to forty times. The stack converts one into the other, and he puts the value multiplier at around ten. This is why he renamed the course and why he wants students thinking as full-stack thinkers rather than only as engineers: if you want to do frontier research sustainably, you have to run the entire loop.
Then he takes on the objection he says he's heard for four years — that compute is a commodity and you should just hand the company money and let them rent from a hyperscaler. Not quite. Using an internal forecasting system his team calls the Amp grid, he shows rental prices for the H100, a chip more than two years old, and asks the room what the last ninety days look like. Going up. Two years ago, when he was calling cloud contacts to rent capacity for their teams, the average H100 hour went for about a dollar seventy-three. He then describes a founder who has raised somewhere between seven hundred million and a billion dollars messaging that morning in a compute crunch, urgently needing H100s, answering the question about quantity with "take them right now," and answering the question about price with — his phrase — "price not a problem." The point is that entire public companies are valued on the assumption that all chips depreciate, while on the ground the opposite is happening.
What he does next is the reason I'd recommend this lecture to someone who isn't in the field. Rather than extrapolate, he goes to history. Steel from 1867 to 1895, with a speculative run-up, the panic of 1873, and then a long plateau once society decided it wanted stable production and consumption. Fiber optics, meaning Cisco, Lucent, Nortel and WorldCom, which is the comparison everyone reaches for when they say AI bubble. DRAM and its violent semiconductor cycles. The Baltic Dry Index for shipping. Uranium and the 1970s nuclear build-out, where government intervention is what eventually stabilized the resource. GPU rental prices sit at the end of that lineup, falling after launch and then climbing steadily from around August 2024.
His reading: infrastructure runs in cycles, and the historical interval between them is about 2.8 years in digital build-outs and 6.3 years in physical ones. What makes AI novel is that it's both at once. You have to marshal enormous physical inputs — land, power, shell, silicon — to produce something made of bits. Atoms in, intelligence out, and the two worlds don't like colliding.
The argument closes on a distinction I'd have appreciated earlier in the hour. Compute is not fungible the way electricity is. A megawatt is a megawatt anywhere on the grid, and we've had roughly seventy-five years of stable forecasting for energy supply. Compute has neither property. Chips from different vendors aren't interchangeable, and neither are chips from the same vendor — an H100 is not a GB200 is not a B300. Worse, demand is nearly impossible to forecast: training is spiky, because you fiddle at small scale and then spike into a hero run when an algorithm starts working, while inference follows a daily cycle, busy during the day and idle at night. He illustrates the contrast with power cuts at the boarding school he attended in rural India, roughly every other week in summer.
That produces hoarding, and hoarding produces the question he wants the class to sit with all quarter. Historically, two things turn a scarce, monopolized input into something productive and broadly accessible: a standard everyone agrees on — AC/DC, TCP/IP — and an institution with enough authority to enforce that standard, because at scale humans don't stay aligned on their own. He puts up the textbook definition of fungibility: a common unit, a standard delivery interface, interconnection and pooling, metering and settlement, and buyers able to substitute one supplier's unit for another. None of that holds for compute today. His punchline is that we're in the pre-standardization era for compute, sitting where railroads were around 1886 or electrification around 1907 — new general purpose technology, an infrastructure explosion, consolidation into three or four players, and then either industry self-regulation or an institution stepping in.
Two assignments follow. What would a peaceful transition on compute actually take over the next couple of years, and what's your part in it? He's insistent that students in the room are participants rather than spectators, and that writing publicly about which standards you wish existed is a real contribution, because the people running these institutions are coming to speak to you and might hear it.
Then the closer: a screenshot of RTX 5090 prices over eighteen months, the chip that was last year's grand prize for the best project, five of which Jensen Huang signed when he visited. This year, he jokes, he won't be asking for more, because they've gotten valuable.
If there's one thing I'd carry forward from the hour, it's the pairing he keeps returning to. The technical story — that the expensive part of the pipeline is now the part that needs a verifiable environment — and the historical story — that every input this valuable eventually gets standardized, messily, after a boom — are the same story told at different scales. Both are arguments about frontier systems as something you build end to end, from the power contract up to the feedback loop, rather than a model you download. And both, on his own telling, are unresolved.
Content Disclaimer:
This article is for general reference only and does not constitute professional R&D guidance, production process advice or quality certification. All material performance data has specific test premises; readers should verify parameters against actual equipment and working conditions.
All contents below are exclusive to the paid Word file, NOT available on this web page

