At a glance
The guest is Neil Movva, co founder of Sail Research, and before that a GPU kernel engineer at NVIDIA starting in 2016, back when the tensor core was still a skunkworks fight for five to ten percent of the die. He is on Invest Like The Best with Patrick O'Shaughnessy to describe a company he calls a token factory: an API that serves open weight models at, in his words, a price that is unbeatable in the market, plus long running sandboxes for agents that work for hours, days or weeks.
The headline number in the title is real but it is the floor of a range and it is an ask, not a measurement. At 1:15:23 he says the world needs "at least three to six orders of magnitude improvement in cost per token," anchored to a specific arithmetic: a trillion tokens at frontier API pricing costs at least five million dollars, and he wants that trillion to cost five thousand. That is the 1000x. The six order version puts it at five dollars.
What makes the interview worth reading rather than clipping is how he decomposes the path, and how much of the conventional list he throws out. Process node does not help him: he says performance per watt of a BF16 multiply barely moved from Hopper to Blackwell to Rubin, and barely moves across TSMC 5, 4, 3 and 2 nanometer. Sparsity does not help him either: frontier models are already about one percent dense. What is left is the KV cache (uncompressed by an order of magnitude or two, his only quantified headroom), serving for throughput instead of latency, buying whatever silicon is mispriced rather than bidding against Anthropic for Blackwell, and scavenging one megawatt sites on intermittent solar and wind that he will happily run at ninety five percent uptime.
The whole thing only closes because of one product bet: background agents. If nobody is waiting on the answer, latency stops being a constraint, and every expensive thing built to protect latency (NVLink, redundant fiber, diesel generators, concentrated gigawatt campuses) becomes optional. This page rebuilds the conversation in its order, keeps every figure he gives, and then puts the 1000x on a table where you can check it.
Building a token factory (0:38)
Patrick opens with the question he always opens with: literally what is the system, and why should it exist.
Movva's answer is compact. Sail Research is a token factory. There is an API. Anyone can send requests against large open source language models for any task, and Sail serves those tokens at a price he claims is unbeatable. On top of that sit sandboxes, which are long running agent virtual machines hosted in the cloud, built specifically for agents that run for hours, days or weeks.
Patrick reframes it back to him: you are a peer to the other inference companies, but you serve one specific kind of inference, and the goal is to be the absolute cheapest provider and enabler of that kind of use of intelligence. Movva takes the framing exactly.
The theme of our company is abundance. We want to deliver this new commodity of intelligence to as many people as possible at a cost that is sustainable for almost every industry.
Then the line that sets the bar for everything after: whenever you make something ten times cheaper, it is a new product category. That is the aspiration for tokens. And underneath it, the thing he keeps returning to for the next eighty minutes, stated plainly: it is profound that the machine can think, and the job now is to make as many machines as possible in the world work toward thinking.
Is cost per token even the right unit? (1:32)
Patrick pushes on the metric itself. Is token cost the right way to think about this, or is there a better frame?
For now, yes, absolutely. His north star is the lowest cost per token in the industry, by a mile. But he does not believe tokens are the final unit of work or of intelligence. They are simply what we use today, and they are straightforward.
The direction after tokens is outcomes, which he admits is a vague word, so he grounds it in something concrete about how agents already behave. When you consume tokens through an agent, you do not actually control how many tokens the agent reasons for. It can reason for some amount of time, or call some number of tools, and you do not set that. Increasingly the shape will be: give the agent a unit of work, let it take as many shots on goal as it can, and however many tokens it burns getting there is a dependent variable of the task.
Patrick's summary, which Movva accepts: agents that self administer a token budget, rather than a company setting a budget for how many tokens its engineers may spend per month.
Why there is an opening at all (2:34)
The obvious objection: the entire world is already oriented around more, better, faster, cheaper tokens. Every serious company is attacking this. Where is the market inefficiency?
Two tailwinds, in his order.
One: the rise of open source. He insists on saying it first. An increasing number of his Customers, and the broader market, care about owning intelligence. They want control and sovereignty over the thing they depend on. That produced a real market for customized models, or even just vanilla open weight models that nobody can ever take away from you: you always have the weights, you always have the right to deploy them however you like.
Two, and this is the actual opening: everyone who serves those weights built for the wrong workload. He names the field directly, Baseten, Fireworks, Together, and says take your pick, they all focus on low latency inference. And they were pulled that way by one very important Customer: Cursor.
His read is that this was the right choice about a year ago. As of six months ago it started to look like low latency was not the only thing you want out of an agent. You wanted persistence, longer horizons. And now, to him, it is obvious:
The future of agentic inference is long horizon tasks. You're going to run the machine for hours or days at a time. It doesn't matter if it spits out tokens at 100 tokens per second. Maybe 10 is just fine. That comes with corresponding advantages and efficiency.
That last clause is the entire thesis in miniature. Ten tokens per second instead of a hundred is not a degradation he tolerates. It is the source of the money.
The future of background agents (4:21)
Patrick pushes back with the honest consumer instinct: I want everything as fast as possible.
Movva's answer is a redirect rather than a denial. When you are waiting on it, you absolutely deserve the fastest answer possible. His trick is that he does not want you waiting on it at all. He wants the work proactive, in the background.
One way to say it is like the best latency is no latency at all. When you wake up in the morning, the work's already been done overnight. You didn't even have to ask for it.
He concedes we are not quite there yet. But the more important point is structural: the more you sit in the loop prompting an agent and waiting for a response, you are the bottleneck on how much work the agent does. What he wants is for agents to operate on human time scales. You do not manage your colleagues every five minutes. You give them a high level task and check in maybe daily, more likely weekly. That, to him, is the shape of human and agent collaboration.
Test time compute is the early indication (5:07)
Asked for the evidence that this is actually happening, he names test time compute scaling: the idea that you can give an agent more time and it gives you a better answer.
His timeline for it is specific. Theorized about two years ago. Not something you could actually bet on until late last year with Opus 4.5, which he calls the first agent at all suitable for longer horizon tasks, and "pretty mediocre when it first came out." Look at more recent models, and at what Sail has done on the open source side, and agents are capable of running for an hour at a time. Not days. But an hour is quite suitable today.
And that is enough to draw the line: average turn or task length keeps getting longer, and it does not take many points to draw out the exponential and conclude agents are worth running for longer periods.
The split he is underwriting: 50/50 now, 90/10 later (6:09)
Patrick asks for a market share number on long running agents in three years. Movva loves the market precisely because it is unbounded: there is no human in the loop, so you can consume as many tokens as you like in the background.
Then the constraint on the other side, and it is a good one:
If you tell me to consume 10x as many tokens at Codex or at Claude Code, I'm actually not sure if I can anymore. I'm already in a loop and locked in coding for most of the day that I'm at the laptop.
Human attention is capped. Background consumption is not. His numbers: we end this year at maybe fifty fifty background versus real time workloads, and he sees it going to ninety ten in favor of background.
What background compute is actually good at (6:40)
He gives categories, not hand waving.
Deep research at a scale humans do not attempt. Not a hundred sources, not a thousand sources, but ten thousand or more, when you want a definitive answer. He names a Customer, Parallel Web Systems, which is building an index over the whole internet and monitoring it in real time for changes. That is exabyte scale work, and it needs a different kind and scale of intelligence.
Cybersecurity, which he sees following the same path. The argument is asymmetric in a way that is easy to miss: yes, there is a lot of code you can generate, but there are exponentially more ways to break that code than there are to generate it. Customers are working hard on agents that break any piece of software so they can proactively patch it. When Fable came out, and when Mythos came out, the security community pushed to run them against every line of code ever written, looking for bugs in twenty different ways: memory errors, business logic errors, network vulnerabilities. These become specialized agents. You do not have the model read the source once. You stand up environments and actually pentest the application.
Out of that comes the joke he repeats as if it were a definition:
At some point, people started to make this joke that security has become proof of work. When you want secure software, it's really a question of how many dollars did you spend on Anthropic's APIs trying to break into your software.
And a genuinely important technical aside: the frontier of intelligence is jagged. It is not the case that the biggest model finds a superset of the bugs. You will find bugs with a very small model that the large model misses. You find bugs with Haiku that you do not find with Fable, and vice versa. That pushes toward diverse sampling across many models rather than always reaching for the largest one, which is itself a cost argument.
Dreaming further: proactive intelligence and the dollar cost of an answer (8:42)
Patrick asks him to get imaginative about the new product category.
For individuals, the thing he wants is proactive intelligent agents. Imagine a Siri running in the background all the time, understanding every email and every text message you receive in a day, holding a much more encyclopedic view of your life. Today's assistants are point solutions, which is why you end up doing so much prompting. Siri is not proactive. That is fixable with abundant inference. And if you trust the machine enough on both axes, reliable and private, the machine can model how you interact with it and surface your next action whenever you open your phone. Can we build a good model of what you are going to do next? His estimate: yes, we totally can.
The key to it is not a model breakthrough. It is price:
You have to be willing to spend tokens without any promise of return. That's the unlock.
Then he zooms out to the version of this argument that matters most. We have a form of intelligence that can tackle any verifiable problem. Any verifiable problem means most software. It means a lot of formal work like math proofs. And it could mean scientific discovery. All of those already have a hidden dollar cost attached: how many tokens could you possibly harness to make this work.
And the number is coming into view. Not millions. Thousands. Maybe hundreds, or even tens of dollars in the near future, for a definitive answer to a scientific question or a research problem.
The limit case, and the one thing he will not claim (11:45)
If that future arrives, Patrick says, we become limited only by the questions people can ask. Movva agrees: the models are on the cusp of taking a high level question and chasing down every possible follow up on their own. The remaining variable is the token budget, and, he says, we will solve the token budget problem.
Asked about non verifiable tasks, he draws the line cleanly and does not hedge it:
I put basically the entire category of human taste into that category. We have not solved human taste yet and I don't know that it fundamentally can be.
He says he is excited to be surprised, but Sail is focused on quantitative problems. The quality of writing and the beauty of art, he leaves to people.
The master plan, layer by layer (12:46)
Patrick sets up the second half of the conversation: the stack of solutions that gets to the giant token factory. Software, hardware, power. Talk through the plan.
Movva's ordering principle is capital efficiency. You always start with software, because the question is where the opportunity is on today's chips, in today's data centers. Software has the highest leverage and the lowest capital requirement, so it is where a startup earns the right to touch anything heavier.
NVIDIA and the GPU stack (13:09)
The first thing Sail did was build the entire LLM software stack around peak GPU efficiency: use NVIDIA GPUs, and squeeze more tokens out of the same chip than anyone else in the world. That starts at the lowest level of programming, kernels, which is his background. NVIDIA was his first job while he was still in college, and he got to watch, in his phrase, the tensor cores earn their right to be on the chip. This is 2016.
What a tensor core is, and why matrix multiply (13:47)
Patrick asks for the layman's version, and Movva gives a clean one. A tensor core is a specialized unit on the GPU that accelerates matrix multiplication. That is it.
Then Patrick asks the better question: why is matrix multiplication so important? Movva does not pretend to a deep answer, which is to his credit:
I cannot say that there is a divine truth of the universe that explains why matrix multiplies seem to be the atomic unit of computation. But one way I've heard it described to me is, well, it's a really succinct way to mix two blocks of numbers together and have them interact in some interesting way.
His only additional comment: it is really convenient that linear algebra turns out to be a compact representation of arbitrary relationships in data.
How the tensor core won its die area (14:18)
The history here is the part a general audience never gets, and he tells it from inside.
NVIDIA is a graphics company with dominant share in GPUs and gaming graphics. In the mid 2010s it began skunkworks projects to make the graphics processor more suitable for machine learning workloads it was tracking. He remembers reading the lab notebooks of his managers: they would go to small ML conferences, ICML and NeurIPS at the time, and take notes. Oh, this deep learning thing seems to be catching on. What is interesting is that these grad students are using gaming NVIDIA GPUs to train their large models. We should double click on this and figure out what is going on.
By 2015 and 2016, Jensen Huang had the conviction that this usage was only going to grow, and started allocating precious silicon die area to a capability that was still emerging.
The internal politics of that decision are the interesting part. You are taking a gaming chip designed for painting pixels on a screen and adapting it to do matrix multiplies. When you ask for silicon area you are competing directly against the graphics teams, and in any chip company die area is guarded ferociously, because investing in the wrong technology is opportunity cost you can never recover.
We kind of fought tooth and nail and got just a tiny bit of die area, maybe like 5, 10%, something like that, for the first generation of these chips.
What that first generation bought was acceleration for basic convolutions, the fundamental operation for the computer vision models of the day. And then a software team, his team, tried to squeeze every drop of performance out of it.
Speed of light (15:51)
That software team is where he says he learned the thing he still runs his company on. NVIDIA has a term: speed of light. They always chase the speed of light for any piece of hardware they make, and it is so ingrained that if the machine can do it, engineers will push the machine to that frontier.
The speed of light is the edge of what's possible. If we think the chip can run at this frequency and produce this many multiplies per cycle, we're going to get there. We're going to break every bottleneck and get to that peak level of performance.
He carries it forward as a management principle, and the distinction he draws matters:
To this day, I tell all my engineers, we're chasing 100% speed of light. I don't care about relative numbers versus the competition. I only care about absolute numbers.
Two more NVIDIA stories: tenure, and the milk club (16:53)
Patrick asks what else stood out about the company. Two things.
Retention. A lot of the people he worked with at NVIDIA in 2015 and 2016 are still there today. He calls them the best engineers on the silicon side he has worked with in his career, extremely motivated, believers in parallel computing as a concept through all its incarnations. This is their life's work.
Frugality, illustrated with one perfect detail. All the Silicon Valley companies cut perks after 2008, so no free lunch. NVIDIA took it a step further: no free milk in the fridge. If you wanted milk in your coffee, you chipped in a dollar a month to the milk club, and the milk club stocked Costco milk. He notes that Sail does not do that. But the frugality permeates the company.
Throughput or latency: the trade nobody gets to break (17:53)
This is the technical heart of the first half, and it is the mechanism that funds everything downstream.
The GPU is fundamentally a throughput machine. It is happiest when you give it a lot of work and let it chew through that work at peak utilization of its compute units. But that is not the direction AI went. We pushed AI to be an interactive chatbot, and in that world you care enormously about spitting answers at the person at the keyboard as fast as possible.
That is hostile to the hardware. It is very difficult to put the GPU in its happy path of being fully compute utilized when you are trying to emit tokens quickly. There is a fundamental trade off on the GPU between being throughput oriented and being latency optimized, and everyone chose latency, because the shape of usage was chatbot shaped. His claim: that is the most profound change coming in the next year. Chatbots give way to proactive background agents, and in that world you build the stack around throughput.
Why the trade is unbreakable (19:00)
Patrick asks the right follow up: explain technically why you cannot have both.
Movva starts general. It is foundational in almost every system: there is always a trade between getting a small amount of data through as fast as possible, leaving a lot of buffer room for it, versus running wide and slow. Narrow and fast or wide and slow is a classic trade in all of computer science.
For the GPU specifically, the thing to focus on is batching. You want to group many users' work into a batch and run it all at once, because that is what the parallel processing hardware wants. The catch:
You're doing net more work when you run a large batch of compute together. And so you might be filling all the units, but every step along the way, as you carry a batch of work through the GPU, there's more work to be done. And so any individual token or any individual user's request in that batch, it's going to spend a longer time on the GPU being carried with other people's traffic.
Then the analogy, which he offers unprompted and which is the cleanest thing in the interview:
If you want to get downtown in SF, you can take the bus or you can take private transit. The private transit is going to have its own direct path, as the crow flies, using exactly the roads that you want from point A to point B. A bus, it's going to have to serve many more people and it has to fundamentally do something that works for everyone.
So it takes a slower path, and it stops and waits for people to get on and off. Patrick lands it: step one of your optimization is to build the best possible bus on top of NVIDIA GPUs. That is exactly right.
NVLink, tensor parallelism, and sublinear scaling (20:24)
His example of what a different parallelism scheme buys you is worth following closely, because it is where NVIDIA's real moat sits.
One of the things NVIDIA has genuinely innovated on is NVLink, the interconnect between GPUs. It is good enough that if you have a large matrix multiply you want to perform faster, you can cut that matrix multiply into pieces and shard it across two or more, up to eight, GPUs, have them each work on a tile, and reduce the results back together at the end.
The arithmetic of what that actually buys is the point:
Each GPU is now doing one eighth as much work, and therefore it can finish faster, but not eight times faster. It's sublinear scaling. You'll use eight times more hardware, but you won't get eight times the speed. You might get like four to five times the speed.
Eight times the hardware, four to five times the speed. Two reasons: communication overhead, and the fact that every GPU is a little less efficient working on a smaller tile than on a larger one. It is the only way to get minimum latency, and, he says flatly, it is not the choice he would make.
What he would use instead: expert parallelism or pipeline parallelism, plus tricks to overlap and hide communication latency in ways a low latency server has far less freedom to do (because hiding latency behind other work is precisely what a latency optimized server cannot afford).
Patrick asks the sharp question: is NVLink therefore a technology that improves latency performance, and only latency performance? Yes. NVLink, he says, is mandatory for low latency inference.
Which is exactly why he does not need NVIDIA (22:22)
And here the whole strategy snaps into focus.
NVIDIA is excellent at low latency inference. Sail does not care much about low latency inference. So where does that leave him? He is not holding his breath for other vendors to figure out NVLink: it is challenging technology, hard to scale, hard to productionize. But if some other vendor's chip is good at the foundational compute components, if it can still do matrix multiplies really well and simply cannot communicate results across peers quickly, then there is room for that chip in his stack as a really good compute per dollar option.
What I actually optimize for in most cases is how many FLOPs does this chip have and how much is it going to cost me per hour to operate, to own and operate.
There are chips that rank higher than NVIDIA on FLOPs per dollar and do not have the interconnect. His job then is to pick a parallelism scheme that makes that chip suitable for inference. It will not be tensor parallelism, where NVIDIA is basically mandatory. Other techniques may work fine.
| Dimension | Training economics (the last two years) | Inference economics (his bet) |
|---|---|---|
| Is the spend speculative? | Yes. You buy compute against demand that may never arrive speculative | No. "You don't hoard tokens, you use them immediately" non speculative |
| Workload superset | The superset. Any training cluster can serve inference | The subset. An inference fleet cannot necessarily train |
| What it demands of networking | High bandwidth between chips, cluster wide | Far less. NVLink is mandatory only for low latency serving |
| Cluster concentration | One place. Nobody wants cross data center training | Distributed. Aggregate power, never concentrated power |
| Viable site size | 100 MW and up, "basically impossible" to build now | 1 MW, which he calls plentiful (about eight racks) |
| Uptime required | Redundant power, redundant fiber, diesel on site, SLAs | 95% is fine, 80% possibly fine at the right price |
| Optimization target | Time to a finished model | FLOPs per dollar per hour, owned and operated |
| Latency posture | Not applicable | Average throughput competitive, P99 explicitly uncontrolled |
| His historical analogy | Dot com networking capex: demand that never came | Not that. Inference spend "monotonically increases" |
Chips, memory and transformers (23:27)
Before leaving latency, Patrick asks about the companies that go the other way: Cerebras, Groq, the very low latency accelerators. What happens to that segment?
Movva's answer becomes a genuinely good lecture on memory, so it is worth following it all the way down.
Cerebras, Groq and a couple of others coming out of stealth made an interesting bet: not to build another GPU, but to build a different kind of accelerator organized around a different memory hierarchy. They want to maximize the amount of SRAM on the chip and use that as very fast memory for weights and KV cache.
SRAM against DRAM (23:58)
There are two ways to make memory for a chip.
On die SRAM. You integrate the memory onto the logic die itself. You tell TSMC you want this many megabytes of storage on your chip, and there is a standard cell library for it: you print out a bunch of SRAM cells. The standard construction is the 6T cell, a six transistor arrangement that is stable. You write a bit to it and it holds that state without any active management, which is what "static" in static RAM means. The problem is area. SRAM eats silicon.
His worked example: take a large die, say NVIDIA Blackwell at 800 square millimeters. If you made that entire die SRAM, you would land in the single digit gigabytes. Not a crazy amount of data storage for the most expensive silicon on Earth.
DRAM, from an entirely different process. Not TSMC, but Micron, SK Hynix and Samsung. DRAM is a whole different way to build memory, focused on capacitors rather than transistor cells. To write data you push a charge onto a capacitor, and the instant you do, the charge starts leaking. That is the "dynamic" part:
You must, every 50 milliseconds or so, refresh every bit you've written. So you're constantly juggling billions of balls in the air. Essentially billions of bits have to be managed by a memory controller which is reading and refreshing every bit on the DRAM.
The payoff for all that juggling is density, on a process so different that the industry split DRAM manufacturing into entirely separate companies. Take DRAM from Micron, SK Hynix or Samsung, stack it into many layers, and print or solder it around the main logic die you got from NVIDIA, and you get HBM, high bandwidth memory.
The numbers, and they are not close (25:59)
Here is where the abstraction turns into figures you can check.
- Blackwell carries 288 GB of HBM capacity around the logic die.
- The logic die itself has maybe 500 MB of SRAM.
- He calls that possibly multiple orders of magnitude, three orders of magnitude, of density difference between DRAM and SRAM.
And bandwidth runs the other way, for a physical reason he states cleanly: SRAM is physically close to the logic gates that do the computation. The arithmetic logic units sit right next to the SRAM they pull from, so the compute units doing the matrix multiplies can pull data from SRAM at mind boggling speeds.
- Cerebras quotes 21 petabytes per second for Wafer Scale Engine 3.
- HBM on an NVIDIA Blackwell is around 10 terabytes per second.
Many orders of magnitude in the other direction. More capacity, proportionally less bandwidth.
What Cerebras actually does, and the path to a thousand tokens per second (27:02)
Given that SRAM density cannot easily be increased on a chip, Cerebras attacks the constraint from the other side.
They refuse the reticle limit. TSMC imposes an 800 square millimeter limit on a die; Cerebras takes the entire wafer, connects every die to every other die over the scribe lines, and gets as much SRAM as it can on the whole thing. That reaches, say, 50 gigabytes of SRAM per wafer. Then stack many wafers together in a pipeline, and you are at up to a terabyte of very, very fast memory.
And you do all that work for one reason: to read data from SRAM at 21 petabytes per second per wafer, so you can move the entire parameter count of a large model, he uses Kimi as the example, on and off the logic cores in about a millisecond.
So there you go, you have a path to a thousand tokens per second.
The KV cache is the thorn (28:02)
His prediction for that segment is a hybrid outcome, and the reason is the KV cache.
You can pair the Cerebras chip where it is strong, extremely fast access to memory, with something that has more memory capacity. It is true that you can take a one trillion parameter model like Kimi and fit it across a large number of Cerebras wafers. What you cannot do easily is handle the KV cache, because the KV cache grows as people use the model and is always dynamic. You do not even know how much you are going to need. It depends on how many users you have and how many you want to serve.
Patrick asks him to explain KV cache from scratch, and the explanation is one of the better plain language ones on record:
Whenever you use a language model, every token you send through the language model stays in the context window for as long as you're having a conversation. So if we talk for 100,000 tokens, the 100,001st token is still in the conversation behind us, and the model is referencing all the past conversation history in order to make better predictions.
That reference material has to live somewhere. You have to store a representation for every token you sent through the model, and, critically:
It frequently gets to be larger than the weights of the model themselves. You have this crystallized knowledge in the model weights, and you have the dynamic knowledge of the exact conversation we're having in the KV cache.
Remember that sentence. It is the setup for the only quantified inefficiency he names later in the show.
Why long conversations degrade (29:34)
Patrick raises the everyday observation: deep in a conversation, things start to degrade. Is that a technical problem?
Movva's answer separates two things people usually conflate. The KV cache is an exact representation of everything that came before. No information is being thrown away. The problem is on the training side:
During training, the model did not get trained primarily on very long context conversations. It got trained primarily on, let's say, 8,000 token conversations or 16,000 token conversations. So if you take the model to 200,000 tokens, there was some training that happened at that context length, but it's not the model's core strength.
So the perennial battle for the frontier labs is making a model exactly as intelligent at 10,000 tokens as you expect it to be at 200,000. He notes that million token context windows have been a concept for years, that he believes Anthropic was first to hit a million token context length, and then he undercuts the whole feature with his own behavior: he still uses /compact in Claude Code well before a million tokens, because he does not think it is actually great to hit the full length.
So the extremely fast, extremely low latency approaches are ultimately limited by this factor. Yes. You can do whatever you want for the weights, and have unbeatable performance on weight storage. The KV cache will be a big thorn in your side.
Where those chips land in three to five years (30:36)
His answer: think of Cerebras and Groq as accelerators, not as GPU replacements. They are really good used in conjunction with a more traditional GPU like device that critically has off chip memory. You want off chip memory for capacity and on chip memory for speed, and the right answer is to hybridize them.
Then he explains why the transformer itself forces this split, and this is the sharpest technical passage in the interview.
Take a transformer to a million token context length. Two things are happening in every layer:
- The MLP, the feed forward part, where most of the model's world knowledge is encoded. At large enough batch size this is compute bound: the actual matrix multiplies.
- The attention layer, where the model dynamically adapts to the current conversation. In the limit this is memory bound.
I would say the original sin of transformers is that you've taken this extremely fundamentally memory bound layer and juxtaposed it right next to a compute bound layer. It is very difficult to have a single chip that is good at both compute operations and memory operations.
The GPU is quite balanced in this regard, but any specialized chip has to choose. Which yields the placement he expects: put the MLP, the weights, on the Cerebras chip, where fast memory access to a matrix multiply is exactly right, and put attention on the GPU, which has the capacity to scale to really long contexts. And he says he believes this is what is happening with NVIDIA and Groq.
A riff on transformers (32:24)
Patrick asks him to riff on the 2017 innovation itself: strengths, weaknesses, and whether it stays dominant.
What it did. Transformers let us learn on unsupervised data really effectively, because what a transformer is about is taking any arbitrary sequence of data and finding patterns in it. The attention operation lets the model dynamically adapt to whatever it thinks is the most relevant component of the sequence. Every step through the transformer, you are reweighting the input you looked at before and figuring out what is most relevant to the next prediction. That makes it extremely amenable to arbitrary sequence data, and the most interesting sequence data we produce on a regular basis is language. Hence dominance in the language regime.
Zooming out, the real answer is that transformers scaled. They make no human prior:
Transformers just say, well, there's going to be a pattern in the sequence of data, and if there is a pattern, I'm going to find it. I'm going to throw more and more parameters at this problem until it works.
The parameter count history, which he gives in full. One of the hard limits in computer vision was going from hundreds of thousands of parameters upward. Linear models like support vector machines and other legacy machine learning had parameter counts in the thousands. Deep learning got computer vision to tens of millions, and around 150 million parameters was a huge model for computer vision. Now we routinely talk about trillions. Transformers are the link that took us from millions to trillions.
Why they stick around. If the units of scaling are data and compute, transformers survive because they are sponges: increase the compute available to a transformer 10x and you get some log improvement somewhere. So far the scaling laws really work, and he calls them quite beautiful.
And the specific property that has resisted replacement. Every attempt to replace attention runs into this: transformers represent any pairwise relationship you want. Any token in the sequence can attend to any other token. If there is any relationship in the sequence at all, you will find it. It may be true that you do not need all to all modeling, that not every token needs to look at every other one, but until we know a better way to prune that space and make information modeling more selective, attention is a very good operation.
The Karpathy rule, from the computer vision days. He cites Andrej Karpathy for the recipe: with a new dataset, your first goal is to overparameterize the model and overfit the data you have, to prove that a relationship exists that you can model or memorize, and that your learning algorithm actually instills knowledge into the model. Once you can overfit, then you compress, and the compression is how you get generalization. You do not want to memorize the data in front of you. You want to work backward and find the general patterns that fit into the smallest parameter count possible.
The future of AI training data (36:14)
Patrick asks for a prediction on data. Movva opens with a framing worth keeping:
I like the phrase that internet was a one time subsidy on data. We got it for free.
The sizes he gives: about 30 trillion tokens of high quality text, or 300 trillion if you take a wider view of what qualifies as good text. And we have basically looked at it all already. Models have seen the entire internet many times over. There is not much more to be done on human data from the internet.
The next phase is model self improvement through RL environment gyms. And he makes a point that is easy to miss and genuinely important: we do not even benefit much from more random user interactions with AI anymore. It used to be that the valuable new data was interaction data from people using ChatGPT and giving thumbs up or thumbs down signals. His argument now:
The median model that we serve is so much more advanced than a random human giving feedback, that the signal you get from unconditioned human preference is not actually worth anything anymore. You want expert human preference at this point.
The model has outgrown the everyday Joe.
So the future of data is: give the model a hard verifiable task, let it run in a gym where it is isolated, where it has a problem it can make progress on and a way to measure whether it made progress. Coding problems are in that category. Math problems are in that category. And increasingly you can just give the agent a computer, have it act like a human worker, and give it feedback on whether it is making progress toward the target outcome. The environment becomes the data. He is careful to say this is not a super differentiated take, but that it has been extremely productive.
Stacking specialized intelligences (38:14)
Patrick asks whether RL environments are just another big block like the internet, one that will have its day and be exhausted. Movva thinks it is more profound than that:
If you want artificial general intelligence, the best way to get there is to just keep stacking specialized intelligences until you have no more gaps to fill.
There is exactly one requirement, and it is a hard gate: your task must be verifiable. You must give the model a self grading system. With that, you have the recipe for self improvement on any task you like.
His evidence is a spend shift, and it is checkable: frontier labs used to spend heavily on data, and now they spend a lot more on RL environments. Those environments capture the recursive self improvement relationship on a verifiable task.
Back to kernels (38:44)
Patrick steers back to the software layer. Keep going on what Sail has done and what it wants to do.
Kernels are not done, which surprises people. There are great people, he names Tri Dao, who write excellent kernels that form the bedrock of modern deep learning. Modern transformers are built on FlashAttention. But deviate from the happy path at all, and it falls over. A new model ships with a slightly different way to embed positional information, a change to the RoPE system, and suddenly the kernel you had is not suitable and needs a patch.
His honest calibration of the state of the art: we are not in the phase where you have to invent new kernels from scratch. The value is in being able to quickly modify existing ones.
And then the definition, dropped mid sentence for the general listener: a kernel is a general term for any program you run on the GPU. Historically kernels are collected in a library where each has a very scoped purpose. You have a kernel for a matrix multiply. You have another kernel for something as simple as addition, because adding two tensors together is its own kernel.
Fusion is the optimization on top of that. If you do a matrix multiply and then want to add the result to another matrix you also multiplied, those two can become one kernel. The reason that matters is memory traffic: instead of writing the intermediate result out to DRAM and reading it back in just to do the addition, you keep it where it is.
Why humans still write kernels, and how they actually do it now (40:15)
Patrick's instinct is that this is exactly the sort of thing AI should be excellent at. If we are not there, why not?
Movva does not want to speak for Tri Dao, but relays what he learned from him: you should not write kernels by hand anymore, necessarily.
I like to say we write kernels on the whiteboard. We go to the whiteboard, we describe what we think the machine should be doing, then we succinctly describe that in natural language to a model. And then the model is able to do the execution of, okay, here is my input and output, here is the strategy of how we want to dispatch this work onto the GPU, I'm going to go implement this.
Humans do the conceptual design. The model does the execution. He is not sure why models are not superb at the conceptual half yet, does not think it is a moat, expects much better kernel engineering models within six months, and notes that the labs would tell you they already do a lot of their kernel engineering in a fully automated way.
Software is a fading edge, and he says so (41:15)
Patrick puts the uncomfortable question directly: if software trends toward speed of light usage of the hardware, then software stops being an advantage for a company like yours over time.
That's right. The rising tide of something like Mythos or GPT 5.6 lifts all boats. It really does.
He goes further and rejects the obvious defensive move: he does not think there is a point in specializing to make a model better at kernel engineering specifically. It is not the most meaningful subset of coding. Maybe you inject some privileged information into the prompt to steer a model toward better kernels, but broadly, everyone is downstream of the frontier on this capability.
This is the most important admission in the interview, and he makes it without being cornered into it. Software is where you start because it is cheap, not because it is durable.
How much of the coal are we burning? (42:15)
Patrick offers an analogy from the history of energy: there is always a pendulum between the raw source, coal, and the fraction of the energy in that lump we can actually harness. A big part of energy history was moving that number from 10% to 95%. So if Blackwell is the lump of coal, what percent are we at?
Movva splits the answer, and the split is the whole point.
On the happy path, we are already efficient. When the GPU is doing what it most likes, a large dimension matrix multiply, that operation runs at 70 to 80% of peak utilization, and it is limited not by software but by power. He adds a detail worth keeping:
The way NVIDIA quotes peak FLOPs is a little optimistic. You never hit that because of power throttling.
Patrick: because of heat. Yes, thermals. Call it 70 to 80%, saturated, pretty good.
Off the happy path is where the loss lives. In practice you do not spend the majority of your time in a transformer doing a large batch matrix multiply. So the job is to build the engine around the chip such that the GPU is fed large batches of work at all times. That is the real efficiency claim: not a faster matmul, but a serving system that keeps the chip in the regime where the matmul is already fast.
Rack scale is the new unit of programming (43:46)
One of the most profound transitions in the GPU world in the last year: you do not program one GPU at a time anymore. You should think about the whole rack, and maybe the whole cluster, the entire data center at a time.
NVIDIA now ships not a single GPU or a single motherboard but a prescribed rack system: NVL72, where the latest Grace Blackwell 300 ships as a rack of 72 units. It is an open race to figure out who can program the whole rack scale computer most efficiently, and his belief is that this shape of compute is the future of both efficiency and speed.
He also gives NVIDIA its due without qualification: right now, if you want the lowest possible latency you should use that chip, and if you want the highest possible throughput you should probably also use that chip.
Chip scarcity and compute arbitrage (44:32)
Patrick raises the thing everyone says in private: the market for Blackwells feels like a drug market. Everybody is short, and the things people do to get allocation are fascinating. React to that, and then tell us what the market looks like a rung or two down the quality ladder.
On NVIDIA's allocation behavior. They take a long term view. They see immense demand and could do what suppliers have always done: crank prices until supply and demand intersect and everyone is technically happier. They do not, and his read on why is strategic rather than generous:
If they just let the most deep pocketed buy all the chips, that maybe hurts them in the long term if that Customer ends up accruing a lot more power. They understand that compute is power today.
On the mechanics of getting allocation. Relationships matter enormously. Nobody wants a huge chip rental order from a new startup saying it will rent 10,000 Blackwells for three or five years when the startup has only been operating for months. Who knows if they are good for the money. Getting access to compute requires either great relationships or incredible financial backing.
No bad chips, only bad pricing (45:49)
For everything that is not the bleeding edge, he refuses the word inferior:
I like to say there's no bad chips. There's really bad pricing. And I will make any chip work at the right price.
On AMD. Great chips overall. The challenge is that people do not understand how to program them well. NVIDIA is pretty good at writing the best kernels for its own chips out of the box, so the alpha there is thin. On other vendors' chips, the vendor does less of that work, so there is much more alpha available. And better still for him, other people carry a perception that AMD is not as good as NVIDIA:
That's music to my ears. I'm very happy for them to sleep on this chip and for me to buy as much as I can.
He then corrects himself in real time, which is a good sign of honest analysis: that is not really true anymore. AMD is somewhat popular among some large buyers. Publicly, Meta and OpenAI have bought a lot of AMD chips, and increasingly all the AMD supply is allocated too.
On the long tail of new entrants. He names Etched, SambaNova and d-Matrix as examples of net new companies popping up, and identifies their single hardest problem precisely: scale. Can they actually get enough wafer allocation from TSMC to pump out chips and make it into the market? His posture toward all of them is the same: if there is a new chip on the horizon, he wants to know as fast as possible and evaluate whether he can buy a good fraction of that supply.
The arbitrage, stated plainly (47:20)
Patrick names the business model: if you are much better at eking performance out of chips that have received less attention, you resell that at a margin, and it is a great business.
Exactly. But he pushes back on the flattering version of it. It is not that everyone else has a skill issue. He thinks Sail is genuinely one of the best teams in the world at using multiple silicon architectures and chasing performance in unlikely places, but the real advantage is speed and the absence of incumbency. He is not locked into data center providers that only carry one class of chip, where deploying something new would be a huge pain. He has creative data center partners willing to move fast. And, in his framing, he simply does not shy away from the work:
That's frankly a big part of this: just saying yes, we love TPUs, we're going to make TPUs work. Yes, we love Trainium, we're going to make Trainium work.
And if it does not work easily, it gets fitted into the heterogeneous serving system anyway. Every chip has a comparative advantage. His job is to find that advantage and squeeze in that direction. Google's TPU and AWS Trainium are both named as things he will make work.
The investor's worry: semiconductors at a fifth of the S&P (48:53)
Patrick describes the chart he keeps coming back to. Historically semiconductors were 2, 3, 4 percent of the S&P 500. Now they are 19, 20, 21 percent. A student of market history has seen that shape before: crazy near term peak, then collapse back to long term norms. Memory stocks are doing the same thing. Everybody made a lot of money in Micron and SK Hynix, everybody acknowledges a huge shortage, and everybody quietly assumes it reverts.
Movva's answer starts with a disclaimer that is also a good line:
I'm less of a student of history, more of a member of history. I was born in 1997.
His mother worked at Intel in the run up to the year 2000, through the dot com boom and crash. He remembers the period when Cisco was the most valuable company in the world and Intel was close behind. He draws the parallel deliberately, and then names the one difference he thinks breaks it:
A lot of the investment in networking equipment historically was speculative. We anticipated this future demand for users that never came. What's interesting about token consumption is that it's no longer speculative. People buy tokens because they're immediately valuable to them. You don't hoard tokens, you use them immediately.
He extends the same test backward to the recent past to show it is not special pleading. The Hopper supply crunch of 2023 and 2024 was training oriented spend, and training is inherently speculative. Today is different in a way you can observe from the outside: everyone is instituting caps on how much you can spend on Claude Code. That is what demand exceeding supply on an immediately consumed good looks like.
I do think inference spend monotonically increases. There's no speculation on inference spend.
That is the load bearing claim of his entire capital story, and it is the one to stress test hardest.
Reinventing the AI data center (52:44)
Back to hardware, now at the unit above the rack.
His theme reappears: the difference between training and inference explains everything downstream. The most conservative players in the entire AI stack are the infrastructure players, data centers, and even more conservative than them, TSMC and the chip infrastructure people.
So AI data centers were built for training, and training is the superset workload. You can make any training cluster serve inference, but not necessarily the other way around. The difference is networking: how much you invest in bandwidth between chips, and how large a cluster you need.
And there is a diseconomy of scale. It is far more expensive and difficult to build 100,000 GPUs in one data center than 10,000, which is harder than 1,000. Everything is now denominated in megawatts and gigawatts, and his ladder of what is actually buildable in the United States is the most useful set of numbers in this segment:
- A gigawatt data center: basically no way to build one easily anymore.
- 100 MW: increasingly hard, basically impossible unless you are a very special set of Customers.
- 10 MW: probably on the edge of what is possible today.
- 1 MW: plentiful.
Out of that comes the arbitrage he cares most about:
There's this incredible lore on the market where you can find lots of aggregate power but it will not be concentrated.
That was uninteresting to anyone building training data centers, because you assume everything is in one spot and nobody wants to deal with cross data center training. And the market still has lag in it: he still hears from data center developers the attitude that the job is to go shake down 100 megawatt and 10 megawatt sites wherever you can find them. But a few new thinkers are realizing that inference is suitable for distributed 1 megawatt data centers, and Sail is very happy to buy small pools of compute across the United States and use that as its inference fleet.
What a megawatt physically looks like now (55:01)
Patrick asks for the physical scale, and the answer is a good illustration of how much liquid cooling changed the picture.
This got wonky with the advent of liquid cooling, because you can now pack insane power density into a single physical rack. A megawatt of compute used to be a massive data hall, a warehouse. Now:
You can actually pack that into around like eight racks worth of compute. Each rack is about the size of a refrigerator. You can just imagine eight of them lined up. That's a megawatt.
Patrick summarizes the full strategy back to him: a whole bunch of different chips used together, bought opportunistically, coupled in very small data centers doing only inference, and those two moves together, random cheap compute plus small units of expression, equal much cheaper intelligence. Movva certainly thinks so.
One of the ways that I describe what we do is we will buy any chip anywhere in the world for any duration of time.
He claims that as a level of flexibility and liquidity nobody else has right now, and says Sail is aggressive about putting money where its mouth is.
Buying 95% uptime (56:33)
Here is the part where his strategy stops sounding like a preference and starts sounding like a genuinely different business.
Set up an army of a thousand small data centers instead of one gigawatt campus, and here is what you give up, in his own list:
- No power redundancy, quite often.
- No backup diesel generators on site. Those are very expensive. He cuts the overhead.
- No redundant networking in a lot of cases.
- One trenched line of fiber, not three lines with redundancy and failover and SLAs.
It's just going to go down sometimes. In fact, I won't be surprised if some of them get down to like 95% uptime.
Patrick: which is bad. Movva: very bad. Fatal, atrocious, for anyone else. You would have basically zero buyers for a data center at 95% uptime.
I'm that first buyer. I will buy 95% uptime.
Two reasons he can, and only one of them is the obvious one.
First, the control plane. Sail has a robust control plane that handles any single failure in any single data center, as long as the failure is not correlated with other data centers, by moving the workload somewhere else. And critically, his utility function is not a cliff:
The failures happen at some rate and I am basically linearly happy with a data center that's 95% uptime versus 98% versus 99%. It's just linearly good or bad for me.
That is the whole thing. For a low latency provider, uptime is a step function tied to an SLA. For him it is a straight line, so he can price it and buy the cheap end of it.
Second, the async workload. When a request fails he has to find a new GPU for it, so for that single turn of the agent's work, an agent that has been running for an hour hits a roadblock because its GPU got pulled, and experiences an extra minute or two or three, maybe even ten, of latency. His argument is that his Customers do not care, because the agent was running for hours and they are asleep.
We tell our Customers, look, our average throughput is going to be very competitive, but our P99, our 99th percentile latency, it's not going to be controlled. It cannot be. And in return, I'll give you unbeatable economics.
That is the trade stated as cleanly as it can be stated: he is selling uncontrolled tail latency in exchange for price, and background agents are the one workload where that is an acceptable purchase.
Power and the scavenger strategy (59:01)
Patrick asks about power as a category. Movva escalates his own position first: he said 95% uptime, but could he take 80% at the right price? Probably.
And then the reason he wants that flexibility:
I'm a son of California. I love solar and wind. I think solar and wind power is way undertapped in the United States.
The challenge with renewables and data centers has always been intermittency: a persistent base load against an intermittent source. His claim is that we are not far from solving that problem, for a specific and narrow definition of "solving" that only works for his workload.
I am totally capable of tolerating an outage for my data center that's measured in even days or weeks, which is the worst case nightmare scenario for a data center.
The worst case is a long term outage because the wind is not blowing and the clouds are in the sky and fog is hanging over the valley for some time. His answer: that is in fact highly predictable. Model the weather, figure out when the data center will be offline, move the workload elsewhere, and it is fine.
The trick is that it's going to give me better access to power that no one else is going to touch, because it is so annoying to deal with that kind of outage.
And then the line that closes the loop on the whole capital argument, in a single conditional:
If my chips are cheap enough, they're probably not going to be NVIDIA racks. And if my chips are cheap enough, I don't mind the capital cost of having idle chips.
That is the crux of the scavenger strategy, and it is worth stating explicitly because he does not. Idle time is only tolerable when the asset is cheap. Nobody idles a Blackwell rack for a week waiting for wind. You can idle a mispriced chip on a stranded solar site for a week, because the depreciation you are wasting is small. Cheap silicon and unreliable power are not two separate savings. They are the same decision, and neither works without the other.
Scavenging, defined (1:00:09)
Patrick asks him to unpack the phrase.
First we scavenge chips and then we scavenge power for those chips. The idea is in both cases I do not want to be bidding against Anthropic or OpenAI for compute capacity. I'm not going to win against them and I don't want to. I want to be more creative and use the supply that they don't find legible today.
Over time he amasses enough aggregate supply. He is explicit that he will never get concentrated supply, only aggregate supply, and that over time the aggregate factory becomes unbeatable on economics.
His own metaphor for it, which is better than the scavenger one:
We are building a factory. We're trying to build the best steel factory in the world. But it will come through mini mills, not through large monolithic steel plants.
How vertically integrated should this be? (1:00:39)
Patrick lays out the spectrum. At one extreme you own everything: the power source, the data centers, your own chip designs, the software that eats the most out of those chips, and you sell the finished token to the end user. That is enormously capital intensive. At the other extreme you own nothing and are purely the coordination plane, the virtual scavenger.
Movva says there are two parts of him that hear that question.
The founder. Wants to do everything. This is his entire life. He has spent it thinking about chips, power and energy, so of course he wants to be maximally ambitious and will never stop until he has built the most efficient system from soup to nuts. Patrick's read, which Movva accepts happily: you are playing real life Factorio.
The CEO. Has to be more pragmatic, because the capital required to own everything is, in his word, insane. Software has high leverage, so you start there. Then the real questions: do you own power generation, or can you get great power purchase agreements with utilities? He is more inclined to let other people specialize in what they are historically good at, and to see if Sail can reach scale.
The framing he uses for that is the right one for a young company:
I want to get to the scale where I earn the right to take this under our wing.
And the justification for eventually breaking other people's assumptions:
If you can break the assumption that people I would be buying from made about who their Customers would be, I maybe break those assumptions.
He grants it is an optimistic view, and says it is only possible because Sail is trying to underwrite the largest market for compute in the history of computing, billions and trillions of dollars of investment into inference. Because of that focus, it makes sense to build a lot of things custom for inference. His job is to find every place where that is possible, get partners to build custom things for him, and if they cannot, do it himself.
Stack ranking the inefficiency (1:03:11)
Patrick asks the best question of the interview: zoom out on the entire system, software, hardware, energy, and rank where we are most inefficient at producing useful intelligent tokens.
The answer is unusually specific, and it is where you should look if you want to check the 1000x.
Compute scaling: efficient. Not the problem. Give him more FLOPs and he will use more FLOPs, and he says the industry is already fairly judicious with them.
Sparsity: mostly harvested already. His numbers: a modern model is rarely more than 10% dense, meaning 10% of the possible experts you could activate are activated, and frontier models are closer to 1%. Mixture of experts has been worked on for a long time and people are pretty good at squeezing it. He does not think we are wasting much on the MoE side.
Attention and its use of memory: this is where we are bad.
Specifically, the KV cache is quite uncompressed right now. If you look at the entropy in a KV cache, it's not earning its keep. We're storing many kilobytes of data in the KV cache per token, and that's probably off by an order of magnitude or two.
He does not know what the frontier labs do internally, but he points at DeepSeek as publishing genuinely interesting work on compressing it further. And his inference from the rate of publication is the useful bit: the fact that they are able to make order of magnitude progress here every year or so signals that there is a lot more room to go.
That is the one place in eighty four minutes where he attaches a multiplier to a specific mechanism. Ten to a hundred times, on the memory footprint of attention.
Zoom out, and the biggest waste is not technical at all: it is allocation.
We actually don't marshal our compute effectively at all. NVIDIA is pumping out 5 million Blackwell chips this year. Where are they all going? Are they all being used at all, all the time? I certainly doubt it.
The obstacle is that much of that compute disappears into private pools that will never see the light of day, and those GPUs sit idle. He says it pains him physically: silicon and power went into that chip and it is sitting there doing nothing. What he wants is better orchestration of the world's compute as a shared resource, packed more efficiently.
And he gives the comparison that makes the point land:
We all make fun of xAI for having some challenges with total FLOP utilization on its clusters, but the reality for the rest of the world is far, far worse. A ton of GPUs just sit in warehouses or sit in private pools allocated to a specific Customer and don't get utilized.
Fabs, and the process corner nobody offers him (1:05:43)
Patrick asks about fabs themselves. If we snapped our fingers and had 100 times the chips, we would have far cheaper tokens, so what is the future of fabrication, and will the expansion happen in the US?
First, the systems answer. Everything grows in balance. Snap your fingers and double all of it, and you might fix a TSMC bottleneck only to hit another bottleneck immediately. Make 20% more chips and something else binds.
Then the genuinely novel idea, and it is the most interesting thing he says about manufacturing. What interests him is what a fab considers a must deliver, an invariant its Customers will always want, versus what he considers a fluid relationship. If a fab exposed more of its trade offs to him, he could make more intelligent decisions.
His example: process corners.
Any fab has a lot of spread between the worst chip that comes out of the production line and the best chip. TSMC works very, very hard to tighten what we call these process corners. They want to keep the worst chip as close in characterization to the best chip.
They are excellent at it. But tightening the corner means adding process controls he may not need. Maybe he is willing to find a home for that worst chip. You do not need to tighten process control as much, which takes more time and cost. Maybe he is willing to take a lot more rejects.
For him it is a holistic optimization across the cost of the dies, the supply of the dies, the cost of power, and the places he can put them. And the goal ties back to the power strategy:
My whole goal is to dramatically expand the supply of power across the United States so that I have a home for a lot of chips that otherwise would not have earned their place in a data center.
That sentence is the scavenger strategy applied one layer further up the supply chain than anyone usually takes it: not just buying the chips nobody wants, but buying the dies a fab would otherwise throw away.
How he runs the team (1:07:45)
Patrick asks about the culture and the structure of a company with this north star.
In the limit thinking. They do not worry about the immediate state. When they start on a model, efficiency is not going to be good. What matters is where they could be in a month, six months, a year.
Nothing is fixed. They do not accept the state of the machines they work on as given. Even for something like Blackwell, if there is a bottleneck holding them back from the performance they think is achievable, it is very important to Movva that they understand and characterize it well and write it down, so they can tell NVIDIA and friends about it, and so they can keep it in mind for future chips they buy. The goal is to learn things that are essentially invariant for the company long term and fold them into future decisions.
Students and teachers. Highly collaborative. One of the most important traits they look for is people who are both good students and great teachers. A lot of the team were TAs in college and loved sharing knowledge that way. They do whiteboard sessions constantly, and he thinks the collegial environment where everyone has something to teach and something to learn is extremely important.
And on hiring for a world where machines do more of the work, his answer is one word.
Curiosity. It's 100% curiosity. The one thing I cannot teach is love for performance, love for digging into every microsecond that the machine is working and understanding what's happening on the machine at that time.
He then throws out the credential everyone would expect him to screen for:
I don't look for lots of AI experience. I don't look for CUDA experience at all. That's actually a huge red herring. CUDA as a concept, GPUs as a concept, have evolved so much in the last five years. There's no point asking for 10 years of experience. I want to teach that. But I cannot teach the love for performance engineering.
Open against closed (1:10:10)
Patrick asks for an assessment of the major labs, and of the relationship between closed source as a category and open source.
On the premium for being ahead. The labs pay an immense premium to be three to six months ahead of everything else, and he thinks that is probably still worth it. It makes perfect sense for OpenAI and Anthropic to do what they do.
On distillation, where he offers a deliberately contrarian view. There is a sense in the industry that distillation is theft, that training on a frontier model's outputs takes something from it. His counter does not argue about ethics. It argues that the category has already dissolved:
An increasingly large percentage of the artifacts we put out on the internet are AI generated. Even if you just look at GitHub alone, what percentage of repos created in the last year do we think were created by Claude Code? Do we consider that to be distillation? Because that's probably all we need.
He goes further: he would not be surprised if you could train a Fable class model only on the outputs of code you consider good on GitHub that is open source. And if we take the position that users own the outputs of their interaction with AI and choose to publish them, then we will have latent distillation for a long time.
It seems fundamentally impossible to prevent the diffusion of information or model capabilities. It will happen. The question is just how fast.
On whether the three to six month lead is defensible. Patrick puts the optimistic case for the closed labs: if scaling and improvement laws hold for a long time, being three to six months ahead has real value and you can charge a huge premium for those tokens relative to a very cheap open source token.
Movva says it is possible, but he does not think the premium lasts, and his reason is about buyers rather than technology:
If you look at enterprise deployments, they don't move at three to six month speed. A lot of enterprises are probably still on Opus 4.6 or Opus 4.7. They don't adopt the bleeding edge rapidly.
We are so early that he does not think there is any way to call a winner, and he does not think this is a race that can be decided ever, because it is a continual process. And on open source specifically, his claim is structural rather than sentimental:
Fundamentally I don't think open source ever goes away. If there's a vacuum because one leader steps out, a new leader will step in. There's too much incentive, and there's a lot of tailwinds too. It just gets easier every day to train a frontier class model.
Abundant tokens and diverse harnesses (1:12:50)
Asked what he hopes the future looks like, he gives a five word answer and then unpacks it:
I want abundant tokens and diverse harnesses.
He wants everyone to build their own harness. Every company. Every user. Make the agent your own. He thinks we are not far from that level of customization and capability, and he expects it to arrive not through weight fine tuning but through in context learning, which he flags as a technical detail but which is a real prediction about where personalization lands.
Then the passage that opens the episode, in its place:
My job is to make the tokens as cheap as humanly possible. I will achieve that and I will do it through every layer in the stack available to me. I love the supply side levers. I will use every chip, I'll use every source of power, and I will use every piece of land in the United States that's suitable for this. And in return, people will have the incentive to explore what it's like to have abundant intelligence. We still treat the agent as a person that is expensive to consult and you should ask them when you have a hard question. That's not the way to think about intelligence.
The views that make his friends look at him funny (1:13:51)
Patrick asks which of his ideas make well informed friends look at him like he has three heads.
Chips, mostly. When he talks about building custom chips and people ask what would be different, the answer is: sidestepping the HBM shortage and focusing on more extreme offload to other forms of memory, such as flash. He is passionate about it. Everyone on his team knows he keeps banging the drum on the same question:
What would we have to change about the model architecture to make offloading KV cache to flash work at a much greater level?
He whiteboards it constantly. Note how neatly it closes the loop: the KV cache is his one quantified inefficiency, HBM is one of his four named supply bottlenecks, and flash offload attacks both at once, at the price of a memory tier that is far slower, which he can only afford because he serves at one to ten tokens per second.
And in the inference community specifically, the divergent view is about what becomes possible if you design the whole system around serving at one to ten tokens per second. That is his north star, stated as a number.
The trillion token arithmetic, and where the 1000x actually comes from (1:14:30)
The broader question he is really asking is: how do people consume a trillion tokens per day? That is the world he wants to create the capability for. Patrick, correctly, asks him to ground it. What is a trillion tokens?
A trillion tokens, well, okay, at OpenAI pricing that's at least $5 million at the very least, for 5.5 or 5.6.
Patrick: so what is the world in which we consume what currently costs five million dollars per person per day?
And here is the number the title is built on, in his own words and with his own units:
We were asking for at least three to six orders of magnitude improvement in cost per token. Get that into $5,000, you probably have some Customers.
So the 1000x is: a trillion tokens going from at least $5 million at frontier API pricing to $5,000. That is three orders of magnitude, the floor of his range. The ceiling of his range, six orders, puts the same trillion tokens at five dollars.
He also gives the current state of the art on his own side, which is the most checkable claim of the whole episode:
In fact, I would argue that for some size of model, we are approaching a trillion tokens being measured in tens of thousands of dollars. And that's something that you can imagine running for a single job.
Read those two sentences together and the shape of the claim becomes clear. He is not saying frontier intelligence gets 1000x cheaper. He is saying a trillion tokens from a smaller open weight model already costs tens of thousands of dollars rather than millions, which is most of the first two orders of magnitude, and that the remaining distance to $5,000 comes from the levers in this interview.
Is there even demand for that much intelligence? (1:15:45)
Patrick asks the deflationary question, and it is a fair one: maybe the average person cannot and will not do that. They do not do it now with their own brain. Maybe there is not that much demand for intelligence in the world.
I never will believe in that. There is always demand for intelligence in the world.
He locates the problem elsewhere: the on ramps to that intelligence are a product challenge, and he is careful to say he is not a product person and cannot claim the best vision there. What he wants is for those people never to be held back by the sense that free tier users cannot use this, or that a company cannot afford to give them this many tokens. He says he hears exactly that from his Customers all the time, and that is what he wants to fix.
The consensus he thinks is wrong: process nodes barely help (1:16:25)
Patrick asks the inverse of the crazy ideas question. Not what is craziest, but what consensus thing do you think is wrong?
He keeps coming back to NVIDIA. He is bullish on NVIDIA in the short term and says you should never bet against them, they always reinvent themselves. But:
One thing that surprises people is when I tell them that if you look at Hopper to Blackwell to Rubin, and you compare like for like, what is the performance per watt of a BF16 multiply, it hasn't improved all that much. Or you take that one step further, go to TSMC, and if you look at TSMC 5 nanometer versus 4 versus 3 versus 2, the performance per watt on these chips doesn't change a dramatic amount.
This is the single most consequential claim in the interview for anyone modeling the cost curve, and it cuts against him. He is saying the lever most people assume is doing the work, process shrink and generational silicon, is nearly flat on the metric that matters for a token factory. Every order of magnitude has to come from somewhere else.
He then follows it to the geopolitical conclusion, which is where he expects the pushback:
People lose their minds over geopolitics, what would happen if we lost access to TSMC for any reason. My contrarian take is that it wouldn't be that bad. Supply would take a shock for sure, but the best processes that we have in the West, like Intel, are not that far behind. At worst, maybe 2x worse performance per watt. The gap is just far smaller than you would make it out to be if you follow the chip world dialogue.
What is not in his path (1:17:26)
Asked what is happening in AI outside his lane that interests him most, the answer is model architecture, because he is fully downstream of it.
The model people decide how to design their architectures, and he has no input into OpenAI or Anthropic. He can only pray they go in a direction that is amenable to him, or do his best to predict where they are going and build his serving architecture accordingly, across both software and hardware choices.
What makes it interesting is how enormous the downstream consequences of small looking decisions are:
How consequential it is to decide to use something like sparse attention versus dense attention. Or how consequential it is to use a different data type. We were training in BF16, but now we can train in FP8 or FP4, lower precision data types. That is just an arbitrary choice, it feels like, but it has profound implications for what chips I can use and how I should build my hardware.
A datatype decision made in a lab reprices his entire fleet. That is the clearest statement of his structural position in the value chain, and it belongs in any honest assessment of the 1000x, because half the levers sit with people who do not work for him.
Advice to a hundred would be chip founders (1:18:26)
Patrick sets up a hypothetical: a hundred entrepreneurs who all want to start a compute company, specifically hardware, chips, systems, racks. What advice on how to orient the business?
It's all about the bottlenecks on supply chain.
You need to convince him, or an investor, that you understand the three to five bottlenecks that dictate modern chip supply. His list:
- TSMC wafer capacity.
- HBM capacity.
- Advanced packaging.
- Power. Where will you get it, and how will you build these racks?
You should have a great answer for each, because in his words it is all arbitrage at the end of the day. You are building a chip because you think NVIDIA has made choices that are difficult for them to change, which is true. NVIDIA makes many choices that are difficult to change. They are not perfect. They are just really well balanced.
So the advice is: be spiky.
You want to pick something and say, I think they've underpriced the impact of how short we're going to be on HBM. We're going to push really hard in this other direction instead.
And as an aside he names HBM as probably the thing to attack most. Why? Because there is no easy way to bring on a lot more memory fabs, and the memory companies have been burned too many times:
The boys in Boise don't love huge capex for cyclical.
Patrick: so conceivably, because of that shortage, the world routes around it by making everything else in the system more efficient. Movva disagrees, bluntly and with a consumer example:
I think they're going to make everything else more expensive. I think that iPhones will cut their memory. iPhones are going to go up in price and we're just going to deal with it.
Why NVIDIA does not just sell tokens (1:20:14)
Patrick asks the obvious vertical integration question about NVIDIA itself. Why not go all the way to the end and sell tokens? Why not even start a neocloud and sell compute out the back door?
NVIDIA is really smart about this. They don't compete with their Customers. NVIDIA takes the long view on everything.
And then the sharper version:
Jensen is really good at making his friends billionaires. He's made CoreWeave a many billion dollar company. There's no need for him to destroy that goodwill.
What Jensen wants is a diverse community of neoclouds and inference providers all jockeying to create demand for NVIDIA, so that if any one of them decides to vertically integrate or go with AMD or any other option, there are three more hungry people ready to fill that position. Patrick's summary: it is great to have competition among his buyers.
The closing question (1:20:58)
Patrick's standard closer: what is the kindest thing anyone has ever done for you?
Movva's first thought is mentors, the rare person who takes time out of their schedule and makes it a personal interest to make sure you understand something, or to instill a value you were on the cusp of understanding and just needed a push over the line. He names the people at NVIDIA who gave him the love of performance engineering, and then tells a story about his college advisor.
He was an impatient sophomore. He showed up at office hours and said he wanted to build AI chips, he knew what he wanted to do, so why was he wasting time on basic classes in networking and operating systems?
He just looked at me and laid out the whole stack and showed me the beauty of understanding every piece in the puzzle. He took my entire path of trying to focus on one piece of the system and said that it's so rare that someone can actually understand the entire stack, from the gate level silicon all the way to building a great internet scale service, and you should aspire to be someone who over the course of your lifetime achieves that level of understanding.
He calls that a rare trait and a noble level of expertise to chase, and says it has stayed with him. Given that the previous eighty minutes covered 6T SRAM cells, kernel fusion, rack scale scheduling, power purchase agreements and fab process corners in one continuous argument, the advice appears to have taken.
Where the 1000x comes from
He never multiplies the factors out on the show. The only number he commits to as a total is "at least three to six orders of magnitude," and the only mechanism he attaches an explicit multiplier to is the KV cache. So the table below carries both: what he quantified, and what the rest of his argument implies once you lay it against the standard list of cost levers.
| Lever | What it actually changes | Multiplier | Whose number |
|---|---|---|---|
| Model choice | Serve a smaller open weight model instead of a frontier API model. Same trillion tokens, different intelligence | Roughly 100x, from at least $5M to "tens of thousands" per trillion tokens | His, both endpoints, and he says it is already banked |
| KV cache compression | Attention's memory footprint. He says we store many kilobytes per token and the entropy does not justify it | 10x to 100x, "off by an order of magnitude or two" | His, explicitly. Cites DeepSeek making order of magnitude progress "every year or so" |
| Throughput serving | Batch wide and slow at 1 to 10 tokens per second instead of protecting latency. Keeps the GPU on its happy path | No multiplier given | Reconstruction. His supporting figures: 70 to 80% of peak on a large matmul, and 8 GPUs buying only 4 to 5x on latency |
| Chip arbitrage | FLOPs per dollar per hour, owned and operated. AMD, TPU, Trainium, Etched, SambaNova, d-Matrix, anything mispriced | No multiplier given | Reconstruction. His claim is only that some chips "rank higher than NVIDIA on FLOPs per dollar" |
| Data center and power | 1 MW distributed sites, no diesel, no redundant fiber, no SLA, intermittent solar and wind, 95% uptime accepted | No multiplier given | Reconstruction. This is a TCO lever, not a FLOPs lever: same silicon, cheaper hour |
| Fleet utilization | Orchestrating the world's idle compute. 5 million Blackwells ship this year and he doubts they run all the time | No multiplier given | Reconstruction. He names it as the largest inefficiency once you zoom out past the chip |
| Kernel and software efficiency | Speed of light on the happy path. Fusion, parallelism scheme, hiding communication behind other work | At most about 1.3x on the happy path, since 70 to 80% of peak is already reached | His 70 to 80% figure, our arithmetic on the headroom. He also says this edge is not durable |
| Quantization | Bytes per parameter and per KV entry. BF16 to FP8 to FP4 is a 4x narrowing of the datatype | No multiplier given | Reconstruction. He names the shift only as an upstream lab decision that reprices his fleet |
| Sparsity and MoE | Fraction of experts activated per token | About 1x. Ruled out | His. Modern models are under 10% dense, frontier closer to 1%. "I don't think we're wasting too much" |
| Process node and silicon generation | Performance per watt of a BF16 multiply, Hopper to Blackwell to Rubin, TSMC 5 to 4 to 3 to 2 nanometer | About 1x. Ruled out | His, and it is his contrarian take. "It hasn't improved all that much" |
| Open weights closing the gap | Keeps the cheap tier near the frontier, via published models and latent distillation from AI generated public code | No multiplier given | Reconstruction. He argues only that the 3 to 6 month frontier premium may not last |
| Timeframe | When any of this lands | Never stated | Not given anywhere in 84 minutes. The 1000x is an ask, not a dated forecast |
Key takeaways
- The guest is Neil Movva, co founder of Sail Research, previously a GPU kernel engineer at NVIDIA from 2016. Sail is an API serving open weight models plus long running agent sandboxes, aiming to be the cheapest token in the market.
- The 1000x is a floor, not a forecast. His actual statement is "at least three to six orders of magnitude improvement in cost per token," anchored to a trillion tokens going from at least $5 million at frontier pricing to $5,000. He never attaches a date.
- The product bet comes first, and everything else follows from it. Background agents that run for hours mean nobody is waiting, which means throughput beats latency, which means NVLink is optional, which means non NVIDIA silicon is viable, which means unreliable distributed power is viable.
- He rules out the two levers most people assume are doing the work. Performance per watt of a BF16 multiply barely improves Hopper to Blackwell to Rubin, or across TSMC 5, 4, 3 and 2 nanometer. And sparsity is largely harvested: frontier models are already about 1% dense.
- The one mechanism he quantifies is the KV cache, which he says is uncompressed by an order of magnitude or two, with DeepSeek publishing order of magnitude gains roughly annually.
- The original sin of transformers is that a memory bound layer (attention) sits directly next to a compute bound layer (the MLP), and no single chip is good at both. That is why he expects Cerebras and Groq to end up as coprocessors holding the MLP while a GPU holds attention and the growing KV cache.
- Memory, in his figures: Blackwell carries 288 GB of HBM at around 10 TB/s, against roughly 500 MB of on die SRAM. A Cerebras WSE-3 wafer holds about 50 GB of SRAM at 21 PB/s, roughly 2,000 times the bandwidth at about a sixth of the capacity.
- The data center ladder: a gigawatt site is basically unbuildable, 100 MW is nearly impossible, 10 MW is on the edge, and 1 MW is plentiful. A megawatt is now about eight liquid cooled racks, each the size of a refrigerator.
- He will buy 95% uptime, and possibly 80%, because his happiness with uptime is linear rather than a cliff, and because a failed turn costs a sleeping Customer's agent an extra few minutes. He sells uncontrolled P99 latency in exchange for price.
- His four supply chain bottlenecks for anyone building chips: TSMC wafer capacity, HBM capacity, advanced packaging, and power. He thinks HBM is the one to attack, and expects the shortage to raise iPhone prices rather than route around itself.
- Software is where you start, not where you win. He says outright that kernel efficiency is a rising tide that lifts all boats, that he is downstream of the frontier on it, and that specializing a model for kernel engineering is not worth it.
- Data is moving from the internet to RL gyms. The internet was a one time subsidy of about 30 trillion high quality tokens, 300 trillion on a loose definition, and it has been read many times over. The path to general intelligence, in his framing, is to stack specialized intelligences on verifiable tasks until there are no gaps left.
Chapters
- 0:00 Intro
- 0:38 Building a “Token Factory”
- 4:21 The Future of Background Agents
- 13:09 Nvidia and the GPU Stack
- 23:27 Chips, Memory, and Transformers
- 36:14 The Future of AI Training Data
- 44:32 Chip Scarcity and Compute Arbitrage
- 52:44 Reinventing the AI Data Center
- 59:01 Power and the “Scavenger Strategy”
- 1:10:10 Open vs. Closed AI
Notable quotes
"My job is to make the tokens as cheap as humanly possible. I will achieve that and I will do it through every layer in the stack available to me. I love the supply side levers. I will use every chip, I'll use every source of power, and I will use every piece of land in the United States that's suitable for this." Neil Movva, 0:00
"We think that whenever you make something 10 times cheaper, it's a new product category, and we aspire to do that for tokens." Neil Movva, 1:32
"The future of agentic inference is long horizon tasks. You're going to run the machine for hours or days at a time. It doesn't matter if it spits out tokens at 100 tokens per second. Maybe 10 is just fine." Neil Movva, 3:36
"One way to say it is like the best latency is no latency at all. When you wake up in the morning, the work's already been done overnight. You didn't even have to ask for it." Neil Movva, 4:21
"At some point, people started to make this joke that security has become proof of work. When you want secure software, it's really a question of how many dollars did you spend on Anthropic's APIs trying to break into your software." Neil Movva, 7:41
"You have to be willing to spend tokens without any promise of return. That's the unlock." Neil Movva, 9:42
"We kind of fought tooth and nail and got just a tiny bit of die area, maybe like 5, 10%, something like that, for the first generation of these chips." Neil Movva on the first tensor cores, 15:19
"The speed of light is the edge of what's possible. To this day, I tell all my engineers, we're chasing 100% speed of light. I don't care about relative numbers versus the competition. I only care about absolute numbers." Neil Movva, 16:22
"If you want to get downtown in SF, you can take the bus or you can take private transit. A bus, it's going to have to serve many more people and it has to fundamentally do something that works for everyone." Neil Movva on batching, 19:54
"You'll use eight times more hardware, but you won't get eight times the speed. You might get like four to fivex the speed. You're not going to get strong scaling." Neil Movva on tensor parallelism over NVLink, 21:24
"I would say the original sin of transformers is that you've taken this extremely fundamentally memory bound layer and juxtaposed it right next to a compute bound layer." Neil Movva, 31:39
"I like the phrase that internet was a one time subsidy on data. We got it for free." Neil Movva, 36:14
"If you want artificial general intelligence, the best way to get there is to just keep stacking specialized intelligences until you have no more gaps to fill." Neil Movva, 38:14
"I like to say we write kernels on the whiteboard. We go to the whiteboard, we describe what we think the machine should be doing, then we succinctly describe that in natural language to a model." Neil Movva, 40:45
"The way Nvidia quotes peak flops is a little optimistic. You never hit that because of power throttling." Neil Movva, 42:46
"There's no bad chips. There's really bad pricing. And I will make any chip work at the right price." Neil Movva, 45:49
"I'm very happy for them to sleep on this chip and for me to buy as much as I can." Neil Movva on the perception of AMD, 46:50
"I'm less of a student of history, more of a member of history. I was born in 1997." Neil Movva, 49:53
"You don't hoard tokens, you use them immediately." Neil Movva on why inference spend is not speculative, 50:24
"I won't be surprised if some of them get down to like 95% uptime. I'm that first buyer. I will buy 95% uptime." Neil Movva, 57:03
"Our average throughput is going to be very competitive, but our P99, our 99th percentile latency, it's not going to be controlled. It cannot be. And in return, I'll give you unbeatable economics." Neil Movva, 58:35
"If my chips are cheap enough, they're probably not going to be Nvidia racks. And if my chips are cheap enough, I don't mind the capital cost of having idle chips." Neil Movva, 59:37
"First we scavenge chips and then we scavenge power for those chips. I do not want to be bidding against Anthropic or Open AI for compute capacity. I'm not going to win against them and I don't want to." Neil Movva, 1:00:09
"We are building a factory. We're trying to build the best steel factory in the world. But it will come through mini mills, not through large monolithic steel plants." Neil Movva, 1:00:39
"We're storing many kilobytes of data in the KV cache per token, and that's probably off by an order of magnitude or two." Neil Movva, 1:04:12
"Nvidia is pumping out 5 million Blackwell chips this year. Where are they all going? Are they all being used at all, all the time? I certainly doubt it." Neil Movva, 1:04:42
"My whole goal is to so dramatically expand the supply of power across the United States that I have a home for a lot of chips that otherwise would not have earned their place in a data center." Neil Movva, 1:07:15
"Curiosity. It's 100% curiosity. The one thing I cannot teach is love for performance." Neil Movva on hiring, 1:09:16
"I don't look for lots of AI experience. I don't look for CUDA experience at all. That's actually a huge red herring." Neil Movva, 1:09:46
"It seems fundamentally impossible to prevent the diffusion of information or model capabilities. It will happen. The question is just how fast." Neil Movva on distillation, 1:11:18
"I want abundant tokens and diverse harnesses. I want everyone to build their own harness." Neil Movva, 1:12:50
"We were asking for at least three to six orders of magnitude improvement in cost per token. Get that into 5,000, you probably have some customers." Neil Movva, 1:15:23
"If you look at Hopper to Blackwell to Reuben and you compare like for like, what is the performance per watt of a BF16 multiply, it hasn't improved all that much." Neil Movva, 1:16:25
"The boys in Boise don't love huge capex for cyclical." Neil Movva on why memory fabs will not expand quickly, 1:19:58
"Jensen is really good at making his friends billionaires." Neil Movva, 1:20:28
"It's so rare that someone can actually understand the entire stack, from the gate level silicon all the way to building a great internet scale service." Neil Movva quoting his college advisor, 1:22:01
Resources mentioned
People
- Neil Movva, the guest. Co founder of Sail Research, formerly a GPU kernel engineer at NVIDIA from 2016.
- Patrick O'Shaughnessy, the host, CEO of Positive Sum.
- Jensen Huang, NVIDIA CEO, credited with the 2015 and 2016 conviction to allocate die area to tensor cores.
- Tri Dao, author of FlashAttention, named as the person who taught him you should not write kernels by hand any more.
- Andrej Karpathy, source of the overfit first, then compress rule from the computer vision era.
Companies and organizations
- Sail Research, the guest's company. The token factory.
- NVIDIA, his first employer and still the reference point for every comparison in the episode.
- Baseten, Fireworks AI and Together AI, the incumbent open model serving companies he says all optimized for low latency.
- Cursor, the Customer he says pulled the whole field toward low latency.
- Parallel Web Systems, a Sail Customer building a real time index over the whole internet at exabyte scale.
- Anthropic and OpenAI, the frontier labs he does not want to bid against for compute.
- TSMC, the logic foundry, its 800 square millimeter reticle limit, its standard cell SRAM library, and its process corner control.
- Micron, SK Hynix and Samsung Semiconductor, the DRAM and HBM makers. The "boys in Boise" line refers to Micron.
- Cerebras and Groq, the SRAM heavy low latency accelerators he expects to end up as coprocessors.
- AMD, the chip he says people underrate and he is happy to buy.
- Meta and OpenAI, named as large public buyers of AMD silicon.
- Etched, SambaNova and d-Matrix, the new entrants whose hard problem is wafer allocation.
- Google Cloud TPU and AWS Trainium, both of which he says Sail will make work.
- DeepSeek, publishing the KV cache compression work he points at.
- Moonshot AI, maker of Kimi, his example of a one trillion parameter model.
- xAI, used as the FLOP utilization punching bag that he says is still better than the rest of the world.
- Intel, where his mother worked around 2000, and the Western process he says is at worst 2x behind on performance per watt.
- Cisco, the most valuable company in the world during the dot com peak he uses as the historical parallel.
- CoreWeave, the neocloud he cites as evidence that Jensen makes his friends billionaires rather than competing with them.
- GitHub, where he thinks enough good AI generated code now sits to train a Fable class model.
- Apple and Siri, his example of an assistant that abundant inference would make proactive.
- The S&P 500, whose semiconductor weighting went from 2 to 4 percent historically to 19 to 21 percent today.
- ICML and NeurIPS, the conferences NVIDIA managers were scouting in the mid 2010s.
- Factorio, Patrick's description of what Movva is doing in real life.
Chips, systems and technologies
- Blackwell, the 800 square millimeter die with 288 GB of HBM and roughly 500 MB of on die SRAM. NVIDIA is shipping 5 million of them this year, on his figure.
- GB300 NVL72, the Grace Blackwell 300 rack of 72 units, and his example of rack scale as the new unit of programming.
- Hopper, the previous generation and the 2023 to 2024 training era shortage.
- Rubin, the next generation, third point in his flat performance per watt line.
- NVLink, the interconnect he calls mandatory for low latency inference and optional for him.
- HBM, high bandwidth memory, one of his four supply chain bottlenecks.
- Tensor cores, the matrix multiply units he watched win 5 to 10 percent of the die in 2016.
- CUDA, which he calls a red herring on a resume.
- FlashAttention, the kernel modern transformers are built on.
- RoPE, rotary positional embeddings, his example of an architecture change that breaks an existing kernel.
- Attention Is All You Need, the 2017 transformer paper the riff at 32:24 is about.
- Claude Code and Codex, the two coding agents he names when explaining that human attention is the cap on foreground token consumption. He uses
/compactin Claude Code well before a million tokens. - Opus 4.5, which he calls the first agent at all suitable for long horizon tasks, plus Opus 4.6 and 4.7 as what enterprises are actually still running.
- Haiku, his example that the intelligence frontier is jagged: small models find bugs large ones miss.
- Fable and Mythos, the Anthropic models the security community ran against every line of code ever written.
- GPT 5.5 and 5.6, the pricing basis for the five million dollars per trillion tokens figure.
Sponsors read in the episode
The episode itself
- Watch on YouTube, and the show's own home at Colossus.
Related on this site
- Alex Wissner-Gross on Kimi K3, for the other side of the open weight argument, including mixture of experts density and linearized attention on the cost versus capability frontier.
- 3Blue1Brown on how transformers actually work, if the attention and MLP split in this page needs a visual foundation.
- Nate B Jones on what a $20 AI plan actually costs, for the demand side of the same economics.
- TheAIGRID on AI prices, which argues the price of tokens is about to go up, not down, and is worth reading directly against this one.
- Jeff Geerling on Intel matching Apple Silicon, a measured performance per watt result on the vendor Movva says is at worst 2x behind.
Where it stands
Now the part the title does not do. Movva is a serious engineer and this is one of the more technically honest interviews in the genre, but he is also selling the thing he is forecasting, and there are places where the argument leans harder than the evidence.
He never gives a timeframe, anywhere. Eighty four minutes, three to six orders of magnitude, and no date attached to any of it. The 1000x is an ask, a statement of what the world would need for abundant intelligence to make sense, not a projection with a curve behind it. That is a legitimate way to talk, but it is not what a headline number implies, and it is the first thing to notice.
The largest quantified step is a model swap, not an engineering result. Going from at least $5 million per trillion tokens at frontier API pricing to "tens of thousands" for a smaller open weight model is roughly two of the three orders of magnitude, and it is mostly the difference between buying frontier intelligence and buying cheaper intelligence, plus the gross margin the frontier labs charge. He is straightforward that this is what he sells. But a trillion tokens from a small open model and a trillion tokens from GPT 5.6 are not the same good, and the chart in Figure 4 puts them on the same axis because that is how he compares them.
His flat performance per watt claim is real, and it is also a choice of metric. Holding BF16 fixed across Hopper, Blackwell and Rubin suppresses the two places NVIDIA has actually banked generational gains: lower precision datatypes (FP8, then FP4) and rack scale integration, both of which he discusses approvingly elsewhere in the same interview. So "silicon is flat" is true on the metric he picked and misleading as a summary of what a generation buys. On the other hand, the underlying point survives, and it is the useful one: if you want a thousand times, waiting for TSMC will not get you there.
The 95% uptime argument has an unpriced hedge inside it. He says he will tolerate an outage measured in days because he can model the weather and move the workload somewhere else. Moving it somewhere else requires that somewhere else to have idle capacity, which means the aggregate fleet has to be oversized relative to steady state demand, which is a real cost he never quantifies. He also qualifies the control plane claim with "as long as it's not correlated with other data centers," and weather is the most spatially correlated failure mode there is. A fleet of solar sites in one region goes dark together.
Inference spend "monotonically increases" is the load bearing capital claim, and it is the least examined. His argument that tokens are not speculative because you use them immediately is a good one at the level of a single purchase. It says less than it appears to about aggregate spend, which depends on the value delivered per token holding up as agents burn thousands of times more of them. The cybersecurity as proof of work framing he offers is exactly a case where spend scales with adversarial pressure rather than with delivered value, which is a different and less durable demand curve.
Against public cost curves, the direction is right and the mechanism is unsettled. Cost per token at a given capability level has fallen very fast across the industry over the last few years, driven by a mix of smaller models trained better, distillation, lower precision, better serving software, and competition compressing margin. That is largely the list he gives, minus the process node he correctly excludes. What is genuinely novel in his version is the bottom of the stack: nobody else is publicly arguing that stranded one megawatt sites on intermittent renewables, bought at 95% uptime, are a competitive place to put inference. If that works it is a real cost advantage and it is not on anyone's curve yet. If it does not work, it fails on operations rather than on physics, and it will fail quietly.
And the commercial interest is total, which he does not hide. Sail Research is a token factory whose entire value depends on two things being true: that tokens get radically cheaper, and that background agents become the dominant workload. The 90/10 background split, the claim that low latency serving was the wrong bet, the claim that NVLink is optional, and the claim that open weights never go away are all forecasts his company needs. He argues them well and he concedes the things that hurt him, most notably that his software edge is temporary and that he is fully downstream of decisions made in labs he has no input into. Take the mechanisms, which are excellent and checkable. Hold the aggregate number at arm's length.


