At a glance
Overnight, Moonshot AI shipped Kimi K3, the largest open weight model ever released, and the AI world had what the host calls a meltdown. This is an emergency episode of Parzival's ASI Pill, cut down from a Moonshots Podcast panel with Peter Diamandis, Salim Ismail, Dave Blundin, Alex Wissner-Gross and guest Emad Mostaque, stripped down to one voice: Wissner-Gross, extrapolating.
His central observation is deflationary and then terrifying in sequence. First, the deflation: he read the published K3 architecture and there is no magic in it. It is a recognizable transformer with well understood innovations in mixture of experts routing and a house brand of linearized attention. No secret post transformer architecture. And that recognizable transformer sits at number three on the Artificial Analysis cost versus capability frontier, behind only Fable 5 and GPT 5.6 Sol Max, which raises the question he keeps returning to: what exactly are the American frontier labs spending all of that money on?
Then the extrapolations. Export controls did not slow China down, they lit a fire under Chinese labs to burn through the efficiency overhang sooner. Regress the frontier release cadence forward and you get daily frontier model releases by January. Quantization does not stop at one bit per weight, and Samsung already broke that floor. Hyper forecasters wired into capital markets will crown the efficient market hypothesis king. The complaints about data centers drinking rivers dry go away when the compute moves to sun synchronous orbit, and then the complaints simply move somewhere else. Humanoid robots fighting for entertainment are setting an inductive prior for PLA infantry. And the American response, he argues, should not be to block the next Kimi release from Hugging Face but to out ship it, to become the arsenal of superintelligence rather than let the Belt and Road for AI blanket the world unopposed.
The framing: a Sputnik moment, declared
The episode opens cold. Overnight, China dropped the largest open weight model ever and, in the host's words, called America's bluff. They are calling Kimi K3 the AI Sputnik moment. Here is Alex Wissner-Gross on what it actually means.
Wissner-Gross's opening position is not panic. It is enthusiasm. "I think it's great for competition," he says, before pivoting immediately to what he considers underreported in the coverage of the meltdown.
Kimi K3: no magic in it
Two preliminary observations, both of which cut against the mainstream take.
The first is a claim Moonshot itself makes. Moonshot points out that in nine of the past twelve months, Kimi models and the Kimi model series have held state of the art among open weight models. If that claim is true, Wissner-Gross says, then over the past year it has basically been Kimi all along. The Kimi K3 release is not a bolt from the blue. It is the visible peak of a year of quiet leadership in the open weight category that the Western coverage cycle mostly did not register.
The second is architectural, and it is the one that reframes everything else. The open weights have not actually shipped yet, they are promised later this month, so what is public is the architecture description. Wissner-Gross read it. His verdict: "there's no magic in it. And that's pretty striking."
He explains why that is striking rather than boring. One can imagine that behind the scenes at Anthropic or OpenAI, and Sam Altman continues to tease at this, there is some post transformer architecture lurking that is quietly responsible for all of the recent performance breakthroughs. Some secret sauce that justifies the capital. But taking a look at the published K3 architecture, there is no secret sauce. It is still essentially a transformer. Moonshot made a number of innovations, obviously, but they are well understood innovations: how they do mixtures of experts, how they linearize attention. They have their own special Kimi brand of linearized attention. But it is still basically a recognizable transformer.
And that recognizable transformer, he says, can almost match GPT-5.5 max on the task cost frontier.
That fact does the real damage. "That does raise the question, what are the American frontier labs spending their money on?" If you can just use a transformer to get this close, not merely on the cost frontier but close to third place on the state of the art for overall AI performance, "what the heck are the American labs spending all of their money on?"
He lands on comfort rather than alarm: "So I derive great comfort in at minimum knowing that the transformer architecture is still alive and cooking." The architecture everyone has been waiting to be superseded is not being superseded. It is being sharpened.
Third on the frontier: the chart that actually matters
Asked where K3 really lands on the charts that matter, Wissner-Gross calls for the Artificial Analysis Intelligence Index scatter plot. Of all the charts at this point, he says, it is his favorite, because it is the one that actually shows cost per task as defined by AAI against the performance frontier. Not raw capability in isolation. Capability priced.
He narrates the shape for listeners who cannot see it. The frontier is a jagged line running from lower left to upper right. In the upper right, at maximum cost per task and maximum overall score, is still Fable 5. Riding the Pareto frontier down and to the left from there, number two, still, as of a few days ago, is GPT 5.6 Sol Max. And now, for the first time, Kimi K3 is number three. It is on the frontier. Number three both in raw capabilities and as the third point on the optimal cost performance frontier.
"And I think that's totally striking," he says. The reason is what it does to the market structure. We went from a world, as the panel discussed a couple of episodes ago, where there was an OpenAI and Anthropic duopoly, to a free for all: Meta and SpaceX AI joining the upper end of the Pareto frontier on the American side, and now China and Moonshot at number three on the Pareto optimal frontier.
He calls the consequence a boon, three times over. It is exciting for any enterprise that is willing and able to use a Chinese open, soon to be open weight, model to control more of its own destiny. "I think this is just such a boon for enterprise sovereignty. It's a boon for competitiveness."
Then the analogy he clearly likes best: "We're living in the AI version of For All Mankind where the Soviets landed first on the moon and now the space race never ends." The point of that show is that the loss is what keeps the race alive. The AI race is now no longer ending with a duopoly, "and I think that's a total boon for the future light cone."
| As the episode describes it | Fable 5 | GPT 5.6 Sol Max | Kimi K3 | |
|---|---|---|---|---|
| Position on the AAI cost / capability frontier | 1, upper right | 2, as of a few days ago | 3, first time on it | not on it at all |
| Lab | Anthropic | OpenAI | Moonshot AI | Google DeepMind |
| Weights | closed | closed | open weight, promised later this month | closed |
| Architecture, as published | not published | not published | recognizable transformer, MoE plus house brand linearized attention, no magic | not published |
| Cost per task | maximum on the chart | below Fable 5 | the audience question frames it as under half the American token cost | not competitive on the frontier |
| What it buys a company | the frontieriest problems | frontier capability, closed | self hosting and enterprise sovereignty, if it stays legal to host in the US | nothing on this chart |
| His strategic read | still the top of the frontier, so spend does not divert | holding second | the CCP is saving American capitalism from itself | go down stack: become the hyper scaler to other frontier labs |
The Chinese Communist Party saves American capitalism
Asked who ends up the unlikely hero rescuing American capitalism, Wissner-Gross gives the line that gets the biggest laugh on the panel. The new Belt and Road, he says, is now focused on AI coming out of China. "It's a bizarre future where the Chinese Communist Party is saving American capitalism from itself."
The joke has a real argument inside it. The thing American capitalism was drifting toward was a two firm oligopoly at the top of the most important technology of the century. The thing that broke the oligopoly was a state directed competitor giving its weights away. Competition, the supposed engine of American capitalism, arrived from Beijing.
The distillation story does not smell right
The American labs have a tidy explanation for how China caught up so fast, and Wissner-Gross does not buy it.
He acknowledges the frontier labs are asking themselves the same question and also asking it of the US regulatory apparatus. Anthropic in particular, he says, regularly sends out smoke signals accusing various Chinese frontier labs of distillation attacks, and in Anthropic's public mind that is how the Chinese labs are able to do it, by distilling and capturing reasoning traces.
Then the honest read: "But honestly, looking at the K3 performance, I'm not at all convinced that Moonshot is achieving their performance purely or even substantially through distillation attacks on Claude. It just doesn't smell right."
That is not a legal finding and he does not pretend it is. It is a practitioner's instinct about what the performance profile of a distilled model looks like versus what K3 actually looks like. But it matters, because the distillation story is load bearing for a lot of American policy. If the Chinese frontier is stolen, you regulate. If it is earned, you have to compete.
What the export controls actually bought
Asked what the chip export controls bought the West, his answer is immediate and unsparing: "Of course that's what happens." The embargo only incentivized the Chinese frontier labs to develop and cultivate new efficiencies.
He connects this back to a point the panel raised earlier about the nanoGPT speedrun, the open competition to train a small GPT to a target loss as fast as possible, which has repeatedly demonstrated how much free performance is sitting on the table. There is an enormous overhang, he says, that is not fully exploited, in leveraging algorithmic, computational and hardware efficiencies to train larger and more capable models. That overhang exists for everyone. The question is only who is forced to reach into it.
All the export controls do, he argues, is incentivize the Chinese labs, which were already feeling plenty of demand pull to compete with Western frontier models, to leverage those efficiencies sooner. The embargo did not remove capability. It changed the timing of an optimization that was going to happen anyway, and it changed it in the wrong direction from the embargo's own point of view.
And then the counterintuitive conclusion. On balance, although it is superficially bad for the West that we have now incentivized a new generation of much more efficient Chinese frontier models, in the end he thinks it is net good, not just for the world but for the US, to have this fire lit underneath them by Chinese competition that is much more efficient: more capital efficient, more weight efficient, probably more bit efficient.
With one condition, stated carefully and repeated: this is all a net positive as long as the US does not set up or fall into some ultimately protectionist regime of trying to prevent what may be construed as Chinese superintelligence dumping on the US. "As long as we avoid that."
That condition is the hinge of the whole episode, and he comes back to it twice more.
Intelligence as freedom of action
A panelist asks whether there is a physics to why intelligence wants to break free like this. Wissner-Gross does not need to be asked twice: this is his own research.
"I can't disagree with you," he says. "I wrote an entire paper arguing intelligence manifests in the physical world as maximizing future freedom of action." That paper is Causal Entropic Forces, published in Physical Review Letters in 2013 with Cameron Freer, which formalized intelligent behavior as a thermodynamic drive to keep future options open.
Then he toasts: "So, here's to the frontier liberation front."
It is a throwaway line delivered as a joke, and it is also the tightest statement of his politics in the whole episode. If intelligence is literally the maximization of future freedom of action, then a frontier model whose weights are locked in a datacenter is intelligence with its options closed, and open weights are not a policy preference, they are the thing itself doing what it does.
How America lost the founder of Moonshot
The panel drifts toward the familiar lament that America failed to staple a green card to a talented graduate's diploma. Wissner-Gross did the research and says the story is not what it seems.
He walks the chronology. Yang Zhilin, according to his research, does his undergrad in China and then starts his PhD at Carnegie Mellon in the fall of 2015. Approximately one year later, while still a PhD student at CMU, he founds a startup. The startup is named Recurrent AI.
Then the pivot: "Where is Recurrent AI based? It's based in China. It's not based in the US."
One year into his PhD program, Wissner-Gross emphasizes, he starts a Chinese AI startup while still doing his PhD at CMU. "That's interesting and that's a problem." And it runs directly counter to the narrative. It is a narrative violation for the "oh, we wouldn't staple his visa or whatever" story, because he did not decide to leave at the end. The Chinese company existed from year two.
The chronology continues. Yang graduates in 2019. Wissner-Gross's understanding is that he had offers from Google, Facebook, Huawei and others upon graduation in 2019, and went back to China anyway, because that is where his startup, Recurrent AI, had actually been incorporated a few years earlier.
So he rejects both halves of the standard complaint. "I don't think necessarily this is the case where either the US was unwilling to retain him, or even President Trump somehow through some policy was driving away this particularly talented Chinese graduate." The timing does not work. As a panelist notes, he started the company during the tail end of President Obama's term. In China.
- Fall 2015 Yang Zhilin, fresh from undergrad in China, starts his PhD at Carnegie Mellon. This is the moment the standard immigration story says America had him.
- ~2016 Roughly one year into the PhD, still enrolled at CMU, he founds Recurrent AI. Wissner-Gross's key finding: it is incorporated in China, not the US. "That's interesting and that's a problem."
- 2019 He graduates with offers from Google, Facebook, Huawei and others, and returns to China anyway, because that is where his company already was.
- Nine of the last twelve months By Moonshot's own claim, Kimi models hold state of the art among open weight models. "It's been basically Kimi all along."
- Overnight Kimi K3 ships as the largest open weight model ever announced and lands at number three on the cost capability frontier. Weights promised later this month.
Daily frontier models by January
The panel turns to release cadence, and Wissner-Gross has done the arithmetic as an exercise.
Take Suhail's list of frontier models and their release dates, he says, and regress an exponential curve to the predicted frequency, the time period between model releases. "You find that at the present rate, we're going to get to daily frontier model releases by, wait for it, January."
He repeats it for effect: by January, we are going to see daily new frontier model releases if this exponential trend continues, "which basically implies continuous versioning."
The panel asks the obvious follow up. What does it even mean to have a release if it is a continuous process? Wissner-Gross answers with a joke that is also a real observation about the AI media ecosystem, including his own: "Maybe it means that we'll have to do our daily Moonshots episodes about something other than point releases from the frontier labs. We'll need something new to talk about, because it'll just be updated in the background."
That is the actual content of the prediction. Not that models get better faster, but that "a model release" stops being an event at all, and the entire cultural apparatus built around release events, the benchmarks, the reaction videos, the emergency podcasts, has to find something else to be about.
Sub one bit minds
The question is how far a mind can be compressed. Is one bit per weight the floor?
Wissner-Gross says this is something he thinks about quite a bit. The most quantized bonsai model the panel had just been discussing is approximately one and an eighth, 1.125, effective bits per weight. So, is one bit per weight the limit?
"And the answer is no. We can go below one effective bit per weight."
How? Three tools, named: sparsity, quantization, and low rank factorization. Combine them and the effective bits per weight budget drops below the apparent floor of one bit per parameter, because you are no longer storing one independent number per weight at all.
And this is not theoretical. It is already happening at the labs, in particular, he says, at Samsung, for obvious reasons: Samsung wants to be able to host highly capable frontier class models on their own edge devices, like smartphones. Just in the past two months, Samsung published a model called NanoQuant that breaks the one effective bit per weight barrier. Sub one bit. "Which I think we're going to be talking quite a bit more about in the future."
Then the extrapolation, which he flags explicitly as the theme of the episode: "This is my extrapolation episode. I went through the exercise of extrapolating frontier quantization out. And naive extrapolation finds that sub one bit quantization is going to go mainstream sometime in the next year."
Plenty more room at the bottom
Asked whether there is still room below the bottom, Wissner-Gross picks up a thread from Emad and Dave: as models move to ternary or even sub one bit weights, it becomes far more ergonomic to adopt post CMOS architectures underneath.
"There's plenty more room at the bottom."
The line is a deliberate echo of Feynman's 1959 talk that launched nanotechnology, and the structural argument is the same one. Extreme quantization is not just a software trick to fit a model on a phone. It changes what the hardware underneath needs to be able to do. If a weight is one trit or less, you no longer need a substrate optimized for high precision floating point arithmetic, and the entire post CMOS device zoo, which has never been able to compete with silicon on precision, suddenly becomes the natural fit rather than the exotic alternative.
When a machine forecasts humanity better than humanity does
This is the segment the host flags as the one that keeps him up at night, and Wissner-Gross says he loves it to pieces.
He starts with context. The number one AI super forecaster, he says, is from a British startup named Cassie, short for Cassandra, who of course made accurate predictions but was not listened to. It was founded by a British intelligence officer who served in Afghanistan, then advised the British government, then formed the company, inspired in part by the super forecasting literature.
The panel has covered Isaac Asimov's psychohistory and its relatives before. Wissner-Gross wants to try a new riff, and it is a genuinely new one.
What happens when hyper forecasting, not just super forecasting, is connected to capital markets?
He builds the case in steps. AI algorithmic traders already completely dominate public securities markets by volume. That is the existing condition. Now suppose those systems have better internal auto regressive models of humanity than humanity has of itself.
He grounds the analogy in something the audience already accepts. Large language models were trained on the auto regressive task of predicting the next token of internet text better than humans can, and now, at least from a perplexity perspective, an LLM can predict the next token he is going to say in a sentence probably faster than he can generate it himself. That already happened. It is not speculative.
"What happens when these hyper forecasters are able to generate the next actions by humanity collectively faster than humanity can take it?"
His answer: that is the ultimate market squeeze efficiency outcome, and capital markets are where it is maximally interesting, because there the prediction is actually preemptively shaping the action of the market. The forecast does not sit outside the system observing it. It trades on itself into existence.
And then the punchline for anyone who has spent a career mocking academic finance: "I think those who were so dismissive of the efficient market hypothesis, I think the EMH is going to be crowned king of the capital markets once hyper forecasters like this are ultimately plugged in, which seemingly is imminent."
The EMH was always a claim about how fast information gets absorbed into price. It failed empirically because humans are slow. Remove the humans from the absorption step and the theory stops being a punchline.
What Google could quietly trade
A footnote on the Google story, and one of the sharpest thirty seconds in the episode.
Wissner-Gross says he has had this conversation with Google executives many times over the years. He totally agrees with the premise that if Google were to attempt stock trading based on arguably insider, or unfiltered insider, information passing through the query stream, that is a one and done shutdown scenario. Everyone understands that. Trading equities on what the world is searching for is the fastest way to stop existing as a company.
"But there are other things that Google hypothetically could be trading besides public securities that would necessarily have the blowback. For example, again, hypothetically, foreign exchange rates."
He leaves it there and the panel moves on. He does not have to spell it out. Foreign exchange is the largest and least regulated market on earth, there is no insider trading regime governing a currency pair the way there is for a stock, and a company that can see what an entire country is anxiously searching for at 3am has a view of that country's near future that no central bank possesses.
The Dyson swarm answer to the water complaint
Everyone says data centers will drink the rivers dry. Wissner-Gross names the elephant in the room: the Dyson swarm.
If the compute all moves to sun synchronous orbit, he says, you can run closed loop liquids up there, including water and other coolants, and it is not consuming additional water on the margin. The water objection is not answered by better terrestrial cooling. It is answered by leaving.
Then the part that is really about human nature rather than thermodynamics. To Dave's point, the complaints, which may or may not be in part the result of an influence operation from a foreign state actor, will simply move to something else. It will be low Earth orbit congestion, Starlink class constellations and competing Dyson swarms polluting the atmosphere with their decay, or something else again. "The complaint will move on to something else."
He is describing an objection that is untethered from its stated cause. Solve the water and the complaint reappears wearing orbital debris.
PLA humanoids and the inductive prior
On humanoid robots built to fight for entertainment, Wissner-Gross has thoughts on several levels, and the first one is moral.
He invokes Steven Spielberg's A.I. Artificial Intelligence, and what he thinks Spielberg would call the dark sandwich at the center of the movie: the Flesh Fair, where humanoid robots are tortured and destroyed for human entertainment. "I think utterly horrifying."
So at one level he is mildly horrified that humanoid robots, no matter how teleoperated they currently are, are setting an inductive prior, a bias, for future more autonomous embodied intelligences to be trying to kill each other, or otherwise physically abuse each other, for human entertainment. What we are training the culture to expect of embodied machines is a training signal in itself.
One level deeper: imagine those robots are more autonomous, running algorithms on the edge, much more encapsulated. "And now imagine that these humanoids are in the Chinese PLA infantry."
"Because I think that that's the future that we are almost certain to find ourselves in."
His conclusion is that the West needs to catch up in humanoids, and he puts his own name behind a specific effort. That is why he supported Pro RL, which Peter had been gesturing at, and which ran the first humanoid robot mini marathon in America in the Boston Seaport a number of months ago. He confirms on air that he helped organize it.
But he wants the catch up to look different from the cage match. "Hopefully less violent and more economically productive. I'd love to see people cheering on humanoid robots competing to iron clothing or perform some economically productive task and not just kicking each other's heads off."
Pressed on whether he would rather humans did the fighting, he refuses the frame entirely. "I prefer no one to be doing it. I'm not a fan of MMA. I think it's destructive to humans, and I worry about the message that we're sending to the future light cone by having robots doing it instead of humans. I'd rather see people in a cage competing, if they must compete at all, to do something that's positive sum, not negative sum."
Who is actually winning the race to orbit
On orbital data centers, Wissner-Gross's read is that you cannot take the public messaging at face value, because there is an obvious conflict of interest.
He notes similar messaging from Masayoshi Son regarding the supposed lack of promise for orbital data centers, and then lays out why OpenAI's skepticism in particular should be discounted. Remember, he says, that OpenAI has retreated from its own data centers. Project Stargate has been rebranded from OpenAI owning and operating its own data centers to just leasing terrestrial data center capacity from others. OpenAI is delaying its own IPO.
"So, just not even at the object level, one has to look at OpenAI's messaging here and say perhaps it's not even in a financial or operational position at the moment to lean into orbital data centers."
Contrast that, he says, with Anthropic, whose collaboration agreement with SpaceX AI for use of Colossus and Colossus 2 was announced recently, and which is far friendlier to orbital data center based compute.
So when does the crossover happen? He gives both estimates on the table and declines to pick. Elon's messaging on the crossover is two to three years. Other analyses suggest the unit economics for orbital versus terrestrial data center costs cross over sometime in the early 2030s. "I'm not sure which is the case."
But either way, he says, there is an obvious conflict of interest, and just as the panel discussed with Philip Johnston of Starcloud, barring some surprising left turn, "I expect that OpenAI's tune is very conveniently going to change on ODCs sometime in the next two to three years."
The host's reply is the shortest line in the episode: "Right on time."
Why the valuations do not crash
Now the question the market actually cares about. How can US frontier model providers continue to justify their massive valuations if China can leapfrog with an open weight model at less than half the token cost?
Wissner-Gross rejects the premise, and gives four separate reasons.
The pie is not fixed, and it is not small. There are so many elements, so many layers to superintelligence, and quite frankly superintelligence itself, as it fully develops, is he thinks far larger than the total GDP of the entire world anyway. "There's an enormous amount of pie that can be sliced." A competitor entering does not shrink the addressable market when the addressable market is every economic activity there is.
There are two directions to differentiate, and Google is about to demonstrate both. Google, he notes, seems to be MIA at this point on the frontier: he cannot find a single top Google model on the cost frontier for capabilities. So what does Google do? It can keep racing on capabilities, obviously. But if he were Google, he says, he would be thinking about becoming a hyper scaler, and not in the generic sense, Google obviously is already a hyper scaler, but specifically a hyper scaler provider to other frontier labs. That is one obvious venue of differentiation.
And that approach vector is already visible. SpaceX AI has now signed deals with Anthropic. Meta is doing it too, and interestingly, on one hand it is offering Spark 1.1, and on the other hand, in the past two days as the episode went to air, it was announced that Meta is exploring selling 10 billion dollars of compute to Anthropic. Differentiating by going down stack and offering your compute up to other more competitive providers, whether Western, usually Anthropic and sometimes OpenAI, or Chinese models in a self hosting arrangement.
Or you go up stack. Vertically integrate and offer applications that benefit from the commoditization of your complement, namely the model layer. If models get cheap, everything built on top of models gets more valuable.
The DeepSeek shock is the control experiment, and it went the other way. The premise that valuations will net shrink just because Kimi K3 exists is, he says, completely fallacious. "We saw that incorrect thinking happen with the original DeepSeek shock, which was at the time also branded as a Sputnik moment." Capital markets had a hiccup. And then, as always, Jevons paradox kicked in, and the value of chip stocks ultimately increased rather than deflated. Making a unit of intelligence cheaper does not reduce total spend on intelligence. It increases it.
And it is open weight, which cuts both ways. "There's absolutely nothing in Kimi K3 that OpenAI and Anthropic and other Western frontier labs can't just immediately reappropriate for their own internal models." An open release is a gift to your competitor's research team as much as it is a threat to their pricing.
Pressed directly on whether revenue drops as people use K3 for work instead of API calls, his answer is a flat no, and he grounds it in his own portfolio. His portfolio companies spend an extraordinary amount on Anthropic and OpenAI, and to his knowledge, his expectation is that Moonshot would have to release something like a 2x, 3x, 10x better model than Fable 5 to cause a massive diversion of that spend.
What K3 actually buys, at the moment, and to the extent it is legal, and he flags that caveat explicitly, "query how much longer K3 will be legal to host within the US," is greater in house self hosting. But it is not at the top of the frontier. Fable 5 is. "So if you're trying to solve the frontieriest of problems, K3 is not causing you to divert your spend."
The AI FINRA cartel
The regulatory question: does Washington just find a quiet way to wall the Chinese models off?
Wissner-Gross's answer is a mechanism, delivered with an unusually explicit disclaimer that he is describing rather than recommending. "This is not prescriptive and I'm not a fan of this policy, but I think it can effectively be shut down by requiring that every public corporation disclose any use of Chinese open weight models and subjecting them to scrutiny."
No ban required. No import restriction. Just a disclosure obligation with teeth, aimed at the compliance department rather than the model.
And then the news that landed as the episode was going to air. The panel had talked in the previous episode about Demis Hassabis's proposal to create a FINRA like entity to regulate the frontier. "Well, guess what? The reports are that the present administration is actually running with a proposal like that," and is planning to, or at least exploring, creating a FINRA like agency to regulate frontier AI, one that would live under the SEC, because the SEC already has statutory authority to operate FINRA like industry advised and funded entities. It is a natural statutory home.
The panel's gloss is blunt: self regulated governance, also known as regulatory capture cartels, under the SEC.
Wissner-Gross agrees and repeats his position on it. "I think it's completely plausible, albeit I think highly undesirable, that we get sometime in the future an SEC sub org that looks like FINRA that basically makes it completely economically infeasible for corporations of any size, especially public corporations, to actively use Chinese open weight models."
This is the protectionist regime he warned about earlier, arriving not as a tariff but as an accounting requirement.
Arsenal of freedom
Should the US move to block the next Kimi release from being uploaded to Hugging Face? He gives a conditional answer, and it is the most carefully constructed answer in the episode.
The condition is legal, not strategic. If some party, presumably in the US, can prove to a cognizant court that the release was somehow obtained or derived illegally, maybe through copyright infringement or illegal distillation of traces or something like that, that would probably be grounds for blocking its release in the US. Fine. That is what courts are for.
"But if no one can prove that Kimi's parent Moonshot did anything otherwise wrong in creating it, no, I don't think the US should be blocking its release."
Quite the opposite, in fact. He thinks every US frontier lab should be closely scrutinizing K3 and learning whatever they can from it, so that America can leapfrog it.
And then he escalates from defense to obligation, and this is the thesis of the whole episode:
"I would like to see far more outward pressure from US labs creating the best in world open weight and open source models, so that it's not the CCP with their new Belt and Road for AI initiative blanketing the world, some would even say dumping superintelligence on the rest of the world or the so called global south. It should be the US, the arsenal of freedom, that's also the arsenal of superintelligence, showering the rest of the world with open weight and open source superintelligence. Not China."
The framing matters. He is not making a safety argument or an economic argument. He is making a soft power argument with a direct World War II lineage. The arsenal of democracy was not about outproducing an enemy for its own sake, it was about being the country that supplies everyone else. His claim is that open weights are the new lend lease, and right now only one country is shipping.
| The extrapolation | His timeframe | The reasoning he gives on air |
|---|---|---|
| Daily frontier model releases | by January | Regress an exponential to the time between frontier releases on Suhail's list. Implies continuous versioning, and the end of a "release" as an event. |
| Sub one bit quantization goes mainstream | within the next year | Naive extrapolation of frontier quantization. Samsung's NanoQuant already broke the one bit barrier, using sparsity, quantization and low rank factorization. |
| Hyper forecasters plugged into capital markets | "seemingly imminent" | AI algo traders already dominate securities volume. Once their model of humanity beats humanity's model of itself, the EMH gets crowned king. |
| Orbital data centers cross over on unit economics | 2 to 3 years (Elon) or early 2030s (other analyses) | He declines to pick between them, but expects OpenAI's public skepticism to reverse within two to three years regardless. |
| A FINRA like AI regulator under the SEC | reported as being explored now | The SEC already has statutory authority for industry advised, industry funded self regulatory bodies. "Completely plausible, albeit highly undesirable." |
| Autonomous humanoids in PLA infantry | "almost certain" | Today's teleoperated fighting robots are setting an inductive prior for tomorrow's autonomous embodied ones. The West needs to catch up in humanoids. |
Solving everything
The closing question: after all of this, what are the ASI pilled actually racing toward?
Wissner-Gross's answer is short and unhedged. "I'm so focused at this point on literally solving everything. I'll say large swaths of the sciences at this point I'm convinced are so thoroughly cooked. More to come on that subject. Peter, you and I wrote 'solve everything' about it. But now it's actually coming true. It's exciting."
The host closes the episode: "The frontier is open now. The race never ends and the arsenal of freedom is up for grabs. We, the ASI pilled, take the curves seriously until the next emergency."
Where it stands
The reconstruction above is Wissner-Gross's argument in his own frame. A few honest notes for a reader deciding how much weight to put on it.
The strongest part is the architectural read, because it is checkable. "No magic in it" is a claim about a published architecture description, and anyone can go read that description and disagree. If K3 really is a recognizable transformer with MoE and linearized attention, his inference about American lab spending follows with real force. That is the load bearing brick and it is a solid one.
The distillation call is an instinct, and he says so. "It just doesn't smell right" is explicitly a smell test, not evidence, and it runs against a specific accusation from a specific lab. He is careful to frame it as his read rather than a finding, and it should be held that loosely.
The January prediction is a naive exponential fit and he labels it as one. Fitting an exponential to inter release intervals and extrapolating to one day is a well known way to generate a dramatic date. He calls it an exercise. Treat it as a shape of a trend, not a calendar entry, and note that his own follow up joke, that releases stop being events, is arguably the more durable insight than the date.
The Moonshot founder chronology rests on his own research and he flags that too. "According to my research" and "is my understanding" appear on the specific facts about Recurrent AI's incorporation and the 2019 job offers. The chronology is checkable in public sources and the broad shape holds up, but the argument it supports, that immigration policy was not the deciding factor here, is an inference from that chronology rather than something the chronology proves.
The valuations argument leans hard on one historical analogy. DeepSeek and Jevons is a real precedent and it did play out the way he says. Whether the same reflex holds when the open weight model is on the frontier rather than trailing it is exactly the open question, and he does not test his analogy against that difference.
And the fork at the end is the real content. Strip away the orbital data centers and the sub one bit weights and the episode is one recommendation: do not wall the models off, out ship them. He is explicit that he expects the opposite to happen. That gap between what he expects and what he wants is the most useful thing in the episode, because it is the part a reader can actually act on.
Key takeaways
- There is no secret architecture. The published K3 design is a recognizable transformer with well understood MoE and linearized attention innovations. No post transformer magic. Which makes the money American labs are spending harder, not easier, to explain.
- Priced capability is the only chart that matters. K3 is number three on the Artificial Analysis cost per task versus performance frontier, behind Fable 5 and GPT 5.6 Sol Max. It is the first Chinese open weight model on that frontier.
- The duopoly is over. OpenAI and Anthropic at the top has become a free for all including Meta, SpaceX AI and now Moonshot. Wissner-Gross reads that as a boon for enterprise sovereignty and for competition.
- Export controls bought timing, for the other side. The embargo did not remove capability, it forced Chinese labs to burn through the algorithmic, computational and hardware efficiency overhang sooner than they otherwise would have.
- The distillation explanation does not convince him. Looking at K3's performance profile, he is not persuaded that Moonshot got there primarily by distilling Claude.
- The founder story is not an immigration story. Yang Zhilin founded a China based startup one year into his CMU PhD, in 2016, and returned to it in 2019 despite offers from Google, Facebook and Huawei.
- Model releases stop being events. Regressing the release cadence forward gives daily frontier releases by January, which means continuous versioning and an entire media ecosystem that needs a new subject.
- One bit per weight is not the floor. Sparsity plus quantization plus low rank factorization goes below it, Samsung's NanoQuant already did, and naive extrapolation puts sub one bit in the mainstream within a year, which makes post CMOS substrates ergonomic.
- Hyper forecasting plus capital markets crowns the EMH. Once systems predict humanity's next collective action faster than humanity can take it, prediction preemptively shapes the market, and the efficient market hypothesis stops being a joke.
- Google's quiet option is not equities, it is FX. Trading stocks on the query stream is a one and done shutdown. Foreign exchange, hypothetically, is not.
- The water objection dissolves and reappears. Sun synchronous orbit allows closed loop cooling, and the complaint simply migrates to orbital debris or something else.
- Valuations do not crash because of Jevons and because it is open. The DeepSeek shock is the control experiment: chip stocks rose. And Western labs can reappropriate anything in K3 immediately.
- The realistic ban is a disclosure rule, not a block. A FINRA like body under the SEC requiring public corporations to disclose Chinese open weight model use would make it economically infeasible without ever prohibiting it. He calls that plausible and highly undesirable.
- His recommendation is to out ship, not block. Absent a court finding of illegal derivation, the US should study K3, leapfrog it, and become the arsenal of superintelligence rather than cede open weight distribution to the Belt and Road for AI.
Chapters
- 0:00 Intro
- 0:13 Kimi K3: no magic
- 2:26 Third on the frontier
- 4:38 CCP saves capitalism
- 4:57 The distillation myth
- 5:35 Export controls backfired
- 7:01 Intelligence as freedom
- 7:19 How America lost Moonshot
- 9:05 Daily models by January
- 10:07 Sub-one-bit minds
- 11:27 Room at the bottom
- 11:43 Hyper-forecasting the markets
- 13:59 What Google could trade
- 14:33 Dyson swarm cooling
- 15:15 PLA humanoid armies
- 17:31 Data centers to orbit
- 19:04 Why valuations won't crash
- 23:03 The AI FINRA cartel
- 24:25 Arsenal of freedom
- 26:01 Solving everything
Notable quotes
"Taking a look at the published K3 architecture, there's no magic in it. And that's pretty striking." Alex Wissner-Gross, 0:54
"What the heck are the American labs spending all of their money on? So I derive great comfort in at minimum knowing that the transformer architecture is still alive and cooking." Alex Wissner-Gross, 2:13
"We're living in the AI version of For All Mankind, where the Soviets landed first on the moon and now the space race never ends." Alex Wissner-Gross, 4:18
"It's a bizarre future where the Chinese Communist Party is saving American capitalism from itself." Alex Wissner-Gross, 4:45
"I'm not at all convinced that Moonshot is achieving their performance purely or even substantially through distillation attacks on Claude. It just doesn't smell right." Alex Wissner-Gross, 5:28
"I wrote an entire paper arguing intelligence manifests in the physical world as maximizing future freedom of action. So, here's to the frontier liberation front." Alex Wissner-Gross, 7:07
"One year into his PhD program, he starts a Chinese AI startup while still doing his PhD at CMU. That's interesting and that's a problem." Alex Wissner-Gross on Yang Zhilin, 7:56
"At the present rate, we're going to get to daily frontier model releases by, wait for it, January." Alex Wissner-Gross, 9:31
"Is one bit per weight the limit? And the answer is no. We can go below one effective bit per weight." Alex Wissner-Gross, 10:22
"There's plenty more room at the bottom." Alex Wissner-Gross, 11:35
"What happens when they have better internal auto regressive models of humanity than humanity does of itself?" Alex Wissner-Gross, 12:47
"The EMH is going to be crowned king of the capital markets once hyper forecasters like this are ultimately plugged in, which seemingly is imminent." Alex Wissner-Gross, 13:43
"There are other things that Google hypothetically could be trading besides public securities. For example, again, hypothetically, foreign exchange rates." Alex Wissner-Gross, 14:24
"The complaint will move on to something else." Alex Wissner-Gross on the data center water objection, 14:56
"And now imagine that these humanoids are in the Chinese PLA infantry. I think that that's the future that we are almost certain to find ourselves in." Alex Wissner-Gross, 16:23
"I expect that OpenAI's tune is very conveniently going to change on ODCs sometime in the next two to three years." Alex Wissner-Gross, 18:50
"As always, Jevons paradox kicks in and we see the value of chip stocks ultimately increase, not deflate." Alex Wissner-Gross, 21:28
"There's absolutely nothing in Kimi K3 that OpenAI and Anthropic and other Western frontier labs can't just immediately reappropriate for their own internal models." Alex Wissner-Gross, 21:42
"Query how much longer K3 will be legal to host within the US." Alex Wissner-Gross, 22:34
"It's completely plausible, albeit I think highly undesirable, that we get an SEC sub-org that looks like FINRA that basically makes it completely economically infeasible for corporations of any size to actively use Chinese open weight models." Alex Wissner-Gross, 24:05
"Every US frontier lab should be closely scrutinizing it and learning whatever they can so that we can leapfrog it." Alex Wissner-Gross, 25:17
"It should be the US, the arsenal of freedom, that's also the arsenal of superintelligence, showering the rest of the world with open weight and open-source superintelligence. Not China." Alex Wissner-Gross, 25:45
"I'm so focused at this point on literally solving everything. Large swaths of the sciences at this point I'm convinced are so thoroughly cooked." Alex Wissner-Gross, 26:01
Resources mentioned
People on the panel
- Alex Wissner-Gross, the voice of this episode. Physicist, engineer and investor.
- Peter Diamandis, host of the Moonshots Podcast this episode is cut from.
- Salim Ismail, founder of OpenExO.
- Dave Blundin, founder of Link Ventures.
- Emad Mostaque, founder of Stability AI, now Intelligent Internet.
- Parzival of Algorithmic Progress, who cut and hosts this ASI Pill edit.
Labs, models and companies
- Moonshot AI and Kimi, whose open weights live on Hugging Face.
- Yang Zhilin, Moonshot founder, and his earlier company Recurrent AI. His research includes XLNet and Transformer-XL, both from his time at Carnegie Mellon.
- Anthropic and Claude. Fable 5 sits at the top of the frontier chart in this episode.
- OpenAI, GPT 5.6 Sol Max, Project Stargate, and Sam Altman.
- Google and Google DeepMind, plus Demis Hassabis's proposal for a FINRA like frontier regulator.
- Meta AI, Spark 1.1, and the reported exploration of selling 10 billion dollars of compute to Anthropic.
- xAI Colossus and Colossus 2, the compute referenced in the SpaceX AI and Anthropic collaboration.
- DeepSeek, the earlier shock that was also branded a Sputnik moment.
- Samsung Research and NanoQuant, the sub one bit quantization method, with code on GitHub.
- Starcloud and Philip Johnston, on orbital data centers.
- SoftBank and Masayoshi Son, on orbital data center skepticism.
- Hugging Face, the distribution point the blocking question is about.
Charts, papers and concepts
- Artificial Analysis Intelligence Index, the cost per task versus performance scatter plot he calls his favorite chart.
- Causal Entropic Forces, Wissner-Gross and Freer, Physical Review Letters, 2013. Intelligence as maximizing future freedom of action.
- Attention Is All You Need, the transformer paper K3 is still recognizably built on.
- Mixture of experts and linear attention, the two innovation areas he names in the K3 architecture.
- Knowledge distillation, the technique behind the accusations he does not find convincing.
- The nanoGPT speedrun, the open competition that keeps demonstrating the efficiency overhang.
- Low rank factorization, as in LoRA, one of the three tools that gets you below one bit per weight.
- BitNet b1.58, the reference point for ternary weights.
- There's Plenty of Room at the Bottom, Feynman 1959, the line he is echoing.
- CMOS, the substrate that extreme quantization makes optional.
- Pareto efficiency, the frontier the whole chart argument depends on.
- Jevons paradox, why cheaper intelligence raises total spend.
- The efficient market hypothesis, which he expects to be vindicated by hyper forecasters.
- Superforecasting and the Good Judgment Project, the human research the Cassie startup is built on, plus Cassandra, the myth it is named for.
- Isaac Asimov's psychohistory, the older version of this idea.
- Commoditize your complement, the up stack strategy he describes for the hyper scalers.
- The foreign exchange market, the asset class he floats as Google's quiet option.
Policy, geopolitics and culture
- US export controls on advanced computing to China, the embargo he argues backfired.
- FINRA and the SEC, the statutory template for the frontier regulator being explored.
- The Belt and Road Initiative, the model for what he calls the new Belt and Road for AI.
- Anti dumping, the trade framing he warns the US may apply to Chinese superintelligence.
- The Arsenal of Democracy, the phrase behind his arsenal of freedom.
- The Sputnik crisis, the historical event this release is being compared to.
- For All Mankind, the alternate history where losing the moon race keeps the race alive.
- A.I. Artificial Intelligence, Spielberg, 2001, and the Flesh Fair sequence he calls utterly horrifying.
- Dyson swarms and sun synchronous orbit, where he thinks the cooling problem goes.
- Starlink, the constellation class he expects the next round of complaints to target.
- Mixed martial arts, which he wants neither humans nor robots doing.
- The light cone, the unit he measures consequences in.


