youtube.nixfred.com nixfred.com

China Just Shocked Everyone With a 10 Trillion Parameter AI Model

Three frontier model stories inside 48 hours, walked in order. The Financial Times reported that ByteDance is training a model with as many as 10 trillion parameters, which would put it above the industry estimates for Anthropic's Mythos 5 at roughly 8 trillion and Fable 5 around 5 trillion, and roughly 3.5x Moonshot's Kimi K3 at 2.8 trillion. Reuters relayed the report and said directly it could not verify it. Meta shipped the beta of Muse Code, a terminal coding agent powered by Muse Spark 1.2, with persistent background agents, a replayable event log, and contributor pricing 12.5x cheaper on input than standard. OpenAI made GPT 5.6 Luna free with unlimited text for roughly a billion users while a much larger flagship, Astra, sits at release candidate under the checkpoint name Mu4. The page rebuilds every number and caveat, and separates what is announced from what is reported, estimated, and leaked.

Published Aug 8, 2026 15:27 video 26 min read Added Aug 8, 2026 Open on YouTube →

At a glance

Three things happened inside about 48 hours, and AI Revolution walks all three in order. The Financial Times reported that ByteDance, TikTok's parent company, is currently training a model with as many as 10 trillion parameters, which would make it the largest model anyone has ever admitted to building. Meta shipped the beta of Muse Code, a terminal based AI coding agent powered by Muse Spark 1.2, aimed squarely at the same developers Claude Code serves. And OpenAI made GPT 5.6 Luna completely free with unlimited text conversations for roughly a billion people, while confirming almost by accident that a much larger model, codenamed Astra, is already sitting at release candidate.

The through line the narrator draws is scale. ByteDance is not building at the Chinese frontier, it is building at the global one: the industry estimates the FT relayed put Anthropic's Mythos 5 at roughly 8 trillion parameters and Fable 5 around 5 trillion, which means a 10 trillion parameter model would be larger than the biggest system Anthropic is known to have. Meta's entry means three of the largest labs on earth are now fighting over the same terminal window. And OpenAI's Astra, rumored at 7 to 10 trillion parameters, is described as the largest pre training run the company has done since GPT 4.5, pointed at one target.

What follows is the whole report rebuilt in the video's order, with every named model, every number, and every caveat the narrator attached to it.

10 trillion parameters, and a 48 hour weekend

The video opens on the number, not the story. Ten trillion parameters, out of the Financial Times that morning, and the narrator's framing is immediate: if it holds up, it is the biggest model anyone has ever admitted to building. ByteDance is in the middle of training it right now. Meta just dropped a coding agent aimed directly at Claude Code's throat. And OpenAI made its flagship free for a billion people while confirming almost by accident that something much larger is already sitting at release candidate. All of that inside roughly two days.

The order is deliberate. ByteDance goes first, in the narrator's words, "because this is the one with real weight behind it."

ByteDance: what the Financial Times actually reported

The FT reported on Friday, citing people with knowledge of the matter, that ByteDance is training a model with as many as 10 trillion parameters. Reuters picked it straight up.

The narrator is careful about what makes this different from the usual Chinese model release. This is not a marginal step over what China has been putting out. It is a different category entirely.

Where the Chinese frontier was before this

To make the jump legible, the report lays out the domestic leaderboard as it stood before the FT story.

Moonshot AI's Kimi K3, which was the headline number over there until basically five minutes ago, sits at 2.8 trillion parameters. ByteDance's model would be more than three times that. Before K3 shipped, the domestic leaders were Meituan's LongCat 2.0 and DeepSeek's V4 Pro, both at 1.6 trillion total parameters, with a handful of other Chinese labs having crossed the trillion mark.

The narrator's summary of the trajectory: the jump from where that ecosystem was six months ago to where ByteDance is aiming is close to an order of magnitude.

total parameters, as stated or estimated in the report ByteDance, in pre training 10T reported Anthropic Mythos 5 ~8T estimated OpenAI Astra 7 to 10T rumored Anthropic Fable 5 ~5T estimated OpenAI GPT 4.5 ~5T estimated Moonshot Kimi K3 2.8T Meituan LongCat 2.0 1.6T DeepSeek V4 Pro 1.6T 0 2T 4T 6T 8T 10T Chinese lab US lab rumored range, not a stated figure
Figure 1. The whole argument of the segment in one picture. Only the four Chinese figures and GPT 4.5 come from disclosed or reported numbers; the Anthropic bars are industry estimates the FT relayed, because Anthropic and OpenAI do not publish parameter counts, and the Astra bar is a leak. The point is not the precision of any single bar. It is that a Chinese lab is now aiming above the top of the estimated global range rather than below it.

The honest caveat: parameters measure scale, not capability

The narrator stops the momentum here on purpose, "because parameter counts get waved around like they settle arguments."

Parameters are the numerical settings a model learns from data in order to recognize patterns, generate answers, and carry out tasks. They are a rough measure of scale, and scale tends to correlate with capability, but the two are not the same thing and never have been. Plenty of bloated models have lost to smaller, better trained ones. So a 10 trillion parameter number, on its own, is not proof of anything.

The comparison that makes it serious

What makes the story serious is not the number in isolation, it is the comparison the FT actually drew.

Direct benchmarking against leading US models is genuinely difficult, because Anthropic and OpenAI do not disclose parameter counts for Fable, Mythos, or GPT 5.5. Nobody publishes those numbers anymore. What the FT has instead are industry estimates, and those estimates put Anthropic's most advanced system, Mythos 5, at roughly 8 trillion parameters, with Fable 5 around 5 trillion.

Which means ByteDance is not building something in the neighborhood of the Chinese frontier. They are building something in the neighborhood of the global one, potentially larger than the biggest model Anthropic is known to have. The narrator flags this as the part worth sitting with.

From efficiency to frontier scale

Here is the reversal the report is really about. For the last couple of years, the framing on Chinese labs has been efficiency: smaller, cheaper, cleverer, squeezing more out of constrained compute. This is the opposite move. This is a lab deciding to go head on at frontier scale against a company that has never been outscaled by anyone in that market.

Unverified, and still in pre training

Two caveats close the segment, and the narrator gives both of them plainly rather than burying them.

First, verification. Reuters said directly that it could not immediately verify the report, and ByteDance did not respond to a request for comment. Keep that in your pocket.

Second, timing. The model is currently in pre training, which typically runs three to six months before it can even be fine tuned and released. Nothing ships tomorrow.

But the direction of travel is unmistakable. Chinese firms keep accelerating their release cycles to keep pace with the global race, and the balancing act they are running is building more powerful systems without making them prohibitively expensive to actually serve. Ten trillion parameters makes that second half a lot harder, which, as the narrator puts it, tells you how badly they want the first half.

The interlude: the gap is not access anymore

The video breaks for the creator's own pitch, and the framing is worth keeping because it is the argument, not just the ad. Everyone is using AI at this point. Almost nobody is getting paid for it. The gap is not access anymore, it is knowing where to point it.

Mark Cuban's version of this gets played as a clip: learn to build agents, go talk to small businesses, and find the boring stuff nobody has time for. Lead follow up, booking appointments, chasing invoices, answering the same five Customer questions 40 times a day. That work is worth real money to the person drowning in it, and businesses are paying thousands to hand it off. The creator's offer is an 11 page report on seven agents companies are paying $3,000 to $10,000 for right now, who buys each one, and how to spot a business that needs one, with no coding, free.

Meta ships Muse Code

Meanwhile, Meta finally decided to stop watching from the sidelines. The day before the video, Meta shipped the beta of Muse Code, an AI coding agent that runs in your terminal, powered by their latest model, Muse Spark 1.2. Reuters covered the launch.

Mark Zuckerberg posted about it personally, saying it can complete complex software engineering tasks in large code repositories, analyzing projects, planning modifications, writing code, running tools, and verifying results. The narrator's read: that is the whole loop. And it means Meta, OpenAI, and Anthropic are now standing on the exact same ground, competing for the exact same developers.

It is macOS and Linux for now, installed with a single terminal command. The pitch is that you hand it a full requirement rather than a snippet. Fix a bug spanning multiple modules, add a feature, refactor a large project. It reads the codebase first, builds a plan, modifies files, runs tests, and keeps adjusting based on what comes back.

The architecture: a simple loop plus persistent background agents

Architecturally, the core is a fairly simple agent loop. What Meta paired it with is the interesting part: a set of asynchronous background agents that boost the main one.

The distinction the narrator says actually matters is persistence. Those background agents run continuously across the entire session instead of being spun up fresh for each task. Because they persist, they are not recollecting the same context over and over. And they can independently carry out follow up steps and decide for themselves when results are worth feeding back to the main agent.

The payoff is lower latency, and a main agent that leans on you a lot less during complex multi step work.

The runtime: an event log as the only trusted source

Then there is the runtime design, which the narrator calls genuinely the smartest piece of this.

Muse Code keeps a local event log where model calls, tool runs, approval operations, and code modifications are all recorded in sequence. That log is the only trusted source of data, which means the entire run can be replayed accurately and safely restored after a restart. If the program crashes, the agent resumes from the exact point of interruption rather than starting over or drifting somewhere weird.

For long running tasks, that is the difference between a tool you trust overnight and one you have to babysit.

Muse Code runtime Main agent loop reads the codebase builds a plan modifies files runs tools and tests adjusts on what comes back spun up per task Background agents asynchronous, many at once persist across the whole session no recollecting the same context carry out follow up steps alone decide when to report back never spun up fresh boost Local event log, in sequence model calls · tool runs · approval operations · code modifications the only trusted source of data replay exactly, resume at the point of interruption
Figure 2. Why the log design is the piece the report singles out. Because every action is written to one ordered log before anything else trusts it, a crash is not a lost session: the run replays and picks up at the exact point it stopped. That is what turns an agent you supervise into one you can leave running.

The built in skills: /plan, /grill, /goal

On top of the runtime sit three built in skills.

/plan decomposes a task into a plan that requires your approval before anything executes. /grill repeatedly stress tests that plan until the solution is reliable enough. And /goal just keeps pushing toward whatever target you specify until the task is actually complete.

Naturally, someone in the replies immediately asked whether Muse Code is getting open sourced, and Zuckerberg answered with "there will be more content to share on this topic soon," which the narrator reads as the kind of answer that usually means yes eventually, on their timeline.

Muse Spark 1.2, the model underneath

The model underneath, Muse Spark 1.2, is an upgrade on 1.1 focused specifically on code generation, debugging complex problems, repository understanding, and the end to end development process.

Meta significantly increased training compute for programming tasks and expanded the diversity of training environments, while claiming the model held on to its general agent performance.

Benchmarks: second everywhere, first nowhere

The narrator calls the benchmark picture consistent and frankly a little pointed.

benchmarkwhat it measureswho finishes ahead of Muse Spark 1.2
Terminal Bench 2.1 How well agents complete tasks in a terminal environment. Opus 5 Max, and nothing else. Beaten by exactly one model.
Deep Suite 1.1 113 tasks across 91 repositories in five languages: TypeScript, Go, Python, JavaScript, and Rust. Every task ships a manually written functional verification program plus regression tests. Opus 5 Max and GPT 5.6 Terra Max.
Meta's internal coding benchmark 440 tasks derived from real internal pull requests, covering bug fixing, feature development, refactoring, and cleanup. Opus 5 Max again, and nothing else.
Figure 3. The three benchmark results the report cites, and the pattern the narrator pulls out of them: second everywhere, first nowhere. The same name sits above Muse Spark 1.2 on all three, which is also the tell for who Meta was aiming at.

Two things fall out of that table. The first is that the Deep Suite result is no joke: hand written functional verification plus regression tests on every one of 113 tasks is a much harder bar than pattern matching a diff. The second is the narrator's line about the shape of the whole picture. Second everywhere, first nowhere. Strong debut. And it also tells you precisely who Meta was benchmarking against in the mirror.

Where the gains came from

Meta attributes the improvement to three things, and the report walks all three.

First, co training. Muse Spark 1.2 and Muse Code were trained together, so the pair performs properly in tandem. Agent operation trajectories were introduced through rejection sampling training, optimized around target execution, context compression, and subagent links, with the actual Muse Code tool set folded into training to improve compatibility between model and framework.

Second, long horizon capability. Large scale training on full repository generation, big end to end projects, and automated research, where the model sequences work through planning, holds direction via goal conditions, and uses context compression to retain the information it needs to keep going.

Third, self improvement. Meta used Muse Spark 1.1 to generate highly difficult programming environments and instruction following templates, then had the model evaluate how well candidate solutions met the requirements. That builds scalable training data for 1.2 and pushed instruction following past the previous generation.

Pricing as a recruitment drive

Pricing is where the narrator says Meta gets aggressive, and the split is the story.

per million tokensstandardcontributor versiongap
Input $1.25 $0.10 12.5x cheaper
Cached input $0.15 two tenths of a cent 75x cheaper
Output $4.25 $0.20 21x cheaper
Figure 4. The two price tiers as read out in the report, with the multiples computed. A discount of this size is not a margin decision, and the narrator says so directly: "That's not a price point. That's a recruitment drive."

GPT 5.6 Luna goes free and unlimited

Which leaves OpenAI, who spent the week doing the one thing nobody else can afford to.

GPT 5.6 Luna is now completely free, with unlimited text conversations, for roughly a billion users worldwide. The Verge covered it. Luna is the smallest model in the GPT 5.6 family, built for ultra fast response and low cost, sitting under Terra and Sol in the lineup. Starting that day it becomes the default for both the free and Go tiers, directly replacing GPT 5.5.

They also added a Think button to free ChatGPT, so when you hit something difficult you can make the model reason longer before answering. That is the first time free users have been able to actively turn up reasoning intensity themselves.

The fine print matters, though, and the narrator does not skip it. Unlimited applies to text conversation only. File uploads, image generation, voice, and everything else still run on the original quotas, and you will still hit caps.

Sol gets the slider, and stops rambling

Sol got upgraded alongside Luna, and the report argues this change is more meaningful than the headline suggests.

In previous ChatGPT versions, instant mode and thinking mode behaved almost like two separate personalities. Different tones, different formatting, and switching between them felt like swapping models entirely. Now Sol handles everything uniformly, with fully adjustable speed, where the only difference is how long it reasons. The disjointed feeling is gone. The five level reasoning slider that was previously exclusive to ChatGPT work is now in the standard chat interface for everyone.

The most immediately obvious change is that Sol stopped rambling, and OpenAI's own example is the best part of the segment. Someone asks whether biking from the Mission to the beach after work would leave them soaked. The old instant mode produced a wall covering rainfall, wind speed, temperature, sea fog, and beach hazard warnings, plus a note that a cotton t shirt might feel sticky. New Sol opens with the conclusion: you will not get soaked, the real issue is the headwind, west wind at 10 to 20 miles per hour, bring a thin windbreaker.

Then the user follows up to say they will be leaving at half past five. The old version repeated the entire forecast. The new one only updates the conclusion that changed.

The accuracy jump underneath

There is a bigger update underneath the personality fix. Factual accuracy across the whole GPT 5.6 family improved substantially, tested on finance, healthcare, and law, the three domains where factual errors are least acceptable.

The narrator emphasizes that the grading was strict: a single factual error anywhere in a response marks the entire response incorrect. Under that bar, Sol's error rate came in 68% lower than GPT 5.5 Instant, with Luna 62% lower.

Astra, and the Mu4 release candidate

The free tier, the report says, was the appetizer.

Industry insider Leo broke the news that OpenAI is preparing to launch its next generation flagship the following week, codenamed Astra. It is a completely new pre trained model, the largest OpenAI has trained since GPT 4.5, and the latest internal checkpoint, codenamed Mu4, has already reached release candidate status. That is the final version before official release.

And the trail, the narrator says, was there all along.

  • Jul 30An OpenAI preview video gets deleted quickly, but users grab a screenshot showing the word Mu3 flashing on screen.
  • Aug 1OpenAI publishes a mathematics blog post stating their model had solved 10 open mathematical problems unsolved for more than a decade. Buried in it is a line crediting an internal build of Astra, described as their next major model. Announced in the open, and nobody caught it.
  • Aug 5Meta ships the beta of Muse Code, powered by Muse Spark 1.2, and Zuckerberg posts about it personally.
  • this weekMore leaks confirm Mu4 is in internal testing and has reached release candidate. Counting from Mu to Mu2 on upward, the series has quietly reached its fourth form.
  • Aug 6OpenAI makes GPT 5.6 Luna free with unlimited text conversations for roughly a billion users, and upgrades Sol with the five level reasoning slider.
  • Aug 7The Financial Times reports that ByteDance is training a model with as many as 10 trillion parameters. Reuters picks it up and says directly that it could not immediately verify the report.
  • next weekPer the leaks, Astra launches. "Whether it wins by a landslide or by a nose, we find out within a week."
Figure 5. The sequence the report assembles, with the Astra breadcrumbs interleaved into the weekend's three announcements. The dates on the deleted preview video and the mathematics blog post are the two the narrator treats as hard evidence; everything under "this week" and "next week" is leak.

What Astra is for, and the run at Fable 5

Current leaks position Astra for long duration multi agent collaboration, where multiple AI instances work together for hours or even days on one large complex problem.

Two rumored specs are circulating. Some media reports say twice the size of GPT 5.6 Sol, while developer Haider estimates GPT 4.5 at around 5 trillion parameters and puts Astra at 7 to 10 trillion. Massive either way. But with stronger infrastructure, optimizations, and a large amount of new compute coming online this year, OpenAI's service cost for running it may actually land below what it costs Anthropic to run Mythos 5.

Haider's read on why this is the moment is the sharpest strategic argument in the video. OpenAI has the strongest post training capability in the industry. That is how GPT 5.5 hit the performance it did. Its only real disadvantage against Mythos was the pre training foundation. Astra upgrades exactly that shortcoming into the largest pre training base in history.

Top tier fine tuning on the biggest base anyone has built, aimed at one target: taking Fable 5 off the top spot.

Key takeaways

Where it stands: confirmed, reported, and rumored

The report itself flags most of this, but it is worth collecting in one place, because the three stories sit at very different levels of confirmation.

Announced and verifiable. Meta's Muse Code beta and Muse Spark 1.2 are a real product launch with Reuters coverage and a personal post from Zuckerberg. OpenAI making GPT 5.6 Luna free and unlimited for text is an announced product change covered by The Verge. Those two are not in question.

Reported but unverified. The 10 trillion parameter figure is a Financial Times report citing anonymous sources. Reuters explicitly said it could not immediately verify it, and ByteDance declined to comment. Treat it as a credible outlet's sourced claim, not as a confirmed spec, and note that a model in pre training can change shape or be abandoned before anyone sees it.

Estimates, not disclosures. The Mythos 5 and Fable 5 numbers are the softest load bearing figures in the video. No lab publishes parameter counts anymore, so 8 trillion and 5 trillion are third party industry estimates the FT relayed. The comparison between ByteDance and Anthropic is therefore an estimate against a report, not a spec against a spec.

Leaks. Astra, the Mu4 release candidate, the 7 to 10 trillion range, and the "next week" launch all come from an industry insider and a developer, not from OpenAI. The two pieces with independent footing are the deleted preview video screenshot showing Mu3 and the line in OpenAI's own mathematics blog post crediting an internal build of Astra. Those make the existence of the program hard to dispute. They do not confirm the size or the date.

And the caveat the video puts on itself. Parameter count correlates with capability, but does not determine it. A larger model that is worse trained, worse tuned, or too expensive to serve loses to a smaller one, and the serving cost problem is exactly the one the narrator says a 10 trillion parameter model makes harder.

One note on names: this is a fast news read and several proper nouns arrive in a rush. Where the transcript is ambiguous, this page uses the standard spellings for the labs and model lines involved.

Chapters

Notable quotes

10 trillion parameters. That's the number that came out of the Financial Times this morning. And if it holds up, it's the biggest model anyone has ever admitted to building. narrator, 0:00

Parameters are the numerical settings a model learns from data in order to recognize patterns, generate answers, and carry out tasks. They're a rough measure of scale and scale tends to correlate with capability, but the two are not the same thing and never have been. narrator, 1:40

They're building something in the neighborhood of the global one, potentially larger than the biggest model Anthropic is known to have. That's the part worth sitting with. narrator, 2:35

For the last couple of years, the framing on Chinese labs has been efficiency, smaller, cheaper, cleverer, squeezing more out of constrained compute. This is the opposite move. narrator, 2:48

10 trillion parameters makes that second half a lot harder, which tells you how badly they want the first half. narrator, 3:45

You know what I would do coming out of college? I would go to small medium sized businesses having learned how to do agents. Mark Cuban, 4:25

That's the difference between a tool you trust overnight and one you have to babysit. narrator, on the Muse Code event log, 7:00

Second everywhere, first nowhere. Strong debut. And it also tells you precisely who Meta was benchmarking against in the mirror. narrator, 8:55

That's not a price point. That's a recruitment drive. narrator, on Meta's contributor pricing, 10:15

You won't get soaked. The real issue is the headwind. West wind at 10 to 20 mph. Bring a thin windbreaker. OpenAI's example of the new Sol answering, 12:15

And the funniest part is that OpenAI announced it in the open and nobody caught it. narrator, on the Astra mention in the mathematics blog post, 13:40

Top tier fine tuning on the biggest base anyone's built, aimed at one target, taking Fable 5 off the top spot. Whether it wins by a landslide or by a nose, we find out within a week. narrator, 15:00

Resources mentioned

The one idea to walk away with

The three stories look unrelated until you notice they are all bets on the same scarce thing. ByteDance is spending it on raw scale, betting that a 10 trillion parameter base buys a seat at the global frontier even though it makes serving the model far harder. Meta is spending it on distribution, pricing a frontier class coding agent at a tenth of the going rate to buy developers rather than margin. OpenAI is spending it on both ends at once, giving away the small model to a billion people while pouring its largest pre training run since GPT 4.5 into a single flagship. Compute is the currency, and what each lab chooses to buy with it tells you what it thinks the next year is actually a race about. Nobody in this video is competing on being clever with less anymore.

Full transcript
======================================== 10 trillion parameters. That's the number that came out of the Financial Times this morning. And if it holds up, it's the biggest model anyone has ever admitted to building. Bite Dance, Tik Tok's parent company, is in the middle of training it right now. Meta just dropped a coding agent aimed directly at Claude Code's throat. And OpenAI made its flagship free for a billion people while confirming almost by accident that something much larger is already sitting at release candidate. All of that happened inside about 48 hours. Let's go. Starting with Bite Dance, because this is the one with real weight behind it. The FT reported Friday, citing people with knowledge of the matter, that Bite [music] Dance is training a model with as many as 10 trillion parameters. Reuters picked it straight up. And the scale here is not a marginal step over what China has been putting out. It's a different category entirely. Moonshot AI's Kim [music] K3, which was the headline number over there until basically five minutes ago, [music] sits at 2.8 trillion. Bite Dance's model would be more than three times that. Before K3 shipped, the domestic leaders were Maitto's Longat 2.0 and Deepseek's V4 Pro, both at 1.6 trillion total parameters, with a handful of other Chinese labs having crossed the trillion [music] mark. So the jump from where that ecosystem was 6 months ago to where bite dance is aiming is close to an order of magnitude. Now the honest caveat because parameter counts get waved around like they settle arguments. Parameters are the numerical settings a model learns from data in order to recognize patterns, generate answers, and carry out tasks. They're a rough measure of scale and scale tends to correlate with capability, but the two are not the same thing and never have been. Plenty of bloated models have lost to smaller, better trained ones. So, this isn't proof of anything on its own. What makes it serious is the comparison the FT actually drew. Direct benchmarking against leading us models is genuinely difficult because Anthropic and OpenAI don't disclose parameter counts for Fable, Mythos, or GPT 5.5. Nobody publishes those numbers anymore. What the FT has instead are industry estimates. And those estimates put Anthropic's most advanced system, Mythos 5, at roughly 8 trillion parameters with Fable 5 around 5 trillion. Which means Bite Dance is not building something in the neighborhood of the Chinese frontier. They're building something in the neighborhood of the global one, potentially larger than the biggest model Anthropic is known to have. That's the part worth sitting with. For the last couple of years, the framing on Chinese labs has been efficiency, smaller, cheaper, cleverer, squeezing more out of constrained compute. This is the opposite move. This is a lab deciding to go head-on at frontier scale with a company that has never been outscaled by anyone in that market. Reuters did say directly that they [music] couldn't immediately verify the report and Bite Dance didn't respond to a request for comment. So, keep that in your pocket. And there's a timeline element, too. The model is currently in pre-training, [music] which typically runs 3 to 6 months before it can even be fine-tuned and released, so nothing ships tomorrow. But the direction of travel is unmistakable. Chinese firms keep accelerating their release cycles to keep pace with the global race, and the whole balancing act they're running is building more powerful systems without making them prohibitively expensive to actually serve. 10 trillion parameters makes that second half a lot harder, which tells you how badly they want the first half. And that's really the whole point here. Everyone's using AI at this point. Almost nobody's getting paid for it. The gap isn't access anymore. It's knowing where to point it. Mark Cuban's version of this. Learn to build agents. Go talk to small businesses. And find the boring stuff nobody has time for. >> You know what I would do coming out of college? I would I would go to small medium-sized businesses having learned how to do agents, >> lead follow-up, booking appointments, chasing invoices, answering the same five customer questions 40 times a day. That work is worth real money to the person drowning in it. Businesses are paying thousands to hand it off. So, I put together an 11-page report. Seven agents companies are paying $3,000 to $10,000 for right now. Who buys each one? and how to spot a business that needs one. [music] No coding. Links below. It's free. Meanwhile, Meta finally decided to stop watching from the sidelines. Yesterday, they shipped the beta of Muse Code, an AI coding agent that runs in your terminal powered by their latest model, Muse Spark 1.2. Zuckerberg posted about it personally, saying it can complete complex software engineering tasks in large code repositories, analyzing projects, planning modifications, writing code, running tools, and verifying results. That's the whole loop. And it means Meta, OpenAI, and Anthropic are now standing on the exact same ground, competing for the exact same developers. It's Mac OS and Linux for now, installed with a single terminal command. The pitch is that you hand it a full requirement rather than a snippet. Fix a bug spanning multiple modules, add a feature, [music] refactor a large project. It reads the codebase first, builds a plan, modifies files, runs [music] tests, and keeps adjusting based on what comes back. Architecturally, the core [music] is a fairly simple agent loop, but Meta paired it with a set of asynchronous background agents that boost the main one. The distinction that matters, those background agents run continuously across the entire session instead of being spun up fresh for each task. Because they persist, they're not recollecting the same context over and over. And they can independently carry out follow-up steps and decide for themselves when results are worth feeding back to the main agent. The payoff is lower latency and a main agent that leans on you a lot less during complex multi-step work. Then there's the runtime design which is genuinely the smartest piece of this. Musecode keeps a local event log where model calls, tool runs, approval operations, and code modifications are all recorded in sequence. That log is the only trusted source of data, which means the entire run can be replayed accurately and safely restored after a restart. If the program crashes, the agent resumes from the exact point of interruption rather than starting over or drifting somewhere weird. for longunning tasks. That's the difference between a tool you trust overnight and one you have to babysit. On top of that, sit the built-in skills. Slash plan decomposes a task into a plan that requires your approval before anything executes. SLGrill repeatedly stress tests that plan [music] until the solution is reliable enough. And slashgoal just keeps pushing toward whatever target you specify until the task is actually complete. Naturally, someone in the replies immediately [music] asked whether Musecode is getting open- sourced, and Zuck answered with there will be more content to share on this topic soon, which is the kind of answer that usually means yes eventually on their timeline. [music] The model underneath, Muse Spark 1.2 is an upgrade on 1.1 focused specifically on code generation, debugging complex problems, repository understanding, and the end-to-end development process. Meta significantly increased training compute for programming tasks [music] and expanded the diversity of training environments while claiming the model held on to its general agent performance. The benchmark picture is consistent [music] and frankly a little pointed. On terminal bench 2.1, which evaluates how well agents complete tasks in a terminal environment, Muse Spark 1.2 is beaten by exactly one model, Opus 5 Max. On Deep Suite 1.1, it comes in behind Opus 5 Max and [music] GPT 5.6 Terramax. And that benchmark is no joke. Defining 113 tasks across 91 repositories in five languages, [music] TypeScript, Go, Python, JavaScript, and Rust [music] with every task shipping a manually written functional verification program plus regression tests. Then on Meta's internal coding benchmark built from 440 tasks derived from real internal pull requests covering bug fixing, feature development, refactoring, and cleanup. It's once again only [music] Opus 5 Max ahead. Second, everywhere first nowhere strong debut and it also tells you precisely who Meta was benchmarking against in the mirror. They attribute the gains to three things. First is co-raining Muse Spark 1.2 2 and Muse code were trained together so the pair performs properly in tandem with [music] agent operation trajectories introduced through rejection sampling training optimized around target execution context compression and subagent links and the actual Muse code tool set folded into training to improve compatibility between model and framework. Second is long horizon capability, meaning large-scale training on full repository generation, big end-to-end projects, and automated research where the model sequences work through planning, holds direction via goal conditions, and uses context compression to retain the information it needs to keep going. Third is self-improvement. Meta used Muse Spark 1.1 to generate highly difficult programming environments [music] and instruction following templates. Then had the model evaluate how well candidate solutions [music] met the requirements, building scalable training data for 1.2 and pushing instruction following past the previous generation. Pricing is where they get aggressive. Standard runs $125 per million input tokens, 15 for cached input, $4.25 25 cents for output, but the contributor version is 10 cents per million input, 2/10 of a cent cashed, 20 cents output. That's not a price point. That's a recruitment drive, which leaves Open AI, who spent this week doing the one thing nobody else can afford to. GPT 5.6 Luna is now completely free with unlimited text conversations for roughly a billion users worldwide. Luna is the smallest model in the GPT 5.6 six family built for ultra fast response and low cost sitting under Terra and Saul in the lineup. Starting today, it becomes the default for both free and go tiers directly replacing GPT 5.5. They also added a think button to free chat GPT. So when you hit something difficult, you can make the model reason longer before answering. The first time free users have been able to actively turn up reasoning intensity themselves. The fine print matters though. Unlimited applies to text conversation only. File uploads, image generation, voice, and everything else still run on the original quotas and you will [music] still hit caps. Saul got upgraded alongside it. And this change is more meaningful than the headline suggests. In previous chat GPT versions, instant mode and thinking mode behaved almost like two separate personalities. Different tones, different formatting, and switching between them felt like swapping models entirely. Now Soul handles everything uniformly with fully adjustable speed where the only difference is how long it reasons and the disjointed feeling is gone. The fivelevel reasoning slider that was previously exclusive to chat GPT work is now in the standard chat interface for everyone. The most immediately obvious change is that Saul stopped [music] rambling. Open AAI's own example was someone asking whether biking from mission to the beach after work would leave them soaked. The old instant mode produced a wall covering rainfall, wind speed, temperature, sea fog, beach hazard warnings, plus a note that a cotton t-shirt might feel sticky. New Saul opens with the conclusion, "You won't get soaked. The real issue is the headwind. West wind at 10 to 20 mph. Bring a thin windbreaker." Then when the user follows up with, "I'll leave at 5:30," the old version repeated the entire forecast, while the new one only updates the conclusion that changed. There's a bigger update underneath that though. Factual accuracy across the whole GPT 5.6 family improved substantially tested on finance, healthcare, and law. The three domains where factual errors are least acceptable. And the grading was strict. A single factual error anywhere in a response marks the entire response incorrect. Under that bar, Saul's error rate came in 68% lower than GPT 5.5 Instant [music] with Luna 62% lower, but the free tier was the appetizer. Industry insider LEO broke the news that OpenAI is preparing to launch its next generation flagship next week, codenamed Astra. It's a completely new pre-trained model, the largest OpenAI has trained since GPT 4.5. and the latest internal checkpoint codeamed mu4 has already reached release candidate status. That's the final version before official release. The trail was there all along. On July 30th, an OpenAI preview video got deleted quickly, but users grabbed a screenshot showing the word MU3 flashing on screen. This week, more leaks confirmed MU4 was in internal testing. Counting from MW to Mew2 on upward, the series has quietly reached its fourth form. And the funniest part is that OpenAI announced it in the open and nobody caught it. On August 1st, they published a mathematics blog post stating their model had solved 10 open mathematical problems unsolved for more than a decade. And buried in it was a line crediting an internal build of Astra described as their next major model. As for what it is, current leaks position it for longduration multi- aent collaboration where multiple AI instances work together for hours or even days on one large complex problem. Two rumored specs are circulating. Some media reports say twice the size of GPT 5.6 SOL, while developer Haidider estimates GPT 4.5 at around 5 trillion parameters and puts Astra at 7 to 10 trillion. massive either way. But with stronger infrastructure, optimizations, and a large amount of new compute coming online this year, OpenAI's service cost for running it may actually land below what it costs anthropic to run Mythos 5. Hiders read on why this is the moment. OpenAI has the strongest post-training capability in the industry. That's how GPT 5.5 hit the performance it did. And its only real disadvantage against Mythos was the pre-training foundation. Astra upgrades exactly that shortcoming into the largest pre-training base in history. Top tier fine-tuning on the biggest base anyone's built, aimed at one target, taking Fable 5 off the top spot. Whether it wins by a landslide or by a nose, we find out within a week. That's the whole situation. Curious where you land on it. Comments below. Thanks for watching and I'll see you in the next one.