At a glance
Tina Huang lays out every piece of computing hardware she owns, sorts it into classes by how much memory it has, and tries to run an AI model on every rung. The bottom rung is an Arduino Uno R4 with 32 kilobytes of RAM, which fits no model at all; the top rung is a rented 8x H100 node with 640 gigabytes, running 700 billion parameter models and generating video faster than it plays back. In between she runs a three model voice assistant on a Raspberry Pi 5, discovers that her iPhone 16 has exactly the same 8 gigabytes as the Pi and yet behaves nothing like it, and uses a restaurant kitchen to explain why. The lesson underneath the whole climb is that memory capacity decides what fits and memory bandwidth decides whether it is usable, and that almost every machine you can buy forces you to pick one. She also hands over a formula for the largest model your own laptop can hold, which turns out to reduce to a single number you can do in your head. A portion of the video is sponsored by Crusoe; the read sits at 12:11 and is covered in full below.
The format: if it fits, it sits
The premise is stated in the first fifteen seconds and never complicated. She has a pile of hardware in front of her, she is going to run AI locally on all of it, and the hope is that you come away with ideas for things to build rather than a shopping list. The organizing principle is memory class, not price and not year of release, and that choice is what makes the video work: once you sort hardware by how much memory it has, the question "what can this run" stops being mysterious.
The first class is the one most hardware videos skip entirely. She calls it chips, and she means it literally.
Chips, as they are called, are literally just chips, just computer chips. They have no operating system, no file system, and no separate memory stick, either. Literally everything is just built into these boards. Tina Huang, 0:16
That has a consequence she turns into the running gag of the whole video. On a chip there is no install step, no package manager, no model directory. You flash the thing and find out.
When we try to run an AI model on these, we're not even installing it. We just literally stick it on here and if it fits, it sits and we can run it, or not. Tina Huang, 0:25
"If it fits, it sits" is the test for every rung above the chips too, and the chart below is that test applied to the full set of hardware in the video.
Chips, class one: 32 kilobytes and 8 megabytes
The Arduino that fits nothing
The smallest board on the table is an Arduino Uno R4 with 32 kilobytes of RAM. She promises the rest of the specifications on screen, then delivers the punchline with a drumroll: it is big enough to fit no models. Nothing. She bought it as part of a set because she is trying to learn hardware, and it turned out to be the one board in the video that cannot hold a language model of any size.
That does not make it useless for local AI, and she flags the reason immediately: the Uno R4 has Wi-Fi, so it can talk to her bigger machines and use the models running on them. She parks the idea at 0:55 and pays it off seventeen minutes later, which is one of the better structural jokes in the video.
The ESP32-S3, the cheapest board in the video that runs a model
The next board up is an ESP32-S3 with 8 megabytes of RAM. She paid around 20 US dollars for hers because it came with a screen; without the screen, she says, you are looking at two to five dollars. That is the price of a sandwich for a board that runs a transformer.
Two models fit, and both are from the TinyStories family, which is exactly what the name suggests: models trained to write tiny stories. She runs what the captions render as the "1 megabyte" and "3 megabyte" models, and then corrects herself out loud to the real names.
So, yes, I can run the Tiny Stories 260K model and the Tiny Stories 3M model very comfortably. Tina Huang, 1:40
One correction worth making explicit, because the captions garble it every time: those numbers are parameter counts, not file sizes. TinyStories-260K is a 260 thousand parameter model, the smallest checkpoint from Andrej Karpathy's llama2.c project, and TinyStories-3M is a 3 million parameter model from Ronen Eldan's original TinyStories release. They are not 1 megabyte and 3 megabyte downloads. Getting this right matters if you go looking for them, and it also sets up the arithmetic later: a 260 thousand parameter model is tiny by any modern standard and still too large for the Arduino sitting next to it.
Both run comfortably on the ESP32-S3, and she is clear about what "running" buys you at this size. These models do the very basics of predicting the next word in order to form a coherent story, which she points out is genuinely impressive given the scale. What they cannot do is hold a conversation like a chat model, and they cannot touch any multimodal work: no audio, no music, no video, no images.
What you actually build with a 3 million parameter model
This is where the video earns its "ideas for cool things to build" framing instead of just benchmarking. Getting the model running, she says, is just the start. You make it cooler by bolting modules onto the board.
Her two examples: a little game called scrambled versus not scrambled, which she suggests might be useful if you are learning English grammar, and LED lights that represent the mood of the story being produced. A 3 million parameter model producing a story, and a strip of lights reading the story's emotional temperature in real time, on a five dollar chip. She floats a whole separate video on building with these little devices, decides there is no time for it here, and puts up a summary slide of the ESP32 models plus project ideas with the instruction to take a screenshot. There is also a quick start guide in the free guide linked from the video description.
Mini computers, class two: two devices with 8 gigabytes that behave nothing alike
The second class is where the video's central puzzle gets set up. She puts two things on the table: a Raspberry Pi 5 and her personal iPhone 16.
What separates this class from the chips is not just more memory, it is being a real computer. These run an operating system, Linux on the Pi and iOS on the phone, which means they manage memory, have a file system, and run multiple programs at once. The practical upshot for local AI is that you can keep several models around and use them together, as long as they fit.
Then the setup lands:
What is really interesting is that both of these actually have 8 gigs of RAM. Yet, they are not able to run the same kinds of models. Tina Huang, 3:15
She promises the hardware lesson that explains it and makes you wait for it, which is the right call, because the demos land harder before the theory than after.
The Raspberry Pi 5: three models chained into a conversation
The Pi 5 demo is the best few minutes in the video. She asks it a question out loud and it answers out loud, with the whole pipeline running on the board.
Tina: What color is an octopus?
The Pi: Octopuses come in a variety of colors, but they are usually camouflage, blending with their surroundings. They can be blue, green, red, or even white, depending on the species. the Raspberry Pi 5 voice assistant, 3:34
The chain is three models deep, and she names every link:
- Whisper, a speech to text model, listens and transcribes.
- Qwen3 1.7B, the large language model, interprets the question and writes the response. (The captions render this as "Quiet and Three 1.7B," which is the speech to text layer of this very page's source material garbling the name of a speech to text demo's language model. It is Qwen3, Alibaba's open weights family, at 1.7 billion parameters.)
- Piper, a tiny text to speech model, speaks the answer through the speakers.
Her own framing: full conversations with a Pi, by chaining together three different AI models, all on one tiny device.
Then she adds a fourth capability without adding a bigger machine. A webcam plugged into the Pi grabs a frame, a tiny visual language model called Moondream describes what is in it, and Piper reads the description out. The output is worth quoting in full, because the level of detail from a vision model running on an 8 gigabyte single board computer is the actual demonstration:
The image features a woman sitting in a chair looking at her cell phone. She is wearing a black shirt and is surrounded by a variety of books and other items. There are at least 13 books scattered around the room, some of which are placed on a bookshelf. A backpack can be seen near the woman. A potted plant is located close to her. The scene suggests that the woman might be engaged in reading or studying, as she is surrounded by numerous books and other items. Moondream, running on the Raspberry Pi 5, 4:34
What the Pi 5 can and cannot do, in her words
She gives the Pi a precise capability report rather than a thumbs up, and the precision is the value:
- Micro to small large language models: very easily.
- Tiny tier visual language models: yes.
- Tiny text to speech models: yes.
- Small but decent speech to text models: yes.
- Mid size large language models: it struggles, but "if you really want to, you kind of can do it."
- Small code tier models: it can "kind of barely" run them.
- Image generation, music generation, video generation: really cannot.
That last line comes with the most important caveat in the section, and it is the first crack in the "memory is destiny" idea:
It really cannot do image generation, music generation, or video generation. Not actually because a small image generation model wouldn't fit on here. It does. The problem is that it doesn't have any GPUs, so it's just like painfully slow to do that. Tina Huang, 5:20
The model fits. It sits. It is still useless, because fitting was never the only question. She defers the GPU explanation and moves on.
What the Pi gives back for that weakness is buildability. You can attach an enormous range of accessories and sensors, which is the whole reason the board exists, and she notes that she believes the Hugging Face robots run off a Raspberry Pi as well. That is correct: the wireless version of Reachy Mini, Hugging Face's open source desktop robot, has a Raspberry Pi 5 inside it doing exactly this job. Summary slide, project ideas, and another pointer to the free quick start guide.
The iPhone 16: same 8 gigabytes, a different machine entirely
The phone has the same 8 gigabytes as the Pi and is, in her words, so different. The first reason has nothing to do with silicon:
Because it is running on iOS, Apple does limit each app to only be able to use around 4 or 5 gigs of RAM. So, even though it technically has 8 gigs of RAM, you actually aren't able to run the mid size level of models from 7 to 14 B that you can't do using the Raspberry Pi, because like Apple just doesn't let you do it, unfortunately. Tina Huang, 6:10
So the phone's effective ceiling for a single model is roughly half its physical memory, by policy rather than by physics. The second Apple constraint is distribution: you cannot put whatever you like on the device, you go through the official system, which means every model arrives inside an app that downloads, installs and runs it. That shapes the whole rest of the section, because each capability she demonstrates comes with a named app.
- Language models: Gemma 4 via the Google AI Edge Gallery, which you install and then use to download models locally onto the phone. (Gemma 4 is the current on device family from Google DeepMind, and the Edge Gallery is Google's own open source app for running it with no account and no network.)
- Visual language models: the small class, 1 to 4 billion parameters, very comfortably, since one of them is only 1.5 gigabytes.
- Image generation: an app called Draw Things, running any small image generation model around the 1 billion parameter range with no problem. Her example is Stable Diffusion 1.5 at 2 gigabytes, which runs very comfortably.
- Speech to text: medium and large models, including Whisper Large, very comfortably.
- Text to speech: all of them.
- Object detection: an app called Ultralytics, running YOLO11n.
The detection model gets the longest aside in the section, and she is right that it is underrated. YOLO11n does object recognition to understand a scene, and she points out it is the underlying technology behind everything from facial recognition on your devices to self driving cars, because you have to be able to recognize objects as you drive. Her verdict on the category:
We don't really talk about this class of models very much, these detection models, but they're actually very, very useful. Tina Huang, 7:30
Note what just happened across these two devices. The Pi, with a full Linux box's freedom, runs a bigger language model than the phone and cannot generate a single image. The phone, locked down to roughly 4 to 5 gigabytes per app, generates images in seconds. Same memory, opposite strengths. That is the puzzle the next section solves.
The hardware lesson: a device is a restaurant kitchen
At 7:50 she stops demoing and teaches, and the analogy she picks is good enough that it is worth rebuilding piece by piece rather than summarizing. Think about any device as a restaurant kitchen.
The spaces. Downstairs there is a walk in freezer and a storage space, where all the ingredients and all the things are kept. Upstairs there is a prep counter, where you load up only the ingredients you need right now for the dish you want to make.
The staff. There is one head chef, very advanced, able to make any sort of dish, but just the one of them. There is a brigade of line cooks, less experienced and less versatile, but there are a lot of them and they can help. And there are specialized machines bolted to the walls: an onion machine that only chops onions, a noodle slicing machine that automatically slices noodles. Her description of that third category is the part that makes the mapping click later:
Basically these machines can only do one thing. It does that one thing really well and really fast and it uses almost no electricity. Tina Huang, 8:35
The connection. Between the prep counter and the chefs, the cooks and the machines, there is a conveyor belt passageway. That is how ingredients get from the counter to whoever is going to process them.
First, make a salad
She runs the kitchen once on a real dish before mapping anything, which is what makes the analogy stick. To prepare a salad: go to deep storage in the freezer and pull the salad things out, carrots, lettuce, olives, chicken, other salad things, and put them on the prep counter. Put them on the conveyor belt. The head chef starts giving out orders, tells the line cooks to chop all the vegetables, starts up the onion machine to chop the onions, and handles only the tricky part himself, mixing the sauce. A few minutes later everything comes together and you have a salad.
Now make a model
Same kitchen, different dish. This time what you are cooking is a model.
The mapping, in her order:
- Walk down to the freezer, except we call it storage, and grab the model.
- Put it on the prep counter, which we call the RAM.
- Pass it along the conveyor belt, which we call the memory bus.
- Hand it to the head chef, the CPU, who also directs the line cooks to start cooking parts of the model.
- The line cooks are the GPU.
- The CPU also starts up the little specialized bolted machines, the onion machine, which we call the NPU.
And after all that cooking, the model produces something. You ask "Hi, how are you?" and it says "Hello, I am good. How are you?"
The point she drives home next is the one most people skip: this is not a one time setup cost.
Every time we want to be working with the model, we need to be going through this entire process of cooking. Tina Huang, 9:55
Why the Pi and the iPhone diverge
With the kitchen built, she goes back and collects the debt from 3:15. Both devices have 8 gigabytes, so both have the same prep counter. What they do not have is the same conveyor belt.
For most large language models, they're bottlenecked by how quickly you can deliver them from the RAM, the prep counter, to the chefs and the cooks and the machines to process the model and get them to be running. Tina Huang, 10:30
On the Pi, the memory bus is narrow and slow, so it physically takes longer. And the Pi has only a CPU, no GPU and no NPU, which she describes as having a single head chef, very capable, but only one pair of hands having to do everything. That is two independent penalties stacked: the ingredients arrive slowly and there is only one person to cook them. The iPhone has a CPU, a GPU and an NPU, which is why it processes things so much faster.
Then the second half of the puzzle, image generation, which is a harder no than "slow":
When it comes to image, video, and music generation, it is really processing heavy. So, if you just only have CPUs, no matter how hard they work, it's just not enough to be able to generate music, images, and videos. Tina Huang, 11:20
That is the real answer to why a model can fit on the Pi and still be unusable. Text generation is bottlenecked on moving weights, so a narrow bus makes it slow. Image, video and music generation are bottlenecked on raw arithmetic, so no amount of patience substitutes for parallel hardware. A fitting model and a working model are two different claims.
She closes the lesson with her own caveat, which is more honest than most explainer videos manage:
Of course, what I said about AI hardware is much more simplified than it actually is. There's concepts like VRAM, unified memory, like so much more, which I'm not going to go into too much more detail about in this video, cuz I'm sure this video is already very long. But, this is like the general introduction that will help you a lot in figuring out what kind of models can run on what kind of devices. Tina Huang, 12:00
Quiz 1
At 11:53 she puts a pop quiz on screen and asks you to answer in the comments to make sure you were paying attention. The questions are on the slide rather than in the audio, so the page cannot reproduce them, but everything needed to answer them is in the lesson above: storage versus RAM, what the memory bus does, which of the three compute units is one capable generalist and which is a crowd, and which kinds of generation cannot be rescued by waiting.
The sponsor read: Crusoe Intelligence Foundry
The sponsored portion sits at 12:11, between the hardware lesson and the personal computer class, and she introduces it off the back of her own habit rather than as a hard cut. She has been tinkering with and running open source models all video, and the next obvious step, she says, is tinkering with and customizing the open source models themselves. For that she names Crusoe Intelligence Foundry.
Her description of what it gives you is two things:
- Serverless inference, to run open models without touching any of the underlying GPU infrastructure.
- Serverless fine tuning, to customize open models on your own data without having to own or manage any hardware.
The workflow she describes is fine tune a model, deploy it in one click, and start calling it with an API in minutes. Her reaction to it is the most quotable line in the read:
Seriously, the first time that I did this, I almost shed a tear. There is no cluster supervision or crazy long setups. I cannot even explain how painful this used to be. Tina Huang, 12:35
Her recommendation is scoped honestly: it is for anybody building with open models who does not want infrastructure overhead eating into actual build time. The offer stated on screen is 5 US dollars of free credit on sign up, with the link in the video description. She thanks Crusoe for sponsoring that portion and goes straight back to the hardware.
For what it is worth, the product matches the description. Serverless fine tuning in Intelligence Foundry is a LoRA based supervised workflow over a curated library of open weight base models, Qwen, DeepSeek, Llama and Gemma among them, priced per million tokens processed rather than per reserved GPU hour. That pricing model is the actual substance of the "no infrastructure overhead" claim: you are not holding a reservation open while you think.
Personal computers, class three: the 32 gigabyte laptop that does everything
Back to hardware, and the class most viewers are actually sitting in front of. Her example is the MacBook Pro she is recording on, with 32 gigabytes of RAM. She is explicit that the class is bigger than her machine: her older MacBook Air counts, and so do Windows laptops and Linux laptops. Portable personal machines, as a class.
The reason she bought 32 gigabytes specifically is the headline of the section:
This is the machine that pretty much allows me to run any category of AI models. Tina Huang, 13:35
"Any category" is not hand waving, and she lists them. On the 32 gigabyte MacBook Pro, locally: large language models, coding models, visual language models for visual analysis, image generation, video generation, text to speech, speech to text, audio generation, voice cloning, and music generation. That is the first rung on the ladder where no category is simply off the table, and it is worth pausing on how far that is from the top of the previous class. One rung down, the phone could not run a 7 billion parameter language model and the Pi could not generate an image at all.
The software layer, not just the models
The other thing that changes at this rung is that models stop being demos and start being plumbing. She runs each of these categories by itself, and she also runs them inside software that uses them:
- Hermes, her agent, which she says she uses to run many aspects of her life and her work. On it she runs pretty much any mid size to large size local model. (Hermes Agent is Nous Research's open source, self hosted agent framework, built around persistent memory and self created skills, and it is model agnostic by design, which is why a local model slots straight in.)
- Her go to models in that role are from the Qwen3 family. She names 14B, what the captions render as 16B, and 32B. A note on that middle number: Qwen3 ships dense models at 0.6B, 1.7B, 4B, 8B, 14B and 32B, plus a 30B-A3B mixture of experts model, and there is no 16B. Either the captions mangled it or she misspoke; the two sizes that definitely exist in that range are 14B and 32B.
- For coding she runs Qwen3 Coder 30B locally inside OpenCode, an open source coding harness, and reports it works pretty good.
- Her favorite for image generation is Flux Dev, Black Forest Labs' open weights image model.
She then does something most hardware channels will not do, which is state the cost honestly and immediately:
Now, I know what you're thinking already, "But Tina, I don't have a 32 gig MacBook Pro." Understandable. It is really expensive. I only got this machine because it is literally my job to be testing out models and things, right? Tina Huang, 14:40
Her answer is two more summary slides, one for 8 gigabyte laptops and one for 16 gigabyte laptops, which she calls the general RAM sizes of laptops that most people have. And then, for anything in between or outside those, the formula.
The formula: how big a model fits on your machine
This is the single most reusable thing in the video, and it takes about twenty seconds to state. Take the total amount of RAM your machine says it has. Subtract 25 percent of it, because that is used up by other stuff. What is left is your actual usable RAM. Divide that by 0.6, and the result is the number of billions of parameters your model can have.
Her worked example is her own machine. The MacBook Pro says 32 gigabytes. Subtract 25 percent and you have 24 gigabytes of usable RAM. Divide 24 by 0.6 and you get 40. So the MacBook Pro fits a model of up to 40 billion parameters. She calls it a very rough way of figuring it out, and offers the lazy alternative of asking a chatbot what fits on your machine, which she notes also works.
Three things are worth adding, because the formula rewards a second look:
It collapses to one multiplication. Taking 75 percent of a number and dividing by 0.6 is the same as multiplying it by 1.25. So the rule is just 1.25 billion parameters per gigabyte of RAM, which is a thing you can do in your head at a store. Her three worked tiers, 32, 64 and 128 gigabytes, give 40, 80 and 160 billion.
The 0.6 is a quantisation assumption in disguise. She never names a quantisation anywhere in the video, which is the one specification missing from an otherwise specific hour of work. But 0.6 gigabytes per billion parameters is 0.6 bytes per weight, which is about 4.8 bits per weight. That is squarely in four to five bit quantised territory, the range most people actually download. A model at full 16 bit precision needs about 2 gigabytes per billion parameters, more than three times her number, so her formula silently assumes you are running quantised weights. If you run anything unquantised, divide her answer by three and start again.
It leaves no room for context. The formula accounts for the weights and nothing else. A long conversation, a large context window, or a coding agent holding a repository in its head all need working memory on top of the weights, and that memory grows with how much you are asking the model to remember. This is almost certainly why, two rungs up, she tells you her 64 gigabyte Mac Studio runs models in the 27 to 32 billion range rather than the 80 billion the formula allows. The formula is a ceiling for what fits, not a recommendation for what to run.
Home servers, class four: 64 and 128 gigabytes, always on
The fourth class is two machines: a Mac Studio with 64 gigabytes and an AMD Ryzen AI Halo with 128. She is clear that plenty of other machines belong in the class, and names the cheapest entry point: a Mac mini, although she says those are really hard to get these days. Her own reason for owning a Studio is pure supply chain comedy. She only got the Mac Studio because she could not get a Mac mini.
The class exists for two reasons, and she gives both before any benchmarking.
One: it is a dedicated always on machine. A laptop you carry around and close, and when you close it, your models stop working if you were in the middle of something. A home server is designed to be always on. It does not shut down, so you can keep running your AI models around the clock.
Two: more capacity and more bandwidth. Generally, she says, these devices have both, so they can run bigger models and run them faster. Note that this is the first rung where she claims both axes improve at once. It does not hold all the way through the class, as the AMD machine demonstrates about ninety seconds later.
Quiz 2, asked in the middle of the video
Before she demos the Mac Studio she stops and gives the viewer the formula back as a test:
It says it has 64 gigs of RAM. How much of it is actually usable? And what is the largest model in billions of parameters that can fit on my Mac Studio? Were you paying attention? Tina Huang, 17:00
Run her own arithmetic: 64 minus 25 percent is 48 gigabytes usable, and 48 divided by 0.6 is 80 billion parameters. That is the answer she is fishing for in the comments, and it is deliberately larger than anything she actually runs on the machine, which is the gap discussed above.
The Arduino payoff
The best moment in the home server section has nothing to do with the server's own capabilities. She picks up the Arduino Uno R4, the board that could run nothing at all, and connects it to the Mac Studio over Wi-Fi.
This Arduino Uno is using the Qwen 3.6 35B model on this Mac Studio in order to display encouraging happy messages. Tina Huang, 17:30
A board with 32 kilobytes of RAM driving a 35 billion parameter model, because the model is not on the board. That is the whole argument for a home server in one object: the Studio is where she runs local models for her local AI agents, and it serves other devices too. She asks viewers to write what the Arduino's display says into the comments.
The model is Qwen3.6-35B-A3B, Alibaba's April 2026 mixture of experts release: 35 billion total parameters with about 3 billion active per token, Apache 2.0, and a natural fit for exactly this job, because a sparse model of that shape is far cheaper to serve continuously than a dense 35 billion would be.
The reframe: from new categories to bigger and faster
Here she makes the structural point of the back half of the video, and it is the sentence to keep:
Really with this Mac Studio and all devices that are bigger and more powerful than this, it's not about unlocking new categories of AI models at this point. It's just about being able to run a bigger models, higher quality things to get higher quality responses, and faster. Tina Huang, 17:40
Everything from the 32 gigabyte laptop up is a question of degree. The last genuine category unlock happened two rungs down, when a GPU first appeared and made image generation possible at all.
What the Mac Studio runs, in her list: large tier large language models between 27 and 32 billion parameters, large code models, large image models, and large music generation models. For video, it can now manage mid tier generation, like Wan 2.2, Alibaba's open weights video model (rendered by the captions as "One 2.2"), but her verdict is blunt: it is slow. It is really slow still.
The AMD Ryzen AI Halo: 128 gigabytes, and the first real trap
The 128 gigabyte machine is where the video's thesis pays off properly, because it is the first device that is clearly better on paper and clearly worse at something.
Her numbers: 128 gigabytes of advertised RAM, around 96 gigabytes of usable RAM. (That is her 25 percent rule applied again, and it checks out: 96 is three quarters of 128, which by the formula allows 160 billion parameters.) What that buys:
- Extra large models in the 70 billion range, comfortably. Llama 3.3 70B runs fine.
- Frontier models in the 100 to 250 billion range. gpt-oss-120b, OpenAI's open weights model, also runs fine.
- Extra large coding models, like DeepSeek Coder V2.
But the capability she singles out as most useful is not the size of any one model. It is holding several at once:
What I find the most useful of having something like this is the ability of having a lot of different models loaded simultaneously and working simultaneously because it has that massive amount of memory, right? So, I could be using a coding agent, the large language models, and doing like image video generation all at the same time. Tina Huang, 19:00
That is a genuinely different reason to buy memory than "fit a bigger model," and it is the one most people miss. A 96 gigabyte working set lets you keep a coding model, a chat model and a diffusion model resident simultaneously instead of paying the load cost every time you switch tasks.
Then the caveat, delivered immediately rather than buried:
I do want to make a caveat here though. This machine, even though it has very large capacity, it actually has pretty small bandwidth, which means that it is slow, especially when it comes to image, video, and music generation. It can do it. It's just going to do it really, really slow. Tina Huang, 19:22
Back to the kitchen. The Halo has an enormous prep counter and a narrow conveyor belt. You can lay out ingredients for four dishes at once, and then you wait.
GPUs, class five: buying cooks who bring their own counter
The final class is the answer to the Halo's problem, and she re explains the GPU before using one, going back to the kitchen for the third time.
A GPU is a brigade of line cooks, and because there are so many of them they work really quickly and do compute very fast. That matters most for music, image and video generation, because producing those needs a lot of processing power, which means a lot of cooks. The bad news she states plainly: every device covered so far does have a GPU, which is why they can generate images and music and video at all, but they do it really slowly because they are not specialized at it.
The good news is more discrete GPUs, and here she extends the analogy in the single most useful way in the video:
A GPU actually isn't just line cooks. It's actually a bundle. When you get another GPU, you get cooks, but they also come with their own private prep counters and a very wide conveyor belt, the memory bus. So, when you get more GPUs, you get more cooks who have their own private counter and their very own very wide conveyor belt. Tina Huang, 20:25
That one correction explains the whole top of the ladder. A discrete GPU is not a compute upgrade bolted onto your existing memory system, it is a second memory system with its own compute attached, and the belt it brings is wider than the one in your machine. It is why adding a GPU does not just make the cooking faster, it removes the delivery bottleneck too. Then, as she puts it, "it's getting really late," and she moves on.
The constraint on all of this is brutal and she does not soften it: you cannot add more GPUs to any device covered earlier in the video. The chips, the Pi, the phone, the laptops, the Mac Studio, the AMD Halo. None of them are built that way. A gaming PC is, which is the one consumer machine in this conversation that can grow. She does not have one and is not about to buy one, because they are expensive.
Cheating with a rented RTX 4090
So she rents. She is explicit that this is a cheat, and explicit about why she is moving fast:
I'm going to cheat a little bit, okay? Because I do want to show you what it's like to run these multi modality generation models on these specialized GPUs. So, I'm actually going to rent an RTX 4090. This is a pretty high end GPU that you might have in a high end gaming set. Let's do this quickly, because it's costing me by the hour. Tina Huang, 21:10
She logs in, runs an image generation model, and it is really, really quick. Then a video generation, also really, really quick. After four classes of hardware where generation was either impossible or painful, the RTX 4090 does it immediately. That is the demonstration, and it needs no numbers to land.
And then the downside, which is the mirror image of the Halo:
The downside of this GPU is that it only has 24 gigs of memory. So, you can think about it almost like the opposite of the AMD device, because the AMD device has a lot of memory, has very big capacity, but low bandwidth. While the RTX 4090 has smaller capacity, only 24 gigs, but much larger bandwidth. Tina Huang, 22:00
So you can generate anything you like, fast, and you cannot fit a large language model worth talking about. Her summary of the trade: you end up being able to generate a lot of these things and run these models, but you cannot really fit very large size models on it.
The final boss: 8x H100, 640 gigabytes
The top rung is a rented 8 way H100 node: 640 gigabytes of memory, massive capacity and massive bandwidth. She has to move fast on this one, because it is costing her money every minute, which she notes is worth it for the viewers.
What runs on it:
- Ultra large tier large language models, up to 700 billion parameters and above.
- Ultra large coding models, 480 billion parameters and above.
- Visual language processing as well.
- All types of multimodal generation.
The demo she picks is video, using MiniMax H3, and her claim is that it generates minutes of video, and does it so quickly. (MiniMax's H3, also branded Hailuo 3.0, is a 33 billion parameter omni modal model whose base weights were opened in August 2026. Worth knowing before you try to reproduce this: each generation from the open weights is documented as a clip in the 4 to 15 second range, so "minutes of video" is best read as total output rather than one continuous take. The open weights also max out at a 768 pixel short edge, and the community license excludes self hosting in the United States, the European Union, the United Kingdom and South Korea, which matters a great deal if you were planning to rent a node and follow along.)
Her framing of where that leaves the viewer is careful:
This is pretty much what you can run assuming that you're not like a literal model development company or just have like a lot of money. But, it really is comparable at this point to the frontier models that you're paying for through cloud APIs. Tina Huang, 23:15
And then the closing argument of the entire video, which retroactively organizes every rung below it:
You see everything up to this GPU is kind of just forcing you to choose between like capacity and bandwidth, right? Like the AMD Halo is all about big capacity but low bandwidth, and the RTX 4090 is all about small capacity and large bandwidth. But, this is when you get both. Tina Huang, 23:30
She calls it the aspirational class, puts up a final summary slide, and signs off by asking viewers to answer the on screen quiz in the comments to help retain what was covered.
Capacity or bandwidth: the choice the whole ladder makes for you
Her closing argument deserves to be drawn, because once you see the two axes separately, every result in the video stops being surprising. Capacity is the prep counter and decides what fits. Bandwidth is the belt and decides whether what fits is usable. Almost every machine you can buy is strong on one and weak on the other.
The whole ladder, as a reference table
| Device and class | Memory she states | Largest model that ran | Speed she reports | Her verdict |
|---|---|---|---|---|
| Arduino Uno R4 Chips | 32 KB | nothing fits | No model to run | Useless on its own for local AI. Useful over Wi-Fi as a display for a model running on a bigger machine |
| ESP32-S3 Chips | 8 MB | TinyStories 260K and TinyStories 3M | Comfortable at this size | Next word story generation only. No chat, no multimodal. Two to five dollars without a screen |
| Raspberry Pi 5 Mini computers | 8 GB | Qwen3 1.7B, chained with Whisper, Piper and Moondream | Slow: narrow bus, CPU only, no GPU or NPU | Micro to small LLMs easily, tiny vision and speech models fine. Mid size LLMs struggle. Image, music and video effectively impossible |
| iPhone 16 Mini computers | 8 GB, roughly 4 to 5 GB per app by iOS policy | Gemma 4, vision models 1 to 4B, Stable Diffusion 1.5 at 2 GB, Whisper Large, YOLO11n | Fast: CPU plus GPU plus NPU | No 7 to 14B models, by Apple's cap rather than by physics. Every model arrives inside an app |
| MacBook Pro Personal computers | 32 GB, 24 usable, 40B by her formula | Qwen3 32B, Qwen3 Coder 30B, Flux Dev | Workable in every category | The first rung where no category of model is off the table. Expensive, and she says she owns it only because testing models is her job |
| Mac Studio Home servers | 64 GB, 48 usable, 80B by her formula | Qwen3.6-35B-A3B served to other devices; LLMs 27 to 32B, large code, image and music | Faster than a laptop, but mid tier video is really slow | Always on and serves other devices. No new categories unlock here, just bigger models and faster answers |
| AMD Ryzen AI Halo Home servers | 128 GB advertised, around 96 usable | Llama 3.3 70B fine, gpt-oss-120b fine, DeepSeek Coder V2 | Big capacity, pretty small bandwidth, so generation is really slow | The real win is many models loaded and working at once, not one bigger model |
| RTX 4090, rented GPUs | 24 GB | Image and video generation, both immediate. No large language model worth the name | Very fast: much larger bandwidth | The exact opposite of the Halo. Cannot be added to any device earlier in the video |
| 8x H100, rented GPUs | 640 GB | LLMs above 700B, coding above 480B, vision language, MiniMax H3 video | Massive capacity and massive bandwidth | Comparable to the frontier models you pay for through cloud APIs. She calls it the aspirational class |
What unlocks where
| Model category | Chips, 8 MB | Mini computers, 8 GB | Laptop, 32 GB | Home server, 64 to 128 GB | Discrete GPU |
|---|---|---|---|---|---|
| Text language models | TinyStories only | Micro to small | Mid to large, 32B | Large to frontier, 120B | Ultra large, 700B+ |
| Coding models | no | Barely, small tier | Coder 30B | Extra large, Coder V2 | 480B+ |
| Vision language | no | Tiny on the Pi, 1 to 4B on the phone | yes | yes | yes |
| Speech to text | no | Small on the Pi, Whisper Large on the phone | yes | yes | yes |
| Text to speech | no | Tiny on the Pi, all of them on the phone | yes | yes | yes |
| Image generation | no | Phone only, 1B class | yes | Large, but slow | Very fast |
| Video generation | no | no | yes | Mid tier, really slow | Fast |
| Music and audio, including voice cloning | no | no | yes | Large, but slow | Fast |
| Object detection | no | YOLO11n on the phone | not named | not named | not named |
Key takeaways
- Sort hardware by memory class and "what can this run" stops being mysterious. That single organizing choice is what makes 24 minutes cover seven orders of magnitude of memory without losing the thread.
- Capacity decides what fits, bandwidth decides whether it is usable, and they are independent. A small image model fits on a Raspberry Pi 5 and is still useless there. A 120 billion parameter model fits on the AMD Ryzen AI Halo and generates images really slowly. Checking only the gigabytes on the box is how people buy the wrong machine.
- Text generation and media generation fail for different reasons. Language models are bottlenecked on moving weights from memory, so a narrow bus makes them slow. Image, video and music generation are bottlenecked on raw arithmetic, so without parallel hardware they do not become slow, they become impractical.
- Her formula is 1.25 billion parameters per gigabyte of RAM. Subtract 25 percent for the system, divide by 0.6, or just multiply by 1.25. It is a ceiling on weights at roughly four to five bit quantisation, not a recommendation, and it leaves nothing for context.
- Only two rungs unlock new categories; everything above is degree. A language model first becomes possible at 8 megabytes. Vision, speech and detection arrive at 8 gigabytes. From the 32 gigabyte laptop up, as she says outright, it is no longer about new categories, just bigger models, better answers and more speed.
- Apple's per app memory cap is a real ceiling, not a technicality. The iPhone 16 has the same 8 gigabytes as the Pi and cannot run 7 to 14 billion parameter models because iOS allows an app roughly 4 to 5 gigabytes, which is policy rather than physics.
- A discrete GPU is a second memory system, not a compute add on. Her correction to her own analogy, that a GPU brings its own prep counter and its own wider belt, is why adding one fixes both the arithmetic and the delivery problem at once, and why no device below the gaming PC in this video can be upgraded that way.
- The most useful thing 96 gigabytes buys is not a bigger model, it is several resident models. A coding agent, a chat model and an image model all loaded at once, with no reload cost when you switch tasks.
- The machine with no compromise is the one nobody owns. Both rungs that report massive capacity and massive bandwidth are rented by the hour, and she says so out loud.
Chapters
The nine entries in bold are Tina's own chapters, reproduced as written. The rest are sub beats added here from the transcript clock, because nine markers across 24 minutes leaves a lot of ground unmarked.
- 0:00 Intro: the premise, the hardware on the table, and the Crusoe disclosure
- 0:16 Chips
- 0:30 The Arduino Uno R4, 32 KB, and the number of models that fit: none
- 1:15 The ESP32-S3, 8 MB, around 20 dollars with a screen
- 1:31 TinyStories 260K and TinyStories 3M, both running comfortably
- 2:01 Project ideas: a grammar game, and LED lights for the story's mood
- 2:31 What a 3 million parameter model can and cannot do
- 2:58 Mini Computers (Raspberry Pi, iPhone 16)
- 3:15 The puzzle is set: same 8 GB, different models
- 3:34 The Pi 5 answers out loud, through Whisper, Qwen3 1.7B and Piper
- 4:34 Moondream describes a webcam frame, read aloud by Piper
- 5:04 The Pi's full capability report, and why it cannot generate an image
- 5:35 Accessories, sensors, and the Hugging Face robot that runs on a Pi
- 6:06 The iPhone 16, and Apple's cap of roughly 4 to 5 GB per app
- 6:37 Gemma 4 through the Google AI Edge Gallery
- 6:55 Draw Things, and Stable Diffusion 1.5 at 2 GB
- 7:08 YOLO11n object detection through the Ultralytics app
- 7:50 Hardware Lesson
- 8:10 The staff: one head chef, a brigade of line cooks, bolted on machines
- 8:41 Cooking a salad, then cooking a model
- 9:11 The mapping: storage, RAM, memory bus, CPU, GPU, NPU
- 10:11 Why the Pi and the iPhone diverge on identical memory
- 11:11 Why no amount of CPU patience will generate an image
- 11:53 Quiz 1
- 12:00 Her own caveat: VRAM, unified memory, and what she left out
- 12:11 Crusoe Intelligence Foundry, the sponsored portion
- 13:18 Personal Computers
- 13:35 Every category of model, on one 32 GB laptop
- 14:14 Hermes, the Qwen3 family, Qwen3 Coder 30B in OpenCode, Flux Dev
- 14:40 On cost, honestly, plus slides for 8 GB and 16 GB laptops
- 15:05 The formula: subtract 25 percent, divide by 0.6
- 15:15 Worked through: 32 GB gives 40 billion parameters
- 16:07 Home Servers (Mac Studio, AMD Ryzen AI Halo)
- 16:17 Why the class exists: always on, with more capacity and bandwidth
- 17:00 The 64 GB quiz, asked in the middle of the video
- 17:30 The Arduino payoff: a 32 KB board driving a 35B model
- 17:40 The reframe: above here it is bigger and faster, not new
- 17:49 What the Mac Studio runs, and Wan 2.2 being really slow
- 18:19 The AMD Ryzen AI Halo, 128 GB advertised, around 96 usable
- 18:50 Llama 3.3 70B, gpt-oss-120b and DeepSeek Coder V2
- 19:00 The real win: many models loaded and working at once
- 19:22 The caveat: large capacity, pretty small bandwidth
- 19:45 GPUs (RTX 4090, 8X W100, 8X H100)
- 20:25 A GPU is a bundle: cooks, their own counter, and a wide belt
- 20:53 You cannot add GPUs to anything covered so far
- 21:10 Renting an RTX 4090, by the hour
- 21:50 Image generation, then video generation, both immediate
- 22:00 The downside: 24 GB, the mirror image of the Halo
- 22:27 The final boss: a rented 8x H100 node, 640 GB
- 22:57 MiniMax H3 generating video faster than it plays back
- 23:30 The closing argument: capacity or bandwidth, until right here
- 24:15 Quiz 2
One oddity in her own chapter titles, left as written above: the GPU chapter is labelled with an "8X W100" alongside the 8X H100, but no such machine appears anywhere in the video. The GPU section covers exactly two rented machines, the RTX 4090 and the 8x H100 node.
Notable quotes
When we try to run an AI model on these, we're not even installing it. We just literally stick it on here and if it fits, it sits and we can run it, or not. Tina Huang, on the chips class, 0:25
Which is big enough to fit, drumroll, please, no models. That's right, nothing fits on here. Alas. Tina Huang, on the Arduino Uno R4's 32 kilobytes, 0:40
I'm able to have full conversations with my Pi by chaining together three different AI models, all on this tiny little device. Tina Huang, on the Raspberry Pi 5 voice assistant, 4:10
Even though it technically has 8 gigs of RAM, you actually aren't able to run the mid size level of models from 7 to 14 B that you can't do using the Raspberry Pi, because like Apple just doesn't let you do it, unfortunately. Tina Huang, on the iPhone 16, 6:15
Basically these machines can only do one thing. It does that one thing really well and really fast and it uses almost no electricity. Tina Huang, on the bolted on kitchen machines that stand in for the NPU, 8:35
It's basically like having a single head chef, very capable, but only a single pair of hands having to do everything to run this model. Tina Huang, on why the Raspberry Pi is slow, 10:50
I only got this machine because it is literally my job to be testing out models and things, right? Tina Huang, on the 32 gigabyte MacBook Pro, 14:40
I actually only got the Mac Studio because I couldn't get a Mac Mini. Tina Huang, on why she owns a Mac Studio, 16:20
Really with this Mac Studio and all devices that are bigger and more powerful than this, it's not about unlocking new categories of AI models at this point. It's just about being able to run a bigger models, higher quality things to get higher quality responses, and faster. Tina Huang, the structural claim of the back half, 17:40
I hope that analogy makes sense. It's getting really late. Tina Huang, after extending the kitchen analogy a third time, 20:50
Now, I really got to move fast on this one, because it's costing me so much money every minute. But, for you guys, it's worth it. Tina Huang, on the rented 8x H100 node, 22:35
You see everything up to this GPU is kind of just forcing you to choose between like capacity and bandwidth, right? But, this is when you get both. Tina Huang, the closing argument, 23:30
Resources mentioned
The hardware, bottom to top
- Arduino Uno R4 WiFi, 32 KB of RAM (documentation). The board that fits nothing and ends up driving a 35 billion parameter model anyway.
- ESP32-S3 from Espressif, 8 MB of RAM, around 20 US dollars with a screen and two to five without.
- Raspberry Pi 5, 8 GB.
- iPhone 16, 8 GB, with iOS capping each app at roughly 4 to 5 GB.
- MacBook Pro, 32 GB, and her older MacBook Air.
- Mac Studio, 64 GB, and the Mac mini she names as the cheapest entry to the class.
- AMD Ryzen AI Halo, 128 GB advertised, built on the Ryzen AI Max+ 395.
- NVIDIA RTX 4090, 24 GB, rented by the hour.
- NVIDIA H100, eight of them, 640 GB in total, also rented.
The models
- TinyStories-260K, the smallest checkpoint from Andrej Karpathy's llama2.c, and TinyStories-3M from the original TinyStories release by Ronen Eldan.
- Whisper for speech to text, including Whisper large v3 on the phone.
- Qwen3 1.7B on the Pi, Qwen3 14B and Qwen3 32B on the laptop, Qwen3 Coder 30B for code, and Qwen3.6-35B-A3B on the Mac Studio, all from Qwen at Alibaba.
- Piper for text to speech.
- Moondream, the tiny vision language model that describes the webcam frame (weights).
- Gemma 4 from Google DeepMind, downloaded through the Google AI Edge Gallery.
- Stable Diffusion 1.5, 2 GB, run on the phone.
- YOLO11n for object detection, via the Ultralytics app.
- FLUX.1 dev from Black Forest Labs, her favorite for image generation.
- Llama 3.3 70B from Meta and gpt-oss-120b from OpenAI, both running fine on 128 GB.
- DeepSeek Coder V2 from DeepSeek.
- Wan 2.2 for mid tier video generation on the Mac Studio.
- MiniMax H3, also branded Hailuo 3.0, from MiniMax, for video on the eight H100 node.
Software, apps and services
- Hermes Agent from Nous Research, the agent she runs much of her life and work on.
- OpenCode, the open source coding harness she runs Qwen3 Coder 30B inside.
- Draw Things, the iOS app for local image generation.
- Crusoe Intelligence Foundry, the sponsor, with serverless fine tuning and serverless inference for open models, and 5 US dollars of free credit on sign up.
- Hugging Face, where most of the above lives, and Reachy Mini, the open source robot whose wireless version runs on a Raspberry Pi 5, which is the robot she is thinking of.
The channel
- Tina Huang on YouTube, where the free guide and the quick start guides she references are linked from the video description.
An honest footnote
Four things are worth saying once, at the end, for anyone planning to act on this video rather than just enjoy it.
There are no speed numbers in it. Not one tokens per second figure, not one seconds per image, across all nine machines. Every speed claim is comparative and verbal: really quick, painfully slow, really really slow. That is a legitimate choice for a video whose argument is about which axis binds rather than by how much, and the argument lands without numbers. But if you are trying to decide between two machines in the same class, this video will tell you which axis to think about and will not tell you what to expect.
No quantisation is ever named, and no runtime either. The 0.6 gigabytes per billion parameters in her formula implies roughly four to five bit weights, which is the only quantisation signal in 24 minutes, and it is implicit. She also never names the software actually loading these models on any device, so if you want to reproduce the Pi 5 voice assistant or the 120 billion parameter model on 96 gigabytes, you are supplying the missing layer yourself. The obvious candidates, neither of them mentioned in the video, are llama.cpp and Ollama, which is where the four and five bit quantised weights her formula assumes actually come from.
The top two rungs are rented, and she says so. "I'm going to cheat a little bit" is her own framing, and renting an RTX 4090 and an eight way H100 node by the hour is a different proposition from everything below it, where the hardware sits on her desk and the electricity is the only running cost. The demos are real and the capability claims are sound; the word "local" is doing lighter work at the top of the ladder than at the bottom. The honest reading is that the top two rungs are a preview of what the other seven cannot do, not a recommendation.
One claim is worth checking before you copy it. Her 8x H100 demo generates "minutes of video" with MiniMax H3. The open weights release of H3 is documented as producing clips in the 4 to 15 second range at up to a 768 pixel short edge, so minutes of output means many clips rather than one long take. More practically, the community license on those open weights excludes self hosting in the United States, the European Union, the United Kingdom and South Korea, so reproducing that particular demo on rented hardware in those places is not simply a question of affording the node.
None of that undercuts the video. The lesson it exists to teach, that capacity and bandwidth are separate questions and almost every machine answers only one of them well, is correct, well taught, and the exact thing people get wrong when they spend money on hardware for local AI. The ladder is the proof, and the Arduino driving a 35 billion parameter model over Wi-Fi is the best single image of it: the smallest thing on the table was never limited by what fit inside it.


