youtube.nixfred.com nixfred.com

Every Size Local AI In 24 Minutes

Tina Huang lines up every computer she owns, sorts it by memory class, and tries to run an AI model on every rung: a 32 kilobyte Arduino that fits nothing, an 8 megabyte ESP32-S3 running TinyStories, a Raspberry Pi 5 chaining Whisper, Qwen3 1.7B and Piper into a talking assistant, an iPhone 16 with the same 8 gigabytes that behaves nothing alike, a 32 gigabyte MacBook Pro, a 64 gigabyte Mac Studio, a 128 gigabyte AMD Ryzen AI Halo, and rented GPUs up to an eight way H100 node with 640 gigabytes. A restaurant kitchen analogy explains why capacity and bandwidth are separate questions, and a one line formula tells you exactly what fits on your own machine.

Published Sep 28, 2026 24:22 video 53 min read Added Oct 8, 2026 Open on YouTube →

At a glance

Tina Huang lays out every piece of computing hardware she owns, sorts it into classes by how much memory it has, and tries to run an AI model on every rung. The bottom rung is an Arduino Uno R4 with 32 kilobytes of RAM, which fits no model at all; the top rung is a rented 8x H100 node with 640 gigabytes, running 700 billion parameter models and generating video faster than it plays back. In between she runs a three model voice assistant on a Raspberry Pi 5, discovers that her iPhone 16 has exactly the same 8 gigabytes as the Pi and yet behaves nothing like it, and uses a restaurant kitchen to explain why. The lesson underneath the whole climb is that memory capacity decides what fits and memory bandwidth decides whether it is usable, and that almost every machine you can buy forces you to pick one. She also hands over a formula for the largest model your own laptop can hold, which turns out to reduce to a single number you can do in your head. A portion of the video is sponsored by Crusoe; the read sits at 12:11 and is covered in full below.

The format: if it fits, it sits

The premise is stated in the first fifteen seconds and never complicated. She has a pile of hardware in front of her, she is going to run AI locally on all of it, and the hope is that you come away with ideas for things to build rather than a shopping list. The organizing principle is memory class, not price and not year of release, and that choice is what makes the video work: once you sort hardware by how much memory it has, the question "what can this run" stops being mysterious.

The first class is the one most hardware videos skip entirely. She calls it chips, and she means it literally.

Chips, as they are called, are literally just chips, just computer chips. They have no operating system, no file system, and no separate memory stick, either. Literally everything is just built into these boards. Tina Huang, 0:16

That has a consequence she turns into the running gag of the whole video. On a chip there is no install step, no package manager, no model directory. You flash the thing and find out.

When we try to run an AI model on these, we're not even installing it. We just literally stick it on here and if it fits, it sits and we can run it, or not. Tina Huang, 0:25

"If it fits, it sits" is the test for every rung above the chips too, and the chart below is that test applied to the full set of hardware in the video.

MEMORY ON EVERY RUNG, AND THE BIGGEST MODEL THAT RAN ON IT Arduino Uno R4, 32 KB nothing fits ESP32-S3, 8 MB TinyStories 3M Raspberry Pi 5, 8 GB Qwen3 1.7B iPhone 16, 8 GB VLM up to 4B RTX 4090 rented, 24 GB 24 GB ceiling MacBook Pro, 32 GB Qwen3 32B Mac Studio, 64 GB Qwen3.6 35B-A3B AMD Ryzen AI Halo, 128 GB gpt-oss-120b 8x H100 rented, 640 GB 700B+ 32 KB 8 MB 1 GB 32 GB 1 TB LOG SCALE: EACH STEP RIGHT IS TEN TIMES MORE MEMORY
Figure 1. Every device in the video, sorted by memory rather than by the order she presents them, which moves the rented RTX 4090 down to where its 24 gigabytes actually put it. Amber bars are the two machines she rents by the hour; blue bars are hardware she owns. On a log scale the whole climb is seven decades wide, and the two rungs that change what is possible, rather than just how big, sit at 8 megabytes and 8 gigabytes.

Chips, class one: 32 kilobytes and 8 megabytes

The Arduino that fits nothing

The smallest board on the table is an Arduino Uno R4 with 32 kilobytes of RAM. She promises the rest of the specifications on screen, then delivers the punchline with a drumroll: it is big enough to fit no models. Nothing. She bought it as part of a set because she is trying to learn hardware, and it turned out to be the one board in the video that cannot hold a language model of any size.

That does not make it useless for local AI, and she flags the reason immediately: the Uno R4 has Wi-Fi, so it can talk to her bigger machines and use the models running on them. She parks the idea at 0:55 and pays it off seventeen minutes later, which is one of the better structural jokes in the video.

The ESP32-S3, the cheapest board in the video that runs a model

The next board up is an ESP32-S3 with 8 megabytes of RAM. She paid around 20 US dollars for hers because it came with a screen; without the screen, she says, you are looking at two to five dollars. That is the price of a sandwich for a board that runs a transformer.

Two models fit, and both are from the TinyStories family, which is exactly what the name suggests: models trained to write tiny stories. She runs what the captions render as the "1 megabyte" and "3 megabyte" models, and then corrects herself out loud to the real names.

So, yes, I can run the Tiny Stories 260K model and the Tiny Stories 3M model very comfortably. Tina Huang, 1:40

One correction worth making explicit, because the captions garble it every time: those numbers are parameter counts, not file sizes. TinyStories-260K is a 260 thousand parameter model, the smallest checkpoint from Andrej Karpathy's llama2.c project, and TinyStories-3M is a 3 million parameter model from Ronen Eldan's original TinyStories release. They are not 1 megabyte and 3 megabyte downloads. Getting this right matters if you go looking for them, and it also sets up the arithmetic later: a 260 thousand parameter model is tiny by any modern standard and still too large for the Arduino sitting next to it.

Both run comfortably on the ESP32-S3, and she is clear about what "running" buys you at this size. These models do the very basics of predicting the next word in order to form a coherent story, which she points out is genuinely impressive given the scale. What they cannot do is hold a conversation like a chat model, and they cannot touch any multimodal work: no audio, no music, no video, no images.

What you actually build with a 3 million parameter model

This is where the video earns its "ideas for cool things to build" framing instead of just benchmarking. Getting the model running, she says, is just the start. You make it cooler by bolting modules onto the board.

Her two examples: a little game called scrambled versus not scrambled, which she suggests might be useful if you are learning English grammar, and LED lights that represent the mood of the story being produced. A 3 million parameter model producing a story, and a strip of lights reading the story's emotional temperature in real time, on a five dollar chip. She floats a whole separate video on building with these little devices, decides there is no time for it here, and puts up a summary slide of the ESP32 models plus project ideas with the instruction to take a screenshot. There is also a quick start guide in the free guide linked from the video description.

Mini computers, class two: two devices with 8 gigabytes that behave nothing alike

The second class is where the video's central puzzle gets set up. She puts two things on the table: a Raspberry Pi 5 and her personal iPhone 16.

What separates this class from the chips is not just more memory, it is being a real computer. These run an operating system, Linux on the Pi and iOS on the phone, which means they manage memory, have a file system, and run multiple programs at once. The practical upshot for local AI is that you can keep several models around and use them together, as long as they fit.

Then the setup lands:

What is really interesting is that both of these actually have 8 gigs of RAM. Yet, they are not able to run the same kinds of models. Tina Huang, 3:15

She promises the hardware lesson that explains it and makes you wait for it, which is the right call, because the demos land harder before the theory than after.

The Raspberry Pi 5: three models chained into a conversation

The Pi 5 demo is the best few minutes in the video. She asks it a question out loud and it answers out loud, with the whole pipeline running on the board.

Tina: What color is an octopus?
The Pi: Octopuses come in a variety of colors, but they are usually camouflage, blending with their surroundings. They can be blue, green, red, or even white, depending on the species. the Raspberry Pi 5 voice assistant, 3:34

The chain is three models deep, and she names every link:

Her own framing: full conversations with a Pi, by chaining together three different AI models, all on one tiny device.

Then she adds a fourth capability without adding a bigger machine. A webcam plugged into the Pi grabs a frame, a tiny visual language model called Moondream describes what is in it, and Piper reads the description out. The output is worth quoting in full, because the level of detail from a vision model running on an 8 gigabyte single board computer is the actual demonstration:

The image features a woman sitting in a chair looking at her cell phone. She is wearing a black shirt and is surrounded by a variety of books and other items. There are at least 13 books scattered around the room, some of which are placed on a bookshelf. A backpack can be seen near the woman. A potted plant is located close to her. The scene suggests that the woman might be engaged in reading or studying, as she is surrounded by numerous books and other items. Moondream, running on the Raspberry Pi 5, 4:34

What the Pi 5 can and cannot do, in her words

She gives the Pi a precise capability report rather than a thumbs up, and the precision is the value:

That last line comes with the most important caveat in the section, and it is the first crack in the "memory is destiny" idea:

It really cannot do image generation, music generation, or video generation. Not actually because a small image generation model wouldn't fit on here. It does. The problem is that it doesn't have any GPUs, so it's just like painfully slow to do that. Tina Huang, 5:20

The model fits. It sits. It is still useless, because fitting was never the only question. She defers the GPU explanation and moves on.

What the Pi gives back for that weakness is buildability. You can attach an enormous range of accessories and sensors, which is the whole reason the board exists, and she notes that she believes the Hugging Face robots run off a Raspberry Pi as well. That is correct: the wireless version of Reachy Mini, Hugging Face's open source desktop robot, has a Raspberry Pi 5 inside it doing exactly this job. Summary slide, project ideas, and another pointer to the free quick start guide.

The iPhone 16: same 8 gigabytes, a different machine entirely

The phone has the same 8 gigabytes as the Pi and is, in her words, so different. The first reason has nothing to do with silicon:

Because it is running on iOS, Apple does limit each app to only be able to use around 4 or 5 gigs of RAM. So, even though it technically has 8 gigs of RAM, you actually aren't able to run the mid size level of models from 7 to 14 B that you can't do using the Raspberry Pi, because like Apple just doesn't let you do it, unfortunately. Tina Huang, 6:10

So the phone's effective ceiling for a single model is roughly half its physical memory, by policy rather than by physics. The second Apple constraint is distribution: you cannot put whatever you like on the device, you go through the official system, which means every model arrives inside an app that downloads, installs and runs it. That shapes the whole rest of the section, because each capability she demonstrates comes with a named app.

The detection model gets the longest aside in the section, and she is right that it is underrated. YOLO11n does object recognition to understand a scene, and she points out it is the underlying technology behind everything from facial recognition on your devices to self driving cars, because you have to be able to recognize objects as you drive. Her verdict on the category:

We don't really talk about this class of models very much, these detection models, but they're actually very, very useful. Tina Huang, 7:30

Note what just happened across these two devices. The Pi, with a full Linux box's freedom, runs a bigger language model than the phone and cannot generate a single image. The phone, locked down to roughly 4 to 5 gigabytes per app, generates images in seconds. Same memory, opposite strengths. That is the puzzle the next section solves.

The hardware lesson: a device is a restaurant kitchen

At 7:50 she stops demoing and teaches, and the analogy she picks is good enough that it is worth rebuilding piece by piece rather than summarizing. Think about any device as a restaurant kitchen.

The spaces. Downstairs there is a walk in freezer and a storage space, where all the ingredients and all the things are kept. Upstairs there is a prep counter, where you load up only the ingredients you need right now for the dish you want to make.

The staff. There is one head chef, very advanced, able to make any sort of dish, but just the one of them. There is a brigade of line cooks, less experienced and less versatile, but there are a lot of them and they can help. And there are specialized machines bolted to the walls: an onion machine that only chops onions, a noodle slicing machine that automatically slices noodles. Her description of that third category is the part that makes the mapping click later:

Basically these machines can only do one thing. It does that one thing really well and really fast and it uses almost no electricity. Tina Huang, 8:35

The connection. Between the prep counter and the chefs, the cooks and the machines, there is a conveyor belt passageway. That is how ingredients get from the counter to whoever is going to process them.

First, make a salad

She runs the kitchen once on a real dish before mapping anything, which is what makes the analogy stick. To prepare a salad: go to deep storage in the freezer and pull the salad things out, carrots, lettuce, olives, chicken, other salad things, and put them on the prep counter. Put them on the conveyor belt. The head chef starts giving out orders, tells the line cooks to chop all the vegetables, starts up the onion machine to chop the onions, and handles only the tricky part himself, mixing the sauce. A few minutes later everything comes together and you have a salad.

Now make a model

Same kitchen, different dish. This time what you are cooking is a model.

WHAT HAPPENS WHEN A DEVICE RUNS A MODEL STORAGE Freezer the model at rest RAM Prep counter only what you need now CAPACITY what fits MEMORY BUS Conveyor belt BANDWIDTH how fast it arrives CPU One head chef capable, and directs the rest GPU A brigade of line cooks many hands, fast in parallel NPU The onion machine one job, fast, almost no power "Hello, I am good. How are you?"
Figure 2. Tina Huang's kitchen analogy, mapped straight onto the hardware. The two annotated boxes are the ones that decide everything on the ladder: the prep counter sets capacity, which is what fits, and the belt sets bandwidth, which is how fast it arrives. Every surprise in the rest of the video is one of those two being the binding constraint while the other looks fine.

The mapping, in her order:

And after all that cooking, the model produces something. You ask "Hi, how are you?" and it says "Hello, I am good. How are you?"

The point she drives home next is the one most people skip: this is not a one time setup cost.

Every time we want to be working with the model, we need to be going through this entire process of cooking. Tina Huang, 9:55

Why the Pi and the iPhone diverge

With the kitchen built, she goes back and collects the debt from 3:15. Both devices have 8 gigabytes, so both have the same prep counter. What they do not have is the same conveyor belt.

For most large language models, they're bottlenecked by how quickly you can deliver them from the RAM, the prep counter, to the chefs and the cooks and the machines to process the model and get them to be running. Tina Huang, 10:30

On the Pi, the memory bus is narrow and slow, so it physically takes longer. And the Pi has only a CPU, no GPU and no NPU, which she describes as having a single head chef, very capable, but only one pair of hands having to do everything. That is two independent penalties stacked: the ingredients arrive slowly and there is only one person to cook them. The iPhone has a CPU, a GPU and an NPU, which is why it processes things so much faster.

Then the second half of the puzzle, image generation, which is a harder no than "slow":

When it comes to image, video, and music generation, it is really processing heavy. So, if you just only have CPUs, no matter how hard they work, it's just not enough to be able to generate music, images, and videos. Tina Huang, 11:20

That is the real answer to why a model can fit on the Pi and still be unusable. Text generation is bottlenecked on moving weights, so a narrow bus makes it slow. Image, video and music generation are bottlenecked on raw arithmetic, so no amount of patience substitutes for parallel hardware. A fitting model and a working model are two different claims.

She closes the lesson with her own caveat, which is more honest than most explainer videos manage:

Of course, what I said about AI hardware is much more simplified than it actually is. There's concepts like VRAM, unified memory, like so much more, which I'm not going to go into too much more detail about in this video, cuz I'm sure this video is already very long. But, this is like the general introduction that will help you a lot in figuring out what kind of models can run on what kind of devices. Tina Huang, 12:00

Quiz 1

At 11:53 she puts a pop quiz on screen and asks you to answer in the comments to make sure you were paying attention. The questions are on the slide rather than in the audio, so the page cannot reproduce them, but everything needed to answer them is in the lesson above: storage versus RAM, what the memory bus does, which of the three compute units is one capable generalist and which is a crowd, and which kinds of generation cannot be rescued by waiting.

The sponsor read: Crusoe Intelligence Foundry

The sponsored portion sits at 12:11, between the hardware lesson and the personal computer class, and she introduces it off the back of her own habit rather than as a hard cut. She has been tinkering with and running open source models all video, and the next obvious step, she says, is tinkering with and customizing the open source models themselves. For that she names Crusoe Intelligence Foundry.

Her description of what it gives you is two things:

The workflow she describes is fine tune a model, deploy it in one click, and start calling it with an API in minutes. Her reaction to it is the most quotable line in the read:

Seriously, the first time that I did this, I almost shed a tear. There is no cluster supervision or crazy long setups. I cannot even explain how painful this used to be. Tina Huang, 12:35

Her recommendation is scoped honestly: it is for anybody building with open models who does not want infrastructure overhead eating into actual build time. The offer stated on screen is 5 US dollars of free credit on sign up, with the link in the video description. She thanks Crusoe for sponsoring that portion and goes straight back to the hardware.

For what it is worth, the product matches the description. Serverless fine tuning in Intelligence Foundry is a LoRA based supervised workflow over a curated library of open weight base models, Qwen, DeepSeek, Llama and Gemma among them, priced per million tokens processed rather than per reserved GPU hour. That pricing model is the actual substance of the "no infrastructure overhead" claim: you are not holding a reservation open while you think.

Personal computers, class three: the 32 gigabyte laptop that does everything

Back to hardware, and the class most viewers are actually sitting in front of. Her example is the MacBook Pro she is recording on, with 32 gigabytes of RAM. She is explicit that the class is bigger than her machine: her older MacBook Air counts, and so do Windows laptops and Linux laptops. Portable personal machines, as a class.

The reason she bought 32 gigabytes specifically is the headline of the section:

This is the machine that pretty much allows me to run any category of AI models. Tina Huang, 13:35

"Any category" is not hand waving, and she lists them. On the 32 gigabyte MacBook Pro, locally: large language models, coding models, visual language models for visual analysis, image generation, video generation, text to speech, speech to text, audio generation, voice cloning, and music generation. That is the first rung on the ladder where no category is simply off the table, and it is worth pausing on how far that is from the top of the previous class. One rung down, the phone could not run a 7 billion parameter language model and the Pi could not generate an image at all.

The software layer, not just the models

The other thing that changes at this rung is that models stop being demos and start being plumbing. She runs each of these categories by itself, and she also runs them inside software that uses them:

She then does something most hardware channels will not do, which is state the cost honestly and immediately:

Now, I know what you're thinking already, "But Tina, I don't have a 32 gig MacBook Pro." Understandable. It is really expensive. I only got this machine because it is literally my job to be testing out models and things, right? Tina Huang, 14:40

Her answer is two more summary slides, one for 8 gigabyte laptops and one for 16 gigabyte laptops, which she calls the general RAM sizes of laptops that most people have. And then, for anything in between or outside those, the formula.

The formula: how big a model fits on your machine

This is the single most reusable thing in the video, and it takes about twenty seconds to state. Take the total amount of RAM your machine says it has. Subtract 25 percent of it, because that is used up by other stuff. What is left is your actual usable RAM. Divide that by 0.6, and the result is the number of billions of parameters your model can have.

HOW BIG A MODEL FITS, IN FOUR STEPS STEP 1 32 GB what the machine says STEP 2: MINUS 25% 24 GB usable the OS takes a cut STEP 3: DIVIDE BY 0.6 24 / 0.6 = 40 GB per billion weights STEP 4: THE ANSWER 40 billion parameters that fit 32 GB gives 40B 64 GB gives 80B 128 GB gives 160B THE SAME RULE IN ONE STEP: 1.25 BILLION PARAMETERS PER GIGABYTE
Figure 3. Her formula, worked exactly as she states it, plus the shortcut that falls out of it. Subtracting a quarter and dividing by 0.6 is the same as multiplying by 1.25, so the whole rule collapses to 1.25 billion parameters per gigabyte of RAM. Run it backwards on the smallest board in the video and it explains the opening joke: 32 kilobytes allows roughly 40 thousand parameters, and the smallest model in her kit has 260 thousand.

Her worked example is her own machine. The MacBook Pro says 32 gigabytes. Subtract 25 percent and you have 24 gigabytes of usable RAM. Divide 24 by 0.6 and you get 40. So the MacBook Pro fits a model of up to 40 billion parameters. She calls it a very rough way of figuring it out, and offers the lazy alternative of asking a chatbot what fits on your machine, which she notes also works.

Three things are worth adding, because the formula rewards a second look:

It collapses to one multiplication. Taking 75 percent of a number and dividing by 0.6 is the same as multiplying it by 1.25. So the rule is just 1.25 billion parameters per gigabyte of RAM, which is a thing you can do in your head at a store. Her three worked tiers, 32, 64 and 128 gigabytes, give 40, 80 and 160 billion.

The 0.6 is a quantisation assumption in disguise. She never names a quantisation anywhere in the video, which is the one specification missing from an otherwise specific hour of work. But 0.6 gigabytes per billion parameters is 0.6 bytes per weight, which is about 4.8 bits per weight. That is squarely in four to five bit quantised territory, the range most people actually download. A model at full 16 bit precision needs about 2 gigabytes per billion parameters, more than three times her number, so her formula silently assumes you are running quantised weights. If you run anything unquantised, divide her answer by three and start again.

It leaves no room for context. The formula accounts for the weights and nothing else. A long conversation, a large context window, or a coding agent holding a repository in its head all need working memory on top of the weights, and that memory grows with how much you are asking the model to remember. This is almost certainly why, two rungs up, she tells you her 64 gigabyte Mac Studio runs models in the 27 to 32 billion range rather than the 80 billion the formula allows. The formula is a ceiling for what fits, not a recommendation for what to run.

Home servers, class four: 64 and 128 gigabytes, always on

The fourth class is two machines: a Mac Studio with 64 gigabytes and an AMD Ryzen AI Halo with 128. She is clear that plenty of other machines belong in the class, and names the cheapest entry point: a Mac mini, although she says those are really hard to get these days. Her own reason for owning a Studio is pure supply chain comedy. She only got the Mac Studio because she could not get a Mac mini.

The class exists for two reasons, and she gives both before any benchmarking.

One: it is a dedicated always on machine. A laptop you carry around and close, and when you close it, your models stop working if you were in the middle of something. A home server is designed to be always on. It does not shut down, so you can keep running your AI models around the clock.

Two: more capacity and more bandwidth. Generally, she says, these devices have both, so they can run bigger models and run them faster. Note that this is the first rung where she claims both axes improve at once. It does not hold all the way through the class, as the AMD machine demonstrates about ninety seconds later.

Quiz 2, asked in the middle of the video

Before she demos the Mac Studio she stops and gives the viewer the formula back as a test:

It says it has 64 gigs of RAM. How much of it is actually usable? And what is the largest model in billions of parameters that can fit on my Mac Studio? Were you paying attention? Tina Huang, 17:00

Run her own arithmetic: 64 minus 25 percent is 48 gigabytes usable, and 48 divided by 0.6 is 80 billion parameters. That is the answer she is fishing for in the comments, and it is deliberately larger than anything she actually runs on the machine, which is the gap discussed above.

The Arduino payoff

The best moment in the home server section has nothing to do with the server's own capabilities. She picks up the Arduino Uno R4, the board that could run nothing at all, and connects it to the Mac Studio over Wi-Fi.

This Arduino Uno is using the Qwen 3.6 35B model on this Mac Studio in order to display encouraging happy messages. Tina Huang, 17:30

A board with 32 kilobytes of RAM driving a 35 billion parameter model, because the model is not on the board. That is the whole argument for a home server in one object: the Studio is where she runs local models for her local AI agents, and it serves other devices too. She asks viewers to write what the Arduino's display says into the comments.

The model is Qwen3.6-35B-A3B, Alibaba's April 2026 mixture of experts release: 35 billion total parameters with about 3 billion active per token, Apache 2.0, and a natural fit for exactly this job, because a sparse model of that shape is far cheaper to serve continuously than a dense 35 billion would be.

The reframe: from new categories to bigger and faster

Here she makes the structural point of the back half of the video, and it is the sentence to keep:

Really with this Mac Studio and all devices that are bigger and more powerful than this, it's not about unlocking new categories of AI models at this point. It's just about being able to run a bigger models, higher quality things to get higher quality responses, and faster. Tina Huang, 17:40

Everything from the 32 gigabyte laptop up is a question of degree. The last genuine category unlock happened two rungs down, when a GPU first appeared and made image generation possible at all.

What the Mac Studio runs, in her list: large tier large language models between 27 and 32 billion parameters, large code models, large image models, and large music generation models. For video, it can now manage mid tier generation, like Wan 2.2, Alibaba's open weights video model (rendered by the captions as "One 2.2"), but her verdict is blunt: it is slow. It is really slow still.

The AMD Ryzen AI Halo: 128 gigabytes, and the first real trap

The 128 gigabyte machine is where the video's thesis pays off properly, because it is the first device that is clearly better on paper and clearly worse at something.

Her numbers: 128 gigabytes of advertised RAM, around 96 gigabytes of usable RAM. (That is her 25 percent rule applied again, and it checks out: 96 is three quarters of 128, which by the formula allows 160 billion parameters.) What that buys:

But the capability she singles out as most useful is not the size of any one model. It is holding several at once:

What I find the most useful of having something like this is the ability of having a lot of different models loaded simultaneously and working simultaneously because it has that massive amount of memory, right? So, I could be using a coding agent, the large language models, and doing like image video generation all at the same time. Tina Huang, 19:00

That is a genuinely different reason to buy memory than "fit a bigger model," and it is the one most people miss. A 96 gigabyte working set lets you keep a coding model, a chat model and a diffusion model resident simultaneously instead of paying the load cost every time you switch tasks.

Then the caveat, delivered immediately rather than buried:

I do want to make a caveat here though. This machine, even though it has very large capacity, it actually has pretty small bandwidth, which means that it is slow, especially when it comes to image, video, and music generation. It can do it. It's just going to do it really, really slow. Tina Huang, 19:22

Back to the kitchen. The Halo has an enormous prep counter and a narrow conveyor belt. You can lay out ingredients for four dishes at once, and then you wait.

GPUs, class five: buying cooks who bring their own counter

The final class is the answer to the Halo's problem, and she re explains the GPU before using one, going back to the kitchen for the third time.

A GPU is a brigade of line cooks, and because there are so many of them they work really quickly and do compute very fast. That matters most for music, image and video generation, because producing those needs a lot of processing power, which means a lot of cooks. The bad news she states plainly: every device covered so far does have a GPU, which is why they can generate images and music and video at all, but they do it really slowly because they are not specialized at it.

The good news is more discrete GPUs, and here she extends the analogy in the single most useful way in the video:

A GPU actually isn't just line cooks. It's actually a bundle. When you get another GPU, you get cooks, but they also come with their own private prep counters and a very wide conveyor belt, the memory bus. So, when you get more GPUs, you get more cooks who have their own private counter and their very own very wide conveyor belt. Tina Huang, 20:25

That one correction explains the whole top of the ladder. A discrete GPU is not a compute upgrade bolted onto your existing memory system, it is a second memory system with its own compute attached, and the belt it brings is wider than the one in your machine. It is why adding a GPU does not just make the cooking faster, it removes the delivery bottleneck too. Then, as she puts it, "it's getting really late," and she moves on.

The constraint on all of this is brutal and she does not soften it: you cannot add more GPUs to any device covered earlier in the video. The chips, the Pi, the phone, the laptops, the Mac Studio, the AMD Halo. None of them are built that way. A gaming PC is, which is the one consumer machine in this conversation that can grow. She does not have one and is not about to buy one, because they are expensive.

Cheating with a rented RTX 4090

So she rents. She is explicit that this is a cheat, and explicit about why she is moving fast:

I'm going to cheat a little bit, okay? Because I do want to show you what it's like to run these multi modality generation models on these specialized GPUs. So, I'm actually going to rent an RTX 4090. This is a pretty high end GPU that you might have in a high end gaming set. Let's do this quickly, because it's costing me by the hour. Tina Huang, 21:10

She logs in, runs an image generation model, and it is really, really quick. Then a video generation, also really, really quick. After four classes of hardware where generation was either impossible or painful, the RTX 4090 does it immediately. That is the demonstration, and it needs no numbers to land.

And then the downside, which is the mirror image of the Halo:

The downside of this GPU is that it only has 24 gigs of memory. So, you can think about it almost like the opposite of the AMD device, because the AMD device has a lot of memory, has very big capacity, but low bandwidth. While the RTX 4090 has smaller capacity, only 24 gigs, but much larger bandwidth. Tina Huang, 22:00

So you can generate anything you like, fast, and you cannot fit a large language model worth talking about. Her summary of the trade: you end up being able to generate a lot of these things and run these models, but you cannot really fit very large size models on it.

The final boss: 8x H100, 640 gigabytes

The top rung is a rented 8 way H100 node: 640 gigabytes of memory, massive capacity and massive bandwidth. She has to move fast on this one, because it is costing her money every minute, which she notes is worth it for the viewers.

What runs on it:

The demo she picks is video, using MiniMax H3, and her claim is that it generates minutes of video, and does it so quickly. (MiniMax's H3, also branded Hailuo 3.0, is a 33 billion parameter omni modal model whose base weights were opened in August 2026. Worth knowing before you try to reproduce this: each generation from the open weights is documented as a clip in the 4 to 15 second range, so "minutes of video" is best read as total output rather than one continuous take. The open weights also max out at a 768 pixel short edge, and the community license excludes self hosting in the United States, the European Union, the United Kingdom and South Korea, which matters a great deal if you were planning to rent a node and follow along.)

Her framing of where that leaves the viewer is careful:

This is pretty much what you can run assuming that you're not like a literal model development company or just have like a lot of money. But, it really is comparable at this point to the frontier models that you're paying for through cloud APIs. Tina Huang, 23:15

And then the closing argument of the entire video, which retroactively organizes every rung below it:

You see everything up to this GPU is kind of just forcing you to choose between like capacity and bandwidth, right? Like the AMD Halo is all about big capacity but low bandwidth, and the RTX 4090 is all about small capacity and large bandwidth. But, this is when you get both. Tina Huang, 23:30

She calls it the aspirational class, puts up a final summary slide, and signs off by asking viewers to answer the on screen quiz in the comments to help retain what was covered.

Capacity or bandwidth: the choice the whole ladder makes for you

Her closing argument deserves to be drawn, because once you see the two axes separately, every result in the video stops being surprising. Capacity is the prep counter and decides what fits. Bandwidth is the belt and decides whether what fits is usable. Almost every machine you can buy is strong on one and weak on the other.

THE TRADE EVERY RUNG MAKES, IN HER OWN TERMS BIG CAPACITY, NARROW BELT BOTH, AT LAST SMALL AND SLOW FAST, BUT SMALL AMD Ryzen AI Halo, 128 GB Mac Studio, 64 GB MacBook Pro, 32 GB Raspberry Pi 5, 8 GB iPhone 16, 8 GB RTX 4090, 24 GB 8x H100, 640 GB low high low high BANDWIDTH: HOW FAST THE MODEL REACHES THE COMPUTE CAPACITY: HOW MUCH MODEL FITS
Figure 4. Positions come from what she says, not from measured figures: she states the Halo's large capacity and small bandwidth, the 4090's small capacity and large bandwidth, the eight H100 node having both, the Pi's narrow bus and missing GPU, and the iPhone's CPU, GPU and NPU on the same 8 gigabytes as the Pi. Amber dots are rented by the hour. The empty top right is the entire point of the video: that corner is the only one that does not ask you to give something up, and nothing on her table reaches it.

The whole ladder, as a reference table

Device and classMemory she statesLargest model that ranSpeed she reportsHer verdict
Arduino Uno R4
Chips
32 KBnothing fitsNo model to runUseless on its own for local AI. Useful over Wi-Fi as a display for a model running on a bigger machine
ESP32-S3
Chips
8 MBTinyStories 260K and TinyStories 3MComfortable at this sizeNext word story generation only. No chat, no multimodal. Two to five dollars without a screen
Raspberry Pi 5
Mini computers
8 GBQwen3 1.7B, chained with Whisper, Piper and MoondreamSlow: narrow bus, CPU only, no GPU or NPUMicro to small LLMs easily, tiny vision and speech models fine. Mid size LLMs struggle. Image, music and video effectively impossible
iPhone 16
Mini computers
8 GB, roughly 4 to 5 GB per app by iOS policyGemma 4, vision models 1 to 4B, Stable Diffusion 1.5 at 2 GB, Whisper Large, YOLO11nFast: CPU plus GPU plus NPUNo 7 to 14B models, by Apple's cap rather than by physics. Every model arrives inside an app
MacBook Pro
Personal computers
32 GB, 24 usable, 40B by her formulaQwen3 32B, Qwen3 Coder 30B, Flux DevWorkable in every categoryThe first rung where no category of model is off the table. Expensive, and she says she owns it only because testing models is her job
Mac Studio
Home servers
64 GB, 48 usable, 80B by her formulaQwen3.6-35B-A3B served to other devices; LLMs 27 to 32B, large code, image and musicFaster than a laptop, but mid tier video is really slowAlways on and serves other devices. No new categories unlock here, just bigger models and faster answers
AMD Ryzen AI Halo
Home servers
128 GB advertised, around 96 usableLlama 3.3 70B fine, gpt-oss-120b fine, DeepSeek Coder V2Big capacity, pretty small bandwidth, so generation is really slowThe real win is many models loaded and working at once, not one bigger model
RTX 4090, rented
GPUs
24 GBImage and video generation, both immediate. No large language model worth the nameVery fast: much larger bandwidthThe exact opposite of the Halo. Cannot be added to any device earlier in the video
8x H100, rented
GPUs
640 GBLLMs above 700B, coding above 480B, vision language, MiniMax H3 videoMassive capacity and massive bandwidthComparable to the frontier models you pay for through cloud APIs. She calls it the aspirational class
Figure 5. The full ladder in her own order, with her own numbers. Read the speed column on its own and the video's thesis appears without any of the prose: the two rows that report no compromise are the two she does not own.

What unlocks where

Model categoryChips, 8 MBMini computers, 8 GBLaptop, 32 GBHome server, 64 to 128 GBDiscrete GPU
Text language modelsTinyStories onlyMicro to smallMid to large, 32BLarge to frontier, 120BUltra large, 700B+
Coding modelsnoBarely, small tierCoder 30BExtra large, Coder V2480B+
Vision languagenoTiny on the Pi, 1 to 4B on the phoneyesyesyes
Speech to textnoSmall on the Pi, Whisper Large on the phoneyesyesyes
Text to speechnoTiny on the Pi, all of them on the phoneyesyesyes
Image generationnoPhone only, 1B classyesLarge, but slowVery fast
Video generationnonoyesMid tier, really slowFast
Music and audio, including voice cloningnonoyesLarge, but slowFast
Object detectionnoYOLO11n on the phonenot namednot namednot named
Figure 6. Every capability claim in the video, arranged by class. Only two columns change what is possible rather than what is fast: the 8 megabyte chip, where a language model becomes possible at all, and the mini computer, where everything except generation arrives. From the laptop rightwards the grid stops filling in and starts speeding up. "Not named" is honest: she lists detection models only on the phone and never returns to them, so the page does not guess on her behalf.

Key takeaways

Chapters

The nine entries in bold are Tina's own chapters, reproduced as written. The rest are sub beats added here from the transcript clock, because nine markers across 24 minutes leaves a lot of ground unmarked.

One oddity in her own chapter titles, left as written above: the GPU chapter is labelled with an "8X W100" alongside the 8X H100, but no such machine appears anywhere in the video. The GPU section covers exactly two rented machines, the RTX 4090 and the 8x H100 node.

Notable quotes

When we try to run an AI model on these, we're not even installing it. We just literally stick it on here and if it fits, it sits and we can run it, or not. Tina Huang, on the chips class, 0:25

Which is big enough to fit, drumroll, please, no models. That's right, nothing fits on here. Alas. Tina Huang, on the Arduino Uno R4's 32 kilobytes, 0:40

I'm able to have full conversations with my Pi by chaining together three different AI models, all on this tiny little device. Tina Huang, on the Raspberry Pi 5 voice assistant, 4:10

Even though it technically has 8 gigs of RAM, you actually aren't able to run the mid size level of models from 7 to 14 B that you can't do using the Raspberry Pi, because like Apple just doesn't let you do it, unfortunately. Tina Huang, on the iPhone 16, 6:15

Basically these machines can only do one thing. It does that one thing really well and really fast and it uses almost no electricity. Tina Huang, on the bolted on kitchen machines that stand in for the NPU, 8:35

It's basically like having a single head chef, very capable, but only a single pair of hands having to do everything to run this model. Tina Huang, on why the Raspberry Pi is slow, 10:50

I only got this machine because it is literally my job to be testing out models and things, right? Tina Huang, on the 32 gigabyte MacBook Pro, 14:40

I actually only got the Mac Studio because I couldn't get a Mac Mini. Tina Huang, on why she owns a Mac Studio, 16:20

Really with this Mac Studio and all devices that are bigger and more powerful than this, it's not about unlocking new categories of AI models at this point. It's just about being able to run a bigger models, higher quality things to get higher quality responses, and faster. Tina Huang, the structural claim of the back half, 17:40

I hope that analogy makes sense. It's getting really late. Tina Huang, after extending the kitchen analogy a third time, 20:50

Now, I really got to move fast on this one, because it's costing me so much money every minute. But, for you guys, it's worth it. Tina Huang, on the rented 8x H100 node, 22:35

You see everything up to this GPU is kind of just forcing you to choose between like capacity and bandwidth, right? But, this is when you get both. Tina Huang, the closing argument, 23:30

Resources mentioned

The hardware, bottom to top

The models

Software, apps and services

The channel

An honest footnote

Four things are worth saying once, at the end, for anyone planning to act on this video rather than just enjoy it.

There are no speed numbers in it. Not one tokens per second figure, not one seconds per image, across all nine machines. Every speed claim is comparative and verbal: really quick, painfully slow, really really slow. That is a legitimate choice for a video whose argument is about which axis binds rather than by how much, and the argument lands without numbers. But if you are trying to decide between two machines in the same class, this video will tell you which axis to think about and will not tell you what to expect.

No quantisation is ever named, and no runtime either. The 0.6 gigabytes per billion parameters in her formula implies roughly four to five bit weights, which is the only quantisation signal in 24 minutes, and it is implicit. She also never names the software actually loading these models on any device, so if you want to reproduce the Pi 5 voice assistant or the 120 billion parameter model on 96 gigabytes, you are supplying the missing layer yourself. The obvious candidates, neither of them mentioned in the video, are llama.cpp and Ollama, which is where the four and five bit quantised weights her formula assumes actually come from.

The top two rungs are rented, and she says so. "I'm going to cheat a little bit" is her own framing, and renting an RTX 4090 and an eight way H100 node by the hour is a different proposition from everything below it, where the hardware sits on her desk and the electricity is the only running cost. The demos are real and the capability claims are sound; the word "local" is doing lighter work at the top of the ladder than at the bottom. The honest reading is that the top two rungs are a preview of what the other seven cannot do, not a recommendation.

One claim is worth checking before you copy it. Her 8x H100 demo generates "minutes of video" with MiniMax H3. The open weights release of H3 is documented as producing clips in the 4 to 15 second range at up to a 768 pixel short edge, so minutes of output means many clips rather than one long take. More practically, the community license on those open weights excludes self hosting in the United States, the European Union, the United Kingdom and South Korea, so reproducing that particular demo on rented hardware in those places is not simply a question of affording the node.

None of that undercuts the video. The lesson it exists to teach, that capacity and bandwidth are separate questions and almost every machine answers only one of them well, is correct, well taught, and the exact thing people get wrong when they spend money on hardware for local AI. The ladder is the proof, and the Arduino driving a 35 billion parameter model over Wi-Fi is the best single image of it: the smallest thing on the table was never limited by what fit inside it.

Full transcript
[00:00:00] Hello. Today we're going to be running every size AI locally with [music] the hardware that we have here and see what we can do with it. Hopefully give you guys some ideas for cool things that you can build with local AI, too. Let's go. A portion of this video is sponsored by Crusoe. So, I'm going to be grouping the hardware that I have here into different classes, different memory classes. So, the first class are chips. So, chips, as they are called, are literally just chips, just computer chips. They have no operating system, no file system, and no [00:00:30] separate memory stick, either. Literally everything is just built into these boards. So, when we try to run an AI model on these, we're not even installing it. We just literally stick it on here and if it fits, it sits and we can run it, or not. And in my case, the smallest one that I have here is an Arduino R4 Uno. It has only 32 kilobytes of RAM. I will put the rest of the stats on screen, which is big enough to fit, drumroll, please, no models. That's right, nothing fits on here. Alas. >> [laughter] >> I actually just got this Arduino Uno R4 [00:01:00] as part of a set cuz I'm trying to learn hardware and it unfortunately does not fit any AI models, but do not be fooled. That does not mean I cannot use this with local AI. I actually can since it has a Wi-Fi functionality, so I can connect with some of my bigger devices. But first, let's actually talk about something that can actually fit an AI model. Isn't this shocking? It is so tiny. This is the ESP32-S3 and it has 8 megabytes of RAM. I will put the rest of the stats on screen and I got this for around $20 USD. Although this one does come with a screen, so if you get it without the screen, you're [00:01:31] looking at like two to five dollars, so cheap. And you can actually run AI models on here, two AI models, actually. Both of them are called Tiny Stories models, which, as its name suggests, allows you to run Tiny Stories. Let me show you what it looks like. So, this one is the Tiny Stories 1 megabyte model and then here is the Tiny Stories 3 megabyte model. You can run both of these very comfortably. So, yes, I can run the Tiny Stories 260K model and the Tiny Stories 3M model very comfortably. And honestly, this is just a start getting a model running. You can make it [00:02:01] so much cooler by adding on different modules to it. For example, you can make it into a little game called scrambled versus not scrambled. Might be useful if you're learning English grammar. And you can attach LED lights that represent the mood of the story that's being produced. So many cool things that you can do. Maybe I will make another video showcasing stuff that you can build using these little devices. But alas, we do not have time in this video. So here is a summary slide for the models that we can run on the ESP32 and some ideas for AI projects that you can build. Take a screenshot. In the free guide in the description, I will also link a little [00:02:31] quick start guide for how to get started building. So these tiny stories models are only able to do the very basics of predicting the next word in order to form a coherent story, which is still very impressive actually given how small it is. And you can build some really cool things on top of that, too. But it is not capable of holding a conversation like a chat. And it's not able to run any multimodality models like audio, music, video, or images. For that, we're going to need to move on to the next class of devices, which is the mini computers category. This is a Raspberry Pi, the Pi 5, and [00:03:03] this is my personal iPhone 16. Now, these are proper computing devices. They run an operating system. The Raspberry Pi runs a Linux, and my iPhone runs the iOS. And they're able to do things like manage memory, file systems, and run multiple programs. And so you can actually have multiple models as long as they fit into these devices. But what is really interesting is that both of these actually have 8 gigs of RAM. Yet, they are not able to run the same kinds of models. I'm going to give you guys a quick lesson on AI hardware in just a little bit. But first, I want to show you what these guys can do. This is the [00:03:34] Raspberry Pi, the Pi 5. And even though it looks very small, it is very mighty because it is a full-blown mini computer. Let me show you what you can do with it. What color is an octopus? >> Octopuses come in a variety of colors, but they are usually camouflage, blending with their surroundings. They can be blue, green, red, or even white, depending on the species. >> First is using a speech-to-text AI model called Whisper in order to listen to it. Then I have the Quiet and Three 1.7B [00:04:04] model. It's a large language model, which is an LLM, to be able to interpret it and then come out with a response. And then the Piper model, which is a tiny text-to-speech model, is able to actually answer it out loud on the speakers. Isn't that crazy? I'm able to have full conversations with my Pi by chaining together three different AI models, all on this tiny little device. But wait, that's not all. I can also connect a webcam that grabs a frame, takes a picture, and I can have a tiny visual language model called Moondream describe it. And the tiny text-to-speech [00:04:34] model Piper is able to describe what it sees. >> Here is the camera, the webcam, attached to the Raspberry Pi. >> The image features a woman sitting in a chair looking at her cell phone. She is wearing a black shirt and is surrounded by a variety of books and other items. There are at least 13 books scattered around the room, some of which are placed on a bookshelf. A backpack can be seen near the woman. A potted plant is located close to her. The scene suggests that the woman might be engaged in reading or studying, as she is [00:05:04] surrounded by numerous books and other items. >> So yeah, you can pretty much run micro-to-small large language models very easily on the Raspberry Pi, as well as the tiny tier visual language models, tiny text-to-speech models, and small but decent speech-to-text models. It does struggle to run mid-size large language models, but if you really want to, you kind of can do it. And it can kind of barely run small code tier models as well. But it really cannot do image generation, music generation, or video generation. Not actually because a small image generation model wouldn't fit on here. It does. The problem is [00:05:35] that it doesn't have any GPUs, so it's just like painfully slow to do that. And I will explain GPUs a little bit later. But it does make up for it because you can build using the Raspberry Pi. You can add so many different types of accessories and sensors to your Raspberry Pi. In fact, I believe the Hugging Face robots are running off a Raspberry Pi as well. So, yes, I'm going to put on a summary slide now, including some ideas of things that you can build with a Raspberry Pi. And please do check out the free guide linked below for a quick start guide. Now, let's talk about the iPhone 16, which also has 8 gigs of [00:06:06] RAM, [music] same as the Raspberry Pi, but it is so different. For one, because it is running on iOS, Apple does limit each app to only be able to use around 4 or 5 gigs of RAM. So, even though it technically has 8 gigs of RAM, you actually aren't able to run the mid-size level of models from 7 to 14 B that you can't do using the Raspberry Pi, because like Apple just doesn't let you do it, unfortunately. Also, because of Apple reasons, you can't just be like putting on things that you wish onto your iPhone. You have to do it through Apple's official system, in which you need to install an app, and the app allows you to download, install, and run [00:06:37] your local AI models. To get the Gemma 4 model, you need to download the AI Edge Gallery from Google, and then download the models locally on your phone. This phone is able to run the small class visual language models, meaning it can analyze images of the 1 to 4 B range very comfortably, since it's only 1.5 gigs. For the image generation, you need to download an app called Draw Things, and you can download and run any small image generation models at around the 1 B range with no problem. The Stable Diffusion 1.5, for example, is 2 gigs. You can run it very comfortably. You can also run medium and large speech text [00:07:08] models like the Whisper Large very comfortably, and all text-to-speech models, too. And finally, you can run a model called detection model. It's actually very cool. People don't really talk about it very much for some reason. You can download an app called Ultralytics, and download and run [music] the YOLO 11 N model. The YOLO 11 N model allows you to do object recognition to understand the scene, and it's the underlying technology behind all things from facial recognition on your devices to self-driving cars, cuz you know, you got to be able to recognize different objects as we're driving. We don't really talk about this class of models very much, these [00:07:39] detection models, but they're actually very, very useful. Isn't that really cool? I'm going to put on screen now a summary slide including some ideas of things that you can build using your phone with local AI. A small lesson on AI hardware. Now, think about a device, any device. It is like a restaurant kitchen. Downstairs, we have a walk-in freezer and a storage space. This is where we store all the ingredients, all of the things. Then we have the prep counter. This is where we load up all the ingredients that we need right now for the dish that we want to make. Now, let's talk about people who work in the [00:08:10] kitchen. We have the head chef. The head chef is very advanced. They can make any sort of dish, but just one head chef. There is also a brigade of line cooks. These line cooks are less experienced, less versatile, but we have a lot of them that can help us out. And finally, we have specialized machines bolted to the walls like say an onion machine that only chops onions. Or like a noodle slicing machine that just automatically sizes noodles. I don't know if you guys have seen those before, [music] but basically these machines can only do one thing. It does that one thing really well and really fast and it uses almost no electricity. Now, between the prep counter and the chefs and the cooks and [00:08:41] the machines, there is a conveyor belt passageway. This is how we get the ingredients from the prep counter to the chefs, cooks, and machines to process them into hopefully a meal. So, say for example, we want to prepare a salad. Well, first got to go get all of the salad stuff from a deep storage in the freezer and put them onto prep counter like your carrots, lettuce, olives, [music] chicken, other salad stuff. Then we put on the conveyor belt and it gets passed through and the head chef starts giving out orders. Tells all the line cooks to be chopping up all the vegetables and also starts up the onion machine to be chopping up onions [music] [00:09:11] while he himself is only say like mixing the sauce because that's like the tricky part. And then after a few minutes, everything comes together and voila, we get a salad. Yay! Amazing. So, this is a great analogy for how AI models work on our devices. So, this time what we want to be cooking is a model. So, we walk down to our freezer downstairs except we call it storage and we grab our model and we put it onto our prep counter which we call the RAM. Then we pass it along the conveyor belt which we call a memory bus, and hand it to our head chef, call it the CPU, who then also [00:09:41] directs the line cooks to start cooking up parts of the model. The line cooks are called the GPU, and also start up the little specialized bolted machines, the onion machine, which we call the NPU. And after all this cooking, your model is able to produce something, like respond to your question of, "Hi, how are you?" The model then can say, "Hello, I am good. How are you?" Okay, make sense? So, every time we want to be working with the model, we need to be going through this entire process of cooking. Great. Now you understand the major components of hardware when it comes to AI models. However, remember [00:10:11] the question that we asked earlier? Why is it that a Raspberry Pi has 8 gigs of RAM and an iPhone has 8 gigs of RAMs, too, but they are so different? Models run on iPhone are much faster, and it's able to do image generation, while the Raspberry Pi cannot. Well, that is [music] because these two devices have the same amount of RAM, which means that the same amount of capacity, but they do not have the same amount of bandwidth, which is referring to the conveyor belt between the prep counter [music] and our cook and our chefs. You see, for most large language models, they're [00:10:41] bottlenecked by how quickly you can deliver them from the RAM, the prep counter, to the chefs and the cooks to and the machines to process the model and get them to be running. On a Raspberry Pi, this conveyor belt, the memory bus, is quite narrow and slow. So, it physically takes longer to be able to be processed. And on top of that, a Raspberry Pi only has a CPU without any GPUs or NPUs. So, it's basically like having a single head chef, very capable, but only a single pair of hands having to do everything to run this model. That's why it is really, really slow. While for the iPhone, it [00:11:11] has CPUs, GPUs, and NPUs. That's why it's able to process things so much faster. So, that explains why it is that things run so slowly on the Pi, but run so quickly on iPhone. But, what about image generation? The Pi simply cannot do image generation, and that again is because of the lack of GPUs. You see, when it comes to image, video, and music generation, [music] it is really processing heavy. So, if you just only have CPUs, no matter how hard they work, it's just not enough to be able to generate music, images, and videos. So, you really got to have some GPUs, like the iPhone, to be able to do that. We [00:11:41] will return to this analogy a little bit later as I explain more devices. But, for now, amazing. You now have an understanding of AI hardware. Yay! I'll put on screen now a little pop quiz. Answer these questions and put them into the comments to make sure that you're paying attention. Of course, what I said about AI hardware is much more simplified than it actually is. There's concepts like VRAM, unified memory, like so much more, which I'm not going to go into too much more detail about in this video, cuz I'm sure this video is already very long. But, this is like the general introduction that will help you a lot in figuring out what kind of [00:12:11] models can run on what kind of devices. So, as you guys can see from this video, I have been a very into tinkering and running open-source models. And the next obvious step is also tinkering and customizing the open-source models themselves. For that, Crusoe Intelligence Foundry has been amazing. It primarily gives you two things: serverless inference to run open models without touching any of the underlying GPU infra. And serverless fine-tuning to customize the open models on your own data without having to own or manage any hardware. You can fine-tune a model, deploy it in one click, and start [00:12:41] calling it with an API in minutes. Seriously, the first time that I did this, I almost shed a tear. There is no cluster supervision or crazy long setups. I cannot even explain how painful this used to be. So, for today, anybody that's building with open models who don't want the infrastructure overhead to eat into actual build time. This is a really clean and easy path to get started. And now, when you sign up, you will get $5 of free credit. You can now try out Crusoe, link is in the description. Thank you so much, Crusoe, for sponsoring this portion of the video. Now, back to the video. Okay, now let's move on to the next category of [00:13:11] devices, which is personal computers, such as my MacBook Pro that I'm using right now. So, my MacBook Pro is 32 gigs, but of course, there are other machines in this category as well, like my older machine, which is a MacBook Air, and of course Windows laptops and Linux laptops, too. They all belong these portable personal computer personal machines category. So, the reason why I got a MacBook Pro with 32 gigs of RAM is because this is the machine that pretty much allows me to run any category of AI models. I can run [00:13:43] large language models, coding models, visual language models for visual analysis, image generation, video generation, text-to-speech, speech-to-text, audio generation, voice cloning, and music generation. >> [music] >> All of this I can run locally. And not only can I run each of these models locally by itself, I'm also able to use these models in different ways by using it with different types software. For example, with my Hermes agent, which I use to run many aspects of my life and [00:14:14] my work, I can run pretty much any mid-size to large-size local model, like the Qwen 3 14B, 16B, 32B, some of my go-tos in the Qwen family. And I can run the Qwen 3 Coder 30B locally also an open-source coding harness like Open Code, and it works pretty good. My favorite for image generation is Flux Dev. I'm going to put a summary sign now of all the different types of models that you can run in this personal computer personal machine category, as well as some ideas of things that you can build with these models. Now, I know what you're thinking already, "But Tina, I don't have a 32 gig MacBook Pro." [00:14:44] Understandable. It is really expensive. I only got this machine because it is literally my job to be testing out models and things, right? Don't worry, I got you. I'm also going to put on screen now the models that you can run if you have 8 gigs of RAM and 16 gigs of RAM. These are the general RAM sizes of laptops that most people have. But, if you do have something that's not like 8 gigs or 16 gigs, literally you have something smaller, bigger, something in the middle, that is okay as well. Because I'm also teach you a formula for calculating roughly the size of model that you can run on your machine giving the amount of RAM that you have. Okay, ready? The equation is that take the [00:15:15] total amount of RAM that your machine says it has, subtract that by 25% of it, because this is used up by like other stuff. And the rest is the actual usable amount of RAM that you have. Divided by 0.6 and that equals the amount of parameters that your model can have. For example, my MacBook Pro says it has 32 gigs of RAM. So I subtract 25% of that, which means I have 24 gigs of usable RAM. Now I do 24 divided by 0.6 to get 40. So my MacBook Pro can fit a model of up to 40 billion parameters. That is a very rough way of figuring it out. Of [00:15:46] course, you can also just ask one of your favorite AI chatbots. Hey, this is what kind of models can I fit on it? You know, that works too. But in any case, I hope that is helpful. I want to now move on to the next category of devices, home servers. The two that I have to demo right now is the Mac Studio and the AMD Ryzen AI Halo. All right, the Mac Studio and the AMD Ryzen. These are examples of the home server class. There are a lot of other machines that fit into this category. I would say the cheapest entry-level here [00:16:17] would be a Mac Mini. Although those are really hard to get these days. I actually only got the Mac Studio because I couldn't get a Mac Mini. Anyways, at this tier, what is the most attractive and why you would get something like this is for two reasons. The first one is that it is a dedicated always-on machine. So unlike your laptop, which you got to carry around and close and stuff like that. And when you close it, your, you know, your model stop working in case you were doing something with it. But the home servers, these machines are always-on. They are designed to be always-on. So they don't close down. You can constantly doing stuff and running your AI models 24/7. That's the first attractive part. And the second reason [00:16:48] is that generally these home server devices do have more capacity and bandwidth. So it can run a bigger models and do it faster. Let's first talk about the Mac Studio here. It has 64 gigs of RAM. Pop quiz. It says it has 64 gigs of RAM. How much of it is actually usable? And what is the largest model in billions of parameters that can fit on my Mac Studio? Were you paying attention? Put it into the comments. Amazing. Hope you got that. My Mac Studio is where I like to run my local models to use for my local AI agents and also allow other devices. Remember this [00:17:18] device, the Arduino Uno R4? That couldn't run anything by itself. I also like to be able to connect this with my Mac Studio, so I can use the models that are being run on my Mac Studio. This Arduino Uno is using the Qwen 3.6 35B model on this Mac Studio in order to display encouraging happy messages. Write into the comments what this says if you can read it. Yay! That's really cool, right? Really with this Mac Studio and all devices that are bigger and more powerful than this, it's not about unlocking new categories of AI [00:17:49] models at this point. It's just about being able to run a bigger models, higher quality things to get higher quality responses, and faster. In the case of the Mac Studio, I can run large tier large language models between a 27 and 32 billion parameters, as well as large code models, large image models, and large music generation models. >> [music] >> For video generation, I can now run mid-tier video generation, like the One [00:18:19] 2.2, but it is slow. It is really slow still. I'm going to put a summary slide now with all the models that you can run with a Mac Studio, as well as some ideas for things we can build. And let's move on to the AMD Halo, which is 128 GB of advertised RAM. Now, with 128 GB of advertised RAM, around 96 GB of usable RAM, it is now able to comfortably run extra large models in the 70B range and frontier models in the 100 to 250B range. We can see that Llama 3.3 70B is [00:18:50] able to run fine, and GPT-OSS 120B is also able to run fine. It's also able to run extra large tier of coding models, like the Deep Seek Coder V2. Great. So, we can see that it can run bigger models now. Makes sense. However, what I find the most useful of having something like this is the ability of having a lot of different models loaded simultaneously and working simultaneously because it has that massive amount of memory, right? So, I could be using a coding agent, the large language models, and doing like image video generation all at the same time. I'm going to put on screen now all the models that you can run with a device like the AMD Halo and [00:19:22] ideas of what you can build. But, I do want to make a caveat here though. This machine, even though it has very large capacity, it actually has pretty small bandwidth, which means that it is slow, especially when it comes to image, video, and music generation. It can do it. It's just going to do it really, [music] really slow. Which is why I want to introduce you to this final category that we're going to cover today, which is GPUs. First, let me explain GPUs a little bit more. Remember that analogy that we had earlier, the restaurant kitchen analogy? We said that a GPU is like having a [00:19:53] brigade of line cooks. Because we have so many of them, they're able to work really, really quickly and they do compute very fast. Now, this matters the most when it comes to music generation, image generation, and video generation. Cuz to be able to produce these, you just need like a lot of processing power. So, you need a lot of line cooks in the kitchen. The bad news is that all the devices that I went through earlier, they do of course have GPUs, which is why they can do things like music generation, image generation, and video generation, but they do it really, really slowly cuz they're not specialized at doing this. But, the good news is that there is a way to be really [00:20:23] good at processing and be able to do really fast and really good image, music, and video generation. And that is by getting more discrete GPUs. You see, a GPU actually isn't just line cooks. It's actually a bundle. When you get another GPU, you get cooks, but they also come with their own private prep counters and a very wide conveyor belt, the memory pass. So, when you get more GPUs, you get more cooks who have their own private counter and their very own very wide conveyor belt, the memory pass. So, they're able to get the ingredients really fast. It's basically [00:20:53] like a boost for the line cooks. So, you're able to process things really, really quickly, cook your models really quickly. I hope that analogy makes sense. It's getting really late. >> [laughter] >> So, there are a few ways to get more GPUs. Unfortunately, for all the devices that we covered earlier, you cannot add more GPUs to them. Just like they're not built that way. But, if you had a gaming PC, for example, that would have a GPU. I I do not have a gaming PC, and I'm not about to go buy one right now, cuz those are expensive. But, if you do have one, good news for you. You have a GPU, and you can actually add more GPUs as well. [00:21:23] So, I'm going to cheat a little bit, okay? Because I do want to show you what it's like to run these multi-modality generation models on these specialized GPUs. So, I'm actually going to rent an RTX 4090. This is a pretty high-end GPU that you might have in a high-end gaming set. Let's do this quickly, because it's costing me by the hour. We're going to log in here, and then I'm going to run an image generation model. You can see that it is really, really quickly. And here is a video generation as well. [00:21:57] Really, really quick. Now, the downside of this GPU is that it only has 24 gigs of memory. So, you can think about it almost like the opposite of the AMD device, because the AMD device has a lot of memory, has very big capacity, but low bandwidth. While the RTX 4090 has smaller capacity, only 24 gigs, but much larger bandwidth. So, in the end, you end up being able to generate a lot of these things and run these models, but you can't really fit very large-size models on it. All right, I'm going to now put a summary slide for the models that you can run on this device, [music] and the things that you can build with it. And that finally leads me to the [00:22:27] final boss, the 8X H100 GPU. >> [music] >> Now, this is top tier, 640 gigs of RAM, massive capacity, and massive bandwidth. Now, I really got to move fast on this one, because it's [music] costing me so much money every minute. But, for you guys, it's worth it. On here, not only can we run ultra-large tier large language models up to 700 billion parameter-plus ultra-large coding models, 480 billion parameter-plus visual language processing as well, and all types of multi-modal generation. [00:22:57] Specifically, let's check out video generation. Using the Minimax H3 model, we can see that it's able to generate minutes of video, and it does this so quickly. >> [music] [music] >> This is pretty much what you can run assuming that you're not like a literal model development company or just have like a lot of money. But, it really is [00:23:28] comparable at this point to the frontier models that you're paying for through cloud APIs. You see everything up to this GPU is kind of just forcing you to choose between like capacity and bandwidth, right? Like the AMD Halo is all about big capacity but low bandwidth, and the RTX 4090 is all about small capacity and large bandwidth. But, this is when you get both. You get very big capacity and very big bandwidth. And now for the 8 H100 rented, I'm going to put a summary slide now of the models that you can run with it and the things that you can build. This is like [00:23:58] aspirational class. Great. Amazing. Wow. Thank you so much for watching until the end of this video. I hope this was insightful, interesting, helpful, and it has inspired you to want to build things with local AI as well. I'm going to put on screen now a little quiz. Please answer these questions in the comments below to help you retain all of the information that we have covered today. Thank you so much, and I will see you guys in the next video or live stream.