youtube.nixfred.com nixfred.com

I'm Obsessed With Local AI. Here's Why

A 39 minute solo case for running open models on hardware you own, rebuilt in full. Isenberg maps local AI into four pieces (the model, Hugging Face as the warehouse, LM Studio and Ollama as the software, the workflow you wrap around it), decodes the vocabulary that keeps non technical founders out (parameters, tokens, context window, quantization, GGUF), then walks the whole Gemma 4 lineup with real sizes and memory figures, three install paths, and a hardware cheat sheet by RAM tier. The back half is the business case: a first workflow on a folder of support tickets, the cheapest useful eval, and three startup ideas specified down to the Customer, the flags, and the go to market.

Published Sep 8, 2026 38:46 video 56 min read Added Oct 8, 2026 Open on YouTube →

At a glance

Greg Isenberg spends 39 minutes making one argument: open models running on hardware you own are about to produce a large number of small, profitable software businesses, and almost nobody outside the developer world has the map yet. He builds the map in four layers (the model, the warehouse you get it from, the software that runs it, and the workflow you wrap around it), decodes the vocabulary that scares non technical founders off (parameters, tokens, context window, quantization, GGUF), walks the Gemma 4 lineup model by model, then gives three concrete installation paths: LM Studio for the first taste, Ollama for wiring a model into your own code, and Google AI Edge with LiteRT-LM for shipping a model inside a real app. The back half is the business half: a hardware cheat sheet by RAM tier, a first workflow you can run tonight on a folder of support tickets, a deliberate argument for workflows before fine tuning, the smallest useful eval, and three startup ideas spelled out to the level of who buys it, what version one does, what it flags, and how he would sell it. Google sponsored the episode, which he states at the top, and Gemma plus AI Edge are the running examples throughout, though the framework is model agnostic by design.

The claim: a 24 month window, and most people do not have the map

He opens cold, before the music, with the thesis:

I think local AI and open models are going to create a ridiculous number of business opportunities over the next 24 months and I don't think most people actually have the map yet.

The diagnosis that follows is about audience, not technology. His viewers have used ChatGPT and Claude. But the moment the words local AI, Hugging Face, Ollama, LM Studio and AI Edge show up in the same sentence, it reads as developer territory, and normal founders conclude they are not supposed to touch it. He thinks that conclusion is a mistake, and that the opportunity behind the vocabulary wall is, in his words, "pretty endless."

So he states the deliverables up front, which is also the shape of the episode: what local AI is, when it matters, how to run open models at work, where Hugging Face fits, which Gemma model he would start with, how he would run a model locally with LM Studio or Ollama, and how all of that turns into real business ideas. Then three startup ideas he would actually consider building, each with the Customer named, the first version specified, the reason local matters, and a sales motion. He calls it "a masterclass around local AI, how to run models, how to build apps, how to make money from it," explained for the average person who is not technical.

The disclosure comes immediately after, at 1:05: Google sponsored the episode, and he thanks them for "caring about local AI and open models for entrepreneurs." He names the consequence of that in the same breath, which is the honest thing to do. Gemma and Google AI Edge are the main examples for the rest of the video. The stated goal is still to hand you a full map "so you can actually understand the space and build with it and use whatever model suits you," and the middle section on competing model families is where he keeps that promise.

Then the show bumper, which is the one piece of pure color in the episode, and which tells you the format: this is a fireside solo episode of The Startup Ideas Podcast, not an interview. "The startup by the fireplace." A beat of music. "It's sipping time, baby."

What local AI actually means, and the only question that matters

The definition is one sentence long and deliberately boring: local AI means the model runs on hardware you control.

Then he enumerates the hardware, and the range is the point. A MacBook. A Windows laptop. An Android phone. An iPhone. A browser. A Raspberry Pi. A workstation sitting in your office. At the top end, he mentions that he just got an NVIDIA DGX Spark, "which is like a high-end one." The DGX Spark is the desktop box built on the GB10 Grace Blackwell superchip, with 128 GB of unified LPDDR5X memory at 273 GB/s and up to one petaFLOP of FP4 compute, which launched at $3,999 and now lists at $4,699. That is the expensive end of the list he just read out, and he returns later to tell you not to buy one yet.

The sentence he actually wants you to keep is the one after the hardware list: "a phone now could run local AI."

Cloud AI, by contrast, means the model runs somewhere else and you reach it through a website or an API. That is the whole technical distinction, and he does not dwell on it, because the interesting version of the question is commercial:

The business question is where should the intelligence live?

His own answer splits cleanly. Deep research, strategy, hard reasoning, anything where he wants the strongest possible model: that is a frontier cloud model. Private files, sensitive Customer data, offline usage, field work, low latency, audio input, or an internal workflow that runs again and again and again: that is where local starts to make a lot of sense. "A smaller model in the right place can actually be very valuable."

And then the reframe that the rest of the episode hangs on. The first question people ask about an open model is whether it is smarter than the biggest cloud model. He says that is the wrong question, and substitutes two better ones:

Is this model good enough for the job, and does running it locally make the product better?

Once you ask it that way, he says, the business opportunities start appearing. That is not a hedge. It is the actual mechanism of the episode: every startup idea in the back half is an answer to those two questions for a specific Customer.

The four pieces: model, warehouse, software, workflow

This is the map, and it is the most reusable thing in the episode. There are four pieces to the local AI landscape, and almost every piece of jargon you will hear belongs to exactly one of them.

1 · THE MODEL The brain file Gemma · Llama · Qwen · DeepSeek · GLM · Mistral · Phi Some reason better, some code better, some are just smaller 2 · THE WAREHOUSE Where you find it Hugging Face: model cards, licenses, file formats, examples, benchmarks, community forks, pre quantized builds 3 · THE SOFTWARE What runs the model LM Studio: a normal desktop app, friendliest first run Ollama: one command, plus a local API on port 11434 4 · THE WORKFLOW The product you build around it A folder of private files in, a reusable artifact out This is the layer that is actually the business WHAT SITS UNDERNEATH llama.cpp Powers a lot of local model inference MLX Matters if you are on Apple silicon Google AI Edge + LiteRT-LM AI Edge is the on-device development world, LiteRT-LM is the runtime for language models Android, iOS, web, desktop, edge
Figure 1. Isenberg's four piece map, rebuilt. Almost every scary word in local AI belongs to exactly one of these four boxes, which is why the map is worth more than any individual tool recommendation. The right hand column is the layer he tells beginners they can safely ignore on day one and will meet the moment they try to ship something.

The model is the brain file

Gemma is a model family. Llama is a model family. Qwen, Mistral and Phi are model families too. The useful mental move is to stop treating them as a leaderboard and start treating them as a catalogue with different shapes: some are better at reasoning, some at coding, some are smaller, some are faster, some are better at images, and some are simply easier to run on your own machine.

Hugging Face is the warehouse

You need somewhere to find these things, and that is Hugging Face. It is the first place he would go, and he calls it the biggest. The best one line explanation he offers is "model warehouse": you go there and you find model cards, licenses, file formats, examples, benchmarks, community versions, and, crucially for anybody with a normal laptop, versions that have already been compressed so they are far easier to run locally.

He drops a valuation aside while introducing it: "I think they're trying to get acquired right now at $13 billion." That number was accurate and then some. Hugging Face was reported in late August 2026 to be fielding offers around $13 billion, and on 3 September 2026, five days before this episode went up, NVIDIA announced a definitive agreement to acquire it for about $12.93 billion, $11.9 billion to shareholders plus roughly $1 billion in retention equity, against a $4.5 billion valuation in 2023. So the warehouse at the centre of his map is on its way to being owned by the company that sells the hardware, which is worth knowing when you are deciding where to anchor a business.

His actual advice about Hugging Face is a reading exercise, and it is the most practical beginner instruction in the episode:

If you're new to local AI, one of the best exercises is actually just to open Hugging Face and read a model card really slowly. You're going to learn a lot.

Then he tells you which half of the card to ignore. At first, skip the scary looking details and answer six questions:

Answer those six and, in his words, the space "gets a lot more intimidating," which is a slip of the tongue he immediately corrects by explaining the opposite: he was overwhelmed by model cards the first time he opened one, and these six questions are what made them readable. He means less intimidating.

The software is what runs it, and two things sit underneath

For most people, start with LM Studio or Ollama. LM Studio feels like a normal desktop app: you download it, you search for a model, you click download, you chat with it. In his opinion it is "probably one of the most friendly first-time user experiences if you're non-technical." Ollama is more builder oriented: you run a command like ollama run gemma4:e4b and you now have a model running locally with an API your apps can talk to.

Underneath both of those sit two names you will start hearing. llama.cpp powers a lot of local model inference. MLX matters if you are on Apple silicon. He does not go deeper than that, and for the audience he is addressing, that is the correct depth: these are the engines, and LM Studio and Ollama are the cars.

And if you are thinking about shipping real on-device apps in the Google ecosystem, that is where Google AI Edge and LiteRT-LM come in. AI Edge is the broader on-device AI development world; LiteRT-LM is the runtime layer for the language models specifically. He frames it as a threshold rather than a tool: "this is what you study when you want to move from I ran a model on my laptop to I want this model inside an iOS app or an Android app or web app, desktop app, whatever it is."

The captions render LiteRT-LM as "light RTLM" and "light RT LM" throughout. The real name is LiteRT-LM, Google's production inference framework for LLMs on edge devices, and it is the successor line to TensorFlow Lite by way of LiteRT. It already ships on-device generative AI in Chrome, Chromebook Plus and Pixel Watch, and it loads models in a .litertlm container, which is the evolution of the older .task format.

Vocab decoder

He stops the tour here to define the six words that do the most gatekeeping, with a promise that by the end of the segment you will know the core basics of local AI vocabulary. The definitions are deliberately plain.

Parameters. The internal weights of the model, the numbers you see quoted as 2 billion or 4 billion. The beginner shortcut: more parameters usually means more capacity, and more capacity helps with harder tasks. The trade is memory, speed and hardware. He then gives the size tiers that the rest of the episode uses:

And immediately a cost warning, because he knows where this impulse goes: "I recommend like not going out there and spending 5, 10, 20 thousand dollars on a workstation just yet." By the end of the episode, he says, you will know how to set this up on your phone, or on a spare laptop you have had since 2021.

Tokens. The chunks of text the model reads and writes. The point he makes about them is economic: locally, you care about speed and memory rather than the per token bill. That is one of the quiet arguments for local AI and he does not belabour it, but it is the same observation that later becomes "high repeated API cost" on his list of hunting grounds.

Context window. How much information the model can work with at once. One sentence, no more.

Quantization. His favourite bit, delivered with the only self deprecating line in the technical half: "The word sounds more technical than it needs to. Honestly, I can barely pronounce it." Quantization is compression for models. It is what lets giant models fit on normal laptops. If the full model is the giant version, the quantized model is the version that actually fits, and you lose a little quality but "suddenly this thing magically runs."

The practical rule he gives is two formats and a default. You will see Q4 and Q8. Q4 is usually easier to run. Q8 keeps more quality but needs more memory. If you are just getting started, Q4 is a reasonable place to begin, so begin there.

GGUF. A common file format for local models that makes inference easier on normal machines. That is all he says about it, and all you need on day one. (The format reference lives on Hugging Face if you want the rest.)

LiteRT-LM. In the Google AI Edge world, this is the model format and runtime path you care about when building on-device apps.

Then he compresses the whole first third of the episode into a six line map, which is worth reading as a unit:

Google Gemma, clearly explained

Now the sponsor's family, and he introduces it as a thing that is "a bit overwhelming" before breaking it down. Gemma is Google's family of open models, and Gemma 4 is the generation built for efficient, local and on-device use.

The model picker, in his order:

Worth adding what he does not say, because it changes how you read that last tier. The two big Gemma 4 models are not two rungs on one ladder, they are different architectures. Gemma 4 31B is a dense model, all 31 billion parameters active on every token, built to bridge server grade quality and local execution. Gemma 4 26B A4B is a Mixture of Experts model that activates only about 4 billion parameters per token, so it behaves far faster than its 26 billion total suggests and is aimed at high throughput reasoning. The E in E2B and E4B stands for effective parameters, which is why those two can be far smaller in practice than their file sizes imply. Every model in the Gemma 4 line also ships with a dedicated draft model for speculative decoding, which is a free speed win with no quality loss. The small models carry a 128K context window and handle text, image and audio; the 12B, 26B A4B and 31B carry 256K.

8 GB RAM 16 GB RAM 32 GB RAM E2B 2B effective 4.6 10 E4B 4B effective 6.6 16 12B 12B dense 8.0 24 26B A4B 26B MoE, 4B active 18 52 31B 31B dense 20 63 q4_K_M, the quantized build bf16, full precision 0 10 20 30 40 50 60 70 download size, GB
Figure 2. What quantization actually buys you, in gigabytes. Sizes are the published download sizes for each tag in Ollama's Gemma 4 library; the dashed lines are the three RAM tiers from Isenberg's own hardware cheat sheet later in the episode. Download size is the floor, not the ceiling: running memory sits higher, and Google's own figures for the full precision models are 11.4 GB for E2B, 17.9 for E4B, 26.7 for 12B, 57.7 for 26B A4B and 69.9 for 31B, before you fill any of that 128K or 256K context window.

The specialist Gemmas most people have never heard of

This is the part of the segment he flags as underrated, and he is right that it is poorly known. Beyond the general purpose picker there is a set of task specific Gemma models:

Then the permission to ignore all of it: "You can leave most of the family alone on day one." The practical path is Gemma 4 E4B, understand the workflow, then move up or down "or sideways actually, depending on what you are building." That word sideways is doing real work. Sideways is where the specialists live, and it is a more common upgrade than up.

The captions mangle a couple of these. "Google for E2B" is Gemma 4 E2B. "Function Gemma" is one word, FunctionGemma, and it is built on a 270M parameter Gemma 3 base specifically as a foundation for fast, private, local agents that turn natural language into executable API calls.

How the Google pieces fit together, and why hybrid wins

His one paragraph summary of the whole Google stack:

AI Edge Gallery deserves the extra sentence he does not give it, because it is the single lowest friction way to do what he describes. It is a free app on Google Play and the App Store, open source under Apache 2.0, that runs Gemma 4, Qwen and Phi models entirely on your phone with no connection, and it reports real time benchmarks for time to first token, decode speed and latency as you use them. If you want the "a phone now could run local AI" claim to stop being abstract, that is the install.

Then the architectural conclusion, which is the most important design claim in the episode: "a lot of big products and serious products are going to use hybrid setup."

His worked example is a local AI tool for a professional services firm, and it is the seed of startup idea three. The local model reads the sensitive drafts and checks for issues. It strips or summarizes the private details and prepares a clean version of the problem. Then, when the Customer wants deeper reasoning, a cloud model works on the sanitized version. A human approves the work before anything important goes out.

That to me feels like a more natural architecture than just putting everything into the cloud, which a lot of people don't want.

Local handles the private files as a first pass, cloud handles the heavy thinking when you need it, a human signs off. "This is how I'm starting to think about building a lot of these products."

Other open model families, with the trade offs stated

This is where he keeps the promise he made after the sponsor read. Six families, each with an upside and a downside, delivered quickly.

Llama, from Meta. He calls it the default open model reference point for a lot of developers, because the ecosystem is big. The upside is the community, the tooling, the examples and the support. The downside is that you still need to read the license and the model card, "especially if you're building a serious commercial product."

Qwen, from Alibaba. He says it has become very strong, specifically around coding, multilingual work, long context and agentic tasks. Then the part most creators skip:

The China thing is real.

And then the distinction that actually matters, which almost nobody draws cleanly. A lot of people use Qwen because it performs really well. But if you are in an enterprise, a government, healthcare, finance, or any sensitive data environment, "you need to separate running open weights locally from sending data to a hosted service." Those are two different risks wearing the same brand name. Downloading open weights and running them on a machine you own sends nothing anywhere. Calling a hosted API with the same model name sends everything. He tells you to check what your company is comfortable with, and what you are comfortable with. The Qwen weights live on Hugging Face if you want to take the first path rather than the second.

DeepSeek. Similar in that it made a lot of people realize how strong Chinese open models really could be, especially for reasoning and coding. Upside: performance and cost, "it's pretty cheap." Trade off: some buyers will have procurement, security or geopolitical concerns. His counsel is to be thoughtful about where you would use it, how you would deploy it, "and even if you want to use it and go down that path."

GLM, which people know as Z.ai. Another one you will see pop up a lot, and he notes he did a whole episode on it. If you spend time on Hugging Face, Ollama "or just local model Twitter," you will see it constantly. The thing to know is that some of these models are really good for specific jobs, so do not ignore them just because they are not the obvious brand name. Test them, read the model card, check the license, play with them. But what you actually deploy for your business may be very different from what you play with. (For the record: Z.ai is the former Zhipu AI, spun out of Tsinghua University, and it publishes most of its GLM weights openly, usually under MIT, which is about as permissive as licenses get.)

Mistral. The European family, based in France, as he correctly recalls. If you care about efficient models and practical developer use cases, they are pretty good: "a strong model with a pretty builder-friendly posture." The downside he names is real and under discussed: "the lineup is a little confusing. Some models are open, some are commercial." That is a license reading problem, not a quality problem, and it is exactly what the model card exercise from earlier is for.

Phi, from Microsoft. Interesting if you care about smaller, faster, lower latency models. And then the bluntest line he gives any vendor in the episode: "for a lot of use cases, I haven't seen it work very well."

He closes the survey by admitting the survey is obsolete the moment he records it. New models show up all the time, "it feels like every other day." Which is the real argument for Hugging Face: you are not going there just to find the big models. You are going to find the weird specialist models, the community fine tunes, the quantized versions, the forks, and the model cards that tell you whether a thing is actually usable for your workflow.

Your job as a founder, or just someone who's playing with these models, is to pick a model family that fits your workflow, that you connect with that company, you like how they do things, and then go from there.

Play with a lot, learn a lot, then pick a family. You do not need to memorise any of it.

Path 1: run Gemma in LM Studio

Three installation paths, in increasing order of commitment. This is the one he recommends to everybody.

Download LM Studio. It is free. Open the app. Search for Gemma 4. If your machine is solid, try E4B; if your machine is slower or older, look for E2B. Then look for the quantized version if you are on the GGUF path, "because you want the model to run just a lot more comfortably."

Once it downloads, open a chat and ask it something simple. And here he makes a deliberate choice about what that something should be, because he wants the value to land immediately rather than abstractly. Not a riddle, not a poem. A business prompt:

Read these customer notes and turn them into a one-page memo about what customers are struggling with, what has changed and what the business should fix this week.

Then paste in some Customer notes. Real ones, or fake ones if you just want to see the shape of it. The point of the exercise is narrow and he says so: the model is now running on your machine, and you are using AI without sending that prompt to a cloud model. "I believe everyone should try that and feel what that is."

Then the step that turns a toy into infrastructure, and it is two clicks deep in the same app. Go to LM Studio's developer section and start the local server. Now other apps can talk to the model on your laptop. "Your computer becomes this little AI server." A script, a prototype, or an internal tool can call the model through localhost and get an answer back.

For the record, that server speaks the OpenAI API dialect on port 1234 by default, exposing the standard chat completions, completions and models endpoints, which is why anything already written against an OpenAI compatible client works against it unmodified. Isenberg does not say that part, but it is the reason this step is a bigger deal than it sounds: you are not learning a new SDK, you are pointing an existing one at your own machine.

I think that's when you start to see how products are going to get built in the modern age.

Path 2: Ollama

Install Ollama. Then:

ollama pull gemma4
ollama run gemma4:e4b

Now you have Gemma running locally from a command line. Ollama also gives you a local API port, and he half remembers the number out loud: "I think it's on 11, 11434." He is right. That is useful because you can connect your own app or script to it.

If you want to test a larger model later, you can try the 12 billion, 26 billion or 31 billion versions, "assuming your hardware can handle it." And the advice on how to find out whether it can is funny and also correct: "you can ask an LLM if your hardware can handle it, or you can do yourself and just suffer through the slowness and the pain of it."

The tag names in the Ollama Gemma 4 library are exactly as he says them: e2b, e4b, 12b, 26b (aliased 26b-a4b) and 31b, with gemma4:e4b as the default latest. There are also MLX builds of each tag for Apple silicon, which connects straight back to the vocabulary section.

Path 3: Google AI Edge and LiteRT-LM

The third path has a gate on it, and he states the gate first: "I would only use this path if I wanted to build an actual app and a model inside of it."

The examples: a mobile app where the model runs on the phone. A browser app where the model runs locally. A desktop app with a private workflow. Something on an edge device. LiteRT-LM is designed for that world, across Android, iOS, web, desktop and edge environments.

That is the path from local AI as a demo to local AI as a product.

That sentence is the right way to think about the three paths as a set. LM Studio is for feeling it. Ollama is for wiring it into something you are writing. AI Edge is for putting it in the hands of a Customer who never sees the model at all.

The hardware cheat sheet

Four RAM tiers and a phone. He runs through them fast, so here they are with the Gemma 4 build that actually fits each one, taken from the published Ollama download sizes.

Your hardwareWhat Isenberg says to doA Gemma 4 that fitsReality check
8 GB of RAM"Start small and keep the first test simple."gemma4:e2b, 4.6 GBfits for short prompts. Do not expect to fill the 128K context window.
16 GB of RAM"You can do some useful experiments with models like E4B and smaller quantized models."gemma4:e4b, 6.6 GBfits comfortably, and E4B is the model he names as the default starting point.
32 GB of RAM"You have way more room to do larger local workflows."gemma4:12b, 8.0 GBfits easily. The 26B Mixture of Experts build at 18 GB also fits, and runs like a much smaller model.
A strong GPU or a workstation"The bigger models become just much more realistic." The one he names is the DGX Spark he just bought.gemma4:31b at 20 GB, or gemma4:26b-a4b at 18 GBfits quantized. Full precision 31B wants about 63 GB of weights and 69.9 GB of working memory, which is why a 128 GB unified memory box exists.
A phone"Think a lot less about model size and more about the job."E2B or E4B in .litertlm form, easiest through AI Edge Gallerywrong axis. The app manages the memory. Your job is picking a task small enough to be worth doing on a phone.
Figure 3. Isenberg's hardware cheat sheet, with a real Gemma 4 build attached to each tier. Download sizes are the published q4_K_M tags in Ollama's Gemma 4 library; the full precision working memory figures come from Google's Gemma 4 model overview.

The phone questions are the useful part of that last row, and they double as a product checklist:

That last one is the only question on the list that is about product design rather than capability, and it is the one that separates a local model from a chatbot. If the work happens before the user asks, there is no prompt box, and the model has stopped being a feature and started being plumbing.

The first workflow to build

He builds it out loud, as a thought experiment you can run tonight. The whole thing is four steps.

One. Make a folder on your desktop called customer notes.

Two. Put 10 support tickets from a specific business in it. His examples of the business: a home health agency, a med spa, a water damage restoration company. His examples of the tickets, which are worth quoting because they are exactly the texture of real ones:

Three. Run a local model like Gemma against the folder and ask it to create a file called What customers are telling us.md. The output should include five things, and the specificity of this list is what makes the exercise work:

Four. Read it. Then do it again.

Why this is the right first workflow, in his framing: it is useful and it is simple. "You have this private messy data, the model runs next to it, and the output is a memo someone can actually use." Nothing left the building, and the artifact is something an operator would act on.

And then the generalisation, which is the actual payload of the segment. Once you have done it once, the shape shows up everywhere:

That last one he will spend three minutes on later as a whole company.

Workflows before fine tuning

This is the most useful piece of discipline in the episode, and he is explicit that he is correcting an instinct he had himself.

People hear "open model" and immediately want to train their own. He gets why. "It sounds really cool, but I feel like that's like an advanced move." The beginner move, the practical move, is to find a repeated workflow first.

The loop he prescribes:

  1. Pick one folder, one model, one output
  2. Run it 10 times
  3. See where it gets confused
  4. Improve the prompt, add examples
  5. Add a checklist
  6. Then build a small eval

The smallest useful eval

He defines the word in one line, because he knows it is another piece of gatekeeping vocabulary: "An eval is just a test that tells you whether the model did the job well enough."

And then he gives the cheapest possible instance of one, which costs you nothing but an afternoon. Take the same 10 Customer notes. Run them through Gemma locally. Run them through a strong frontier cloud model. Compare the outputs, and ask four questions:

The comparison actually teaches you where local's already useful and where you still want that stronger cloud model.

That is a genuinely good suggestion and it is the honest counterweight to everything else in the episode. He is not asking you to take the local model's quality on faith. He is telling you to measure it against the best thing available, on your own data, before you build a business on it. The frontier model is the ruler, not the competitor.

Local, cloud, or both

The decision rule, stated plainly:

AxisRun it locallySend it to the cloudThe hybrid he argues for
The kind of workRepetitive, high volume, device native, the same job every dayDeep reasoning, strategy, broad research, hard one off problemsLocal does the pass that repeats, cloud does the pass that thinks
The dataPrivate files, sensitive Customer data, nothing leaves the machineWhatever you are willing to transmitLocal reads the sensitive original, strips or summarizes it, and hands a sanitized version upward
Latency and connectivityLow latency, works offline, works in the fieldNeeds a connection, pays round trip costThe offline path still works when the escalation path cannot
Cost shapeHardware once, then free at the margin. No per token billPer token, forever, scaling with volumeThe high volume pass is the free one, so only the rare pass is metered
Capability ceilinglower, and he never pretends otherwise. Good enough for the job is the barhighest available, and it changes answer quality on hard problemsYou get the ceiling only on the work that needs it
Who it reassuresThe buyer who "feels better when the model is just close to them"The buyer who wants the best known answerBoth, which is why he thinks serious products land here
Where the human sitsReviewing an artifact the model producedReviewing an answer the model producedApproving, explicitly, before anything important goes out
Figure 4. The decision ledger, built on the axes Isenberg actually argues rather than a generic local versus cloud table. The third column is the one he says serious products will land in, and the professional services example he gives earlier in the episode is exactly that row read top to bottom.

The filter: what makes a local AI business

Before the ideas, the screen he runs them through. The kind of business he is looking for is "niche, useful, cash-flowing businesses that you don't need to raise venture for, and tied to a painful workflow." That is a specific thesis, and it rules out most of what gets called an AI startup.

The filter itself is five conditions, and he wants them stacked:

  1. A Customer with sensitive data
  2. Repeated review work
  3. Bad software, usually
  4. Expensive mistakes, the kind that cost real money
  5. A workflow that happens close to the device

"That combination is like the interesting zone for me."

Then the invitation, which he means literally: "I want you to steal these ideas, and at the very least it'll get your creative juices flowing with how you can use local AI to run model, build apps, and make money."

Startup idea 1: a local QA reviewer for home health agencies

The Customer. Home health agencies employ nurses and caregivers who go into people's homes. They write visit notes, update care plans, and deal with billing and compliance. The paperwork is a pain and it takes a lot of time, "if you've ever witnessed it in person, but it matters so so much."

Why it matters so much. Three specific failure modes, which is how you know he has actually looked at this workflow:

Version one. A local desktop app for the agency. The agency drops in visit notes, care plans and dictated transcripts. The model reviews them before submission and flags problems. His three example flags, which read like they came off a real chart review:

What the buyer is actually buying. Not AI. Fewer documentation problems before the billing run, the audit, or the supervisor review. "So, if you solve that, you have their attention."

How he would grow it, and this is the part that lifts the segment above a list of ideas:

I would actually start it as a service.

Find five small home health agencies. Offer to review a batch of notes. Do the review with local AI helping behind the scenes, but inspect everything manually, with human beings, himself first. Write down the 20 issues that keep showing up. Those 20 issues become the checklist. The checklist eventually becomes the product.

That is a services to software path, and it solves the hardest problem in vertical AI, which is that you do not yet know what to flag. You cannot write the checklist from the outside. You buy it, one batch of notes at a time, by doing the work by hand until the pattern is obvious. The wedge, in his words, is "catching documentation problems before they cost the agency time or money, and then you build from there."

I love this business and totally would start it.

Startup idea 2: an offline field report copilot for restoration contractors

The Customer. Restoration contractors. Water damage, fire damage, mold remediation. He has first hand knowledge here, and says so: "I unfortunately had this, so I know a little bit about it."

The workflow. Those teams are in the field taking photos, recording notes, documenting damage, and creating reports for homeowners and insurance adjusters. The job is visual, it is physical, and it happens away from a desk. And the report is not paperwork, it is the hinge of the whole job: it is the handoff between the technician, the Customer, the office and the insurance process.

Version one. A mobile app. A technician walks through the property, takes photos, records voice notes, and the app drafts the report before they leave the site. The genuinely clever part is what it does with the time while the technician is still standing in the house: it flags what is missing, while that is still fixable.

He singles out that last one: "the last part of that is underrated. In a stressful home damage situation, clear communication is part of the product."

He is right, and it is the strongest product instinct in the episode. Two of those flags are compliance. One is completeness. The fourth is empathy, delivered as a rewrite, to someone whose house is wet and who has never read an insurance scope of work before. That is the thing a human reviewer in an office never gets to do, because by the time they read the report the conversation is over.

Why local matters here specifically, and he does not spell this out, so it is worth making explicit: a technician in a flooded basement is the canonical no signal environment. Every item on his earlier local AI list is present at once. Camera input, audio input, offline operation, low latency, a repeated workflow, and photographs of the inside of someone's home.

How he would grow it. Pick one niche first; do not go after everything. Say water damage restoration. Talk to owner operators. Look at their current report templates. Study the software they already use, "which is some old stack." Build around the checklist that is already in their head. And the demo writes itself: "send me three old jobs and I'll show you how fast your techs could create reports."

Where it expands. QA, estimates, insurance packets, Customer updates, and training new technicians. But start with the field report "because it's specific and obviously super annoying."

The anecdote that produced the thesis. He recently had water damage in his apartment, saw the software the contractors were using, and found it "antiquated. Like, it's stuff from like the early 2000s." Which is where the headline window comes from:

I think there's a 24-month window and opportunity to do some of these products.

Startup idea 3: a local pre-send reviewer for professional services

The Customer. Every professional services firm, "or 99.9% of them," runs a version of the same loop. Someone writes a client email, a proposal, a memo, a contract summary, an investment note, an HR note. Then they ask someone else to check it before it goes out. It happens constantly, at law firms, accounting firms, wealth advisors, recruiting firms, and consultancies.

Version one. A local desktop app that reviews outbound drafts before they leave the company. The flags are vertical specific, and the list is the best part of the segment because each one is a real liability in that profession:

The pitch, in his words. "The product is basically a second set of eyes for sensitive work. It's basically schmuck insurance is the way I think about it." He then floats schmuckinsurance.com as the company name on the spot and asks the audience to tell him whether it is taken, which is as close as this episode gets to a bit.

How he would grow it. One vertical, one document type. Email review for independent wealth advisors. Not everyone, and "probably not the big banks to start." Interview 10 advisors and ask them which emails make them nervous. Collect anonymized examples. Turn their real concerns into a review checklist. Build a local tool that checks drafts against that checklist.

Why he thinks it sells itself. The buyer already understands the behavior, because they already ask someone to check the draft. You are not selling a new habit. "You're just basically giving them a faster first pass that lives closer to their client data and internal rules."

I love this idea and hope a few of you take it.

1 · Home health QA reviewer2 · Offline field report copilot3 · Pre-send reviewer
CustomerSmall home health agenciesRestoration contractors: water, fire, moldProfessional services firms, starting with independent wealth advisors
The painful workflowVisit notes, care plans and billing documentation reviewed before submissionOn site damage documentation turned into a report for homeowners and insurance adjustersSomeone checking an outbound draft before it leaves the firm
Version oneLocal desktop app. Drop in notes, care plans, dictated transcriptsMobile app. Photos and voice notes in, draft report out before the tech leaves the siteLocal desktop app that reviews outbound drafts
What it flagsDizziness mentioned but vitals missing; unclear medication follow up; note that may not support the billed service levelBasement mentioned but unphotographed; ceiling damage without moisture readings; explanations too technical for the homeownerImplied guaranteed returns; over definitive legal language; employee data in the wrong thread; promises the scope does not support; numbers that disagree with the attachment
Why localPatient records, and a desktop the agency already controlsNo signal on site, camera and audio input, latency matters while the tech is still in the houseClient data and internal rules, reviewed without transmitting the draft
His go to marketStart as a service. Five agencies, review batches by hand, let the recurring 20 issues become the checklist, then the productOne niche. Study their templates and their old software, build around the checklist already in the owner's head, demo on three old jobsOne vertical, one document type. Interview 10 advisors on which emails make them nervous, turn that into the checklist
Where it expandsFrom documentation QA outward into the agency's billing and audit prepQA, estimates, insurance packets, Customer updates, technician trainingOther document types, then other verticals
Figure 5. The three businesses side by side. Every cell is his, not inferred: the Customer, the flags, and the go to market are all stated in the episode. Read the "his go to market" row across and you get the same move three times, which is to buy the checklist by hand before you try to sell it as software.

Build your own local AI lab

Then a turn that makes the episode better than a list of business ideas:

By the way, if you're not building one of these ideas tomorrow, I still think you should learn local AI because it does change how you work with your own files.

From a personal productivity angle, he says, it is still "super super helpful." And his stated reason for wanting you to be more productive is the most human line in the episode: so you have more time "to scroll TikTok or watch movies or hang with your family."

The exercise is the Customer notes workflow pointed at yourself. Make a folder called local AI lab. Put 10 files that matter to your work in it. His examples: sales calls, meeting transcripts, old tweets, ideas you have had. Then run Gemma, or whatever model you chose, and make it produce one useful artifact. Four prompts he offers:

And then the criterion that separates this from using a chatbot, which is the single most transferable idea in the back half of the video:

The key, basically, is to produce a file, a memo, a checklist, a brief, a report, or a review that you can reuse. A chat answer is nice, but a useful artifact changes that workflow.

The loop he recommends as your first repetition: a model reads the folder, the model writes the file, you inspect it, you improve the workflow, then you run it again. (The captions render this as "the first wrap I would recommend"; he means the first rep.)

Do that a few times and, he says, your brain starts connecting the dots. Specifically, you start noticing three things, and they map one to one onto the filter he gave before the startup ideas:

Closing thoughts: the hunting ground

Once you see the pattern, you start spotting local AI businesses everywhere.

And the permission slip that goes with it: "You can learn enough of the map to spot where these models belong without turning yourself into a local engineer overnight."

His summary position on the architecture question: some AI belongs in the cloud, some belongs on the device, "and a lot of the best products of the next couple years are going to combine them both."

If he were starting today, in order:

  1. Run Gemma locally and read model cards on Hugging Face
  2. Learn the difference between LM Studio and Ollama
  3. Play with Google AI Edge
  4. Then look for one boring workflow where local AI actually makes the product better

That word boring is load bearing. Nothing in this episode is pointed at a frontier capability. Every idea in it is pointed at a form someone fills in badly.

The hunting ground, as he lists it:

Local AI is just way easier to understand once you stop treating it like a model benchmark conversation and start treating it like a product conversation.

And the five questions he leaves you with, which are the whole episode compressed into something you can run on any business you walk into:

"And then you answer those questions and you just start seeing the idea."

He signs off on a personal note rather than a pitch. He does not see many non technical people playing with local AI, and over the last two months or so he has gone deeper and deeper into it: "it's been connecting the dots and I'm grateful for it." Then: he reads every comment on YouTube and responds to most, share it with a friend who would benefit from understanding local AI clearly, and "happy building."

Key takeaways

Chapters

One note on that list. These are the creator's own chapter boundaries, and every timestamp is a real section break in the video, but the published titles sit one slot early from the vocabulary segment through the first workflow, so the chapter named "Vocab Decoder" on YouTube actually opens the four pieces segment and so on down the line. The labels above are moved back onto the boundaries where their content actually starts, checked against the caption track, with one extra entry at 24:35 for the fine tuning argument that the published list leaves unnamed.

Notable quotes

I think local AI and open models are going to create a ridiculous number of business opportunities over the next 24 months and I don't think most people actually have the map yet. Greg Isenberg, the opening line, 0:00

It's sipping time, baby. Greg Isenberg, the show bumper, 1:32

The business question is where should the intelligence live? Greg Isenberg, 2:11

Is this model good enough for the job, and does running it locally make the product better? Greg Isenberg, on the two questions that replace the benchmark question, 2:59

If you're new to local AI, one of the best exercises is actually just to open Hugging Face and read a model card really slowly. Greg Isenberg, 4:40

The word sounds more technical than it needs to. Honestly, I can barely pronounce it. Greg Isenberg, on quantization, 8:43

It is part of the path from the model gave me an answer to the model help the product do the next step. Greg Isenberg, on FunctionGemma and structured tool calling, 11:53

That to me feels like a more natural architecture than just putting everything into the cloud, which a lot of people don't want. Greg Isenberg, on the local first pass and cloud escalation pattern, 13:50

The China thing is real. Greg Isenberg, on Qwen, 15:05

Your computer becomes this little AI server. Greg Isenberg, on starting LM Studio's local server, 19:55

You can ask an LLM if your hardware can handle it, or you can do yourself and just suffer through the slowness and the pain of it. Greg Isenberg, on trying a bigger model in Ollama, 20:53

That is the path from local AI as a demo to local AI as a product. Greg Isenberg, on Google AI Edge and LiteRT-LM, 21:30

An eval is just a test that tells you whether the model did the job well enough. Greg Isenberg, 25:08

I would actually start it as a service. Greg Isenberg, on how he would launch the home health QA reviewer, 28:47

In a stressful home damage situation, clear communication is part of the product. Greg Isenberg, on rewriting a technician's report for the homeowner, 30:45

It's basically schmuck insurance is the way I think about it. Greg Isenberg, on the pre-send reviewer, 33:33

A chat answer is nice, but a useful artifact changes that workflow. Greg Isenberg, 35:45

Local AI is just way easier to understand once you stop treating it like a model benchmark conversation and start treating it like a product conversation. Greg Isenberg, 37:34

Resources mentioned

Models and model families

Where you get models

Software that runs them

Hardware

The creator

Where this lands

The framework is the durable part of this episode, and it holds up. Four pieces, six vocabulary words, three install paths, and a filter for finding the Customer. Anybody could run that on a business they already know and come out with a real idea. A few things are worth adding before you act on it.

It is a sponsored episode, and it says so. Google paid for it, Gemma and AI Edge are the running examples, and he discloses that at 1:05. The survey of competing families in the middle is where he earns the "full map" claim, and it is not a soft survey: he tells you to read Meta's license carefully, that Mistral's lineup is confusing, and that he has not seen Phi work well. But the specific recommendation at every decision point in the episode is the sponsor's model, and that is worth holding in mind.

The Hugging Face line was already out of date when it aired. He says Hugging Face is "trying to get acquired right now at $13 billion." Five days before this went up, NVIDIA announced a definitive agreement to buy it for about $12.93 billion. That is a bigger fact than a correction. The neutral warehouse at the center of his map is becoming part of the company that sells the accelerators, and if you are building a business whose supply chain runs through it, that is a dependency worth watching rather than a trivia update.

The memory tiers are optimistic once context enters the picture. His RAM cheat sheet is about loading the weights, and the numbers work: gemma4:e4b is a 6.6 GB download, which is fine on a 16 GB machine. But the reason you would want a 128K or 256K context window is to throw a whole folder at the model, and the key and value cache for a long context sits on top of the weights, not inside them. A workflow that reads 10 support tickets is comfortable. A workflow that reads 300 PDFs is a different hardware question than the one the cheat sheet answers.

Two of the three ideas are regulated, and he never says the words. A home health documentation reviewer handles protected health information, which in the United States means HIPAA, business associate agreements, audit logging, and a security review before an agency can legally hand you a batch of notes to review by hand. A pre-send reviewer for wealth advisors sits inside SEC and FINRA recordkeeping and advertising rules, which is precisely why "sounds like a guaranteed return" is the flag he reaches for first. Running the model locally genuinely removes one category of risk, which is the vendor data transmission question, and that is a real and underrated advantage. It does not remove the compliance program. The services first go to market he recommends is actually the right shape for this, because it forces you to solve the paperwork on five accounts before you try to solve it on five hundred, but budget for it.

The 24 month window is a claim, not a forecast. He offers one piece of evidence for it, which is that the incumbent restoration software he saw after his own flood looked like it was from the early 2000s. That is a genuine observation and a fair signal about a specific vertical. It is not a timetable. The useful version of his thesis does not depend on the window being 24 months: bad software in a document heavy, privacy sensitive trade is an opportunity whether or not it closes on schedule.

The eval advice is the part to take most seriously. He tells you to measure your local model against a frontier model on your own 10 files before you believe anything. That is the single most protective instruction in the episode, it costs almost nothing, and it is the step people skip. Do that first and the rest of the map tells you where to go next.

A couple of caption notes, since the automatic transcript mangles names the way it always does. "Light RTLM" and "light RT LM" are LiteRT-LM. "Google for E2B" is Gemma 4 E2B. "Function Gemma" is one word, FunctionGemma. And when he says answering the six model card questions makes the space "a lot more intimidating," the sentence immediately after it makes clear he means the opposite.

Full transcript
======================================== I think local AI and open models are going to create a ridiculous number of business opportunities over the next 24 months and I don't think most people actually have the map yet. >> [music] >> They've used ChatGPT, they've used Claude, but when they hear local AI, Hugging Face, Ollama, LM Studio, AI Edge, it sounds like it's for this developer world and that normal founders are just not supposed to touch it. And I think that's a mistake because the opportunity here is actually pretty endless. By the end of today's episode, you're going to understand what local AI is, when it matters, [music] how to run open models at work, where Hugging Face fits in here, which Gemma model I'd start with, how I'd run a model locally with LM Studio or Ollama, and how this turns into real business ideas. And I'll give you three startup ideas I'd actually consider building using local AI, including who the customer is, what the first version does, why local matters, and how I'd sell it. Basically, this is going to be a masterclass around local AI, how to run models, how to build apps, how to make money from it, and I'm going to explain it for the average person who isn't technical. Quick shoutout to Google for sponsoring today's episode and for caring about local AI and open models for entrepreneurs. Today's episode, I'm going to use Gemma and Google AI Edge as the main examples, but the goal is to give you a full map so you can actually understand the space and build with it and use whatever model suits you. Okay, let's dive in. >> The startup by the fireplace. >> It's sipping time, baby. >> So, put simply, local AI means the model runs on hardware you control. The hardware could be your MacBook, it could be your Windows laptop, an Android phone, an iPhone. It could be a browser, Raspberry Pi. It could be in a workstation in your office. I just got a DGX Spark, which is like a high-end one. But, the the important part to note is a phone now could run local AI. Cloud AI means the model runs somewhere else and you access it through a website or an API and that's the basic difference. The business question is where should the intelligence live? If I'm doing deep research and strategy and hard reasoning or something where I want the strongest possible model, I'm probably going to be using a frontier cloud model. If the work involves private files, like sensitive customer data, offline usage, field work, low latency, audio input, or an internal workflow that runs again and again and again, local AI starts to make a lot of sense. A smaller model in the right place can actually be very valuable. That is the idea I want you to keep in your head. The first question most people ask is is this model smarter than the biggest model in the cloud? The actual more useful question to ask actually is is this model good enough for the job and does running it locally make the product better? Once you ask it that way, you start seeing these business opportunities which we'll go into. So, there's four pieces to the local AI landscape. The model, which is the brain file, that could be something like Gemma, Llama, or Mistral. The warehouse, which is where you find the model, you might have heard of Hugging Face. I think they're trying to get acquired right now at $13 billion. Um that's what they do. The software, which is what runs the model, that's something like LM Studio or Ollama. And then the workflow, which is the product you're building around all of it. And those are the real four pieces. Uh the model is the brain file. Gemma's a model family. Llama's a model family. Qwen or Mistral, you might have heard of Phi. These are model families, too. Some of these are actually better at reasoning, and some of them are better at coding, some of them are smaller, some of them are faster, some of them are better for images, some are easier to run on your own machine. And then, you need somewhere to find these models. That's what Hugging Face is. That's the first place I would go. They're the biggest uh at it. The easiest way to explain Hugging Face is that it's a model warehouse. You go there, and you can find model cards, licenses, file formats, examples, benchmarks, community versions, and sometimes versions that have already been compressed, so they're way easier to run locally. If you're new to local AI, one of the best exercises is actually just to open Hugging Face and read a model card really slowly. You're going to learn a lot. At first, though, you can ignore half the scary-looking details and just look for a few basic things, in my opinion. What is the model for? How big is it? What license does it use? What hardware are people running it on? Does it support text, images, audio, tool use, or embeddings? Are there quantized files available? Once you can answer those questions, the space gets a lot more intimidating. Um cuz I know when I first looked at these cards uh initially, I was like overwhelmed. So, just those are the key questions to ask. Then, you need software that runs the model. For most people, I would just start with LM Studio or Ollama. LM Studio feels like a normal desktop app. You download it, you search for the model, you click download, and then you can just chat with it. Um my opinion is it's probably one of the most friendly first-time user experiences if you're non-technical. Ollama is a little more builder-oriented or developer-oriented. You To it, you run a command like Ollama run Gemma 4 colon E4B. And now you have a model running locally with an API your apps could talk to. Then underneath those tools, uh you're going to start hearing about things like llama.cpp and MLX. And I'll explain what those two things are. powers a lot of the local model inference. MLX matters if you're on Apple silicon. And if you're thinking about shipping real on-device apps in the Google ecosystem, that's where Google AI Edge and Light RTLM come in. Basically, Google AI Edge is the broader on-device AI development world, and Light RTLM is the runtime layer for the language models. This is what you study when you want to move from I ran a model on my laptop to I want this model inside an iOS app or an Android app or web app, desktop app, whatever it is. We got to talk about some key vocabulary just about the most important things you need to know about these words that uh that come up time and time again in local AI. I'm just going to give you simple, clear definitions of what they are. By the end of this part, you you'll know, you know, just the core basics of of local AI vocab. So, I'm sure you've heard this one before of uh parameters, like 2 billion, 4 billion. These are what's called the internal weights of the model, and more parameters just usually means more capacity for harder tasks, but it does require more memory. So, parameters are the internal weights of the model. The beginner shortcut is that more parameters usually means more capacity, and more capacity can help with the harder tasks. So, the trade-off is usually memory, speed, uh and hardware. So, a 2 billion or 4 billion model is the kind of thing you might use for edge devices, phones, fast workflows, and smaller tasks. A 12-billion uh parameter model is more of a middle ground, and a 26-or-31-billion model is getting to the stronger uh workstation territory. Depending on your hardware and how the model is built, um I I recommend like not going out there and spending 5, 10, 20 thousand dollars on a workstation just yet. Uh by the end of this episode, you're going to understand how to just, you know, set up some of these things on your phone or on a laptop, a spare laptop that you have from 2021. Then, there are tokens. So, tokens are the chunks of text that the model reads and writes. Locally, you care about speed and memory rather than the per-token bill. Then, there is the context window. The context window is basically how much information the model can work with at once. Then, there is quantization. The word sounds more technical than it needs to. Honestly, I can barely pronounce it. Quantization is the compression for models. It allows giant models to fit on normal laptops. For example, you might have heard of Q4, Q8 formats. That's quantize quantization. If the full model is the giant version, the quantized model is the version that can actually fit on a normal laptop. So, you might lose a little quality, but suddenly this thing magically runs. You will see things like Q4 or Q8. And as a beginner rule, Q4 is just usually easier to run, and Q8 keeps more quality, but it needs more memory. If you're just getting started, Q4 is just a reasonable place to begin, so I would start there. Then, there is GGUF. It's a common file format for local models that make inference easier on normal machines like you and I have. And in the Google AI Edge world, you'll see something called the light RTLM. This is the model format and runtime path you care about when building on-device apps with light RTM. So, the simple map is this: Hugging Face helps you find and understand models. Gemma is Google's open model family, and Google's a trusted brand. Uh I run my business on top of Google, so it just makes sense. LM Studio helps you try models locally without much friction. Ollama helps you run models locally in a way that a bit more technical people can plug into apps. GGUF is a common local model format, and Google AI Edge and light RTLM are the path toward shipping on-device AI products. That's what you need to know. So, let's talk about Google's open model family, because I feel like there's a lot here. It's a bit overwhelming, and I'm just going to break it down so you understand what you need to know about the whole Google AI open model family. So, Gemma is Google's family of open models, and Gemma 4 is built for the efficient, local, and on-device use. So, they have Google for E2B, which is the smaller edge model for phone workflows. You have a bigger uh E4B, Gemma 4 E4B, which is the It's pretty much the most practical starting point for most local tasks. Then you have Gemma 4 12B, which is a middle ground with more capability for laptops. And then you have Gemma 4 26B/31B, which is, you know, the stronger local workstation territory. That is the main model picker. Then, and a lot of people don't know this, there's specialized Gemma models that are just really useful to know. So, you have things like embedding Gemma, which is just for search. So, specifically, it helps you turn text into embeddings, which lets you search by meaning. If [snorts] you want to search your own docs or customer notes or support tickets, sales calls, or knowledge base locales, embeddings matter a lot. Then they have something called function Gemma, and that's a tool use in structured function calling. That means the model can help software take actions in a way more structured way. It is part of the path from the model gave me an answer to the model help the product do the next step. Then you have a few more like Pali Gemma, which is more vision focused. You have Shield Gemma, which is more safety focused. Then you have Gemma scope, which is more understanding how models work under the hood. You can leave most of the family alone on day one. The practical path like on day one, if you're a beginner, start with Gemma 4E4B, understand the workflow, then you can move up or down or sideways actually, depending on what you are building. So, the way I understand the whole Google AI ecosystem is you have Gemma as the open model family. You have Google AI Edge, which is the on-device AI development ecosystem. You have Light RTLM, which is the runtime for running languages models across all the devices. You have AI Edge Gallery, which lets you try on-device models and see the experience just more directly. And if you need huge scale, you know, things like strong managed infrastructure or frontier level cloud reasoning, you still have Gemini and Google Cloud that you can use or another out frontier LLM that you can use. The The reality is uh a lot of big products and serious products are going to use hybrid setup. They're going to use cloud for certain things and you're going to use uh local for other things. As an example, imagine a local AI tool for a professional service firm. So, the local model is going to read the sensitive drafts, you know, checking for the issues. It's going to strip or summarize all the private details and prepare a clean version of the problem. Then, when the customer wants deeper reasoning, a cloud model can help with the sanitized version. That to me feels like a more natural architecture than just putting everything into the cloud, which a lot of people don't want. You basically have local handling the private files as a first pass and then cloud handles the heavy thinking when you need it. A human can approve the work before anything important goes out. This is how I'm starting to think about building a lot of these products. Beyond Google Gemma, I'll give you a quick primer on the other families or other open model families you'll hear about and some of the pros and cons. Llama is a Meta's model family and it's probably the default and it's probably the default open model reference point for a lot of developers because it's a pretty big ecosystem. The upside is the community, the tooling, the examples, support. The downside is you still need to read the license and the model card, especially if you're building a serious commercial product. Qwen is Alibaba's model family and has become very strong, especially around coding, multilingual work, long context and agentic tasks. The China thing is real. A lot of people use Qwen because it performs really well, but if you're in an enterprise, a government, health care, finance or sensitive data environment, you need to separate running open weights locally from sending data to a hosted service and you need to check what your company is comfortable with or what you're comfortable with. DeepSeek is similar in the sense that it's was made it's made a lot of people realize how strong Chinese-based open models really could be. Especially for reasoning and coding. The upside is performance and cost. It's pretty cheap. The trade-off is that some buyers will have procurement, security, or geopolitical concerns. Uh so I'd be thoughtful about where I'd use it, how I deploy it, and even if you want to even if you want to use it and go down that path. There's also GLM or people know it as Z.ai. Um it's another one you'll see pop up a lot a lot. I actually did an episode on it. Uh especially if you spend time on Hugging Face and Ollama or just local model Twitter, you're going to see it a lot. Um the thing to know is that some of these models can be really good for specific jobs. So I wouldn't ignore them just because they're not the obvious brand name. You can test them. You can read the model card. You can check the license. And just play with them. Um but what you might deploy in the sense of for your business or for what you're doing might be very different. There's also uh Mistral which is the European model family. I think they're based in France. Um if you care about efficient models and you know they do a lot of releasing uh a lot of practical developer use cases, they're pretty good. Um it's a strong model with a pretty builder-friendly posture. Um but the downside is the lineup is a little confusing. Some models are open, some are commercial. Um so you know some question marks there. Um Microsoft also has uh their open model family. It's called Phi. Um I think it's interesting if you care about smaller, faster, lower latency models. Um but for a lot of use cases, uh I haven't seen it work very well. Um and honestly, there are new models showing up all the time. It feels like every other day. Um and that's why Hugging Face matters. Um you're not going there just to find the big models. You're going to find these like weird specialist models, these like community fine-tunes, quantized versions of stuff, these forks, and model cards that tell you whether something's actually use- usable for the workflow. So, you don't know you don't need to memorize uh all of this, um but the takeaway basically is that there's these ecosystems, and your job as a founder uh or just, you know, someone who's playing with these models is to pick a model family that fits your workflow, that you connect with that company, uh you like how they do things, um and then go from there. You can play with a lot, learn a lot, and then, you know, pick a family. So, how do we make this whole thing real? Like, if you actually want to run Gemma, here's how I would do it. I would start with LM Studio. I would download LM Studio. It's free to download. You open the app. You search for Gemma 4. If your machine is solid, try E4B, but if your machine is a bit slower, older, I would look for E2B. And then, I would look for the quantized version uh if you're using the GGUF path, because you want the model model to run just a lot more comfortably. Once it downloads, open a chat and just ask it something really simple. Um you know, I would use like a business prompt, because I want you to feel the value immediately. It's sort of an aha moment. Maybe it's something like, "Read these customer notes and turn them into a one-page memo about what customers are struggling with. What has change and what the business should fix this week. And then just paste like some customer notes or just fake customer notes just if you want to see the value. The point of this exercise is just basic. Uh the model is now running on your machine and you're using AI without sending that prompt to a cloud model. I believe everyone should try that and feel what that is because I do think that it's it's just going to be a lot more common and there's just it's going to unlock your brain in a completely new way. After that, go to LM Studio's developer section and start the local server. Uh because that that just gets a lot more interesting because other apps can talk to the model on your laptop. Your computer becomes this little AI server. So, you can have a script or a prototype or an or or just an internal tool that can call the model through localhost and you get an answer back. I think that's when you start to see how products are going to get built in the modern age. The second path is Ollama. So, install Ollama and run uh Ollama pull Gemma 4. Then run Ollama run Gemma 4:E4B. Now you have Gemma running locally from a command line. Ollama also gives you uh local API port. I think it's on 11 uh 11434. It is useful because you can connect your own app or script to it. If you want to test a larger model later, you can try the 12 billion, 26 billion, or 31 billion versions assuming your hardware can handle it. And you can ask uh an LLM if your hardware can handle it or you can do yourself and just suffer through the slowness and the pain of it. >> [gasps] >> The third path is Google AI Edge and light RT LM. I would only use this path if I wanted to build an actual app and a model inside of it. For example, maybe I'm building a mobile app and the model is running on the phone. Or could be like a browser app where the model runs locally. Um or it maybe it's a desktop app with a private workflow. Um or something on an edge device. Um light RT LM is designed for that world. Um Android, iOS, web, desktop, and edge environments. Um that is the path from local AI as a demo to local AI as a product. So, here's the hardware cheat sheet that I would use. If you have 8 GB of RAM, start small and keep the first test simple. But, if you have something like 16 GB of RAM, you can do some useful experiments with models like E4B and smaller quantized models. If you have 30 GB 32 GB of RAM, you have way more room to do, you know, larger local workflows. If you have a strong GPU or a workstation, uh like a DGX Spark, uh the bigger models become just much more realistic. And for phones, I would think a lot less about model size and more about the job. So, can the model understand a photo? Can it summarize audio? Can it classify something quickly? Can it help a worker in the field? Can it run without a strong connection? Can it do something useful inside the app before the user even thinks to ask? Now, let's build the first workflow in our heads. So, I would make a folder uh on your desktop called customer notes. And inside that folder, I'd put 10 support tickets for a specific business. Let's say it's a home health agency or med spa or water damage restoration company. The notes might say something like, "I tried to reschedule but couldn't find the link." or "The technician didn't explain what happens next." or "Hey, no one actually confirmed my appointment." or "I was charged twice here." Then, I would run a local model like Gemma and ask it to create a file called "What customers are telling us .md" the markdown file. The output should include the repeated complaints, the exact customer language, the likely root cause, the part of the business that seems broken, and the one thing the operator should test this week, the high priority stuff. This is a good first local AI workflow because it's useful and it's simple. What do you have here, right? You have this private messy data, the model runs next to it, and the output is a memo someone can actually use. And then once you actually go and, you know, you're going to go and do this and and get the output, you're going to like the unlock I was talking uh before, like it's going to unlock something in your brain. You're going to see this pattern everywhere. A folder of customer calls become a market research memo. A folder of support tickets become a product roadmap signal. A folder of PDFs become like a risk checklist. A folder of drafts become a pre-send reviewer. This is why I always start with uh workflows before I'm fine-tuning anything. People here, you know, open model and immediately want to train their own model and I get it. I get why. I was actually the same way. Um it sounds really cool, but I feel like that's like an advanced move. The practical move, the beginner move, where you should start is just to find a repeated workflow first. You pick one folder, one model, one output, and you run it like 10 times. You see where it gets confused. You see where you can improve the prompt and add examples. You add a checklist and then you create like a small eval. You know, what's an eval? An eval is just a small It's just a test that tells you whether the model did the job well enough. For this workflow, for example, the this the eval could be like really simple. It could be like, you know, take the same 10 customer notes and run them through Gemma locally and then run them through a strong cloud model, a frontier model, and then just compare the outputs. And then you you know, you ask, "Did Gemma, you know, catch the same complaints? Did Gemma pull the right quotes?" And did it follow the format? Did it miss something? The comparison actually teaches you where locals are already useful and where you still want that stronger uh cloud model and what you know, how you should think about the hybrid model I was talking about. That's really how I think about local versus cloud decisions. Use uh local for private, repetitive, fast, offline, device native, and high-volume workflows. Stuff that you want to run all the time. You use cloud for deep reasoning, giant context, broad research, in cases where the strongest model changes the quality of the answer. So, you use both when the product has sensitive data and hard reasoning. A lot of valuable products will work that way. You know, local first pass, you do the cloud escalation, human approval for anything important. I think that's the way work's going to get done. So, I want to give you three startup ideas where local AI actually matters and these are the kind of businesses I would look for, niche, useful, cash-flowing businesses that you don't need to raise venture for, and tied to a painful workflow. The filter is pretty straightforward. So, I look for a customer with sensitive data, repeated review work, bad software usually, uh expensive mistakes, like and mistakes that will cost them a lot, and a workflow that happens close to the to the device. That combination is like the interesting zone for me. So, let's go through the three ideas. Uh I want you to steal these ideas, and at the very least it'll get your creative juices flowing with how you can use uh local AI to run model, build apps, and make money. Idea number one is a local QA reviewer for home health agencies. So, home health agencies have nurses and caregivers, and they go into people's homes, and they write, you know, visit notes, and updating care plans, and dealing with billing and compliance. The paperwork is a pain. It takes a lot of time if you've ever witnessed it in person, but it matters so so much. Like, a missing detail can create a billing delay, and a vague note can create extra admin work, and a mismatch between the visit and the care plan can create a ton of risk, and we don't want that. So, the first version is a local desktop app for the agency. The agency drops in visit notes and care plans and dictated transcripts, and then the model is going to review them before the submission, and it should look for flags. So, it's going to flag things like this note mentions dizziness, but vitals are missing, or the caregiver described a medication change, but the follow-up instructions is pretty unclear, or the note may not support the billed service level. The buyer mostly cares about fewer documentation problems before the billing or the audit or a supervisor review. So, if you solve that, you have their attention. Now, I don't want to just give you the idea. I mean, how would you actually grow this? If I was starting this business, how would I grow this business? I would actually start it as a service. So, I would find five small home health agencies and then would offer to review a batch of notes. I would do the review with AI helping behind the scenes with the local AI. And I would inspect everything manually with like human beings, myself first. I would write down the 20 issues that keep showing up and those issues become the checklist and then the checklist eventually becomes the product. So, you have this wedge, it's pretty simple, where you're catching documentation problems before they cost the agency time or money and then you build from there. I love this business and totally would start it. The second startup idea is an offline field report co-pilot for restoration contractor. So, think water damage or fire damage or mold remediation, things like that. Those teams are out there field taking photos, recording notes, documenting damage and creating reports for homeowners and insurance adjusters. I unfortunately had this, so I know a little bit about it. The job is actually pretty visual. Um it's also physical, right? They're It happens like away from a desk and the report matters because the report becomes the handoff between the technician, the customer, the office and the insurance process. So, how would we build a product here? The first version is a mobile app. So, a technician walks through the property, takes photos, record voice notes and the app drafts the report before they leave the site. So, it can flag missing pieces while the technician is is still there walking around. You mentioned the basement, but there are no basement photos. You took a photo of ceiling damage, but there are no moisture meeting reading, things like that or or maybe like the affected room is like missing. Could be the homeowner explanation is way too technical. Here's a clearer version they can understand. And the last part of that is underrated. In a stressful home damage situation, clear communication is part of the product, right? Um so, if you had that, that would be key. How would I grow this business? Well, I would pick one niche first. I wouldn't go after everything. So, say I'm going after, you know, water damage restoration. I would talk to owner-operators. I'd look at their current report templates, study the software they use, which is some old stack, and I'd build around the checklist that's already in their head. The demo is actually the easy part. You know, send me three old jobs and I'll show you how fast your techs could create reports. If that works, then the product could expand from there. That's just the wedge, right? Uh it can go into QA and estimates and insurance packets, customer updates, and training new technicians. But, I would start with the field report because it's specific and obviously super annoying. And, you know, I just think that there's uh when you look at some of these old softwares that, you know, these people are using, I recently had some water damage in uh at my at my apartment, and I I was seeing some of the software, and it's it's antiquated. Like, it's stuff from like the early 2000s. So, I think that there's just opportunity to create local AI-native software, uh and and and wedge now. And that's why I said in the beginning, like, I think there's a 24-month uh window and opportunity to do some of these products. Let's go into startup idea number three. So, startup idea number three is a local pre-send reviewer for professional services. So, every professional service firm, or 99.9% of them, has a version of this workflow. Someone writes a client email, a proposal, a memo, a contract summary, an investment note, an HR note, and then someone and And ask someone else to check it out before it goes out like a review. And it happens constantly. Law firms, accounting firms, wealth advisors, uh recruiting firms, um even consultants have a version of this. So, the first version is you build a local desktop app that reviews outbound drafts before they leave the company. So, for a wealth advisor, it could be flagging language that sounds like a guaranteed return, which is a definite no-no. For a law firm, it'll flag a sentence that sounds too definitive. For HR, it's going to flag sensitive employee information that should stay out of the threat. For an agency, it flags a promise that the scope does not support. And for an accountant, it flags a number that doesn't match the attached file. You'd be surprised how often that happens. The product is basically a second set of eyes for sensitive work. It's basically schmuck insurance is the way I think about it. And maybe that that would be the name, schmuckinsurance.com. Someone tell me if that's taken. How would I grow the business? I would start with one vertical and one document type. For example, I would do email review for independent wealth advisors. Not everyone, probably not the big banks to start, uh independent wealth advisors. I would interview 10 advisors and then ask them which emails make them nervous. I would collect uh anonymized examples. I would turn their real concerns into a review checklist, and I would build a local tool that checks drafts against that checklist. Uh obviously, this is so sellable because the buyer understands this behavior, and they already asked someone to check the draft. So, you're just basically giving them a faster first pass that lives closer to their client data and internal rules. I love this idea and hope I hope a few of you take it. By the way, if you're not building one of these ideas tomorrow, I still think you should learn local AI because it does change how you work with your own files. So, I think just like from a personal productivity perspective, uh it's still super super helpful. So, you know, if you're working uh at a company, say, and and you just want to be more productive, so you have more time to scroll TikTok or watch movies or hang with your family, make a folder called local AI lab and then put 10 files that matter to your work in that folder. It could be anything from sales calls or meeting transcripts, old tweets, ideas that you have. Then, run Gemma, whatever model you choose, to make it produce one useful artifact. And then ask it to create a weekly business pulse or ask it to find what's changed in customer conversations or meeting notes. Uh ask it to group feature requests by the actual pain behind it. Ask it to review drafts and tell you what your audience keeps responding to. The key, basically, is to produce a file, a memo, a checklist, uh a brief, a report, or a review that you can reuse. A chat answer is nice, um but a useful artifact changes that workflow. This is the first wrap I would recommend. A model reads the folder, the model writes the file, you inspect it, you improve the workflow, then you run it again. If you do that a few times, your brain really starts to connect the dots. You start noticing where private data is trapped in folders. You notice which reviews happen over and over again. And you notice which workflows depend on someone checking a form, reading a note, comparing two files, cleaning up a report, or writing the same kind of memo week after week. Hopefully, this episode got your creative juices flowing because once you see the pattern, you start spotting local AI businesses everywhere. You can learn enough of the map to spot where these models belong without turning yourself into a local engineer overnight. I believe some AI local AI belongs in the cloud and some some AI belongs in the device and a lot of the best products of the next couple years are going to combine them both. So if I was starting today, what I would do is I'd run Gemma locally and read model cards on hugging face. I'd learn the difference between LM Studio and Ollama. I'd play with Google AI Edge and then look for one boring workflow where local AI actually makes the product better. Those categories are things like private data, offline work, camera audio context, or low latency, or if there's a high repeated API cost. If there's a buyer who feels better when the model is just close to them. A workflow where a small agent team could recheck, summarize, and prepare work every day. That's like the hunting ground. Local AI is just way easier to understand once you chop stop treating it like a model benchmark conversation and start treating it like a product conversation. You have to ask yourself, where is the work happening? Where is the data? Where is the device? Where is the trust issue? Where is an annoying review loop? And then you answer those questions and you just start seeing the idea. So overall, I hope you understand a little about, you know, the the the core things you need to understand about local AI, some of the models, some of the apps you need to download, some of the workflows that you can build, and some of the business opportunities that exist. I just don't see that many non-technical people playing with local AI and the last 2 months or so, I've I've gotten deeper and deeper into it and it's just like I said, it's been connecting the dots and I'm grateful for it. Uh I hope you have a creative day. I read every single comment in on YouTube and respond to most. So, I'll see you in there. Share share this with a friend who you think could benefit from understanding local AI in a clear way and I'll see you next time. Happy building.