youtube.nixfred.com nixfred.com

I Gave an AI Zelda and One Rule: Don't Cheat.

A coding agent spent ten days writing a program to beat The Legend of Zelda on the NES under one rule: no cheating. The model never touches the controller, because Link needs a decision sixty times a second and a language model takes seconds to think, so instead it authored 25,000 lines of Python, a Lua bridge inside the emulator, and a four message protocol. Six emulators grind all 341 rooms in parallel, every frame and heart and bomb and moment of hesitation gets priced in one currency, and the entire overworld is decoded straight out of the ROM. The first finish took 1 hour 42 minutes. The final run takes 37 minutes and 2 seconds, verified byte for byte.

Published Sep 23, 2026 48:17 video 51 min read Added Oct 8, 2026 Open on YouTube →

At a glance

Bears Gaming Den spent ten days pointing a coding agent at The Legend of Zelda on the NES with one constraint attached: learn the whole game, beat it from power on, and do it without cheating. The twist in the setup is that the AI never touches the controller. Link needs a decision sixty times a second and a language model takes seconds to think, so instead of playing, the model spent those ten days writing a player: 25,000 lines of Python, a Lua script living inside the emulator, a priced decision function that scores nine possible futures every eight frames, and a route planner that decoded all 128 overworld screens straight out of the ROM. Along the way it kept a journal, 45 entries, each one a problem and what it changed, and the narration is largely that journal read back with the footage next to it. The first finish took 1 hour 42 minutes; the run that ships with the video takes 37 minutes and 2 seconds, and the final 38 minutes of the video are that run playing out, uncut and without commentary.

One note on the shape of this video before the rebuild starts. The narration runs from the opening to roughly the eleven minute mark, and all twelve of the creator's own chapters sit inside that stretch, ending at "Conclusion and full run." Everything after it, about 38 of the 48 minutes, is the complete final playthrough with game audio and music only. This page reconstructs the narrated section in full and then tells you exactly what is in the run section and where to look in it.

The pitch: ten days, six emulators, one rule

The opening is plain about the ambition. "I wanted to see if I could actually get AI to learn an old school video game like Legend of Zelda and have it actually learn the game start to finish," he says, "and not only learn it and beat it, but actually potentially beat a world record. That was what I set out to do. And I just wanted to see what would happen."

The montage that follows is the whole project compressed into one image: six copies of Zelda running at once, all in the same room, all trying the same fight over and over. What gets recorded off each attempt is three numbers. How many frames it took. How many hearts it had left. Whether it won. Ten days of that, and at the end of it, in his words, "the whole game is faster than most humans will ever play it."

That is the frame for everything below: not a neural network learning to press buttons, but a coding agent writing, watching, reading its own logs and rewriting, with a brute force search bolted underneath it to grind out the last frames.

Why the model never plays the game

This is the constraint that shapes the entire architecture, and it lands in the first minute. "AI actually never plays the game live. It can't. Link wants a decision 60 times a second and the language model takes seconds to think."

The NES runs at roughly 60 frames per second, which means the gap between what the game needs and what a language model can deliver is three or four orders of magnitude. There is no prompt engineering fix for that. So the model does not try to be the player. It becomes the author of the player. On day one it builds itself what he calls "a set of hands."

SLOW LOOP: TEN DAYS, ONCE PER ITERATION The coding agent Writes the player, reads the logs, writes the journal, rewrites the code 25,000 lines of Python Pathfinding, lookahead, search, route planner, boss players 45 journal entries, each ending in a rule FAST LOOP: EVERY FRAME, FOREVER Python side The player the agent wrote, driving the game over a socket 4 messages only THE WHOLE NETWORK PROTOCOL 1. Hold buttons for N frames 2. Tell me the game state 3. Save a bookmark 4. Load a bookmark Lua script in the emulator Presses the buttons, reads the 2 KB of RAM, saves state in about 1 ms Zelda, at 60 fps One decision due every 16.7 milliseconds THE MISMATCH 60 per second what Link needs seconds each what the model costs
Figure 1. The architecture in one picture. The model lives only in the slow loop on the left, where a round trip can take seconds and nobody cares. The fast loop on the right is ordinary deterministic code talking to a Lua script over a socket, which is why it can keep up with a 60 fps console. Every capability in the rest of this page is built out of those four messages and nothing else.

The set of hands: four messages and nothing else

The interface is almost comically small. "A tiny script that lives inside the emulator and listens to the network port. And on the other side, Python sending it four kinds of messages. Hold these buttons for this many frames. Tell me the game state. Save a bookmark. Load a bookmark. That's the entire interface. Everything you're about to see is built out of those four things."

The transcript does not name the emulator, but the project's own files do. The companion repository, bearsgaming-ui/AIBeatsZelda, names BizHawk 2.11.1 on Windows x64, with bridge.lua as the in emulator script ("hold buttons, read RAM, save/load a bookmark, over a socket") and zelda/emulator.py on the Python end. That choice is not incidental. BizHawk is the emulator the tool assisted speedrun community standardizes on, its Lua API ships a comm library with socket and HTTP support so a script inside the emulator can talk to an outside process, and it records input as a replayable .bk2 movie, which is what makes the "don't cheat" claim checkable later. Pick an emulator without scriptable sockets and savestates and none of this project exists.

Then comes the part that is the actual story. "For 10 days, it wrote the player, watched what the player did, read the logs, and rewrote it. By the end of it, 25,000 lines of Python, and it kept a journal, 45 entries. Every one of them is a problem, what it tried, what went wrong, what it changed. A lot in this video is just reading that journal back with the footage next to it."

That journal is the unusual artifact here. It is not a training log and it is not a diff history. It is a sequence of failure writeups, each one ending in a rule, and several of the rules get names: "passive is losing," "read the manual," "measure, don't assume." The repository ships the journal directory publicly, where the count is 46 entries rather than 45, presumably one more written after the video was cut.

It never looks at the screen

Here is the second structural decision, and it is the one that quietly makes the project tractable. "One important thing to note, though, is it never really looks at the screen. It really just looks at memory. Everything exists in 2 kilobytes worth of memory."

Two kilobytes is the NES's entire work RAM, and that is not a figure of speech. The console has 2,048 bytes of internal RAM, and the complete observable state of Zelda at any instant lives in there. So no vision model, no screenshot pipeline, no OCR of the heart meter. The player just reads bytes.

He walks the ones that matter: "One of them is Link's X, the other is Y, one is hearts, one is keys, and one is bomb. 12 of them are where the monsters are, and 12 of them are what they are."

That last pair is the important bit of design. Zelda tracks up to twelve active enemies per screen as two parallel arrays, one slot per enemy for position and one for type. Read both and you have a full tactical picture of the room in twenty four bytes: where everything is and what everything is. The public RAM map on Data Crystal documents the same addresses the project relies on, including Link's X at 0x70, Link's Y at 0x84, the heart byte at 0x066F (low nibble filled, high nibble containers), keys at 0x066E and bombs at 0x0658.

And then the first journal entry is about getting burned by exactly that kind of documentation. "The journal from day one: the wiki's table of game modes has two values swapped, and the code is waiting for the wrong number." Day one, hour one, and the lesson is already that a community reference can be subtly wrong and your code will sit there politely waiting for a state that never arrives. The game mode byte is the one that tells you whether you are on the title screen, in normal play, in a transition, or in a cave, so getting it wrong means nothing downstream fires at the right moment.

Measuring the walls instead of trusting them

The map is in memory too, and that part comes free: "Every screen is a grid of tiles, and the game keeps that grid decoded." But there is a hole in it. "What the grid doesn't say is which tile Link can walk on, so it measures that."

The measurement procedure is pure physical experiment, run inside a video game. "It walks Link into things. Walks up a column until he stops. Note which tile stopped him. Do it again for every tile type."

Two hard facts came out of it, and they are the sort of thing you only get by measuring:

MEASURED BY WALKING LINK INTO EVERY TILE TYPE UNTIL HE STOPPED 16 x 6 Y as stored in RAM top of the box, Y + 3 bottom of the box, Y + 9 16 px wide 8 px 16 x 16 sprite footprint 6 highlighted tiles are the ones a single 16 wide strip can be touching at once: 3 columns, 2 rows
Figure 2. The two facts the AI established by experiment rather than by reading a wiki. The sprite is 16 by 16, but the part that collides is a 16 by 6 strip starting 3 pixels below the stored Y, and solidity is asked per 8 by 8 tile. Together those explain every stopping position Link had been mysteriously landing on.

All of it gets written down. "All that goes into a file. And that being said, when it's all said and done, we know which tiles you can walk over and which tiles will stop you."

Then there is one more inference rule, and it is the cleverest small thing in the whole build: "When a plan step fails and exactly one of the tiles that would have entered is unknown, then the tile is now solid forever on every screen and every future run."

Read that again, because the word doing the work is exactly one. If a move fails and several candidate tiles are unknown, you have learned nothing you can safely act on, because you cannot tell which one blocked you. But if precisely one unknown tile was in the way, the failure uniquely identifies the culprit, and the conclusion is permanent and global: that tile type is solid, on every screen, in every future run. It is single candidate blame assignment, the same move a human does when a process of elimination leaves one suspect, and it turns a stream of ordinary failures into a monotonically growing map of the world.

Walking is a search, and passive is losing

With a solidity map in hand, movement becomes a graph problem. "Walking is a path search," he says, and then lists the rules it runs under. The captions garble one clause here, but the sense is clear from the surrounding list: the search steps in eight pixel increments, which is exactly the resolution collision is resolved at, so the search space and the physics agree.

The other two rules are the interesting ones.

Unknown tiles are assumed walkable until proven otherwise. The planner is an optimist. It will happily route through a tile it has never classified, find out the hard way, and then apply the permanent solid rule from the section above. Pair those two and you get a system that explores aggressively and never makes the same mistake twice, which is a much better trade in a game with a fixed finite tile set than cautious avoidance would be.

Every monster on the screen adds a cost to the squares around it. Not a hard no go zone, a cost field. "So it will always choose the lowest cost value that it's assigned to each tile." Danger is priced, not forbidden, which means the planner will walk right past a monster when the detour costs more than the risk, and will take the long way around when it does not. That single design choice is why Link can thread rooms rather than freeze at the first thing with teeth.

Except that on day two, he froze at the first thing with teeth. "On day two, Link was just standing in the doorway because a slime was in its way." A single slime in a doorway, and the cost field said the safest thing available was to not move at all.

The correction came from the creator, in plain English, and it is the clearest example in the video of how the human half of this loop actually works. "I said, you don't have to kill all the slimes, but you do have to make progress."

That night's journal was called passive is losing. And that's a rule that came out of it. Standing still is never safety. Detour or kill what's in the way and keep moving.

Three things are worth noticing about that exchange. The fix was not a code patch handed down, it was a goal stated at the level of intent. The AI turned it into a named rule in its own journal. And the rule is a general principle about the cost function, not a special case about slimes in doorways, which is why it stuck. "Standing still is never safety" is the kind of sentence that later shows up as a positive price on standing around, which is exactly where it reappears in the combat section below.

Finding level three by wandering

The other thing the project deliberately did not do is hand the AI the answer key. "One thing we didn't do is actually load like any kind of world map to find level three. It literally just wandered around until it found it. So it would go left, right, up, down, and it would just do this a million times until it found the best route to go."

The payoff is a route in compass directions, learned from scratch: "20 screens later, it had the route west, north, west three times, south, east, up, and around the river."

Twenty screens of overworld, found by exhaustive wandering, with no prior knowledge of where dungeons live. Keep that in mind, because an hour of project time later the same system will have the entire overworld decoded out of the ROM and will never need to wander again. The video shows both phases, in order, which is a more honest portrait of how this kind of build actually progresses than a finished demo would be.

The emulator as a time machine: 341 rooms, 341 searches

This is the section where the project stops being clever and starts being brutal. The enabling fact is a performance number: "The emulator can save its entire state in about a millisecond."

A millisecond per savestate changes what is possible. It means you can rewind reality thousands of times per minute, which means you can treat a video game as a sequence of independent optimization problems rather than one long fragile attempt. So that is exactly what happens. "The game is cut into 341 segments, roughly one per room."

And every one of those 341 segments gets ground:

Every segment is a search. Load the bookmark at the door. Try the room with a different random seed. Log how many frames it took and how many hearts came out. Do that 60 times or 90 across six emulators at once. Keep the best. Bookmark the end of it. Next room.

(The auto captions mangle two words in that passage: "random seat" is random seed, and "how many cards came out" is how many hearts came out. The numbers are his.)

That is the six copies in one room from the opening montage, explained. The parallelism is not there to make the AI smarter, it is there to make the clock cheaper: six instances times sixty to ninety attempts means every single room in Zelda gets played several hundred times, and only the best run of each survives into the final route. The repository notes that six instances want a reasonably modern CPU and that a full pass takes four to six hours, which is the real cost of this approach.

THE PER ROOM SEARCH LOOP Load the bookmark at the door A savestate restores in about 1 ms SIX EMULATORS, SAME FIGHT, AT ONCE EMU 1 EMU 2 EMU 3 EMU 4 EMU 5 EMU 6 60 to 90 attempts per room, a different random seed each time, losing attempts killed mid run WHAT GETS LOGGED Frames taken Hearts remaining Won or lost Keep the best attempt Bookmark the end of it NEXT ROOM 341 segments, roughly one per room, from power on to the end of the game
Figure 3. Why a one millisecond savestate is the whole ballgame. Each room becomes an independent search problem with a known start state and a scored outcome, six emulators grind it in parallel, the best attempt is frozen as the start state of the next room, and the process repeats 341 times. Nothing about this is intelligent; it is a time machine pointed at a fixed cost function.

Everything becomes currency

A search needs something to sort by, and this is where the design gets its spine. "To pick the best attempt, it needs one number per attempt. So everything gets converted into currency."

The unit of that currency is the frame, because the frame is what the project is minimizing:

That bomb price is the clearest window into the method. A bomb is not valued because bombs are good. It is valued at exactly the number of frames of detour it will save you later, which means a resource in your inventory and a shortcut in the map and a near death heart pickup are all denominated in the same unit and can be traded against each other by a sort function. "So everything gets converted into currency whether it's good or whether it's bad. And every decision is based on that."

And the prices were never claimed to be correct. "These prices were guesses at first and every one of them got changed when the footage disagreed with it." That is the loop in one sentence: guess a price, run it, watch the footage, and when Link does something stupid, the price was wrong, not the game.

One more efficiency detail, and it is a classic: "any attempt that can no longer beat the best one gets cut off in the middle so the emulators don't waste time finishing losing runs." That is branch and bound. Once an attempt's accumulated cost exceeds the incumbent best, there is nothing left to learn from watching it lose politely, so it dies and the emulator slot goes to the next seed. With six instances and ninety attempts a room, killing hopeless runs early is the difference between a four hour pass and an overnight one.

The one rule: no cheating, and how it gets proved

The title's rule is not decoration, and the enforcement is specific. The final run is not a stitched together highlight reel of the 341 best segments. It is one input log, replayed cold.

The whole input log, every button, every frame from power on is played back from a fresh emulator with no bookmarks. A human record is set live in one sitting. And that's exactly what we're going to do here, where AI is going to basically have this thing played out all in one sitting as well.

So savestates are a practice tool, not a run tool. During the search they are everywhere, because that is the entire mechanism. In the run that counts, they are gone: the bookmarks were only ever a way to discover which buttons to press, and what survives is the button sequence itself, which has to work from frame zero on a fresh machine with nothing loaded.

The repository spells out the rest of the rule set that the narration only implies. Unmodified game, controller inputs only, no memory editing, no cheat codes, no glitches (specifically no screen scrolling and no block clips), one continuous session, no human input during the run, and the replay verified byte for byte against the original recording with a published SHA-1. The run ships as a BizHawk .bk2 movie plus a plain text inputs.txt with one line per frame, a search_log.txt showing every room's attempts, and a VERIFICATION.txt demonstrating that the same inputs in a fresh emulator reach the same RAM state. That is a higher bar of evidence than most human submissions carry, and it is the part of the project that makes the word "run" mean something.

Note what reading memory is not. The player reads RAM constantly, and that is not on the forbidden list, because reading is observation. Writing to RAM is what "memory editing" means, and that is forbidden. Any tool assisted run makes the same distinction.

Day five: thirteen hearts, the best sword, and zero bombs

In between the architecture, the video drops one perfect illustration of what route planning is actually for. "In day five, it reached the last dungeon with 13 hearts, the best sword in the game, and zero bombs in a dungeon made for bombable walls."

Thirteen hearts and the top tier sword is a strong position by any measure. Zero bombs at the door of a dungeon built around bombable walls is a soft lock in disguise: every wall that should have been a shortcut becomes a long way around, and the frame cost compounds across the whole level. The planner had optimized combat capability and ignored a consumable, which is precisely the failure that the 220 frame bomb price exists to prevent. This is the kind of thing the journal was for.

Where the first fighting approach died

Routing was solvable. Fighting was not, at least not the way it was first written. "Fighting is actually where the first approach died. In level three, we entered a room with five of the red knights, and the handwritten reaction code lost it hundreds and hundreds of times."

The red knights are Darknuts, which are the single most instructive enemy in Zelda for an AI to fail against, because they are armored on the front and can only be hit from the side or behind. Five of them in one room is a geometry problem disguised as a fight.

What killed it was not that the reaction code was dumb. It was that it was two pieces of reaction code. "The journal's frame trace found that the get out of its way logic and get behind it logic disagreed every frame. So Link just jittered in place until a red knight walked into it."

That is a textbook oscillation. One rule said move away from the thing in front of you; the other said move around behind it; those point in opposite directions; and every frame the winner flipped. The result was a Link vibrating on the spot while an armored knight strolled into him. And notice how it was diagnosed: not by watching the footage and guessing, but by a frame by frame trace in the journal that showed the two subsystems contradicting each other on every tick. The journal was not documentation. It was the debugger.

The best idea in the build: nine futures, every eight frames

Then the fix, and he flags it himself as the high point. "The coolest part of the whole build though happened now. It's when it stopped reacting and it started looking ahead."

The mechanism is simple enough to state in one breath. Every eight frames, take a snapshot of the game. From that snapshot, try all nine things Link can do:

"Now it has all nine possible futures and has to pick one."

This is the move from a reactive controller to a planner, and it dissolves the oscillation problem by construction. Two rules can no longer disagree, because there are no longer two rules. There is one score, nine candidates, and an argmax. If "get behind it" and "get out of the way" point different directions, the scorer simply tells you which of those two futures is worth more, and Link does that one.

The eight frame cadence is the tuning knob that makes it affordable: at 60 frames per second it replans roughly seven and a half times a second, which is fast enough to react to a Darknut and slow enough that the simulation cost stays sane. The repository puts this logic in zelda/lookahead.py, described as combat planning with eight frame lookahead.

And the scoring is the same trick as before, applied one level down. "Each frame gets turned into a number the same way the attempts did. A list of things that matter each with a price added up."

Priced thingPriceScopeWhat the price actually encodes
Every frame that passescost 1Attempt scoreTime is the objective. Existing is expensive, so nothing is free.
A heart picked up while low on lifepays 100Attempt scoreSurvival is not a separate constraint. It is bought with the same money as speed.
A bomb in handpays 220 framesAttempt scoreExactly the detour it saves later by opening a wall. A resource priced as a shortcut.
Every frame that passescost 0.5Combat lookaheadFighting is scored at finer resolution than routing, so a half frame of advantage still registers.
Dyingcost 100,000Combat lookaheadIn his words, "which just means never." Large enough that no payout can outbid it.
Losing half a heartcost in the hundredsCombat lookaheadExpensive but finite, so Link will trade health for time when the trade is good.
Taking one hit point off an enemypays 60Combat lookaheadPartial progress counts, so the planner commits to a fight instead of flinching.
Killing an enemypays 150Combat lookaheadMore than the sum of its hits, which rewards finishing rather than chipping.
Standing in front of a knightcost 60Combat lookaheadDarknuts block from the front. The shield's facing is read straight out of RAM.
Standing beside or behind the nearest enemya gentle pullCombat lookaheadThe only geometry where a sword hit actually lands. Encoded as attraction, not as a rule.
Standing around doing nothingpriced highBothThe "passive is losing" rule from day two, converted into money.
Figure 4. Every price the video states, at both scopes. The attempt score sorts whole room attempts; the combat lookahead scores single frames inside a fight. Amber is a cost, green is a payout. Nothing in this table is a rule or a condition; it is all just money, which is why nine candidate futures can be compared with a sort.

Two details in that list are worth pulling out.

The shield facing comes from memory, not from the sprite. "It knows which way the shield is facing because the byte is in memory, too." A vision based agent would have to infer a Darknut's facing from pixels. This one reads it. Which is the whole argument for the memory first design in one example: the thing that is hardest to see is trivial to look up.

The hittable side is a pull, not a rule. "There's a gentle pull towards the one place a sword hit can happen, beside or behind the nearest enemy." He does not write "always attack from behind." He adds a mild gradient toward the position where a hit is possible and lets it compete with everything else on the list. If getting behind the knight would cost half a heart and three hundred frames, the pull loses, and that is correct behavior.

Then the line that is the best summary of the entire project: "Those prices are basically Link's entire personality."

Tuning a personality, and the two ways it goes wrong

The prices did not come out of an optimizer. "All those prices had to be set by hand and argued with." And he is specific about the two failure modes, which are mirror images of each other:

"Both those things happened in the run." And the fix is not to lower the first two prices, it is to add a third: "those fixes need to happen by making standing around have a high price as well." You cannot make cowardice unattractive by making danger cheap. You make it unattractive by charging rent on inaction. That is the same rule the journal named on day two, now expressed as a term in the objective function.

His own description of the resulting system is better than any paraphrase: "All these things are fighting with each other to try to get the most logical way to complete the game as quickly as possible."

Every boss: measure, don't assume

Bosses got the same experimental treatment as walls. "Every boss got the same treatment. Measure, don't assume." Which, in practice, means firing every weapon in the inventory at each one and recording the damage number, instead of trusting a strategy guide.

The findings he reads off are the real Zelda 1 boss rules, and they are a nice demonstration of why assuming would have failed:

A caption note, flagged once: the auto captions mangle the boss names here. "Petra" is Patra, "GMA" is Gohma, and "Dongo" is Dodongo, all of which are unambiguous from the stated weaknesses. The last one is read as "the flying glee here," which is either Gleeok, whose heads detach and fly, or a second pass at Patra, whose satellites orbit a core. The behavior he describes, outer before inner, fits Patra exactly, so that is the reading used above, with the ambiguity noted rather than hidden.

The real point of the section is the rule that came out of it, and it is the one that sounds least like an AI and most like an engineer who has wasted a week. "This journal entry is called read the manual, and that's a rule that came out of it. Research first, experiment second."

Which is worth sitting with for a second, because it is a reversal. The project's earlier posture was measure everything from scratch, because the wiki had its game mode table wrong on day one. Measuring every boss with every weapon is expensive, and most of the answers were already written down somewhere. So the rule became research first and then experiment to confirm, which is a better policy than either pure trust or pure measurement, and it took getting burned in both directions to arrive at it.

Reading the ROM: the ninth kill and the whole overworld

"Research first" starts with walkthroughs and then goes somewhere most people would not. "It reads walkthroughs, and then it went one step further and read the game's own code. And there's a complete disassembly of Zelda on the cartridge itself."

The game's logic is 6502 assembly sitting in a ROM that is small by any modern standard, and complete annotated disassemblies of it exist publicly. Which means the authoritative answer to "what does this game actually do" is readable, and anything a wiki says about it is a secondary source that can be wrong. Once you have decided to trust the code over the documentation, two things fall out immediately.

The first is a free bomb exploit that is not a glitch. Here is the mechanic, exactly as he states it: "There's a kill counter, and if you don't get hit, every 10th kill is a guaranteed drop. Five rupees, or if the killing blow is a bomb, bombs. So the planner counts to the ninth kill, holds that sword back and pulls out a bomb to take the 10th with it. So, hey, free bombs."

This is real and it is well documented in the speedrunning community as a forced item drop: the game runs a kill counter, a clean streak with no damage taken makes the tenth kill's drop deterministic, and the type of drop depends on what delivered the killing blow. Bomb the tenth enemy and you get bombs back. It is not an exploit of a bug, it is an exploit of the game's actual intended drop table, which is exactly why it survives the no glitches rule. And the behavior it produces is wonderfully specific: a planner that counts enemies, deliberately withholds a sword swing it could have taken, and swaps to a bomb for the tenth, purely to refill the resource it priced at 220 frames a few sections ago.

The second is the entire map. "Then it decoded the entire overworld map straight out of the ROM. All 128 screens, every tile, every cave, every shop and what it sells, every secret and what opens it."

That is the complete Zelda overworld, 16 screens across and 8 down, extracted as data rather than explored. Compare it to the level three search twenty minutes of project time earlier, where finding one dungeon took twenty screens of blind wandering. Same system, same goal, and the difference is entirely about where the knowledge came from.

Route optimization: shuffling the order of the whole game

With the map as data, the problem changes shape again. "With that whole map, it could plan the entire game."

First it had to teach the router how Link actually moves across the overworld, which is not a straight line: "Built a route that walks any leg of the map the way Link actually walks it, with the raft or the ladder, the mazes, the hidden road, the whirlwind, which takes you all over the place."

Every one of those is a different kind of edge in the graph. The raft only works at specific docks. The ladder only crosses single tile water. The mazes are the screens that loop you back unless you walk a secret sequence of directions. The whirlwind, summoned by the recorder, is a warp whose destination you do not fully control. A router that treats the overworld as an open grid will produce a route that is physically impossible, so the movement model has to encode all of it.

Then comes the part that is genuinely a planning problem rather than a pathing one. "Then it shuffled the order of everything."

"The best order for everything." This is a combinatorial ordering problem over a graph with conditional edges, and the AI is solving it by search against the same frame denominated cost function that scores everything else in the project.

The run was still weird

Here is where the video earns a lot of credibility, because the next thing shown is not a victory lap. It is the creator watching a working run and being unsatisfied with it.

But that run was still weird. I'm watching it and looking at it. I'm like, this pathing is still strange. You're running around things that you shouldn't be running around. You should just be killing things instead of wandering around. Sometimes it just sits there and seems to think or not know what to do. So it's just weird.

And then the specific failure, which is the best bug in the video: "For 160 frames, Link zigzagged between two Darknuts that were exactly one sword length away from each other with their sides exposed. Every swing it simulated came back as a miss."

Unpack that. Two Darknuts, one sword length apart, sides exposed. That is the single most favorable arrangement those enemies can present: both of them in range, both of them vulnerable, no shield in the way. And Link stood there for 160 frames, roughly two and a half seconds of a game that is being optimized to the frame, shuffling between them.

The cause is visible in the design. The lookahead's hit prediction was returning a miss for every swing it tried, so the only positive term left on the board was the gentle pull toward the hittable side of the nearest enemy. With two equally valid nearest enemies in opposite directions, that pull tied, and a tie in an argmax over nine futures looks exactly like a coin flip every eight frames. Which is the same oscillation that killed the handwritten reaction code, reappearing one level up. Replacing two arguing rules with one score does not eliminate oscillation; it just moves it into the score, where a wrong simulation and a symmetric situation can still produce a Link who dithers.

The times: from 1 hour 42 minutes to 37 minutes and 2 seconds

Then the moment: "1 hour 42. The first finish, first time Link actually beats it. 1 hour and 42 minutes."

An hour and forty two minutes is slow for Zelda, and he says so immediately: "There is a ton that we can improve on this." The improvement loop that follows is the same pattern as the slime in the doorway, which is to say it is a human pointing at a symptom in plain language and letting the system find the cause.

What I'm doing to have it improve is say, "Hey, you are taking way too much time on these screens. Like, why are you doing that?" So it actually just kind of goes in, looks, and says, "Why are we taking so much time on this screen? Why are we taking so much time on that screen?" And it just does more brute force attacks. It just will keep trying and doing new things and trying new things until it actually improves that time based on all of those rules that we had.

Notice what the human is contributing there. Not a fix, not a price, not a code change. A pointer to where the waste is. The system already has the machinery to grind a room several hundred times; what it did not have was a reason to spend that grinding budget on screen 212 rather than spreading it evenly. Attention allocation is the scarce thing, and that is what the creator supplies.

COMPLETION TIME, POWER ON TO THE END, AS STATED IN THE VIDEO Run 1 1h 42m Run 2 57m Run 3 39m Final run 37m 02s human world record, 27m 40s 0 30 min 60 min 90 min A 2.8x improvement across the runs he names, ending 9m 22s short of the human record.
Figure 5. The four times stated in the narration. He also mentions an intermediate pass at about 37 minutes before the 37 minutes 2 seconds result, and the repository's final movie file is named run6, so there were more passes than the four he calls out. The world record line is the figure the project's own repository cites for the any percent No Up plus A category, and the repository is explicit that it does not claim to beat it.

The progression he reads out: "First run, hour 42. Second run, 57 minutes. Third run, 39 minutes. We got it down to 37 minutes, and then 37 minutes and 2 seconds. So that's currently where we stand."

The shape of that curve is the real story. The first improvement is enormous, nearly cutting the time in half, because the first run is full of outright waste. The second is large. After that it flattens hard, because what remains is not waste but optimization, and each additional minute costs a lot more grinding than the last. That is what a search against a fixed cost function looks like when it starts converging.

The thesis, stated at the end: the AI is the author, not the player

The closing argument is the most important sentence in the video, and it is easy to skate past because it sounds like a caveat. "As I mentioned earlier, you don't need AI to actually play this. AI did all the work to build these files and build these scripts to actually play the game on its own."

What exists at the end of ten days is not a model. It is a program. Twenty five thousand lines of Python, a Lua bridge, a pile of extracted knowledge files, and an input log, all of it ordinary deterministic code that runs on a machine with no network connection and no API key. Everything the model contributed happened in the authoring loop and is now frozen into source. "I'll put all these files in GitHub in the description if you want to actually download them yourself. You can have your own computer play a live version of Zelda."

And he is explicit that the authoring loop never stopped improving: "I'm sure that you could have your own AI bots if you were into that kind of thing, improve on it and actually get better and better, because every single time I had AI working on it, it got stronger and stronger every time through."

The glitch question, and the record he did not chase

Then the honest coda on the rule in the title. He does not claim the record, and he does say what he thinks would happen without the rule.

I have actually very little doubt that if I allowed it to use glitches, it would have done some even crazier stuff and beaten the world record by a human. That wouldn't surprise me one bit. I did watch the world record human run through and it was extremely impressive, don't get me wrong. So maybe it couldn't.

That is a sharp framing of what the constraint cost. Zelda speedrunning at the top level runs on movement exploits, and the human record category this project measures itself against still permits screen scrolling and block clips. Forbidding those while keeping frame perfect execution is a strange hybrid: superhuman precision inside subhuman rules. The 37 minutes 2 seconds result is best read not as "slower than a human" but as "the fastest anyone has gone under a rule set nobody competes under."

He then hands it to the audience, both the question and the next game: "If you guys wanted to see this kind of thing get done again, is there maybe a game that you would want to see that you couldn't beat as a kid, that maybe we can train AI to beat? And we can watch that happen too."

And the sign off, which is the best line in the video and belongs on a sticker: "I know half of you hate AI, half of you think it's amazing, and most of you think it's going to kill us. And we're all probably right."

The last 38 minutes: the run itself

Narration ends at roughly 10:50. Everything after it, about 38 of the video's 48 minutes, is the full final run playing out with no commentary at all: game audio and the Zelda soundtrack only.

A few things worth knowing before you scrub into it:

Key takeaways

Chapters

The creator's own twelve chapters all sit inside the narrated first ten minutes. The final entry is the run.

Notable quotes

I wanted to see if I could actually get AI to learn an old school video game like Legend of Zelda and have it actually learn the game start to finish, and not only learn it and beat it, but actually potentially beat a world record. Bears Gaming Den, stating the goal, 0:06

It was running six copies of Zelda, all in the same room, all trying the same fight over and over. How many frames it took, how many hearts it had left, whether it won. Ten days of this. Bears Gaming Den, on the opening montage, 0:30

AI actually never plays the game live. It can't. Link wants a decision 60 times a second and the language model takes seconds to think. Bears Gaming Den, on the constraint that shapes everything, 0:36

Hold these buttons for this many frames. Tell me the game state. Save a bookmark. Load a bookmark. That's the entire interface. Everything you're about to see is built out of those four things. Bears Gaming Den, 1:06

By the end of it, 25,000 lines of Python, and it kept a journal, 45 entries. Every one of them is a problem, what it tried, what went wrong, what it changed. Bears Gaming Den, 1:06

One important thing to note, though, is it never really looks at the screen. It really just looks at memory. Bears Gaming Den, on reading memory instead of pixels, 1:36

I said, you don't have to kill all the slimes, but you do have to make progress. That night's journal was called passive is losing. Standing still is never safety. Detour or kill what's in the way and keep moving. Bears Gaming Den, on the day two journal entry, 3:08

A bomb in hand is worth 220 frames because it opens a wall later that would otherwise be a detour. Bears Gaming Den, on pricing a bomb, 4:09

The coolest part of the whole build happened now. It's when it stopped reacting and it started looking ahead. Bears Gaming Den, on the switch to lookahead, 5:09

Those prices are basically Link's entire personality. Bears Gaming Den, 6:12

If you set being hit too high, he's never going to leave the doorway. If you price distance too high, he never attacks. Both those things happened in the run. Bears Gaming Den, 6:12

Every boss got the same treatment. Measure, don't assume. This journal entry is called read the manual. Research first, experiment second. Bears Gaming Den, on the boss testing pass, 6:42

For 160 frames, Link zigzagged between two Darknuts that were exactly one sword length away from each other with their sides exposed. Every swing it simulated came back as a miss. Bears Gaming Den, on the best bug in the run, 8:13

You don't need AI to actually play this. AI did all the work to build these files and build these scripts to actually play the game on its own. Bears Gaming Den, on what actually got built, 9:14

I have actually very little doubt that if I allowed it to use glitches, it would have done some even crazier stuff and beaten the world record by a human. That wouldn't surprise me one bit. Bears Gaming Den, on the rule in the title, 9:44

I know half of you hate AI, half of you think it's amazing, and most of you think it's going to kill us. And we're all probably right. Bears Gaming Den, signing off, 10:45

Resources mentioned

Everything the video points at, plus the project files it links in its own description:

Context this page added, for anyone who wants to go further: tool assisted speedrunning and speedrunning generally, 6502 assembly (the NES CPU, and the language the Zelda ROM is written in), save states, branch and bound (the mid run pruning of losing attempts), model predictive control (the closest formal name for the snapshot, simulate nine futures, pick one, replan loop), A* search (the shape of the tile pathfinder with a monster cost field), and random seeds (what gets varied between the sixty to ninety attempts per room). The repository credits Claude as the coding agent that did the authoring; the video itself says only "AI."

Where it stands

A few things are worth saying plainly once the rebuild is done.

By construction, this is a tool assisted run, and the repository says so. Savestates everywhere, 341 segments searched independently, sixty to ninety attempts per room, six emulators in parallel, losing attempts killed mid run. That is the definition of TAS, and the project's own documentation calls it "tool assisted speedrunning rather than human competition records." The "don't cheat" rule set is about what the inputs are allowed to exploit, not about whether the discovery process was assisted.

The 37 minutes 2 seconds and the 27 minutes 40 seconds are not comparable in either direction, and the project does not pretend otherwise. The human record permits screen scrolling and block clips, which this run forbids. The AI run permits frame perfect execution and thousands of retries per room, which a human does not get. Reading it as "AI is 9 minutes slower than a human" is as wrong as reading it as "AI nearly beat the world record." The honest version is the one in the repository: a run under a stricter rule set than any competitive category uses, verified more thoroughly than most submissions are.

The word "learn" is doing a lot of work. No model is learning Zelda in the machine learning sense. There is no policy network and no reward gradient. There is a coding agent reading logs and rewriting Python, a brute force search over random seeds, and a hand tuned cost function, all of it supervised by a person. That is a legitimate and genuinely impressive thing to build, and it is a different thing from what "AI learns to play a game" usually implies.

The two decisive unlocks came from the human. "You don't have to kill all the slimes, but you do have to make progress" produced the passive is losing rule and fixed the frozen Link. "You are taking way too much time on these screens, why are you doing that?" produced the grind that took an hour and forty two minutes down to thirty seven. In both cases the person supplied the pointer and the system supplied the fix, which is a more interesting division of labor than either "the AI did it" or "the human did it."

The evidence standard is the best part. A published input log, a search log per room, a SHA-1 verified byte for byte replay, and a movie file anyone can load into the same emulator version. Most AI plays a game demos ship a highlight reel. This one ships something falsifiable.

Small discrepancies, noted rather than smoothed over. The video says 45 journal entries; the repository has 46. The video names four run times; the repository's final movie is run6. The auto captions garble several proper nouns, which is why this page flags Patra, Gohma, Dodongo and the Darknuts explicitly rather than quietly fixing them. And the emulator is never named out loud in the narration; BizHawk 2.11.1 comes from the project files, not from the video.

Full transcript
[00:00:06] What's up, guys? I wanted to see if I could actually get AI to learn an old school video game like Legend of Zelda and have it actually learn the game start to finish and not only learn it and beat it, but actually potentially beat a world record. That was what I set out to do. And I just wanted to see what would happen. And there was a lot that kind of is crazy about this video and how it learns and all this stuff. So, it's pretty interesting. And yeah, hang on. Let's take a look. Uh, it was running six [00:00:36] copies of Zelda, all in the same room, all trying the same fight over and over. How many frames it took, how many hearts it had left, whether it won 10 days of this at the end of the run. The whole game is faster than most humans will ever play it. So, let's go through how it actually works from [music] the very first hour. First thing is AI actually never plays the game live. It can't. Link wants a decision 60 times a second and the language model takes seconds to think. So on day one, it builds itself a set of hands, a tiny script that lives [00:01:06] inside the emulator and listens to the network port. And on the other side, Python sending it four kinds of messages. Hold these buttons for this many frames. Tell me the game state. Save a bookmark. Load a bookmark. That's the entire interface. Everything you're about to see is built out of those four things. For 10 days, it wrote the player, watched what the player did, read the logs, and rewrote it. By the end of it, 25,000 lines of Python, and it kept a journal, 45 entries. Every one of them is a problem, what it tried, what went wrong, what it changed. A lot [00:01:36] in this video is just reading that journal back with the footage next to it. One important thing to note, though, is it never really looks at the screen. It really just looks at memory. Everything exists in 2 kilobytes worth of memory. One of them is Link's X, the other is Y, one is hearts, one is keys, and one is bomb. 12 of them are where the monsters are, and 12 of them are what they are. The journal from day one, the wiki's table of game modes has two values swapped, and the code is waiting for the wrong number. The map is in memory, too. Every screen is a grid of tiles, and the game keeps that grid [00:02:07] decoded. What the grid doesn't say is which tile Link can walk on, so it measures that. It walks Link into things. Walks up a column until he stops. Note which tile stopped him. Do it again for every tile type. And two facts came out of that. Link's collision box is 16x6 and starts three pixels below the Y value the game stores, which explains every stopping position. And collision is per 8x8 tile. All that goes into a file. And that being said, when it's all said and done, we know which tiles you can walk over and which tiles [00:02:38] will stop you. When a plan step fails and exactly one of the tiles that would have entered is unknown, then the tile is now solid forever on every screen and every future run. Walking is a path search. Monsters over eight pixel steps. Unknown tiles assume walkable until proven otherwise. And every monster on the screen adds a cost to squares around it. So, it will always choose the lowest cost value that it's assigned to each tile. This last part was a rule that had to be learned the hard way. And on day two, Link was just standing in the doorway because a slime was in its way. [00:03:08] And I said, "You don't have to kill all the slimes, but you do have to make progress." And that [music] night's journey was called passive is losing. And that's a rule that came out of it. Standing still is never safety. Detour or kill what's in the way and keep moving. One thing we didn't do is actually load like any kind of [music] world map to find level three. It literally just wandered around until it found it. So it would go left, right, up, down, and it would just do this a million times until it found the best route to go. 20 screens later, it had the route west, north, west three times, [00:03:38] south, east, up, and around the river. The emulator can save its entire state in about a millisecond. So the game is cut into 341 segments, roughly one per room. Every segment is a search. Load the bookmark at the door. Try the room with a different random seat. Log how many frames it took and how many cards came out. Do that 60 times or 90 across six emulators at once. Keep the best. Bookmark the end of it. Next room. To pick the best attempt, it needs one number per attempt. So everything gets converted into currency. Frame time is being minimized. So every frame costs a [00:04:09] point. But getting a heart when you're low on life actually gives you 100 points. So everything gets converted into currency whether it's good or whether it's bad. And every decision is based on that. A bomb in hand is worth 220 frames because it opens a wall later that would otherwise be a detour. These prices were guesses at first and every one of them got changed when the footage disagreed with it and any attempt that can no longer beat the best one gets cut off in the middle so the emulators don't waste time finishing losing [music] runs. The whole input log, every button, [00:04:39] every frame from power on is played back from a fresh emulator with no bookmarks. a human record is set live in one sitting. And that's exactly what we're going to do here, where AI is going to basically have this thing played out all in one sitting as well. In day five, it reached the last dungeon with 13 hearts, the best sword in the game, and zero bombs in a dungeon made for bombable walls. Fighting is actually where the first approach died. In level three, we entered a room with five of the red knights, and the handwritten reaction code lost it hundreds and hundreds of times. The journal's frame trace found [00:05:09] that the get out of its way logic and get behind it logic disagreed every frame. So Link just jittered in place until a red knight walked into it. The coolest part of the whole build though happened now. It's when it stopped reacting and it started looking ahead every eight frames and it took a snapshot of the game and then tried nine moves of the snapshot. Walk up, walk down, walk left, walk right or wait or swing up, swing left, swing right or swing down. Now it has all nine possible futures and has to pick one. So each frame gets turned into a number the same [00:05:40] way the attempts did. A list of things that matter each with a price added up. Time costs half a point for per frame. Dying minus 100,000 which just means never. Losing half a heart cost hundreds. And when Link takes a hit point off of an enemy, it's plus 60 and a kill is plus 150. Standing in front of a knight is minus 60. It knows which way the shield is facing because the bite is in memory, too. And there's a gentle pull towards the one place a sword can hit, a sword hit can happen beside or behind the nearest enemy. Those prices [00:06:12] are basically Link's entire personality. And all those prices had to be set by hand and argued with. But if you set something too high, if you set being hit too high, he's never going to leave the doorway. If you price distance too high, he never attacks. Both those things happened in the rule. And those fixes need to be h need to happen by making standing around have a high price as well. So all these things are fighting with each other to try to get the most logical way to complete the game as quickly as possible. Every boss got the same treatment. Measure, don't assume. [00:06:42] Petra, bombs, arrows, boomerang, all tested at zero damage. So GMA only an arrow and while the eyes open. Dongo eats bombs, two of them. Dragon, hacken slash. The flying glee here. You just have to kill all the outer ones before you can hit the inside one. This journal entry is called read the manual and that's a rule that came out of it. research first, experiment second. So, it reads walkthroughs and then it went one step further and read the game's own code. And there's a complete disassembly of Zelda on the cartridge itself. There's a kill counter and if you don't [00:07:12] get hit, every 10th kill is a guaranteed drop. Five rupees or if the killing blow is a bomb, bombs. So, the planners count to the ninth kill, holds that sword back and pulls out a bomb to take the 10th with it. So, hey, free bombs. Then it decoded the entire overworld map straight out of the ROM. All 128 screens, every tile, every cave, every shop, and what it sells, every secret, and what opens it. With that whole map, it could plan the entire game. Built a route that walks any leg of the map the [00:07:43] way Link actually walks it with the raft or the ladder, the mazes, the hidden road, the whirlwind, which takes you all over the place. It figures everything out here based on this map and actually optimizes its route based on this. Then it shuffled the order of everything. Which dungeon came first? When to fetch the white sword? When which of the three candle shop should it go to? What of the 100 rupee caves should should you go to and win? When to ride the wind? The best order for everything. But that run was [00:08:13] still weird. I'm watching it and looking at it. I'm like, this pathing is still strange. You're running around things that you shouldn't be running around. you should just be killing things instead of wandering around. Sometimes just sits there and seems to think or not know what to do. So, it's just weird. And for 160 frames, Link zigzagged between two dark nuts that were exactly one sword length away from each other with their sides exposed. Every swing, it simulated it that them came back as a miss. 1 hour 42. The first finish, first time Link actually beats it. 1 hour and 42 minutes. There [00:08:43] is a ton that we can improve on this. And what I'm doing to have it improve is say, "Hey, you are taking way too much time on these screens. Like, why are you doing that?" So, it actually just kind of goes in, looks, and says, "Why are we taking so much time on this screen? Why are we taking so much time on that screen?" And it just does more brute force attacks. It just will keep trying and doing new things and trying new things until it actually improves that time based on all of those rules that we had. So, first run hour 42, second run 57 minutes, third run 39 minutes. We got [00:09:14] it down to 37 sec 37 minutes and then 37 minutes and 2 seconds. So that's currently where we stand and what you're going to actually look at at the end of this video if you want to stick around and watch the entire playthrough on its own. So as I mentioned earlier, you don't need AI to actually play this. AI did all the work to to build these files and build these scripts to actually play the game on its own. I'll put all these files in GitHub in the description if you want to actually download them yourself. You can have your own computer [00:09:44] play a live version of Zelda. And I'm sure that you could have your own AI bots if you were into that kind of thing. Improve on it and actually get better and better cuz every single time I had AI working on it, it got stronger and stronger every time through. I have actually very little doubt that if I allowed it to use glitches, it would have done some even crazier stuff and uh beaten the world record by a human. I that doesn't that wouldn't surprise me one bit. I did watch the world record human run through and it was extremely impressive, don't get me wrong. So maybe [00:10:15] it couldn't, but I cannot see a world in which uh this could not be extremely extremely difficult to to try to stop [music] from cheating and world record stuff and all that all that kind of crap. Anyway, we're going to keep going here into the full final run. Hopefully you guys enjoy it. There's still a handful of mistakes in there, but overall it's extremely impressive. And if you guys wanted to see this kind of thing get done again, is there maybe a game that you would want to see that uh you couldn't beat as a kid that maybe we [00:10:45] can train AI to beat it? And uh we can watch that happen, too. So, all right. Enjoy it, guys. I know half of you hate AI, half of you think it's amazing, and most of you think it's going to kill us. And we're all probably right. >> [music] [music] [00:11:51] >> Heat. Heat. [00:12:52] Heat. Heat. N. [00:13:46] >> Heat. Heat. N. [00:14:20] >> [music] [music] [00:15:38] >> Heat. Heat. Heat. [00:17:28] >> Heat. Heat. [00:19:42] >> Heat up here. >> [music] [00:20:15] [music] >> Heat. Heat. [00:21:21] >> Heat. [00:25:56] >> Heat. Hey, heat. Hey, heat. >> [music] [music] [00:26:47] Heat. [00:27:20] >> Heat. [00:28:44] >> Heat. Heat. Heat. >> [music] [00:29:21] [music] [00:30:22] >> Heat. Heat. [00:31:13] >> Heat. Heat. N. [00:31:55] >> Heat. Heat. [music] Heat. Heat. [music] [music] [00:32:29] Heat. Heat. [00:33:03] >> [music] >> Heat. [00:33:43] Heat. [music] [00:34:30] >> Hey. >> [music] [00:35:38] >> Heat. Heat. Heat. Heat. [00:36:36] >> Heat. [music] Heat. [00:37:11] >> [music] [00:39:20] >> Heat. Heat. [00:40:36] Heat. Heat. [music] [00:41:14] >> Heat. Heat. >> [music] [00:42:47] >> Heat. Heat. [music] [00:43:36] >> Hey, heat. Hey. Heat up [00:44:11] Heat. Hey, heat. Hey, heat. [music] [00:44:55] >> [music] [music] >> Heat. Heat. Heat. Heat. [00:47:13] >> Heat. Heat. [music] >> [music] [00:47:47] [music] [bell]