At a glance
Bears Gaming Den spent ten days pointing a coding agent at The Legend of Zelda on the NES with one constraint attached: learn the whole game, beat it from power on, and do it without cheating. The twist in the setup is that the AI never touches the controller. Link needs a decision sixty times a second and a language model takes seconds to think, so instead of playing, the model spent those ten days writing a player: 25,000 lines of Python, a Lua script living inside the emulator, a priced decision function that scores nine possible futures every eight frames, and a route planner that decoded all 128 overworld screens straight out of the ROM. Along the way it kept a journal, 45 entries, each one a problem and what it changed, and the narration is largely that journal read back with the footage next to it. The first finish took 1 hour 42 minutes; the run that ships with the video takes 37 minutes and 2 seconds, and the final 38 minutes of the video are that run playing out, uncut and without commentary.
One note on the shape of this video before the rebuild starts. The narration runs from the opening to roughly the eleven minute mark, and all twelve of the creator's own chapters sit inside that stretch, ending at "Conclusion and full run." Everything after it, about 38 of the 48 minutes, is the complete final playthrough with game audio and music only. This page reconstructs the narrated section in full and then tells you exactly what is in the run section and where to look in it.
The pitch: ten days, six emulators, one rule
The opening is plain about the ambition. "I wanted to see if I could actually get AI to learn an old school video game like Legend of Zelda and have it actually learn the game start to finish," he says, "and not only learn it and beat it, but actually potentially beat a world record. That was what I set out to do. And I just wanted to see what would happen."
The montage that follows is the whole project compressed into one image: six copies of Zelda running at once, all in the same room, all trying the same fight over and over. What gets recorded off each attempt is three numbers. How many frames it took. How many hearts it had left. Whether it won. Ten days of that, and at the end of it, in his words, "the whole game is faster than most humans will ever play it."
That is the frame for everything below: not a neural network learning to press buttons, but a coding agent writing, watching, reading its own logs and rewriting, with a brute force search bolted underneath it to grind out the last frames.
Why the model never plays the game
This is the constraint that shapes the entire architecture, and it lands in the first minute. "AI actually never plays the game live. It can't. Link wants a decision 60 times a second and the language model takes seconds to think."
The NES runs at roughly 60 frames per second, which means the gap between what the game needs and what a language model can deliver is three or four orders of magnitude. There is no prompt engineering fix for that. So the model does not try to be the player. It becomes the author of the player. On day one it builds itself what he calls "a set of hands."
The set of hands: four messages and nothing else
The interface is almost comically small. "A tiny script that lives inside the emulator and listens to the network port. And on the other side, Python sending it four kinds of messages. Hold these buttons for this many frames. Tell me the game state. Save a bookmark. Load a bookmark. That's the entire interface. Everything you're about to see is built out of those four things."
The transcript does not name the emulator, but the project's own files do. The companion repository, bearsgaming-ui/AIBeatsZelda, names BizHawk 2.11.1 on Windows x64, with bridge.lua as the in emulator script ("hold buttons, read RAM, save/load a bookmark, over a socket") and zelda/emulator.py on the Python end. That choice is not incidental. BizHawk is the emulator the tool assisted speedrun community standardizes on, its Lua API ships a comm library with socket and HTTP support so a script inside the emulator can talk to an outside process, and it records input as a replayable .bk2 movie, which is what makes the "don't cheat" claim checkable later. Pick an emulator without scriptable sockets and savestates and none of this project exists.
Then comes the part that is the actual story. "For 10 days, it wrote the player, watched what the player did, read the logs, and rewrote it. By the end of it, 25,000 lines of Python, and it kept a journal, 45 entries. Every one of them is a problem, what it tried, what went wrong, what it changed. A lot in this video is just reading that journal back with the footage next to it."
That journal is the unusual artifact here. It is not a training log and it is not a diff history. It is a sequence of failure writeups, each one ending in a rule, and several of the rules get names: "passive is losing," "read the manual," "measure, don't assume." The repository ships the journal directory publicly, where the count is 46 entries rather than 45, presumably one more written after the video was cut.
It never looks at the screen
Here is the second structural decision, and it is the one that quietly makes the project tractable. "One important thing to note, though, is it never really looks at the screen. It really just looks at memory. Everything exists in 2 kilobytes worth of memory."
Two kilobytes is the NES's entire work RAM, and that is not a figure of speech. The console has 2,048 bytes of internal RAM, and the complete observable state of Zelda at any instant lives in there. So no vision model, no screenshot pipeline, no OCR of the heart meter. The player just reads bytes.
He walks the ones that matter: "One of them is Link's X, the other is Y, one is hearts, one is keys, and one is bomb. 12 of them are where the monsters are, and 12 of them are what they are."
That last pair is the important bit of design. Zelda tracks up to twelve active enemies per screen as two parallel arrays, one slot per enemy for position and one for type. Read both and you have a full tactical picture of the room in twenty four bytes: where everything is and what everything is. The public RAM map on Data Crystal documents the same addresses the project relies on, including Link's X at 0x70, Link's Y at 0x84, the heart byte at 0x066F (low nibble filled, high nibble containers), keys at 0x066E and bombs at 0x0658.
And then the first journal entry is about getting burned by exactly that kind of documentation. "The journal from day one: the wiki's table of game modes has two values swapped, and the code is waiting for the wrong number." Day one, hour one, and the lesson is already that a community reference can be subtly wrong and your code will sit there politely waiting for a state that never arrives. The game mode byte is the one that tells you whether you are on the title screen, in normal play, in a transition, or in a cave, so getting it wrong means nothing downstream fires at the right moment.
Measuring the walls instead of trusting them
The map is in memory too, and that part comes free: "Every screen is a grid of tiles, and the game keeps that grid decoded." But there is a hole in it. "What the grid doesn't say is which tile Link can walk on, so it measures that."
The measurement procedure is pure physical experiment, run inside a video game. "It walks Link into things. Walks up a column until he stops. Note which tile stopped him. Do it again for every tile type."
Two hard facts came out of it, and they are the sort of thing you only get by measuring:
- Link's collision box is 16 by 6 pixels, and it starts three pixels below the Y value the game stores. As he puts it, that "explains every stopping position." Link looks like a 16 by 16 sprite, but the thing that actually bumps into walls is a short wide strip near his feet, offset down from the number in RAM. Every off by a few pixels mystery in the pathfinder resolves once you know that.
- Collision is resolved per 8 by 8 tile, not per 16 by 16 block. Zelda's visible blocks are 16 pixels, but the solidity question is asked at half that resolution, which means a 16 pixel wide collision strip can be touching three different 8 pixel tile columns at the same time.
All of it gets written down. "All that goes into a file. And that being said, when it's all said and done, we know which tiles you can walk over and which tiles will stop you."
Then there is one more inference rule, and it is the cleverest small thing in the whole build: "When a plan step fails and exactly one of the tiles that would have entered is unknown, then the tile is now solid forever on every screen and every future run."
Read that again, because the word doing the work is exactly one. If a move fails and several candidate tiles are unknown, you have learned nothing you can safely act on, because you cannot tell which one blocked you. But if precisely one unknown tile was in the way, the failure uniquely identifies the culprit, and the conclusion is permanent and global: that tile type is solid, on every screen, in every future run. It is single candidate blame assignment, the same move a human does when a process of elimination leaves one suspect, and it turns a stream of ordinary failures into a monotonically growing map of the world.
Walking is a search, and passive is losing
With a solidity map in hand, movement becomes a graph problem. "Walking is a path search," he says, and then lists the rules it runs under. The captions garble one clause here, but the sense is clear from the surrounding list: the search steps in eight pixel increments, which is exactly the resolution collision is resolved at, so the search space and the physics agree.
The other two rules are the interesting ones.
Unknown tiles are assumed walkable until proven otherwise. The planner is an optimist. It will happily route through a tile it has never classified, find out the hard way, and then apply the permanent solid rule from the section above. Pair those two and you get a system that explores aggressively and never makes the same mistake twice, which is a much better trade in a game with a fixed finite tile set than cautious avoidance would be.
Every monster on the screen adds a cost to the squares around it. Not a hard no go zone, a cost field. "So it will always choose the lowest cost value that it's assigned to each tile." Danger is priced, not forbidden, which means the planner will walk right past a monster when the detour costs more than the risk, and will take the long way around when it does not. That single design choice is why Link can thread rooms rather than freeze at the first thing with teeth.
Except that on day two, he froze at the first thing with teeth. "On day two, Link was just standing in the doorway because a slime was in its way." A single slime in a doorway, and the cost field said the safest thing available was to not move at all.
The correction came from the creator, in plain English, and it is the clearest example in the video of how the human half of this loop actually works. "I said, you don't have to kill all the slimes, but you do have to make progress."
That night's journal was called passive is losing. And that's a rule that came out of it. Standing still is never safety. Detour or kill what's in the way and keep moving.
Three things are worth noticing about that exchange. The fix was not a code patch handed down, it was a goal stated at the level of intent. The AI turned it into a named rule in its own journal. And the rule is a general principle about the cost function, not a special case about slimes in doorways, which is why it stuck. "Standing still is never safety" is the kind of sentence that later shows up as a positive price on standing around, which is exactly where it reappears in the combat section below.
Finding level three by wandering
The other thing the project deliberately did not do is hand the AI the answer key. "One thing we didn't do is actually load like any kind of world map to find level three. It literally just wandered around until it found it. So it would go left, right, up, down, and it would just do this a million times until it found the best route to go."
The payoff is a route in compass directions, learned from scratch: "20 screens later, it had the route west, north, west three times, south, east, up, and around the river."
Twenty screens of overworld, found by exhaustive wandering, with no prior knowledge of where dungeons live. Keep that in mind, because an hour of project time later the same system will have the entire overworld decoded out of the ROM and will never need to wander again. The video shows both phases, in order, which is a more honest portrait of how this kind of build actually progresses than a finished demo would be.
The emulator as a time machine: 341 rooms, 341 searches
This is the section where the project stops being clever and starts being brutal. The enabling fact is a performance number: "The emulator can save its entire state in about a millisecond."
A millisecond per savestate changes what is possible. It means you can rewind reality thousands of times per minute, which means you can treat a video game as a sequence of independent optimization problems rather than one long fragile attempt. So that is exactly what happens. "The game is cut into 341 segments, roughly one per room."
And every one of those 341 segments gets ground:
Every segment is a search. Load the bookmark at the door. Try the room with a different random seed. Log how many frames it took and how many hearts came out. Do that 60 times or 90 across six emulators at once. Keep the best. Bookmark the end of it. Next room.
(The auto captions mangle two words in that passage: "random seat" is random seed, and "how many cards came out" is how many hearts came out. The numbers are his.)
That is the six copies in one room from the opening montage, explained. The parallelism is not there to make the AI smarter, it is there to make the clock cheaper: six instances times sixty to ninety attempts means every single room in Zelda gets played several hundred times, and only the best run of each survives into the final route. The repository notes that six instances want a reasonably modern CPU and that a full pass takes four to six hours, which is the real cost of this approach.
Everything becomes currency
A search needs something to sort by, and this is where the design gets its spine. "To pick the best attempt, it needs one number per attempt. So everything gets converted into currency."
The unit of that currency is the frame, because the frame is what the project is minimizing:
- Every frame costs a point. Time is the thing being minimized, so existing is expensive.
- Getting a heart while low on life pays 100 points. Survival is not a separate objective with its own constraint. It is bought and sold in the same money as everything else.
- A bomb in hand is worth 220 frames. And he gives the reason, which is the thing that makes the whole scheme work: "because it opens a wall later that would otherwise be a detour."
That bomb price is the clearest window into the method. A bomb is not valued because bombs are good. It is valued at exactly the number of frames of detour it will save you later, which means a resource in your inventory and a shortcut in the map and a near death heart pickup are all denominated in the same unit and can be traded against each other by a sort function. "So everything gets converted into currency whether it's good or whether it's bad. And every decision is based on that."
And the prices were never claimed to be correct. "These prices were guesses at first and every one of them got changed when the footage disagreed with it." That is the loop in one sentence: guess a price, run it, watch the footage, and when Link does something stupid, the price was wrong, not the game.
One more efficiency detail, and it is a classic: "any attempt that can no longer beat the best one gets cut off in the middle so the emulators don't waste time finishing losing runs." That is branch and bound. Once an attempt's accumulated cost exceeds the incumbent best, there is nothing left to learn from watching it lose politely, so it dies and the emulator slot goes to the next seed. With six instances and ninety attempts a room, killing hopeless runs early is the difference between a four hour pass and an overnight one.
The one rule: no cheating, and how it gets proved
The title's rule is not decoration, and the enforcement is specific. The final run is not a stitched together highlight reel of the 341 best segments. It is one input log, replayed cold.
The whole input log, every button, every frame from power on is played back from a fresh emulator with no bookmarks. A human record is set live in one sitting. And that's exactly what we're going to do here, where AI is going to basically have this thing played out all in one sitting as well.
So savestates are a practice tool, not a run tool. During the search they are everywhere, because that is the entire mechanism. In the run that counts, they are gone: the bookmarks were only ever a way to discover which buttons to press, and what survives is the button sequence itself, which has to work from frame zero on a fresh machine with nothing loaded.
The repository spells out the rest of the rule set that the narration only implies. Unmodified game, controller inputs only, no memory editing, no cheat codes, no glitches (specifically no screen scrolling and no block clips), one continuous session, no human input during the run, and the replay verified byte for byte against the original recording with a published SHA-1. The run ships as a BizHawk .bk2 movie plus a plain text inputs.txt with one line per frame, a search_log.txt showing every room's attempts, and a VERIFICATION.txt demonstrating that the same inputs in a fresh emulator reach the same RAM state. That is a higher bar of evidence than most human submissions carry, and it is the part of the project that makes the word "run" mean something.
Note what reading memory is not. The player reads RAM constantly, and that is not on the forbidden list, because reading is observation. Writing to RAM is what "memory editing" means, and that is forbidden. Any tool assisted run makes the same distinction.
Day five: thirteen hearts, the best sword, and zero bombs
In between the architecture, the video drops one perfect illustration of what route planning is actually for. "In day five, it reached the last dungeon with 13 hearts, the best sword in the game, and zero bombs in a dungeon made for bombable walls."
Thirteen hearts and the top tier sword is a strong position by any measure. Zero bombs at the door of a dungeon built around bombable walls is a soft lock in disguise: every wall that should have been a shortcut becomes a long way around, and the frame cost compounds across the whole level. The planner had optimized combat capability and ignored a consumable, which is precisely the failure that the 220 frame bomb price exists to prevent. This is the kind of thing the journal was for.
Where the first fighting approach died
Routing was solvable. Fighting was not, at least not the way it was first written. "Fighting is actually where the first approach died. In level three, we entered a room with five of the red knights, and the handwritten reaction code lost it hundreds and hundreds of times."
The red knights are Darknuts, which are the single most instructive enemy in Zelda for an AI to fail against, because they are armored on the front and can only be hit from the side or behind. Five of them in one room is a geometry problem disguised as a fight.
What killed it was not that the reaction code was dumb. It was that it was two pieces of reaction code. "The journal's frame trace found that the get out of its way logic and get behind it logic disagreed every frame. So Link just jittered in place until a red knight walked into it."
That is a textbook oscillation. One rule said move away from the thing in front of you; the other said move around behind it; those point in opposite directions; and every frame the winner flipped. The result was a Link vibrating on the spot while an armored knight strolled into him. And notice how it was diagnosed: not by watching the footage and guessing, but by a frame by frame trace in the journal that showed the two subsystems contradicting each other on every tick. The journal was not documentation. It was the debugger.
The best idea in the build: nine futures, every eight frames
Then the fix, and he flags it himself as the high point. "The coolest part of the whole build though happened now. It's when it stopped reacting and it started looking ahead."
The mechanism is simple enough to state in one breath. Every eight frames, take a snapshot of the game. From that snapshot, try all nine things Link can do:
- Walk up, walk down, walk left, walk right
- Wait
- Swing up, swing down, swing left, swing right
"Now it has all nine possible futures and has to pick one."
This is the move from a reactive controller to a planner, and it dissolves the oscillation problem by construction. Two rules can no longer disagree, because there are no longer two rules. There is one score, nine candidates, and an argmax. If "get behind it" and "get out of the way" point different directions, the scorer simply tells you which of those two futures is worth more, and Link does that one.
The eight frame cadence is the tuning knob that makes it affordable: at 60 frames per second it replans roughly seven and a half times a second, which is fast enough to react to a Darknut and slow enough that the simulation cost stays sane. The repository puts this logic in zelda/lookahead.py, described as combat planning with eight frame lookahead.
And the scoring is the same trick as before, applied one level down. "Each frame gets turned into a number the same way the attempts did. A list of things that matter each with a price added up."
| Priced thing | Price | Scope | What the price actually encodes |
|---|---|---|---|
| Every frame that passes | cost 1 | Attempt score | Time is the objective. Existing is expensive, so nothing is free. |
| A heart picked up while low on life | pays 100 | Attempt score | Survival is not a separate constraint. It is bought with the same money as speed. |
| A bomb in hand | pays 220 frames | Attempt score | Exactly the detour it saves later by opening a wall. A resource priced as a shortcut. |
| Every frame that passes | cost 0.5 | Combat lookahead | Fighting is scored at finer resolution than routing, so a half frame of advantage still registers. |
| Dying | cost 100,000 | Combat lookahead | In his words, "which just means never." Large enough that no payout can outbid it. |
| Losing half a heart | cost in the hundreds | Combat lookahead | Expensive but finite, so Link will trade health for time when the trade is good. |
| Taking one hit point off an enemy | pays 60 | Combat lookahead | Partial progress counts, so the planner commits to a fight instead of flinching. |
| Killing an enemy | pays 150 | Combat lookahead | More than the sum of its hits, which rewards finishing rather than chipping. |
| Standing in front of a knight | cost 60 | Combat lookahead | Darknuts block from the front. The shield's facing is read straight out of RAM. |
| Standing beside or behind the nearest enemy | a gentle pull | Combat lookahead | The only geometry where a sword hit actually lands. Encoded as attraction, not as a rule. |
| Standing around doing nothing | priced high | Both | The "passive is losing" rule from day two, converted into money. |
Two details in that list are worth pulling out.
The shield facing comes from memory, not from the sprite. "It knows which way the shield is facing because the byte is in memory, too." A vision based agent would have to infer a Darknut's facing from pixels. This one reads it. Which is the whole argument for the memory first design in one example: the thing that is hardest to see is trivial to look up.
The hittable side is a pull, not a rule. "There's a gentle pull towards the one place a sword hit can happen, beside or behind the nearest enemy." He does not write "always attack from behind." He adds a mild gradient toward the position where a hit is possible and lets it compete with everything else on the list. If getting behind the knight would cost half a heart and three hundred frames, the pull loses, and that is correct behavior.
Then the line that is the best summary of the entire project: "Those prices are basically Link's entire personality."
Tuning a personality, and the two ways it goes wrong
The prices did not come out of an optimizer. "All those prices had to be set by hand and argued with." And he is specific about the two failure modes, which are mirror images of each other:
- Price being hit too high and "he's never going to leave the doorway." If damage is catastrophic enough, standing still dominates everything, because standing still is the only action with zero risk. This is the slime in the doorway from day two, reappearing as a numerical pathology instead of a logic bug.
- Price distance too high and "he never attacks." If closing the gap is expensive enough, the planner keeps its distance forever and the fight never resolves.
"Both those things happened in the run." And the fix is not to lower the first two prices, it is to add a third: "those fixes need to happen by making standing around have a high price as well." You cannot make cowardice unattractive by making danger cheap. You make it unattractive by charging rent on inaction. That is the same rule the journal named on day two, now expressed as a term in the objective function.
His own description of the resulting system is better than any paraphrase: "All these things are fighting with each other to try to get the most logical way to complete the game as quickly as possible."
Every boss: measure, don't assume
Bosses got the same experimental treatment as walls. "Every boss got the same treatment. Measure, don't assume." Which, in practice, means firing every weapon in the inventory at each one and recording the damage number, instead of trusting a strategy guide.
The findings he reads off are the real Zelda 1 boss rules, and they are a nice demonstration of why assuming would have failed:
- Patra: bombs, arrows and the boomerang all measured at zero damage. The core of a Patra is simply not a valid target while its ring of satellites is alive, so every weapon reads as useless until the geometry changes.
- Gohma: an arrow, and only while the eye is open. Everything else is immune, and the window is timed, so this is a boss you cannot brute force with a sword at all.
- Dodongo: "eats bombs, two of them." Feed it, do not fight it.
- The dragon (Aquamentus, the level one boss): "hack and slash." The simplest one in the game, and the only one on the list where the obvious approach is the right approach.
- The flying one: "you just have to kill all the outer ones before you can hit the inside one," which is the outer pieces before the master piece.
A caption note, flagged once: the auto captions mangle the boss names here. "Petra" is Patra, "GMA" is Gohma, and "Dongo" is Dodongo, all of which are unambiguous from the stated weaknesses. The last one is read as "the flying glee here," which is either Gleeok, whose heads detach and fly, or a second pass at Patra, whose satellites orbit a core. The behavior he describes, outer before inner, fits Patra exactly, so that is the reading used above, with the ambiguity noted rather than hidden.
The real point of the section is the rule that came out of it, and it is the one that sounds least like an AI and most like an engineer who has wasted a week. "This journal entry is called read the manual, and that's a rule that came out of it. Research first, experiment second."
Which is worth sitting with for a second, because it is a reversal. The project's earlier posture was measure everything from scratch, because the wiki had its game mode table wrong on day one. Measuring every boss with every weapon is expensive, and most of the answers were already written down somewhere. So the rule became research first and then experiment to confirm, which is a better policy than either pure trust or pure measurement, and it took getting burned in both directions to arrive at it.
Reading the ROM: the ninth kill and the whole overworld
"Research first" starts with walkthroughs and then goes somewhere most people would not. "It reads walkthroughs, and then it went one step further and read the game's own code. And there's a complete disassembly of Zelda on the cartridge itself."
The game's logic is 6502 assembly sitting in a ROM that is small by any modern standard, and complete annotated disassemblies of it exist publicly. Which means the authoritative answer to "what does this game actually do" is readable, and anything a wiki says about it is a secondary source that can be wrong. Once you have decided to trust the code over the documentation, two things fall out immediately.
The first is a free bomb exploit that is not a glitch. Here is the mechanic, exactly as he states it: "There's a kill counter, and if you don't get hit, every 10th kill is a guaranteed drop. Five rupees, or if the killing blow is a bomb, bombs. So the planner counts to the ninth kill, holds that sword back and pulls out a bomb to take the 10th with it. So, hey, free bombs."
This is real and it is well documented in the speedrunning community as a forced item drop: the game runs a kill counter, a clean streak with no damage taken makes the tenth kill's drop deterministic, and the type of drop depends on what delivered the killing blow. Bomb the tenth enemy and you get bombs back. It is not an exploit of a bug, it is an exploit of the game's actual intended drop table, which is exactly why it survives the no glitches rule. And the behavior it produces is wonderfully specific: a planner that counts enemies, deliberately withholds a sword swing it could have taken, and swaps to a bomb for the tenth, purely to refill the resource it priced at 220 frames a few sections ago.
The second is the entire map. "Then it decoded the entire overworld map straight out of the ROM. All 128 screens, every tile, every cave, every shop and what it sells, every secret and what opens it."
That is the complete Zelda overworld, 16 screens across and 8 down, extracted as data rather than explored. Compare it to the level three search twenty minutes of project time earlier, where finding one dungeon took twenty screens of blind wandering. Same system, same goal, and the difference is entirely about where the knowledge came from.
Route optimization: shuffling the order of the whole game
With the map as data, the problem changes shape again. "With that whole map, it could plan the entire game."
First it had to teach the router how Link actually moves across the overworld, which is not a straight line: "Built a route that walks any leg of the map the way Link actually walks it, with the raft or the ladder, the mazes, the hidden road, the whirlwind, which takes you all over the place."
Every one of those is a different kind of edge in the graph. The raft only works at specific docks. The ladder only crosses single tile water. The mazes are the screens that loop you back unless you walk a secret sequence of directions. The whirlwind, summoned by the recorder, is a warp whose destination you do not fully control. A router that treats the overworld as an open grid will produce a route that is physically impossible, so the movement model has to encode all of it.
Then comes the part that is genuinely a planning problem rather than a pathing one. "Then it shuffled the order of everything."
- Which dungeon came first. The nine dungeons are not strictly ordered, and the items inside them unlock overworld shortcuts, so dungeon order and route length are the same question.
- When to fetch the white sword. It sits in a cave that will not hand it over until you have enough heart containers, so it is gated on progress elsewhere.
- Which of the three candle shops to visit. Same item, three locations, three different detours depending on where you already are.
- Which 100 rupee cave to use, and when. Money is a resource with a location and a timing, not a number.
- When to ride the wind. A free teleport is only free if you were going there anyway.
"The best order for everything." This is a combinatorial ordering problem over a graph with conditional edges, and the AI is solving it by search against the same frame denominated cost function that scores everything else in the project.
The run was still weird
Here is where the video earns a lot of credibility, because the next thing shown is not a victory lap. It is the creator watching a working run and being unsatisfied with it.
But that run was still weird. I'm watching it and looking at it. I'm like, this pathing is still strange. You're running around things that you shouldn't be running around. You should just be killing things instead of wandering around. Sometimes it just sits there and seems to think or not know what to do. So it's just weird.
And then the specific failure, which is the best bug in the video: "For 160 frames, Link zigzagged between two Darknuts that were exactly one sword length away from each other with their sides exposed. Every swing it simulated came back as a miss."
Unpack that. Two Darknuts, one sword length apart, sides exposed. That is the single most favorable arrangement those enemies can present: both of them in range, both of them vulnerable, no shield in the way. And Link stood there for 160 frames, roughly two and a half seconds of a game that is being optimized to the frame, shuffling between them.
The cause is visible in the design. The lookahead's hit prediction was returning a miss for every swing it tried, so the only positive term left on the board was the gentle pull toward the hittable side of the nearest enemy. With two equally valid nearest enemies in opposite directions, that pull tied, and a tie in an argmax over nine futures looks exactly like a coin flip every eight frames. Which is the same oscillation that killed the handwritten reaction code, reappearing one level up. Replacing two arguing rules with one score does not eliminate oscillation; it just moves it into the score, where a wrong simulation and a symmetric situation can still produce a Link who dithers.
The times: from 1 hour 42 minutes to 37 minutes and 2 seconds
Then the moment: "1 hour 42. The first finish, first time Link actually beats it. 1 hour and 42 minutes."
An hour and forty two minutes is slow for Zelda, and he says so immediately: "There is a ton that we can improve on this." The improvement loop that follows is the same pattern as the slime in the doorway, which is to say it is a human pointing at a symptom in plain language and letting the system find the cause.
What I'm doing to have it improve is say, "Hey, you are taking way too much time on these screens. Like, why are you doing that?" So it actually just kind of goes in, looks, and says, "Why are we taking so much time on this screen? Why are we taking so much time on that screen?" And it just does more brute force attacks. It just will keep trying and doing new things and trying new things until it actually improves that time based on all of those rules that we had.
Notice what the human is contributing there. Not a fix, not a price, not a code change. A pointer to where the waste is. The system already has the machinery to grind a room several hundred times; what it did not have was a reason to spend that grinding budget on screen 212 rather than spreading it evenly. Attention allocation is the scarce thing, and that is what the creator supplies.
The progression he reads out: "First run, hour 42. Second run, 57 minutes. Third run, 39 minutes. We got it down to 37 minutes, and then 37 minutes and 2 seconds. So that's currently where we stand."
The shape of that curve is the real story. The first improvement is enormous, nearly cutting the time in half, because the first run is full of outright waste. The second is large. After that it flattens hard, because what remains is not waste but optimization, and each additional minute costs a lot more grinding than the last. That is what a search against a fixed cost function looks like when it starts converging.
The thesis, stated at the end: the AI is the author, not the player
The closing argument is the most important sentence in the video, and it is easy to skate past because it sounds like a caveat. "As I mentioned earlier, you don't need AI to actually play this. AI did all the work to build these files and build these scripts to actually play the game on its own."
What exists at the end of ten days is not a model. It is a program. Twenty five thousand lines of Python, a Lua bridge, a pile of extracted knowledge files, and an input log, all of it ordinary deterministic code that runs on a machine with no network connection and no API key. Everything the model contributed happened in the authoring loop and is now frozen into source. "I'll put all these files in GitHub in the description if you want to actually download them yourself. You can have your own computer play a live version of Zelda."
And he is explicit that the authoring loop never stopped improving: "I'm sure that you could have your own AI bots if you were into that kind of thing, improve on it and actually get better and better, because every single time I had AI working on it, it got stronger and stronger every time through."
The glitch question, and the record he did not chase
Then the honest coda on the rule in the title. He does not claim the record, and he does say what he thinks would happen without the rule.
I have actually very little doubt that if I allowed it to use glitches, it would have done some even crazier stuff and beaten the world record by a human. That wouldn't surprise me one bit. I did watch the world record human run through and it was extremely impressive, don't get me wrong. So maybe it couldn't.
That is a sharp framing of what the constraint cost. Zelda speedrunning at the top level runs on movement exploits, and the human record category this project measures itself against still permits screen scrolling and block clips. Forbidding those while keeping frame perfect execution is a strange hybrid: superhuman precision inside subhuman rules. The 37 minutes 2 seconds result is best read not as "slower than a human" but as "the fastest anyone has gone under a rule set nobody competes under."
He then hands it to the audience, both the question and the next game: "If you guys wanted to see this kind of thing get done again, is there maybe a game that you would want to see that you couldn't beat as a kid, that maybe we can train AI to beat? And we can watch that happen too."
And the sign off, which is the best line in the video and belongs on a sticker: "I know half of you hate AI, half of you think it's amazing, and most of you think it's going to kill us. And we're all probably right."
The last 38 minutes: the run itself
Narration ends at roughly 10:50. Everything after it, about 38 of the video's 48 minutes, is the full final run playing out with no commentary at all: game audio and the Zelda soundtrack only.
A few things worth knowing before you scrub into it:
- It is the real run, in real time. The playable stretch spans roughly 11:00 to 47:47, which is about 37 minutes of wall clock, matching the stated 37 minutes 2 seconds result. Nothing is cut and nothing is sped up, which is the whole point given the one unbroken input log rule.
- The route is the optimized one. The dungeon order, the shop visits, the white sword pickup and the whirlwind warps are all the output of the route shuffling described above, so watching it is the only way to see what "the best order for everything" actually decided.
- He warns you about the quality. "There's still a handful of mistakes in there, but overall it's extremely impressive." The 160 frame Darknut dither is the class of thing to watch for.
- The auto captions in this stretch are noise, not dialogue. YouTube's caption track fills the run with repeated "Heat" fragments and
[music]markers. Those are the recognizer trying to transcribe Link's attack grunt and the game's sound effects. There is no speech in this section, so if you are reading the transcript at the bottom of this page, that is what those lines are. - It ends on a bell. The last caption event in the file is
[bell]at 47:47, which is the end of the run.
Key takeaways
- The model's job was authoring, not playing. A language model cannot act sixty times a second, so it wrote a program that can. The deliverable is deterministic code plus an input log, and it runs with no AI in the loop.
- The whole interface was four messages. Hold these buttons for N frames, tell me the game state, save a bookmark, load a bookmark. Everything else in the project is built on top of that.
- Read memory, not pixels. Two kilobytes of NES RAM holds the complete game state, including a Darknut's shield facing. Reading it is exact, instant and free; inferring it from the screen is none of those things.
- A one millisecond savestate turns a game into 341 independent optimization problems. That single performance fact is what makes six emulators times ninety attempts per room possible.
- Convert everything into one currency and the hard problems become sorts. Frames, hearts, bombs, danger and inaction all priced in the same unit, so a search can compare them. A bomb was worth 220 frames because that is the detour it saved.
- The prices are the behavior. "Those prices are basically Link's entire personality." Price damage too high and he never leaves the doorway; price distance too high and he never attacks. The fix for both is to charge rent on standing still.
- Measure what you can, but research first. Day one burned on a wiki with two swapped values, so the policy became measure everything. Measuring every boss with every weapon was expensive, so the policy became research first, experiment second.
- Reading the ROM beats reading the wiki. The game's own code gave up the forced drop on every tenth kill and the full 128 screen overworld, which replaced blind wandering with planning.
- Oscillation is the recurring failure, at every level. Two reaction rules disagreeing every frame, then two tied futures in the lookahead, then a planner that sat and thought. A single score does not prevent it, it relocates it.
- The constraint is what makes the result interesting. No glitches, no memory editing, no human input, one unbroken log from power on, verified byte for byte. Without that rule set the number would be meaningless.
Chapters
The creator's own twelve chapters all sit inside the narrated first ten minutes. The final entry is the run.
- 0:00 Project introduction
- 1:03 Interface and architecture
- 1:42 Memory and collision logic
- 2:47 Pathfinding and progress
- 3:42 Emulator automation
- 4:42 Fighting and tactical AI
- 5:17 Predictive planning
- 6:38 Boss mechanics
- 7:00 ROM and map analysis
- 7:57 Route optimization
- 8:39 Results and final runs
- 9:28 Conclusion and full run
- 10:50 Narration ends, and the complete unnarrated playthrough begins
- 47:47 The closing bell on the run
Notable quotes
I wanted to see if I could actually get AI to learn an old school video game like Legend of Zelda and have it actually learn the game start to finish, and not only learn it and beat it, but actually potentially beat a world record. Bears Gaming Den, stating the goal, 0:06
It was running six copies of Zelda, all in the same room, all trying the same fight over and over. How many frames it took, how many hearts it had left, whether it won. Ten days of this. Bears Gaming Den, on the opening montage, 0:30
AI actually never plays the game live. It can't. Link wants a decision 60 times a second and the language model takes seconds to think. Bears Gaming Den, on the constraint that shapes everything, 0:36
Hold these buttons for this many frames. Tell me the game state. Save a bookmark. Load a bookmark. That's the entire interface. Everything you're about to see is built out of those four things. Bears Gaming Den, 1:06
By the end of it, 25,000 lines of Python, and it kept a journal, 45 entries. Every one of them is a problem, what it tried, what went wrong, what it changed. Bears Gaming Den, 1:06
One important thing to note, though, is it never really looks at the screen. It really just looks at memory. Bears Gaming Den, on reading memory instead of pixels, 1:36
I said, you don't have to kill all the slimes, but you do have to make progress. That night's journal was called passive is losing. Standing still is never safety. Detour or kill what's in the way and keep moving. Bears Gaming Den, on the day two journal entry, 3:08
A bomb in hand is worth 220 frames because it opens a wall later that would otherwise be a detour. Bears Gaming Den, on pricing a bomb, 4:09
The coolest part of the whole build happened now. It's when it stopped reacting and it started looking ahead. Bears Gaming Den, on the switch to lookahead, 5:09
Those prices are basically Link's entire personality. Bears Gaming Den, 6:12
If you set being hit too high, he's never going to leave the doorway. If you price distance too high, he never attacks. Both those things happened in the run. Bears Gaming Den, 6:12
Every boss got the same treatment. Measure, don't assume. This journal entry is called read the manual. Research first, experiment second. Bears Gaming Den, on the boss testing pass, 6:42
For 160 frames, Link zigzagged between two Darknuts that were exactly one sword length away from each other with their sides exposed. Every swing it simulated came back as a miss. Bears Gaming Den, on the best bug in the run, 8:13
You don't need AI to actually play this. AI did all the work to build these files and build these scripts to actually play the game on its own. Bears Gaming Den, on what actually got built, 9:14
I have actually very little doubt that if I allowed it to use glitches, it would have done some even crazier stuff and beaten the world record by a human. That wouldn't surprise me one bit. Bears Gaming Den, on the rule in the title, 9:44
I know half of you hate AI, half of you think it's amazing, and most of you think it's going to kill us. And we're all probably right. Bears Gaming Den, signing off, 10:45
Resources mentioned
Everything the video points at, plus the project files it links in its own description:
- AIBeatsZelda, the project repository and the files he promises on screen:
bridge.lua(the in emulator script),zelda/emulator.py(the Python side),zelda/overworld.py(tile maps and pathfinding),zelda/lookahead.py(the eight frame combat planner),zelda/search.py(attempt ranking by frames, health and bombs),zelda/runner.py(segment management and verification),zelda/owmap.pyandzelda/owroute.py(the overworld decoded from ROM),route_planner.py,zelda/boss.py,zelda/combat.py,zelda/secrets.py, thejournal/directory of dated entries, and theknowledge/directory of learned patterns - The run artifacts:
zelda_ai_run6_37m02s.bk2(the BizHawk movie),runs/run6/inputs.txt(one line per frame),runs/run6/search_log.txt(every room's attempts), andruns/run6/VERIFICATION.txt(the SHA-1 replay check) - The Legend of Zelda on the Nintendo Entertainment System
- BizHawk, the emulator (source on GitHub), and its Lua function reference, including the
commsocket library the bridge depends on - Lua, the language the in emulator script is written in
- The Legend of Zelda RAM map on Data Crystal, the community reference whose game mode table had two values swapped on day one, and which documents Link's X at
0x70, Link's Y at0x84, hearts at0x066F, keys at0x066Eand bombs at0x0658 - A complete annotated disassembly of The Legend of Zelda, the kind of source that makes "read the game's own code" possible
- Forced item drops and the item drops chart at ZeldaSpeedRuns, which document the kill counter and the guaranteed tenth kill drop
- Zelda 1 boss mechanics and technical information at Red Candle, covering Patra, Gohma, Dodongo, Aquamentus and Gleeok
- Darknut, the red knight that killed the first fighting approach, and the full Zelda 1 enemy list
- The complete overworld map, all 128 screens, the thing the AI eventually decoded out of the ROM
- The Zelda 1 item list, covering the white sword, the blue candle, the raft, the stepladder and the recorder that summons the whirlwind
- The ZeldaSpeedRuns hub and the Zelda 1 speedrunning reference at Red Candle
- The human record he says he watched: the 27 minutes 40 seconds first quest any percent run, and the Guinness World Records entry for the category
- Bears Gaming Den, the channel
Context this page added, for anyone who wants to go further: tool assisted speedrunning and speedrunning generally, 6502 assembly (the NES CPU, and the language the Zelda ROM is written in), save states, branch and bound (the mid run pruning of losing attempts), model predictive control (the closest formal name for the snapshot, simulate nine futures, pick one, replan loop), A* search (the shape of the tile pathfinder with a monster cost field), and random seeds (what gets varied between the sixty to ninety attempts per room). The repository credits Claude as the coding agent that did the authoring; the video itself says only "AI."
Where it stands
A few things are worth saying plainly once the rebuild is done.
By construction, this is a tool assisted run, and the repository says so. Savestates everywhere, 341 segments searched independently, sixty to ninety attempts per room, six emulators in parallel, losing attempts killed mid run. That is the definition of TAS, and the project's own documentation calls it "tool assisted speedrunning rather than human competition records." The "don't cheat" rule set is about what the inputs are allowed to exploit, not about whether the discovery process was assisted.
The 37 minutes 2 seconds and the 27 minutes 40 seconds are not comparable in either direction, and the project does not pretend otherwise. The human record permits screen scrolling and block clips, which this run forbids. The AI run permits frame perfect execution and thousands of retries per room, which a human does not get. Reading it as "AI is 9 minutes slower than a human" is as wrong as reading it as "AI nearly beat the world record." The honest version is the one in the repository: a run under a stricter rule set than any competitive category uses, verified more thoroughly than most submissions are.
The word "learn" is doing a lot of work. No model is learning Zelda in the machine learning sense. There is no policy network and no reward gradient. There is a coding agent reading logs and rewriting Python, a brute force search over random seeds, and a hand tuned cost function, all of it supervised by a person. That is a legitimate and genuinely impressive thing to build, and it is a different thing from what "AI learns to play a game" usually implies.
The two decisive unlocks came from the human. "You don't have to kill all the slimes, but you do have to make progress" produced the passive is losing rule and fixed the frozen Link. "You are taking way too much time on these screens, why are you doing that?" produced the grind that took an hour and forty two minutes down to thirty seven. In both cases the person supplied the pointer and the system supplied the fix, which is a more interesting division of labor than either "the AI did it" or "the human did it."
The evidence standard is the best part. A published input log, a search log per room, a SHA-1 verified byte for byte replay, and a movie file anyone can load into the same emulator version. Most AI plays a game demos ship a highlight reel. This one ships something falsifiable.
Small discrepancies, noted rather than smoothed over. The video says 45 journal entries; the repository has 46. The video names four run times; the repository's final movie is run6. The auto captions garble several proper nouns, which is why this page flags Patra, Gohma, Dodongo and the Darknuts explicitly rather than quietly fixing them. And the emulator is never named out loud in the narration; BizHawk 2.11.1 comes from the project files, not from the video.


