At a glance
Julian Eisenkirchner opens by burying the workflow he used for years: import every A-roll clip into Adobe Premiere Pro, cut out the mistakes and the pauses by hand, go hunting for B-roll on a stock site, download it, lay it over the cuts to hide them, go hunting again for music and sound effects, and then, because subtitles were "either a manual nightmare or I had to export my video and then use a third-party tool," export the finished file into something else entirely just to burn in captions. In ten minutes he replaces all of it with a single browser tab.
The tool is Riverside, the online recording platform most people know as a podcast studio. Julian says he has used it for podcast recording for years, and this video, which Riverside sponsored, is specifically about the editing half of the product rather than the recording half. His framing is blunt: no downloads, no third-party tools, no export-and-reimport, everything in one tab.
The demo uses a control. He takes one raw talking-head clip, the same clip he had already edited in Premiere Pro under a tight deadline (where, he admits, "I was just cutting out the mistakes and that was basically the video, so not the most engaging video on the entire planet"), and uploads that identical file to Riverside. Everything that follows is measured against that lazy Premiere cut.
What the page below rebuilds, click by click and prompt by prompt: the four short-form clips with virality scores that Riverside generates before you touch anything, plus the hooks, social posts and snapshots that come with them; the text-based editor where you cut video by deleting words; the co-creator prompt box and the three exact prompts he types into it; the AI tools panel and auto-generate B-roll; the music, sound effects and caption presets; the multi-clip workflow driven by one sentence; and the part he calls "welcome to 2026," describing a B-roll shot in plain English and having Google Veo 3 generate it inside the editor. Every tool, every setting, every prompt verbatim, and every claim he makes about the time it saves.
The workflow he is burying
The cold open is a list of grievances, and it is worth keeping in full because every item on it comes back later as something Riverside claims to absorb.
He "used to spend absolute bloody ages editing and perfecting my videos on my laptop and in Premiere Pro." The sequence went: import all the A-roll clips. Cut out all the mistakes and all the pauses. Then go searching for B-roll clips, download them, and lay them over the cuts "so that I can hide my mistakes and the cuts." Then search for music and sound effects. And subtitles were not even in that list, because subtitles were their own separate problem: "to be honest, either it was a manual nightmare or I had to export my video and then use a third-party tool."
That last grievance is the one that shapes the whole video. Every other step at least happened inside one application. Captions forced a round trip out of the editor and back, which is exactly the friction the "one browser tab" pitch is aimed at.
He states the promise plainly at 0:29: he is going to show an AI-based online video editor, how it works, "everything that can be done, and ultimately we're going to find out, should you give it a try or not." He answers his own question immediately, which tells you the shape of the video: "I can tell you you should definitely give it a try, because you can try it out for free."
At 0:45 he does the channel intro. He calls this "an absolute game-changer to my content creation workflow," and describes the audience he is talking to: people "interested in making better videos, turning your knowledge into a sellable digital asset." That is the channel's beat, and it explains the angle of the whole tutorial. This is not a colorist's or an editor's video. It is a video for someone whose bottleneck is publishing volume.
Riverside, and the half of it nobody talks about
At 1:07 he introduces the tool on screen and next to him: Riverside. He is upfront that his audience already knows it, and knows it as the wrong thing. "You probably know it as this online podcasting tool where you can maybe also record a video."
The pivot: "today we're not going to talk about the podcasting side of things, but we're actually going to talk about editing videos online without needing any third-party tools or without needing to download anything."
That framing is the entire thesis, and it has two halves that are easy to run together. The first half is "online," meaning no local application, no render farm, no laptop fans. The second half is "without needing any third-party tools," meaning no Epidemic Sound tab for music, no Pexels or Storyblocks tab for B-roll, and no separate captioning app. The second half is the harder claim and it is the one he spends the rest of the video demonstrating.
The control: one clip, two edits
At 1:31 he sets up the comparison honestly, which is the best part of the video's structure.
He has uploaded "a raw clip of me, you know, sitting in this very spot and talking about a topic." He had already edited that same clip in Premiere Pro. And he tells you exactly how much care that Premiere edit got: "I had a pretty tight deadline, so I was not putting in a lot of effort. So, what I ended up doing, I was just cutting out the mistakes and that was basically the video. So, not the most engaging video on the entire planet."
Then: "I have also uploaded the exact same clip into Riverside."
So the benchmark is not a polished professional edit. It is the edit a busy creator actually ships when the deadline is tight, which is a fair and more useful baseline than a showreel. The question the rest of the video answers is whether the automated pass beats the rushed manual pass, not whether it beats a good one.
What one upload produces before you click anything
This is the part of the demo that does the most work in the video, because it happens without any user action at all.
He scrolls down on the uploaded file's page: "just by scrolling down, I can already see that I have four short-form clips cut out of my main piece of content before I even have started."
The clips, and the virality score
Each of the four clips carries a score. He calls it "a score here, like a virality score," and Riverside uses it to rank them: "Riverside is telling me this is the top choice." He does not claim to know how the score is computed, and neither does the video. It is presented as a ranking signal, not a prediction.
He plays the top-ranked clip back inline and you hear the raw audio of his own footage, the phrase "failing and failing," and he narrates over it: "the difference between failing and failure. So, very good."
Around each clip he gets a small control surface:
- A description of the clip, generated with it.
- An edit path, if you want to take that clip further rather than accept it.
- A thumbs up and thumbs down, which he calls out specifically as a training signal: "I can say I like it, I don't like it, so that the AI is actually learning."
- Three more options beyond the top choice, which is where the count of four comes from.
He repeats the point twice because it is the point: "even before I started off doing anything at all, this is just done automatically just by uploading a raw clip."
Hooks
Below the clips is a hooks section. He gets four of them, playable in place, and the workflow is the same: listen and "decide if I like one of these." These are candidate opening lines pulled out of the footage, which for short-form is the single highest-leverage second of the video.
Posts
Next, written posts. His explanation of the mechanism is the useful bit: "it's basically just transcribing everything and turning these into posts." Everything downstream in this product runs on one transcription pass, and the posts are simply another consumer of it. That single observation explains why captions later cost him one click and why the co-creator can answer questions about content it was never shown.
Snapshots
Finally, stills pulled from the footage. He names the uses directly: "a couple of really cool snapshots that I can use for a thumbnail, for example, or for a LinkedIn posting, an Instagram posting, whatever it might be."
At 3:03 he closes the block with the three exits available on any of these assets: "if I want to use this, I can just download it, or I could edit it even further, or I could also share it straight away."
Opening the editor
At 3:09 he moves from the asset view into the editor proper. The mechanic is one click: "all I need to do is just click on edit, and then the editing window is going to open up in parallel." The word "parallel" is doing real work there. The editor opens alongside rather than replacing, so the clip page and its generated assets stay where they were.
The text-based editor: cut video by deleting words
The editor's left side is the full transcription. He describes what that enables in one sentence, and it is the sentence that defines the whole category of tool: "this is a text-based editor, so I could, in theory, just cut out a couple of things just by selecting the words I don't like. I don't need to watch anything back at all."
That last clause is the actual saving. Not the click count, the scrubbing. Traditional editing requires you to watch your own footage in real time to find the bad take. Text-based editing turns that into reading, which is several times faster than playback and does not require headphones.
He is candid that the text editor alone is not the sell. "But this is very basic. I can also add a couple of clips if I wanted to." Then he names what he thinks the real feature is: "I think the most important thing and the coolest part is actually the co-creator and the AI tool."
The co-creator: editing by prompt
At 3:42 the video reaches its center. The co-creator is a prompt box inside the editor, and Julian's description of how to think about it is deliberately plain: "it's basically just as if I was using any type of AI tool, I could just give it a prompt."
He types three prompts in a row. All three are worth having verbatim, because the phrasing is the interface.
Prompt one: build me a short
create a ready-made vertical 30-second reel from the most important part
He submits it and narrates the wait: "now in the background, the AI is doing its magic. It's going through the transcript, and within just a couple of seconds, I'm going to have my viral reel ready-to-post video."
Note what the prompt carries. Three separate instructions are folded into one sentence: an aspect ratio (vertical), a duration (30 seconds), and an editorial judgment (the most important part). The third one is the interesting request, because it is the only one a human normally has to make.
At 4:12 the result arrives, and he times it: "after like 15 seconds or so, I now have my video." He plays it back, says "it's about just getting started," and lists the exits: "I can also download it straight away or just use it as it is."
Fifteen seconds is the only hard number he attaches to any single operation in the entire video, so it is worth holding onto.
Prompt two: write the copy
give me a description for this
His explanation of why this works is the same architectural point from earlier, stated a second time: "Again, it has all of the information. It has all of the transcripts, and you basically just need to tell it what you would like to get. It's as simple as that."
This is the co-creator behaving as a plain text assistant rather than an editor. The video does not linger on it, but it demonstrates that the prompt box is not restricted to timeline operations. It writes as well as it cuts, because both actions read from the same transcript.
He closes the pair with a claim about the category, not the product: "I think it has never been any easier to create any type of short-form content from a long-form piece of content."
Prompt three: remove pauses and filler words
remove pauses and filler words
This is the one that most directly replaces a specific chunk of the old Premiere workflow. He describes the pass as it runs: "it's going through my transcript. It's going through the video, and it's just going to cut out all of the oohs and ahs and ums, and all of the pauses."
The result: "within just a couple of seconds, it has cut out all of the pauses and the filler words, and I don't need to do this manually."
Notice the ordering in his own description. The transcript first, then the video. Filler word removal is a text operation whose result is applied to frames, which is exactly the model Figure 2 shows. It also explains why it can finish in seconds on a long file: nothing is being re-encoded to find the ums, the ums were already located when the file was transcribed on upload.
Why he picked Riverside: the AI tools panel
At 5:05 he does something a sponsored video does not have to do, which is acknowledge the competition: "everything we have covered so far has been pretty cool, but I know there are also other tools out there that can do similar things."
He is right, and the honest reading is that clip generation, text-based editing and filler removal are now table stakes. Descript, Opus Clip, Submagic, Vizard and CapCut all do some subset of it. So he sets up the differentiator: "let me blow your mind right now. And this is also one of the reasons like why I personally decided to go with Riverside."
The differentiator lives in a panel he had not shown yet. Under AI tools, below the co-creator, is a list of named operations you can run without writing a prompt at all. Two ways in, same engine: "you could go again into the co-creator and just type in what you want, but I also wanted to show you that down here within AI tools, you can also see, like, you know, other options as well."
Auto-generate B-roll
His favorite item in that panel: auto-generate B-roll.
He frames the pain first, and it is the single most laborious item on the original grievance list: "if I want to add B-roll clips, again, in the past, either I had to film them myself or I had to go to platforms that offer, you know, stock videos, and then I had to download and insert into Premiere Pro."
Three separate steps there, and two of them involve leaving the editor. The replacement is two clicks. "What I can do here is just click on auto-generate B-roll, click on apply."
What happens then: "it's going to go through the entire transcript, and wherever it thinks it's going to make sense to just add some B-roll clips, it's going to do all of that for me."
That is a meaningful shift in what the tool is deciding. Earlier features executed instructions you gave. This one makes an editorial call about where B-roll belongs, which is a judgment about pacing and emphasis. He accepts it without qualification here, though he walks it back into a "check it" caveat a few minutes later.
He also drops the forward reference to the feature the last third of the video builds to: "There is, you know, a library in the background, and I could even create videos with AI right directly within Riverside, which is just absolutely mind-blowing."
The B-roll pass runs for a while, and he uses the wait to cover two other tabs, which is a nice piece of demo pacing: the generation time becomes the segue.
Music and sound effects, without leaving the tab
At 6:05, while the B-roll generates, he opens the audio tab. "You can also add music and sound effects right directly within Riverside. All we need to do is just go here to audio."
What is in there:
- Recently used. The tab opens on "the ones that we recently have been using," which matters more than it sounds for anyone building a consistent channel sound.
- Categories. He names three by example and gestures at more: "we can search for different, you know, styles of lifestyle, news and politics, education, and so on and so forth."
- Your own uploads. "We could also upload our own clips if we had a few that we like." So the library does not lock you out of tracks you already license elsewhere.
He auditions a few options out loud and picks by feel: "what are we liking the most, and what would we like to use for this very video." His choice for this edit is education, used for the intro.
Adding it is one control: "I can just click on this plus icon. You can see I already have it in here, and then we're basically good to go."
Captions: the grievance that started the video
At 6:56 he returns to the item he singled out in the cold open as the worst of the lot. The fix is anticlimactic on purpose.
Two routes in, again: "if I wanted to add captions to this video, I just click on captions, or I could also go to the co-creator and tell it to add captions."
And then the reason it is nearly instant: "since everything is already transcribed, basically all I need to do is just pick a preset that I like the most, click on it, and then it's going to add it automatically."
The presets vary by tone rather than by feature. "If you want something a little more playful, if you want something a little, you know, less playful and more serious, basically, you're going to find something that you are going to like." He picks one and moves on.
Set against the opening complaint (a manual nightmare, or an export into a third-party tool), this is the cleanest single win in the video. The transcription that happened on upload has now paid for itself four separate times: clips, hooks, posts, and captions.
The claim: ten minutes instead of hours
At 7:25 he takes stock, and he repeats himself in the excitement, which is how you can tell it is unscripted: "just with the simple tap of a button, and just with a tap of a button, I now already have music, I have my captions, I already have B-roll, and I have already cut out all of the mistakes and the ums and ahs."
Then, at 7:38, the headline number: "this now has taken me like, I don't know, 10 minutes, where usually this was a task of like hours spending in the editing room."
That is his claim, stated with his own hedge included. The "like, I don't know" is doing honest work, and it is the reason the number is worth quoting rather than dismissing. What follows is the task-by-task ledger behind it.
| The task | His Premiere Pro workflow | Inside Riverside | Where |
|---|---|---|---|
| Get the footage in | Import all the A-roll clips into a local project on the laptop | Upload the raw file to the browser | Upload |
| Cut mistakes and pauses | Scrub the timeline, find every bad take, mark in and out by hand | Prompt remove pauses and filler words, or select the words in the transcript and delete them | Co-creator, or text editor |
| Find B-roll | Film it yourself, or search a stock video platform, download it, drag it into Premiere, lay it over the cut | Click auto-generate B-roll, click Apply, and it decides where B-roll belongs from the transcript | AI tools panel |
| B-roll that does not exist yet | Shoot it, or settle for the closest stock match | Describe the shot in a sentence and generate it with Veo 3 or Pika 2.2 | Videos tab |
| Music and sound effects | Search a separate library, license, download, import | Browse by category (lifestyle, news and politics, education), audition, click the plus icon | Audio tab |
| Subtitles | "A manual nightmare", or export the finished video and run it through a third-party tool | Click a preset. The file was transcribed on upload, so there is nothing to generate | Captions tab |
| Short-form clips | Re-edit the long video into verticals by hand, once the main edit is done | Four already exist with virality scores, before the first click | Automatic on upload |
| Hooks, descriptions, posts, thumbnails | Write them yourself, and pull stills manually | Four hooks, social posts and snapshots generated from the same transcript | Automatic, or co-creator |
| His stated total | "Hours spending in the editing room" | "Like, I don't know, 10 minutes" | His claim, at 7:38 |
Multiple clips, one sentence
At 7:46 he covers the case his single-clip demo has been dodging: a real edit usually has more than one take.
The mechanic: "you can actually upload all of the clips in here. You just go to your media, upload more clips, and then again, all of the clips are going to be transcribed."
Then a single prompt does the whole assembly. The one he suggests is deliberately vague, and that is the demonstration:
make this video look more interesting in every way possible
His description of the result: "it's going to add the B-rolls for you. It's going to add the music, transitions, and everything will be done for you."
That is the most ambitious claim in the video, and it is also the only one he does not show finishing on camera. He describes the outcome rather than playing it back.
The caveat he volunteers
Immediately after, at 8:11, and without being prompted by anything, he pulls back:
"Of course, I would still recommend to not let it do everything for you and never look back at it. I would still recommend to take your time, check if the B-rolls are right, check if the music is good, check if everything is still inside, if the information is there."
Three specific review checks, and the third one is the sharpest: confirm the information is still there. An automated pass that cuts pauses, cuts filler, and reflows a multi-clip edit can remove a clause that mattered, and a text-driven cut will do it silently because it looks correct in the transcript.
Then his own hit rate: "from experience, I can tell you in nine out of 10 times it does an absolutely fabulous job."
Nine out of ten is a good number for a first pass and a terrible number to publish on unreviewed, which is exactly why he says both sentences together.
The B-roll lands
At 8:29 the auto-generated B-roll pass from six minutes earlier finishes and he scrolls the timeline: "a couple of B-roll clips, so these parts right here have been added to my edit."
If the automatic pass is not enough, there is a manual escalation path in the videos tab, and it has two levels.
Level one: transcript-driven suggestions. "In case I wanted to add even more, in case I'm not super happy with what I got, I can also go down here to videos, and automatically, again based on the transcript, Riverside is going to give me a couple of recommendations on what I could use." Same index, different surface. The transcript that drove the automatic placement also drives the shortlist you browse by hand.
Level two is the one he has been building to all video.
Welcome to 2026: generating B-roll that does not exist
At 8:48 he delivers the line the whole thing is titled after: "this is absolutely mind-blowing and welcome to 2026, I can also just go in here and describe the video I want to have, and then AI is going to generate something for me."
His prompt, verbatim:
I want to have a B-roll clip of a man sitting with his laptop on the beach working and enjoying life, and the sun is shining
That is not a stock library search string. A stock search would be "man laptop beach." The prompt is written as a scene description with a mood attached, which is how you write for a video generation model rather than for a keyword index.
The model picker
Before generating, he shows the choice. Riverside gives him two engines inside the editor:
- Pika 2.2
- Google Veo 3 (he calls it "V03" on screen and "VO3" out loud, which is how everyone says it)
His pick and his reason are one sentence: "I like Veo 3 a lot, so I'm just clicking here on generate."
He comes back to it after the clip lands, and this is his only stated technical opinion in the video: "the technology in the background is Google Veo 3, in my opinion, one of the best AI video generation tools."
The result
"Within just a couple of seconds, I have my clip right here."
It drops into the editor as an asset like any other, and it needs the same handling as any other: "I can also, like, change the scale, so it actually fits." Worth flagging, because it is the one moment in the video where a generated asset does not arrive perfectly formed. Aspect ratio and framing are still your problem.
If you go to reproduce this, note that the interface has moved since June. Riverside's current documentation puts generation behind Stock in the editor toolbar rather than a Videos tab: click Stock, click the Stock textbox at the bottom of the panel, type your prompt, toggle Add audio on or off, hit Generate, and confirm the media usage agreement. The clip lands in Uploads, and the three dots menu gives you Retry and Delete. From there you either drop it in as a video overlay or insert it straight into the timeline. The Add audio toggle is the one control the video does not show, and it is the interesting one, because native synchronized audio is the whole reason to pick Veo over the alternatives.
He describes what he got: "this is the video of a man sitting on the beach enjoying online business and how you can really succeed."
And then the best twenty seconds of the video happen. He starts narrating over the generated clip, gets distracted by what is actually on screen, and the audio from the clip and his own commentary collide into "'cause you know, we just released ... life, drinking something. So, yeah." He recovers with "yeah, it's actually really cool" and moves on. It is left in the edit, which is a small and appropriate joke in a video about a tool that removes exactly this kind of thing.
The wrap
At 9:47 he closes on the outcome rather than the features: "the entire editing process just is so, so much faster. It's a lot less of a hassle. So, now I can create a lot more content and it's still very engaging."
Read that carefully, because it is a more specific claim than "editing is faster." The unit of value he is optimizing is content volume at a fixed engagement level. He is not claiming better videos. He is claiming the same quality bar, reached more times per week. For his audience, which he defined at 0:59 as people turning knowledge into sellable digital assets, that is the right metric.
And the second half of the payoff, which is easy to miss because it arrives so early in the workflow: "the very cool thing is that I also have lots of short-form clips that I can use at the same time." One upload produced both the long-form edit and the distribution assets for it.
The offer
At 10:02, the sponsor read, stated plainly: check the first link below the like button, and "with the code Julian, you can get one month completely for free. So, you can try everything out when it comes to recording podcasts or editing your videos just as you have seen in this video."
The link in the description is his Riverside creators link, and the code is JULIAN. He closes with "huge thanks to Riverside for supporting the channel," which is where the sponsorship gets its on-camera disclosure.
Every prompt he typed, in one place
The video's real teachable content is the prompt set, so here is all of it verbatim, with what each one produced and every timing he stated on camera.
| Prompt or control | Where you enter it | What it produced | His stated time |
|---|---|---|---|
| (none, just upload) | Drop a raw file in | Four short-form clips with virality scores and a ranked top choice, four hooks, social posts, and snapshots for thumbnails | Ready before he clicked anything |
| create a ready-made vertical 30-second reel from the most important part | Co-creator | A finished vertical short, playable in the editor, downloadable as is | "After like 15 seconds or so" |
| give me a description for this | Co-creator | Written copy for the clip, drawn from the transcript it already holds | Not timed on camera |
| remove pauses and filler words | Co-creator | Every um, ah and pause cut from the transcript and the video together | "Within just a couple of seconds" |
| auto-generate B-roll then Apply | AI tools panel (no typing) | B-roll inserted wherever the transcript suggested it belonged | Long enough to cover two other tabs |
| make this video look more interesting in every way possible | Co-creator, after uploading multiple clips to Media | B-roll, music and transitions across the whole multi-clip edit | Described, not shown finishing |
| I want to have a B-roll clip of a man sitting with his laptop on the beach working and enjoying life, and the sun is shining | Videos tab, generate, with Pika 2.2 or Veo 3 selected | A brand new generated clip dropped into the editor, needing a manual scale adjustment to fit | "Within just a couple of seconds" |
| Pick a caption preset | Captions tab, or ask the co-creator | Styled burned-in captions across the video | One click, nothing to generate |
| Pick a music bed | Audio tab, category browse, plus icon | Licensed music on the timeline (he chose education for the intro) | One click |
What the stack costs, which the video never says
The one number the video leaves out is the price. Julian mentions cost exactly twice, both times as an offer rather than a figure: "you can try it out for free" in the cold open, and "with the code Julian, you can get one month completely for free" at the end. So here is the ledger he skipped, read off Riverside's own pricing page.
Almost every AI feature demonstrated in this video sits on the Pro tier. Riverside's Pro listing names them close to how Julian shows them: "unlimited text-based editing," "AI editing and repurposing agent," "AI B-roll, eye contact, remove silences and filler words," "unlimited transcriptions," and "Magic clips and show notes." The free tier does include editing, but caps you at 720p output with a Riverside watermark on everything, which rules it out for the published-video use case the whole video is about.
The one exception is the feature that opens the video. Magic Clips is free on every plan, watermark and all. Riverside's own page says so in as many words: "Magic Clips is free but feels like a Hollywood-grade producer." So the single most impressive moment in the demo, four scored short-form clips waiting for you before you click anything, is the part you can try without paying.
The part the video does not mention at all: AI credits
There is a second currency underneath the subscription, and the video never says a word about it. Riverside's help center is explicit: "AI credits are used to power premium AI features on Riverside, such as AI B-roll and AI translation." Credits are available on Pro and up, and two details in that document matter more than the rest:
- "AI credits are purchased separately from your subscription and do not renew monthly." They are a top-up, not an allowance that refills. You buy a pool, you spend it, and when it runs out you buy more.
- Only account owners and admins can purchase them, which is a real friction point on a team.
Riverside's own list of what consumes credits is short: AI B-roll and AI translation. Everything else Julian demonstrates, the Co-Creator prompts, Magic Clips, text-based editing, filler removal, captions and the music library, is flat-rate once you are on the plan. So the ledger is simpler than it first looks. One subscription covers the entire ten minute workflow except the finale, and the finale is metered.
For the finale specifically, Riverside's AI B-roll documentation adds the shape of what a credit buys: each Veo 3 generation "produces ~8 seconds of high-quality video per prompt," and "Veo3 supports text prompts only," while Pika accepts an image as well. Worth knowing that the exact credit cost per generation is missing from that page: the sentence that should carry the number reads "Each Veo3 generation costs produces ~8 seconds," a copy error on Riverside's side where the figure used to be. Check your balance in Settings under Subscription before planning a video around generated B-roll.
| Tool or model | Its official name | What it does in this workflow | Cost |
|---|---|---|---|
| Riverside | Riverside Editing, the browser editor | The container for everything below. Records, transcribes, edits, captions and publishes without a local install | Pro: $29/mo monthly, or $24/mo billed annually ($288/yr). Free tier is 720p with a watermark and 2 hours of multi-track recording |
| The four scored clips | Magic Clips | Cuts short-form verticals out of the long video on upload and ranks them by score | Free on every plan, including the free tier. Riverside's own page says "Magic Clips is free" |
| The prompt box | AI Co-Creator | Takes plain-English instructions: build a reel, write a description, remove filler words | Included in Pro |
| Cutting by deleting words | Text-based editing | The transcript panel doubles as the timeline, so edits are made by reading rather than scrubbing | Included in Pro (listed as unlimited) |
| Filler and pause removal | Clean Up Speech | Strips ums, ahs and silences from the transcript and the video in one pass | Included in Pro |
| Auto B-roll | AI B-roll, in the AI tools panel | Reads the transcript and decides where B-roll belongs, then places it from the library | Included in Pro |
| Music and sound effects | The Riverside audio library | Categorized beds (lifestyle, news and politics, education) plus your own uploads | Included in Pro |
| Styled subtitles | Captions | Preset caption styles applied in one click, using the upload-time transcript | Included in Pro |
| Generating new B-roll | Google Veo 3 | Turns a written scene description into a brand new clip, roughly 8 seconds per generation, text prompt only. Julian's pick and his stated favorite | Pro and up, and it burns AI credits, which are bought separately and do not renew monthly. Google also sells Veo directly through Gemini, Flow and Vertex AI |
| The other generator | Pika 2.2 | The alternative model in the same picker. He does not demo it. Unlike Veo 3 it accepts an image as input, so it can animate a still you already have | Same AI credit pool as Veo 3. Pika sells its own plans directly too |
| What he left | Adobe Premiere Pro | The desktop editor that carried this workflow before, and still the reference for anything frame-accurate | A separate Adobe subscription, billed on its own |
| What auto B-roll replaces | Stock video platforms such as Pexels and Storyblocks | The search, download and import loop he describes as the old way of getting B-roll | Per clip or a separate subscription, plus the round trip out of the editor |
Set against that, Julian's offer is a real discount rather than a trial extension in name only. Riverside's public offer is a 14-day free trial; the code JULIAN through his affiliate link is a full month. He is an affiliate and the video is sponsored, both of which he discloses on camera, so price the recommendation accordingly and price the software honestly.
Key takeaways
- One transcription pass is the whole architecture. Uploading the file transcribes it, and every feature in the video is a consumer of that transcript: the clip picker, the hooks, the posts, the text editor, filler word removal, B-roll placement, B-roll suggestions, and captions. This is why the operations finish in seconds and why captions cost one click.
- Four short-form clips, four hooks, social posts and snapshots exist before you click anything. The clips carry a virality score and a ranked top choice, and thumbs up or thumbs down feeds back as training signal.
- Text-based editing removes the scrubbing, not just the clicking. Julian's line is "I don't need to watch anything back at all," and that is the largest non-generative saving in the workflow.
- The co-creator takes plain sentences. His three working prompts:
create a ready-made vertical 30-second reel from the most important part,give me a description for this, andremove pauses and filler words. It writes copy as readily as it cuts video, because both read the same transcript. - Auto-generate B-roll is two clicks and one editorial judgment. Click it, click Apply, and Riverside decides where B-roll belongs based on the transcript. This is the first feature that makes a taste decision rather than executing an instruction.
- Music, sound effects and captions never leave the tab. The audio tab has a categorized library (lifestyle, news and politics, education) plus your own uploads. Captions are preset picks, ranging from playful to serious.
- Multi-clip edits run on one vague sentence. Upload everything to Media, then
make this video look more interesting in every way possible, and it adds B-roll, music and transitions across the lot. - Generated B-roll is real and it is inside the editor. Describe a shot in a sentence and pick Pika 2.2 or Google Veo 3. Julian prefers Veo 3 and calls it "one of the best AI video generation tools." The generated clip may still need a manual scale adjustment to fit your frame.
- His hit rate and his warning belong together. "Nine out of 10 times it does an absolutely fabulous job," and also "check if the B-rolls are right, check if the music is good, check if everything is still inside, if the information is there."
- His headline claim, in his own hedged words: "like, I don't know, 10 minutes, where usually this was a task of like hours spending in the editing room."
- The cost picture the video omits. Magic Clips is free on every plan; everything else demonstrated here needs Pro ($29/mo, or $24/mo billed annually). Generated B-roll is the only metered piece: it runs on AI credits that are bought separately and "do not renew monthly," and each Veo 3 generation buys about eight seconds of video from a text prompt.
Chapters
0:00 The workflow he is burying 0:29 What this video is going to show 0:45 Welcome back to the channel 1:07 Riverside, and the half nobody talks about 1:31 The control: one clip, two edits 1:55 Four short clips and a virality score, before you click 2:33 Hooks 2:42 Posts written from the transcript 2:50 Snapshots for thumbnails and social 3:09 Opening the editor 3:21 The text-based editor 3:42 The co-creator 3:53 Prompt one: a ready-made vertical 30-second reel 4:09 Fifteen seconds later 4:27 Prompt two: give me a description 4:44 Prompt three: remove pauses and filler words 5:05 Why he picked Riverside: the AI tools panel 5:30 Auto-generate B-roll 6:05 Music and sound effects 6:56 Caption presets 7:25 Ten minutes instead of hours 7:46 Multiple clips, one prompt 8:11 The caveat: still check it 8:29 The B-roll lands, and the videos tab 8:48 Welcome to 2026: generating B-roll from a description 9:10 Pika 2.2 or Veo 3 9:22 The result, and a blooper 9:47 The wrap 10:02 The offer
Notable quotes
- 0:00 "I used to spend absolute bloody ages editing and perfecting my videos on my laptop and in Premiere Pro."
- 0:20 "I'm not even talking about adding subtitles, because to be honest, either it was a manual nightmare or I had to export my video and then use a third-party tool."
- 1:41 "I had a pretty tight deadline, so I was not putting in a lot of effort. So, what I ended up doing, I was just cutting out the mistakes and that was basically the video. So, not the most engaging video on the entire planet."
- 1:57 "Just by scrolling down, I can already see that I have four short-form clips cut out of my main piece of content before I even have started."
- 2:19 "I can say I like it, I don't like it, so that the AI is actually learning."
- 2:58 "This is just done even before we started doing anything at all."
- 3:28 "This is a text-based editor, so I could, in theory, just cut out a couple of things just by selecting the words I don't like. I don't need to watch anything back at all."
- 4:12 "After like 15 seconds or so, I now have my video."
- 4:32 "It has all of the transcripts, and you basically just need to tell it what you would like to get. It's as simple as that."
- 5:09 "I know there are also other tools out there that can do similar things. So, let me blow your mind right now."
- 5:50 "It's going to go through the entire transcript, and wherever it thinks it's going to make sense to just add some B-roll clips, it's going to do all of that for me."
- 7:38 "This now has taken me like, I don't know, 10 minutes, where usually this was a task of like hours spending in the editing room."
- 8:12 "I would still recommend to not let it do everything for you and never look back at it."
- 8:25 "From experience, I can tell you in nine out of 10 times it does an absolutely fabulous job."
- 8:48 "This is absolutely mind-blowing and welcome to 2026, I can also just go in here and describe the video I want to have, and then AI is going to generate something for me."
- 9:43 "The technology in the background is Google Veo 3, in my opinion, one of the best AI video generation tools."
Resources mentioned
The tool the video is about
- Riverside, the platform, and its pricing page.
- Riverside Editing, the browser based text editor Julian works in for the whole demo.
- Magic Clips, the feature that produces the four scored short-form clips on upload.
- AI Co-Creator, the prompt box at the center of the video.
- Clean Up Speech, Riverside's name for filler word and silence removal.
- Captions, the preset caption styles he applies in one click.
- Transcription, the pass that runs on upload and that every other feature reads from.
- The Riverside music library, where the education bed for his intro comes from.
- All Riverside AI tools and the free tools page, for everything the video does not cover.
- Magic Audio, AI Show Notes and AI Translation, the neighboring AI features in the same panel.
- Julian's Riverside link, the offer in the description, with the code JULIAN for one month free.
- AI credits, what happens when they run out, and how AI B-roll generation actually works, the three help center pages that answer the cost questions the video skips.
- AI at Riverside, their own accounting of which AI features are opt-in and which third parties process your data.
The generation models in the picker
- Google Veo 3, the model he selects and calls "one of the best AI video generation tools." Also reachable directly through Gemini, Google Flow and Vertex AI.
- Pika 2.2, the alternative model offered alongside Veo 3 in the same dropdown, and Pika's own pricing. Pika now fronts version 2.5, and Google's current video docs lead with Veo 3.1, so both models in the video have since been superseded.
The old workflow he is replacing
- Adobe Premiere Pro, the desktop editor he used for the control version of this exact clip.
- Stock video platforms of the kind he describes downloading from, such as Pexels and Storyblocks, and music libraries such as Epidemic Sound.
Tools in the same category, for comparison
- Descript, the original text-based video editor, which pioneered cutting video by deleting words.
- Opus Clip, the clip generator best known for popularizing a virality score.
- Submagic and Vizard, short-form clip and caption tools.
- CapCut, the free editor with its own auto-captioning and templates.
The channel
- Julian Eisenkirchner on YouTube, 139,000 subscribers, run by Julian and his partner Alex, documenting what it takes to build an online business: digital products, AI tools and workflows, personal branding, community and paid ads.
- Julian's Filmora Master Class for Wondershare, a paid relationship with a competing video editor, worth knowing when reading his tool preference.
Where it stands
Taken as a tutorial, this is a clean and honest walkthrough. Julian shows real operations on real footage, times the ones that can be timed, volunteers a caveat nobody asked him for, and names his competition out loud in a sponsored video. That is more integrity than the format usually gets. A few things a viewer should carry away with the workflow.
It is a sponsored video by an affiliate, and the framing follows. Riverside paid for it, he uses an affiliate link, and both facts are disclosed on camera at the end. It is also worth knowing what channel this is. Julian Eisenkirchner is not an editing channel; it is a 139,000 subscriber online business channel he runs with a partner, Alex, about digital products, personal branding and paid ads, where AI tooling is the means rather than the subject. Nothing shown is fake, but you are watching a demo run by someone with an interest in it working, on footage he chose, with results he was happy enough to keep. The failures, retries and rejected outputs that any generative workflow produces are not in this cut.
The ten minute claim is a comparison, not a benchmark. His own words are "like, I don't know, 10 minutes," measured against "hours spending in the editing room." Two things inflate the gap. The Premiere baseline he compares against is one he explicitly rushed, and it did not include B-roll, music or captions at all, so part of the difference is work he simply never did before. And the ten minutes is his active time in the tab, not wall clock: the auto B-roll pass ran long enough for him to cover two whole other tabs while waiting. Both numbers are still directionally right. Text-based editing plus one-click captions genuinely collapses hours of manual work. Just do not expect a stopwatch to agree.
The virality score is unexplained. Riverside ranks the four clips and names a top choice, and neither the video nor the product tells you what the score measures or what it was trained on. Treat it as a shortlist ordering, not a prediction. The thumbs up and thumbs down are more interesting, because they imply the ranking adapts to you over time, which also means it will drift toward whatever you already publish.
Automatic B-roll placement is a taste decision, not a lookup. It is the one feature in the demo where the tool decides something an editor normally decides: where emphasis belongs and where the viewer needs a break from a talking head. Julian's own three-point checklist covers this, and his third check is the one people will skip. An automated pass that removes filler, removes pauses, and reflows a multi-clip edit can drop a qualifier or a caveat and leave a transcript that still reads clean. His nine out of ten hit rate is his, on his content, with his eye on the result.
Generated B-roll is metered, and the video never says so. The Veo 3 clip lands in seconds and needs a manual scale fix, which is the honest version of what generated assets are like. The bigger omission is economic. Riverside's own help center says AI B-roll runs on AI credits that are "purchased separately from your subscription and do not renew monthly," and that each Veo 3 generation buys you about eight seconds. Nothing about that is hidden, but it is nowhere in the ten minutes, and it is the difference between "describe any shot you want" as a capability and as a budget. Everything else in the workflow is flat-rate; this one thing has a meter on it.
The models he names have already moved on. This is a hazard of a video with 2026 in the title, not a criticism of it. Julian picks Veo 3 in June 2026, and Google's current video generation documentation now leads with Veo 3.1, described as "a model for generating video with native audio." Pika has moved too: the picker in the video offers Pika 2.2, while pika.art now fronts Pika 2.5. Expect the dropdown to look different from the one on screen, and expect the output to be better than what he demos.
He has a paid relationship with a competing editor too. Julian fronts a Filmora Master Class for Wondershare, a directly competing video editor. That does not make anything in this video false, and he is not pretending to run a neutral review. It does mean the framing of "why I personally decided to go with Riverside" is a sponsored creator's preference rather than the verdict of someone with no stake in any of them.
The category is moving, and lock-in is real. Julian is right that other tools do pieces of this: Descript has led text-based editing for years, Opus Clip built its business on scored clips, and CapCut gives away captioning. What Riverside is selling is that they are all in one tab with one transcript underneath. That is a genuine convenience and it is also the shape of a switching cost, since a project living inside a browser editor is harder to take somewhere else than a Premiere timeline on a drive. Worth knowing going in, and it does not make the workflow any less real.


