This morning I asked Claude Code a question I'd been putting off: with the video editing it can do now, do I even need Descript? I pay about $25 a month for it, and its core feature, editing a video by editing its transcript, is exactly the kind of work coding agents have become good at.
I already had evidence. Last week I edited a client video testimonial entirely in Claude Code, and this week I edited a 9-minute teardown the same way. Both times, though, the workflow lived in one-off scripts and a long conversation. So I asked Claude Code to turn it into something I could run on any recording, and by the end of the day it was a skill that takes a raw recording, cuts the retakes and false starts, shortens long pauses, sets the audio for YouTube, adds captions, a lower third and an end card in the SuperMarketers brand, and checks its own work before handing anything back.
It's free and open source, in the marketing-agents repo alongside the rest of the system I've published. This post covers how it works, the free tools underneath it, what broke while I built it, and why I think it does a better job than the software and the agency I used to pay for.
Why this is possible now
Video editing used to need a timeline editor because the only way to tell software what to cut was to drag clips around with a mouse. Coding agents like Claude Code and Codex changed that. They can do the dragging for you, as long as the job can be described in words and run as a command.
Editing turns out to be easy to describe. "Cut the first take of this sentence and keep the second" is a sentence, and once you have a transcript with a timestamp on every word, it's also a precise instruction. The tools that do the heavy lifting have been free and open source for years. What was missing was something that could read a transcript, make a judgment call about what to cut, and then operate those tools without me learning their syntax.
The tools, in plain English
You don't need to know any of these to use the skill, but it helps to know what's doing the work.
| Tool | What it is | What it does here | Cost |
|---|---|---|---|
| OBS Studio | Free, open-source recording software used by most streamers | Records your camera, screen and mic | Free |
| Claude Code | Anthropic's coding agent, running in your terminal | Reads the transcript, decides what to cut, runs every other tool, checks the result | Included in a paid Claude plan |
| Whisper (whisper.cpp) | OpenAI's open-source speech recognition model, running on your own machine | Turns the recording into a transcript with a timestamp on every word | Free, and nothing leaves your machine |
| ffmpeg | The command-line tool underneath most video software you've used | Makes the cuts, cleans up the audio, sets YouTube loudness | Free |
| HyperFrames | HeyGen's open-source framework that turns HTML and CSS into video | Renders captions, lower thirds and end cards over the footage | Free to run locally |
The scripts that tie it together are plain Python, so Codex or any other coding agent can drive them too. The skill file that tells the agent what to do is written for Claude Code.
How the workflow runs
You point it at a recording, usually just "edit my latest recording," and it works through seven steps.
- Transcribe. Whisper transcribes the recording locally. A 4-minute recording takes about a minute and 45 seconds on my M2 Pro.
- Decide what to cut. Claude reads the transcript the way an editor would and lists every retake, false start and "scratch that," keeping the last complete take of each sentence. This is the judgment step, and it's the part Descript leaves to you.
- Show me the cut list. Before anything renders, I get a document with every removed phrase struck through in context, the old and new length, and anything it left in because it wasn't sure. I approve it or tell it what to put back.
- Render a preview. ffmpeg makes the cuts and renders a quick 720p preview, which takes about 16 seconds for a 4-minute video.
- Check its own work. It transcribes the preview again and compares it with what should be there, which catches removed phrases that survived and words that got clipped. Then I watch the preview once and give notes with timestamps.
- Brand it. HyperFrames adds word-by-word captions, a lower third with my name and title, the logo and an end card, all generated from the brand settings rather than designed by hand each time.
- Final render. A full-resolution master with the audio at YouTube's loudness standard of -14 LUFS.
On a 4-minute test recording, the cut came out at 3:37 after shortening 29 long pauses and removing two false starts. The branded version, with captions, lower third and end card, rendered in about two and a half minutes.
Here's what the cut list looks like for a short test clip where the same line was said twice:
Source 00:26.97 -> cut 00:03.95 (85% removed, 2 segments) Removed: 1 retake; 3 pauses over 0.45s shrunk to 0.25s (22.1s total) Removed speech - 00:02.24-00:02.83 [retake] ...This is a testthis is a test let's see... Coverage check 0 kept words clipped, 0 removed words leaking through.
Two videos made this way
The Intryc testimonial was the first video I edited in Claude Code. Alex Marantelos, Intryc's founder, recorded one 10-minute take on his webcam, and it became the 8:41 case study below, a two-minute cut and four LinkedIn clips. The full write-up of that edit walks through every step.
This week's Flodesk teardown started as two files, a face recording and a screen recording, 10:21 long. Claude Code turned them into an 8:44 YouTube master at 1440p with an end card, a 1:45 LinkedIn cut in 4:5 with burned-in captions, a corrected caption file and a thumbnail.
The editing is only part of it. The YouTube titles, descriptions, chapters, tags and captions for both videos came out of Claude Code too, written from the transcript in the same session as the edit.
The parts that broke
The timestamp problem I wrote about last week came back immediately: Whisper's word timings drifted by up to two seconds, so cutting on them slices words in half. The skill now lets the audio decide where speech starts and stops, and uses the transcript only to decide what was said. Packaging the workflow for any recording turned up new failures I hadn't hit on a single video.
Whisper made things up. In silence it heard the word "you," a well-known Whisper habit. When I gave it a list of names to spell correctly, like Claude and n8n, it sometimes typed the list back as if I'd said it. When I prompted it to keep the "ums" it normally drops, it invented an "uh" that ran from 3.95 seconds to 30 seconds in a 4-second clip. Each of those now has a specific check, and the "um" detection only trusts a filler if it lands between two words that a second, unprompted pass agrees on.
A cut clipped the first word of a sentence. The re-transcription check caught it: the preview said "Generating the LinkedIn post now" where the speaker had said "It's generating the LinkedIn post now." The "It's" started softly enough that the speech detector missed its first tenth of a second. That check is the reason I trust the output, because it found the problem before I did.
It's conservative about "ums" on purpose. Two of the "uh"s in the 4-minute test ran straight into the next word, so cutting them would have clipped a real word. The skill leaves those in and lists them. A missed "uh" is a much smaller problem than a chopped word.
Keeping every video on-brand
Once the editing worked, I asked how to make sure every video matched the SuperMarketers brand, and that we could do the same for clients.
That question surfaced something embarrassing. I had two brand kits. One was a document from February with a slightly different yellow (#DFFF00), pure black and one font. The other was the live website, which uses #DFFF05, a near-black ink, four fonts and a different overall look. My LinkedIn image generator was still reading the February kit, and its prompts were asking for "hand-drawn typography" because it had picked up a setting meant for diagrams. While checking the logos I also found the yellow logo file had no color set at all, so the favicon on 51 pages of the site had been rendering black.
All of that is fixed, and the fix is what makes the video branding work. The website's stylesheet is now the single source of truth. A brand block in each client's config copies its exact colors, fonts, logos and video rules, and a check fails if the two ever drift apart. The captions, lower third and end card are generated from that block, never styled by hand. Before anything renders, a lint step scans the video for any color, font or logo file that isn't in the brand. When I planted five off-brand values to test it, including the old #DFFF00, it caught all five:
FAIL: compositions/brand-lower-third.html:10 color #DFFF00 is not a brand color
FAIL: compositions/brand-lower-third.html:10 named color 'black' is not a brand color
FAIL: compositions/brand-lower-third.html:10 font 'montserrat' is not a brand font
FAIL: compositions/brand-lower-third.html:21 color rgb(255, 0, 0) is not a brand color
FAIL: compositions/brand-lower-third.html:21 logo oyster-logo.svg is not from brand/logos
FAIL: brand lint, 4 files, 5 errors, 0 warningsFor clients it's the same block in their config. The skill won't render a client video until their brand passes the check, so nobody's video ships in my colors by accident.
There were two things the lint couldn't see, and I only found them by looking at frames from the render. A white logo vanished over a cream-colored screen recording, so the skill now measures how bright the footage is under the logo and picks the light or dark version. Captions also ran underneath a facecam in the corner of a screen recording, so there's now a setting to keep them out of that area.
Why I'd rather have this than Descript, or an agency
The obvious win is that it removes Descript. That's $25 a month, or $300 a year, and nothing else in the workflow costs anything beyond the Claude plan I already pay for.
The bigger win is that the editing now lives in the same place as everything else I publish. My positioning, my voice, my brand and the questions buyers ask AI engines all sit in one Claude Code repo, the brand brain I use for blog posts, LinkedIn and AEO work. When the video is edited there, the title, the description, the chapters, the captions and the on-screen text are written from the same source as the blog post and the AEO pages, so they say the same thing in the same words. That consistency is what helps a brand get described accurately, by people and by AI search engines.
I've paid a YouTube agency before. They did this kind of packaging, but never to this extent and never fully aligned with the rest of my content, because they were working from a brief rather than from the source material. Having everything centralized in one place is what makes the alignment automatic instead of something I have to police.
What it still can't do
It can't hear your delivery. It works from the transcript and the audio levels, so it can tell that you said a sentence twice, but not which take sounded better. When two complete takes differ only in tone, it asks.
It also can't feel pacing. A cut can be correct on paper and still land too abruptly, which is why the preview step exists. You still watch the video once and give notes like "cut 2:14 to 2:20" or "the jump at 3:05 is too fast," and it applies them.
Descript still has features this doesn't: Studio Sound for cleaning up a bad room, eye-contact correction and voice cloning. If you depend on those, keep it. For solo talking-head videos, interviews and screen walkthroughs recorded with a decent mic, this replaces it.
Running it on your own recordings
The skill, its scripts and a demo brand are in marketing-agents/tools/edit-recording. You need a Mac with Homebrew, ffmpeg and whisper.cpp, Python with PyYAML, Node for the captions, and Claude Code. The README covers setup in a few commands, including the one-time Whisper model download. After that, you record in OBS and tell Claude Code to edit your latest recording.
Two recording habits make it work much better. If you flub a line, pause, say "scratch that," and redo the whole sentence, because that phrase makes the retake obvious in the transcript. And leave a second of silence between takes, so the cuts land in silence and come out clean.
I asked Claude Code whether I still needed Descript, and a day later it had built the replacement, found three brand problems I didn't know I had, and shown me exactly where its own edits went wrong before I watched a single frame. That last part is why I'd trust it with a client video.