I Made a Music Video in 8 Hours. The AI Was the Easy Part.

I was working on a video I wanted to make, and it wasn’t pulling me anywhere. I wasn’t have creator’s block and there was no particular crisis. The work just sat there while I stared at it. I just wasn’t feeling it.
So I asked a different question. Not what should I be working on. What would actually engage me right now?
The answer was a song. I’d been creating music in Suno 5.5 for weeks and loving what came out of it, so I opened ChatGPT and typed a prompt at 4:51 PM. Funk-soul. Syncopated bassline, tight horns, clavinet. A song about who I am and what I do.
Eight hours later I exported the finished music video. One in the morning, and it was on my hard drive.
The song stopped being about me
The first version of the prompt was a profile piece. TRS-80, WorldVillage, ClassicGames to Yahoo, iFart, the Barnes & Noble register. My story, set to a groove.
I asked for it and then didn’t want it.
A song about my résumé is a song nobody needs. But the arc underneath the résumé is the thing I say from stages and on my podcast, and that one belongs to the audience, not to me. Six waves came through. People laughed at every one of them before they piled in. If you can recognize a disruption while you’re standing in it, you already know where you are on the curve.
You are not behind. You are early. Again.
That’s the war cry, and the song took its name from it. “You’re Early Again.” Once it pointed outward instead of inward, the writing got easy. Thirty minutes with ChatGPT to find the story and land the lyrics. Thirty minutes, total, for the part I assumed would take the longest.
The funk-soul angle didn’t survive contact with my own ear. The groove was fine. But it just wasn’t me. What I actually wanted was a Huey Lewis and the News 80’s vibe, arena rock with a Hammond B3 and horn stabs and a crowd you can hear. Heartland pop-rock, 118 BPM.
Sixty-two versions
Then two hours in Suno, and sixty-two iterations to get one keeper.
Suno is fast enough that you can’t walk away from it. The results land in a minute, so you sit there and listen and fix and go again. Most of the fixing was structural. Section tags that collapsed. A pre-chorus that ran into a chorus. Lyrics that read fine on the page and landed wrong in a mouth.
Some of it was stranger. I wanted a live audience feel, and Suno decided the crowd should sing my lines back at me. Not the chorus. My lines. I ended up putting “The crowd does NOT sing along” into the style box in capital letters, the same way I write negative prompts for image tools, and it still fought me. The two lines the crowd end up singing in the final version actually work out pretty nicely.
I also fed it my own voice. I don’t have an amazing singing voice, but I’m okay on karaoke night. Well, Suno held my voice for about a minute, sounded like me, and then drifted off into something Suno preferred. Every time. If you’re waiting on voice cloning to be a solved problem in music generation, wer’re going to need some more advanced technology. I’m sure we’ll have it soon.
The thing I keep coming back to: the style got locked but the output changed. Same prompt, same lyrics, different song. That’s why sixty-two happens even after you’ve already made every creative decision. You’re not steering toward a target. You’re rolling until something shows up that checks most of the boxes.
Where the vision broke
Revid.ai came next, and this is where the day got real.
I’d written a long visual spec. Sold-out 2,500-seat theater, warm amber and deep blue, charcoal sport coat over a black tee, vintage chrome microphone, full heartland band, three backup singers, a three-piece horn section. Netflix concert documentary. Subtle film grain, realistic skin tones, natural camera movement, no lasers.
I saw the first clips come back and knew immediately that the film in my head wasn’t the film I was going to get.
Three things missed. My character stood too still, which is a problem when the whole premise is a frontman working a room. The shot selection came back flat, none of the drama I’d pictured. And there’s no real mechanism in Revid for holding continuity between scenes, so the musicians drifted. Different band, clip to clip, no matter how precisely I described them.
That last one is the state of AI video right now, and the continuity guidelines I wrote didn’t fix it. You can specify wardrobe and stage layout and exactly how many horn players are standing where. Nothing enforces it. The specification isn’t a constraint, it’s a suggestion.
So I stopped chasing the version in my head and started building the version I could actually get.
Chunking, and about seventy-five dollars
Ten primary clips, roughly 35 seconds each, generated with deliberate overlap so I’d have material on both sides of every cut. Three of those I threw out and regenerated from scratch. Then eight b-roll pieces around ten seconds, no audio, just musicians and crowd and hands and faces.
The overlap wasn’t waste. It was insurance against holes I couldn’t predict yet.
Total spend on Revid for the video: about $75.
Meanwhile, dinner with Erin. Then an hour of a board game while renders ran. Video generation is slow enough that you get your evening back, which nobody mentions when they talk about AI speed. The bottleneck wasn’t effort. It was waiting, and waiting is time you can spend living.
The part the tools don’t do
Filmora at the end, and this is the part I love.
Two jobs. Line up every segment so the lip sync and the captions hit exactly, because that alignment is the entire difference between footage that reads as real and footage that reads as generated. Then cut out everything that looked wrong and cover it with something else.
That’s what the overlap and the b-roll were for. When the band morphed mid-phrase, I had a crowd shot to go to. When a face went strange, there was another angle sitting in the bin.
I’ve been editing video for decades. It’s a skill I stacked long before any of these tools existed, and it’s the one that saved this project. The AI handed me raw material. The decades told me what to do with the material when it came back broken.
Six skills, stacked
Nobody made this video. I made this video, using five things at once.
- Ideating, which is knowing what to build and why.
- Visualizing, which is holding a picture clearly enough to measure the output against it, even when the output loses.
- Prompting, which is a writing skill wearing a technical costume.
- Storytelling, which decided the song should point at you instead of at me.
- Chunking, which is knowing to build with overlap before you know where the seams will fall.
- Editing, which is the one that turns a pile of clips into something a person will watch.
Take any one of those away and there’s no video. Take the tools away and there’s still a person who knows how to make something.
“You’re Early Again” says exactly that. I wrote the line about AI and about six waves before it, and then spent one evening proving it out with a stalled project, a $75 render bill, and a band that kept changing faces.
Go make the thing you can’t quite see yet. You’ll find out what it actually is on the way there.
Here’s my music video. I hope you sing along and leave a comment below!



