45 seconds. 30 shots. 179 render jobs. One digital double, a tiny render farm, and a crew of AI agents that never once asked for a lunch break. Here's the whole messy, fun process.
The pitch was one sentence: put me in the music video.
Not a face filter. Not a green screen. I wanted the actual camera moves, the dancing, the energy and my co-star exactly where they were, with one change: the guy in the scene becomes a version of me that belongs in that pink, loud, studded-leather world.
LEFT: THE ORIGINAL · RIGHT: MY DIGITAL DOUBLE · FULL 45 S, ORIGINAL AUDIO
Step one: who's actually walking into this scene? I held auditions. All of them were me.
TOO MUCHROCK STAR
TOO COWBOYTHE HAT
CLOSER…STREET CAP
NAHBRAIDSFIRST RUN · THE ROCK STAR LOOK, ACTUALLY IN THE SCENE
Leather everything, glasses, maximum flair. It worked, and that was the problem: it felt like a costume. Then a studded hat went full cowboy. Tokyo street style with a backwards cap got closer. Braids? I tried them. Nah.
Where we landed: my own loose black hair with silver streaks, sunglasses, a tank top, tattoos and layered chains. Something I'd actually (almost) wear.
THE LOOK
THE FACEThe pink-studio character sheet locks the costume and hair (with and without glasses, because the glasses come off in some shots). My real portrait sits next to it so the styling can go wild while the face stays… actually mine.
Forty-five seconds sounds short. It's not.
Drums. The floor. Close-ups. A whip pan that lasts eight frames. Every cut changes the rules, so we sliced the video at every edit and turned each shot into its own job: its own prompt, its own references, and a list of things that are not allowed to change.
Woman-only shots? Untouched. If a close-up lost my face, we fixed that close-up without throwing away a good take somewhere else. Directing one shot at a time.

EVERY SHOT IN THE FINAL CUT
Here's the cool part.
SAM3 finds the guy in every frame and flips the colors inside his outline. That trippy negative video becomes the motion guide for MiniMax H3: it says exactly who to redraw, while the original motion, timing and camera carry straight through.
Add the character sheet and my portrait as references, and H3 paints me into the gap.
Think of the mask as the most literal stage direction ever written: this person. If it points at the wrong person, you get a gorgeous, completely wrong take. (Foreshadowing.)
Rendering thirty shots, over and over, is a lot for one graphics card.
So I built two apps. Pocket AI is my ComfyUI front end: every workflow (H3 video, Qwen Image, Krea 2, voice) in one place. Tensor Master is the dispatcher: it queues the jobs, schedules them onto whatever GPU is free, and has GPU rental built in. My RTX 3090 at home plus rented RTX 5090s on Vast all pull from one queue. Renting is meant to be a one- or two-click thing: prepared profiles bring the models and environment along, so I'm not rebuilding a ComfyUI box by hand every time.
And the whole thing works from my phone. Check the queue, watch a worker, look at a returned take, ask for a repair, add capacity. From the couch. 🛋️
Rentals have an idle cutoff and a max lifetime so nothing keeps billing after I walk away. Which also means the agents need the next job ready before a worker goes idle.
One agent doing everything in order was a bottleneck. A rented GPU can finish its work while everyone's still staring at yesterday's problem. So: four roles, shared production notes, me in the director's chair.

keeps the shot list, my latest notes and the selected takes

turns notes into prompts + references, queues the jobs

babysits the machines: pickup, progress, failures

checks the frames and writes very specific notes
QC went frame by frame on anything headed for the edit. Every problem became a note for Prep. Sometimes the fix was a sharper prompt. Sometimes a better reference. Sometimes the guide itself was wrong, and no amount of persuasive wording would save it.
In one whip pan the mask grabbed the wrong person, and the AI just… added me as a third performer between the original two. We accidentally cast an extra. A hand-tracked mask on the man on the left fixed it, and the pan survived.
In a dark shot the mask vanished for a few frames, so I flickered in and out like a ghost and bits of the original costume leaked through. Filling the nine empty masks gave us a steady take.
One "shot" secretly contained a cut from a duo to a solo singer. The model happily carried the woman and the drum kit into the solo part. Splitting it at the internal cut and generating both halves fixed the scene.
Too much "match the original expression" and the model rebuilt the original actor's face on my character. Loosening that and leaning on likeness kept me recognizable, with the performance still borrowed.
SAME SHOT, SAME MOMENT · HALF SPEED · SILENT LOOP
The hardest shot in the whole video is two seconds long.
In the original, my co-star pulls the glasses off his face mid-performance. Swap the guy, keep the handoff, keep her, keep the room. It took 11 takes.
For comparison I ran the same source through Genjutsu on Higgsfield, a commercial model. It left the glasses on his face, recast my co-star with dark hair, and quietly swapped the whole set and its colors. Pretty video. Wrong scene.
Tested with and without. No clear enough likeness boost.
Crop the face, regenerate, merge back. Pose, glasses and seams fought us.
More room for the performance, but more frames, memory and waiting.
Saved one take, then added me instead of replacing him in the pan.
● 158 done ● 18 cancelled ● 3 failed
The cost nobody puts in the GPU math: the crew ran all night, and my whole weekly token budget went with it.
Shot 11 (the star glasses) logged 11 completed takes; the sunglasses handoff in shot 3 took 11 submissions. Peak concurrency was two workers. About 6h50m of that GPU time ran on my 3090 at home and 2h48m on rented 5090s.
Counts are recorded jobs across tests and earlier versions, not clips in the final cut. Rental cost is modeled from recorded rates × lifetime (about $6.25, up to ~$7.87 with one uncertain rental), not an invoice. GPU time is summed per-job duration, not wall clock.
I set out to put myself in a music video. Somewhere along the way I accidentally built a small studio around it.