AI Video Guides 11 min read

How we made a 2-minute AI music video in 7 days

How to make an AI music video became clearer once we saw the hard number: 205 generations gave us 20 minutes of footage, and 1:58 made the cut. We made the full 2-minute animated music video in Vynzo for an original track, with two recurring young adult characters, Ava and Jordan, across 10-16 August 2026.

What we made, in numbers

We made a 2-minute animated music video for an original track, with Ava and Jordan recurring from start to finish. The full production ran for 7 calendar days, from 10-16 August 2026.

The funnel was the real story. We produced 205 generations, got 20 minutes of footage, and kept 1 minute 58 seconds in the final edit, which worked out to roughly a 10:1 ratio.

That ratio, not the quality of any one clip, was what people underestimate when they make a music video with AI. A good AI music video generator can give strong clips, but the finished video still came from selection, timing, reuse, and cuts.

StageResult
Video generations run205
Completed successfully203
Footage generated20 min 10 s
Clips kept after review152
Clips used in the final cut~90
Shots in the final cut142
Final runtime1 min 58 s
Calendar time7 days

Frames, motion and lip sync for your own music video — in one studio.

Open the studio

The look came first: concept art, then frames

We did not animate anything until the look existed as still images. Concept art fixed the palette, wardrobe, and both faces before we moved into shot production.

After that, every shot got its own starting frame generated in Nano Banana 2 from the concept art. A portion of those frames were refined afterward in an image editor, and the project used 151 frames total.

This order mattered because image-to-video inherited composition, lighting, and character from the still. A weak frame could not be rescued by a good motion prompt.

We also gave the characters names, Ava and Jordan, and used those names in the prompts. Twenty-two prompts referred to them by name, and that carried identity better than re-describing hair and clothes every time; it became one of our clearest character consistency AI video lessons.

Concept — the look
Concept — the look
Concept — the palette
Concept — the palette
Concept — the leads
Concept — the leads

Concept art, made before a single shot was animated. It fixed the palette, the wardrobe and both faces — everything downstream had something to match.

What it took: engines, settings and credits

For video generation, 204 of 205 runs used Seedance 2.0 image-to-video at 1080p. Within that work, 155 were 5-second clips at 15 credits each, and 35 were 10-second clips at 30 credits each.

The average render took 94 seconds, the longest took 19 minutes, and the project used 5.3 hours of machine time in total. For the video ledger, the number was exact: 3,613 credits charged, 30 refunded for two failed runs, and 3,583 net credits.

The full project came to roughly 4,500 credits, but the two halves came from different sources. The video number was measured from the account ledger; the frames and lip sync were priced from the public rate card because those tools were not billed on the account we used.

Nano Banana 2 images cost 2 credits from a prompt or 4 through the editor, and ours were a mix of both, so 151 frames came to about 453 credits. Kling 2.6 lip sync billed 2 credits per second of audio: the 27 lip-synced clips we kept measured 209 billed seconds, which came to 402 credits, while all 33 lip-sync runs worked out to roughly 491 credits.

Frames

Engine
Nano Banana 2
Runs
151 images
Settings
Text-to-image and editor passes off the concept art
Credits
~453

Motion

Engine
Seedance 2.0
Runs
204 generations
Settings
5 s and 10 s at 1080p, image-to-video
Credits
3,608

One shot

Engine
Kling 2.6
Runs
1 generation
Settings
5 s, first frame + last frame
Credits
5

Lip sync

Engine
Kling 2.6
Runs
33 runs, 27 kept
Settings
2 credits per second of audio
Credits
~491

Music

Engine
Suno
Runs
1 track
Settings
2 min 37 s, trimmed to 1:58
Credits
separate subscription

Total

Engine
Runs
Settings
Video measured; frames and lip sync at list price
Credits
≈ 4,500

Hard part one: the emotion you asked for vs the one you get

This was the single most repeated fight. When we asked for a chorus performance, the model kept drifting toward its default: wide-eyed, sincere, and slightly sad, while we wanted cool, self-aware, and a bit amused.

Positive description alone did not move it enough. What helped was writing the expression as an arc inside one shot, so the face changed over time instead of holding one adjective for ten seconds; 26 of our prompts used that kind of arc.

The other fix was naming the unwanted default and banning it outright. We used phrases like "no deer eyes, no innocent wide-eyed expression, no ballad sadness," and 19 prompts carried a ban like it; when a model had a strong prior, a negative was a stronger instrument than an adjective.

Seedance 2.0 — performing the chorus

Two things are doing the work here. The expression is written as an arc inside a single shot — playful first, sincere by the end — rather than as one adjective held for ten seconds. And the last line reserves the mouth: motion stays relaxed because lip sync is applied later. The companion take went further and banned the default outright ("no deer eyes, no innocent wide-eyed expression, no ballad sadness"), which is what finally moved the read.

Reveal the prompt

10-second 16:9 animated music-video shot. Preserve the exact girl, face, bright blue eyes, long straight blonde hair, silver hair clip, denim jacket, black top, proportions and illustrated visual style from the reference image. She performs the chorus while sitting at her school desk, looking directly toward camera. Start playful and self-confident: soft knowing smile, slightly raised eyebrow. Then her expression gradually becomes more sincere and vulnerable, like she suddenly means the words more than she expected. Small head tilts, subtle shoulder rhythm, fingers lightly tapping the desk to the beat. Gentle hair movement. Slow cinematic push-in from medium shot to medium close-up. Background students stay soft and secondary. Keep mouth movement subtle and relaxed for later lip sync. No exaggerated gestures, no camera shake, no text, no style change.

The fix for a wrong emotional read is a list of what NOT to do.

Hard part two: the same set from another angle

Once a location existed, returning to it for a second angle became one of the most expensive parts of the workflow. Rewriting the prompt often changed too much, so we got more value from re-rolling the identical prompt.

In total, 21 prompts were submitted more than once byte-for-byte, creating 29 extra generations. We also locked both ends of a move with a first frame and a last frame, which was what the single Kling 2.6 generation in the project was for.

Depth of field helped cover the remaining mismatch. Thirty-four prompts pushed the background soft and secondary, so a re-generated set did not need to match perfectly in every detail.

Seedance 2.0 — the punch bowl

This exact prompt was submitted four times in one hour, unchanged. Nothing was wrong with the wording — the model simply had to be asked repeatedly before the boy approached from behind without the two of them merging. Re-rolls, not rewrites, carried 29 of our 205 generations.

Reveal the prompt

cinematic shot based on the reference image. The blonde girl stands beside the punch bowl, relaxed and confident, holding a plastic cup. She raises the cup and takes one large sip of the party punch. At the same time, the dark-haired boy quietly approaches her from behind, naturally closing the distance through the party crowd. Just as she finishes the sip, he gently places one hand on her shoulder to get her attention. She has not turned around yet. Background guests continue talking, drinking and moving naturally. Warm amber party lighting, subtle camera push-in, realistic body movement. Preserve both characters, outfits, room layout and anime-cinematic style. No sudden motion, no morphing, no extra limbs.

Seedance 2.0 — the almost-kiss

Two versions were generated 14 minutes apart: one where they kiss, one where they lean in and stop. Keeping both cost 60 credits and made the ending an editing decision rather than a generating one — which is the cheaper place to change your mind.

Reveal the prompt

Animate the provided image, preserving the exact faces, hairstyles, clothing and anime style of both characters. They stand very close, hold soft eye contact, slowly lean in but don't kiss. At the exact moment the camera begins a smooth cinematic 360° orbit around them, keeping their faces and upper bodies centered and sharp. The couple stays almost still with subtle breathing, blinking and slight hair movement. As the camera circles them, the entire party background gradually becomes heavily blurred with warm golden bokeh from candles and string lights, as if the world disappears around them. Shallow depth of field, natural parallax, warm amber rim light, romantic pacing, stable camera, high character consistency. Avoid face morphing, anatomy errors, camera shake, fast motion, focus loss, or characters rotating instead of the camera. /

Returning to a set you already generated is the most expensive thing in the whole workflow.

Lip sync goes on last

Our order was simple: generate the shot first, then apply lip sync in Kling 2.6, driven by the vocal from the finished track. We made 33 lip-synced clips that way.

The counter-intuitive part was the motion prompt. It had to keep mouth movement subtle and relaxed because the mouth got replaced later; 10 prompts said this explicitly, and shots with expressive mouth articulation fought the lip sync applied on top.

Seedance 2.0 → Kling 2.6 lip sync

Note what the source prompt does NOT ask for: it keeps the motion subtle and the mouth relaxed, because the mouth gets replaced later. Lip sync goes on last, in Kling 2.6, driven by the vocal from the finished track. This shot ended up in the final cut four separate times.

Reveal the prompt

Animate this image into a subtle cinematic anime-style video. Preserve the blonde woman’s identity, face, blue eyes, freckles, hair, outfit, headphones, bedroom, lighting, colors, and composition. She lies on her stomach along the bed, hands gently supporting her face, looking into the camera and softly singing along to music in her headphones. Add natural lip-sync, small mouth and jaw movements, gentle blinking, slight eyebrow motion, soft breathing, head movements to the rhythm, and minimal hair motion. Keep her gaze mostly on the camera. Fairy lights softly flicker, candle flame moves naturally, neon heart glow subtly varies. Camera stays almost static with a very slow push-in. Keep motion smooth and restrained. No face changes, warped hands, extra fingers, body distortion, outfit changes, camera cuts, or exaggerated movement.

The part nobody talks about: the edit

The strongest lesson of the project was that generation skill without editing skill produced a showreel, not a music video. The track ran at 130 BPM, and the finished video had 142 shots in 118 seconds.

The median shot was 0.70 seconds, which matched a dotted quarter. Eighty-five shots were exactly 1.5 beats long, 31 were exactly 2 beats, and 117 of 141 cuts landed within half a frame of the half-beat grid.

The density also came from where the cuts went. Eighty of 141 cuts jumped to a different location, so the video cross-cut between school, cafe, bedroom, and party instead of playing each scene as a block.

We reused material on purpose: 82 source clips covered 123 shots, and some clips appeared up to four times. The tempo was also built into generation, not only the edit, with 19 prompts naming a BPM and 3 scripting choreography beat by beat; the two long holds around 77-84 seconds were the only breather, and they were deliberate.

Seedance 2.0 — choreography scripted in beats

The choreography is written as beats, at the same 130 BPM the track runs at: "Beats 1-2: dancers step outward. 3-4: pom-poms rise in alternating diagonals." Motion generated on the tempo grid cuts cleanly on the tempo grid — 19 of our prompts name a BPM for exactly this reason.

Reveal the prompt

High-angle wide shot. Keep exactly the same five Westfield cheerleaders, uniforms and gym. Animate geometric synchronized choreography at 130 BPM. Beats 1-2: dancers step outward, expanding the formation. 3-4: pom-poms rise in alternating diagonals. 5-6: two quick steps inward while everyone rotates about 45°. 7-8: formation opens into a star-like shape and hits a strong final pose. Maintain precise spacing and synchronization. Camera slowly cranes downward and pushes toward center court. Floor reflections and pom-pom highlights move naturally. No added dancers, morphing, character drift or costume changes.

Six seconds of the finished cut

Nine shots in six seconds, with sound. The median shot in this video is 0.70s — a dotted quarter at 130 BPM — and 117 of 141 cuts land within half a frame of the half-beat grid. This is what a face has to survive: 0.7 seconds at a time, 142 times.

Generate on the grid, cut on the grid.

What still breaks

Two failure modes still broke through. On long, near-static holds, the model could hallucinate extra figures, with a second person fading in out of nothing; lettering also degraded whenever the camera moved, so WESTFIELD on the cheer uniforms drifted into nonsense across a shot.

The practical mitigations were not glamorous, but they helped. We kept holds short, which the fast cut already did, kept readable text out of frame or out of focus, and re-rolled rather than rewrote when a clip was close.

Seedance 2.0 — another take, same bedroom

Watch over her shoulder near the end: a second figure fades in out of nothing. Long holds on a near-static frame are where this shows up most, which is one more reason the edit runs short. Lettering fails the same way — the WESTFIELD on the cheer uniforms drifts into nonsense whenever the camera moves.

The workflow that worked

This was the sequence we would repeat, in order.

  1. 1Finish the track first. The edit is built on its tempo, so you need final audio before you generate a single shot.
  2. 2Design the look on paper. Concept art that fixes palette, wardrobe and both faces gives every later prompt something to match.
  3. 3Name your characters and use the names in every prompt — "Ava" and "Jordan" held identity across generations better than physical description alone.
  4. 4Generate one frame per shot in Nano Banana 2, working from the concept art rather than from scratch.
  5. 5Animate each frame in Seedance 2.0. Five seconds at 1080p is the workhorse; reach for ten only when the action genuinely needs it.
  6. 6Script motion in beats at the track tempo, so generated movement lands on the same grid the edit will use.
  7. 7Say what you do not want. Wrong emotional defaults are the most common reason a technically fine shot reads wrong.
  8. 8Add lip sync last, in Kling 2.6, and only on the shots that hold the camera on a face.
  9. 9Cut on the beat, and cross-cut between locations instead of playing each scene as a block.

Frequently asked questions

How do you make an AI music video?

We made ours by building the look first, then generating frames, motion, lip sync, and the edit. Concept art fixed Ava, Jordan, palette, and wardrobe before any animation; then 151 starting frames drove image-to-video generation, lip sync went on last, and the 130 BPM edit shaped the final 1:58.

How long does it take to make an AI music video?

Our 2-minute video took 7 calendar days, from 10-16 August 2026. That included 205 generations, 20 minutes of footage, 5.3 hours of machine render time, and the edit that reduced the material to 1:58.

How do you keep the same character in every shot?

We kept character consistency by starting from fixed concept art and named characters. Ava and Jordan appeared by name in 22 prompts, and each shot started from a Nano Banana 2 frame based on the established look instead of relying on text alone.

Which AI engine is best for a music video?

For this project, Seedance 2.0 image-to-video did nearly all the video work. We used it for 204 of 205 video generations at 1080p, while Nano Banana 2 handled frames and Kling 2.6 handled lip sync.

Do you still need to know how to edit?

Yes. The edit turned generated clips into a music video instead of a showreel. Our final cut used 142 shots in 118 seconds, with most cuts landing on the half-beat grid and many cuts jumping between locations for density.

Keep exploring

Bring your photos to life

Frames, motion and lip sync for your own music video — in one studio.

Make your music video