How to Prepare Your Song Before Generating a Music Video

Export the finished, full-length master as MP3 or WAV, under 50 MB, straight out of your DAW or distributor — never a re-encode of a re-encode, and never a recording of a speaker. Melodious reads that file to find your tempo and your section boundaries, then plans the storyboard against what it finds, so the file is the plan. Three things move the result most: dynamic contrast between quiet and loud passages, because sections are found by tracking level over time; an audible vocal, because the lyrics are read out of the mix; and uploading the version you will actually release, since the timeline is built from the file you send. The first run plans a 30-second video centered on your track's loudest sustained passage, so know which part of the song that is before you press send.
What does preparing your song actually involve?
The short answer: Export the finished, full-length master as MP3 or WAV, under 50 MB, straight out of your DAW or distributor — never a re-encode of a re-encode, and never a recording of a speaker. Melodious reads that file to find your tempo and your section boundaries, then plans the storyboard against what it finds, so the file is the plan. Three things move the result most: dynamic contrast between quiet and loud passages, because sections are found by tracking level over time; an audible vocal, because the lyrics are read out of the mix; and uploading the version you will actually release, since the timeline is built from the file you send. The first run plans a 30-second video centered on your track's loudest sustained passage, so know which part of the song that is before you press send.
Most guides to making a music video start after the song is already uploaded. That skips the step where first attempts most often go quietly wrong.
The whole process of turning your song into a music video hangs off one input: the audio file. Everything downstream — how many shots there are, where they change, what the plan thinks your chorus is — is derived from what the analysis reads out of it. The rest of the pipeline is covered in how to make an AI music video; this post is only about the file you hand it. Give it two minutes of care and you skip a redo.
Why the file matters more than it used to
Melodious used to interview you before it showed you anything. It asked what look you wanted, made you pick a length and a pacing style, then wrote out a storyboard and waited.
That is gone. Now you attach a song, type what you want to see, and press send — and it writes the storyboard and generates every keyframe image in the same run, reading the song's tempo and section structure to make those calls itself. There is no interview and nothing to approve before the images exist. (The full behavior is documented in one-message storyboards.)
Two consequences that matter for how you prepare:
- The audio is now doing work on the very first message. Under the old flow, your answers to a round of questions carried most of the plan. Now the analysis carries a chunk of it, so the quality of what it has to read matters more than it used to.
- Pressing send is the decision to spend. The credit cost for the storyboard and the keyframe images is stated under the composer before you send it. The video render — the expensive stage — is still a separate, explicit approval you click when the stills look right. Nothing renders because you sent a chat message.
So a sloppy file is not just a worse read, it is a run you paid for and want to do again. Prep is cheap insurance.
Export the right file: format, length, and what quietly degrades it
Melodious accepts MP3, WAV, FLAC, M4A, AAC and OGG, up to 50 MB. Within that list, the format is close to a non-issue. A good MP3 and a WAV both give the analysis plenty to work with, and agonizing over which one is a way of avoiding the decisions that actually matter.
What does matter:
- Export the finished, full-length track. Not a rough bounce, not a trimmed preview, not a two-minute radio edit unless that is the release. A partial file is analyzed as though it were the whole song, so you get a plan for a structure your record does not have.
- Export the audio, not a recording of the audio. A screen capture of a player, a phone pointed at a monitor, a voice memo of a rehearsal — all of these add room noise and clipping, and often lose the first half-second. The low end takes the worst of it, and the low end is where a lot of the level contrast lives.
- Do not re-encode a lossy file. MP3, AAC and OGG all discard information to save space. Decoding an MP3 and re-encoding it as another MP3 discards a second round from what is already an approximation — and the artifacts compound. If your song only exists as an MP3, upload that MP3 as-is rather than "cleaning it up" through another export. If you have the WAV, use the WAV.
- Upload the version you are actually releasing. If you are re-cutting vocals tomorrow, wait. The section boundaries, the tempo and the shot plan are all built from the file you send, and a new mix means a new plan.
- Bounce a stereo mixdown, not stems. One file, both channels. The analysis converts to mono internally, so a mono export is not wrong — it just does not help either.
If your song came out of a generative tool rather than a DAW, the export rules are the same but a couple of other things change; the Suno song to video guide covers those specifically.
Dynamics: give the analysis a shape it can read
This is the least obvious lever and the one worth understanding, because it explains a whole class of "why did it plan it there" reactions.
Your song's sections are not read from tags or metadata. They are found by tracking how loud the track is across its length and splitting it where the level shifts — quiet stretches, mid-level stretches, loud stretches, in sequence. Those bands are worked out relative to your own track, not against an absolute standard, so every song produces sections. The question is whether the boundaries land where you would put them.
What that means in practice:
| Your master | What the section read looks like |
|---|---|
| Clear arrangement contrast — sparse verse, big chorus | Boundaries land close to where you hear them |
| Heavily limited, near-constant level throughout | Sections still appear, but boundaries drift from the musical ones |
| A long quiet intro or ambient outro | Read as its own low-energy section, which is usually right |
| Sections shorter than about five seconds | Absorbed into the neighbouring section rather than kept separate |
Be honest with yourself about two things here. The labels are inferred from energy rather than from music theory — the loudest sustained passage gets treated as the chorus. If your actual chorus is the quiet one, expect the plan's peak to land somewhere else. And you do not need to un-master your song or hand over a dynamic "pre-limiter" version to fix this. Modern masters are loud; that is fine. The fix is to read the storyboard against your song's real structure and edit the scenes that sit in the wrong place, which is exactly what the storyboard panel is for.
If you already know your song's map — where the pre-chorus lifts, where the bridge drops out — writing it down before you upload is worth more than any export setting. Storyboarding against the song's sections walks through how to do that.
Vocals: can the mix be heard clearly enough to read?
The lyrics are read out of the mix you upload — there is no lyric sheet to paste in and no metadata carrying them across. So how present the vocal is in your master directly affects how much of it comes back.
A vocal sitting where a finished pop or hip-hop master would put it comes back cleanly. A vocal buried under a wall of guitars, drenched in reverb, heavily pitch-processed, or double-tracked in a dense arrangement is harder to read, and you will see that in the result — words missing, words wrong.
Which leads somewhere useful rather than somewhere depressing:
- Check the extracted lyrics before you rely on them. They are shown to you and they are editable. A misheard word caught there never reaches the finished file.
- If the video is meant to be lyric-led, vocal clarity is the whole ballgame. Melodious has a separate path that puts the words on screen timed to the vocal — see the lyric video guide — and its accuracy tracks how clearly the vocal sits in the mix.
- Do not remix your song for the AI. Pushing the vocal 3 dB above where it belongs to help the analysis is a bad trade: you would be publishing a video whose audio is not your record. Upload the master, correct what comes back.
Which 30 seconds will you actually get?
The first run is always a 30-second clip. That is deliberate — it is the cheapest thing worth looking at — and longer versions are available once you have seen the storyboard, which is a much easier judgement than picking a length blind.
The part worth knowing before you upload: that 30-second window is not the first 30 seconds of your song. It is centered on the longest, loudest sustained passage the analysis finds — the part it treats as your chorus — and then nudged onto the nearest detected beat if one is within about a second, so it tends not to open mid-hit.
So before you press send, ask yourself which passage in your track is loudest for longest. If that is the moment you want a video of, you are already aligned. If it is not — if the moment that matters is a quiet second verse, or a bridge, or the a cappella intro — say so in your first message. Your written prompt reaches the director alongside the analysis, so "build the whole thing around the stripped-back bridge, not the big chorus" is a steer worth spending a sentence on. Treat it as direction rather than a timestamp control, and reshape the shot list in the storyboard panel if it lands wide.
What not to do is trim the file down to the 30 seconds you want. A trimmed upload is analyzed as if that excerpt were the entire song, which throws off the section structure and the pacing that comes out of it.
What to do if your track is an instrumental
Instrumentals work. The analysis detects that there is no vocal and plans without lyrics; tempo and section structure still drive the shot list.
Two things change, and both are on you rather than the file:
- Your prompt carries more of the meaning. With a vocal, there is a lyric giving the director something concrete to picture. Without one, whatever you type is the only semantic input. "Instrumental post-rock, a lighthouse keeper through one winter, no people in the wide shots" is a brief. "Instrumental track, make it cinematic" is not. How to write a brief for your AI music video director goes deeper on this.
- The on-screen-lyrics path does not apply. There are no words to time, so the video is carried entirely by the images and the edit.
Instrumentals also tend to have more dynamic range than vocal-led pop masters, which usually means clearer section boundaries. That is a small advantage worth taking.
The pre-flight checklist
Two minutes, once, before the first send:
| Check | Why it matters |
|---|---|
| Finished master, not a rough mix | The plan is built from this file; a new mix means a new plan |
| Full length, not a trimmed excerpt | A trim is analyzed as if it were the whole song |
| MP3 / WAV / FLAC / M4A / AAC / OGG, under 50 MB | Anything else is rejected at upload |
| Exported, not recorded off a speaker | Room noise and clipping obscure the level contrast |
| Not re-encoded from another lossy file | Lossy-on-lossy artifacts compound |
| You know which passage is loudest for longest | That is where the first 30 seconds is centered |
| You know whether the vocal is clearly audible | It determines how much of the lyric comes back |
| One or two sentences written about what you want to see | It is the only semantic input the director gets |
Mistakes that cost you a run
| Mistake | What happens | Fix |
|---|---|---|
| Uploading a rough mix "just to try it" | You plan against a song that will not ship | Wait for the master, or accept you will redo it |
| Trimming to the section you want | The excerpt is read as the whole song | Upload full length, say which part in the message |
| A phone recording of a playback | Noise and clipping blur the level contrast | Export the file properly |
| Re-exporting an MP3 as another MP3 | Compounded lossy artifacts | Upload the original file untouched |
| Sending with no text, just the song | The director gets tempo and structure and nothing about intent | Write a sentence or two before sending |
| Assuming the extracted lyrics are correct | A misheard word reaches the finished video | Read and correct them before rendering |
| Expecting the first 30 seconds of the song | The window centres on the loudest sustained passage | Check where that is, or ask for a different focus |
What you do not need to bother with
Just as useful as the list above:
- Perfect loudness targets. Section detection uses thresholds relative to your own track, so there is no number to hit.
- Choosing WAV over MP3 on principle. Within the accepted formats the difference is not what stands between you and a plan you like.
- Stripping metadata, renaming files, or removing artwork. None of it is read.
- Providing a BPM or a key. Tempo is detected from the audio.
- Supplying a lyric sheet. The lyrics are read from the mix and shown to you to correct.
- Separate stems. One stereo mixdown is what is wanted.
Two minutes now, or a redo later
The compressed version of all of this: upload the release master, at full length, exported rather than recorded, and know two things about it before you press send — where its loudest sustained passage is, and how clearly the vocal sits in the mix. Those two facts predict most of what comes back in the storyboard.
Everything after that is editable. The scenes, the pacing, the length, the look — all of it can be changed in the storyboard panel or asked for in chat, and nothing renders into video until you approve it. But the analysis only ever gets one thing to read, and it is the file you hand it. Make that the good one.
Got a finished master sitting on your drive? Start with your song and see what the storyboard makes of it.
Frequently asked questions
What audio format should I upload for an AI music video?
MP3 or WAV are both fine, and Melodious also accepts FLAC, M4A, AAC and OGG, up to 50 MB. Export straight from your DAW, distributor or the tool that made the song. What matters far more than the format is that it is the finished full-length mix rather than a rough bounce, a trimmed preview, or a lossy file that has been re-encoded several times over.
Does the audio quality of my song affect the music video?
It affects what the analysis can read, which is what the storyboard is planned against. A clean export gives a clear picture of tempo, section boundaries and vocal. A phone recording of a speaker adds room noise, clipping and a wrecked low end, which blurs the level contrast the section detection depends on. A good MP3 and a WAV are close enough that the difference is not worth agonizing over; a recording of a playback is not.
Should I upload a mastered version or a rough mix?
The mastered version — the one you are releasing. Section boundaries are found by tracking loudness over time, and a finished master has the intended shape. A rough mix with an unbalanced vocal or a missing arrangement section produces a plan for a song that will not exist by release day. If your master is not ready, it is worth waiting rather than re-running against a second file later.
Do I need to trim my song to 30 seconds first?
No, and it is usually the wrong move. Upload the full track. A trimmed file is analyzed as if it were the whole song, so the section structure it reports is the structure of the excerpt, not of your record. The first run plans a 30-second video anyway, centered on the loudest sustained passage it finds, and you can reshape the shot list in the storyboard panel afterwards.
What if my track is an instrumental?
It works. Melodious detects that there is no vocal and simply plans without lyrics — tempo and section structure still drive the storyboard. Two things change: there are no lyrics to lean on, so your written prompt carries more of the meaning, and the on-screen-lyrics path does not apply. Say what the piece is about in your first message rather than assuming the music will imply it.
Does a heavily compressed, loud master hurt the storyboard?
It can flatten the thing sections are detected from. Boundaries are found where the level shifts, using thresholds relative to your own track, so a master squashed to a near-constant level still produces sections — they just land less where you would put them. You do not need to un-master anything. Read the storyboard's shot list against your song's real structure and edit any scene that sits in the wrong place.
Upload the master, not the rough
The storyboard is planned against the file you send — tempo, section boundaries, vocal and all. Send the version you'd release.
Start with your song