How to Write an AI Character Brief That Holds

An AI character brief holds across scenes when every word in it maps to something a frame can physically show at a distance. Lead with silhouette — hair shape and colour, one distinctive garment described by length and volume, and a two-or-three-colour palette — then add one hard accessory and one prohibition. Cut age, beauty adjectives, mood words and facial micro-detail: nothing in the image corresponds to them, so the model re-interprets them scene by scene. Test the brief with three changes: move the character to a different setting, change the light from day to night, and shrink them to a hundred pixels tall. Whatever survives all three is what actually identifies them. In Melodious the brief is uncapped free text saved alongside one reference image, and both travel with the character when you mention it in a prompt.
What makes an AI character brief hold across scenes?
The short answer: An AI character brief holds across scenes when every word in it maps to something a frame can physically show at a distance. Lead with silhouette — hair shape and colour, one distinctive garment described by length and volume, and a two-or-three-colour palette — then add one hard accessory and one prohibition. Cut age, beauty adjectives, mood words and facial micro-detail: nothing in the image corresponds to them, so the model re-interprets them scene by scene. Test the brief with three changes: move the character to a different setting, change the light from day to night, and shrink them to a hundred pixels tall. Whatever survives all three is what actually identifies them. In Melodious the brief is uncapped free text saved alongside one reference image, and both travel with the character when you mention it in a prompt.
Character drift is usually blamed on the model. A large share of it is written into the brief before generation starts: the description says almost nothing, so the model fills the gap, and fills it differently every time.
This post is the deep end of one row in the video brief template — the row where you name who recurs on screen. For the mechanics of saving and reusing a character, see reusable AI characters for music videos; for the end-to-end flow, how to make an AI music video. What follows is only the craft of the description itself.
A brief that holds has five slots and runs thirty to fifty words:
- Hair — shape and colour
- One garment — named, with length and volume
- Palette — two or three colours, and no others
- One accessory — small, hard, specific
- One prohibition — something the character never wears or does
The rest of this post is why those five and not others, and what each one looks like written well.
Why does "beautiful woman, 25" fail?
Run that phrase through the test that matters: could a stranger use it to pick your character out of a line-up of ten people? No. Every word in it is either a judgement, an estimate, or a category so large it excludes nobody.
- "Beautiful" is a judgement with no pixel behind it. The model resolves it against whatever else is in the prompt — one thing in a neon alley, another on a beach at noon.
- "25" is an estimate that moves with lighting and styling.
- "Woman" narrows the field from billions to billions.
- "Stylish outfit" delegates the single most useful identifying detail — clothing — straight back to the model, which re-invents it per scene.
The phrase isn't wrong, it's just not discriminative, and a brief that isn't discriminative gives the model nothing to hold between scenes. So it re-samples a plausible person each time.
Which character details hold across scenes, and which drift?
Not all descriptors are equal. The ones that hold have a physical correlate that survives a change of angle, light and distance.
| Detail | Holds? | Why |
|---|---|---|
| Hair shape + colour | Strong | High contrast, readable from across a room, unaffected by wardrobe |
| One distinctive garment (length, volume, colour) | Strong | A single unusual item identifies faster than a full outfit described |
| A two-or-three-colour palette | Strong | Colour is the first thing a viewer matches from shot to shot |
| Silhouette + posture (tall, stooped, coat always open) | Strong | It's the outline, so it survives distance and back-lighting |
| One hard accessory (a silver ear cuff, round glasses, a case) | Medium | Anchors medium shots; too small to read in wides |
| Facial structure in words ("square jaw") | Weak | Only readable in close-ups — where the reference image is already doing the work |
| Age as a number | Drifts | Re-estimated per scene from light and styling |
| Beauty or mood adjectives | Drifts | No pixel corresponds; re-interpreted from surrounding context |
| Expression ("smiling") | Drifts — and shouldn't be fixed | Different sections of the song need different faces |
| A named real artist | Never use | Trademark and brand-safety risk, and the likeness drifts anyway |
The pattern: the further a detail is from the outline, the less it survives. Most music-video shots are wides and mediums, not close-ups, so a brief built out of face detail describes the one thing the frame usually can't show.
Before and after: one character brief, rewritten
Here is a brief of the kind that produces four different people across four scenes.
A beautiful woman, 25, long hair, stylish outfit, moody vibe. Very cinematic.
Twelve words, none of them load-bearing. Rewritten:
Tall, cropped platinum bleach with dark roots grown out an inch.
Oversized oxblood leather trench, past the knee, always open and always
moving. One thick silver hoop in the left ear.
Palette: oxblood, platinum, black — no other bright colour on her, ever.
Forty-two words, and every one of them is something a camera can see. Line by line, here is what changed and why:
| Was | Became | Why it holds |
|---|---|---|
| "long hair" | "cropped platinum bleach, dark roots grown out an inch" | Shape and two colours, both readable in a wide. Note this changes the length, not just the wording — "long" was a placeholder, and committing to a definite cut is part of the job |
| "stylish outfit" | "oversized oxblood leather trench, past the knee" | One named garment with length and volume — a silhouette, not a mood |
| "moody vibe" | deleted | That's the director style's job, not the character's |
| "25" | deleted | The reference image carries age far better than a number does |
| "beautiful" | deleted | No pixel corresponds to it |
| — | "tall" | Silhouette before anything else — height reads at any distance, in any light |
| — | "always open and always moving" | Behaviour that reads at distance: a coat that moves is visible when a face isn't |
| — | "one thick silver hoop in the left ear" | The hard accessory — one specific object that anchors the medium shots |
| — | "no other bright colour on her, ever" | The prohibition — the single highest-leverage line in the brief |
That last row does work no positive description can, because it constrains every scene you haven't thought about yet. "No other bright colour" is what stops the character turning up in a yellow jacket in the bridge because the scene description happened to mention sunflowers.
Strip this particular person out and the shape is reusable:
[Hair: shape + colour]. [Garment: item, length, volume, how it moves].
[One accessory: small, hard, specific].
Palette: [two or three colours] — never [the prohibition].
If your character isn't a person, the five slots still hold — swap hair for surface or material, and the garment for whatever breaks the outline. A cracked brass visor, a torn left wing, one dented chrome panel: same job, different noun.
The three-change test
Before you save a character, put the brief through three changes and see what's left standing.
- Change the setting. Delete anything in the brief that depends on where the character is standing. "Leaning on a neon sign" is a scene, not a character — it belongs in the storyboard. A brief that only makes sense in one location fights every other location in the video.
- Change the light. Anything describing how the character looks when lit — "glowing skin", "sun-kissed", "backlit hair" — dies the moment you cut to a night interior. Describe intrinsic colour instead: hair, garment, palette.
- Change the distance. Shrink the character to a hundred pixels tall, roughly what a wide shot gives you. What survives? If the honest answer is "nothing", the brief is face-only, and most of a music video is not close-ups.
Whatever survives all three is doing real work. Whatever fails one is decoration — and decoration is exactly what the model re-invents per scene.
What does the reference image do that words can't?
The brief steers; the image anchors. When you type @ in the composer and pick a saved character, two things travel with it: your brief is appended to the message as a note telling the director to feature that character and keep it consistent, and the saved image is passed to the image model that draws your keyframes as an actual reference, with an instruction to preserve the subject's likeness. That second one is a far stronger signal than any sentence.

A saved character in Melodious holds one image, so choose it deliberately: sharp, evenly lit, roughly front-on, and — this is the part people skip — showing the load-bearing features you named. If the brief says "oversized oxblood trench, past the knee", a headshot is the wrong reference. You want the coat in frame.

One clarification, because the two fields get confused: the 200-character limit you may have read about applies to a custom director style, which sets the look of the whole video. The character brief is uncapped free text. That isn't licence to write a paragraph — thirty to fifty words of physical detail beats three hundred words of atmosphere, because the model weights everything you give it and atmosphere dilutes the features.
Faces, close-ups, and one thing that isn't built yet
Two constraints should shape the brief.
First, there is no lip-sync in Melodious. A character can't credibly deliver the hook to camera, so don't design the video around singing close-ups — plan wides, silhouettes, backs of heads, hands on the guitar, a figure walking out of frame. If the face is rarely the subject, a face-first brief is describing the wrong thing anyway.
Second, be honest about the ceiling. A saved reference gets you far closer than re-typing a description, and it is the strongest lever available today, but holding one performer genuinely identical across every scene isn't reliable yet — we'd rather say that than sell it. The hedge is the craft above: build a character recognisable by hair, coat, colour and posture, so when the face wavers between shot four and shot nine the audience still reads one person. Wardrobe and silhouette survive a scene change more dependably than a jawline does.
What are the most common character brief mistakes?
| Mistake | What you get | Fix |
|---|---|---|
| Adjectives where nouns belong | A different plausible person each scene | Name the hair, the garment, the colours |
| No prohibition line | A stray colour or prop creeping in from the scene | Add one "never" |
| The scene written into the character | The character fights every other location | Move it to the storyboard |
| A brief written for a close-up | Identity collapses in every wide | Write for a hundred pixels tall |
| A reference image that's a headshot | The named garment never appears | Reference the whole look |
| A real artist named for the look | Legal risk, and it drifts anyway | Describe the aesthetic in features |
From brief to storyboard
Save the character once and it stops being something you re-type. Mention it in your first message so the storyboard is planned around a real figure rather than a placeholder, and mention it again on the scenes that matter — you can ask the director in chat to use a saved character on a specific scene, and it re-does that scene's keyframe with the reference attached.
Then read the storyboard back against the brief, scene by scene. Is the coat in the chorus? Has a second bright colour appeared? Did the bridge quietly become a close-up? Fixing that on the storyboard costs a text edit; fixing it after generating costs a re-render. If you're building a figure for a whole release rather than one video, building an artist persona with AI video picks up where this leaves off.
Open the studio with a track and one good reference image, and write the forty words that keep your character the same person from verse to chorus.
Frequently asked questions
What should an AI character brief include?
Five things, in this order: hair (shape and colour), one distinctive garment described by length and volume, a palette of two or three colours, one hard accessory, and one prohibition — something the character never wears or does. That is roughly thirty to fifty words. Everything in it should be visible in a frame; everything that isn't visible is decoration the model will re-interpret differently in every scene.
What is character drift in AI video?
Character drift is when the same character is rendered as a visibly different person from scene to scene — the face changes, the hair length changes, the clothes change. It is usually blamed on the model, but a large share of it is written into the character description before generation starts: if the description is vague, the model has nothing to hold on to between scenes, so it re-samples a plausible person each time. The fix is a description built from features a frame can physically show — hair, one distinctive garment, a fixed colour palette — plus a reference image to anchor it.
Why does "beautiful woman, 25" produce a different person in every scene?
Because none of those words correspond to pixels. "Beautiful" is a judgement, not a feature, and the model resolves it differently depending on the setting, the lighting and the genre cues around it. "25" is an estimate that shifts with styling and light. "Woman" narrows a population of billions to a population of billions. A description only holds identity if it would let a stranger pick your character out of a line-up of ten people.
How long should a character brief be?
Around thirty to fifty words. There is no hard limit in Melodious — the character brief is free text with no character cap — but long briefs dilute rather than sharpen, because the model weights a paragraph of atmosphere against the handful of features that actually identify someone. Separately, a 200-character limit does apply to a custom director style in Melodious, which is a different field: it sets the look of the whole video, not who is in it.
Do I need a reference image, or is a good description enough?
Both, and they do different jobs. The description steers; the reference image anchors. When you mention a saved character in a prompt, the saved image is passed to the image model that draws your keyframes as an actual reference, with an instruction to preserve the subject's likeness — that is a far stronger signal than any sentence. Words alone will drift. A reference alone can't tell the model which parts of the picture are the character and which are that day's outfit.
How many reference images can I save per character?
One. A saved character in Melodious holds a name, a handle, one image and one brief. So choose that image the way you'd choose a passport photo that also has to show the wardrobe: sharp, evenly lit, roughly front-on, and showing the load-bearing features you named in the brief — the hair, the distinctive garment, the colours — not just the face.
Can I get the same face in every shot?
Not reliably, not yet. A saved reference gets you much closer than re-typing a description, and it is the strongest lever available today, but holding one performer genuinely identical across every scene of a video isn't something we'd promise. That's the practical reason to make your character recognisable by silhouette, hair and wardrobe as well as by face — those read even when the face doesn't.
Save the character once, use it everywhere
Bring a track and one reference image. Save your character, mention it in the composer, and it conditions every keyframe.
Open the studio