MiniMax H3 Prompting Guide
MiniMax H3 generates video and synchronized stereo audio together. In Dopamine Girl, H3 currently works as image-to-video: upload a first frame, optionally add a last frame, choose a 3–5 second duration, and describe everything that should happen between those anchors.
This guide adapts the official MiniMax H3 model card, prompt-writing skill, and base prompt guide to the controls available in Dopamine Girl.
Use H3's Three-Part Prompt
H3 prompts work best when they keep these fields in this exact order:
integrated_multimodal_description:
[Describe the visible action, camera, timing, dialogue, and matching sound effects.]
overall_soundscape:
[Describe the continuous environmental sound and spatial audio.]
non_diegetic_music:
[Describe background score that does not come from the scene, or write N/A.]
The first field is the main timeline. It should connect motion, camera direction, and diegetic sound—the sound produced inside the scene—rather than treating picture and audio as unrelated lists.
The other two fields have narrower jobs:
overall_soundscapedescribes the continuous ambient bed, such as rain, traffic, room tone, wind, or crowd noise. One to four sentences is normally enough.non_diegetic_musicdescribes a background score heard by the audience but not produced inside the scene. Keep it to one to three sentences, or useN/Awhen you do not want music.
First Frame Only
With only a first frame, treat the uploaded image as the composition at 0.00 seconds. Start your prompt with this alignment instruction:
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
Then describe motion that grows naturally out of that frame. The image already establishes the subject, clothing, environment, color palette, and initial camera angle, so spend most of the prompt on what changes:
- the subject's action and expression;
- movement in hair, fabric, particles, weather, or background objects;
- camera direction, speed, and travel distance;
- sound effects timed to visible events;
- the state in which the shot should finish.
Example: first frame only
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description:
[Shot 1] The woman looks up from the letter and takes one measured step toward the rain-streaked window. The camera performs a very slow, short dolly-in while keeping her face centered. Her sleeve brushes the wooden desk with a soft rustle; distant thunder follows as blue light briefly brightens the room. She ends still, watching the rain beyond the glass.
overall_soundscape:
Steady rain taps against the window in a quiet interior, with low distant city traffic and a faint room echo.
non_diegetic_music:
Sparse, restrained piano notes enter softly and remain beneath the rain.
First and Last Frame
After uploading a first frame, Dopamine Girl lets you add a last frame. Use both images as endpoints and explain the continuous path between them. For a five-second clip, begin with:
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 5.00-second mark of the target video.
Replace 5.00 with the duration you selected. Keep both references in the same shot unless a cut is genuinely necessary. Describe the intermediate movement explicitly: what initiates the change, how the subject and camera travel, and how they settle into the last frame. This helps avoid a clip that holds the first image and then abruptly jumps to the second.
Example: first and last frame
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 5.00-second mark of the target video.
integrated_multimodal_description:
[Shot 1] Starting from Picture 1, the skateboarder pushes off and glides down the wet alley. The camera tracks backward at the same moderate speed, then arcs gently to the right as she turns. Reflections stretch beneath the board; each wheel clicks lightly over pavement seams. During the final second she slows, rotates her shoulders, and settles precisely into the stance and framing of Picture 2.
overall_soundscape:
Close rolling wheels and light splashes sit over a broad nighttime bed of rain, ventilation fans, and distant traffic.
non_diegetic_music:
A minimal electronic pulse rises subtly through the shot without overpowering the wheel sounds.
Direct Motion as a Short Timeline
A 3–5 second clip has little room for setup. One coherent shot with one main action is usually stronger than several scene changes.
Use [Shot 1] without a timestamp. If you intentionally add a cut, give each later shot an increasing timestamp:
[Shot 1] A cyclist accelerates toward the intersection...
[Shot 2] [2.80s] Cut to a close side profile as she passes the camera...
Keep every described event inside the selected duration. If a five-second prompt asks a character to wake up, cross a room, leave a building, drive away, and reach another city, the model has no clear motion to prioritize.
Specify camera movement precisely
Combine three details:
- Type — pan, tilt, dolly, orbit, handheld track, crane, or locked-off shot.
- Amplitude — slight, short, half-circle, close follow, or wide sweep.
- Speed — very slow, steady, brisk, or rapidly accelerating.
For example, a slow, short dolly-in from a medium shot to a close-up is more actionable than cinematic camera movement.
Direct Dialogue, Sound, and Music
H3 can generate speech and scene audio together with the picture. Assign stable speaker IDs when more than one person speaks, and place dialogue at the point where it occurs:
The shopkeeper (S1) leans closer and whispers <d>[English] The door was already open.</d>
The customer (S2) turns toward the entrance without answering.
Preserve the exact wording and language you want spoken. Do not repeat dialogue in overall_soundscape; use that field for ambience. Put visible signs, captions, or title text in double quotes so they are clearly distinguished from speech.
Keep sound sources spatially consistent. A nearby object should sound close and distinct, while traffic outside a closed room should be muffled and distant. Time one-off sounds alongside the action that causes them: a cup lands and clicks, a door closes and thuds, or thunder follows a flash.
Copy-Ready Templates
First frame template
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description:
[Shot 1] Starting from Picture 1, [subject] [main action]. The camera [type, amplitude, and speed]. [Environment or secondary motion]. [Sound timed to a visible event]. The shot ends with [clear final state].
overall_soundscape:
[Continuous ambience, spatial position, and intensity.]
non_diegetic_music:
[Background score and mood, or N/A.]
First and last frame template
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the [DURATION]-second mark of the target video.
integrated_multimodal_description:
[Shot 1] Starting from Picture 1, [subject and camera begin moving]. [Describe the continuous intermediate action and synchronized sounds]. During the final second, [describe how motion slows or resolves] and settles precisely into Picture 2.
overall_soundscape:
[Continuous ambience, spatial position, and intensity.]
non_diegetic_music:
[Background score and mood, or N/A.]
Final Checklist
- Keep the three field names in the required order.
- Match all timing to the selected 3–5 second duration.
- Describe motion away from the first frame, not just what the image contains.
- With a last frame, explain the physical transition and final settling motion.
- Prefer one main action and one shot for a short clip.
- Define camera type, amplitude, and speed instead of relying on abstract words such as “cinematic.”
- Put synchronized effects beside the action that creates them.
- Keep ambience out of the music field and music out of the soundscape field.
- Use stable speaker labels and preserve dialogue exactly.
- Use concrete visible and audible details instead of broad mood words alone.