Choose one job
Define the clip's single purpose: a product reveal, dialogue beat, motion study, transition, edit, or narrative payoff.
Turn an idea into a shot-by-shot video brief. Choose your workflow, define visible action, direct camera and sound, then copy a template built for MiniMax H3.
Work from intent to constraints. This order makes conflicts easier to spot and revisions easier to evaluate.
Define the clip's single purpose: a product reveal, dialogue beat, motion study, transition, edit, or narrative payoff.
Name the identity, product geometry, wardrobe, logo, location, or layout that must not drift during motion.
Use active verbs in time order. Describe what the viewer can see instead of internal thoughts or abstract mood alone.
Specify framing, camera movement, speed, lighting, dialogue, effects, ambience, and non-diegetic music.
State where the subject, camera, and key object finish. A clear landing point gives the clip a readable payoff.
Write a video brief, not a still-image caption. Tell the model what changes, when it changes, what the camera observes, and what the viewer hears.
Assign one clear job to every image, video, or audio file.
Describe a starting state, visible actions, cuts, and the ending state.
Set framing, movement, speed, lighting, texture, and capture style.
Name the speaker, exact line, effects, ambience, and music direction.
State what must remain fixed and what should not appear.
Reference roles + timeline of visible action + camera and visual language + dialogue and sound + preservation rules + final frame
Choose the mode that matches your inputs. Image-to-video frame roles and reference-to-video roles cannot be mixed in the same V2 API request.
Text-to-video needs the whole world: subject, setting, visible action, camera, lighting, sound, and ending. Keep a short clip focused on one central idea.
integrated_multimodal_description: [Shot 1] A [specific subject] in [specific environment]. [Starting state]. The subject [visible action in chronological order]. The camera begins with [framing] and [camera movement + speed]. [Lighting, material, and capture style]. End with [final composition or resolved action]. overall_soundscape: [ambient sound], [event-linked sound effects], and [dialogue with named speaker if needed]. non_diegetic_music: [instrumentation, mood, tempo, and when the cue changes—or state no music].
API settingFor text-only generation, choose a concrete ratio such as 16:9 or 9:16. The V2 API does not allow adaptive for T2V.
These templates show the level of specificity to aim for. The media illustrates relevant MiniMax H3 use cases; generation results can vary with references, settings, and provider implementation.
A compact narrative beat with a clear path, one spoken line, environmental motion, and a defined final frame.
A courier in a dark green coat crosses a crowded night market in light rain. The camera tracks backward at chest height as she moves through steam and red lantern reflections. At 00:04.500, she stops under an awning, turns toward camera, and says in a low, urgent voice: "We leave before dawn." End on a tight close-up as a scooter passes behind her. overall_soundscape: rain, distant scooters, crowd murmur, and fabric movement. non_diegetic_music: restrained low strings, fading under the dialogue.
An image-to-video template that prioritizes product geometry, readable branding, controlled lighting, and a clean landing frame.
Preserve the exact bottle shape, label position, cap, glass material, colors, typography, and proportions from Picture 1. The bottle remains centered while the camera makes a slow clockwise orbit at label height. A narrow cyan light travels across the glass as fine mist drifts behind it. Keep the label sharp and facing the camera whenever visible. End on a clean front-facing hero shot. Sound: a soft glass tap and restrained room tone. No dialogue.
A multimodal brief that separates identity, movement, environment, and voice so each reference has a distinct purpose.
Image 1 fixes the lead character's face, black bob haircut, silver jacket, and red gloves. Image 2 fixes the rainy rooftop environment and blue-magenta lighting. Video 1 supplies the running motion and handheld camera rhythm only. Audio 1 supplies the lead's voice timbre. Create a 9:16 chase beat. She runs toward the roof edge, looks back once, and says: "Keep the signal alive." Preserve her identity and wardrobe throughout. Do not copy the performer or location from Video 1.
A motion-design prompt with explicit text behavior, transition order, interaction states, and a ban on garbled interface copy.
Create a 16:9 product interface demo on a dark graphite background. The headline "DESIGN IN MOTION" slides down and holds for two seconds. A mobile dashboard enters from the right, then three data panels expand in sequence as the cursor selects "Preview." Keep all requested English text clean and legible. Use crisp mask reveals and hard cuts, not liquid morphs. Add subtle interface clicks and a low electronic pulse. End with the dashboard centered and the headline unchanged.
Specific direction beats adjective stacks. Use one primary camera idea, connect sound to visible events, and keep the action possible within the selected duration.
For multi-beat clips, state what happens first, what changes, and where the scene ends. Introduce later shots with a clear timestamp or transition.
[Shot 2] At 00:04.500, cut to a tight close-up as the bridge shakes.
"Camera moves" is vague. Name the framing and one controlled motion: a slow low-angle dolly, a fast handheld push, or a wide orbit with restrained speed.
Begin in a medium profile, then make a slow clockwise orbit at shoulder height.
Name who speaks, how the voice sounds, and whether the speaker is on screen. Keep lines short enough to fit the scene without rushing the action.
The mechanic, on screen, whispers: "Let's see if you remember the sky."
Use the soundscape for dialogue, footsteps, impacts, and room tone. Use non-diegetic music for score, instrumentation, tempo, and cue changes.
overall_soundscape: light rain and distant traffic. non_diegetic_music: no music.
Pair a shot size with one primary movement, then add speed or direction only when it changes the result you want. This gives the model clearer staging than a loose request to “make it cinematic.”
A useful default is to name the shot size before the action.
Use familiar cinematography terms, then clarify pace and direction.
Combined example
Extreme close-up of coffee being poured into a white ceramic cup on a marble countertop, morning light through frosted glass. Camera slowly orbits counterclockwise while rack focusing from steam to the cup surface. Rich coffee pour sounds, quiet kitchen ambience. Warm golden color grade. 2K.
Reference-to-video works best when each file controls a different part of the result. State the role in plain language before the timeline.
| Input | Useful role | Prompt language | Documented V2 API limit |
|---|---|---|---|
| Reference image | Identity, product shape, wardrobe, location, layout, or art direction | "Image 1 fixes the lead's face and wardrobe." | Up to 9; use reference_image |
| Reference video | Movement, performance, camera path, timing, edit rhythm, or source footage | "Video 1 supplies movement and camera timing only." | Up to 3; each 2–15 seconds; total video duration up to 15 seconds |
| Reference audio | Voice timbre, delivery, pacing, ambience, music, or rhythm | "Audio 1 supplies the lead's voice timbre only." | Up to 3; each 2–15 seconds; total audio duration up to 15 seconds |
| First frame | Exact opening composition for image-to-video | "At 0.00 seconds, Picture 1 is fully referenced." | One first_frame image |
| Last frame | Exact landing composition or transition endpoint | "At the end, Picture 2 is fully referenced." | One last_frame image paired with the first frame |
Each rewrite turns a general idea into visible action, a camera decision, and an ending the clip can plausibly reach.
Beautiful cinematic video of a woman walking.
Medium tracking shot of a woman in a cream linen dress walking along a coastal cliff path at golden hour. Camera tracks alongside at walking pace. 2K.
Close-up of coffee being poured. Morning light. Warm.
Extreme close-up of espresso being poured into a white ceramic cup on a marble surface, morning window light. Camera holds static then slowly pushes in at 5s. Audio: rich pour sounds, quiet kitchen ambience. 2K.
A woman running through a forest at dawn. 2K.
Medium tracking shot of a woman trail running through a dense green forest at dawn, mist rising. Camera tracks alongside at running pace, handheld energy. Audio: rhythmic footsteps on earth, birdsong, gentle instrumental score rising — no dialogue. 2K.
Shot 1: Wide shot of city. Cut to Shot 2: Medium shot of character entering. Then close-up of their face.
Shot 1 [0–4s]: Wide — city skyline at dusk. Shot 2 [4–10s]: Medium — character enters a building lobby. Shot 3 [10–15s]: Close-up — face in the elevator reflection, determined expression.
A futuristic city at night with neon signs and rain.
Wide establishing shot of a futuristic city at night. Rain begins to fall, neon signs flicker on in sequence, and a hovercar descends through the frame. Camera slowly cranes up to reveal the full skyline. Audio: rain, distant traffic, electronic hum. 2K.
Use Image 1, Video 1, and Audio 1 to make a video.
Image 1 fixes the lead character's face, black bob haircut, and silver jacket. Video 1 supplies the running motion and handheld camera rhythm only. Audio 1 supplies the lead's voice timbre. Create a chase scene through a rain-lit alley. Preserve her identity throughout.
Change one variable at a time. Keep the parts that worked, revise the smallest conflicting instruction, and compare the next generation against the same goal.
| Problem | Likely cause | Prompt fix | Example correction |
|---|---|---|---|
| The first beat consumes most of the clip | Every shot has a range, or later transitions are vague | Describe Shot 1 without a timestamp; introduce later shots at exact transition times | "[Shot 2] At 00:04.500, cut to…" |
| The wrong character says the line | The dialogue is not bound to a visible speaker | Name the character, on/off-screen state, voice quality, and exact line | "Mara, on screen, whispers: 'Wait.'" |
| The face or product drifts | Motion direction appears before preservation rules | Lock identity, shape, wardrobe, label, and colors before the action | "Preserve the exact face, jacket, logo, and bottle geometry." |
| Reference media conflicts | Two files compete for identity, style, or movement | Assign a single job to each reference and remove redundant files | "Video 1 supplies camera rhythm only; do not copy its subject." |
| Unexpected background music appears | Music direction is missing or ambiguous | State the soundscape and non-diegetic music separately | "non_diegetic_music: no music." |
| The shot feels static | The prompt describes appearance but no visible change | Add active verbs, environmental reaction, camera movement, and an ending state | "The doors open, steam rolls out, and the camera pushes through." |
| The result tries to do too much | Too many locations, actions, styles, or camera ideas for the duration | Reduce the clip to one purpose and one primary camera idea | Move the second location into a separate generation. |
Use these terms as building blocks, not as a keyword pile. Choose only the language that clarifies what the viewer should see or hear.
Start with reference roles, then describe timed visual beats, camera direction, sound, constraints, and the final frame. Use active verbs and visible changes instead of a list of style words.
The MiniMax video generation V2 API documents a maximum of 7,000 characters for each required text prompt. Shorter prompts can work when the shot is simple.
The V2 API lists integer durations from 4 through 15 seconds and resolution options of 768P and 2K. Available ratio behavior depends on the workflow.
Yes. MiniMax H3 produces audiovisual output and accepts audio references for reference-to-video. Direct the speaker, exact line, effects, room tone, and music separately.
No. The official V2 API describes image-to-video roles and reference-to-video roles as mutually exclusive within one request.
The V2 API documents up to nine reference images, three reference videos, and three reference audio files. File formats, dimensions, sizes, durations, and the total request size also apply.
No. You can use these templates with a compatible hosted MiniMax H3 interface. Local ComfyUI workflows may require separate software, model files, and suitable hardware.
Describe the intended room tone and effects in overall_soundscape, then state non_diegetic_music: no music. Results can still vary, so confirm the output before production use.
Commercial use depends on the license and terms of the service you use. Review current provider terms and confirm that you hold the necessary rights to every uploaded image, video, voice, and audio reference.
Credit pricing is provider-specific and can change. Check the current MiniMax H3 pricing page before generating; this guide does not publish an unverified credit rate.
Move from the core guide into a narrower task when you need more examples, setup details, or pricing context.
Choose a template, tailor the visible action and reference roles, then test the prompt in the MiniMax H3 generator.
Start with MiniMax H3