Playback order beats keyword clouds
If the model must choose between five aesthetics and three camera moves, the shot usually collapses. Sequence the decisions the viewer will actually see.
A useful MiniMax H3 prompt is more than a mood board. Give one shot a clear subject, a readable setting, a controlled movement, a camera path, optional sound intent, and an ending—then remove anything that competes with the main action.
Write decisions in playback order. One dominant movement and one clear camera path usually communicate more than a stack of conflicting instructions.
MiniMax H3 can invent a full scene, animate a locked frame, bridge first and last images, or follow omni references. The safest way to guide that mix is a short production brief in playback order: scene, subject, action, camera, sound intent, ending.
If the model must choose between five aesthetics and three camera moves, the shot usually collapses. Sequence the decisions the viewer will actually see.
Text-only asks H3 to invent the world. First-frame asks it to move a known world. Reference mode asks it to inherit signals. Write for the job you selected.
The content can look similar, but the job is different. Pick the mode before polishing language.
Use when MiniMax H3 must invent subject, environment, and opening composition from text. Ratio is required in pure text mode.
[Duration and aspect intent]. [Subject and defining appearance] in [specific scene and time]. [Primary action or change over the clip]. [Camera path and framing evolution]. [Lighting, palette, material, atmosphere]. [Optional dialogue or sound intent]. End on [final composition or emotional beat].Use when the first frame already establishes identity, composition, and the initial aesthetic.
The [main subject visible in the image] [performs one clear action]. [Environment or secondary elements change]. [Camera path] while preserving the subject’s defining appearance and the original visual style. Optional sound: [voice/ambience cue]. End with [final pose or framing].Describe how the first controlled image evolves into the last rather than re-describing both frames in isolation.
Beginning from the first frame, [subject action and environmental transition]. [Camera movement and timing]. Preserve [identity, object, or style constraints] as the composition evolves naturally into the supplied final frame.Map each reference before restating style adjectives. Do not mix reference roles with first/last-frame roles.
Use reference image(s) for [character / product / style]. Use reference video for [motion / camera / rhythm]. Use reference audio for [voice timbre / delivery] if provided. Create a [duration] shot in [scene]. Subject action: [one readable action]. Camera: [path]. Keep referenced identity consistent. End on [final beat].Add detail only when it earns control. Extra adjectives are not free—they compete with motion.
Name the main person, product, creature, or form. Add appearance details only when they affect recognition.
Place the subject in a concrete environment: night street, studio, forest edge, bridge interior, stylized world.
Describe the action and how the scene evolves. Video needs a change over time, not a list of static visual tags.
Choose framing and travel: static, track, push, pull, rise, pan, orbit, or a restrained combination in sequence.
If dialogue, room tone, or voice reference matters, state it. Pair audio references with visual media in r2va mode.
State where the shot should settle: a close expression, a wide reveal, a completed turn, or a still product frame.
MiniMax H3 accepts 4–15 integer seconds. A five-second product reveal and a twelve-second character beat need different action density. Do not pack a three-act story into one short clip.
The words can look similar, but MiniMax H3’s job changes with the media roles you attach.
Spend tokens on appearance, environment, and opening composition. Ratio cannot be left adaptive in pure text generation.
Do not re-describe every pixel. Direct motion, camera, light change, and what must stay consistent.
Write the transition logic: what moves, what morphs, what the camera does between the supplied ends.
Map which asset controls face, wardrobe, motion, camera language, or voice. Avoid first/last-frame roles in the same request.
Uploading media is not enough. The prompt should assign each reference a job.
Say it explicitly when multiple images could compete.
Keep reference clips short and focused—documented limits are 2–15s per clip.
Audio cannot travel alone; pair it with image or video references.
If yes, split into two tasks. The public contract treats those modes as mutually exclusive.
Trim low-value adjectives before trimming the reference map.
If not, the motion often drifts instead of resolving.
Time words help MiniMax H3 understand priority. Use them to stage the shot, not to micromanage every frame.
First second: who we are watching and where we are. Avoid starting mid-chaos unless the first frame already holds the read.
Middle window: the movement that justifies generating video instead of a still.
If you need two camera ideas, sequence them—track then push—rather than stacking them as equal priorities.
Final beat: expression settle, product still, wide reveal, or completed gesture.
These are newly written structural examples. Swap subject and setting while preserving hierarchy.
One subject action, one camera path, and a final emotional beat.
Create an 8-second 16:9 cinematic night shot. A courier in a reflective jacket cycles through a rain-lit city intersection. Water sprays from the tires as neon reflections stretch across the road. The camera tracks beside the bicycle, then gently closes on the courier’s focused expression. Cool blue practicals with restrained magenta accents, wet-surface realism. Soft tire hiss and distant traffic. End on a readable close three-quarter face.Let the first frame own the product design while the prompt directs light and camera motion.
The watch remains centered as a narrow band of warm light travels across the brushed metal case. Fine mist moves slowly behind the product. Camera trucks slightly right while preserving watch geometry and the dark marble surface. No redesign of the dial. End on a clean three-quarter product angle with the dial fully readable.Describe the evolution between two controlled compositions.
Beginning from the first frame, the model walks forward as studio softboxes dissolve into rainy street practicals. Fabric motion stays natural and identity remains consistent. Camera slowly pulls back while the environment completes the transition into the supplied final frame. End exactly on that last composition with settled posture.Map references before restating style adjectives.
Use reference images for the character’s face and wardrobe. Use reference video for relaxed walking rhythm. Use reference audio for warm mid voice timbre. Create a 10-second medium shot on a windy coastal boardwalk at golden hour. The character walks toward camera, smiles once, and says a short free-spirited line. Camera tracks backward smoothly. Keep identity and wardrobe consistent. End on a stable medium close-up.Most weak MiniMax H3 prompts do not need more words. They need a clearer subject, action, mode, or camera priority.
Replace “cinematic, beautiful, 4K, dramatic” with a subject performing one visible action inside a defined scene.
If the first frame already exists, stop rebuilding wardrobe and background from scratch unless change is intentional.
Map which asset controls face, motion, or voice. Unmapped references become noise.
Choose one dominant path. If a second is necessary, place it later as a sequence.
Give the shot one job. Build longer sequences from several generated clips.
Split the work. The public MiniMax-H3 contract treats those mode families as mutually exclusive.
Read the prompt once as a director and once as a timeline.
Text-only, first frame, first+last, or omni reference—pick one contract.
A viewer should know what to watch before the camera starts moving.
Name place, time, weather, or production context when it changes the visual result.
Describe an action, transformation, expression, or environmental event across time.
Sequence secondary moves; do not stack contradictions.
A clear last beat helps motion resolve instead of drifting.
Natural-language camera direction is enough for many MiniMax H3 shots. When precision matters, keep commands sparse and sequential.
truck lefttruck rightMove the camera laterally while preserving viewing direction.
pan leftpan rightRotate left or right from a relatively fixed position.
push inpull outPhysically move toward or away from the subject.
riselowerRaise or lower the camera without forcing a tilt.
tilt uptilt downRotate the view upward or downward.
tracking shotfollowMove with the subject so environment changes while framing stays readable.
zoom inzoom outChange apparent focal length rather than traveling through space.
static shotlocked offHold the camera steady and let performance carry the clip.
subtle handheldrestrained shakeUse sparingly so subject motion remains readable.
Length is not the goal. Use enough detail to identify subject, setting, movement, camera path, optional sound intent, and ending—then remove instructions that compete with the primary action. The documented limit is 7,000 characters.
The first frame already establishes subject, composition, and style. Focus the prompt on motion, camera behavior, environmental change, and the features that must remain consistent.
Map each reference to a job—face, wardrobe, motion, camera rhythm, or voice—before adding style adjectives. Keep total media within documented caps and never mix reference roles with first/last-frame roles.
You can state spoken intent in the prompt, and reference audio can guide timbre when paired with visual media in reference mode. Keep speech short enough for the chosen duration.
Use the MiniMax H3 API guide on this site and MiniMax’s official V2 documentation for content roles, duration, ratio, and status handling.
Start with the mode, add one useful camera decision, and compare the result before increasing complexity.