Create Video
Model introductionMiniMax H3 · MiniMax-H3 · Open multimodal video

An overview of MiniMax H3, the multimodal video model behind directed moving images

MiniMax H3 is MiniMax’s next-generation open-weights, general-purpose multimodal video model. It understands text, image, video, and audio inputs in a unified way—and turns them into controllable 2K clips instead of one-off silent sketches.

Official capability signal

Generation, reference creation, and editing in one multimodal family.

MiniMax documents MiniMax H3 as a unified model for inventing shots from text, locking first/last frames, and guiding results with reference media.

T2VA
Text to video
I2VA
First / last frame
R2VA
Omni references
API
Video generation V2
Model ID
MiniMax-H3
Output
Native 2K
Duration
4–15 seconds
Inputs
Text · Image · Video · Audio
Prompt budget
Up to 7,000 chars
Omni refs
Up to 12 media files
Visual direction

Shot ideas that match how MiniMax H3 is meant to be used

These local covers stand in for the kinds of directed clips H3 is built for: performance, product motion, environment, fashion, and camera-led storytelling.

Close shot
Performance

Portrait motion study

A single subject, readable expression change, and restrained camera path.

Cinematic street
City night

Tracking through rain light

Environment, practical lights, and one dominant movement through the frame.

I2VA style
Product

Controlled material reveal

First-frame identity stays stable while light and camera do the work.

Landscape motion
World building

Wide environment energy

Scale, weather, and a final wide beat instead of a crowded plot.

Editorial
Fashion

Wardrobe and body line

Silhouette continuity matters more than stacking five style tags.

Narrative unit
Story beat

One scene, one job

Short clips work best when they resolve a single visual intention.

Short answer

Not just another 1080p clip model.

Earlier Hailuo generations excelled at short, physics-aware motion. MiniMax H3 widens the contract: longer windows, native 2K, and multimodal references that can carry face, motion, camera, style, or voice into the next take.

01

Multimodal by design

Text is the director’s brief. Images lock composition. Video references can guide motion or rhythm. Audio references can influence voice texture when paired with visual media.

02

Control without rebuilding identity

Reference generation is for the shots where consistency is the product: recurring characters, product lines, branded motion language, or a known camera feel.

03

A longer useful window

Documented durations run from 4 to 15 integer seconds—enough for a setup, action, and exit without forcing every idea into a six-second punchline.

04

Open general-purpose positioning

MiniMax presents H3 as an open-weights general-purpose multimodal video model, not a single-purpose effect filter.

Capability map

What the public MiniMax H3 release emphasizes

Read these as production levers, not marketing adjectives.

01

Text-to-video from a clean brief

Invent subject, setting, motion, and camera from language when you do not yet have a locked frame.

02

First-frame animation

Start from a designed still and spend the prompt on movement, light change, and camera path.

03

First and last frame control

Direct a transition when the ending composition is as important as the opening.

04

Omni reference generation

Combine up to 9 images, 3 videos, and 3 audio files (12 total) to guide identity, motion, style, or voice.

05

Native 2K delivery

Documented output resolution is 2K, so masters can enter edit timelines without an automatic upscale step.

06

Flexible framing

Common aspect ratios are supported, with adaptive geometry when media defines the frame.

Workflow ownership

Choose the starting point before writing the first line

The MiniMax H3 interface language on this site groups creative starts. The official API encodes them as roles inside one V2 content array.

Starting pointWhat you provideOfficial modeBest for
Text to videoPrompt onlyt2vaInventing a scene from written direction
Image to videoFirst frame + prompti2vaAnimating a controlled visual opening
First and last frameTwo images + prompti2va + last_frameDirecting a transition toward a chosen ending
Omni referenceImages / videos / audio + promptr2vaCarrying identity, motion, style, or voice

First/last-frame roles and reference roles are mutually exclusive in the public V2 contract.

Model notes

Useful constraints before designing a prompt

These numbers are summarized from MiniMax’s public MiniMax H3 documentation and should be re-checked before a permanent product promise.

Model name

MiniMax-H3

Use this exact ID in API requests.

Output resolution

2K

Documented native delivery class for H3.

Duration

4–15 seconds

Integer values only.

Prompt length

≤ 7,000 characters

Enough for a full shot brief, not a novel.

Reference images

≤ 9

Each within documented size and aspect limits.

Reference videos

≤ 3 clips

2–15s each, total ≤ 15s.

Reference audio

≤ 3 clips

Must be paired with image or video input.

Mixed media

≤ 12 files

Request body preferably under 64 MB.

Where it fits

Good jobs for MiniMax H3 short-form video

The strongest workflows give one clip a focused job, then assemble multiple results through editing or campaign systems.

Previsualization

Film and story concepts

Explore camera, blocking, light, and atmosphere before a physical shoot.

Campaign production

Product and social motion

Turn a controlled product frame into short movement studies and hooks.

Character systems

Consistent performers

Use references when the same face, wardrobe, or voice needs to reappear.

Transition design

First-to-last storytelling

Direct openings and endings when continuity is the creative point.

Prompt framing

A practical MiniMax H3 prompt structure

Write the shot in playback order. Details earn their place only when they change what the viewer sees or hears.

01
Scene

Name the format, place, and mood

A night street ad, a quiet portrait, a product hero, or a stylized world—before dialogue or camera flourishes.

02
Subject

Define who or what carries the shot

Appearance only where it affects identity. Reference media when consistency matters more than prose.

03
Motion + camera

Describe the change over time

One dominant action and one clear camera path usually beat a stack of competing moves.

04
Sound + ending

Close the beat

If voice, ambience, or an exit frame matters, say so. End on a composition the viewer can remember.

Model questions

A precise way to describe MiniMax H3

What is MiniMax H3?

MiniMax H3 is MiniMax’s open general-purpose multimodal video model. The public API model ID is MiniMax-H3. It supports text-to-video, first/last-frame image-to-video, and reference-based generation.

Is MiniMax H3 the same as Hailuo 2.3?

No. Hailuo 2.3 is an earlier public video family known for strong short-clip motion around 1080p and shorter durations. MiniMax H3 is the newer multimodal generation with documented 2K output, 4–15s windows, and omni references.

Is MiniMax H3 also called Hailuo 3.0?

Community and partner writeups often use Hailuo 3.0 as an informal alias. For integration, trust the official model name MiniMax-H3 and the V2 documentation.

How do I generate MiniMax H3 videos from my product?

Use the MiniMax H3 API on minimaxh3.video for production tasks, or start in the browser generator when you want to explore shot direction first.

Start with a focused shot

Choose the source. Direct the movement.

Pick text, a first frame, a first-and-last pair, or omni references—then write one clear job for the clip.