TryH3

Reference to video

Guide the paid full render with image, video and audio references while keeping one written direction.

Output controls

Set the model, duration and frame before previewing and rendering.

768P

Hailuo 3 · 768P or native 2K full video with synchronized audio and multimodal reference control.

Reference limits depend on the selected model. Audio cannot be used by itself.

0/7000

Verified H3 outputs

Start from a real MiniMax example

Six MiniMax-published prompt and output pairs. Reference templates state exactly how many source images you still need to add.

4s5 sec15s

Visual review uses 0 credits · Full render uses 25 credits after purchase

Describe the scene before generating.

How MiniMax H3 reference to video works

Reference to video is what separates H3 from single-input models. Instead of treating an image, a video clip and an audio file as three separate tools, H3 reads them as one context alongside your written direction. That means you can hold a character's face from a photo, borrow camera behaviour from a video, and match the rhythm of an audio track — in the same generation.

What each reference type controls

  1. 1

    Image reference: identity and style

    Use it to lock a face, an outfit or a visual style. This is the reference that makes a character survive across multiple shots and multiple generations.

  2. 2

    Video reference: motion and camera

    Supply a clip whose movement you want echoed. H3 borrows the camera behaviour and pacing without copying the subject.

  3. 3

    Audio reference: rhythm and timing

    Give the model a track and it will time cuts and motion against it. Useful for music-driven edits where the picture has to land on the beat.

Combining references without confusing the model

References compete when they contradict each other. These rules keep them working together:

  • One reference per job. Image for identity, video for motion, audio for timing — do not use two images fighting over the same face.
  • Keep the written prompt as the tiebreaker. When a reference and the prompt disagree, the prompt should state which one wins.
  • Crop image references tight on what you actually want held. Background detail in the reference leaks into the output.
  • For multi-shot character consistency, reuse the same image reference across every generation rather than re-describing the character in words.

A worked example

With a tight portrait supplied as the image reference, this prompt produces a consistent character across a new scene:

The referenced character walks through a rain-soaked night market, slow tracking shot from the side, neon signage reflected in wet pavement, same face and jacket as the reference, ambient rain and distant vendor calls.

Naming the reference explicitly ("the referenced character", "same face and jacket") tells H3 which parts to carry over and which to generate fresh. Without that, the model has to guess how much of the reference is instruction and how much is inspiration.

Frequently asked questions

What reference types does MiniMax H3 accept?
Image, video and audio. Image references hold identity and style, video references guide motion and camera behaviour, and audio references drive timing and rhythm. They can be combined in one generation.
How do I keep the same character across multiple videos?
Reuse the same image reference in every generation rather than describing the character in words each time. A tight, well-lit portrait crop works better than a full scene.
Can I use several references at once?
Yes, but give each one a distinct job — image for identity, video for motion, audio for timing. Two references competing over the same attribute produces a blurred compromise.