TryH3
Sign In

Reference to video

Build from references

Guide the result with image, video and audio references while keeping one written direction.

Output controls

Set the model, duration and frame before rendering.

2K

Hailuo 3 · Native 2K video with synchronized audio and multimodal reference control.

Reference limits depend on the selected model. Audio cannot be used by itself.

0/7000
4s5 sec15s
2KFixed

Watermark

Add the provider watermark to the exported clip.

Credits required: 40

Describe the scene before generating.

Recent renders

New jobs update here while they render.

Sign in to view your creations

Sign in to generate videos and see your history here.

Sign In

How MiniMax H3 reference to video works

Reference to video is what separates H3 from single-input models. Instead of treating an image, a video clip and an audio file as three separate tools, H3 reads them as one context alongside your written direction. That means you can hold a character's face from a photo, borrow camera behaviour from a video, and match the rhythm of an audio track — in the same generation.

What each reference type controls

  1. 1

    Image reference: identity and style

    Use it to lock a face, an outfit or a visual style. This is the reference that makes a character survive across multiple shots and multiple generations.

  2. 2

    Video reference: motion and camera

    Supply a clip whose movement you want echoed. H3 borrows the camera behaviour and pacing without copying the subject.

  3. 3

    Audio reference: rhythm and timing

    Give the model a track and it will time cuts and motion against it. Useful for music-driven edits where the picture has to land on the beat.

Combining references without confusing the model

References compete when they contradict each other. These rules keep them working together:

  • One reference per job. Image for identity, video for motion, audio for timing — do not use two images fighting over the same face.
  • Keep the written prompt as the tiebreaker. When a reference and the prompt disagree, the prompt should state which one wins.
  • Crop image references tight on what you actually want held. Background detail in the reference leaks into the output.
  • For multi-shot character consistency, reuse the same image reference across every generation rather than re-describing the character in words.

A worked example

With a tight portrait supplied as the image reference, this prompt produces a consistent character across a new scene:

The referenced character walks through a rain-soaked night market, slow tracking shot from the side, neon signage reflected in wet pavement, same face and jacket as the reference, ambient rain and distant vendor calls.

Naming the reference explicitly ("the referenced character", "same face and jacket") tells H3 which parts to carry over and which to generate fresh. Without that, the model has to guess how much of the reference is instruction and how much is inspiration.

Frequently asked questions

What reference types does MiniMax H3 accept?
Image, video and audio. Image references hold identity and style, video references guide motion and camera behaviour, and audio references drive timing and rhythm. They can be combined in one generation.
How do I keep the same character across multiple videos?
Reuse the same image reference in every generation rather than describing the character in words each time. A tight, well-lit portrait crop works better than a full scene.
Can I use several references at once?
Yes, but give each one a distinct job — image for identity, video for motion, audio for timing. Two references competing over the same attribute produces a blurred compromise.