TryH3
Sign In

Text to video

Direct a scene from text

Describe the subject, camera, movement and sound, then choose the model that fits the shot.

Output controls

Set the model, duration and frame before rendering.

2K

Hailuo 3 · Native 2K video with synchronized audio and multimodal reference control.

0/7000
4s5 sec15s
2KFixed

Watermark

Add the provider watermark to the exported clip.

Credits required: 40

Describe the scene before generating.

Recent renders

New jobs update here while they render.

Sign in to view your creations

Sign in to generate videos and see your history here.

Sign In

How MiniMax H3 text to video works

Text to video is the fastest way into MiniMax H3. You write one written direction and H3 renders the whole shot — subject, camera movement, lighting and a matching stereo audio track — in a single pass at native 2K. There is no separate sound step and no upscaling stage. The model reads your prompt as a director's brief rather than a keyword list, so the more precisely you describe motion and framing, the closer the result lands.

Writing a prompt H3 can actually direct

  1. 1

    Name the subject and the setting

    Start with who or what is on screen and where they are. "A woman in a neon-yellow jacket at a night market" gives H3 a concrete anchor; "a person outside" does not.

  2. 2

    Describe the camera, not just the scene

    H3 responds directly to camera language. Say "slow dolly forward", "handheld tracking shot" or "locked-off wide". Without it the model picks its own movement and results vary between runs.

  3. 3

    Add the sound you expect to hear

    Audio is generated with the picture, so name it. "Rain on canvas awnings, distant crowd chatter" produces a very different track than silence-by-default.

What changes the output most

Prompts have a 7,000-character budget, but length is not what drives quality. These four choices move the result more than any amount of extra adjectives:

  • Duration: 4-6 seconds holds one continuous action cleanly. 10-15 seconds gives H3 room to cut between shots, which is where multi-shot consistency shows.
  • Aspect ratio: pick before generating, not after. 16:9 for landscape and cinema, 9:16 for Shorts and Reels, 21:9 for widescreen framing.
  • Lighting and time of day: "golden hour backlight" or "overhead fluorescent" changes the entire grade and is cheaper to specify than to fix later.
  • One action per clip. Two unrelated events in one prompt usually produces a blurred compromise rather than both.

A worked example

This prompt produces a consistent multi-shot result rather than one pretty frame:

A young woman in a neon-yellow jacket races through a lantern festival across four shots, urgent handheld tracking, close-ups between wide crowd views, consistent face and wardrobe, vivid night ambience with distant drums and crowd noise.

It works because it fixes four things at once: the subject and wardrobe (so the face stays stable across cuts), the shot count, the camera style, and the audio bed. Remove any one of them and H3 fills the gap with its own choice.

Frequently asked questions

How long should a MiniMax H3 prompt be?
The limit is 7,000 characters, but most strong results sit between 30 and 80 words. Beyond that, extra description tends to compete with itself. Prioritize subject, camera movement, lighting and sound over stacked adjectives.
What video lengths can H3 generate from text?
Between 4 and 15 seconds in a single generation. Shorter clips hold one continuous action; 10-15 seconds gives the model room to cut between shots while keeping the same character.
Does text to video include sound?
Yes. MiniMax H3 renders synchronized stereo audio together with the picture in the same pass — ambience, effects and motion-matched sound. Describe the audio in your prompt to control what you get.