Describing a photo instead of a video
Beautiful scene, zero verbs — the model has nothing to animate, so it invents drift. Fix: every shot gets one observable change, stated in playback order.
MiniMax H3 prompting guide
MiniMax H3 turns one prompt into a finished shot — picture, motion, dialogue and stereo audio together. This guide covers the prompting rules the model actually rewards, and the free Prompt Agent applies every one of them for you in a short chat.
The shortcut
Most ideas don’t fail as ideas — they fail in translation to camera moves, timing and audio fields. The Prompt Agent closes that gap: describe your video in one plain sentence and it hands back a production-ready MiniMax H3 prompt.
Agent conversations never consume credits. Credits are only used after you continue to the generator and confirm the render — with the cost shown first.
The Agent asks only when the answer would change the result. Everything else — camera, lighting, sound, pacing — it decides with professional defaults.
Attach photos, video or audio. Each file is assigned one clear job in the prompt, and the Agent routes you to the right mode: text, frames or full reference.
Start from a cinematic trailer, UGC product demo, talking head, beat edit or twelve more proven structures instead of a blank box.
A free account keeps your briefs, prompts and renders together.
Plain language is enough — no cinematography vocabulary required. Attach references if you have them.
If something would genuinely change the shot — who the subject is, what a reference is for — the Agent asks once, with quick options.
The finished prompt, mode, model and duration carry over in one click, with the credit estimate shown before anything is spent.
The formula
You won’t need all eight every time. But when a render disappoints, the missing element is almost always on this list.
Subject
Name who or what the shot is about — one main subject with a concrete identity. “A gray-haired watchmaker in a leather apron” gives the model something to hold; “a person” does not.
Action
Give the subject an observable change: what they do, what happens to them, what looks different by the final second. A prompt without change describes a photo, not a video.
Environment
Anchor the scene in a place the camera can show — surfaces, weather, time of day, background life. Environment is what makes motion read as real.
Camera
One move per shot, said in camera language: push in, pan, truck, tracking, static — with amplitude and speed when they matter. “The camera pushes in with small amplitude at slow speed.”
Lighting
Say where light comes from and how it behaves: “a single warm desk lamp, shadows swinging as it rocks.” Light direction does more than a stack of mood adjectives.
Style
Set the look once — film stock, animation style, color palette — at the start of the prompt, instead of sprinkling adjectives through every sentence.
Sound
H3 generates stereo audio natively. Keep dialogue with its speaker, put ambience and physical sounds in their own field, and only request background music you actually want.
Ending
Decide the final frame — a payoff, a reveal, a held composition. Prompts that end nowhere produce videos that end nowhere.
Failure modes
Every one of these comes from real prompts. The fix is always cheaper than a re-render.
Beautiful scene, zero verbs — the model has nothing to animate, so it invents drift. Fix: every shot gets one observable change, stated in playback order.
A push-in that also pans while orbiting reads dramatic on paper and renders as jitter. Fix: one camera instruction per shot; change shots when you need a new move.
“A red or blue dress”, “something like…” — the model picks for you, randomly. Fix: decide before you render. A final prompt contains decisions, not options.
H3 ships sound whether you direct it or not — undirected audio means a random soundtrack over your shot. Fix: separate dialogue, ambience and music; write N/A for silence.
A full story arc crammed into a few seconds renders as a blur of half-finished beats. Fix: budget actions to the duration, and only cut when the cut reveals new information.
Worked example
What most people type
A cool cinematic video of a barista making latte art, beautiful lighting, high quality, 4k, trending
The directed version
integrated_multimodal_description: [Shot 1] Warm 35mm film look. A close-up frames a barista’s hands tilting a steel pitcher over an espresso cup; white foam blooms into a rosetta as she finishes the pour. The camera pushes in with small amplitude at slow speed. [Shot 2] At 00:05.000, cut to a static counter-level shot: she slides the cup toward the lens and the rosetta settles. overall_soundscape: Espresso machine hiss fading out, milk swirling against steel, ceramic sliding on wood. non_diegetic_music: N/A
This is the rewrite the Prompt Agent performs on your idea — free, in chat, in under a minute.
Rewrite my ideaWrite it like a directing brief, not a caption: one concrete subject, an observable action, a real environment, one camera move per shot, and explicit audio. If you’d rather not learn the craft, the free Prompt Agent builds the brief from a plain-language sentence.
MiniMax’s guide splits a prompt into three fields: integrated_multimodal_description for the visual timeline, overall_soundscape for physical sound, and non_diegetic_music for score. Short natural-language prompts also work — use the structure when timing, cuts or audio control matter.
No. Camera vocabulary makes results more predictable, but the Prompt Agent exists precisely so you don’t have to learn it — it converts plain language into shot scales, camera moves and audio fields.
Yes — chatting with the Agent costs nothing and never touches your credits. Credits are only spent when you continue to the generator and confirm a render, and the estimate is shown before you do.
Our verified prompt library pairs MiniMax-published H3 videos with the exact prompts behind them, each linked to its primary source — the fastest way to calibrate what a given prompt style produces.
Open the Prompt Agent, describe the shot in your own words, and walk into the generator with a production-ready MiniMax H3 prompt.