Native 2K output
H3 renders at up to 2K resolution natively. Frames stay crisp on the big screen without upscaling tricks.
MiniMax H3 video generator
Create a free static first-frame concept from text or an image. When the direction feels right, unlock MiniMax H3 motion, native 2K and synchronized audio.
Nine hand-picked motion studies. See why each made the cut, play the original sound when available, then reuse the direction in MiniMax H3.
A production-focused AI video model for turning direction, frames and references into finished shots.
H3 renders at up to 2K resolution natively. Frames stay crisp on the big screen without upscaling tricks.
Every clip includes synchronized ambience and effects, matched to the movement in the scene.
Start from a prompt, a first frame, or both first and last frames. H3 animates everything in between.
Render for cinema, feeds and Shorts in 16:9, 9:16, 1:1, 4:3, 3:4 or 21:9.
Multimodal reference keeps faces, outfits, and styles stable across shots and scenes.
Videos made with subscription or purchased pack credits are yours to use in ads, client work and monetized content.
Built for real production work
Switch the lane to see the visual language, practical controls and a ready-to-run prompt recipe.
MiniMax H3 / film
Use MiniMax H3 to explore camera movement, atmosphere and action before committing to a production setup.
From idea to finished clip in three steps.
Write a prompt or upload a first frame. Pick a ratio and a duration from 4 to 15 seconds.
1Review a free static concept of the opening composition before spending credits on video rendering.
2When the direction looks right, unlock the full H3 render with motion and synchronized audio, then download the MP4.
3MiniMax H3 video generator guide
MiniMax H3 is an open, general-purpose multimodal video model released by MiniMax on July 31, 2026. It reads text, images, video and audio as one context instead of treating every input as a separate tool. TryH3 turns those verified model capabilities into a focused browser workflow: begin with text to video, animate an image, control a first or last frame, or direct a new shot with multimodal references. The sections below explain what each mode actually does, which controls are available, and where every technical claim comes from.
Open the MiniMax H3 generatorRead the full H3 prompt guideText-to-video mode starts with one required prompt and no media input. Describe the subject, action, environment, camera behavior, visual style and sound you want; then choose a duration and aspect ratio. MiniMax’s H3 documentation lists six fixed ratios for text generation: 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16. That makes the same MiniMax H3 AI video generator useful for widescreen concepts, square feeds and vertical shorts. Clear relationships matter more than a pile of adjectives, so explain who does what, where the camera moves and what the audience should hear.
Image-to-video mode combines a prompt with a first frame, a last frame, or both. A first frame gives H3 the opening composition to bring to life. A last frame defines where the shot should arrive. Using both asks the model to create a natural transition between two supplied moments. The official API treats this as a distinct mode from reference generation, and the source image determines the output aspect ratio automatically. Use a clean, well-composed image and write about movement rather than redescribing every visible detail: subject motion, camera direction, environmental motion and the intended end state are the most useful instructions.
Reference-to-video is for direction that needs more context than a single opening frame. Official MiniMax documentation says references can guide a character, motion, camera style, visual style, voice or editing rhythm. H3 accepts up to nine reference images, three reference videos and three audio clips, with a maximum of twelve files in one request. Reference audio cannot be used alone; it must be accompanied by an image or video. Tell the model exactly how each reference relates to the requested result—for example, use one image for identity, one video for camera movement and one audio clip for voice timbre. This relationship-based prompting matches H3’s unified multimodal design.
MiniMax’s current H3 API documents two output resolutions: 768P and 2K. TryH3 defaults to 768P for a lower-credit full render; choose 2K when you need the final native output. Duration is selectable in whole-second values from 4 through 15 seconds. Text-to-video supports six fixed aspect ratios; image-to-video follows the supplied frame; reference generation can use a fixed ratio or adaptive selection. Every request needs a non-empty text prompt, which may contain up to 7,000 characters. The model’s launch paper says video can be generated with native stereo sound and highlights instruction following, accurate text and brand rendering, and video-to-video motion transfer. These are model capabilities, not guarantees that every generation will be perfect, so plan to review and iterate.
A useful prompt reads like compact direction: subject and setting first, action second, camera and composition third, then lighting, style, pacing and sound. State timing or shot order when it matters. For image to video, describe how the existing frame changes instead of asking H3 to rebuild it. For reference generation, name the role of each asset in plain language and avoid conflicting instructions. If readable words or a logo matter, quote the exact text and explain when it appears, while still checking the result before publishing. Generate one clear idea first, identify the weakest element, and revise only the instruction that controls it.
A signed-in user can create a free static first-frame concept preview before purchasing a full video render. In text-to-video mode, TryH3 creates an opening-frame image at the selected aspect ratio. In image-to-video mode, the uploaded opening frame becomes the preview. The preview is not a playable video and no MiniMax H3 video job is submitted at this stage. Selecting the locked play control shows the paid render options; after Stripe confirms payment, TryH3 submits the saved settings and opening frame to MiniMax H3. The studio shows the full-render credit quote before checkout, and failed paid video jobs follow the normal automatic credit-refund flow.
Start with a one-time credit pack: no subscription, no auto-renewal, and 12 months to use your credits. Subscriptions remain available for recurring use.
One-time purchase • No auto-renewal • Credits valid for 12 months
For a few extra renders without a subscription
First purchase only · One-time payment · No auto-renewal
For a focused production sprint
First purchase only · One-time payment · No auto-renewal
For larger one-off projects
First purchase only · One-time payment · No auto-renewal
Everything you need to know before your first generation.
TryH3 is an AI video generator powered by MiniMax H3. Type a prompt or upload an image and H3 renders a native 2K clip with synchronized audio in about 1-3 minutes.
Subscription credits renew each billing period and expire at that paid period’s end. One-time credit packs remain usable for 12 months from purchase. Each generation is quoted from its model, duration, resolution, audio tier and reference inputs; failed jobs are refunded automatically.
Clips run 4-15 seconds at up to native 2K resolution, with sound. Choose from 16:9, 9:16, 1:1, 4:3, 3:4 and 21:9.
Yes. Switch to Image to Video and upload a first frame. H3 animates it while your text prompt directs the motion.
Yes. Videos generated with subscription or purchased pack credits include commercial use for ads, client work and monetized content.
Usually 1-3 minutes depending on duration and queue load. Your video appears on the generator page as soon as it is ready.
No card is required for the static concept preview. Unlock the full MiniMax H3 render only when the direction feels right.
Create free preview