MiniMax H3 Video Model
The one model here that will not hand you a silent clip — H3 generates stereo sound as part of the picture, not after it.
- 32 kHz stereo, always on
- 4–15 seconds at 768p or 2K
- Six ratios, 21:9 included
- End frame on its own
MiniMax H3 is a multimodal generation model — MiniMax describes it as one that “understands unified context across text, images, video, and audio, generating video with native stereo sound.” In practice that means the audio is not a second pass: dialogue, effects and music are modelled together with the picture and arrive in the same clip, at 32 kHz stereo and 24 fps.
On Labnana it runs from four to fifteen seconds at 768p or 2K. 768p sets the short side to 768 pixels, so a 16:9 clip lands at 1344×768. 2K is not an upscale — MiniMax regenerates the clip in-context at the higher resolution rather than running a super-resolution pass over it, which is why small text and fine detail survive the jump.
What you can do with MiniMax H3 on Labnana
Get a clip that already has sound
Describe what should be heard as well as seen — a line of dialogue, a room tone, a piece of score — and it comes back in the clip. Dialogue is stable in eleven languages including English, Chinese, Japanese, Korean, French, German, Spanish, Italian, Portuguese, Russian and Arabic. There is no audio toggle in the generator because the model has no silent mode.
Pin the shot at either end
Give a first frame, an end frame, or only an end frame. Every other model here needs the first frame; H3 will work backwards from the last one, which is the way to land on a specific closing composition.
Carry a face, a product or a voice across shots
Attach up to five reference images to keep a character or an object consistent. It also accepts reference video and reference audio — up to three clips each, two to fifteen seconds per video clip — so motion and voice can come from material you already have.
Cut for an ultrawide frame
21:9 is on the list alongside 16:9, 4:3, 1:1, 3:4 and 9:16. It is the only model on Labnana that offers it, which matters if the clip has to sit in a letterboxed edit rather than be cropped into one.
MiniMax H3 at a glance
The generator filters itself to whatever the model supports. Credit cost depends on resolution and duration, and is shown live before you submit — so it is not duplicated here.
| Tier | Duration | Resolution | Aspect ratios | End frame | Prompt limit | Best for |
|---|---|---|---|---|---|---|
| MiniMax H3MiniMax | 4–15s, whole seconds | 768p / 2K, 24 fps | 6 (21:9, 16:9, 4:3, 1:1, 3:4, 9:16) | Supported, and works without a first frame | 2,500 characters | Clips that need sound, ultrawide framing |
How it works
Describe the shot and what it sounds like
Write the picture and the audio in the same prompt. A prompt is required for text-to-video and stays required when you attach references, up to 2,500 characters.
Attach frames or references if you have them
A first frame, an end frame, or both. Up to five reference images, three reference videos and three reference audio clips. With an input image the aspect ratio follows that image, so the ratio control steps aside.
Pick length and resolution, then generate
Four to fifteen whole seconds, 768p or 2K. The credit cost updates as you change either one; you can leave the page and come back to the result.
Compare with another model
Where to use MiniMax H3
All modelsThe video generator is one input box — switch models there without losing what you have written. If the clip does not need sound, or needs to run longer than fifteen seconds, the other two models are worth a look.
Frequently asked questions
Generate with MiniMax H3
Four to fifteen seconds, 768p or 2K, sound included. Costs are shown before you submit.