MiniMax has released H3, and the short version is that it isn’t really a video generator. It generates video, but it also takes video in, edits video that already exists, and produces sound as part of the same output. The category label hasn’t caught up with the thing.
That distinction matters more than any individual feature, so it’s worth walking through what the model actually does before deciding whether the framing is marketing or accurate.
It accepts four kinds of input at once
Text, images, video, and audio all enter the same context, and the model reasons across them together rather than processing each separately.
In practice that means responsibilities get split across sources instead of crammed into a prompt. Stills define who a character is and what a product looks like. A reference clip defines how a body moves and how the camera moves with it. An audio sample defines voice and delivery. An existing edit can carry cutting rhythm and visual grade across to new material.
The result is that you specify a shot instead of describing one. Anyone who has spent an afternoon writing increasingly desperate adjectives to keep a character’s face consistent will understand why this is the headline capability rather than a footnote.
It generates sound natively
Every output from Minimax H3 arrives with native stereo audio generated alongside the picture, in the same pass — dialogue, ambient tone, footsteps and impacts, music, and the sync relationships between all of them.
This is the difference between a clip and a plate. Silent output from a video model is one layer of a deliverable; you still need a voice, a sound library, and time spent aligning them. Sync produced as a property of the generation, rather than achieved afterwards, is also why lip movement holds up on hard material like fast rap delivery, where a two-frame error is visible.
It edits existing footage
This is the capability that separates H3 from most of its competition, and it’s why the model currently leads the Artificial Analysis video editing leaderboard ahead of Seedance 2.0.
The supported operations cover most of what a revision request actually consists of:
- Replacing, adding, or removing objects and people
- Changing backgrounds, environments, and lighting
- Adding or adjusting visual effects
- Modifying a character’s motion or performance
- Replacing dialogue, with the mouth re-forming around the new line
- Cloning or transferring vocal timbre
The important property isn’t the list, it’s the locality. Change what you specified and leave everything else — the other actors, the framing, the camera move, the grade — where it was. Models that regenerate the whole scene with your change applied haven’t edited anything; they’ve made a second video that resembles the first, and every approval you’d already won is back in play.
It’s unusually good at screens and type
A quieter capability, and a genuinely differentiating one. Interfaces have historically been the worst-case input for video models: HUDs melt, buttons multiply, labels dissolve into letter-shaped noise as soon as the camera moves.
H3 holds them. Game UI keeps its geometry through a transition, app and web interfaces survive a pan, product screens stay readable at 1440p, and display typography stays intact through a push-in. That makes it viable for a set of jobs previously off the table — game UI and interaction demos, product feature walkthroughs, character PVs and game CG, MV cuts built around large kinetic lyric type, animated posters, e-commerce creative where the packaging has to be legible.
The specs
- Duration: 5–15 seconds
- Frame rate: 24 FPS
- Resolution: up to 1440p
- Ratios: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16
- Prompt limit: 7,000 characters
- Reference set: up to 12 mixed files, with video and audio references capped at 15 seconds each
- Modes: first/last frame (preserves the input image’s ratio) and all-purpose reference (infers the output ratio)
Two of those deserve a comment. 24 FPS reads as a limitation to people used to 60 and isn’t — it’s the frame rate of film and commercials, so output conforms into a real timeline without conversion. And the 7,000-character prompt ceiling is more consequential than it sounds, because it’s enough to direct a shot properly: performance beat, camera timing, lighting, audio intent, what’s in frame at second three versus second nine.
What it doesn’t do
Fifteen seconds is a shot, not a film. A finished piece is still several generations assembled by a person deciding order and rhythm.
There’s no engine export, so UI and motion work produces a target rather than an implementation. Local editing is very good but not unlimited — large edits disturb more of the frame than small ones, and past a certain threshold you’re effectively regenerating with extra steps. And it won’t photograph a real product exactly as it exists, which for some categories is the entire brief.
The honest positioning is that this is a strong tool for a wide band of commercial work and not a replacement for production. Anyone testing MinimaxH3 Free should find the edges within a day or two of real use rather than trusting a launch post about it.
Cost
The last piece is pricing, which does a lot of the persuading. Per-second cost sits substantially below comparable models, Seedance 2.0 included, and since every good clip is the survivor of several discarded ones, cheap iteration functions as a quality mechanism rather than only as a saving. Current rates are on the Minimax h3 pricing page.
The summary
H3 is a video-centred multimodal model that understands text, images, video, and sound together, generates and edits, and outputs audio natively. It’s aimed squarely at commercial production rather than at benchmark scores, and the editing behaviour is the part that will decide whether teams keep it after the first week.






