We're already testing H3, MiniMax's new video model, in GPTunneL.
The first tests give a clear impression: H3 is built for controllable video, not just a pretty frame. For AI video, that's exactly what matters — not in a demo reel, but in real production work.
MiniMax H3 is set to launch in early August and will compete with ByteDance's Seedance 2.0. H3 is expected to be noticeably cheaper, which could quickly make the model attractive to teams that generate video regularly.
What MiniMax H3 is
MiniMax H3 is a multimodal video generation model. In MiniMax's materials it's labeled the H3 Multimodal Video Generation Model, with the API model listed as MiniMax-H3.
The core idea is one API for several scenarios: text-to-video generation, keyframe control, multimodal reference support, and video or audio continuation.
MiniMax puts particular emphasis on natively synchronized audio — sound is meant to be part of the output itself, not a layer bolted on after generation.
First impressions
The first impression is good: H3 feels more controllable than many video models on first contact. It loses the intent of a prompt less often and holds a scene together within the given idea more consistently.
This shows up especially in tasks where the video needs to be understandable, not just pretty — for example, when a character needs to perform a specific action, an object needs to stay recognizable, and a scene needs to develop without abrupt breaks in meaning.
This is still a first look, not a final verdict. But H3 already looks like a model worth watching.
What H3 can do
H3 comes with four key modes:
- Text to Video: a single prompt generates a dynamic clip with synchronized sound, 5-15 seconds long;
- Keyframe Control: you can upload the first and last frame to lock in the start and end of a scene, and the model fills in the motion between them naturally;
- Multimodal Reference: the model supports up to 9 images, 3 videos, and 3 audio references to keep an object or character consistent across scenes;
- A/V Continuation: you can continue an uploaded video or audio track, with the result meant to stitch seamlessly onto the source material; the claimed duration is up to 20 seconds.
That makes H3 interesting for more than just generating clips from scratch. The model should also work well for more controlled tasks: working with references, a start and end frame, continuing existing footage, and synchronized sound.
Why the comparison with Seedance 2.0 matters
Seedance 2.0 remains one of the main benchmarks in AI video right now. It's often discussed for its dynamic scenes, motion physics, character stability, and the quality of short commercial clips.
That's why comparing H3 specifically to Seedance 2.0 says a lot about where MiniMax is aiming. This isn't just another video model for experiments — it's a bid to be a working tool for creative and production tasks.
Price matters here almost as much as quality. In AI video, a good clip often only comes after several attempts, so the cost of each generation directly affects whether a model can be used regularly.
For H3, that's important context: if the quality holds up in practice, a more accessible price could end up being not a side note, but one of the model's main selling points.
Where H3 could be useful
H3 looks especially promising for creative production:
- short ad clips;
- social video for TikTok, Reels, Shorts, and similar formats;
- visual concepts for brands and products;
- storyboards and previsualization;
- generating scenes from a text description;
- working with image, video, and audio references;
- continuing existing video or audio;
- clips where syncing visuals and sound matters;
- quick tests of different visual directions.
It's worth paying particular attention to motion physics, human actions, and object consistency within a scene — these are often exactly what separates a pretty demo from a model you can actually use every day.
What this means for the market
H3 is interesting as more than just another video generation model. MiniMax is positioning it right where Seedance 2.0 is currently strong: short dynamic clips, scene control, reference-based workflows, and treating sound as part of the result.
If H3 lives up to its first impressions and keeps its more accessible price, competition in AI video is set to get noticeably more practical. For teams, this isn't just about frame quality — it's about the cost of every iteration: the cheaper it is to try different variants, the easier it is to fold video generation into regular work.
For GPTunneL users, this will be a convenient way to try H3 alongside other video models and quickly see where it actually delivers.
