MiniMax H3 Max: video faster than real time

MiniMax H3 Max: video faster than real time

MiniMax H3 Max is available in GPTunneL from day one — August 28. Not a week after the hype, but on the day fal announced it: the model is open in Creative Lab, as a block in Workflow, and through the CreativeLab API. It is a version of the open-weight MiniMax H3 that fal post-trained on new data and accelerated alongside its own inference stack. The headline number is a five-second 768p video generated in under three seconds — faster than the clip takes to play.

The record itself is not the most important part to me. If generation really fits into a few seconds, working with AI video stops being a queue of slow attempts and becomes a short loop: change the prompt, see the result, adjust the scene, run the next version.

It is worth separating the sources up front. The three-second speed and the wins in internal comparisons come from fal, not from our own testing. The parameters, prices and modes below are ours, taken from the live platform catalog. There is also an independent Artificial Analysis ranking: it supports the case that H3 Max is a top-tier model, but it does not make the model the absolute leader in every mode.

How MiniMax H3 Max differs from the original H3

The original MiniMax H3 is an open-weight multimodal video model with text-to-video, image-to-video, and natively synchronized sound. H3 Max is not a new MiniMax generation. It is a post-trained and purpose-optimized version of that model.

According to fal's release post, the team introduced a substantial amount of new data during post-training and focused on two areas: prompt adherence and visual quality. At the same time, the engineers optimized inference — the model is trained and served on NVIDIA GB200 NVL72 systems. They kept a speed optimization only if the model held its position in internal preference evaluations.

That is different from taking a finished model and optimizing it in isolation. With H3 Max, the model and the serving system were tuned together. The claimed speed therefore belongs to the complete combination of model, optimized engine, and fal hardware, not to the weights alone.

The H3 Max model and inference system are optimized together to preserve both speed and video quality

This official text-to-video example shows the model building a character, environment, camera movement, and soundscape from one detailed prompt:

What faster than real time actually means

fal says H3 Max generates a five-second 768p clip in under three seconds. By its calculation, that is roughly 35 times the throughput of the official MiniMax H3 endpoint and, on average, 15 times faster than models of comparable quality in fal's internal comparison. The independent Design Arena, which fal cites, also ranks H3 Max first and puts the speed gap with the original H3 at 50x.

Throughput is not the same as the wait a user sees. Queueing, uploading a source image, transferring the finished file, and interface processing can add time. The three-second figure should be read as inference time in fal's infrastructure, not a millisecond-accurate guarantee for every end-to-end request.

Even with that caveat, the order of magnitude changes the product experience. A single idea used to take minutes before you could discover that a character performed the wrong action. In the same time, you can now check several prompt formulations, camera angles, or scene variants.

H3 Max's rapid iteration loop turns one scene into several video variations

In this official image-to-video demo, the model states the release's core idea itself:

This is not streaming video or an endless scene generated continuously in real time. H3 Max still produces a discrete finished clip. The loop between a decision and its result is simply short enough to feel almost interactive.

What the quality evidence says

fal ran its own comparison against twelve leading video models — the named ones include Veo 3.1, Kling 3, Seedance 2.5, Wan 3.0, Gemini Omni Flash, and the original H3. Human evaluators made blind choices across three dimensions: overall preference, prompt understanding, and aesthetics. Results were computed as Bayesian Elo ratings with 95% confidence intervals. H3 Max ranked first in all three categories in that internal study and won the majority of head-to-head matchups.

That should be treated as a vendor result. The methodology is described, but fal selected the prompts, comparison set, and infrastructure.

The independent picture is strong too, but more specific:

  • on the Artificial Analysis image-to-video leaderboard with audio, H3 Max ranks first at the time of publication with an Elo around 1,200, ahead of Seedance 2.0, the original H3, and Gemini Omni Flash;
  • on the text-to-video leaderboard with audio, the model ranks third with an Elo around 1,240. Its confidence interval overlaps the leaders, which puts them in the same statistical top group rather than showing a clear H3 Max win;
  • Seedance 2.5 is not yet present in these independent tables, so Artificial Analysis cannot currently support a comparison between it and H3 Max.

The accurate conclusion is that H3 Max is already among the best video models with audio and currently leads image-to-video. Calling it the best video model overall would go beyond the evidence.

The practical consequence matters more to us than the ranking. Almost the entire list fal compared against sits in our own catalog: Veo 3.1, Kling 3, Seedance 2.5, Wan 3.0, Gemini Omni, and the original H3. So you do not have to take the leaderboard on faith: run the same working prompt through four models from one balance and see which one wins on your material. That is the only comparison that is actually about your task.

Another official example tests fast movement, cloth inertia, and hands interacting with a small object, the kind of scene where aggressive speed optimizations can expose quality loss:

What H3 Max can do in GPTunneL

We did not ship a trimmed-down version — Creative Lab and Workflow expose the full parameter set:

ParameterValues
Resolution480p, 768p (768p by default)
Duration5 to 15 seconds
Aspect ratioauto, 1:1, 3:4, 4:3, 9:16, 16:9, 21:9
Inputtext, first frame, last frame
Promptup to 7,000 characters, with the prompt improver in Lab
Input imageJPEG, PNG, WebP up to 30 MB, 256–5,760 px

Two details that save time. First, there are three modes, not two. Besides text-to-video and animating an image, you can set the first and last frame at once and let the model build the transition between two images. That is the most controllable way to get the motion you want when the final state of the scene matters as much as the opening one.

Second, as soon as you upload a first frame, the aspect ratio switches to auto and is taken from that image. This is a model requirement, not an interface quirk — crop the source in advance if you need a specific format.

Native audio carries over from H3 and is generated together with the visuals. In Workflow, H3 Max is an ordinary block: it can take an image from a previous step — say, a frame generated by Grom Art or Seedream — and pass the finished clip on to lip sync or upscaling. In that kind of chain the model's speed stops being a nice number and starts shortening the whole pipeline, not one step of it.

How much H3 Max costs in GPTunneL

We offer H3 Max on our usual pay-per-use model, with no subscription:

  • $0.10 per second of finished video at 480p;
  • $0.16 per second at 768p.

A five-second clip costs $0.50 at 480p or $0.80 at 768p; the full 15 seconds cost $1.50 and $2.40. For iteration, I would start at 480p: it costs less to check composition, action, and the prompt. Once the scene works, render the final version at 768p.

A separate question is when to pick H3 Max over the regular H3. H3 costs $0.13 per second but outputs 2K. So H3 Max is more expensive per second and lower in resolution: you are paying for iteration speed, not for the picture. The logic is simple — explore ideas and draft on H3 Max, render the final shot for a large screen on H3 or another 1080p model.

These are GPTunneL prices, not a conversion of fal's tariff. You can always check the current total for your duration on the video pricing page.

The limits that still matter

H3 Max speed depends on server-side optimization. Even if fal publishes the weights, faster-than-real-time performance will not automatically transfer to a consumer GPU. A Reddit discussion points to a statement from fal's CTO about plans to release the weights, but at publication time they are not available and there is no release date. The original H3 is open weight; H3 Max is not yet.

The second boundary is resolution and length. A 768p, 15-second ceiling covers ad shots, product animation, social video, storyboards, and exploring visual directions very well. It does not replace editing a longer story or guarantee that a complex multi-character scene will work on the first attempt.

Finally, a leaderboard measures average preference on its own prompt set. Your task may depend on exact dialogue, specific physics, character consistency, or a long continuous action. Comparing the same working prompt across models is still more useful than choosing from rank alone.

What we are doing with our own model

H3 Max set a bar, and we have accepted it for Grom TV, our video model running on our own hardware. We are working on three things at once: bringing generation time down to something comparable with H3 Max, raising image quality, and separately improving prompt adherence — the most painful part, where the gap between models shows up earlier than in aesthetics.

I am not naming dates until there is something to show in a fair comparison. The point is different: our own model exists not to replace the others, but so that a class of scenarios — longer clips with sound, lip sync, your own voice as input — has a foundation that does not depend on someone else's price list and someone else's queue.

Who I would recommend H3 Max to

H3 Max is a good choice when you need short clips with sound and many fast iterations: ad concepts, motion from a source image, a transition between two frames, social variants, previsualization, and prompt testing. The real gain is not that one clip arrives a minute earlier. It is that you can test many more hypotheses in one working session.

Open MiniMax H3 Max in Creative Lab or build a scenario around it in Workflow. There is no mandatory subscription: top up your balance and pay only for the video you generate.