MiniMax H3 Max: video faster than real time

MiniMax H3 Max: video faster than real time

MiniMax H3 Max is now available in GPTunneL. It is a version of the open-weight MiniMax H3 that fal post-trained on new data and accelerated alongside its own inference stack. The headline number is a five-second 768p video generated in under three seconds. The model produces the finished clip faster than the clip takes to play.

The record itself is not the most important part to me. If generation really fits into a few seconds, working with AI video stops being a queue of slow attempts and becomes a short loop: change the prompt, see the result, adjust the scene, run the next version.

It is worth separating the sources up front. The three-second speed and the wins in internal comparisons come from fal, not from our own testing. There is also an independent Artificial Analysis ranking. It supports the case that H3 Max is a top-tier model, but it does not make the model the absolute leader in every mode.

How MiniMax H3 Max differs from the original H3

The original MiniMax H3 is an open-weight multimodal video model with text-to-video, image-to-video, and natively synchronized sound. H3 Max is not a new MiniMax generation. It is a post-trained and purpose-optimized version of that model.

According to fal's release post, the team introduced a substantial amount of new data during post-training and focused on two areas: prompt adherence and visual quality. At the same time, the engineers optimized inference. They kept a speed optimization only if the model held its position in internal preference evaluations.

That is different from taking a finished model and optimizing it in isolation. With H3 Max, the model and the serving system were tuned together. The claimed speed therefore belongs to the complete combination of model, optimized engine, and fal hardware, not to the weights alone.

The H3 Max model and inference system are optimized together to preserve both speed and video quality

GPTunneL supports the two main H3 Max workflows: generating video from text and animating a source image. Native audio carries over from H3 and is generated together with the visuals.

This official text-to-video example shows the model building a character, environment, camera movement, and soundscape from one detailed prompt:

What faster than real time actually means

fal says H3 Max generates a five-second 768p clip in under three seconds. By its calculation, that is roughly 35 times the throughput of the official MiniMax H3 endpoint and, on average, 15 times faster than models of comparable quality in fal's internal comparison.

Throughput is not the same as the wait a user sees. Queueing, uploading a source image, transferring the finished file, and interface processing can add time. The three-second figure should be read as inference time in fal's infrastructure, not a millisecond-accurate guarantee for every end-to-end request.

Even with that caveat, the order of magnitude changes the product experience. A single idea used to take minutes before you could discover that a character performed the wrong action. In the same time, you can now check several prompt formulations, camera angles, or scene variants.

H3 Max's rapid iteration loop turns one scene into several video variations

In this official image-to-video demo, the model states the release's core idea itself:

This is not streaming video or an endless scene generated continuously in real time. H3 Max still produces a discrete finished clip. The loop between a decision and its result is simply short enough to feel almost interactive.

What the quality evidence says

fal ran its own comparison against twelve leading video models. Human evaluators made blind choices across overall preference, prompt understanding, and aesthetics. H3 Max ranked first in all three categories in that internal study and beat the original H3.

That should be treated as a vendor result. The methodology is described, but fal selected the prompts, comparison set, and infrastructure.

The independent picture is strong too, but more specific:

  • on the Artificial Analysis image-to-video leaderboard with audio, H3 Max ranks first at the time of publication with an Elo around 1,200, ahead of Seedance 2.0, the original H3, and Gemini Omni Flash;
  • on the text-to-video leaderboard with audio, the model ranks third with an Elo around 1,240. Its confidence interval overlaps the leaders, which puts them in the same statistical top group rather than showing a clear H3 Max win;
  • Seedance 2.5 is not yet present in these independent tables, so Artificial Analysis cannot currently support a comparison between it and H3 Max.

The accurate conclusion is that H3 Max is already among the best video models with audio and currently leads image-to-video. Calling it the best video model overall would go beyond the evidence.

Another official example tests fast movement, cloth inertia, and hands interacting with a small object, the kind of scene where aggressive speed optimizations can expose quality loss:

How much H3 Max costs in GPTunneL

We offer H3 Max on our usual pay-per-use model, with no subscription:

  • $0.10 per second of finished video at 480p;
  • $0.16 per second at 768p.

A five-second clip costs $0.50 at 480p or $0.80 at 768p. For iteration, I would start at 480p: it costs less to check composition, action, and the prompt. Once the scene works, render the final version at 768p.

These are GPTunneL prices, not a conversion of fal's tariff. You can always check the current total for your duration on the video pricing page.

The limits that still matter

H3 Max speed depends on server-side optimization. Even if fal publishes the weights, faster-than-real-time performance will not automatically transfer to a consumer GPU. A Reddit discussion points to a statement from fal's CTO about plans to release the weights, but at publication time they are not available and there is no release date. The original H3 is open weight; H3 Max is not yet.

The second boundary is short-form video. A fast five-second generation is a strong fit for ad shots, product animation, social video, storyboards, and exploring visual directions. It does not replace editing a longer story or guarantee that a complex multi-character scene will work on the first attempt.

Finally, a leaderboard measures average preference on its own prompt set. Your task may depend on exact dialogue, specific physics, character consistency, or a long continuous action. Comparing the same working prompt across models is still more useful than choosing from rank alone.

Who I would recommend H3 Max to

H3 Max is a good choice when you need short clips with sound and many fast iterations: ad concepts, motion from a source image, social variants, previsualization, and prompt testing. The real gain is not that one clip arrives a minute earlier. It is that you can test many more hypotheses in one working session.

Open MiniMax H3 Max in the GPTunneL lab. There is no mandatory subscription: top up your balance and pay only for the video you generate.