AI

MiniMax H3: Hailuo's Open Video Model Is Back in the Game

August 1, 2026
MiniMax H3 title graphic in white text over a purple jellyfish background, with the tagline "Next-Generation Open-Weights General-Purpose Multimodal Video Model."

After a long quiet stretch, MiniMax has come roaring back with H3, the latest generation of its Hailuo video model. It has been roughly nine months since the previous 2.3 release, so H3 is effectively a brand-new arrival rather than a minor update. Officially launched at the end of July 2026, it lands as a general-purpose, omni-modal generation model that many are already calling a serious challenger to Seedance.

What Is MiniMax H3?

H3 is a single model that understands and generates across text, images, video, and audio together, rather than treating each of those as a separate task. It can produce video with native stereo sound at up to 2K resolution and up to 15 seconds in length. MiniMax positions it as commercial-grade, built for real content work in advertising, branding, e-commerce, product design, UI and UX, and gaming.

An Omni-Modal Approach

The headline idea is that H3 is "omni-modal." It can take in a first frame and a last frame, work from up to 12 image references, accept video or audio as input, or combine all of those at once. Instead of picking a narrow tool for each job, you describe the relationship between your reference material and the result in plain language, and the model works out how to blend it. That flexibility is what separates H3 from the more specialized video models that came before it.

An Open Model on the Way

MiniMax also plans to release the model weights in the coming days, subject to applicable laws and regulations. Opening up the weights is meant to support the wider community, improve compatibility with a broader range of AI hardware, and let people build their own customized versions. Video generation has long been dominated by closed systems that iterate slowly, so an open release at this quality level is notable in itself.

Hands-On First Impressions

Early testing suggests H3 is already strong, even while it remains in early access ahead of a wider rollout. A few areas stand out.

Dialogue and Acting

One of the most striking results is how well H3 handles dialogue and subtle performance. In a diner scene, the model delivered a full multi-line exchange between two characters without clipping the dialogue, and it even added small acting touches like a finger tap on the table. That kind of natural, unprompted "business" is exactly what makes a shot feel directed rather than generated.

Multilingual Support

H3 can also perform dialogue in more than one language. The trick is to write the prompt in the language you are targeting rather than in English, at which point the model produces convincing spoken delivery. It is a small but meaningful sign of how much the underlying instruction following has matured.

Motion, Fights, and Consistency

Action sequences reveal a deliberate design tradeoff. Fight scenes in H3 tend to run a touch slower and less frantically than the kinetic output some rival models produce. In exchange, the results stay coherent and stable. Characters do not morph, teleport around the frame, or dissolve into artifacts. That choice, favoring consistency over raw speed, is likely intentional and, for most creative work, a welcome one.

Image to Video

Feeding a still image into H3 shows just how far the model has come in nine months. Running an older reference image through the new system produced a clear jump in quality over the 2.3 version, and this time with sound, which the previous generation could not generate at all.

Working With References

The Omni Model in Practice

The omni side of H3 is where its multi-reference abilities shine. It accepts up to 12 image references and up to three video references, and those video inputs can serve as extensions of an existing shot. In testing, extending a scene worked, though it took careful prompting to keep spatial details consistent, such as where a character sits or what props appear on the table. The takeaway is that the capability is there, and precision in the prompt is what unlocks it cleanly.

Reference Maxing

Pushing the model hard by loading in character sheets, locations, and full storyboards at once is more demanding, and quality can dip. Interestingly, H3 appears to take storyboards very seriously and tends to follow them faithfully rather than overriding them the way some competitors do. That makes careful, intentional storyboarding more rewarding, and it rewards creators who plan their shots deliberately.

Pricing

Official pricing was not fully confirmed at launch, but the early signals point to an aggressive position. MiniMax states that at 2K, H3's per-second price is less than a third of mainstream models, and at 768p it comes in at under half the price of the 720p output from mainstream competitors. Hands-on observation lines up with that: a single 15-second generation at 2K ran around 150 credits, a figure that suggests MiniMax is aiming to undercut rivals meaningfully. Native 2K is offered by default, which makes the value proposition sharper still.

How H3 Was Built

From Specialized to General-Purpose

H3 follows two earlier generations, Hailuo 01 and Hailuo 02. With H3, MiniMax set out to break down the boundaries that used to fragment generative models into dozens of narrow tasks, separate systems for text-to-image, editing, motion reference, voice, sound effects, music, and more. The guiding principle was to unify and generalize across all of those, using natural language as the bridge that ties every task into one open, describable form.

The Core Technologies

Underneath the simple idea sits a genuinely complex system. Without going into the weeds, H3 leans on a handful of named advances: Contextual Omni Representation for understanding mixed inputs, an overhauled H3-VAE that improves efficiency and enables native 2K, an H3-Omni Transformer designed for task generalization, and an In-Context Regeneration approach that produces high-resolution output by having the model refine its own results rather than relying on a separate upscaler. A full technical report is promised soon.

The Bigger Picture

MiniMax frames language, images, video, and audio as the fundamental modalities of human experience, and it sees language as the connective tissue that lets a model reason across all of them. Looking ahead, the team wants to strengthen multimodal understanding by drawing on its M-series models, scale the model up to unlock more capability, and keep pushing toward higher resolution and finer visual detail.

Final Thoughts

H3 marks a real return to form for Hailuo. It is not a versus story so much as a widening of choice: between H3, Seedance, and other recent state-of-the-art models, creators now have a growing set of genuinely capable cameras to reach for. H3 trades a little speed for stability, leans hard on strong prompting and storyboarding, and pairs commercial-grade output with pricing that looks set to pressure the whole field. With open weights on the horizon, its arrival is good news for anyone making video, and a clear sign that the competition at the frontier is heating up.

Credits

No items found.