Key Takeaways
- MiniMax H3 is a general-purpose, open-weight multimodal video generation model released July 31, 2026. It generates native 2K video at 2560×1440 resolution and 24fps, with synchronized stereo audio produced in the same generation pass as the video. It accepts up to 12 references per generation (9 images, 3 video clips, 3 audio clips) alongside a text prompt. Clip duration ranges from 5 to 15 seconds. Supported aspect ratios are 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16.
- H3 is priced at approximately $0.13 per second for native 2K output and $0.09 per second for 768p output across major API gateways including OpenRouter, Vercel AI Gateway, EvoLink, and MiniMax’s own platform. MiniMax positions this at roughly one-third of comparable frontier video model pricing. A 15-second 2K clip costs approximately $1.95.
- The stereo audio generation is produced in-pipeline, not added in post. Dialogue, sound effects (foley), and ambient room tone are generated simultaneously with the video frames and locked to them. This eliminates the separate audio post-production pass that closed video pipelines without native audio require. Generated audio should still be reviewed before publication for wording, timing, and artifacts.
- H3’s open-weight release separates it from closed video models like Google Veo 3.1. Open weights allow organizations with sufficient GPU capacity to self-host the model, audit it, fine-tune it, and eliminate per-second API fees at high volumes. Weights are released in stages by region subject to local regulation; availability varies by jurisdiction.
- Independent community consensus on Reddit (r/LocalLLaMA, r/aivideo) places H3 quality at roughly 98% of Seedance 2.5-class output, with strong temporal consistency and effectively solved in-frame text rendering. The remaining quality gap shows up at the top end of cinematic photorealism and complex physical interaction, where Seedance 2.5 and Veo 3.1 hold a marginal edge.
- H3 is the third generation of MiniMax’s frontier video models and a direct successor to the previous MiniMax generation. The previous generation output at 768p or 1080p with no native audio and a simpler text-plus-image input surface. H3 adds native 2K, synchronized stereo audio, full multimodal input (text plus image plus video plus audio), conversational editing of generated clips, and open weights. For any workflow previously using MiniMax video, H3 is a strict capability superset.
MiniMax H3 lands in the middle of a video generation market that is stratifying fast. At the top, closed frontier models from Google and ByteDance compete on peak cinematic quality. At the bottom, fast-turnaround models compete on speed and cost. H3 occupies a specific position: native 2K fidelity, in-pipeline stereo audio, a 12-reference multimodal input system, open weights, and pricing at roughly one-third of the frontier tier. This review covers the full technical picture and the practical decision of where H3 belongs in a production workflow.
What Is MiniMax H3?
MiniMax H3 is a transformer-based, open-weight, general-purpose multimodal video generation model from MiniMax, the Beijing-founded AI lab. It is the third generation of the company’s frontier video models and is also referred to as Hailuo 3.0, the consumer-facing name for the same model available through MiniMax’s Hailuo platform.
H3’s defining design choice is unification. Rather than maintaining separate models for text-to-video, image-to-video, audio, and editing tasks, H3 ingests text, images, reference video, and reference audio in a single context window, generates the video, and produces synchronized stereo audio in the same pass. This means dialogue, sound effects, and ambient sound are created with the video frames rather than added afterward. The model belongs to the same architectural generation as other omni-modal foundation models that process multiple modalities in a unified context, applied here specifically to video production.
H3 was released July 31, 2026, and replaces the previous MiniMax video generation line. The previous generation supported text-plus-image input at up to 1080p with no native audio. H3 is a complete rebuild rather than an iteration: native 2K output, full multimodal input, in-pipeline audio, open weights, and a 12-reference input budget are all new capabilities.
Core Technical Specifications
| Specification | Value |
|---|---|
| Native resolution | 2560×1440 (2K / 1440p) |
| Frame rate | 24fps |
| Clip duration | 5 to 15 seconds |
| Audio | Native synchronized stereo (generated in-pipeline) |
| Image references per generation | Up to 9 |
| Video references per generation | Up to 3 |
| Audio references per generation | Up to 3 |
| Total references per generation | Up to 12 (combined) plus text prompt |
| Supported aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 |
| Weights | Open-weight, staged regional release |
| Editing | Conversational instruction-based revisions on generated clips |
| API price (2K) | ~$0.13 per second |
| API price (768p) | ~$0.09 per second |
Native 2K Output: What It Actually Means
The distinction between native 2K output and upscaled 2K output matters more than the numbers suggest. Many video models that advertise 2K or 4K resolution apply a post-hoc upscaling step to a lower-resolution base generation. This sharpens edges and increases pixel count, but it cannot synthesize detail that the model never produced internally. Upscaled output at fine scales, including hair strands, skin texture, fabric weave, small product labels, and in-frame text, reveals the original lower-resolution base under examination at delivery size.
H3 renders at 2560×1440 pixels inside the model. Fine texture, hair, fabric, small UI elements in screencast-style shots, and on-screen text are generated at the native resolution rather than interpolated from lower-resolution output. At a 24fps frame rate that matches cinema and high-end commercial production rather than social-media 30fps, H3 clips can be used in high-resolution advertising and commercial deliverables without an additional upscale step.
Stereo Audio: In-Pipeline Generation
Native stereo audio is H3’s most operationally significant differentiator. The model generates dialogue, sound effects (foley), and ambient room tone simultaneously with the video frames and locks the audio to the frames during generation. A speaking character has voice and lip movement generated together. An object hitting a surface has the impact sound generated with the visual contact rather than dubbed in afterward.
This changes the post-production workflow in a concrete way. Closed video pipelines without native audio require a separate audio production pass after generation: recording or synthesizing voiceover, adding foley, mixing and locking to the video cut. H3 produces a roughly finished audio-visual package from a single generation. For commercial and advertising workflows where a hero shot needs to deliver dialogue and sound design, this eliminates a production step that previously required either additional AI audio tools or human studio production.
The practical caveat: generated audio still requires human review before publication. Wording, pronunciation, timing, and audio artifacts all need checking. The key is that review and approval of a generated audio track is a substantially lighter lift than producing the audio from scratch. For high-volume workflows generating many clips, the reduction in per-clip audio production time is material.
The 12-Reference Multimodal Input System
H3’s input budget allows a single generation to combine a text prompt with up to 9 reference images, 3 reference videos, and 3 reference audio clips, for 12 total references plus text. Understanding how to use this budget effectively is the most leveraged skill in working with H3.
Each reference slot works best when it has a single job. Common reference assignments in commercial production include: one image for character identity, a second for wardrobe, a third for the target product with correct geometry and materials, a fourth for brand color palette, a fifth for composition or set environment, and a video reference for motion timing or performance style. Audio references can specify voice character, musical mood, or sound design pacing.
The most common failure mode with multimodal models is feeding conflicting references: two character images that depict different identities, or a motion reference video that contradicts the specified first and last frame. When references conflict, the model has to resolve the contradiction during generation, which typically produces a blend that satisfies neither input. Treating each reference slot as a single-purpose control, and auditing references for conflicts before generating, produces more consistent output.
H3 also supports first-frame and last-frame control through the image reference slots. A first-frame image anchors the shot’s opening composition; a last-frame image specifies where the shot ends. The model generates the motion between them. This is particularly useful for structured commercial shots where the start and end state are defined by a brand brief, and the motion needs to connect them naturally.
After generation, H3 supports conversational editing: instruction-based revisions to specific elements of the generated clip without rebuilding the entire shot. A revision prompt can change a character’s action, adjust lighting, modify background elements, or alter pacing without discarding the successful parts of the original generation. For iterative production workflows, this reduces the compute cost of reaching a finished clip significantly.
Open Weights and Self-Hosting
H3’s open-weight release separates it from closed video generation models including Google Veo 3.1, Runway Gen-4, and most commercial consumer video tools. Open weights mean an organization with sufficient GPU capacity can download and run the model on its own infrastructure, eliminating per-second API fees and retaining full control over data, outputs, and fine-tuning.
The release is staged by region and subject to local regulation. MiniMax is releasing weights in sequence rather than as a simultaneous global release. Any team planning a self-hosting deployment should confirm current weight availability in their jurisdiction through official MiniMax channels before committing to infrastructure investment, rather than relying on third-party mirrors that may not reflect the current authorized distribution.
For most teams, the operational path is to start on the hosted API or a gateway and revisit self-hosting when sustained volume makes the math compelling. The crossover point depends on GPU costs and utilization in the team’s environment. At very high volumes, say, thousands of clips per month, the per-second API cost dominates and self-hosting becomes economical. For typical studio or agency volumes, the hosted API is the right default.
Pricing and Value
H3’s rate cards across the major API gateways (OpenRouter, Vercel AI Gateway, EvoLink) and on MiniMax’s own platform cluster at $0.13 per second for native 2K output and $0.09 per second for 768p output. The table below shows the dollar cost at each pricing tier by clip length.
| Output tier | Rate | 5-second clip | 10-second clip | 15-second clip |
|---|---|---|---|---|
| Native 2K (2560×1440) | ~$0.13/sec | ~$0.65 | ~$1.30 | ~$1.95 |
| 768p | ~$0.09/sec | ~$0.45 | ~$0.90 | ~$1.35 |
MiniMax positions H3’s rate card at roughly one-third of comparable frontier video models. The price difference compounds at production volume. A studio generating 1,000 15-second 2K clips pays on the order of $1,950 at H3 pricing, versus approximately three times that at a frontier-priced model’s rate card for equivalent output. For campaigns that require many variants, multiple aspect ratios, multiple product colorways, platform-specific crops, the volume of clips needed can be substantial, and H3’s pricing is designed for that scale.
How to Access MiniMax H3
There are four practical paths to using H3, suited to different kinds of teams.
The first is Hailuo.ai, MiniMax’s consumer-facing platform, which provides a no-code interface for generating H3 clips through a browser without API integration. This is the right starting point for individuals and small teams who want to experiment with the model before committing to a technical integration.
The second is the MiniMax platform API at platform.minimax.io. Developers calling MiniMax directly get access to the full parameter surface and the canonical pricing, and first access to new features as they roll out. This requires MiniMax account setup and, where applicable, KYC verification.
The third is aggregator gateways including OpenRouter, Vercel AI Gateway, and EvoLink. These expose H3 through a unified billing and SDK surface alongside other video and image models. They are the right integration point for teams whose existing infrastructure already speaks the gateway’s protocol, or who want one integration layer across multiple model providers. Gateway rates include a small margin above the model cost.
The fourth is self-hosted weights, available once weight access is confirmed for the team’s region. Self-hosting eliminates per-second fees but requires GPU hardware, inference infrastructure, and operational ownership. It is appropriate for high-volume operations with data residency requirements or whose compute math justifies the overhead.
Quality: How H3 Compares
The most honest source for AI video quality comparisons in 2026 is community benchmarking on r/LocalLLaMA and r/aivideo, where independent creators and technical users run the same prompts across models and compare outputs without promotional framing. The consensus on H3 from those communities places it at approximately 98% of Seedance 2.5-class quality, with two specific strengths called out: temporal consistency (characters, objects, and scenes hold together across frames better than most alternatives) and in-frame text rendering (which has historically been a persistent failure mode for video generation models and which H3 handles well).
The remaining 2% quality gap shows up at the upper end of cinematic photorealism and in complex physical interactions: hands, fast rotating objects, occlusion, intricate contact between multiple objects. These are hard for the entire model class and not specific to H3. For productions where the quality ceiling is the brief and budget is not the primary constraint, Seedance 2.5 and Veo 3.1 hold a marginal edge at the top of photorealistic quality.
Against Kling 3.5 and PixVerse V6, H3 is the stronger choice for cinematic and commercial work where native 2K fidelity and native audio matter. Kling 3.5 is the stronger choice for multi-shot sequences and 4K output; its AI Director and 6-shot-per-generation multi-shot capability have no equivalent in H3. PixVerse V6 is faster for quick social iteration. For hero commercial production, H3’s combination of native 2K and stereo audio is the stronger fit.
Best Use Cases
H3 fits cleanly into workflows where native 2K fidelity, synchronized audio, or multimodal reference control matter more than raw speed or the absolute ceiling of photorealistic quality.
Cinematic commercial and hero product films are H3’s primary target. Native 2K combined with stereo audio means a single H3 generation can deliver a close-to-finished shot for a commercial brief without requiring a separate audio production pass. For advertising campaigns where a product shot needs to look authoritative and sound designed, H3 eliminates a production step.
Dialogue-led character scenes work well because H3 generates lip movement and voice in the same pass. Short narrative scenes, explainer video segments, and character-driven brand content all benefit from the fact that the generated character speaks with a voice that was generated alongside the visual rather than dubbed over it.
Reference-led campaign continuity is where the 12-reference budget is most valuable. A brand campaign that requires the same character in the same wardrobe in multiple settings, or the same product with identical geometry and label placement across multiple shots, can use the reference slots to lock those invariants across many separate generations, producing visual consistency that would otherwise require extensive manual review and rejection of off-brief outputs.
High-volume multivariate creative, including social media A/B testing variants and platform-specific delivery crops, is where H3’s pricing advantage is most concrete. At roughly one-third of frontier model pricing, producing thirty 10-second variants of a creative direction costs approximately $39 at native 2K, versus around $117 at frontier pricing for equivalent output. The volume economics favor H3 for any workflow where many variants of the same creative direction are required.
Previsualization (previz) pipelines for film and large-scale production benefit from both the low API cost and the open-weight availability. Studios that generate large volumes of previz clips to communicate shot intent before principal photography can use H3 at a cost that does not strain pre-production budgets.
Limitations
Generated audio requires human review before publication. Wording, pronunciation, timing between audio and visual, and audio artifacts are all possible failure modes that automated generation cannot fully eliminate. The review step is lighter than producing audio from scratch, but it should be planned into any production workflow that publishes H3 audio to audiences.
Complex physical interaction remains a limitation of the model class. Hands, fast rotating objects, scenes with intricate contact or occlusion between multiple objects, and physically accurate simulations of fast motion are all areas where H3 can drift. For shots that require precise physical realism in these scenarios, Pika 2.5’s physics simulation or Seedance 2.5 may produce more consistent results.
Identity and brand elements need final rights and accuracy review. Faces generated from reference images, on-screen text, product logos, and licensed characters all require clearance and accuracy checking before commercial publication. The model’s ability to render in-frame text cleanly does not guarantee the text is accurate or authorized.
Open weight availability varies by region and changes over time. Teams planning self-hosting deployments should verify current weight access through official MiniMax channels rather than assuming availability based on general announcements.
Prompting Tips
H3 rewards prompts that are structured as shot descriptions rather than scene descriptions. The distinction matters: a scene description covers what is happening in the world of the story; a shot description specifies what the camera sees, how it moves, and what sound accompanies the image. H3, like all video models, generates what the camera would record, not an abstract description of events.
Each generation works best with one main action and one camera movement. Stacking three character transformations or two simultaneous camera movements in a five-second clip forces the model to compress too much into too little time, which produces rushed output where each element is undercooked. One subject performing one action while the camera makes one move produces the cleanest results.
Anchor the invariants first before describing the action. Identify what must not change across the clip (character identity, product geometry, wardrobe, color palette, set environment) and assign reference slots to encode those invariants before writing the action description. Treat each reference slot as a single-purpose control. References that encode two things at once, like an image that shows both a character’s face and a background that should change, create conflicts the model has to resolve during generation.
Write audio into the shot description rather than leaving it implicit. Because H3 generates audio in the same pass as the video, describing the sound environment alongside the visual action produces better audio-visual alignment. Specify dialogue in quotes, describe ambient texture (leather interior, outdoor wind, café background), and note where silence is intentional.
Verdict
MiniMax H3 is the practical default for high-quality AI video production in 2026 for any workflow where native 2K fidelity, synchronized stereo audio, and multimodal reference control are requirements, and where per-clip cost needs to survive high production volume. It is not the quality ceiling (Seedance 2.5 and Veo 3.1 hold a marginal edge at peak photorealism) and it is not the fastest route for quick social iteration (Kling Turbo and PixVerse V6 are faster for that use case). But for the broad middle of commercial, advertising, e-commerce, and branded content production, H3’s combination of open weights, 12-reference multimodal input, native 2K resolution, in-pipeline stereo audio, and roughly one-third of frontier pricing puts it at the center of where professional AI video production belongs in 2026.
Frequently Asked Questions
What is MiniMax H3?
MiniMax H3, also called Hailuo 3.0, is MiniMax’s third-generation open-weight multimodal video model, released July 31, 2026. It generates native 2K video at 2560×1440 resolution and 24fps with synchronized stereo audio produced in the same generation pass. It accepts up to 12 references per generation (9 images, 3 video clips, 3 audio clips) plus a text prompt. Clip duration ranges from 5 to 15 seconds. It is designed for advertising, branding, e-commerce, product design, UI/UX, gaming, and film previsualization workflows.
How much does MiniMax H3 cost?
On the major third-party gateways (OpenRouter, Vercel AI Gateway, EvoLink) and on MiniMax’s own platform, H3 is priced at approximately $0.13 per second for native 2K output and $0.09 per second for 768p output. A 15-second 2K clip costs approximately $1.95. MiniMax positions this at roughly one-third of comparable frontier video model pricing.
Is MiniMax H3 open source?
MiniMax H3 is open-weight, which means MiniMax has released the model weights for download and self-hosting. This is distinct from fully open-source release (which would include training code, data, and pipeline). Open weights allow organizations to run the model on their own GPU infrastructure, fine-tune it, and audit its outputs without relying on MiniMax’s API. Weight availability is released in stages by region and is subject to local regulation; teams should confirm availability in their jurisdiction through official MiniMax channels.
Does MiniMax H3 generate audio?
Yes. H3 generates native synchronized stereo audio in the same generation pass as the video, covering dialogue, sound effects (foley), and ambient room tone. The audio is locked to the video frames during generation rather than dubbed in afterward. Generated audio requires human review before publication for wording, pronunciation, timing, and artifacts.
What resolution and frame rate does MiniMax H3 output?
H3 outputs native 2K video at 2560×1440 resolution and 24fps. The 2K output is generated natively inside the model rather than upscaled from a lower resolution. The 24fps frame rate matches cinema and high-end commercial production standards. A 768p output tier is also available at a lower per-second price of approximately $0.09 per second.
How does MiniMax H3 compare to Kling 3.5 and Veo 3.1?
Against Kling 3.5, H3 is the stronger choice for commercial production requiring native stereo audio and tight multimodal reference control. Kling 3.5 is the stronger choice for multi-shot cinematic sequences (up to 6 camera cuts per generation) and for productions requiring 4K native output at 60fps. Against Veo 3.1, H3 is positioned at roughly 98% of Seedance/Veo-class quality according to independent community benchmarking, with the remaining gap at the top end of photorealistic quality and complex physical interaction. H3’s primary advantage over Veo 3.1 is price (roughly one-third) and open-weight availability.
What inputs does MiniMax H3 accept?
A single H3 generation accepts a text prompt plus up to 9 reference images, 3 reference video clips, and 3 reference audio clips, for 12 total references. First-frame and last-frame image control are supported through the image reference slots. The model uses all inputs in a unified context to generate the video and audio together. Conversational editing of generated clips is also supported, allowing instruction-based revisions without rebuilding from scratch.
Where can I access MiniMax H3?
H3 is accessible through four main paths. Hailuo.ai is MiniMax’s consumer platform with a no-code browser interface. The MiniMax platform API at platform.minimax.io provides direct API access for developers. Aggregator gateways including OpenRouter, Vercel AI Gateway, and EvoLink provide H3 through a unified multi-model billing interface. Self-hosted weights are available in stages by region for organizations with GPU infrastructure who want to eliminate per-second API fees.
What are MiniMax H3’s main limitations?
Generated audio requires human review before publication for wording, timing, pronunciation, and artifacts. Complex physical interactions (hands, fast rotating objects, occlusion, multi-object contact) can produce inconsistent output, which is a limitation of the model class as a whole. Identity elements, brand logos, on-screen text, and licensed characters all require final clearance and accuracy review before commercial use. Open-weight availability varies by region and should be confirmed through official MiniMax channels before planning a self-hosting deployment.




