<aside> 💡
H3 is a next-generation open-weight, general-purpose multimodal video model.
H3 represents a new paradigm that moves beyond specialized task models toward general multimodal intelligence.
Rather than being limited to specialized tasks such as generating, editing, or referencing, H3 understands multimodal contexts that bring together text, images, video, and audio. This enables it to interpret creative intent in a unified way and deliver more natural, coherent generation and expression.
H3, toward more general multimodal intelligence.
</aside>
| Core specification | MiniMax H3 |
|---|---|
| Model | MiniMax H3 |
| Output duration | 5-15 seconds |
| Output aspect ratio | • First/Last Frame: follows the aspect ratio of the uploaded image |
| • Text-to-Video: supports user-selected 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 aspect ratios | |
| • Omni Reference: supports user-selected 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 aspect ratios, or Auto mode, which lets H3 determine the best ratio | |
| Resolution | 768p mode (coming soon) |
| • For aspect ratios from 16:9 to 9:16, the short edge is 768 pixels. | |
| • For wider formats, total resolution is approximately 1 megapixel—for example, 1536×672 at 21:9. | |
| • 768p results can be upscaled to 1440p. | |
| 1440p(2K) mode | |
| • For aspect ratios from 16:9 to 9:16, the short edge is 1440 pixels. | |
| • For wider formats, total resolution is approximately 3.7 megapixels—for example, 2976×1248 at 21:9. | |
| Frame rate | 24 FPS |
| Audio output | Every generation includes native stereo audio. |
| Input modes | |
| First/Last Frame | • Images: 0, 1, or 2; dimensions: [256, 5760]; aspect ratio: 5:2 to 2:5 |
| • With no image input, H3 runs in Text-to-Video mode. | |
| Omni Reference | • Images: up to 9; dimensions: [256, 5760] |
| • Videos: up to 3 clips; 2-15 seconds per clip; 15 seconds total; dimensions: [256, 5760]; aspect ratio: 5:2 to 2:5 | |
| • Audio: up to 3 clips; 2-15 seconds per clip; 15 seconds total. Audio must be paired with at least one image or video and cannot be used on its own. | |
| • Mixed inputs: up to 12 files total | |
| • With no image, video, or audio input, H3 runs in Text-to-Video mode. | |
| Supported input formats | • Video: H.264/AVC, H.265/HEVC; embedded audio: AAC, MP3 |
| • Image: JPG, JPEG, PNG, WEBP, HEIC, HEIF | |
| • Audio: WAV, MP3 | |
| File size limits | • Video: 50 MB per file; |
| • Image: 30 MB per file; | |
| • Audio: 15 MB per file. | |
| • API request body: 64 MB. For API workflows, URL-based media input is recommended. | |
| Note: Limits apply per file, not to the combined upload size. | |
| Prompt length | Up to 7,000 characters |
| Input | |
|---|---|
| Input Type | List Price |
| Video | Billed based on the duration of the input video (in seconds) and the resolution of the output video: 2K at $0.13/s, 768P at $0.09/s |
| Image | First 5 images are free; each additional image is charged at $0.04 per image |
| Audio | Free |
| Output | |
| Resolution | List Price |
| 768P | $0.09/s |
| 2K | $0.13/s |
| Capability | What it does |
|---|---|
| Native Multimodal Understanding & Generation | Built on unified multimodal understanding and generation capabilities, MiniMax H3 supports combined input across various texts, images, audios, and videos. |
| It can interpret characters, motion, sound, emotion, cinematography, visual style, and creative intent across comprehensive multimodal references, then seamlessly combine them to generate coherent, complete audio-visual content through an integrated end-to-end workflow. | |
| Precise Multimodal Editing & Control | Enables precise control and editing across characters, objects, scenes, sound, and rhythm, with strong instruction-following capabilities. |
| Creators can continuously refine and iterate on existing content, accurately adjusting visual, audio, and narrative details to match their creative vision. | |
| Production-Ready Content Creation Across Use Cases | Built for real-world commercial production across film and entertainment, advertising, branding, e-commerce, and gaming. |
| It handles dynamic typography, VFX, product showcases, UI/UX motion design, game visuals, and stylized content. | |
| Enterprise can move faster from concept validation and visual pitching to final production. |
As video models become more capable, they can support a much broader range of real-world production workflows.
H3 can render text, subtitles, brand assets, UI/UX, game content, and e-commerce visuals across film, advertising, gaming, brand, and retail use cases.
Use it to explore short-form concepts, test visual directions, create more expressive product showcases, bring IP characters into new actions and settings, and support concept validation, storyboard previews, and visual pitches.
Trailers, TV commercials, and premium brand films
3.1.1.1 Vintage Binocular Brand Film
Prompt
Input

image.png

image.png

image.png

image.png
Output
[Multimodal Understanding Example 1.mp4](attachment:750d6f42-9487-458f-8b63-05e9fdc77cee:750d6f42-9487-458f-8b63-05e9fdc77cee.mp4)
Multimodal Understanding Example 1.mp4
3.1.1.2 Epic Space Opera Teaser
Prompt
Input

image.png
Output
[campaign film Movie 1.mp4](attachment:6de2e13b-944f-423a-9190-e12b5cdb71b9:6de2e13b-944f-423a-9190-e12b5cdb71b9.mp4)
campaign film Movie 1.mp4