1. Model Overview

<aside> 💡

H3 is a next-generation open-weight, general-purpose multimodal video model.

H3 represents a new paradigm that moves beyond specialized task models toward general multimodal intelligence.

Rather than being limited to specialized tasks such as generating, editing, or referencing, H3 understands multimodal contexts that bring together text, images, video, and audio. This enables it to interpret creative intent in a unified way and deliver more natural, coherent generation and expression.

H3, toward more general multimodal intelligence.

</aside>

2. Model Specifications

Core specification MiniMax H3
Model MiniMax H3
Output duration 5-15 seconds
Output aspect ratio • First/Last Frame: follows the aspect ratio of the uploaded image
• Text-to-Video: supports user-selected 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 aspect ratios
• Omni Reference: supports user-selected 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 aspect ratios, or Auto mode, which lets H3 determine the best ratio
Resolution 768p mode (coming soon)
• For aspect ratios from 16:9 to 9:16, the short edge is 768 pixels.
• For wider formats, total resolution is approximately 1 megapixel—for example, 1536×672 at 21:9.
• 768p results can be upscaled to 1440p.
1440p(2K) mode
• For aspect ratios from 16:9 to 9:16, the short edge is 1440 pixels.
• For wider formats, total resolution is approximately 3.7 megapixels—for example, 2976×1248 at 21:9.
Frame rate 24 FPS
Audio output Every generation includes native stereo audio.
Input modes
First/Last Frame • Images: 0, 1, or 2; dimensions: [256, 5760]; aspect ratio: 5:2 to 2:5
• With no image input, H3 runs in Text-to-Video mode.
Omni Reference • Images: up to 9; dimensions: [256, 5760]
• Videos: up to 3 clips; 2-15 seconds per clip; 15 seconds total; dimensions: [256, 5760]; aspect ratio: 5:2 to 2:5
• Audio: up to 3 clips; 2-15 seconds per clip; 15 seconds total. Audio must be paired with at least one image or video and cannot be used on its own.
• Mixed inputs: up to 12 files total
• With no image, video, or audio input, H3 runs in Text-to-Video mode.
Supported input formats • Video: H.264/AVC, H.265/HEVC; embedded audio: AAC, MP3
• Image: JPG, JPEG, PNG, WEBP, HEIC, HEIF
• Audio: WAV, MP3
File size limits • Video: 50 MB per file;
• Image: 30 MB per file;
• Audio: 15 MB per file.
• API request body: 64 MB. For API workflows, URL-based media input is recommended.
Note: Limits apply per file, not to the combined upload size.
Prompt length Up to 7,000 characters

2.1 Pricing

Input
Input Type List Price
Video Billed based on the duration of the input video (in seconds) and the resolution of the output video: 2K at $0.13/s, 768P at $0.09/s
Image First 5 images are free; each additional image is charged at $0.04 per image
Audio Free
Output
Resolution List Price
768P $0.09/s
2K $0.13/s

3. Core Capabilities

Capability What it does
Native Multimodal Understanding & Generation Built on unified multimodal understanding and generation capabilities, MiniMax H3 supports combined input across various texts, images, audios, and videos.
It can interpret characters, motion, sound, emotion, cinematography, visual style, and creative intent across comprehensive multimodal references, then seamlessly combine them to generate coherent, complete audio-visual content through an integrated end-to-end workflow.
Precise Multimodal Editing & Control Enables precise control and editing across characters, objects, scenes, sound, and rhythm, with strong instruction-following capabilities.
Creators can continuously refine and iterate on existing content, accurately adjusting visual, audio, and narrative details to match their creative vision.
Production-Ready Content Creation Across Use Cases Built for real-world commercial production across film and entertainment, advertising, branding, e-commerce, and gaming.
It handles dynamic typography, VFX, product showcases, UI/UX motion design, game visuals, and stylized content.
Enterprise can move faster from concept validation and visual pitching to final production.

3.1 Production-Ready Content Across Commercial Use Cases

As video models become more capable, they can support a much broader range of real-world production workflows.

H3 can render text, subtitles, brand assets, UI/UX, game content, and e-commerce visuals across film, advertising, gaming, brand, and retail use cases.

Use it to explore short-form concepts, test visual directions, create more expressive product showcases, bring IP characters into new actions and settings, and support concept validation, storyboard previews, and visual pitches.

3.1.1 Brand Films & Cinematic Content

Trailers, TV commercials, and premium brand films

3.1.1.1 Vintage Binocular Brand Film

Prompt

Input

image.png

image.png

image.png

image.png

image.png

image.png

image.png

image.png

Output

[Multimodal Understanding Example 1.mp4](attachment:750d6f42-9487-458f-8b63-05e9fdc77cee:750d6f42-9487-458f-8b63-05e9fdc77cee.mp4)

Multimodal Understanding Example 1.mp4


3.1.1.2 Epic Space Opera Teaser

Prompt

Input

image.png

image.png

Output

[campaign film Movie 1.mp4](attachment:6de2e13b-944f-423a-9190-e12b5cdb71b9:6de2e13b-944f-423a-9190-e12b5cdb71b9.mp4)

campaign film Movie 1.mp4