
Direct from a prompt
Describe the scene, camera motion and mood; MiniMax H3 builds 2K video with synced sound from text alone.
Generate with MiniMax H3Open Multimodal Video Model
MiniMax's open-weight multimodal video model. Generate native 2K video with synchronized stereo audio from text, images, reference video, and reference audio — 4 to 15 seconds, in one pass.
Online generator
Pinned to the minimax-h3-video route. Inputs, duration, aspect ratio and credits follow the live generator.
Switch between text, a single image, or multimodal references to control the shot.

Describe the scene, camera motion and mood; MiniMax H3 builds 2K video with synced sound from text alone.
Generate with MiniMax H3One model, every modality
MiniMax H3 understands text, images, reference videos, and reference audio in a single context — so character identity, motion timing, and soundscape stay coherent in one generation instead of being stitched across tools.
Generate with MiniMax H3
Native 2K
2560×1440 output is produced natively rather than upscaled, keeping fine detail and clean edges for campaign-ready frames.
Generate with MiniMax H3
Synchronized audio
Dialogue, sound effects, and ambient audio are generated together with the picture — no separate voiceover or foley pass required.
Generate with MiniMax H3
Character consistency
Combine multiple reference images to keep the same character and style coherent across shots and edits.
Generate with MiniMax H3
Conversational editing
After generating, instruct changes to characters, objects, scenes, sound, or pacing without redoing the whole shot.
Generate with MiniMax H3
Model comparison
What changed from the previous generation.
| Capability | Hailuo 02 | MiniMax H3 |
|---|---|---|
| Resolution | 768p / 1080p | Native 2K (2560×1440) |
| Duration | 6s / 10s | 4–15 seconds |
| Native audio | No | Synchronized stereo |
| Inputs | Text + image | Text + image + video + audio |
| Video editing | No | Conversational |
| Weights | Closed | Open-weight (staged) |
Hailuo 02 figures reflect MiniMax's prior 768p/1080p line. MiniMax H3 capability per MiniMax official docs, verified 2026-07-31.
Write a prompt, add a start image or first/last frame, and optionally attach reference images, video, or audio.
Pick 4–15s and the aspect ratio (21:9 to 9:16) that fits your platform.
Run the model, then use conversational editing to revise characters, sound, or pacing until it lands.
Where native 2K, synced audio, and multimodal reference pay off.

Native 2K tabletop and hero shots with synced sfx and voiceover for TVC and launch films.

Lock a character across shots with reference images and direct micro-expressions and motion.

Multi-shot action sequences with camera motion direction at 24fps cinematic cadence.

9:16 vertical hooks up to 15s with native audio — ready for Reels, Shorts and TikTok.

Test shot connections, camera moves, and scene mood before expensive production.

Anime, illustration, ink-wash, and game-CG looks driven by reference and prompt.
Full 2K pixel density is rendered natively for cleaner frames.
No separate dub or foley step — audio is generated with the picture.
Text, image, video, and audio references in one coherent generation.
Revise characters, scenes, sound, and pacing by instruction.
Define the motion arc and final composition precisely.
MiniMax's open-weight line, staged by region.
MiniMax H3 is MiniMax's open-weight, general-purpose multimodal video model. It generates native 2K video with synchronized stereo audio from text, images, reference videos, and reference audio in one unified pipeline.
Text-to-video, image-to-video (including first/last frames), and multimodal reference — up to 9 reference images, 3 reference videos, and 3 reference audio clips combined.
Yes. It produces native synchronized stereo audio — dialogue, sound effects, and ambient sound — generated together with the picture in a single pass.
Native 2K (2560×1440), 4–15 second clips, and 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 aspect ratios.
Hailuo 02 output 768p/1080p with no native audio. H3 adds native 2K, synchronized stereo audio, full multimodal input (image+video+audio references), in-generation multi-shot, and conversational video editing.
Yes. MiniMax H3 supports conversational editing — you can instruct changes to characters, objects, scenes, sound, or pacing on an already generated clip.
MiniMax has announced MiniMax H3 as an open-weight model. Weight availability is staged by region and subject to local regulations; check MiniMax's official channels for the current status.
Use the generator above — it is pinned to the minimax-h3-video route. Write a prompt (or add an image / references), pick duration and aspect ratio, then run. Credits are billed per the live generator.
Model fact sources
Generate native 2K video with synchronized audio and multimodal reference on PixMind — pinned to the minimax-h3-video route.
Generate with MiniMax H3