
Direct from a prompt
Describe the scene, camera motion and mood; MiniMax H3 builds 2K video with synced sound from text alone.
Try MiniMax H3 FreeOpen Multimodal Video Model
Try MiniMax H3 free on PixMind. Generate native 2K video with synchronized stereo audio from text, images, reference video, and reference audio — 4 to 15 seconds in one pass. Free-trial availability and remaining uses follow the live generator.
Online generator
Pinned to the minimax-h3-video route. Inputs, duration, aspect ratio and credits follow the live generator.
Switch between text, a single image, or multimodal references to control the shot.

Describe the scene, camera motion and mood; MiniMax H3 builds 2K video with synced sound from text alone.
Try MiniMax H3 FreeOne model, every modality
MiniMax H3 understands text, images, reference videos, and reference audio in a single context — so character identity, motion timing, and soundscape stay coherent in one generation instead of being stitched across tools.
Try MiniMax H3 Free
Native 2K
2560×1440 output is produced natively rather than upscaled, keeping fine detail and clean edges for campaign-ready frames.
Try MiniMax H3 Free
Synchronized audio
Dialogue, sound effects, and ambient audio are generated together with the picture — no separate voiceover or foley pass required.
Try MiniMax H3 Free
Character consistency
Combine multiple reference images to keep the same character and style coherent across shots and edits.
Try MiniMax H3 Free
Conversational editing
After generating, instruct changes to characters, objects, scenes, sound, or pacing without redoing the whole shot.
Try MiniMax H3 Free
Model comparison
What changed from the previous generation.
| Capability | Hailuo 02 | MiniMax H3 |
|---|---|---|
| Resolution | 768p / 1080p | Native 2K (2560×1440) |
| Duration | 6s / 10s | 4–15 seconds |
| Native audio | No | Synchronized stereo |
| Inputs | Text + image | Text + image + video + audio |
| Video editing | No | Conversational |
| Weights | Closed | Open-weight (staged) |
Hailuo 02 figures reflect MiniMax's prior 768p/1080p line. MiniMax H3 capability per MiniMax official docs, verified 2026-07-31.
Write a prompt, add a start image or first/last frame, and optionally attach reference images, video, or audio.
Pick 4–15s and the aspect ratio (21:9 to 9:16) that fits your platform.
Run the model, then use conversational editing to revise characters, sound, or pacing until it lands.
Where native 2K, synced audio, and multimodal reference pay off.

Native 2K tabletop and hero shots with synced sfx and voiceover for TVC and launch films.

Lock a character across shots with reference images and direct micro-expressions and motion.

Multi-shot action sequences with camera motion direction at 24fps cinematic cadence.

9:16 vertical hooks up to 15s with native audio — ready for Reels, Shorts and TikTok.

Test shot connections, camera moves, and scene mood before expensive production.

Anime, illustration, ink-wash, and game-CG looks driven by reference and prompt.
Full 2K pixel density is rendered natively for cleaner frames.
No separate dub or foley step — audio is generated with the picture.
Text, image, video, and audio references in one coherent generation.
Revise characters, scenes, sound, and pacing by instruction.
Define the motion arc and final composition precisely.
MiniMax's open-weight line, staged by region.
MiniMax H3 is MiniMax's open-weight, general-purpose multimodal video model. It generates native 2K video with synchronized stereo audio from text, images, reference videos, and reference audio in one unified pipeline.
Text-to-video, image-to-video (including first/last frames), and multimodal reference — up to 9 reference images, 3 reference videos, and 3 reference audio clips combined.
Yes. It produces native synchronized stereo audio — dialogue, sound effects, and ambient sound — generated together with the picture in a single pass.
Native 2K (2560×1440), 4–15 second clips, and 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 aspect ratios.
Hailuo 02 output 768p/1080p with no native audio. H3 adds native 2K, synchronized stereo audio, full multimodal input (image+video+audio references), in-generation multi-shot, and conversational video editing.
Yes. MiniMax H3 supports conversational editing — you can instruct changes to characters, objects, scenes, sound, or pacing on an already generated clip.
MiniMax has announced MiniMax H3 as an open-weight model. Weight availability is staged by region and subject to local regulations; check MiniMax's official channels for the current status.
MiniMax H3 Eco is the budget line of H3 on PixMind: 480p/720p output with text, image, and first/last-frame input plus synced audio — good for fast drafts. Full multimodal reference (image + video + audio) and higher resolutions stay on the standard H3 route.
Use the generator above — it is pinned to the minimax-h3-video route. Write a prompt (or add an image / references), pick duration and aspect ratio, then run. Credits are billed per the live generator.
Model fact sources
Generate native 2K video with synchronized audio and multimodal reference on PixMind — pinned to the minimax-h3-video route.
Try MiniMax H3 Free