🔥Minimax H3 is live — save 45% on annual membershipGet 45% Off
Pixmind
All video models

Open Multimodal Video Model

MiniMax H3

MiniMax's open-weight multimodal video model. Generate native 2K video with synchronized stereo audio from text, images, reference video, and reference audio — 4 to 15 seconds, in one pass.

  • Native 2K (2560×1440) output
  • Synchronized native stereo audio
  • Full multimodal input: image + video + audio references
  • First / last frame and up to 9 reference images
  • Conversational video editing
PixMind Veo 3.1 model showcase; autoplay is muted.

Online generator

Online MiniMax H3 generator

Pinned to the minimax-h3-video route. Inputs, duration, aspect ratio and credits follow the live generator.

Loading...

Three ways to direct a shot

Switch between text, a single image, or multimodal references to control the shot.

Text-to-video shot

Direct from a prompt

Describe the scene, camera motion and mood; MiniMax H3 builds 2K video with synced sound from text alone.

Generate with MiniMax H3

One model, every modality

A unified pipeline for image, video and audio

MiniMax H3 understands text, images, reference videos, and reference audio in a single context — so character identity, motion timing, and soundscape stay coherent in one generation instead of being stitched across tools.

Generate with MiniMax H3
Multimodal unified pipeline

Native 2K

Renders at full 2K inside the model

2560×1440 output is produced natively rather than upscaled, keeping fine detail and clean edges for campaign-ready frames.

Generate with MiniMax H3
Native 2K detail

Synchronized audio

Native stereo sound in one pass

Dialogue, sound effects, and ambient audio are generated together with the picture — no separate voiceover or foley pass required.

Generate with MiniMax H3
Synchronized stereo audio

Character consistency

Lock identity with up to 9 references

Combine multiple reference images to keep the same character and style coherent across shots and edits.

Generate with MiniMax H3
Character consistency

Conversational editing

Revise a clip by instruction

After generating, instruct changes to characters, objects, scenes, sound, or pacing without redoing the whole shot.

Generate with MiniMax H3
Conversational editing

Model comparison

MiniMax H3 vs Hailuo 02

What changed from the previous generation.

MiniMax H3 vs Hailuo 02
CapabilityHailuo 02MiniMax H3
Resolution768p / 1080pNative 2K (2560×1440)
Duration6s / 10s4–15 seconds
Native audioNoSynchronized stereo
InputsText + imageText + image + video + audio
Video editingNoConversational
WeightsClosedOpen-weight (staged)

Hailuo 02 figures reflect MiniMax's prior 768p/1080p line. MiniMax H3 capability per MiniMax official docs, verified 2026-07-31.

How to generate

01

Choose your inputs

Write a prompt, add a start image or first/last frame, and optionally attach reference images, video, or audio.

02

Set duration & ratio

Pick 4–15s and the aspect ratio (21:9 to 9:16) that fits your platform.

03

Generate & refine

Run the model, then use conversational editing to revise characters, sound, or pacing until it lands.

What MiniMax H3 is best for

Where native 2K, synced audio, and multimodal reference pay off.

Product film still

Brand & product films

Native 2K tabletop and hero shots with synced sfx and voiceover for TVC and launch films.

Character performance still

Character performance

Lock a character across shots with reference images and direct micro-expressions and motion.

Action sequence still

Action & cinematic chase

Multi-shot action sequences with camera motion direction at 24fps cinematic cadence.

Social short video still

Social short video

9:16 vertical hooks up to 15s with native audio — ready for Reels, Shorts and TikTok.

Previsualization still

Previsualization

Test shot connections, camera moves, and scene mood before expensive production.

Stylized content still

Stylized content

Anime, illustration, ink-wash, and game-CG looks driven by reference and prompt.

Why choose MiniMax H3

2K in the model, not upscaled

Full 2K pixel density is rendered natively for cleaner frames.

Sound on the first pass

No separate dub or foley step — audio is generated with the picture.

All-modal context

Text, image, video, and audio references in one coherent generation.

Directable editing

Revise characters, scenes, sound, and pacing by instruction.

First / last frame control

Define the motion arc and final composition precisely.

Open-weight

MiniMax's open-weight line, staged by region.

MiniMax H3 FAQs

What is MiniMax H3?

MiniMax H3 is MiniMax's open-weight, general-purpose multimodal video model. It generates native 2K video with synchronized stereo audio from text, images, reference videos, and reference audio in one unified pipeline.

What inputs does MiniMax H3 accept?

Text-to-video, image-to-video (including first/last frames), and multimodal reference — up to 9 reference images, 3 reference videos, and 3 reference audio clips combined.

Does MiniMax H3 generate audio?

Yes. It produces native synchronized stereo audio — dialogue, sound effects, and ambient sound — generated together with the picture in a single pass.

What resolution, duration and aspect ratios?

Native 2K (2560×1440), 4–15 second clips, and 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 aspect ratios.

How is H3 different from Hailuo 02?

Hailuo 02 output 768p/1080p with no native audio. H3 adds native 2K, synchronized stereo audio, full multimodal input (image+video+audio references), in-generation multi-shot, and conversational video editing.

Can I edit a video after generating it?

Yes. MiniMax H3 supports conversational editing — you can instruct changes to characters, objects, scenes, sound, or pacing on an already generated clip.

Is MiniMax H3 open-weight?

MiniMax has announced MiniMax H3 as an open-weight model. Weight availability is staged by region and subject to local regulations; check MiniMax's official channels for the current status.

How do I generate with MiniMax H3 on PixMind?

Use the generator above — it is pinned to the minimax-h3-video route. Write a prompt (or add an image / references), pick duration and aspect ratio, then run. Credits are billed per the live generator.

Direct a 2K film with synced sound

Generate native 2K video with synchronized audio and multimodal reference on PixMind — pinned to the minimax-h3-video route.

Generate with MiniMax H3