🎬 MiniMax-H3 Demo

Omni-modal video + native stereo audio from MiniMaxAI/MiniMax-H3

Showcase official samples · craft H3-style prompts · generate via MiniMax API

Official MiniMax-H3 sample reels

Weights + research card: MiniMaxAI/MiniMax-H3

H3 is a general-purpose omni-modal system: text / image / video / audio in → video with native stereo audio out (4–15s, many aspect ratios, 24 fps, up to 2K via regenerate).

T2VA · text → video+audio (768p)

Text-only prompt → joint video and stereo soundtrack.

T2VA · 2K regenerate

Same story after H3-Regenerate-2K.

FL2VA · first/last frame

Keyframe-conditioned generation with audio.

I2VA · image → video+audio

Single image drives first-frame motion + sound.

I2VA · 2K

Image-conditioned clip regenerated at 2K.

Ref2VA · multi-reference

Images / video / audio references in one packed context.

R2VA · 2K

Reference workflow at higher fidelity.

Direct API · 768p

Open Platform direct 768p reference output.

Direct API · 2K

Open Platform direct 2K reference output.

Local open weights deploy via SGLang / vLLM / diffusers / ComfyUI — see the model card. Unquantized ZeroGPU split demo: multimodalart/minimax-h3.

Built for demoing MiniMaxAI/MiniMax-H3 · not affiliated with MiniMax · samples © MiniMax via the public model repo