NewOpen weights

MiniMax H3

A general-purpose omni-modal video generation system that accepts text, image, video, and audio context and produces video with native stereo audio.

Primary taskVideo generationOfficial
Context inputsText · image · video · audioOfficial
Maximum output15 seconds · up to 2KOfficial release claim
Audio outputNative stereo audioOfficial
What it is

A video model built around multimodal context

MiniMax describes H3 as a unified generative model rather than a single text-to-video endpoint. Its published workflows include text-to-video, image-guided generation, reference conditioning, video continuation, and audio-aware generation.

Evidence note

This page summarizes linked first-party documentation. ModelRun Lab has not independently reproduced the performance or quality claims.

Setup paths

Choose the route that matches your goal

Hardware evidence

What we can and cannot say yet

QuestionCurrent answerEvidence
Can it run locally?Yes, supported workflows exist● Official
Does it run on consumer GPUs?Community reports exist across several tiers● Community
Is 12 GB consistently sufficient?No reliable universal answer● Unknown
Exact generation speed?Depends on build, resolution, duration, and offload● Variable

Community reports are discovery evidence, not a guarantee that the same configuration will work on another system.

Source log

Primary references