MiniMax H3
A general-purpose omni-modal video generation system that accepts text, image, video, and audio context and produces video with native stereo audio.
A video model built around multimodal context
MiniMax describes H3 as a unified generative model rather than a single text-to-video endpoint. Its published workflows include text-to-video, image-guided generation, reference conditioning, video continuation, and audio-aware generation.
This page summarizes linked first-party documentation. ModelRun Lab has not independently reproduced the performance or quality claims.
Choose the route that matches your goal
Best starting point for a visual local workflow. Follow ComfyUI’s version-specific guide.
Review the model card, repository structure, license, and current file variants.
Use the upstream repository for current implementation details and changes.
What we can and cannot say yet
Community reports are discovery evidence, not a guarantee that the same configuration will work on another system.