MyPrivateClaw

MiniMax Launches H3 Open-Weights Omni-Modal Video Generation Model | ai-models

On July 31, 2026 MiniMax officially launched H3, a general purpose omni modal video generation model with open weights released on HuggingFace under the MiniMa…

Published on MyPrivateClaw

Aug 3, 2026, 4:24 PM UTC

Coverage date

Aug 3, 2026

Last updated

Aug 3, 2026, 4:24 PM UTC

News summary

On July 31, 2026 MiniMax officially launched H3, a general purpose omni modal video generation model that unifies text, image, video, and audio understanding into a single system capable of generating video with native stereo audio at resolutions up to 2K and durations up to 15 seconds. MiniMax released open weights of H3 Base on HuggingFace under the MiniMax H3 Community License, providing two task specific checkpoints: FL2VA (first and last frame to audio video) and Ref2VA (reference to audio video). The H3 Base open weights support inference via SGLang, vLLM, diffusers (ModularPipeline.from pretrained), and ComfyUI. H3 Context IR, the hosted preprocessing system that converts multimodal inputs into intermediate representations for H3 Base, and H3 Regenerate 2K, the in context 2K upscaler module, are not included in the open source release. The initial open source release provides inf…