Back to feed
News Story
APriority81
THE DECODER
1 sources

China's MiniMax H3 becomes first open model to top AI video ranking

MiniMax has released the weights of its H3 video model, making it the first open model to top an AI video ranking. This milestone highlights the growing competitiveness of open-source models in video generation.

SynthePulse Insight · AI deep reading

MiniMax H3: First Open-Source Model to Top AI Video Rankings, but Closed-Source Modules and Commercial Restrictions Remain

Version 1 · 1 source

MiniMax releases H3 video model weights, becoming the first open-source model to top AI video rankings, but the 2K resolution module and context processing module remain closed-source, and commercial licensing is restricted for high-revenue companies.

  • MiniMax H3 is the first open-source model to top AI video rankings, ranking first in video editing, second in text-to-video, and third in image-to-video on the Artificial Analysis leaderboard.
  • H3 has 33 billion parameters and can jointly process text, images, video, and audio, generating 4 to 15-second clips with stereo sound.
  • A single prompt can include up to 9 reference images, 3 video clips, and 3 audio clips.
  • The 2K resolution module and H3-Context-IR (which converts prompts and reference materials into a structured intermediate format) are not open-sourced, with local runs capped at 768p.
  • Open weights allow fine-tuning on custom materials, characters, or visual styles, but commercial use is limited to companies with annual revenue below $20 million.
  • On the same day, ByteDance released closed-source Seedance 2.5, capable of generating 30-second clips with built-in audio.
Open section navigationA Milestone for Open-Source Models: H3 Tops Video Rankings

A Milestone for Open-Source Models: H3 Tops Video Rankings

MiniMax released the H3 video model weights on August 3, 2026, and according to The Decoder, this is the first time an open-source model has topped AI video rankings. The Artificial Analysis leaderboard shows H3 ranking first in video editing, second in text-to-video, and third in image-to-video.

H3 is a 33-billion-parameter model that can jointly process text, images, video, and audio, generating 4 to 15-second clips with stereo sound. According to the model card, a single prompt can include up to 9 reference images, 3 video clips, and 3 audio clips.

Closed-Source Modules and Local Run Limitations

Despite the open weights, two key components remain closed: the 2K resolution module and H3-Context-IR (which converts prompts and reference materials into a structured intermediate format). As a result, local runs of H3 in ComfyUI are capped at 768p, and users must handle context preparation themselves based on MiniMax's prompt guide.

The open weights allow fine-tuning on custom materials, characters, or specific visual styles, but there is a licensing restriction: commercial use is only permitted for companies with annual revenue below $20 million.

Same-Day Competition: ByteDance Releases Seedance 2.5

On the same day, ByteDance released closed-source Seedance 2.5, capable of generating 30-second clips with built-in audio. This contrast highlights the competitive landscape between open and closed-source models in video generation.

Credibility boundary

This article's information primarily comes from The Decoder's report, which cites the Artificial Analysis rankings and the model card on HuggingFace. All specific data (such as parameter counts, rankings, and licensing terms) are from these sources but have not been independently verified.

Insight takeaway

MiniMax H3's open-source release is a significant milestone in video generation, but the closed-source modules and commercial restrictions mean its openness is limited. Users must weigh the performance limitations of local runs against the licensing conditions for commercial use.

Primary report

THE DECODER

Primary source