On August 3, MiniMax officially open-sourced its first multimodal generation model, MiniMax H3. On the day the news was announced, MoLeThread quickly completed the rapid adaptation and efficient operation of the model by leveraging its AI training and inference integrated smart computing card MTT S5000 and MUSA software stack.

As a new masterpiece from MiniMax in the field of multimodal, H3 breaks the traditional single-task boundaries, accurately understands the creative intention in diverse contexts, and comprehensively supports comprehensive input of text, images, audio, and video. According to the official introduction, the model can directly output native audio-visual content with a resolution of 2K and a maximum length of 15 seconds, possessing commercial-grade multi-scenario generation capabilities in terms of performance and practicality.

Notably, H3 achieved the top global score for "video editing capability" in the Artificial Analysis video model ranking. The cost of generating 2K resolution video is as low as 0.8 yuan per second, offering high commercial cost-effectiveness. Meanwhile, its open-source nature significantly lowers the application barrier, supporting flexible local deployment and custom data for enterprises, effectively meeting security and compliance requirements, and accelerating the widespread adoption of multimodal productivity.

Facing the new release of the model, the MoLeThread team demonstrated an extremely high response efficiency. On the day the model was open-sourced, its R&D team quickly conducted model architecture analysis, core technology investigation, and typical operator sorting, completing the full adaptation path from the SGLang-MUSA inference framework to the MATE and muDNN high-performance operator libraries within just 3 hours, achieving fast deployment and stable operation of H3 on the MTT S5000.

MoLeThread stated that it will continue to deepen cooperation with MiniMax in the future, jointly exploring stronger context understanding capabilities and larger model scales in the multimodal field, and fully accelerating the innovation and industrialization of AGI technology.