Recently, the WeChat Visual Team of Tencent officially announced the open-source of the general-purpose multimodal embedding model WeMM-Embedding. This model fully supports multimodal inputs such as text and images, aiming to provide solid technical support for AI application developers through advanced technical architecture and training strategies, further promoting the innovative applications of artificial intelligence technology in more vertical fields.

In terms of version layout and technical architecture, WeMM-Embedding has launched three versions with different scales: 2B, 4B, and 9B, to meet diverse deployment needs. The model adopts a unified architecture design combined with core technologies such as two-stage training and knowledge distillation, significantly enhancing multimodal understanding capabilities while effectively balancing deployment efficiency in practical applications.

As a technically matured result that has been rigorously tested in real business scenarios, the model is now widely deployed in WeChat's recommendation and search systems, with a daily online call volume reaching billions. The full open-source release not only marks further openness of WeChat in building a multimodal AI technology ecosystem but will also provide strong technical empowerment for AI application development in the industry.