The Ant Open Source team officially released and open-sourced the first native multimodal large model of the Bailing series, Ling-3.0-flash-VL. This model is based on the MoE architecture of Ling-3.0-flash and has a total parameter count of 124B. It activates 5.5B parameters per inference and natively supports image, text, and video input with a context window of 256K Tokens.
During the development process, the team focused on exploring how to make large models more reliable and efficient in completing real-world tasks. In response to the industry's common concern that "adding visual capabilities to large models will reduce text intelligence," the actual training practices provided the opposite conclusion: native multimodal joint training not only expanded the application boundaries but also further enhanced text intelligence.
To achieve more accurate execution results, Ling-3.0-flash-VL introduces an innovative visual feedback mechanism. Through a complete closed-loop process of "observing the execution result, comparing it with the goal, identifying deviations, and continuously correcting," the model transforms the traditional one-time "generation" into a dynamic process with self-correction capabilities, significantly improving the reliability of the final result.
Additionally, this model inherits the core advantage of Ling-3.0-flash as an efficient execution node in Agent workflows. In the closed-loop operation with visual feedback, it can perfectly balance output quality and execution efficiency, advancing and completing complex tasks with lower costs and shorter time.
Join Now