On September 16, ByteDance Engine officially announced that the Doubao large model 2.1Pro (Doubao-Seed-2.1-pro) has been upgraded to a new 0915 version. The API of this model is now fully launched on ByteDance's Zhuanyu. At the same time, the Doubao work client has also been updated synchronously. Users only need to download or upgrade to the latest version and manually select "Doubao 2.1Pro (new 0915 version)" to experience it directly. In addition, TRAE has also been synchronized with this version, and the Doubao-Seed-Evolving released in June this year has also been upgraded to the same version, allowing enterprises to seamlessly migrate without changing the API access node.
During the upgrade targeting enterprise-level core needs, the Agent task delivery capability has been significantly enhanced. Official information shows that the new version strengthens evidence traceability, authoritative source retrieval, timeliness judgment, and data verification capabilities, greatly reducing model hallucination, making it more fact-based when executing tool calls. In long report processing tasks such as financial research and office automation, the model demonstrates an analyst-level analytical framework, capable of independently breaking down research requirements, querying high-quality data sources, and conducting in-depth modeling analysis, producing well-structured expert-level modeling drafts, ensuring that final conclusions are traceable and the process is more stable.

In terms of multimodal coding (programming assistance), the model's code engineering understanding capability has been significantly enhanced. Facing complex cross-file and long-term programming issues, the new version can clearly clarify the overall project structure and internal relationships between modules, accurately locate root causes and fix them, covering the entire end-to-end development process from requirement understanding to operation validation. At the same time, by leveraging powerful visual language model (VLM) capabilities, the model can directly read design drafts, engineering drawings, and operation recordings, converting visual information into front-end code and game logic, thereby substantially improving Web3D development and game scene presentation effects.
In the field of multimodal understanding, the new version's video reasoning capability has performed impressively. It can not only precisely locate evidence in videos and integrate information across frames to answer questions, but also compare SOPs (Standard Operating Procedures) to identify actual operations and explain their underlying principles. This upgrade benefits scenarios such as video annotation, content quality inspection, long video editing, and location summaries. It also demonstrates higher-level understanding in professional fields such as physical education, engineering practice, scientific experiments, and medical applications. Additionally, image understanding has been enhanced in 3D objects and dense text-image directions. Whether identifying CAD parts, game engine 3D elements, and AR spaces, or processing dimension annotations and tolerance symbols on engineering drawings, PDF figure captions and headers, financial reports and research reports tables, and cross-page footnotes, the processing accuracy has improved significantly.
Notably, the new version has also achieved significant results in Token efficiency optimization, effectively reducing the number of inference rounds and tool calls in practical business scenarios. Compared to the previous generation product, the Token consumption for image reasoning and video reasoning has decreased by more than 30%, further reducing the operational costs of frequent application scenarios such as image problem-solving, image quality inspection, and video review.
Join Now