DeepSeek has officially released the smallest member of its new model architecture series today - DeepSeek V4.1Flash. As a lightweight model with native multi-modal visual understanding capabilities, it is designed to achieve higher capability limits, faster inference speeds, larger throughput, and the potential for expansion to larger parameter models.
In terms of core architecture, DeepSeek V4.1Flash adopts a 552B parameter MoE (Mixture of Experts) architecture and innovatively introduces a new Causal-Encoder-Decoder structure. Due to its asymmetric input and output design, the input activation is only 8B, and the output activation is 16B, significantly reducing the overall cost compared to known models of the same size. Combined with a new pre-training method and larger-scale reinforcement learning post-training, this model has successfully surpassed a range of flagship models including DeepSeek V4Pro in various benchmark tests.

In terms of efficiency optimization and cost control, the new model has significantly reduced the size of the KV Cache. Compared to the previous generation, V4.1Flash's demand for high-bandwidth memory (HBM) has been reduced to 1/4, and the demand for solid-state drives (SSD) has been reduced to 1/8. Especially in Agent use scenarios, where the cost of context storage and cache hits is relatively high, this compression greatly reduces the actual usage cost for related tasks. According to statistics, the KV Cache has been reduced to 1/437 of the original size compared to the initial model.
With the official launch of the new model, DeepSeek has made systematic adjustments to its product line and billing strategy. DeepSeek V4.1Flash is now available on the DeepSeek API, and users can simply change the model name to deepseek-flash to call it. The old versions of V4Flash and V4Flash Vision Exp have been officially discontinued, while the old model names (deepseek-v4-flash and deepseek-v4-flash-vision-exp) are currently temporarily routed to V4.1Flash. Given that V4.1Flash has comprehensively surpassed V4Pro in performance, cost, speed, and total usage time, the official plans to gradually discontinue the V4Pro model: after 12:00 on September 14, 2026, until the release of V4.1Pro, all requests to access deepseek-v4-pro will be routed to V4.1Flash and charged at its unit price. Currently, official partners such as WorkBuddy, CodeBuddy, and OpenCode under Tencent have fully integrated this model.
Join Now