CLUE team's latest released SuperCLUE-Terminal Chinese AI terminal programming evaluation list shows that DeepSeek-V4-Pro-0813 scored 51.52, ranking second among all tested models, only behind the top-ranked Kimi K3. More notably, it maintains an extremely low call cost while achieving this score, making its cost-effectiveness particularly prominent among high-scoring models. This list is specifically designed for domestic developers, using all local real programming tasks and a unified running framework of Claude Code, fairly comparing major models' complete abilities in understanding Chinese requirements, planning tasks, calling tools, and debugging code.

The realism of the evaluation comes from its task settings: each question requires the model to complete 50 to 110 rounds of interaction, with a full run time of half an hour to one and a half hours, which highly replicates complex workflows in daily development. In this round of the list, Kimi K3 ranked first with 60.61 points, followed closely by DeepSeek-V4-Pro-0813, scoring higher than GLM-5.2's 48.48 and the old version of DeepSeek-V4-Flash's 46.46, firmly placing it in the top tier.
1.42 yuan per question cost, competing with top-tier computing power
The most impressive is DeepSeek-V4-Pro-0813's cost control: the call cost per question is only about 1.42 yuan, roughly one fourteenth of Kimi K3's and one sixteenth of GLM-5.2's. At a time when the scores of top models often differ by just a few points, this order-of-magnitude cost gap means that small and medium-sized teams can afford "nearly top-tier" coding capabilities without paying exorbitant computing costs for every interaction.
From actual testing scenarios, the model covers real demands such as front-end pages, 3D games, and engineering script modifications, able to fully close the game development logic, with collision, scoring, and settlement functions operating normally. The page color and interaction feedback also have moved away from the template-like AI style. There may be minor structural issues in some creative drawing questions, but the overall completion is still leading. For budget-sensitive small and medium enterprises and independent developers, DeepSeek-V4-Pro-0813 redefines the cost-effectiveness benchmark of programming models with a "second-place score and one-fourteenth price"—when top-tier capabilities no longer mean top-tier expenses, the entry barrier for AI programming is truly being lowered.
