For the first time, domestic computing power has firmly supported a model with trillions of parameters. Hai Guang Information announced that it has completed full-stack adaptation and performance verification of Kimi K3 on its DCU platform - this means that running such massive models is no longer limited to just a few overseas flagship chips.

Hai Guang DCU has optimized from the underlying operators to the inference engine, allowing the model code to run directly without any modifications, almost eliminating the migration cost. Developers can deploy the model immediately once they get the computing power, and the compatibility is very smooth. For the KDA attention mechanism and the MoE architecture composed of 896 experts, Hai Guang has done operator-level customization; even when 896 experts run in parallel, the inference remains efficient and stable under heavy scenarios with millions of context, without lagging.

image.png

In practical scenarios, K3 on Hai Guang DCU can support advanced applications such as long-range intelligent agent projects, visual closed-loop creation, and even self-designed chips. Domestic computing power users can finally run open-source models with trillions of parameters by themselves. It's worth noting that before the release of K3, Yueling Dark Face had already collaborated with Hai Guang and CAS Sustech to complete full compatibility optimization: Sustech, based on Hai Guang DCU hardware stack, integrated the distributed training and multi-card parallel inference chain for Kimi MoE architecture; Hai Guang DCU achieved stable long-text inference on Sustech 8000 cluster through collaborative optimization of hardware and software stacks at the bottom level.

Official tests have provided reassurance: the performance loss of domestic cards is controlled within 12%, which has reached the threshold for commercial deployment. When a gap was opened between trillion-parameter models and domestic GPUs, the landscape of computing power in the intelligent era is quietly shifting.