Alibaba Cloud said its Zhenwu M890 super-node instance has been Day‑0 adapted to
Kimi K3, making it the first domestic super-node to run a nearly 3
trillion‑parameter model. Alibaba Cloud, chip unit T‑Head and the Kimi team
performed deep joint optimizations across the software stack and operators;
Qianwen AI platform and Alibaba Cloud Bailian will offer Kimi K3 model APIs.
Core operator optimizations for the Kimi K3 architecture reportedly raise
inference compute and bandwidth efficiency, enabling stable support for up to
1M‑token context; tests show first‑token latency fell about 35% and single‑GPU
decoding throughput rose 1.8x.