Zhipu has open-sourced GLM-5.3-Flash (320B-A18B). Moore Threads rapidly ported
and optimized the model on its AI train/inference card MTT S5000 and MUSA
software stack, enabling efficient, high-speed inference. The company said the
work marks a key step for domestic compute to support ultra-large multimodal
models. GLM-5.3-Flash, anonymised as codename Ox-Alpha during public testing on
OpenCode and OpenRouter, showed strong inference speed and demand, becoming the
week’s most-used model and setting new call-volume records on both platforms.