At an Aug. 31 H1 2026 results briefing, Zhipu AI founder Tang Jie rejected
criticism that the firm only performs fine‑tuning and said the next‑generation
base model will expand in size while explicitly controlling activated parameters
to avoid slower inference and higher costs. Zhipu said it has deployed scaled
inference on roughly 100,000 Chinese chips, cutting per-token inference cost
about 80% since the start of the year. GLM‑5.3‑Flash is the company’s first
model to run entirely on domestic chip clusters under large real‑world traffic;
end‑to‑end service performance is about 3x the initial hardware baseline. Tang
said GLM‑6.0 will target “self‑evolution,” with the primary research challenge
being the model’s ability to judge when to stop and when to self‑correct rather
than simply increasing parameter count.