Rising use of AI agents has driven rapid growth in token consumption. Firms are
deploying compute platforms to orchestrate multiple leading LLMs and compress
token costs, while accelerating China-wide buildout of large-scale compute
clusters. Industry participants say hyperscale clusters powered by China-made
compute chips are the next key lever to cut costs. Over the next 3–5 years,
photonic/optoelectronic chips and other frontier technologies could further
reduce per-token compute costs by roughly 50% or more by lowering latency and
power per operation.