Citigroup says enterprise AI user costs are dropping even as underlying compute
prices hold up. Its usage-weighted token expenditures are down about 37.5% from
the May peak, while NVIDIA Blackwell GPU rental rates have risen roughly 15.2%
over the past three months. Citigroup attributes the unit-cost decline mainly to
model selection, smart routing and inference-efficiency gains rather than GPU
price cuts. The shift suggests AI commercialization is moving from raw compute
scaling toward higher utilization—firms can lower per-call costs by invoking
lighter models for simpler tasks despite tight compute supply. NVIDIA has
demonstrated quantization and other optimizations on Blackwell that materially
raise inference throughput and reduce hardware occupancy, highlighting software
and model-layer improvements as a key lever to cut inference costs.