A Citi report points out that while the cost of enterprise AI usage is declining, the price of underlying computing power has not followed suit. The usage-weighted token expenditures it tracks have decreased by approximately 37.5% from their May peak

2026-08-11

A Citi report points out that while the cost of enterprise AI usage is declining, the price of underlying computing power has not followed suit. The usage-weighted token expenditures it tracks have decreased by approximately 37.5% from their May peak, while Blackwell GPU rental prices have actually increased by approximately 15.2% over the past three months. This means that the current decrease in the unit cost of AI is not primarily due to GPU price reductions, but rather to improvements in model selection, intelligent routing, and inference efficiency. This change is noteworthy because it signifies that AI commercialization may be shifting from simply "piling up computing power" to improving computing power utilization efficiency: enterprises can call different models based on task complexity, reducing the actual cost of each AI service even when underlying computing power remains tight. Nvidia recently demonstrated significant improvements in inference throughput and reduced hardware footprint on the Blackwell architecture through optimization techniques such as quantization, indicating that software and model layer optimizations are becoming important sources of reduced inference costs.