Citigroup believes the capability gap between open-source and closed-source cutting-edge models has narrowed significantly. In the latest Artificial Analysis benchmark, the leading closed-source model, Claude Opus 5, scored 63, while the open-weight

2026-09-02

Citigroup believes the capability gap between open-source and closed-source cutting-edge models has narrowed significantly. In the latest Artificial Analysis benchmark, the leading closed-source model, Claude Opus 5, scored 63, while the open-weight model, Kimi K3, achieved 60, a difference of only 3 points. Practical use is validating this catch-up. As of September 1st, open-weight models accounted for 60.9% of Vercel AI Gateway's token volume over the past three months; however, in June, they carried 29% of the tokens with less than 4% of the expenditure, while the four major closed-source labs in the US still accounted for 95% of the expenditure. This means that open models are taking over high-frequency, fault-tolerant tasks due to their cost-effectiveness, while closed-source models continue to hold the fort for complex and high-risk workflows. As enterprises adopt multi-model routing, closed-source labs must continuously demonstrate performance premiums; otherwise, token prices and gross margins will be under pressure, and more value may flow to inference infrastructure, model routing, and the application layer.