A recent community test has drawn attention: using the same GPT-6 model, Codex's inference budget may be significantly lower than that of directly calling the OpenAI API, even with the same inference intensity.
At the highest inference level (Max), GPT-6 Astra's inference budget via API calls reaches 832, while Codex's is only 128, a difference of approximately 6.5 times. The difference is even more pronounced with GPT-6 Sol, where the API version reaches 768, while Codex's is only 62, a difference of approximately 12.4 times. The community test also shows that this difference is not limited to the highest level; at other comparable inference intensity settings, Codex's inference budget is generally lower than that of the API version.
Inference budget typically refers to the computational budget available for the model's internal inference process, and is not equivalent to the actual number of tokens consumed. OpenAI's official documentation confirms that inference intensity affects the model's depth of thought, response speed, and token consumption, but has not yet confirmed the specific budget values of 832, 128, 768, and 62.
This difference could impact performance on complex software engineering tasks. For simple code generation and routine modifications, a lower inference budget is beneficial for improving response speed and controlling computational costs; however, in large codebase refactoring, cross-module dependency analysis, complex fault diagnosis, and long-chain task planning, a more sufficient inference budget may help the model explore more solutions, check for potential errors, and complete more in-depth verification.