Estimated intensity per 1,000 tokens at the Mid scenario. Reasoning models rank highest: they think at length, out of sight.
- ChatGPT (thinking / o3)35 mL
- DeepSeek-R135 mL
- ChatGPT (standard)10 mL
- Claude Opus10 mL
- Gemini Flash1.5 mL
- Llama (small)1.5 mL
mL per 1,000 tokens, Mid scenario
Full model ranking