Agentic · benchmark vs real usage
Best AI model for agents: benchmarks vs. real agentic traffic
On agentic benchmarks, Anthropic: Claude Fable 5.1 (batch) leads with an agentic index of 57.9. But agent benchmarks are widely known to be gameable - in real production traffic across tool dispatch, workflow execution, planning, memory, and web search, Z.ai: GLM 5.3 Flash (batch) handles the largest share at 19.3%.
- #1 Z.ai: GLM 5.3 Flash (batch) - 19.3% of agentic traffic · agentic index 50.9
- #2 DeepSeek: DeepSeek V4 Flash 0423 - 18.0% of agentic traffic · agentic index 41
- #3 DeepSeek: DeepSeek V4.1 Flash (batch) - 16.4% of agentic traffic
- #4 OpenAI: GPT-5.6 Luna (batch) - 14.6% of agentic traffic · agentic index 42.1
- #5 Tencent: Hy4 preview - 10.8% of agentic traffic
- #6 Tencent: Hy3 - 5.3% of agentic traffic
- #7 Xiaomi: MiMo-V2.5 - 5.3% of agentic traffic · agentic index 15.8
- #8 Z.ai: GLM 5.2 (free) - 3.3% of agentic traffic · agentic index 38.4
- #9 Google: Gemini 3.8 Flash (batch) - 2.2% of agentic traffic · agentic index 40.2
- #10 OpenAI: GPT-6 Astra (batch) - 1.6% of agentic traffic · agentic index 51
- #11 Meta: Muse Spark 1.3 Contributor - 0.9% of agentic traffic
- #12 Google: Gemma 4 26B A4B (free) - 0.6% of agentic traffic
- #13 Google: Gemini 2.5 Flash Lite (batch) - 0.5% of agentic traffic
- #14 OpenAI: GPT-4o-mini (batch) - 0.3% of agentic traffic
- #15 Google: Gemini 3.5 Flash Lite (batch) - 0.3% of agentic traffic · agentic index 14.3
| # | Model | Agentic traffic share | Agentic index |
|---|---|---|---|
| 01 | Z.ai: GLM 5.3 Flash (batch) | 19.3% | 50.9 |
| 02 | DeepSeek: DeepSeek V4 Flash 0423 | 18.0% | 41 |
| 03 | DeepSeek: DeepSeek V4.1 Flash (batch) | 16.4% | - |
| 04 | OpenAI: GPT-5.6 Luna (batch) | 14.6% | 42.1 |
| 05 | Tencent: Hy4 preview | 10.8% | - |
| 06 | Tencent: Hy3 | 5.3% | - |
| 07 | Xiaomi: MiMo-V2.5 | 5.3% | 15.8 |
| 08 | Z.ai: GLM 5.2 (free) | 3.3% | 38.4 |
| 09 | Google: Gemini 3.8 Flash (batch) | 2.2% | 40.2 |
| 10 | OpenAI: GPT-6 Astra (batch) | 1.6% | 51 |
| 11 | Meta: Muse Spark 1.3 Contributor | 0.9% | - |
| 12 | Google: Gemma 4 26B A4B (free) | 0.6% | - |
| 13 | Google: Gemini 2.5 Flash Lite (batch) | 0.5% | - |
| 14 | OpenAI: GPT-4o-mini (batch) | 0.3% | - |
| 15 | Google: Gemini 3.5 Flash Lite (batch) | 0.3% | 14.3 |
| 16 | OpenAI: GPT-5.6 Terra (batch) | 0.2% | 43.2 |
| 17 | Perplexity: Sonar | 0.1% | - |
| 18 | MiniMax: MiniMax M3 | 0.1% | 29.5 |
Agentic traffic share is a weighted rollup of OpenRouter's real classified traffic across five agentic-task categories (tool dispatch, workflow execution, multi-step planning, memory extraction, web search), weighted by each category's own share of total classified traffic.
Frequently asked questions
What is the best AI model for agents?
On benchmarks, Anthropic: Claude Fable 5.1 (batch) leads agentic tasks with an agentic index of 57.9. In real usage, Z.ai: GLM 5.3 Flash (batch) handles the largest share of actual agentic traffic (19%).
Are agent benchmarks reliable?
Agent benchmarks are widely criticized as gameable and inconsistent across labs. Real production usage - what developers actually route agentic workloads to - is a harder-to-game complement to benchmark scores.
