Agentic · benchmark vs real usage
Best AI model for agents: benchmarks vs. real agentic traffic
On agentic benchmarks, Claude Opus 5 (batch) leads with an agentic index of 59.2. But agent benchmarks are widely known to be gameable - in real production traffic across tool dispatch, workflow execution, planning, memory, and web search, DeepSeek: DeepSeek V4 Flash 0423 handles the largest share at 38.3%.
- #1 DeepSeek: DeepSeek V4 Flash 0423 - 38.3% of agentic traffic · agentic index 48.4
- #2 Tencent: Hy3 - 14.5% of agentic traffic
- #3 OpenAI: GPT-5.6 Luna (batch) - 8.8% of agentic traffic · agentic index 46.9
- #4 DeepSeek: DeepSeek V4 Pro - 7.7% of agentic traffic · agentic index 37.8
- #5 Xiaomi: MiMo-V2.5 - 7.6% of agentic traffic · agentic index 24.4
- #6 Z.ai: GLM 5.2 (batch) - 6.8% of agentic traffic · agentic index 45.7
- #7 Google: Gemini 2.5 Flash Lite (batch) - 4.3% of agentic traffic
- #8 Google: Gemini 3.6 Flash (batch) - 3.0% of agentic traffic · agentic index 40.5
- #9 NVIDIA: Nemotron 3 Ultra (free) - 2.7% of agentic traffic · agentic index 27.5
- #10 DeepSeek: DeepSeek V3.1 Terminus - 1.8% of agentic traffic · agentic index 18.1
- #11 OpenAI: gpt-oss-20b (free) - 1.2% of agentic traffic · agentic index 3.1
- #12 MoonshotAI: Kimi K3 - 0.9% of agentic traffic · agentic index 54.3
- #13 Google: Gemini 3 Flash Preview (batch) - 0.8% of agentic traffic
- #14 Poolside: Laguna S 2.1 (free) - 0.5% of agentic traffic
- #15 Google: Gemini 3.5 Flash Lite (batch) - 0.4% of agentic traffic · agentic index 27.2
| # | Model | Agentic traffic share | Agentic index |
|---|---|---|---|
| 01 | DeepSeek: DeepSeek V4 Flash 0423 | 38.3% | 48.4 |
| 02 | Tencent: Hy3 | 14.5% | - |
| 03 | OpenAI: GPT-5.6 Luna (batch) | 8.8% | 46.9 |
| 04 | DeepSeek: DeepSeek V4 Pro | 7.7% | 37.8 |
| 05 | Xiaomi: MiMo-V2.5 | 7.6% | 24.4 |
| 06 | Z.ai: GLM 5.2 (batch) | 6.8% | 45.7 |
| 07 | Google: Gemini 2.5 Flash Lite (batch) | 4.3% | - |
| 08 | Google: Gemini 3.6 Flash (batch) | 3.0% | 40.5 |
| 09 | NVIDIA: Nemotron 3 Ultra (free) | 2.7% | 27.5 |
| 10 | DeepSeek: DeepSeek V3.1 Terminus | 1.8% | 18.1 |
| 11 | OpenAI: gpt-oss-20b (free) | 1.2% | 3.1 |
| 12 | MoonshotAI: Kimi K3 | 0.9% | 54.3 |
| 13 | Google: Gemini 3 Flash Preview (batch) | 0.8% | - |
| 14 | Poolside: Laguna S 2.1 (free) | 0.5% | - |
| 15 | Google: Gemini 3.5 Flash Lite (batch) | 0.4% | 27.2 |
| 16 | OpenAI: gpt-oss-120b | 0.3% | 13.4 |
| 17 | Perplexity: Sonar | 0.2% | - |
| 18 | MiniMax: MiniMax M3 (batch) | 0.2% | 36.1 |
Agentic traffic share is a weighted rollup of OpenRouter's real classified traffic across five agentic-task categories (tool dispatch, workflow execution, multi-step planning, memory extraction, web search), weighted by each category's own share of total classified traffic.
Frequently asked questions
What is the best AI model for agents?
On benchmarks, Claude Opus 5 (batch) leads agentic tasks with an agentic index of 59.2. In real usage, DeepSeek: DeepSeek V4 Flash 0423 handles the largest share of actual agentic traffic (38%).
Are agent benchmarks reliable?
Agent benchmarks are widely criticized as gameable and inconsistent across labs. Real production usage - what developers actually route agentic workloads to - is a harder-to-game complement to benchmark scores.
