SmophyAI

Published daily · Next update 02:00 UTC

Agentic · benchmark vs real usage

Best AI model for agents: benchmarks vs. real agentic traffic

On agentic benchmarks, Anthropic: Claude Fable 5.1 (batch) leads with an agentic index of 57.9. But agent benchmarks are widely known to be gameable - in real production traffic across tool dispatch, workflow execution, planning, memory, and web search, Z.ai: GLM 5.3 Flash (batch) handles the largest share at 19.3%.

  1. #1 Z.ai: GLM 5.3 Flash (batch) - 19.3% of agentic traffic · agentic index 50.9
  2. #2 DeepSeek: DeepSeek V4 Flash 0423 - 18.0% of agentic traffic · agentic index 41
  3. #3 DeepSeek: DeepSeek V4.1 Flash (batch) - 16.4% of agentic traffic
  4. #4 OpenAI: GPT-5.6 Luna (batch) - 14.6% of agentic traffic · agentic index 42.1
  5. #5 Tencent: Hy4 preview - 10.8% of agentic traffic
  6. #6 Tencent: Hy3 - 5.3% of agentic traffic
  7. #7 Xiaomi: MiMo-V2.5 - 5.3% of agentic traffic · agentic index 15.8
  8. #8 Z.ai: GLM 5.2 (free) - 3.3% of agentic traffic · agentic index 38.4
  9. #9 Google: Gemini 3.8 Flash (batch) - 2.2% of agentic traffic · agentic index 40.2
  10. #10 OpenAI: GPT-6 Astra (batch) - 1.6% of agentic traffic · agentic index 51
  11. #11 Meta: Muse Spark 1.3 Contributor - 0.9% of agentic traffic
  12. #12 Google: Gemma 4 26B A4B (free) - 0.6% of agentic traffic
  13. #13 Google: Gemini 2.5 Flash Lite (batch) - 0.5% of agentic traffic
  14. #14 OpenAI: GPT-4o-mini (batch) - 0.3% of agentic traffic
  15. #15 Google: Gemini 3.5 Flash Lite (batch) - 0.3% of agentic traffic · agentic index 14.3
Ranked by real agentic-task usage share
#ModelAgentic traffic shareAgentic index
01Z.ai: GLM 5.3 Flash (batch)19.3%50.9
02DeepSeek: DeepSeek V4 Flash 042318.0%41
03DeepSeek: DeepSeek V4.1 Flash (batch)16.4%-
04OpenAI: GPT-5.6 Luna (batch)14.6%42.1
05Tencent: Hy4 preview10.8%-
06Tencent: Hy35.3%-
07Xiaomi: MiMo-V2.55.3%15.8
08Z.ai: GLM 5.2 (free)3.3%38.4
09Google: Gemini 3.8 Flash (batch)2.2%40.2
10OpenAI: GPT-6 Astra (batch)1.6%51
11Meta: Muse Spark 1.3 Contributor0.9%-
12Google: Gemma 4 26B A4B (free)0.6%-
13Google: Gemini 2.5 Flash Lite (batch)0.5%-
14OpenAI: GPT-4o-mini (batch)0.3%-
15Google: Gemini 3.5 Flash Lite (batch)0.3%14.3
16OpenAI: GPT-5.6 Terra (batch)0.2%43.2
17Perplexity: Sonar0.1%-
18MiniMax: MiniMax M30.1%29.5

Agentic traffic share is a weighted rollup of OpenRouter's real classified traffic across five agentic-task categories (tool dispatch, workflow execution, multi-step planning, memory extraction, web search), weighted by each category's own share of total classified traffic.

Frequently asked questions

What is the best AI model for agents?

On benchmarks, Anthropic: Claude Fable 5.1 (batch) leads agentic tasks with an agentic index of 57.9. In real usage, Z.ai: GLM 5.3 Flash (batch) handles the largest share of actual agentic traffic (19%).

Are agent benchmarks reliable?

Agent benchmarks are widely criticized as gameable and inconsistent across labs. Real production usage - what developers actually route agentic workloads to - is a harder-to-game complement to benchmark scores.