SmophyAI

Published daily · Next update 02:00 UTC

Themed ranking

Fastest and most reliable

Ranked by real measured throughput (tokens/sec) at each model's best-performing provider, tracked daily.

As of September 23, 2026, OpenAI: gpt-oss-120b ranks #1 for fastest and most reliable at throughput (tok/s) 622 tok/s.

  1. #1 OpenAI: gpt-oss-120b - Throughput (tok/s) 622 tok/s
  2. #2 OpenAI: gpt-oss-20b (batch) - Throughput (tok/s) 453 tok/s
  3. #3 Google: Gemma 4 31B (free) - Throughput (tok/s) 304 tok/s
  4. #4 MiniMax: MiniMax M2.7 - Throughput (tok/s) 228 tok/s
  5. #5 Z.ai: GLM 5.2 (free) - Throughput (tok/s) 182 tok/s
  6. #6 Thinking Machines: Inkling (free) - Throughput (tok/s) 164 tok/s
  7. #7 Qwen: Qwen3.6 35B A3B - Throughput (tok/s) 158 tok/s
  8. #8 DeepSeek: DeepSeek V4.1 Flash (batch) - Throughput (tok/s) 154 tok/s
  9. #9 Google: Gemini 3.7 Flash (batch) - Throughput (tok/s) 144 tok/s
  10. #10 MiniMax: MiniMax M3 - Throughput (tok/s) 136 tok/s
50 eligible models, sorted by throughput (tok/s)
#ModelThroughput (tok/s)
01OpenAI: gpt-oss-120b622 tok/s
02OpenAI: gpt-oss-20b (batch)453 tok/s
03Google: Gemma 4 31B (free)304 tok/s
04MiniMax: MiniMax M2.7228 tok/s
05Z.ai: GLM 5.2 (free)182 tok/s
06Thinking Machines: Inkling (free)164 tok/s
07Qwen: Qwen3.6 35B A3B158 tok/s
08DeepSeek: DeepSeek V4.1 Flash (batch)154 tok/s
09Google: Gemini 3.7 Flash (batch)144 tok/s
10MiniMax: MiniMax M3136 tok/s
11Z.ai: GLM 5.3 (batch)134 tok/s
12inclusionAI: Ling 3.0 Flash Fin (free)132 tok/s
13Google: Gemini 3.5 Flash (batch)131 tok/s
14NVIDIA: Nemotron 3 Ultra (free)130 tok/s
15OpenAI: GPT-5.6 Luna (batch)128 tok/s
16OpenAI: GPT-6 Luna (batch)128 tok/s
17Meta: Muse Spark 1.2 Contributor126 tok/s
18Google: Gemini 2.5 Flash Lite (batch)118 tok/s
19Poolside: Laguna XS 2.1 (free)114 tok/s
20Qwen: Qwen3.8 Flash112 tok/s
21DeepSeek: DeepSeek V3.2108 tok/s
22Meta: Muse Spark 1.2105 tok/s
23inclusionAI: Ling-2.6-flash101 tok/s
24NVIDIA: Nemotron 3.5 Lightning (free)99 tok/s
25OpenAI: GPT-5.6 Luna Pro (batch)99 tok/s
26OpenAI: GPT-5 Mini (batch)99 tok/s
27StepFun: Step 3.7 Flash96 tok/s
28NVIDIA: Nemotron 3 Super (free)96 tok/s
29Nex AGI: Nex-N2-Mini94 tok/s
30OpenAI: GPT-5.6 Terra (batch)94 tok/s
31Google: Gemini 3.8 Flash (batch)89 tok/s
32Tencent: Hy388 tok/s
33DeepSeek: DeepSeek V4 Flash Vision Exp88 tok/s
34Meta: Muse Spark 1.3 Contributor87 tok/s
35Qwen: Qwen3.8 27B (free)84 tok/s
36DeepSeek: DeepSeek V4 Flash 042384 tok/s
37Z.ai: GLM 5.3 Flash (batch)84 tok/s
38OpenAI: GPT-5.6 Sol (batch)80 tok/s
39Google: Gemini 2.5 Flash (batch)78 tok/s
40Anthropic: Claude Opus 5 (batch)75 tok/s
41Google: Gemini 3.5 Flash Lite (batch)74 tok/s
42Poolside: Laguna S 2.1 (free)74 tok/s
43Google: Gemini 3.1 Pro Preview (batch)74 tok/s
44Anthropic: Claude Haiku 4.5 (batch)73 tok/s
45Anthropic: Claude Opus 4.8 (batch)72 tok/s
46OpenAI: GPT-5.5 (batch)71 tok/s
47DeepSeek: DeepSeek V4 Pro 042370 tok/s
48Google: Gemini 3.6 Flash (batch)70 tok/s
49Meta: Muse Spark 1.369 tok/s
50MoonshotAI: Kimi K3 (batch)66 tok/s

Source: OpenRouter (openrouter.ai/rankings), as of September 23, 2026.

Benchmark scores: Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings).

Methodology: www.smophy.ai/benchmark/methodology

Token counts originate from each provider's own tokenizer and are not directly comparable across providers.