Themed ranking
Lowest latency
Ranked by real measured median (p50) response latency at each model's best-performing provider, tracked daily. Lower is faster to first meaningful output.
As of September 23, 2026, NVIDIA: Nemotron 3.5 Lightning (free) ranks #1 for lowest latency at latency p50 187ms.
- #1 NVIDIA: Nemotron 3.5 Lightning (free) - Latency p50 187ms
- #2 OpenAI: gpt-oss-120b - Latency p50 219ms
- #3 Google: Gemma 4 31B (free) - Latency p50 244ms
- #4 MiniMax: MiniMax M2.7 - Latency p50 267ms
- #5 Poolside: Laguna XS 2.1 (free) - Latency p50 268ms
- #6 NVIDIA: Nemotron 3 Ultra (free) - Latency p50 279ms
- #7 Thinking Machines: Inkling (free) - Latency p50 333ms
- #8 DeepSeek: DeepSeek V4.1 Flash (batch) - Latency p50 333ms
- #9 Poolside: Laguna S 2.1 (free) - Latency p50 380ms
- #10 Mistral: Mistral Nemo - Latency p50 405ms
Source: OpenRouter (openrouter.ai/rankings), as of September 23, 2026.
Benchmark scores: Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings).
Methodology: www.smophy.ai/benchmark/methodology
Token counts originate from each provider's own tokenizer and are not directly comparable across providers.
