AI Value Over Time · tracked daily
See what AI actually costs - over time
Every AI benchmark site shows you a snapshot. We track quality-per-dollar and real usage share every day - so you can see whether a model is getting better, cheaper, or more popular, not just where it stands today. No invented index: every number here is a real, sourced quantity or a plain ratio.
- Models tracked
- 353
- Best value (top 10 usage)
- DeepSeek: DeepSeek V4 Flash 0423
- History
- 80 days
- Latest data
- September 22, 2026
Quality and usage aren't the whole picture - these track what running a model actually feels like in production, updated every day.
Uptime & latency
Provider reliability
- 1. DeepSeek: DeepSeek V4 Flash 0423100.0%
- 2. Xiaomi: MiMo-V2.599.8%
- 3. Tencent: Hy3100.0%
DeepSeek: DeepSeek V4 Flash 0423 tops the list at 100.0% best-provider uptime.
Full reliability tracker →Real measured throughput
Fastest models
- 1. OpenAI: gpt-oss-120b622 tok/s
- 2. OpenAI: gpt-oss-20b (batch)453 tok/s
- 3. Google: Gemma 4 31B (free)304 tok/s
OpenAI: gpt-oss-120b leads at 622 tokens/second via Cerebras.
Full throughput ranking →Daily rank changes
Latest leaderboard movement
- 1. Anthropic: Claude Sonnet 4.6 (batch)↓ 9
- 2. Xiaomi: MiMo-V2.6-Pronew
- 3. Xiaomi: MiMo-V2.6-Flashnew
September 22, 2026: Z.ai: GLM 5.3 Flash (batch) led real usage share. Xiaomi: MiMo-V2.6-Pro, Xiaomi: MiMo-V2.6-Flash, OpenAI: GPT-6 Luna (batch), Z.ai: GLM 5.3 FlashX, SpaceXAI: Grok 4.7 entered the leaderboard. Anthropic: Claude Sonnet 4.6 (batch) moved 9 places down.
Full daily changelog →Quality per dollar, over time
Top 8 by usage · tracked daily since August 9, 2026
As of September 23, 2026, DeepSeek: DeepSeek V4 Flash 0423 delivers the best measured quality per dollar among widely-used models, at 180.53 intelligence-index points per dollar ($0.19/1M tokens blended). This chart is tracked daily - first-party history begins August 9, 2026 and grows by one point every night; it is not a static snapshot.
- DeepSeek: DeepSeek V4 Flash 0423 - 180.53 pts/$ · $0.19/1M
- Z.ai: GLM 5.3 Flash (batch) - 176.00 pts/$ · $0.24/1M
- Xiaomi: MiMo-V2.5 - 127.43 pts/$ · $0.18/1M
- OpenAI: GPT-5.6 Luna (batch) - 82.89 pts/$ · $0.45/1M
- MiniMax: MiniMax M3 - 55.62 pts/$ · $0.52/1M
- Z.ai: GLM 5.2 (free) - 33.78 pts/$ · $1.00/1M
- NVIDIA: Nemotron 3 Ultra (free) - 21.81 pts/$ · $1.05/1M
- DeepSeek: DeepSeek V4 Pro 0423 - 18.18 pts/$ · $1.98/1M
- DeepSeek: DeepSeek V4 Flash 0423
- DeepSeek: DeepSeek V4 Pro 0423
- MiniMax: MiniMax M3
- NVIDIA: Nemotron 3 Ultra (free)
- OpenAI: GPT-5.6 Luna (batch)
- Xiaomi: MiMo-V2.5
- Z.ai: GLM 5.2 (free)
- Z.ai: GLM 5.3 Flash (batch)
Index date September 22, 2026Data through 00:00 UTC September 23, 2026Calculated 02:00 UTC
How this is calculated
Quality per dollar= Artificial Analysis Intelligence Index ÷ blended price per 1M tokens (75% prompt / 25% completion). No weighting, no composite - a plain ratio of two real, independently sourced numbers, computed fresh from that day's pricing and benchmark snapshot. The Intelligence Index already measures per-response quality (accuracy across a benchmark suite of reasoning, coding, and knowledge tasks) - it is not response volume, so this ratio isn't "cheap model, more tries." A small quality gap next to a huge price gap is what drives a cheap model to the top.
What this ratio can't tell you:
- Consistency - whether a model reliably hits the same quality run to run, not just on average.
- Cost of a wrong answer in your specific use case - a bad response in production code is not the same as a bad response in casual chat.
- Whether a task needs multiple attempts to reach a usable result.
This is a real, industry-standard way to pre-filter candidates by cost-efficiency - not a final verdict. For high-stakes tasks, weigh the absolute Intelligence Index (shown alongside every ratio below) more heavily than the ratio itself, and check a model's measured endpoint uptime for its consistency track record.
Top 15 models by quality per dollar
Intelligence Index points per dollar of blended price - the 15 best-value models among all 94 with a published quality score. This is a ratio, not a capability ranking: an ultra-cheap model with modest absolute quality can outscore a much stronger, pricier one, purely because price sits near zero. Each bar shows its own absolute Intelligence Index and price alongside the ratio, so you can see which is driving the number.
Open vs. closed AI, tracked daily
80 days of real usage history
Open-weight models now handle 78% of real token volume on OpenRouter - up 8.9 points since July 5, 2026. No other site tracks this as a living, dated trend.
- Open-source
- Closed
Best AI model for…
Ranked by real usage share for that specific task, not overall popularity.
code
Code Generation
- 1. Z.ai: GLM 5.3 Flash (batch)19%
- 2. DeepSeek: DeepSeek V4.1 Flash (batch)16%
- 3. OpenAI: GPT-5.6 Luna (batch)14%
Z.ai: GLM 5.3 Flash (batch) leads with 19% of task usage, 3 points ahead of DeepSeek: DeepSeek V4.1 Flash (batch).
Full code generation ranking, 10 models →general
Content Writing
- 1. DeepSeek: DeepSeek V4 Flash 042331%
- 2. OpenAI: GPT-5.6 Luna (batch)5%
- 3. Google: Gemini 2.5 Flash Lite (batch)4%
DeepSeek: DeepSeek V4 Flash 0423 leads with 31% of task usage, 26 points ahead of OpenAI: GPT-5.6 Luna (batch).
Full content writing ranking, 9 models →code
Debugging
- 1. DeepSeek: DeepSeek V4.1 Flash (batch)17%
- 2. Z.ai: GLM 5.3 Flash (batch)15%
- 3. OpenAI: GPT-5.6 Luna (batch)10%
DeepSeek: DeepSeek V4.1 Flash (batch) leads with 17% of task usage, 2 points ahead of Z.ai: GLM 5.3 Flash (batch).
Full debugging ranking, 10 models →general
Translation
- 1. Tencent: Hy-MT2-1.8B19%
- 2. Tencent: Hy-MT2-7B16%
- 3. Tencent: Hy-MT2-30B-A3B12%
Tencent: Hy-MT2-1.8B leads with 19% of task usage, 3 points ahead of Tencent: Hy-MT2-7B.
Full translation ranking, 9 models →general
Summarization
DeepSeek: DeepSeek V4 Flash 0423 leads with 19% of task usage, 13 points ahead of OpenAI: GPT-5.6 Luna (batch).
Full summarization ranking, 9 models →general
Customer Support
- 1. Google: Gemini 2.5 Flash (batch)16%
- 2. DeepSeek: DeepSeek V4 Flash 042313%
- 3. OpenAI: GPT-5.6 Luna (batch)9%
Google: Gemini 2.5 Flash (batch) leads with 16% of task usage, 2 points ahead of DeepSeek: DeepSeek V4 Flash 0423.
Full customer support ranking, 9 models →general
Roleplay & Fiction
- 1. DeepSeek: DeepSeek V4 Flash 042332%
- 2. DeepSeek: DeepSeek V4.1 Flash (batch)8%
- 3. Mistral: Mistral Nemo6%
DeepSeek: DeepSeek V4 Flash 0423 leads with 32% of task usage, 24 points ahead of DeepSeek: DeepSeek V4.1 Flash (batch).
Full roleplay & fiction ranking, 9 models →general
Q&A & Knowledge
- 1. DeepSeek: DeepSeek V4 Flash 042322%
- 2. Z.ai: GLM 5.3 Flash (batch)10%
- 3. DeepSeek: DeepSeek V4.1 Flash (batch)6%
DeepSeek: DeepSeek V4 Flash 0423 leads with 22% of task usage, 12 points ahead of Z.ai: GLM 5.3 Flash (batch).
Full q&a & knowledge ranking, 9 models →general
Classification
- 1. DeepSeek: DeepSeek V4 Flash 042317%
- 2. Google: Gemini 2.5 Flash Lite (batch)6%
- 3. OpenAI: GPT-5.6 Luna (batch)6%
DeepSeek: DeepSeek V4 Flash 0423 leads with 17% of task usage, 11 points ahead of Google: Gemini 2.5 Flash Lite (batch).
Full classification ranking, 9 models →Rankings by price and specialty
Comparing a 7B model against a flagship on one scale is meaningless. Each of these is a self-contained "best model" claim within a comparable price tier or use case.
By price
Under $0.50/1M
- 1. Z.ai: GLM 5.3 Flash (batch)Quality 41.8
- 2. DeepSeek: DeepSeek V4.1 Flash (batch)Quality 39.5
- 3. OpenAI: GPT-5.6 Luna (batch)Quality 37.3
$0.50-2/1M
- 1. Xiaomi: MiMo-V2.6-ProQuality 46.3
- 2. Z.ai: GLM 5.3 (batch)Quality 44.8
- 3. Google: Gemini 3.8 Flash (batch)Quality 40.9
$2-10/1M
- 1. Anthropic: Claude Opus 5.5 (batch)Quality 57.6
- 2. Qwen: Qwen3.8 MaxQuality 53.4
- 3. OpenAI: GPT-6 Sol (batch)Quality 47.5
By specialty
Best for coding
- 1. Z.ai: GLM 5.3 Flash (batch)24.7%
- 2. DeepSeek: DeepSeek V4.1 Flash (batch)21.4%
- 3. OpenAI: GPT-5.6 Luna (batch)13.1%
Best for agentic work
- 1. Z.ai: GLM 5.3 Flash (batch)19.3%
- 2. DeepSeek: DeepSeek V4 Flash 042318.0%
- 3. DeepSeek: DeepSeek V4.1 Flash (batch)16.4%
Best value
- 1. inclusionAI: Ling 3.0 Flash654.0 pts/$
- 2. inclusionAI: Ling 3.0 Flash VL (free)273.3 pts/$
- 3. inclusionAI: Ling 3.0 Flash Fin (free)251.1 pts/$
Daily snapshots
Every published day gets its own permanent, unchanging permalink - a verifiable point in time for citations, not a page that silently drifts.
How this is calculated
No composite index. No invented weights. Two plain, independently-checkable numbers, each plotted as a time series instead of a single snapshot:
Quality per dollar
Artificial Analysis Intelligence Index ÷ blended price per 1M tokens (75% prompt / 25% completion). A plain division, recomputed from that day's real pricing and benchmark snapshot.
Usage share
Real token volume through OpenRouter, as a share of that day's total - no sampling, no estimation.
We tried composite indices three times this year (a weighted-sum composite, a weighted-geometric-mean composite, and a usage-vs-quality divergence index) and killed all three after finding real defensibility problems - see the full methodology page for that history. The lesson: publish real numbers over time, not invented ones.
Source: OpenRouter (openrouter.ai/rankings), as of September 23, 2026.
Benchmark scores: Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings).
Methodology: www.smophy.ai/benchmark/methodology
Token counts originate from each provider's own tokenizer and are not directly comparable across providers.
Journalist or researcher? Pre-computed citable stats and CSV/JSON exports →
