SmophyAI

Published daily · Next update 02:00 UTC

AI Value Over Time · tracked daily

See what AI actually costs - over time

Every AI benchmark site shows you a snapshot. We track quality-per-dollar and real usage share every day - so you can see whether a model is getting better, cheaper, or more popular, not just where it stands today. No invented index: every number here is a real, sourced quantity or a plain ratio.

Models tracked
353
Best value (top 10 usage)
DeepSeek: DeepSeek V4 Flash 0423
History
80 days
Latest data
September 22, 2026

Quality and usage aren't the whole picture - these track what running a model actually feels like in production, updated every day.

Uptime & latency

Provider reliability

  1. 1. DeepSeek: DeepSeek V4 Flash 0423100.0%
  2. 2. Xiaomi: MiMo-V2.599.8%
  3. 3. Tencent: Hy3100.0%

DeepSeek: DeepSeek V4 Flash 0423 tops the list at 100.0% best-provider uptime.

Full reliability tracker

Real measured throughput

Fastest models

  1. 1. OpenAI: gpt-oss-120b622 tok/s
  2. 2. OpenAI: gpt-oss-20b (batch)453 tok/s
  3. 3. Google: Gemma 4 31B (free)304 tok/s

OpenAI: gpt-oss-120b leads at 622 tokens/second via Cerebras.

Full throughput ranking

Daily rank changes

Latest leaderboard movement

  1. 1. Anthropic: Claude Sonnet 4.6 (batch)↓ 9
  2. 2. Xiaomi: MiMo-V2.6-Pronew
  3. 3. Xiaomi: MiMo-V2.6-Flashnew

September 22, 2026: Z.ai: GLM 5.3 Flash (batch) led real usage share. Xiaomi: MiMo-V2.6-Pro, Xiaomi: MiMo-V2.6-Flash, OpenAI: GPT-6 Luna (batch), Z.ai: GLM 5.3 FlashX, SpaceXAI: Grok 4.7 entered the leaderboard. Anthropic: Claude Sonnet 4.6 (batch) moved 9 places down.

Full daily changelog

Quality per dollar, over time

Top 8 by usage · tracked daily since August 9, 2026

As of September 23, 2026, DeepSeek: DeepSeek V4 Flash 0423 delivers the best measured quality per dollar among widely-used models, at 180.53 intelligence-index points per dollar ($0.19/1M tokens blended). This chart is tracked daily - first-party history begins August 9, 2026 and grows by one point every night; it is not a static snapshot.

  1. DeepSeek: DeepSeek V4 Flash 0423 - 180.53 pts/$ · $0.19/1M
  2. Z.ai: GLM 5.3 Flash (batch) - 176.00 pts/$ · $0.24/1M
  3. Xiaomi: MiMo-V2.5 - 127.43 pts/$ · $0.18/1M
  4. OpenAI: GPT-5.6 Luna (batch) - 82.89 pts/$ · $0.45/1M
  5. MiniMax: MiniMax M3 - 55.62 pts/$ · $0.52/1M
  6. Z.ai: GLM 5.2 (free) - 33.78 pts/$ · $1.00/1M
  7. NVIDIA: Nemotron 3 Ultra (free) - 21.81 pts/$ · $1.05/1M
  8. DeepSeek: DeepSeek V4 Pro 0423 - 18.18 pts/$ · $1.98/1M
1118.9839.2559.4279.70.0Intelligence index per dollarAug 9Aug 31Sep 23
  • DeepSeek: DeepSeek V4 Flash 0423
  • DeepSeek: DeepSeek V4 Pro 0423
  • MiniMax: MiniMax M3
  • NVIDIA: Nemotron 3 Ultra (free)
  • OpenAI: GPT-5.6 Luna (batch)
  • Xiaomi: MiMo-V2.5
  • Z.ai: GLM 5.2 (free)
  • Z.ai: GLM 5.3 Flash (batch)

Index date September 22, 2026Data through 00:00 UTC September 23, 2026Calculated 02:00 UTC

How this is calculated

Quality per dollar= Artificial Analysis Intelligence Index ÷ blended price per 1M tokens (75% prompt / 25% completion). No weighting, no composite - a plain ratio of two real, independently sourced numbers, computed fresh from that day's pricing and benchmark snapshot. The Intelligence Index already measures per-response quality (accuracy across a benchmark suite of reasoning, coding, and knowledge tasks) - it is not response volume, so this ratio isn't "cheap model, more tries." A small quality gap next to a huge price gap is what drives a cheap model to the top.

What this ratio can't tell you:

  • Consistency - whether a model reliably hits the same quality run to run, not just on average.
  • Cost of a wrong answer in your specific use case - a bad response in production code is not the same as a bad response in casual chat.
  • Whether a task needs multiple attempts to reach a usable result.

This is a real, industry-standard way to pre-filter candidates by cost-efficiency - not a final verdict. For high-stakes tasks, weigh the absolute Intelligence Index (shown alongside every ratio below) more heavily than the ratio itself, and check a model's measured endpoint uptime for its consistency track record.

Top 15 models by quality per dollar

Intelligence Index points per dollar of blended price - the 15 best-value models among all 94 with a published quality score. This is a ratio, not a capability ranking: an ultra-cheap model with modest absolute quality can outscore a much stronger, pricier one, purely because price sits near zero. Each bar shows its own absolute Intelligence Index and price alongside the ratio, so you can see which is driving the number.

01inclusionAI: Ling 3.0 Flash654.0pts/$Quality 20.6 · $0.03/1M
02inclusionAI: Ling 3.0 Flash VL (free)273.3pts/$Quality 24.6 · $0.09/1M
03inclusionAI: Ling 3.0 Flash Fin (free)251.1pts/$Quality 22.6 · $0.09/1M
04OpenAI: gpt-oss-20b (batch)250.0pts/$Quality 9.0 · $0.04/1M
05DeepSeek: DeepSeek V4.1 Flash (batch)197.5pts/$Quality 39.5 · $0.20/1M
06OpenAI: GPT-6 Luna (batch)186.5pts/$Quality 37.3 · $0.20/1M
07DeepSeek: DeepSeek V4 Flash 0423180.5pts/$Quality 34.3 · $0.19/1M
08Z.ai: GLM 5.3 Flash (batch)176.0pts/$Quality 41.8 · $0.24/1M
09NVIDIA: Nemotron 3.5 Lightning (free)125.9pts/$Quality 12.9 · $0.10/1M
10IBM: Granite 4.2 8B103.3pts/$Quality 11.1 · $0.11/1M
11NVIDIA: Nemotron 3 Nano 30B A3B101.7pts/$Quality 8.9 · $0.09/1M
12Google: Gemma 4 31B (free)101.0pts/$Quality 15.4 · $0.15/1M
13Tencent: Hy3 preview88.8pts/$Quality 25.3 · $0.29/1M
14Xiaomi: MiMo-V2.6-Pro85.1pts/$Quality 46.3 · $0.54/1M
15OpenAI: GPT-5.6 Luna (batch)82.9pts/$Quality 37.3 · $0.45/1M

Open vs. closed AI, tracked daily

80 days of real usage history

Open-weight models now handle 78% of real token volume on OpenRouter - up 8.9 points since July 5, 2026. No other site tracks this as a living, dated trend.

100%75%50%25%0%Share of usageJul 5Aug 13Sep 22
  • Open-source
  • Closed

Best AI model for…

Ranked by real usage share for that specific task, not overall popularity.

code

Code Generation

  1. 1. Z.ai: GLM 5.3 Flash (batch)19%
  2. 2. DeepSeek: DeepSeek V4.1 Flash (batch)16%
  3. 3. OpenAI: GPT-5.6 Luna (batch)14%

Z.ai: GLM 5.3 Flash (batch) leads with 19% of task usage, 3 points ahead of DeepSeek: DeepSeek V4.1 Flash (batch).

Full code generation ranking, 10 models

general

Content Writing

  1. 1. DeepSeek: DeepSeek V4 Flash 042331%
  2. 2. OpenAI: GPT-5.6 Luna (batch)5%
  3. 3. Google: Gemini 2.5 Flash Lite (batch)4%

DeepSeek: DeepSeek V4 Flash 0423 leads with 31% of task usage, 26 points ahead of OpenAI: GPT-5.6 Luna (batch).

Full content writing ranking, 9 models

code

Debugging

  1. 1. DeepSeek: DeepSeek V4.1 Flash (batch)17%
  2. 2. Z.ai: GLM 5.3 Flash (batch)15%
  3. 3. OpenAI: GPT-5.6 Luna (batch)10%

DeepSeek: DeepSeek V4.1 Flash (batch) leads with 17% of task usage, 2 points ahead of Z.ai: GLM 5.3 Flash (batch).

Full debugging ranking, 10 models

general

Translation

  1. 1. Tencent: Hy-MT2-1.8B19%
  2. 2. Tencent: Hy-MT2-7B16%
  3. 3. Tencent: Hy-MT2-30B-A3B12%

Tencent: Hy-MT2-1.8B leads with 19% of task usage, 3 points ahead of Tencent: Hy-MT2-7B.

Full translation ranking, 9 models

general

Summarization

  1. 1. DeepSeek: DeepSeek V4 Flash 042319%
  2. 2. OpenAI: GPT-5.6 Luna (batch)6%
  3. 3. Mistral: Mistral Nemo6%

DeepSeek: DeepSeek V4 Flash 0423 leads with 19% of task usage, 13 points ahead of OpenAI: GPT-5.6 Luna (batch).

Full summarization ranking, 9 models

general

Customer Support

  1. 1. Google: Gemini 2.5 Flash (batch)16%
  2. 2. DeepSeek: DeepSeek V4 Flash 042313%
  3. 3. OpenAI: GPT-5.6 Luna (batch)9%

Google: Gemini 2.5 Flash (batch) leads with 16% of task usage, 2 points ahead of DeepSeek: DeepSeek V4 Flash 0423.

Full customer support ranking, 9 models

general

Roleplay & Fiction

  1. 1. DeepSeek: DeepSeek V4 Flash 042332%
  2. 2. DeepSeek: DeepSeek V4.1 Flash (batch)8%
  3. 3. Mistral: Mistral Nemo6%

DeepSeek: DeepSeek V4 Flash 0423 leads with 32% of task usage, 24 points ahead of DeepSeek: DeepSeek V4.1 Flash (batch).

Full roleplay & fiction ranking, 9 models

general

Q&A & Knowledge

  1. 1. DeepSeek: DeepSeek V4 Flash 042322%
  2. 2. Z.ai: GLM 5.3 Flash (batch)10%
  3. 3. DeepSeek: DeepSeek V4.1 Flash (batch)6%

DeepSeek: DeepSeek V4 Flash 0423 leads with 22% of task usage, 12 points ahead of Z.ai: GLM 5.3 Flash (batch).

Full q&a & knowledge ranking, 9 models

general

Classification

  1. 1. DeepSeek: DeepSeek V4 Flash 042317%
  2. 2. Google: Gemini 2.5 Flash Lite (batch)6%
  3. 3. OpenAI: GPT-5.6 Luna (batch)6%

DeepSeek: DeepSeek V4 Flash 0423 leads with 17% of task usage, 11 points ahead of Google: Gemini 2.5 Flash Lite (batch).

Full classification ranking, 9 models

Rankings by price and specialty

Comparing a 7B model against a flagship on one scale is meaningless. Each of these is a self-contained "best model" claim within a comparable price tier or use case.

By price

By specialty

Daily snapshots

Every published day gets its own permanent, unchanging permalink - a verifiable point in time for citations, not a page that silently drifts.

How this is calculated

No composite index. No invented weights. Two plain, independently-checkable numbers, each plotted as a time series instead of a single snapshot:

Quality per dollar

Artificial Analysis Intelligence Index ÷ blended price per 1M tokens (75% prompt / 25% completion). A plain division, recomputed from that day's real pricing and benchmark snapshot.

Usage share

Real token volume through OpenRouter, as a share of that day's total - no sampling, no estimation.

We tried composite indices three times this year (a weighted-sum composite, a weighted-geometric-mean composite, and a usage-vs-quality divergence index) and killed all three after finding real defensibility problems - see the full methodology page for that history. The lesson: publish real numbers over time, not invented ones.

Source: OpenRouter (openrouter.ai/rankings), as of September 23, 2026.

Benchmark scores: Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings).

Methodology: www.smophy.ai/benchmark/methodology

Token counts originate from each provider's own tokenizer and are not directly comparable across providers.

Journalist or researcher? Pre-computed citable stats and CSV/JSON exports →