SmophyAI

Published daily · Next update 02:00 UTC

Real data comparison

OpenAI: GPT-5.6 Luna (batch) vs Tencent: Hy3

On price, Tencent: Hy3 is cheaper ($0.23 vs $0.45 per 1M). In real production usage, Tencent: Hy3 is used more (rank #7 vs #8 of 353). Across 6 real task categories where both models have classified traffic, OpenAI: GPT-5.6 Luna (batch) handles more of the traffic in 5 of them.

  • OpenAI: GPT-5.6 Luna (batch)
  • Tencent: Hy3
Price ($/1M blended)$0.45 / $0.23
Average usage share6.26% / 8.77%
Best provider uptime100.0% / 100.0%
Context window1,050,000 / 262,144

Head-to-head by real task

OpenAI: GPT-5.6 Luna (batch) leads 5 of 6 shared task categories by real classified traffic share; Tencent: Hy3 leads 1. Not a benchmark score - this is what developers actually route to each model for, per OpenRouter's task classification.

TaskOpenAI: GPT-5.6 Luna (batch)Tencent: Hy3
Tool Dispatch31.6%1.6%
Multi-step Planning7.4%1.9%
Research & Reports7.4%2.1%
Translation5.8%2.4%
File I/O3.2%5.7%
Workflow Execution6.8%4.4%

Usage share over time

25.35%19.01%12.68%6.34%0.00%Share of usageJul 6Aug 14Sep 22
  • OpenAI: GPT-5.6 Luna (batch)
  • Tencent: Hy3

Frequently asked questions

OpenAI: GPT-5.6 Luna (batch) vs Tencent: Hy3: which is better quality?

Quality data is unavailable for one or both models.

OpenAI: GPT-5.6 Luna (batch) vs Tencent: Hy3: which is cheaper?

Tencent: Hy3 is cheaper at $0.23/1M vs $0.45/1M.

OpenAI: GPT-5.6 Luna (batch) vs Tencent: Hy3: which is used more?

Tencent: Hy3 has higher real OpenRouter usage share (rank #7 vs #8 of 353).

OpenAI: GPT-5.6 Luna (batch) vs Tencent: Hy3: which wins more real-world task categories?

OpenAI: GPT-5.6 Luna (batch) handles more traffic in 5 of 6 shared task categories where both models have classified OpenRouter usage.

Source: OpenRouter (openrouter.ai/rankings), as of September 23, 2026.

Benchmark scores: Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings).

Methodology: www.smophy.ai/benchmark/methodology

Token counts originate from each provider's own tokenizer and are not directly comparable across providers.