Real data comparison
Z.ai: GLM 5.2 (batch) vs StepFun: Step 3.7 Flash
On quality, Z.ai: GLM 5.2 (batch) leads (52.6 vs 30.9). On price, Z.ai: GLM 5.2 (batch) is cheaper ($0.11 vs $0.44 per 1M). In real production usage, Z.ai: GLM 5.2 (batch) is used more (rank #4 vs #11 of 291). Across 6 real task categories where both models have classified traffic, Z.ai: GLM 5.2 (batch) handles more of the traffic in 5 of them.
- Z.ai: GLM 5.2 (batch)
- StepFun: Step 3.7 Flash
Head-to-head by real task
Z.ai: GLM 5.2 (batch) leads 5 of 6 shared task categories by real classified traffic share; StepFun: Step 3.7 Flash leads 0. Not a benchmark score - this is what developers actually route to each model for, per OpenRouter's task classification.
| Task | Z.ai: GLM 5.2 (batch) | StepFun: Step 3.7 Flash |
|---|---|---|
| Frontend & UI | 9.7% | 2.6% |
| SQL & Database | 6.6% | 2.6% |
| DevOps & Config | 6.2% | 2.8% |
| Shell Execution | 5.8% | 3.1% |
| File I/O | 5.0% | 2.9% |
| DevOps | 4.4% | 4.4% |
Usage share over time
- Z.ai: GLM 5.2 (batch)
- StepFun: Step 3.7 Flash
Frequently asked questions
Z.ai: GLM 5.2 (batch) vs StepFun: Step 3.7 Flash: which is better quality?
Z.ai: GLM 5.2 (batch) scores higher on Artificial Analysis's Intelligence Index (52.6 vs 30.9).
Z.ai: GLM 5.2 (batch) vs StepFun: Step 3.7 Flash: which is cheaper?
Z.ai: GLM 5.2 (batch) is cheaper at $0.11/1M vs $0.44/1M.
Z.ai: GLM 5.2 (batch) vs StepFun: Step 3.7 Flash: which is used more?
Z.ai: GLM 5.2 (batch) has higher real OpenRouter usage share (rank #4 vs #11 of 291).
Z.ai: GLM 5.2 (batch) vs StepFun: Step 3.7 Flash: which wins more real-world task categories?
Z.ai: GLM 5.2 (batch) handles more traffic in 5 of 6 shared task categories where both models have classified OpenRouter usage.
Source: OpenRouter (openrouter.ai/rankings), as of August 9, 2026.
Benchmark scores: Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings).
Methodology: smophy.ai/benchmark/methodology
Token counts originate from each provider's own tokenizer and are not directly comparable across providers.
