Real work · First-party evidence
SmophyAI Field Tests
How do frontier AI models perform when they have a job to do? SmophyAI Field Tests send the same frozen prompt to six models side by side in Multi-Chat, with three runs per model and verifiable checks wherever possible. We publish the results, method and raw answers for coding, analysis, research and writing. These first-party tests complement the benchmark tracker, which reports scores, prices and usage from other sources.
Field Test Board
No published results yet. Method v1.
Models whose means are within five points of each other and whose ranges overlap share a tier; a model needs at least six tasks across three categories before it receives a tier.
| Model | Web | Tasks | Runs | Mean score (0 to 100) | Range | Coding | Analysis | Research | Writing | Tier |
|---|---|---|---|---|---|---|---|---|---|---|
| The board will appear here after the first verified Field Test is published. | ||||||||||
Published tests
| Test id | Task | Date (UTC) | Models | Category | Web | Results |
|---|---|---|---|---|---|---|
| No Field Tests have been published yet. Each verified test will be listed here with its results and raw answers. | ||||||
Monthly board snapshots
The board is frozen after each month closes, so its results remain citable.
No monthly snapshots yet.
