SmophyAI

SmophyAI Field Tests · Method v1

Real work · First-party evidence

SmophyAI Field Tests

How do frontier AI models perform when they have a job to do? SmophyAI Field Tests send the same frozen prompt to six models side by side in Multi-Chat, with three runs per model and verifiable checks wherever possible. We publish the results, method and raw answers for coding, analysis, research and writing. These first-party tests complement the benchmark tracker, which reports scores, prices and usage from other sources.

Field Test Board

No published results yet. Method v1.

Models whose means are within five points of each other and whose ranges overlap share a tier; a model needs at least six tasks across three categories before it receives a tier.

Field Test Board: model versions and web search settings, scored from 0 to 100
ModelWebTasksRunsMean score (0 to 100)RangeCodingAnalysisResearchWritingTier
The board will appear here after the first verified Field Test is published.

Published tests

Every published Field Test and every test with no material difference
Test idTaskDate (UTC)ModelsCategoryWebResults
No Field Tests have been published yet. Each verified test will be listed here with its results and raw answers.

Monthly board snapshots

The board is frozen after each month closes, so its results remain citable.

No monthly snapshots yet.