Tests explained · Test of real work
Building features in real software projects
113 original tasks across 91 real code projects in five languages, checked with hand-written tests.
Run by
How to read it
The share of tasks the model got right. The thin line, where shown, is the likely range.
Used for
1 task on this site.
How the models did
Share of tasks done correctly · Higher is better
GPT-6 Astra leads, followed by Gemini 3.8 Flash and GPT-5.6 Sol.
Source: Datacurve via Epoch AI · CC BY 4.0
Where we use it
Each task page labels this test by how closely it matches the task.
| Task | How close |
|---|---|
| Best AI for coding and building apps | Tests this task |