Tests explained · Test of real work
Would a real maintainer accept its code?
Coding tasks in 36 major open-source projects, judged on whether the code is good enough to merge, not just whether it runs.
Run by
How to read it
The share of tasks the model got right. The thin line, where shown, is the likely range.
Used for
1 task on this site.
How the models did
Share of tasks done correctly · Higher is better
GPT-6 Astra leads, followed by Claude Fable 5.1 and GPT-5.6 Sol.
Source: Cognition via Epoch AI · CC BY 4.0
Where we use it
Each task page labels this test by how closely it matches the task.
| Task | How close |
|---|---|
| Best AI for coding and building apps | Tests this task |