Tests explained · Test of real work
Research-level maths
New, unpublished maths problems that take specialists hours or days.
Run by
How to read it
The share of tasks the model got right. The thin line, where shown, is the likely range.
Used for
1 task on this site.
How the models did
Share of tasks done correctly · Higher is better
GPT-6 Astra and Claude Fable 5.1 are neck and neck at the top, followed by GPT-5.6 Sol and GPT-5.6 Luna.
Source: Epoch AI via Epoch AI · CC BY 4.0
Where we use it
Each task page labels this test by how closely it matches the task.
| Task | How close |
|---|---|
| Best AI for maths and science | Tests this task |