Tests explained · Test of real work
Humanity's Last Exam
Very hard questions written by experts across many subjects.
How to read it
The share of tasks the model got right. The thin line, where shown, is the likely range.
Used for
2 tasks on this site.
How the models did
Share of tasks done correctly · Higher is better
GPT-6 Astra leads, followed by Claude Fable 5.1 and Gemini 3.1 Pro.
Source: Center for AI Safety, Scale AI via Epoch AI · CC BY 4.0
Where we use it
Each task page labels this test by how closely it matches the task.
| Task | How close |
|---|---|
| Best AI for research | Tests a related skill |
| Best AI for academic research | Tests a related skill |