Tests explained · Test of real work
Answering short factual questions correctly
Short questions that each have one checkable answer.
Run by
How to read it
The share of tasks the model got right. The thin line, where shown, is the likely range.
Used for
2 tasks on this site.
How the models did
Share of tasks done correctly · Higher is better
GPT-6 Astra and Gemini 3.1 Pro are neck and neck at the top, followed by Claude Fable 5.1 and Gemini 3.8 Flash.
Source: Google DeepMind via Epoch AI · CC BY 4.0
Where we use it
Each task page labels this test by how closely it matches the task.
| Task | How close |
|---|---|
| Best AI for research | Tests a related skill |
| Best AI for academic research | Tests a related skill |