PodcastIntelligence Snacks · Weekly conversations about AI, software and business, from the team behind Normie Mode.Listen →
Tests explained · Test of real work

Research-level maths

New, unpublished maths problems that take specialists hours or days.

Run by

Epoch AI

How to read it

The share of tasks the model got right. The thin line, where shown, is the likely range.

Used for

1 task on this site.

How the models did

Share of tasks done correctly · Higher is better

GPT-6 Astra and Claude Fable 5.1 are neck and neck at the top, followed by GPT-5.6 Sol and GPT-5.6 Luna.

GPT-6 Astra · OpenAI94%
Claude Fable 5.1 · Anthropic90%
GPT-5.6 Sol · OpenAI89%
GPT-5.6 Luna · OpenAI82%
Kimi K3 · Moonshot AI72%
Gemini 3.8 Flash · Google DeepMind68%
Claude Sonnet 5 · Anthropic66%
DeepSeek V4 Pro · DeepSeek65%
Gemini 3.1 Pro · Google DeepMind60%
Gemini 3.6 Flash · Google DeepMind59%
Gemini 3.5 Flash-Lite · Google DeepMind26%
Claude Haiku 4.5 · AnthropicNot tested yet
Claude Opus 5.5 · AnthropicNot tested yet
DeepSeek V4.1 Flash · DeepSeekNot tested yet
GPT-6 Luna · OpenAINot tested yet
GPT-6 Sol · OpenAINot tested yet
Source: Epoch AI via Epoch AI · CC BY 4.0

Where we use it

Each task page labels this test by how closely it matches the task.

TaskHow close
Best AI for maths and scienceTests this task
Intelligence Snacks newsletter

The big AI ideas each week, in your inbox