PodcastIntelligence Snacks · Weekly conversations about AI, software and business, from the team behind Normie Mode.Listen →
Tests explained · Test of real work

Competition maths

Problems in the style of a top US high-school maths competition.

Run by

Epoch AI

How to read it

The share of tasks the model got right. The thin line, where shown, is the likely range.

Used for

2 tasks on this site.

How the models did

Share of tasks done correctly · Higher is better

Claude Fable 5.1, GPT-5.6 Sol and GPT-6 Astra are level at the top.

Claude Fable 5.1 · Anthropic100%
GPT-5.6 Sol · OpenAI100%
GPT-6 Astra · OpenAI100%
Gemini 3.8 Flash · Google DeepMind99%
DeepSeek V4 Pro · DeepSeek99%
GPT-5.6 Luna · OpenAI98%
Kimi K3 · Moonshot AI97%
Gemini 3.1 Pro · Google DeepMind96%
Claude Sonnet 5 · Anthropic95%
Gemini 3.6 Flash · Google DeepMind94%
Gemini 3.5 Flash-Lite · Google DeepMind71%
Claude Haiku 4.5 · Anthropic36%
Claude Opus 5.5 · AnthropicNot tested yet
DeepSeek V4.1 Flash · DeepSeekNot tested yet
GPT-6 Luna · OpenAINot tested yet
GPT-6 Sol · OpenAINot tested yet
Source: Epoch AI via Epoch AI · CC BY 4.0

Where we use it

Each task page labels this test by how closely it matches the task.

TaskHow close
Best AI for studying and homeworkTests a related skill
Best AI for maths and scienceTests this task
Intelligence Snacks newsletter

The big AI ideas each week, in your inbox