PodcastIntelligence Snacks · Weekly conversations about AI, software and business, from the team behind Normie Mode.Listen →
Tests explained · Test of real work

Answering short factual questions correctly

Short questions that each have one checkable answer.

How to read it

The share of tasks the model got right. The thin line, where shown, is the likely range.

Used for

2 tasks on this site.

How the models did

Share of tasks done correctly · Higher is better

GPT-6 Astra and Gemini 3.1 Pro are neck and neck at the top, followed by Claude Fable 5.1 and Gemini 3.8 Flash.

GPT-6 Astra · OpenAI76%
Gemini 3.1 Pro · Google DeepMind74%
Claude Fable 5.1 · Anthropic71%
Gemini 3.8 Flash · Google DeepMind70%
GPT-5.6 Sol · OpenAI70%
Gemini 3.6 Flash · Google DeepMind66%
DeepSeek V4 Pro · DeepSeek53%
Kimi K3 · Moonshot AI51%
GPT-5.6 Luna · OpenAI41%
Claude Sonnet 5 · Anthropic34%
Claude Haiku 4.5 · Anthropic13%
Claude Opus 5.5 · AnthropicNot tested yet
DeepSeek V4.1 Flash · DeepSeekNot tested yet
Gemini 3.5 Flash-Lite · Google DeepMindNot tested yet
GPT-6 Luna · OpenAINot tested yet
GPT-6 Sol · OpenAINot tested yet
Source: Google DeepMind via Epoch AI · CC BY 4.0

Where we use it

Each task page labels this test by how closely it matches the task.

TaskHow close
Best AI for researchTests a related skill
Best AI for academic researchTests a related skill
Intelligence Snacks newsletter

The big AI ideas each week, in your inbox