PodcastIntelligence Snacks · Weekly conversations about AI, software and business, from the team behind Normie Mode.Listen →
Tests explained · Test of real work

Building features in real software projects

113 original tasks across 91 real code projects in five languages, checked with hand-written tests.

Run by

Datacurve

How to read it

The share of tasks the model got right. The thin line, where shown, is the likely range.

Used for

1 task on this site.

How the models did

Share of tasks done correctly · Higher is better

GPT-6 Astra leads, followed by Gemini 3.8 Flash and GPT-5.6 Sol.

GPT-6 Astra · OpenAI74%
Gemini 3.8 Flash · Google DeepMind74%
GPT-5.6 Sol · OpenAI73%
Kimi K3 · Moonshot AI69%
GPT-5.6 Luna · OpenAI67%
Claude Sonnet 5 · Anthropic54%
Gemini 3.6 Flash · Google DeepMind47%
Gemini 3.1 Pro · Google DeepMind12%
Claude Fable 5.1 · AnthropicNot tested yet
Claude Haiku 4.5 · AnthropicNot tested yet
Claude Opus 5.5 · AnthropicNot tested yet
DeepSeek V4.1 Flash · DeepSeekNot tested yet
DeepSeek V4 Pro · DeepSeekNot tested yet
Gemini 3.5 Flash-Lite · Google DeepMindNot tested yet
GPT-6 Luna · OpenAINot tested yet
GPT-6 Sol · OpenAINot tested yet
Source: Datacurve via Epoch AI · CC BY 4.0

Where we use it

Each task page labels this test by how closely it matches the task.

TaskHow close
Best AI for coding and building appsTests this task
Intelligence Snacks newsletter

The big AI ideas each week, in your inbox