PodcastIntelligence Snacks · Weekly conversations about AI, software and business, from the team behind Normie Mode.Listen →
Tests explained · Test of real work

Real paid freelance projects

Projects originally done by freelancers for pay, scored on whether the AI's result would be acceptable to the client.

How to read it

The share of tasks the model got right. The thin line, where shown, is the likely range.

Used for

1 task on this site.

How the models did

Share of tasks done correctly · Higher is better

GPT-6 Astra · OpenAI21%
Claude Fable 5.1 · Anthropic18%
Claude Haiku 4.5 · AnthropicNot tested yet
Claude Opus 5.5 · AnthropicNot tested yet
Claude Sonnet 5 · AnthropicNot tested yet
DeepSeek V4.1 Flash · DeepSeekNot tested yet
DeepSeek V4 Pro · DeepSeekNot tested yet
Gemini 3.1 Pro · Google DeepMindNot tested yet
Gemini 3.5 Flash-Lite · Google DeepMindNot tested yet
Gemini 3.6 Flash · Google DeepMindNot tested yet
Gemini 3.8 Flash · Google DeepMindNot tested yet
GPT-5.6 Luna · OpenAINot tested yet
GPT-5.6 Sol · OpenAINot tested yet
GPT-6 Luna · OpenAINot tested yet
GPT-6 Sol · OpenAINot tested yet
Kimi K3 · Moonshot AINot tested yet

Where we use it

Each task page labels this test by how closely it matches the task.

TaskHow close
Best AI for businessTests a related skill
Intelligence Snacks newsletter

The big AI ideas each week, in your inbox