PodcastIntelligence Snacks · Weekly conversations about AI, software and business, from the team behind Normie Mode.Listen →
Best AI for… / maths and science

Best AI for engineering

Engineering calculations, concepts and problem-solving.

Best overall

Claude Fable 5.1

Get it with Claude Max, $100 a month (for up to half your weekly allowance)

Top of 12 models for this task

Best for $25 a month or less

Gemini 3.8 Flash

Get it with Google AI Pro, $19.99 a month

2nd of 12 models for this task

Best freeToo close to call

3rd and 5th of 12 models for this task

Based on the tests for maths and science. The picks follow a fixed rule using the charts below and each app's plans.

Already paying for one? See what each plan gives you here
App and planPriceBest model on it for thisPosition
ChatGPT FreefreeGPT-5.6 Luna8th of 12
ChatGPT Go$8 a monthGPT-5.6 Luna8th of 12
ChatGPT Plus$20 a monthGPT-6 Astra
Only in ChatGPT's Work mode and Codex, not normal chat
4th of 12
ChatGPT Pro$100 a monthGPT-6 Astra
Called GPT-6 Pro in the app
4th of 12
Claude FreefreeClaude Sonnet 59th of 12
Claude Pro$20 a monthClaude Sonnet 5
Its main model, Claude Opus 5.5, is too new to score yet.
9th of 12
Claude Max$100 a monthClaude Fable 5.1
For up to half your weekly allowance
Its main model, Claude Opus 5.5, is too new to score yet.
1st of 12
Gemini FreefreeGemini 3.6 Flash5th of 12
Google AI Plus$4.99 a monthGemini 3.6 Flash5th of 12
Google AI Pro$19.99 a monthGemini 3.8 Flash2nd of 12
Google AI Ultra$99.99 a monthGemini 3.8 Flash2nd of 12
Kimi FreefreeKimi K33rd of 12
Kimi Moderato$19 a monthKimi K33rd of 12

How the models compare

Every model we could score for engineering, combining the tests below.

Overall for engineering

Position on a combined score from 5 tests. The longer the bar, the further ahead · Higher is better

On the tests we have, Claude Fable 5.1 comes out on top, followed by Gemini 3.8 Flash and Kimi K3.

Claude Fable 5.1 · Anthropic1st
Gemini 3.8 Flash · Google DeepMind2nd
Kimi K3 · Moonshot AI3rd
GPT-6 Astra · OpenAI4th
Gemini 3.6 Flash · Google DeepMind5th
GPT-5.6 Sol · OpenAI6th
Gemini 3.1 Pro · Google DeepMind7th
GPT-5.6 Luna · OpenAI8th
Claude Sonnet 5 · Anthropic9th
DeepSeek V4 Pro · DeepSeek10th
Gemini 3.5 Flash-Lite · Google DeepMind11th
Claude Haiku 4.5 · Anthropic12th

The evidence

Competition maths

Tests a related skill · Problems in the style of a top US high-school maths competition. · Higher is better

Claude Fable 5.1, GPT-5.6 Sol and GPT-6 Astra are level at the top.

Claude Fable 5.1 · Anthropic100%
GPT-5.6 Sol · OpenAI100%
GPT-6 Astra · OpenAI100%
Gemini 3.8 Flash · Google DeepMind99%
DeepSeek V4 Pro · DeepSeek99%
GPT-5.6 Luna · OpenAI98%
Kimi K3 · Moonshot AI97%
Gemini 3.1 Pro · Google DeepMind96%
Claude Sonnet 5 · Anthropic95%
Gemini 3.6 Flash · Google DeepMind94%
Gemini 3.5 Flash-Lite · Google DeepMind71%
Claude Haiku 4.5 · Anthropic36%
Source: Epoch AI via Epoch AI · CC BY 4.0

Research-level maths

Tests a related skill · New, unpublished maths problems that take specialists hours or days. · Higher is better

GPT-6 Astra and Claude Fable 5.1 are neck and neck at the top, followed by GPT-5.6 Sol and GPT-5.6 Luna.

GPT-6 Astra · OpenAI94%
Claude Fable 5.1 · Anthropic90%
GPT-5.6 Sol · OpenAI89%
GPT-5.6 Luna · OpenAI82%
Kimi K3 · Moonshot AI72%
Gemini 3.8 Flash · Google DeepMind68%
Claude Sonnet 5 · Anthropic66%
DeepSeek V4 Pro · DeepSeek65%
Gemini 3.1 Pro · Google DeepMind60%
Gemini 3.6 Flash · Google DeepMind59%
Gemini 3.5 Flash-Lite · Google DeepMind26%
Source: Epoch AI via Epoch AI · CC BY 4.0

Graduate-level science questions

Tests a related skill · Biology, physics and chemistry questions written so they can't be answered by searching the web. · Higher is better

GPT-6 Astra and Gemini 3.8 Flash are neck and neck at the top, followed by Gemini 3.1 Pro and Gemini 3.6 Flash; GPT-6 Astra costs about 13 times as much.

GPT-6 Astra · OpenAI96%
Gemini 3.8 Flash · Google DeepMind95%
Gemini 3.1 Pro · Google DeepMind94%
Gemini 3.6 Flash · Google DeepMind94%
GPT-5.6 Sol · OpenAI93%
Kimi K3 · Moonshot AI93%
DeepSeek V4 Pro · DeepSeek92%
GPT-5.6 Luna · OpenAI92%
Claude Sonnet 5 · Anthropic91%
Gemini 3.5 Flash-Lite · Google DeepMind83%
Claude Haiku 4.5 · Anthropic60%
Source: NYU and others via Epoch AI · CC BY 4.0

People's votes on maths questions

Tests a related skill · People compared two anonymous answers to maths questions. · Higher is better

Gemini 3.8 Flash and Claude Fable 5.1 are neck and neck at the top, followed by Gemini 3.6 Flash and Kimi K3; Claude Fable 5.1 costs about 13 times as much.

Gemini 3.8 Flash · Google DeepMind1st
Claude Fable 5.1 · Anthropic2nd
Gemini 3.6 Flash · Google DeepMind3rd
Kimi K3 · Moonshot AI4th
DeepSeek V4.1 Flash · DeepSeek5th
Gemini 3.1 Pro · Google DeepMind6th
GPT-5.6 Sol · OpenAI7th
Claude Sonnet 5 · Anthropic8th
GPT-6 Astra · OpenAI9th
GPT-5.6 Luna · OpenAI10th
DeepSeek V4 Pro · DeepSeek11th
Gemini 3.5 Flash-Lite · Google DeepMind12th
Claude Haiku 4.5 · Anthropic13th
Source: LMArena · votes as of 25 Sept 2026

People's votes on science questions

Tests a related skill · People compared two anonymous answers to questions in the life, physical and social sciences. · Higher is better

Claude Opus 5.5 and Claude Fable 5.1 are neck and neck at the top, followed by Gemini 3.8 Flash and Kimi K3; Claude Fable 5.1 costs about 3 times as much.

Claude Opus 5.5 · Anthropic1st
Claude Fable 5.1 · Anthropic2nd
Gemini 3.8 Flash · Google DeepMind3rd
Kimi K3 · Moonshot AI4th
Gemini 3.1 Pro · Google DeepMind5th
Gemini 3.6 Flash · Google DeepMind6th
DeepSeek V4.1 Flash · DeepSeek7th
Claude Sonnet 5 · Anthropic8th
DeepSeek V4 Pro · DeepSeek9th
GPT-6 Astra · OpenAI10th
GPT-5.6 Sol · OpenAI11th
Gemini 3.5 Flash-Lite · Google DeepMind12th
GPT-5.6 Luna · OpenAI13th
Claude Haiku 4.5 · Anthropic14th
Source: LMArena · votes as of 25 Sept 2026

If you build with it

What the makers charge developers who use the models directly. Using an app? The plans above are what you pay.

Cost per 1,000 typical requests

List price, about 1,500 words in and 500 out per request · Lower is better

GPT-6 Luna is cheapest, followed by DeepSeek V4.1 Flash and GPT-5.6 Luna.

GPT-6 Luna · OpenAI$0.55
DeepSeek V4.1 Flash · DeepSeek$0.72
GPT-5.6 Luna · OpenAI$1.24
DeepSeek V4 Pro · DeepSeek$1.48
Gemini 3.5 Flash-Lite · Google DeepMind$2.35
Gemini 3.8 Flash · Google DeepMind$4.13
Gemini 3.6 Flash · Google DeepMind$4.13
Claude Haiku 4.5 · Anthropic$5.50
Claude Sonnet 5 · Anthropic$11
GPT-6 Sol · OpenAI$11
Gemini 3.1 Pro · Google DeepMind$12
Kimi K3 · Moonshot AI$17
GPT-5.6 Sol · OpenAI$22
Claude Opus 5.5 · Anthropic$22
GPT-6 Astra · OpenAI$55
Claude Fable 5.1 · Anthropic$55
Source: models.dev · MIT
Intelligence Snacks newsletter

The big AI ideas each week, in your inbox