PodcastIntelligence Snacks · Weekly conversations about AI, software and business, from the team behind Normie Mode.Listen →
Tests explained · People's votes

People's votes on science questions

People compared two anonymous answers to questions in the life, physical and social sciences.

Run by

LMArena

How to read it

A rating from thousands of head-to-head votes. Higher means people preferred its answers more often. The thin line is the likely range.

Used for

3 tasks on this site.

How the models did

Rating from blind votes · Higher is better

Claude Opus 5.5 and Claude Fable 5.1 are neck and neck at the top, followed by Gemini 3.8 Flash and Kimi K3.

Claude Opus 5.5 · Anthropic1st
Claude Fable 5.1 · Anthropic2nd
Gemini 3.8 Flash · Google DeepMind3rd
Kimi K3 · Moonshot AI4th
Gemini 3.1 Pro · Google DeepMind5th
Gemini 3.6 Flash · Google DeepMind6th
DeepSeek V4.1 Flash · DeepSeek7th
Claude Sonnet 5 · Anthropic8th
DeepSeek V4 Pro · DeepSeek9th
GPT-6 Astra · OpenAI10th
GPT-5.6 Sol · OpenAI11th
Gemini 3.5 Flash-Lite · Google DeepMind12th
GPT-5.6 Luna · OpenAI13th
Claude Haiku 4.5 · Anthropic14th
GPT-6 Luna · OpenAINot tested yet
GPT-6 Sol · OpenAINot tested yet
Source: LMArena · votes as of 25 Sept 2026

Where we use it

Each task page labels this test by how closely it matches the task.

TaskHow close
Best AI for health questionsTests a related skill
Best AI for maths and scienceTests a related skill
Best AI for chemistryTests this task
Intelligence Snacks newsletter

The big AI ideas each week, in your inbox