PodcastIntelligence Snacks · Weekly conversations about AI, software and business, from the team behind Normie Mode.Listen →
Tests explained · People's votes

People's votes on expert questions

People compared two anonymous answers to questions needing expert knowledge.

Run by

LMArena

How to read it

A rating from thousands of head-to-head votes. Higher means people preferred its answers more often. The thin line is the likely range.

Used for

6 tasks on this site.

How the models did

Rating from blind votes · Higher is better

Claude Opus 5.5 and Gemini 3.8 Flash are neck and neck at the top, followed by Claude Fable 5.1 and Kimi K3.

Claude Opus 5.5 · Anthropic1st
Gemini 3.8 Flash · Google DeepMind2nd
Claude Fable 5.1 · Anthropic3rd
Kimi K3 · Moonshot AI4th
GPT-5.6 Sol · OpenAI5th
DeepSeek V4.1 Flash · DeepSeek6th
Claude Sonnet 5 · Anthropic7th
Gemini 3.6 Flash · Google DeepMind8th
GPT-6 Astra · OpenAI9th
Gemini 3.1 Pro · Google DeepMind10th
GPT-5.6 Luna · OpenAI11th
DeepSeek V4 Pro · DeepSeek12th
Claude Haiku 4.5 · Anthropic13th
Gemini 3.5 Flash-Lite · Google DeepMind14th
GPT-6 Luna · OpenAINot tested yet
GPT-6 Sol · OpenAINot tested yet
Source: LMArena · votes as of 25 Sept 2026

Where we use it

Each task page labels this test by how closely it matches the task.

TaskHow close
Best AI for essays and academic writingTests a related skill
Best AI for studying and homeworkTests a related skill
Best AI for medical and nursing studentsTests a related skill
Best AI for teachersTests a related skill
Best AI for researchTests a related skill
Best AI for academic researchTests a related skill
Intelligence Snacks newsletter

The big AI ideas each week, in your inbox