PodcastIntelligence Snacks · Weekly conversations about AI, software and business, from the team behind Normie Mode.Listen →
Tests explained · People's votes

People's votes on following instructions

People compared two anonymous answers to requests with specific instructions.

Run by

LMArena

How to read it

A rating from thousands of head-to-head votes. Higher means people preferred its answers more often. The thin line is the likely range.

Used for

6 tasks on this site.

How the models did

Rating from blind votes · Higher is better

Claude Opus 5.5 and Claude Fable 5.1 are neck and neck at the top, followed by Gemini 3.8 Flash and Kimi K3.

Claude Opus 5.5 · Anthropic1st
Claude Fable 5.1 · Anthropic2nd
Gemini 3.8 Flash · Google DeepMind3rd
Kimi K3 · Moonshot AI4th
GPT-5.6 Sol · OpenAI5th
DeepSeek V4.1 Flash · DeepSeek6th
Gemini 3.6 Flash · Google DeepMind7th
Gemini 3.1 Pro · Google DeepMind8th
GPT-6 Astra · OpenAI9th
Claude Sonnet 5 · Anthropic10th
DeepSeek V4 Pro · DeepSeek11th
GPT-5.6 Luna · OpenAI12th
Gemini 3.5 Flash-Lite · Google DeepMind13th
Claude Haiku 4.5 · Anthropic14th
GPT-6 Luna · OpenAINot tested yet
GPT-6 Sol · OpenAINot tested yet
Source: LMArena · votes as of 25 Sept 2026

Where we use it

Each task page labels this test by how closely it matches the task.

TaskHow close
Best AI for writingTests a related skill
Best AI for essays and academic writingTests a related skill
Best AI for teachersTests a related skill
Best AI for CVs, job applications and interviewsTests a related skill
Best AI for notes and meeting notesTests a related skill
Best AI for presentations and PowerPointTests a related skill
Intelligence Snacks newsletter

The big AI ideas each week, in your inbox