Tests explained · People's votes
People's votes on back-and-forth conversations
People compared two anonymous answers in conversations with several back-and-forth turns.
Run by
How to read it
A rating from thousands of head-to-head votes. Higher means people preferred its answers more often. The thin line is the likely range.
Used for
3 tasks on this site.
How the models did
Rating from blind votes · Higher is better
Claude Opus 5.5 and Gemini 3.8 Flash are neck and neck at the top, followed by Claude Fable 5.1 and Kimi K3.
Source: LMArena · votes as of 25 Sept 2026
Where we use it
Each task page labels this test by how closely it matches the task.
| Task | How close |
|---|---|
| Best AI for mental health support | Tests a related skill |
| Best AI for planning and personal admin | Tests a related skill |
| Best AI for learning a language | Tests a related skill |