Tests explained · People's votes
People's votes on long, detailed requests
People compared two anonymous answers to long requests with lots of detail to take in.
Run by
How to read it
A rating from thousands of head-to-head votes. Higher means people preferred its answers more often. The thin line is the likely range.
Used for
3 tasks on this site.
How the models did
Rating from blind votes · Higher is better
Claude Opus 5.5 and Claude Fable 5.1 are neck and neck at the top, followed by Gemini 3.8 Flash and Kimi K3.
Source: LMArena · votes as of 25 Sept 2026
Where we use it
Each task page labels this test by how closely it matches the task.
| Task | How close |
|---|---|
| Best AI for writing a book | Tests a related skill |
| Best AI for documents, PDFs and notes | Tests a related skill |
| Best AI for notes and meeting notes | Tests a related skill |