Bookshelf Benchmark

Can AI chat models read a bookshelf? We gave Claude, Gemini and GPT the same six real bookshelf photos and asked them to list every book whose title they could read, then checked their answers against a hand-verified list.

The photos

Six real bookshelf photos: two easy, two medium and two hard. The number is how many books are in the answer key for that photo. Click a photo to open it larger.

Experiment 1: all six photos in one message

Each row is one model and mode, given all six photos at once in the normal chat app. The scale beside the buttons shows which shades mean a high score. Rows are ordered by overall F1. Hover or focus a cell for the details.

Photos a and b are easy, c and d medium, e and f hard. Book counts are the number of books in the answer key for that photo.

Experiment 2: the hard photos, one per chat

The two hard photos, sent on their own instead of with the other five. Each pair of columns compares the same model on the same photo: six photos at once versus one photo per chat.

What we found

  1. Precision is high almost everywhere; recall is what separates the models. When a model names a book, it is nearly always really there. The difference is how many of the books each model manages to list.
  2. Claude reads crowded shelves far better. On the two hard photos, every Claude setup found more books than every Gemini and GPT setup. With one photo per chat, even the weakest Claude run found about twice as many as the best Gemini run.
  3. Gemini 3.8 Flash is best on the easy shelves. On the two easy photos it scored higher than every other model, including all the Claude runs.
  4. Sending six photos at once held Claude back. With one photo per chat, every Claude run found at least as many books, and Claude Opus 5.5 at low effort caught up with Opus 5.5 at high effort. Most of the low-versus-high effort gap in experiment 1 came from the six-photo setup.
  5. Gemini's limit is how well it reads, not the batching. Sent one hard photo at a time, Gemini still listed about the same number of books. It just named fewer wrong ones.
  6. Newer models improved. Opus 5.5 beat Opus 5, and Sonnet 5.5 beat Sonnet 5 by a wide margin, at the same effort.

Where models went wrong

Real examples from experiment 1. The model's answer is struck through; the book on the shelf follows.

Right author, wrong book

The model reads the author and names their better-known title.

  • The Magic of the Lost Temple→The Magic of the Lost Earrings · 5 runs
  • Behave→Determined · Sapolsky
  • Bit of a Blur→All Cheeses Great & Small · Alex James

Missed by every model

41 of the 295 books were found by none of the 12 runs.

  • Hindi titles · जादू, मेघनाद, डेरे वाले
  • Upside-down or thin spines · Silas Marner, Hunted
  • Partly hidden covers · The Great Gatsby, Wild Fictions

Too vague to count

Part of the title, not enough to say which book.

  • Harry Potter→Harry Potter and the Goblet of Fire
  • Diary of a Wimpy Kid→Diary of a Wimpy Kid: The Long Haul
  • British Art→The Thames and Hudson Encyclopaedia of British Art

Marked wrong, but really there

On the shelf, but left out of the answer key as too hard to confirm.

  • The Secret Life of Bees · named in 6 runs
  • Jiggy McCue · 3 runs
  • The Bitcoin Standard · 2 runs

Model vs harness: same model, very different result

The answer key itself was built with Claude Opus 5.5, the same model that scored best here, and together with a person it found nearly every readable book. Yet in a single chat, Opus 5.5 at high effort found about 77% of the books on the hardest photo. The model was the same; what changed was how it was used.

Building the answer keyA single chat
Image detailFull-resolution originals, cropped into small regions and enlargedThe whole photo at once; chat apps usually shrink large images before the model sees them
PassesOver a hundred zoomed crops, shelf by shelf, some rotated so spines read uprightOne look at the whole shelf
TaskCheck and extend an existing listList everything from scratch
CheckingA person confirmed every addition against a cropThe model alone, in one reply

The biggest factor is resolution: in a shrunk 24-megapixel photo, a spine title can end up only a few pixels tall and simply can't be read. Effort mattered less: in experiment 2, low effort matched high once the six-photo batching was removed. In other words, the surrounding system (zooming, multiple passes, verification) can matter more than the model or its settings.

Possible next experiments

How it was measured