Method
How this was built
Enough detail to check the numbers or repeat the work.
Where the complaints came from
Three public places where people talk about Google Photos: reviews on Google Play (US, India and UK storefronts, newest and most relevant), reviews on the App Store (ten countries, recent and most helpful), and threads on the Google Photos Help Community that match 25 phrases about finding photos. That gave 15,266 items. A word filter kept the 4,070 that mention searching, finding, remembering, scrolling, faces, dates or screenshots.
Reddit and YouTube were planned but could not be reached from the tools used for this project, so they are not in the data.
How each complaint was labelled
A language model (Gemini Flash Lite) read each item and filled in the same fields every time: is it about finding a photo; what kind of photo; what the person remembered and what they had forgotten; how they searched and the exact words if they gave them; where it broke (couldn’t find the words, search misread them, the photo never showed up, too many look-alikes, no way to narrow down, a face or place was never tagged, they thought the photo was gone, or it used to work and an update broke it); what they did next; how upset they were; and one quote in their own words.
2,080 items turned out to be about finding a photo (1,154 from Google Play, 520 from the App Store, 406 from the Help Community). The labels are the model’s reading of the text, not the person’s own answers. Spot checks of a few hundred looked right, but expect some noise at the edges.
Two things skew the data. App reviews are mostly written by people who are annoyed, and many recent ones are about the 2025 to 2026 change to AI search in Google Photos. Help Community threads are mostly “my photos are missing”. Both are labelled and can be hidden with one filter.
How “Ask” answers
The question goes to a language model that can run four kinds of lookup over the labelled data: count items by a field, cross two fields, search the text, or pull examples for a category. It uses the results to write an answer and marks each quote with the id of the review it came from, so you can check it in the explorer.
Talking to people, and what didn’t happen
A 15-question survey asked about library size, the last time someone couldn’t find a photo, what they remembered, what they typed, and what happened. Two people answered. Interviews were planned but did not happen in the time available. That is a real gap, and the deck says so.
To get closer to how people describe the problem, the 40 most detailed first-person accounts in the public data were read with the interview sheet: what they remembered, the exact words they typed, what came back, what they did next. These are public posts, not interviews, and they are labelled that way wherever they are used.
The test library behind Recall
Google no longer lets outside apps read someone’s full Google Photos library, so Recall searches a stand-in. It has 154 photos: openly licensed photographs from Unsplash arranged as one person’s timeline from 2019 to 2026 (a Goa trip, a wedding, a trek, a dog, food, work, home), plus screenshots and receipts drawn with made-up details. A vision model wrote a short description of each photo (what is in it, who, where, any visible text, time of day) and each one carries a date. Photographers are credited in the library file.
When you type a memory, one model call turns it into soft constraints (years, indoors or out, how many people, kind of photo, keywords). Those score every photo. A second call looks at the top 24, picks six with a reason each, and chooses one question that would split them best.
Code
The scrapers, the labelling script, the library indexer and this site are in one repository, linked from the deck. Nothing here is made by or with Google.