Turning on web search is what makes AI brand recommendations unstable
I'm Thorsten, co-founder of Findabl. We launch here on Aug 14. Before that I wanted to share what we found, because it changed what we built.
We wrote 26 buyer questions ahead of time and locked the list, so we couldn't quietly tidy it up later. None of them mention a brand by name. That matters. Ask an engine "is Salesforce any good" and you're measuring politeness. Ask "what's the best CRM" and you're measuring the shortlist. Four engines, three setups, 2,580 runs.
Then we asked each question twice in a row and checked whether the top brand stayed the same.
With web search switched off: ChatGPT 78%, Claude 87%, Gemini 78%
With web search on, which is what a real user gets: ChatGPT 48%, Claude 50%, Gemini 54%
So about half the time, asking twice gets you a different winner. And it isn't the model's randomness setting doing it. It's the live web lookup. Only 27% of our questions gave the same top brand across all four engines, and a typical brand turned up in 1 answer out of 10.
We weren't first to notice this. Rand Fishkin and Patrick O'Donnell published a study in January: ask ChatGPT or Google's AI the same thing a hundred times, and the chance any two answers return the same list of brands is under 1 in 100. Schulte and colleagues made the case in April that visibility has to be read as a range, not a single reading. Ronald Sielinski argues it in plain statistical terms: one scan is a sample, so it deserves a margin of error rather than a hard number.
What we think is new is narrower. We pinned the instability on web search rather than on the model itself, and we measured the gap between who gets recommended and who gets cited as the source.
That second part is what makes this useful rather than just depressing.
The answers move around. The sources behind them barely do. Across 376 different sites that got cited, the top 10 accounted for 28.6% of all citations, and one directory alone accounted for 8.2%.
So there's a lot of churn sitting on top of a very small set of sources. Chasing the answer is pointless, it changes while you watch. Getting into the handful of sources the engines keep returning to is not pointless, and that's a list you can actually write down.
That's what we ended up building. Questions that never mention your name. Every answer kept, so you read what the engine said instead of trusting a score. Enough repeat runs to put a margin of error on the number, so nobody celebrates a 67% jump that was noise. And a clear split between who got recommended and which sites the engine leaned on to say it.
For anyone here building on top of retrieval: when your output depends on what the search step pulled back, how do you tell a real problem from run-to-run noise?
Replies