How do you check whether your docs/site are what AI assistants actually learn from?
by•
Something I've been noticing: when people ask an AI assistant about a product, the answer often comes from old blog posts, third-party listicles, or a pricing page from 2 years ago, not from the current site.
For those of you running SaaS products: have you changed how you write your docs, changelog, or pricing page because of this? Things like structured FAQs, llms.txt, or a public changelog. Did any of it make a difference you could measure?
25 views
Replies
I have not measured a difference, so this is only what we did, plus one thing I found this morning while checking it to answer you.
The copies an assistant quotes are mostly the ones nobody rereads when a page changes: the meta description, the share card text, a directory listing built out of them. We learned that when a directory printed our homepage share text word for word as our tagline, after the page itself had moved on. Since then a sentence in a meta tag gets checked like the page, because it is copy that gets republished.
We do serve an llms.txt. It is hand written and short on purpose, nine links with a sentence each, the free tools, the template hubs, the guides and the pricing page, not the two hundred and fifty pages the sitemap already lists. A test fails if one of those links points at a page we no longer publish, and another fails if the file grows past twenty five links, because a generated dump of every page is still a valid llms.txt, just a useless one. The one number in it, how many free invoices a new account gets, is read from the same constant the signup code uses, so it cannot drift.
What no test checks is the sentences. Rereading ours today, the summary line says credits are spent as you send, and the code spends one when an invoice is issued, whether or not it is ever emailed. Small, and exactly the kind of thing an assistant repeats with total confidence. We fixed the same wording in another description last week and missed this copy of it. So my answer is that the file written for AI is also the easiest one to stop reading yourself.
Two habits that help regardless of what the models do: keep prices on one page and point everything else at it, so there is one place to go stale, and keep year specific figures out of evergreen guides, linking to the source instead. Ours carry a reviewed date rather than a year in the text.
Has anyone seen an assistant quote their llms.txt back, or is it still mostly the listicles
@alex_goldwyn On the checking part: the simplest thing I've done is ask ChatGPT, Claude, Gemini and Perplexity about the product directly and look at what they pull. Perplexity always lists its sources, the others do when they search the web. That shows you fast whether they're pulling your current site or an old third-party page.
Measuring it is the hard part for me. GA4 won't tell you which page someone landed on when they came from an AI answer, but Bing Webmaster Tools does show it, and on Bing, Copilot was citing new pages within a few days of publishing. We also just had someone sign up through ChatGPT, yet when I ask ChatGPT myself we don't come up, so I have no idea what they asked it.