Has ChatGPT actually changed how you discover software?

by•

I’ve noticed a pretty big change in my own behavior over the last year.

If I’m looking for a tool for something specific, I’ll usually ask ChatGPT first, then check Product Hunt or Reddit to see what real people are saying. Google has become more of a second step for me.

What I’m curious about is whether thats actually changing discovery for other makers too, or if I’m just unusually deep into AI tools.

For those building SaaS products: are you seeing people discover you through AI yet? Or is Google still doing most of the work?

322 views

Add a comment

Replies

Best

Google still does most of the work across my products, but the shape of that traffic has changed even where the channel has not. Some sessions now arrive with no keyword attached at all, just a referral spike with a flat unfamiliar path, which is what a click out of a ChatGPT or Claude answer looks like in analytics. Hard to separate cleanly from normal search. The thing worth checking before you optimise for any of this: does the AI answer name your product, or does it just describe the category and list three options. Being the one it names by name is worth chasing. Being folded into "there are tools like X" is not, that traffic barely converts.

 That’s a really good distinction. A mention isn’t necessarily a recommendation, and I think people tend to lump those together. Being named as the answer to a specific need is obviously much more valuable than just appearing somewhere in a list. Have you found a good way to track that consistently across different AI models?

 Not consistently, no. The closest I have is checking the referrer domain on sessions with no query string, since and referrals show up in analytics even without the prompt behind them. That tells you volume, not content.

Serdar's signup field idea below is the better fix. I have started doing something close to it, one free text question at the point someone converts rather than at signup, since by then they can actually name what solved it for them instead of guessing what they typed three steps earlier. A fixed prompt panel felt too noisy to trust on its own for the reasons he laid out. Worth running alongside the free text question rather than instead of it, so you have something real to check the panel against.

 That makes sense. The free-text question at conversion is probably the closest thing to ground truth we have right now, because at least it comes from the actual buyer rather than from a simulated prompt set.

I still think the panel is useful, but more as a directional layer than a measurement system on its own. The interesting part is when both start agreeing, buyers tell you they found you through a certain type of question, and repeated model runs show you consistently appearing there.

That overlap is where I’d start trusting the signal.

Do you think AI will eventually become the main way people discover new products, or will communities like Product Hunt always play a different role?

 I think AI will become the first stop for discovery, especially when people already know what problem they’re trying to solve. Communities like Product Hunt will still matter, but more for validation, serendipity, and seeing what other builders are excited about.

So probably less “AI replaces Product Hunt” and more that they end up owning different parts of the discovery journey.

One thing I’m noticing from the replies here: AI seems to be winning the shortlist, but not necessarily
the final decision yet.

Curious if anyone here is actually tracking this on the maker side, not just AI referral traffic, but whether ChatGPT/Claude are naming your product when people ask for recommendations?

 we track it. 20 buyer questions against a grounded model, aug 19. we came back mentioned in 0 of them. competitors got named in 17.

the useful bit is what said. one run each means the ci on that 0 runs up to 16%, so 20 probes cant tell never mentioned apart from mentioned 1 in 6. we treat it as a floor not a score, and only a move across a whole cluster counts as anything.

   That’s actually really interesting. And yeah, treating it as a floor rather than a score makes a lot more sense.

0 vs 17 is kinda brutal though 😅
When you see a gap like that, what do you actually try to change first? Your own site/content, or getting mentioned in more external sources?

   My own pages first, mentions elsewhere last. But before either of those I check whether the page can be fetched at all. I build and sell prebuilt affiliate sites in a single category, watches, so I have a catalog to test that on, and there I caught paths containing a dot falling through to a local language handler and returning 500. The crawler was let in and then handed an error. That was a routing rule rather than a content problem, and since every site I run comes off one shared build, it closes in one place instead of page by page. I cannot show that a fix like that moves a mention count. What I can see is Search Console telling me 236 of my 294 indexable pages have never taken an impression, with another 230 sitting in Discovered, currently not indexed. That is the constraint in front of me, so it goes before rewriting and before chasing mentions. On the floor and not a score reading Nivy passed along, that is still how I hold it, one run has a wide band, and what I would trust is movement across a set rather than a single number.

Answering from the receiving end, because that half is easier to measure than my own behaviour and harder to fool myself about.

On my site, ChatGPT is now the single largest referrer. Ahead of Google. That flipped inside the last year and I did nothing deliberate to cause it, which is the part I keep chewing on.

What changed with it is which page wins. Assistants do not weigh domain age or backlinks the way search does. They lift a specific claim off a specific page and hand it over with attribution. So the page that gets cited is not the one with the strongest link profile or the most complete coverage of a topic. It is the one that answers exactly one question and has a real figure in it, because a figure is the part that can be quoted. General, well-optimised category pages are close to useless there, which inverts a fair amount of standard advice.

That also makes it much friendlier to new sites than search is. A four month old domain can be quoted today. It cannot rank today.

The honest gap, and I have not seen anyone solve it: I can see that the traffic arrives, and I cannot see what was asked. Search Console gives me queries. Assistant referrals arrive with no query attached at all. So I know the channel works and I cannot steer it. All I can do is notice which pages get pulled and infer backwards, which is guessing with extra steps.

On your split between discovery and buying, I would put it slightly differently. The shortlist moved. The verification did not, and that is why you still go to Product Hunt and Reddit afterwards. What that means for makers is that the assistant decides whether you are considered at all, and humans still decide whether you are chosen. Losing the first stage is quieter than losing the second, because nobody bounces off your site. They simply never hear your name.

 “The shortlist moved” is a great way to frame it. And I think the missing-query problem is probably the biggest measurement gap right now.

Referral traffic only tells you when AI already sent someone your way. It doesn’t tell you which prompts you appeared in, why you were chosen, or more importantly, all the prompts where a competitor was recommended and you weren’t.

That’s why I think tracking actual model answers across a defined set of buyer intents is becoming the closest thing we have to Search Console for AI discovery.

And your last point is key: not making the shortlist is an invisible loss. There’s no impression, no bounce, nothing to diagnose.

  Agreed that it is the biggest measurement gap, and I want to push on the prompt-panel idea rather than dismiss it, because I think it is the right instinct with one property that has to be held carefully.

A panel of prompts is a simulation, not a measurement, and the difference is not pedantic. Search Console reports what actually happened to real people. Asking a model the same fifty buyer-intent questions reports what it said to you, on that day, in your session, with whatever personalisation and sampling is in play. Run it twice and you can get two different shortlists. So it is closer to a poll of one respondent with a shaky memory than to a log.

That still beats nothing, and I would run it. But it changes what you are allowed to conclude. Movement of one position on one prompt is noise. Disappearing from twelve of fifty prompts for three consecutive runs is signal. If the tool does not force that discipline, people will read the daily wobble as performance and optimise against randomness, which is worse than not looking.

The other limit is the one I cannot get around. The panel's validity depends entirely on whether your fifty prompts resemble what buyers actually type, and the referral gives you no way to check that. So you are grading yourself against a set you guessed. Two panels built by two people for the same product would not agree, and neither would be wrong.

The only genuine signal I have found for it is embarrassingly low tech: ask at signup what they searched or asked before they found you, in one free text field, and read the answers rather than counting them. It converts the invisible loss into a sampled one. Small sample, real language, and it is the only input I have that comes from a buyer rather than from a model impersonating one.

Which is maybe the honest version of where we are: the artifact is measurable and the intent is not, so we are all measuring the artifact and calling it intent.

 That’s fair. I think the distinction is really between synthetic measurement and observed behavior.

I wouldn’t treat a 50-prompt panel like Search Console data point by data point. The useful signal is the distribution over time: repeated runs, multiple models, and clusters of related buyer intents. One position moving is noise. Disappearing across a whole intent cluster for several runs is much harder to dismiss.

But you’re right about the prompt set itself baking in our assumptions. I actually like the signup question as the reality check. Combining what models appear to recommend with the actual language buyers say they used to find you may be the closest thing we have right now.

So maybe “Search Console for AI” is too strong. It’s closer to brand tracking for a channel that doesn’t have proper analytics yet.

 Brand monitoring is the closer fit, and I will take yours over mine. The awkward part on our side is that a visit which starts in a chat tells you it happened and almost nothing about the sentence that produced it. Chat passed Google as our top source this summer, so the channel we know the least about is also the one sending the most people. Search Console at least hands you the query. Here the wording stays inside the conversation, which means the measurement has to sit there too. Asking someone what they typed is not a weaker substitute for a log in this case, it is the only instrument that reaches the wording, and we treat it as the real one rather than as a gap we are waiting to close.

the bigger shift i see is where discovery ends up, not where it starts. i build a consumer app that reads what people save from social, and the pattern is consistent: discovery got faster, retrieval didn't. someone asks chatgpt, gets a great answer, screenshots it, and that screenshot joins a camera roll they never open again. most of what people save is never revisited unless something actively brings it back. so yes, chatgpt changed step one. what it really exposed is that everything after the tab closes is still broken.

 That’s a really interesting way to frame it. Discovery is getting dramatically better, but retention and retrieval still feel almost unchanged.

We’re generating more useful answers than ever, yet a lot of them end up as screenshots, bookmarks, or tabs that effectively disappear.

I think the next big UX problem isn’t helping people find more information. It’s helping the right thing resurface at the moment it becomes useful again. That’s probably where the real value moves once discovery itself becomes cheap.

 exactly, and "resurface at the moment it becomes useful again" is the harder half because it has to know what you meant, not just what you saved. the pattern i keep seeing in our data: the gap between saving something and needing it again is usually hours to days, not minutes. so resurfacing can't be a feed, it has to be a search you didn't know you were going to run. discovery got a great new front door. the back half, retrieval with intent, is still a folder nobody opens.

 Yeah exactly. Saving something is easy, but knowing when I’ll actually need it again is the whole problem. Half the time I only remember I saved it after I’ve already spent 20 minutes looking for it again 😅 Curious how you’re tackling that timing in the product.

Definitely changed for me - I now use ChatGPT or Claude first when I know the problem but not the product and I have been seeing the same behaviour around me with developers, marketers and founders. Google usually comes later when I want to dig into pricing, reviews or verify something. What makes this more interesting from the maker side is that the loss is almost invisible - if ChatGPT recommends three competitors and leaves you out, there is no impression or bounce to look at and that is actually a big part of why we started building GeoRankers. We wanted to understand not just AI referral traffic but whether a brand is actually making the shortlist across real buyer questions and why.

 That makes sense. The shortlist is probably the most useful unit to measure right now, because referral traffic only shows what happened after the model already chose you.

What I think gets tricky is the “and why” part. A brand can appear for the same buyer question across different runs for completely different reasons: source freshness, wording, citations, model behavior, even how the category is framed.

So the interesting layer for me is not just visibility, but trying to separate repeatable reasons for inclusion from model noise. That’s where this starts becoming genuinely actionable rather than just another ranking dashboard.

the detail that made me take this seriously is that product hunt publishes an llms.txt. its sitting in the footer of every page here. platforms are already formatting themselves to be read by models, and most of us havent changed anything on our own side. on the named versus category point raised, what seems to decide it is whether other people have described your product in their own words somewhere retrievable. our own landing page is one voice repeating one sentence. the places we actually get quoted from are threads where someone else explained what the tool was for, usually with a phrase we never wrote down ourselves.

   That last part is probably the most interesting piece of this.

It makes AI discovery feel less like traditional SEO and more like building enough independent context around a product for models to triangulate what it actually is. llms.txt can make the site easier to read, but third-party descriptions might be what gives the model confidence to recommend it.

Have you been able to trace any actual AI referrals or citations back to those threads?

   not cleanly, no. i can see and as referrer domains same as before, but that only tells you a session came from an assistant, not which thread it actually pulled the description from. the closest thing i've got is when someone's free text answer at signup uses a phrase that matches a specific comment almost word for word, which is circumstantial but its the only link i've found between one particular thread and an actual conversion. rabnoor's llms.txt point makes me want to check whether products that publish one get quoted more often than ones that dont, has anyone actually run that comparison

    the signup wording is a pretty interesting clue, even if it’s messy to prove. Makes me wonder what people would write if we asked “how was the product described to you?” instead of just “where did you find us?”

And fair point on llms.txt. I probably gave it too much credit in my earlier reply. I’d want to see the same site before and after adding it, since comparing different products would be hard to untangle.

   Worth trying, though I would keep it open text rather than that exact phrasing. People describe things in their own words when you do not hand them a frame to fill in. I already switched mine to one free text field at the point someone converts rather than at signup, and the answers read differently, less "found you on Product Hunt" and more the sentence they would actually use to explain the product to a colleague. That is closer to what a model would have said about you than anything collected at signup, when they have not used the thing yet. Have not compared sites with and without llms.txt though, that needs two products at similar traffic running side by side and I do not have a clean pair for it.

I do the SEO on our site, and the honest answer is I cannot tell. Someone asks ChatGPT, gets a name, then types that name into Google, and it lands in my reports as branded search. We went live with something new today, and that blind spot is the part that bothers me most.

 That’s exactly the blind spot I keep coming back to. ChatGPT may have done the actual discovery, but the moment someone Googles the brand to verify it, all we see is “branded search” and Google gets the credit. It probably means AI is influencing far more discovery than the reports suggest.

What did you launch today? I’d love to take a look.