I’ve previously balanced manual prompt testing against enterprise cloud data platforms and legacy SEO tracking software like Semrush or Ahrefs. Manual testing (manually asking various web assistants the same string of questions) is an exhausting, unscalable process heavily corrupted by localized browser cache and personalization bias. On the other hand, classic SEO tools are fundamentally blind to generative spaces; they measure keywords and page-rank layers but fail to capture how an unindexed LLM weights your entity authority inside a chat response.
Enterprise-grade AEO data setups offer deeper multi-model validation pipelines, but they frequently require complex custom scripts or lock their dashboards behind steep corporate sales cycles and pre-paywalls. AI Visibility carves out a highly practical, middle-ground alternative—functioning as a fast, lightweight diagnostic bridge that exposes your brand's AI search gaps without the subscription friction or architectural bloat.
AEO (Answer Engine Optimization): The technical practice of structuring online content, schema data, and entity signals so AI platforms can easily read, synthesize, and cite your brand as a primary reference source.
@baliy Hi there, we started checking this by hand a few months back and what surprised me was how differently each model described us, and how fast it drifted after one blog post got indexed. Do you track the delta over time so you can tie a shift in what the models say back to something specific you published? The one-off snapshot is interesting, but the trend is what would actually change how we write.
@artem_fedorovich Really appreciate this, that drift is exactly what got me building it. Each run is a snapshot, but your checks are saved and you can refresh a brand once a month to watch the score, mentions and cited sources move. Monthly is about the window where something you publish starts showing up in the answers. It's manual for now (no auto-refresh yet, maybe later), and it won't tie a shift to a specific post automatically, but running a check right after you publish lets you line the two up. That trend loop is where I want to take it, so this is really useful. What would make it actually change how you write?
We’ve been using AI Visibility for our SaaS, and it’s one of the few GEO/AI visibility tools that provides insights we can actually act on. We especially like seeing which prompts we’re already visible for, where competitors are ahead, and how our visibility changes over time. The reports have already inspired new content ideas, including comparison pages, integration guides, and content focused on real customer questions.
A couple of features we’d love to see: email notifications when a scheduled report is ready or when visibility changes significantly, plus a “content opportunities” section suggesting blog topics based on prompts where competitors are cited but we aren’t. A directory of potential guest post partners or industry blogs would also be a great addition.
@spiri7 Thanks, this really means a lot. Glad it's already sparking content ideas.
On the features:
Email alerts when a report's ready or your visibility moves a lot - yeah, I want that too.
"Content opportunities" from prompts where competitors show up and you don't - the gap view kind of does this already, and turning it into actual topic suggestions is an easy next step.
A list of guest-post blogs - bit further out, but I like the idea. Which ones matter in your space?
Really helpful, thanks for taking the time to write all this.
Interesting direction. Have you noticed cases where an AI model confidently gives outdated or incorrect brand information, and if so, how do you distinguish hallucinations from missing public data?
@amjad_shaik Yeah, all the time. Models will state old pricing, an old tagline, even the wrong founder, with full confidence.
Rough way I split it: if a model names you but cites no real source, that's usually its memory talking, and where hallucinations creep in. If it cites a page but the page is stale, that's outdated public data. And if it never mentions you at all, that's just a gap, not a wrong answer.
So the tool shows what each model says next to the sources it leans on, and you judge from there. It won't certify what's true, but seeing the claim beside its source usually makes it clear which of the three you've got. Honestly the "confidently wrong" case is the scariest, since publishing more doesn't fix it, you have to go correct the sources it trusts.
@baliy That’s a useful way to think about it. The confidently incorrect answers are exactly the ones that can hurt user trust the most. Have you come across cases where different AI models gave completely different answers about the same brand even though they were looking at similar public information?
@amjad_shaik Constantly. Though "similar public information" is doing a lot of work in that question, because in practice they almost never look at the same slice of it. One leans on a Reddit thread, another on a listicle, another mostly on the brand's own site, and you get three different characterisations of the same company out of the same open web.
The clearest version I see: the same product described as an enterprise tool by one model and a cheap option for solo users by another, purely because one picked up the pricing page and the other picked up an old forum thread. Neither is really hallucinating, they just read different sources and weighted them differently.
It's worst for smaller brands, where there's no strong consensus out there to anchor on, so each model fills the gaps with whatever it happened to find. That's actually a useful signal in itself: if the three disagree wildly about you, it usually means the public story about you is thin or scattered, not that the models are broken.
the "assistants change their answers between runs" honesty is the right instinct, but it raises the obvious question: how many samples per assistant go into one score? a single ChatGPT call is basically one noisy draw, and if the tool only fires once per assistant per check, two people running it an hour apart on the same brand could get meaningfully different scores and think something changed when it's just sampling variance.
@galdayan Exactly, one call tells you nothing, so we don't score off a single request. We run prompts in clusters and show the average per topic cluster, which smooths out most of the single-draw randomness.
Even then the data still moves around, so where we can we also run from geolocations close to the project's main target audience, since answers shift by region too.
It's still not perfectly precise, especially for newer brands where there isn't much out there for the models to go on. But it gives you a usable ballpark, and re-scanning once a month is where the real signal is, you watch the trend instead of chasing one number.
@baliy clusters plus geo-matching makes sense as the fix. re-scanning monthly and watching the trend rather than the absolute number is probably the right mental model for anyone using this - a single scan looks precise but isn't really something you should anchor on
@galdayan Exactly that. And the "looks precise" part is the honest design problem with any score: the moment you print a number, people anchor on it, no matter how many caveats you add. That's on me to fix in the UI rather than in a disclaimer, probably by showing a range or a confidence band instead of one clean figure. Until then, trend-first is exactly the right mental model.
@baliy weather apps solved basically this exact problem years ago - "70% chance of rain" reads as uncertain by design, nobody expects it to be exactly right. a range or band framed the same way, as the actual unit of measurement rather than an error bar bolted onto a number, might land better than a number-with-caveats ever will
Really like that you're measuring AI visibility as a competitive landscape instead of just another SEO score.
One thing I'm curious about: as LLMs update constantly, how do you distinguish between a temporary ranking fluctuation and a genuine shift in a brand's AI visibility? That seems like the difference between founders reacting to noise versus making better strategic decisions.
Congrats on the launch! 🚀
@aryan787544 Thanks, that competitive-landscape angle is exactly the framing I care about, glad it lands.
On noise vs a real shift: honestly you can't tell from one check. A single run wobbles within a band, so if something moves on one prompt one time, I treat it as noise. What I trust is a move that shows up across the whole cluster of prompts, holds on the next scan, and ideally shows on more than one assistant at once. That pattern is usually real, and it almost always has a cause you can point to, a page that got indexed, a competitor's push, or a model update.
So the rule I'd give a founder is simple: don't react to any single number, react to a direction that repeats. That's the whole reason the tool leans on monthly re-scans and trends instead of one snapshot, and it's the part I most want to sharpen, only surfacing moves big and sustained enough to actually mean something, not every fluctuation.
@baliy That's actually the part I keep thinking about.
I think that principle has a second-order effect on how founders eventually use a product like this. If you get that right, I suspect you end up changing decision-making habits, not just giving people better visibility data.
@aryan787544 That's a sharper way to put it than I've managed. And there's a real tension in it: most tools quietly train the opposite habit, a dashboard you check daily, a number that twitches, an alert for every wiggle, because that's what engagement metrics reward. Doing this properly means deliberately making the product less twitchy, fewer signals rather than more, and saying "nothing meaningful changed, go publish something" more often than not.
Which is a strange thing to optimise for, but I think you're right that it's the actual product. If someone ends up checking less and acting better, that's the win. Better data is table stakes, the habit is the hard part.
@baliy I'm really glad the discussion was useful.
I'd genuinely be interested in following how your thinking around this evolves over the next few months. If you're open to it, what's the best email to reach you on?
the "who they name instead" part of this is what actually matters to me, way more than just "does it mention us." I've checked ChatGPT manually for my own stuff before and the bigger surprise was seeing which competitor kept getting recommended over me and having zero idea why. does the audit try to explain why a competitor wins a mention, or just flag that they do?
@omri_ben_shoham1 Yeah, that "who wins instead, and why" gap is the whole reason I built past just a yes/no mention.
It does more than flag them. For each answer you see the sources the model leaned on, so when a competitor keeps getting named you can usually see why: they're sitting on the pages the model trusts, a Reddit thread, a listicle, a review roundup, or a page structured cleanly enough for the model to quote. Nine times out of ten "why do they keep winning" comes down to being present on those trusted sources and being easy to parse, and the citations make that visible.
What it won't do is read the model's mind and hand you one definitive reason, no tool honestly can. But seeing who's cited instead of you, on which prompts, from which domains, gets you most of the way to "right, that's the thread I'm not in, the roundup I'm missing from", which is usually the real lever.
Congrats on the launch. I run several models side by side in my own product and they disagree with each other far more than people expect, so I can believe brand answers vary a lot between ChatGPT, Gemini and Claude. How often do the three actually contradict each other about the same company? And is the check point in time, or do you track how the answers drift week to week?
@henry_s_jung Thanks. And yeah, you already know the punchline then, they disagree way more than people expect.
On how often: for the plain "does it mention us" they're fairly consistent, but the moment you look at who gets named alongside you, the three often return pretty different sets. For big established brands they mostly line up; for smaller or newer ones the overlap can be surprisingly small, sometimes only one name in common across all three. That's exactly why I show them per-assistant instead of blending into one number, an average would hide that you might be strong on Gemini and invisible on ChatGPT.
On timing: each check is point-in-time, a snapshot. Your checks are saved though, so you re-scan (monthly works well) and watch the answers drift. It's a manual re-scan today, not automatic weekly tracking, that's the direction, but I didn't want to fake continuous monitoring before it's solid.