The complaint the market keeps repeating
We didn't go looking for a trust problem. It's the top of the pull. The highest-engagement thread we found, titled roughly “How are people actually handling AI-search visibility? Not sure these tools are worth it,” has the OP stating it plainly: the honest problem is not trusting the numbers, and not knowing what to do with them even when you do (r/SEO, 11 points, 20 comments).
A second, unrelated thread names the sharper version of the same failure: a brand gets mentioned in an AI answer, but the surrounding context is wrong (r/TechSEO). That is not a scoring complaint. That is a factual-accuracy complaint about a product whose entire job is reporting facts about a business.
Independent reviews back up the anecdote. Semrush's AI Visibility Toolkit has been called probabilistic guesswork for brands with a thin AI presence. Independent testing of Ahrefs' Brand Radar found gaps between what it reported as citations and what manual inspection of the same AI answers actually showed.
This is not a gap in the market's tooling. It's the gap this Observatory entry's own “next” step already named.
Zero of nineteen disclosed how the number was made
The sharpest single data point in the pull is a GitHub audit of 19 Chinese-market GEO (generative-engine-optimization) vendors, checked against one question: does the vendor disclose how its score is made? Zero of 19 disclosed per-prompt sample count, noise floor, confidence interval, or seat pricing.
elmo's own explainer names the three levers that make two vendors disagree on the same brand's score: which prompts get asked, how each prompt is weighted, and how a raw AI answer becomes a counted “mention.” Every one of those levers is a private methodology choice, invisible to the buyer, and none of the tools we surveyed publish theirs.
A score with no visible methodology is a guess with a decimal point.
What we found pointing a real audit at a real client
We ran our own check the same way the Reddit thread described the failure: not by reading the rollup number, but by reading what the AI model actually wrote. The client is a real, live home-services business in Lakewood, Ohio, and the prompts were the ones we already track for them.
Most of what the model named checked out — real businesses, sometimes under a different public name or handle than we expected, which we confirmed by hand. One competitor named in the answer did not. It does not resolve to any verifiable, currently operating business in that service area that we could find. We are not publishing the name; the point is not to call out one business, it's that an AI answer stated a fact-shaped claim with no visible uncertainty, and the fact underneath it did not hold up to five minutes of checking.
That is the exact shape of the second Reddit complaint above, reproduced on our own audit, on a real client, this month. We did not have to go looking for a caveat to publish — the caveat found us.
The market's complaint about AI answers stating wrong things with total confidence is not theoretical. It showed up in our own audit the first time we read the raw text instead of the score.
Where the open lane actually is
Vendor tiering in this category is well established: Profound owns enterprise, Peec AI targets agencies, Otterly.ai is the cheap entry point, Scrunch AI differentiates on how AI agents read a site rather than how chat answers cite it, Rankscale claims the most AI responses per dollar for small businesses. None of the five claim a truth-check layer.
The open-source side tells the same story from a different angle. Well-starred tools exist for the observation layer — elmo, Auriti-Labs' geo-optimizer-skill, aigclink's geolook — with full diagnosis-to-strategy-to-verification workflows. elmo's own positioning line, “every metric auditable in code,” is the identical complaint the Reddit threads open with, solved the identical way we're solving it: transparency over a better-sounding score. None of the surveyed tools, open-source or commercial, check a claim against a canonical source of facts about the business being described.
The truth-check layer is not a crowded lane. Nobody surveyed — paid or open-source — claims it.
What this changes about what we're building
Two things follow directly from this month's research, and both are already in motion rather than proposed. First: prioritize the truth-check loop — Business Truth checked against what an AI model actually said, verdicted SUPPORTED, CONFLICT or UNKNOWN, with UNKNOWN treated as a real answer rather than a failure to hide — as the thing we lead with, over a raw mention count. The market asked for exactly this, unprompted, five separate times.
Second: the Lakewood finding is not a one-off to note and move past. An AI answer naming a competitor we can't verify as real is a distinct, checkable failure mode. The Opportunity kind `ai_named_unverifiable` now flags it — quote, execution id, and recorded verification attempts — with "could not verify" treated as a real answer, not proof the name is fake.
Publish the methodology instead of a better score. Flag what can't be verified instead of repeating it.