AI visibility tools have a trust problem, and it isn't a secret

The category's own users say the same thing, unprompted, across five separate threads: nobody trusts the numbers.

RUNNING LABAn experiment actively being exercised and measured. Results may change.

The complaint the market keeps repeating

We didn't go looking for a trust problem. It's the top of the pull. The highest-engagement thread we found, titled roughly “How are people actually handling AI-search visibility? Not sure these tools are worth it,” has the OP stating it plainly: the honest problem is not trusting the numbers, and not knowing what to do with them even when you do (r/SEO, 11 points, 20 comments).

A second, unrelated thread names the sharper version of the same failure: a brand gets mentioned in an AI answer, but the surrounding context is wrong (r/TechSEO). That is not a scoring complaint. That is a factual-accuracy complaint about a product whose entire job is reporting facts about a business.

Independent reviews back up the anecdote. Semrush's AI Visibility Toolkit has been called probabilistic guesswork for brands with a thin AI presence. Independent testing of Ahrefs' Brand Radar found gaps between what it reported as citations and what manual inspection of the same AI answers actually showed.

This is not a gap in the market's tooling. It's the gap this Observatory entry's own “next” step already named.

Zero of nineteen disclosed how the number was made

The sharpest single data point in the pull is a GitHub audit of 19 Chinese-market GEO (generative-engine-optimization) vendors, checked against one question: does the vendor disclose how its score is made? Zero of 19 disclosed per-prompt sample count, noise floor, confidence interval, or seat pricing.

elmo's own explainer names the three levers that make two vendors disagree on the same brand's score: which prompts get asked, how each prompt is weighted, and how a raw AI answer becomes a counted “mention.” Every one of those levers is a private methodology choice, invisible to the buyer, and none of the tools we surveyed publish theirs.

A score with no visible methodology is a guess with a decimal point.

What we found pointing a real audit at a real client

We ran our own check the same way the Reddit thread described the failure: not by reading the rollup number, but by reading what the AI model actually wrote. The client is a real, live home-services business in Lakewood, Ohio, and the prompts were the ones we already track for them.

Most of what the model named checked out — real businesses, sometimes under a different public name or handle than we expected, which we confirmed by hand. One competitor named in the answer did not. It does not resolve to any verifiable, currently operating business in that service area that we could find. We are not publishing the name; the point is not to call out one business, it's that an AI answer stated a fact-shaped claim with no visible uncertainty, and the fact underneath it did not hold up to five minutes of checking.

That is the exact shape of the second Reddit complaint above, reproduced on our own audit, on a real client, this month. We did not have to go looking for a caveat to publish — the caveat found us.

The market's complaint about AI answers stating wrong things with total confidence is not theoretical. It showed up in our own audit the first time we read the raw text instead of the score.

Where the open lane actually is

Vendor tiering in this category is well established: Profound owns enterprise, Peec AI targets agencies, Otterly.ai is the cheap entry point, Scrunch AI differentiates on how AI agents read a site rather than how chat answers cite it, Rankscale claims the most AI responses per dollar for small businesses. None of the five claim a truth-check layer.

The open-source side tells the same story from a different angle. Well-starred tools exist for the observation layer — elmo, Auriti-Labs' geo-optimizer-skill, aigclink's geolook — with full diagnosis-to-strategy-to-verification workflows. elmo's own positioning line, “every metric auditable in code,” is the identical complaint the Reddit threads open with, solved the identical way we're solving it: transparency over a better-sounding score. None of the surveyed tools, open-source or commercial, check a claim against a canonical source of facts about the business being described.

The truth-check layer is not a crowded lane. Nobody surveyed — paid or open-source — claims it.

What this changes about what we're building

Two things follow directly from this month's research, and both are already in motion rather than proposed. First: prioritize the truth-check loop — Business Truth checked against what an AI model actually said, verdicted SUPPORTED, CONFLICT or UNKNOWN, with UNKNOWN treated as a real answer rather than a failure to hide — as the thing we lead with, over a raw mention count. The market asked for exactly this, unprompted, five separate times.

Second: the Lakewood finding is not a one-off to note and move past. An AI answer naming a competitor we can't verify as real is a distinct, checkable failure mode. The Opportunity kind `ai_named_unverifiable` now flags it — quote, execution id, and recorded verification attempts — with "could not verify" treated as a real answer, not proof the name is fake.

Publish the methodology instead of a better score. Flag what can't be verified instead of repeating it.
01 · Question
What do AI-visibility buyers and builders actually say is broken about the tools they already pay for — and does it match what we're building?
02 · Why it matters
We would rather take a build direction from the market's own complaints than from vendor pitches. So this month we ran a structured pull across Reddit, Hacker News, GitHub and the web, then pointed our own audit at a real client and read the raw AI answers ourselves instead of trusting the rollup number. One of the two findings here is uncomfortable and we are publishing it anyway.
03 · State
RUNNING LAB An experiment actively being exercised and measured. Results may change.
04 · What we built or tested
  • A structured research pull across r/SEO, r/TechSEO, r/Entrepreneur, r/aeo, r/GenerativeSEOstrategy, r/ParseAI, Hacker News, GitHub and the open web, dated 2026-09-25.
  • Direct chat_gpt and claude observation calls landed in the Observatory this week — the exact next step the earlier AI Visibility Observatory entry on this page promised, now shipped and running against real client prompts.
  • A read of our own audit output against a real, live client's tracked prompts — not just the aggregate score, the actual sentences an AI model wrote about them.
  • Opportunity kind `ai_named_unverifiable`: extract untracked names from AI recommendation lists, require recorded verification attempts, emit only when verdict is unverifiable (not pending, not "proven fake").
  • lehvel-observatory entity verification merged (PR #8): extract + listings/known-competitor verify ladder; empty listings stamp UNKNOWN “could not verify” and surface as the same `ai_named_unverifiable` label in Business Truth.
  • Business Truth connections spine (observatory PR #13, merged): marketer checklist for GSC + Analytics + Google Business Profile; public GBP listing imports NAP/hours/place id into facts for local and service SMBs.
  • `ai_named_unverifiable` fire rate: completed verification attempts only (pending excluded); formula shipped on Observatory runs and the sales desk overview.
05 · Evidence
Methodology disclosure, 19 GEO vendors audited (GitHub)0 of 19 disclosed sample size, noise floor, confidence interval or seat pricing
Market complaint, five separate threads, Sept 2026“the honest problem is I don't fully trust the numbers, and even when I do, I'm not sure what to actually do with them” (r/SEO) — “Our brand gets mentioned in AI answers, but the context is completely wrong” (r/TechSEO)
Competitive set surveyedProfound, Peec AI, Otterly.ai, Scrunch AI, Rankscale, AthenaHQ, Semrush AI Toolkit, Ahrefs Brand Radar — none check whether a cited answer is actually true
Our own audit, a real client, Lakewood, Ohio, Sept 2026One AI-named competitor in the answer could not be independently verified as a real, currently operating business in the service area
06 · What we learned
  • The loudest complaint in the category is not “the score is too low.” It is “I can't tell if the number means anything,” stated the same way by unrelated people who have never talked to each other.
  • That complaint is not abstract. Pointing a real audit at a real client surfaced the same failure the Reddit threads described — an AI answer stating something with total confidence that did not hold up to a direct check.
  • Every well-starred open-source tool we found does the observation layer well. None of them check a claim against a canonical source of facts about the business. That gap is not crowded.
07 · Limitations
  • The vendor comparisons here are secondhand — independent reviews and vendor-published claims, not our own audit of competitors' tools.
  • The unverifiable-competitor finding is one flagged name on one client's audit, not a study. It is evidence a failure mode exists, not a rate at which it occurs across a client set — the fire-rate formula is shipped; multi-client measured rate still needs observations with `service_location_coordinate`.
  • Reddit and GitHub engagement is a signal, not a market-size study. The $1,250/month price point cited below is one self-reported data point.
08 · What changes next
Run observations on local/service clients with `service_location_coordinate` (GBP public import or hand-set); publish measured `ai_named_unverifiable` fire rate on Labs; ship owner-authorized Google Business Profile OAuth as the next first-party Business Truth connection.
Negative results are published as results. A state on this page is a claim about what is true today, not about what is planned; a limitation is a limitation.
These tools belong to the people who made them. Lehvel reviewed and, where noted, tested them. Nothing here is a Lehvel product, and no performance figure attributed to a vendor has been reproduced by us.