research_paper
14 Confirmed, 9 Unknown, 3 Wrong: An AI Visibility Audit Graded Against Its Own Record
September 11, 2026 · 8 min read · Scout7
A rigorous framework for grading AI visibility audits by artifact traceability, internal-record validation, and declared unknowns.
A credible AI visibility audit does three things: it shows the public artifact behind each claim, leaves hidden metrics blank when they cannot be known, and proves its method on a subject whose internal record can judge what was true, unknowable, and wrong.
Key takeaways:
- Artifact-backed findings are stronger than polished estimates
- Unknowns should stay blank when proof is unavailable
- A tested error rate matters more than a neat dashboard
- 14 confirmed, 9 unknowable, 3 wrong is a usable grade
Introduction
You receive an AI visibility report with clean charts, exact language, and no visible doubt. It looks finished, which is exactly why it deserves inspection.
The useful question is not whether the dashboard looks sophisticated. It is whether each claim can be reopened and checked by the reader.
Most reports present certainty first and method second.
- The missing number is confidence, not visibility
- A claim without a reopenable artifact is weaker than it appears
- A report becomes believable when it marks where evidence ends
- The practical goal here is to help founders grade what they were sold
That leads to the first pressure point: what the report admits it could not know.
Reports Are Sold Before They State Their Error
And that is usually where the problem starts. Many AI visibility reports are sold as if uncertainty were a formatting flaw rather than part of the work.
A report with no empty cells often signals that estimates have replaced proof. Precision on the page can hide ambiguity in the method.
Teams still need a way to defend what they buy.
- Unknowns matter because hidden metrics often cannot be settled from public evidence
- Method comes before confidence when a result will be questioned internally
- Budget pressure magnifies risk when teams buy reporting before they buy verification discipline
- A complete-looking report can still be unfinished work
If a report does not state its limits, treat every precise claim inside it as provisional. The next question is how any method earns the right to sound certain at all.
The Only Test Is Running the Method on a Subject You Can Check
That right is earned only one way. You run the method on a company whose internal record can tell you which findings were true, false, or genuinely unknowable.
We ran the method on ourselves for that reason. Our own internal record is the only place where each external finding could be truth-tested instead of merely defended.
Based on what we saw when we checked our own public-artifact findings against our internal record, the only serious proof of audit quality is a method willing to publish its unknowns and its misses.
Internal validation and external benchmarking do different jobs.
- Internal validation tests truth for a specific finding
- External benchmarking tests pattern frequency across a market
- Those jobs are not interchangeable and should never be blended
- A validated method publishes misses instead of only showcasing hits
This is the line founders can use when grading a vendor: ask where the method was checked against records the public cannot see. That check produced the only grade that matters here.
The Grade: 14 Confirmed, 9 Unknowable, 3 Wrong
When we forced the method against our own record, the result was clear: 14 confirmed, 9 unknowable, 3 wrong. That is not messy bookkeeping; it is what an auditable result looks like.
The confirmed findings were the ones that traced cleanly to public artifacts a reader could reopen themselves. The unknowables were named and left empty rather than inflated into authority.
The grade breaks down into three classes.
- 14 confirmed from reopenable public evidence
- 9 unknowable because spend, signups, and per-post impressions could not be honestly known
- 3 wrong and retained in the record as tested misses
- The empty cells matter because they show where evidence stopped
A finding is trustworthy exactly to the extent that it traces back to a public artifact the reader can reopen: publish timestamps, UTM strings, sitemap files, public ad libraries, robots.txt, archive captures, and historical page copy.
Once that standard is clear, the next issue is which artifact classes can actually carry a claim.
What Artifact Class Can Actually Carry a Finding
That distinction matters because not all findings stand on the same ground. Some claims rest on durable artifacts; others collapse the moment you ask for proof.
In our own check, the confirmed work came from artifact classes that survive reopening. That included a launch date recovered from a UTM string, four repositionings in five months, a 31-day content silence, a machine-set publishing cadence detected from timestamps alone, paid activity across three platforms with 160 creatives in public ad libraries, and 269 sitemap URLs against roughly seven findable pages.
The stronger artifact classes were consistent.
- Publish timestamps support cadence and silence findings
- UTM strings can expose chronology and launch sequencing
- Sitemap files reveal declared site structure versus findable reality
- Public ad libraries surface campaign presence and creative volume
- Robots.txt, archive captures, and historical page copy preserve changes over time
The same discipline applied across the broader study: 83 companies, 249 buyer-intent queries, and 2,045 first-page results. In that set, 58 of 83 were not named on a single first-page result for any of their three own buyer questions, 67 of 83 were never mentioned by anyone outside their own company, 1,038 of 2,045 cited pages were listicles, roundups, and comparisons, and only 13 of 83 ranked for even one of their own three questions.
Strong artifacts can carry a finding. Hidden metrics usually cannot, which is why the three errors matter so much.
The Three Errors and the Rule Each One Bought
This is where the method becomes readable. The most useful part of the exercise was not the confirmed work but the places where the method broke and had to buy stricter rules.
The three wrong findings stayed in the grade because a serious method improves by turning each miss into a guardrail against overclaiming.
Each wrong conclusion bought one rule.
- Wrong one: no social presence; the volume lived on the founder's personal handles, so inspect personal accounts before declaring absence
- Wrong two: the 8,449 follower count looked implausible; plausibility doubt is not evidence, so verify before dismissing a number
- Wrong three: no analytics of any kind from two pages of HTML; never convert a negative observation into an absolute absence
Those are not side notes. They are operating rules for any AI visibility audit that wants to be more than theater.
The same discipline also explains why a broad market scan can surface patterns without pretending to settle hidden truth for each company.
The Three Questions to Ask About Any AI Visibility Audit
So how do you audit brand visibility in AI search? You grade the report itself before you trust its conclusions.
A founder does not need to master every AI retrieval system to do that. Three questions expose whether the work rests on evidence or on presentation.
Ask these in order.
- What public artifact can I reopen? If none, the claim is weaker than it looks
- What did you mark unknowable? If nothing is unknown, the method is likely overclaiming
- Where has this method been checked against internal truth? Without that, accuracy is asserted, not demonstrated
This also answers the second core question. Are AI visibility reports accurate? Some findings can be accurate, but only when they are artifact-backed, explicit about unknowns, and tested against a record that can expose error.
The standard is simple: show me the artifact, not the estimate. That leaves one final point—what changes once you adopt that standard.
What A Truth Audit Changes
A Truth Audit does not promise a prettier score. It changes what counts as acceptable evidence.
Once you apply this lens, polished certainty starts to look cheap and rigorous incompleteness starts to look useful. A report becomes more valuable when it can show its footing, name its blanks, and state where it has already been proven wrong.
Key takeaways:
- Trust the artifact trail before you trust the dashboard
- Prefer declared unknowns to invented precision
- Ask for a tested error rate before accepting a visibility claim
- Use the 14/9/3 model as a grading standard
The opening image was a neat report with no visible doubt. The better version is less polished and more credible: it can point to the underlying artifact, explain why some cells are empty, and show the reader where the method has already failed and improved.
For founders and operators, that is the practical shift. Run your growth loop on findings that can be reopened, challenged, and checked—not on outputs that only look complete.
Use the three-question test on the next AI visibility report you receive.
Frequently asked questions
What makes an AI visibility audit credible?
A credible AI visibility audit shows the public artifact behind each claim, leaves hidden metrics blank when they cannot be known, and tests its method against an internal record that can verify what was true, unknowable, and wrong. Credibility comes from reopenable evidence and a published error rate, not polished reporting.
Why are some findings marked unknowable?
Some metrics cannot be honestly settled from public evidence alone, including things like spend, signups, and per-post impressions. Those cells should stay empty rather than be filled with estimates that look precise but cannot be proven.
Why keep the three wrong findings in the grade?
The three wrong findings matter because they show the method was actually tested and improved. Each error produced a stricter rule: check personal accounts before declaring no social presence, verify implausible numbers before dismissing them, and never convert a limited observation into an absolute absence.
How should a founder evaluate an AI visibility report?
Use a three-question test: ask what public artifact you can reopen, what the report marked unknowable, and where the method was checked against internal truth. If a report cannot answer those clearly, its certainty is stronger than its evidence.