The Five Checks I Run Before I Believe Any AI Search Stat

Bar chart: 81 percent versus 36 percent

Three vendor stats landed inside the same week, and every one of them flatters the company that published it. Here are the five checks I run before a number like that goes anywhere near something I write.

Three vendor stats, one week
81%
Organizations integrating SEO and AI visibility work that saw AI-driven traffic gains, versus 36% managing them separately. Semrush, June 26, 2026
92.4%
ChatGPT’s share of standalone AI referral traffic, up from 84% in December. Previsible, July 6, 2026
8 hours
Average time to first AI citation across 8,000 press releases. Notified, June 2026
Tested against my own audit dataset: 446 evaluations across 192 sites, as of July 27, 2026.

Every number sells something.

Not always the number itself. Sometimes it’s the company that ran the study. Sometimes it’s the tool that conveniently explains the gap the study found. Sometimes it’s just attention for the brand name attached to the release. That’s not an accusation. It’s just where these numbers come from, and pretending otherwise is how bad stats end up in decks.

This week handed me three of them at once. Semrush’s AI Visibility Index says integrating SEO and AI visibility work more than doubles your odds of seeing traffic gains. Previsible says ChatGPT now owns 92.4% of standalone AI referral traffic, up from 84% in December. Notified and PRWeek say a press release hits its first AI citation within 8 hours on average. All three are plausible. None of them get to walk into my next post unchecked. Here’s the process.

Check one: who sells the thing the stat proves works

Semrush is owned by Adobe, and Semrush sells the exact kind of integrated SEO-plus-AI-visibility tooling its own stat says works better. That doesn’t make the finding false. It means I read it the way I’d read a mattress company’s sleep study: probably measuring something real, definitely angled toward a conclusion that helps the vendor. I still cite it below. I just say who’s selling what, in the same sentence, every time.

Check two: is the sample specific enough to check

Previsible names its sample: 6.77 million sessions, 166 GA4 properties, 19 months. That’s a number I could theoretically go verify against public GA4 behavior. Compare that to a stat I refuse to cite anywhere, the one claiming the GEO market grew from $848 million to $33.7 billion in some short window. Nobody can point to where that figure originated. It gets repeated in enough slide decks that people assume someone, somewhere, checked it. Nobody did. If a source can’t tell me the sample, the number doesn’t get to appear in anything with my name on it.

Check three: what does the source admit it doesn’t know

This is the check that took my own data down a notch, and it should. Of my 446 evaluations as of July 27, 2026, only 75 (17%) are high-reliability audits, meaning the institution’s site let my audit bot through for a direct crawl. Another 99 (22%) are low-reliability, meaning the site blocked the bot and the score came from a workaround estimate instead. The remaining 272 haven’t been sorted into either bucket yet. I publish that breakdown every time, because quietly averaging all 446 into one clean number would be exactly the move I’m criticizing vendors for above. None of this week’s three studies published an equivalent confidence split for their own numbers. Worth noticing. Not automatically disqualifying, just worth noticing.

Check four: would it survive a repeat run

Last week I wrote about a Search Engine Land finding that ChatGPT swaps its primary citation source on 11.6% of repeated identical prompts. If a number depends on which day, which prompt set, or which four-month window you happened to sample, it’s fragile by definition. Semrush’s window is January through April 2026. Run the same 126-million-prompt analysis on May through August and I’d bet the number moves some. Not because Semrush did anything wrong. Because the system underneath these numbers doesn’t hold still long enough for one measurement to be the last word on it.

Check five: does it match what I actually see

Here’s where the Semrush integration stat earns some trust back. Across my own audits, the average overall score sits at 60 out of 100, and the weakest category by a wide margin is accessibility at 51. The sites that score well across categories, not just one, tend to be the ones treating structure as a coordinated project instead of a single fix somebody knocked out in an afternoon. That’s not a controlled experiment and I won’t dress it up as one. But it’s the same shape Semrush describes: piecemeal effort loses to coordinated effort. When an outside number matches a pattern I already see in my own dataset, I trust it more. When it contradicts what I see, I say so in public instead of quietly filing it away.

What this actually buys you

None of this means throw out vendor research. Semrush, Previsible, and Notified all published something real in the past several weeks, and I cited all three above, by name and by date. It means every one of those numbers gets a name, a date, and a sentence about who benefits before it goes anywhere near a claim I’m making. That’s not a lot to ask of a stat before you repeat it to a board or a client. It’s usually the one step people skip, because the number, sitting there on its own, already sounds finished.

I write the applied version of this same discipline for boards and marketing teams over at Atlas Instinct, including this week’s piece on why a credit union’s own website is only one layer of the AI visibility problem, not the whole thing.

Based on the audit platform I built for Atlas Instinct, 446 evaluations across 192 sites as of July 27, 2026; Semrush’s 2026 AI Visibility Index (July 2026); Previsible’s July 2026 AI traffic report; and Notified/PRWeek’s GlobeNewswire citation study (July 22, 2026).

FAQ

Why not just trust a vendor’s AI search stat at face value?
Because the vendor usually sells the thing the stat proves works. That doesn’t make the number false, but it means the stat needs a name, a date, and a disclosed sample before it gets repeated as fact.

What’s the biggest red flag in an AI search study?
An impressive-sounding number with no disclosed sample or method behind it, the kind that gets repeated in enough decks that everyone assumes someone already verified it. Often nobody did.

How does this apply to your own audit platform’s data?
The same way it applies to anyone else’s numbers. Of my 446 evaluations as of July 27, 2026, only 17% are high-reliability direct crawls. Another 22% are low-reliability estimates. I publish that split every time instead of averaging it away.

What’s the one check most people skip?
Asking whether the number would survive a repeat run. Citation behavior shifts between individual AI queries, so a single measurement window, even a large one, is a snapshot, not a verdict.

Scroll to Top