Skip to main content

That AI Visibility Percentage Your Vendor Sold You May Be Noise

Profound and Evertune are publicly arguing over how to measure whether your brand shows up in ChatGPT or Claude. The fight reveals something worse: most AI-visibility numbers don't say how they were sampled.

That AI Visibility Percentage Your Vendor Sold You May Be Noise

Ricardo Argüello

Ricardo Argüello
Ricardo Argüello

CEO & Founder

AI in Marketing 4 min read

Profound and Evertune, two of the most-cited vendors for measuring whether a brand shows up in ChatGPT, Claude, or Perplexity answers, are having a public methodology disagreement. It’s not a lawsuit or a scandal, it’s more useful than that: it’s a technical dispute over how to measure something most marketing teams are already buying as if it were a solid number.

The disagreement, reported by The State of Brand on July 20, has a far more concrete and verifiable origin on Profound’s side.

Profound’s own data reveals the problem

On July 8, Profound economist Jennifer Zou published a direct analysis on the company blog titled “Is once a day enough?” They ran 753 prompts across 7 AI platforms, comparing once-daily sampling (129,000 measurements) against ten-times-daily sampling (860,000 measurements) over two weeks in June. The result: measured visibility shifted by about 2 percentage points depending on sampling frequency, and citation share shifted by about 0.3 points.

That’s not a competitor’s accusation. It’s the company itself acknowledging, with its own numbers, that sampling frequency moves the result. According to Evertune CEO Brian Stempeck, the gap is worse still: with a single daily measurement, the margin of error can reach 11-12 points, dropping to about 6 points only after roughly 100 repeated measurements.

The two methodologies are different by design, not by oversight. Profound runs a broad portfolio, roughly 753 prompts across 7 platforms, about once a day each, combined with real-user front-end data capture in production. Evertune uses a smaller, curated query set but repeats each one about 100 times, paired with what it calls “EverPanel,” a proprietary panel the company describes as having roughly 25 million participants. Neither methodology is “the correct one” in the abstract. Each makes a different tradeoff between broad coverage and deep repetition, and that tradeoff is exactly what determines the margin of error each one reports.

This isn’t a two-company problem, it’s a methodology problem

Here’s what gives this disagreement real weight: it isn’t exclusive to Profound or Evertune. Rand Fishkin at SparkToro found that asking an AI model the same question twice produces the same list of recommended brands less than 1% of the time, and the same ranking order less than 0.1% of the time. That doesn’t invalidate measuring aggregated visibility across many repeated queries, but it does invalidate any claim based on a single query or a handful of them.

A University of St. Gallen study and Ron Sielinski’s “IQRush” paper, covered by Search Engine Journal, arrived at the same place through a different route: 33 to 94 repetitions of the same query are needed for the result to stabilize statistically, and in 3 of 30 tests the result never stabilized, no matter how many times it was repeated. That study doesn’t even mention Profound or Evertune. It confirms sampling noise is a property of how language models work, not a defect specific to one vendor.

What this means for any visibility report you’re buying

Translated into a purchasing decision: if an AEO vendor hands you a number like “your brand appeared in 40% of responses this month” without telling you how many times they asked, how often, and across how many platforms, that number doesn’t tell you whether your brand actually improved or the model’s noise simply landed on a different side that month. The question that should be in every AI-visibility measurement contract isn’t “what’s my percentage?” It’s “how many times did you repeat the query to get that percentage?”

This connects to something we already wrote about the difference between being seen and being chosen in AI responses: aggregated visibility without method context can look like a strong signal and actually be nothing more than the normal variation of a system that doesn’t answer the same way twice.

How we handle this in our own monitoring

At IQ Source we run recurring AI-visibility monitoring for clients as part of AI Maestro work, and we apply the same discipline this disagreement just exposed: we report how often measurement happened, across how many platforms, and we explicitly separate a real trend from a fluctuation inside the expected noise margin. We’ve already written about why your AI marketing needs a verifier, and this Profound-versus-Evertune disagreement is exactly the same principle applied to how you measure whether your brand shows up in AI answers, not just how you generate the content.

Ask about the method before trusting an AI-visibility number

Frequently Asked Questions

AEO AI visibility Profound Evertune brand measurement answer engine optimization AI Maestro

Related Articles

The AI Quality Check Your Marketing Team Doesn't Have
AI in Marketing
· 7 min read

The AI Quality Check Your Marketing Team Doesn't Have

Legawrite.AI built a panel of AI judges for litigation, where disagreement is the signal. Marketing generates AI content at scale with no equivalent check.

AI judge panel LLM-as-judge AI marketing content
21 Free Skills Turn Claude Into an MBB-Style Strategist
AI in Marketing
· 4 min read

21 Free Skills Turn Claude Into an MBB-Style Strategist

Oria published 21 free Claude skills for market mapping, competitive intel, and pricing. What it means for a marketing team with no in-house strategist.

Claude skills Oria marketing strategy