That AI Visibility Percentage Your Vendor Sold You May Be Noise
Ricardo Argüello — August 2, 2026
CEO & Founder
General summary
Profound and Evertune, two of the most-cited AI-visibility measurement vendors, are having a public methodology disagreement. Profound's own data shows that sampling once a day versus ten times a day shifts the measured result by about 2 percentage points. Independent research from SparkToro and the University of St. Gallen confirms this is real, and not exclusive to these two vendors: language models don't give the same answer twice, which turns any visibility figure that doesn't disclose its sampling method into noise.
- Profound economist Jennifer Zou published her own analysis on July 8: across 753 prompts on 7 AI platforms, sampling once a day versus ten times a day shifted measured visibility by about 2 percentage points
- A July 20 The State of Brand article reported that Evertune CEO Brian Stempeck challenges Profound's methodology, citing margins of error up to 11-12 points with a single daily measurement
- SparkToro (Rand Fishkin) found that asking ChatGPT the same question twice produces the same brand list less than 1% of the time
- A University of St. Gallen study and Ron Sielinski's IQRush paper found that 33 to 94 repeated queries are needed for statistical stability, and 3 of 30 tests never stabilized at all
- The problem isn't unique to Profound or Evertune, it's a property of sampling AI models that almost no visibility report discloses
Imagine a vendor tells you 'your brand appeared in 40% of ChatGPT responses this month' without telling you how many times they asked, on which days, or with which prompts. It's like reporting the year's average temperature without saying whether they measured once in January or every single day. The number can be real and still be useless for making a decision.
AI-generated summary
Profound and Evertune, two of the most-cited vendors for measuring whether a brand shows up in ChatGPT, Claude, or Perplexity answers, are having a public methodology disagreement. It’s not a lawsuit or a scandal, it’s more useful than that: it’s a technical dispute over how to measure something most marketing teams are already buying as if it were a solid number.
The disagreement, reported by The State of Brand on July 20, has a far more concrete and verifiable origin on Profound’s side.
Profound’s own data reveals the problem
On July 8, Profound economist Jennifer Zou published a direct analysis on the company blog titled “Is once a day enough?” They ran 753 prompts across 7 AI platforms, comparing once-daily sampling (129,000 measurements) against ten-times-daily sampling (860,000 measurements) over two weeks in June. The result: measured visibility shifted by about 2 percentage points depending on sampling frequency, and citation share shifted by about 0.3 points.
That’s not a competitor’s accusation. It’s the company itself acknowledging, with its own numbers, that sampling frequency moves the result. According to Evertune CEO Brian Stempeck, the gap is worse still: with a single daily measurement, the margin of error can reach 11-12 points, dropping to about 6 points only after roughly 100 repeated measurements.
The two methodologies are different by design, not by oversight. Profound runs a broad portfolio, roughly 753 prompts across 7 platforms, about once a day each, combined with real-user front-end data capture in production. Evertune uses a smaller, curated query set but repeats each one about 100 times, paired with what it calls “EverPanel,” a proprietary panel the company describes as having roughly 25 million participants. Neither methodology is “the correct one” in the abstract. Each makes a different tradeoff between broad coverage and deep repetition, and that tradeoff is exactly what determines the margin of error each one reports.
This isn’t a two-company problem, it’s a methodology problem
Here’s what gives this disagreement real weight: it isn’t exclusive to Profound or Evertune. Rand Fishkin at SparkToro found that asking an AI model the same question twice produces the same list of recommended brands less than 1% of the time, and the same ranking order less than 0.1% of the time. That doesn’t invalidate measuring aggregated visibility across many repeated queries, but it does invalidate any claim based on a single query or a handful of them.
A University of St. Gallen study and Ron Sielinski’s “IQRush” paper, covered by Search Engine Journal, arrived at the same place through a different route: 33 to 94 repetitions of the same query are needed for the result to stabilize statistically, and in 3 of 30 tests the result never stabilized, no matter how many times it was repeated. That study doesn’t even mention Profound or Evertune. It confirms sampling noise is a property of how language models work, not a defect specific to one vendor.
What this means for any visibility report you’re buying
Translated into a purchasing decision: if an AEO vendor hands you a number like “your brand appeared in 40% of responses this month” without telling you how many times they asked, how often, and across how many platforms, that number doesn’t tell you whether your brand actually improved or the model’s noise simply landed on a different side that month. The question that should be in every AI-visibility measurement contract isn’t “what’s my percentage?” It’s “how many times did you repeat the query to get that percentage?”
This connects to something we already wrote about the difference between being seen and being chosen in AI responses: aggregated visibility without method context can look like a strong signal and actually be nothing more than the normal variation of a system that doesn’t answer the same way twice.
How we handle this in our own monitoring
At IQ Source we run recurring AI-visibility monitoring for clients as part of AI Maestro work, and we apply the same discipline this disagreement just exposed: we report how often measurement happened, across how many platforms, and we explicitly separate a real trend from a fluctuation inside the expected noise margin. We’ve already written about why your AI marketing needs a verifier, and this Profound-versus-Evertune disagreement is exactly the same principle applied to how you measure whether your brand shows up in AI answers, not just how you generate the content.
Ask about the method before trusting an AI-visibility numberFrequently Asked Questions
Profound published its own analysis on July 8, 2026, showing that sampling frequency shifts measured visibility by about 2 percentage points. Per The State of Brand's July 20 report, Evertune challenges Profound's once-daily sampling methodology, citing margins of error up to 11-12 points.
Per SparkToro, asking an AI model the same question twice produces the same list of recommended brands less than 1% of the time, and the same ranking order less than 0.1% of the time. Language models generate responses with inherent variation, not a fixed answer each time.
A University of St. Gallen study and the IQRush paper found that 33 to 94 repetitions of the same query are needed for the result to stabilize statistically, and in 3 of 30 tests the result never fully stabilized, no matter how many times it was repeated.
Ask specifically how many times each query is repeated, how often measurement happens, and across how many distinct AI platforms. A visibility report that doesn't disclose its sampling method makes it impossible to tell a real trend apart from normal statistical noise in the model.
Related Articles
The AI Quality Check Your Marketing Team Doesn't Have
Legawrite.AI built a panel of AI judges for litigation, where disagreement is the signal. Marketing generates AI content at scale with no equivalent check.
21 Free Skills Turn Claude Into an MBB-Style Strategist
Oria published 21 free Claude skills for market mapping, competitive intel, and pricing. What it means for a marketing team with no in-house strategist.