Skip to main content

That AI Visibility Percentage Your Vendor Sold You May Be Noise

Profound and Evertune are publicly arguing over how to measure whether your brand shows up in ChatGPT. Most visibility numbers never say how they sampled.

That AI Visibility Percentage Your Vendor Sold You May Be Noise

Ricardo Argüello

Ricardo Argüello
Ricardo Argüello

CEO & Founder

AI in Marketing 4 min read

Profound and Evertune, two of the most-cited vendors for measuring whether a brand shows up in ChatGPT, Claude, or Perplexity answers, are having a public methodology disagreement. It’s not a lawsuit or a scandal, it’s more useful than that: it’s a technical dispute over how to measure something most marketing teams are already buying as if it were a solid number.

The disagreement, reported by The State of Brand on July 20, has a far more concrete and verifiable origin on Profound’s side.

Profound’s own data reveals the problem

On July 8, Profound economist Jennifer Zou published a direct analysis on the company blog titled “Is once a day enough?” They ran 753 prompts across 7 AI platforms, comparing once-daily sampling (129,000 measurements) against ten-times-daily sampling (860,000 measurements) over two weeks in June. The result: measured visibility shifted by about 2 percentage points depending on sampling frequency, and citation share shifted by about 0.3 points.

That’s not a competitor’s accusation. It’s the company itself acknowledging, with its own numbers, that sampling frequency moves the result. According to Evertune CEO Brian Stempeck, the gap is worse still: with a single daily measurement, the margin of error can reach 11-12 points, dropping to about 6 points only after roughly 100 repeated measurements.

The two methodologies are different by design, not by oversight. Profound runs a broad portfolio, roughly 753 prompts across 7 platforms, about once a day each, combined with real-user front-end data capture in production. Evertune uses a smaller, curated query set but repeats each one about 100 times, paired with what it calls “EverPanel,” a proprietary panel the company describes as having roughly 25 million participants. Neither methodology is “the correct one” in the abstract. Each makes a different tradeoff between broad coverage and deep repetition, and that tradeoff is exactly what determines the margin of error each one reports.

This isn’t a two-company problem, it’s a methodology problem

Here’s what gives this disagreement real weight: it isn’t exclusive to Profound or Evertune. Rand Fishkin at SparkToro found that asking an AI model the same question twice produces the same list of recommended brands less than 1% of the time, and the same ranking order less than 0.1% of the time. That doesn’t invalidate measuring aggregated visibility across many repeated queries, but it does invalidate any claim based on a single query or a handful of them.

A University of St. Gallen study and Ron Sielinski’s “IQRush” paper, covered by Search Engine Journal, arrived at the same place through a different route: 33 to 94 repetitions of the same query are needed for the result to stabilize statistically, and in 3 of 30 tests the result never stabilized, no matter how many times it was repeated. That study doesn’t even mention Profound or Evertune. It confirms sampling noise is a property of how language models work, not a defect specific to one vendor.

What this means for any visibility report you’re buying

Translated into a purchasing decision: if an AEO vendor hands you a number like “your brand appeared in 40% of responses this month” without telling you how many times they asked, how often, and across how many platforms, that number doesn’t tell you whether your brand actually improved or the model’s noise simply landed on a different side that month. The question that should be in every AI-visibility measurement contract isn’t “what’s my percentage?” It’s “how many times did you repeat the query to get that percentage?”

This connects to something we already wrote about the difference between being seen and being chosen in AI responses: aggregated visibility without method context can look like a strong signal and actually be nothing more than the normal variation of a system that doesn’t answer the same way twice.

How we handle this in our own monitoring

At IQ Source we run recurring AI-visibility monitoring for clients as part of AI Maestro work, and we apply the same discipline this disagreement just exposed: we report how often measurement happened, across how many platforms, and we explicitly separate a real trend from a fluctuation inside the expected noise margin. We’ve already written about why your AI marketing needs a verifier, and this Profound-versus-Evertune disagreement is exactly the same principle applied to how you measure whether your brand shows up in AI answers, not just how you generate the content.

Ask about the method before trusting an AI-visibility number

Frequently Asked Questions

AEO AI visibility Profound Evertune brand measurement answer engine optimization AI Maestro

Related Articles

85.5% Earned Media, 2% Overlap With What You Pitch
AI in Marketing
· 7 min read

85.5% Earned Media, 2% Overlap With What You Pitch

The stat every PR team is sharing and the 2% overlap between pitches and AI citations come from different measurements. Here is what each one is good for.

AEO earned media AI citations
Your Shorts are deposits into the AI that cites you
AI in Marketing
· 6 min read

Your Shorts are deposits into the AI that cites you

Gary Vaynerchuk says YouTube Shorts became his number one platform. Not for the views, but because every video is a deposit into the AEO fight ahead.

AEO AI marketing YouTube Shorts
Your marketing asks AI what to cut. Ask this instead.
AI in Marketing
· 7 min read

Your marketing asks AI what to cut. Ask this instead.

Box created 13 new AI roles, one to market to industries it could not staff before. The question is not what AI lets you cut, but what it makes possible.

AI marketing marketing strategy Box