Skip to main content

Pew Measured It: A Third of the New Web Is AI-Written

Pew ran 490,000 pages through a detector. The useful part is the list of stylistic markers it published, because you can audit your own archive against it.

Pew Measured It: A Third of the New Web Is AI-Written

Ricardo Argüello

Ricardo Argüello
Ricardo Argüello

CEO & Founder

AI in Marketing 5 min read

Here is the number from the Pew study that a marketing director should read first, and it is not the headline.

Government pages sit at about 1% AI authorship. University pages, also about 1%. Nonprofits, 4.6%. Commercial domains sit near 10%, and among everything published since ChatGPT launched, more than a third of pages show signs of machine authorship.

The machine text pooled exactly where the selling happens.

Your buyers are asking a system that reads that pool

Someone evaluating your company this quarter is going to ask a model before they ask you. The model composes an answer out of web pages, and Pew just told us what that set of pages is made of.

So the practical problem moved. For two years the argument inside marketing teams was whether to use AI to produce content, and the market settled that one without waiting for the argument to finish. What stays open is how an answer engine decides who to cite when a large share of the candidates read the same way.

We looked at the measurement side of this recently and found it messy: measuring AI visibility produces more sampling noise than signal at the volumes most companies can afford to run. Pew adds the other half of the picture. Tracking whether you get cited is hard, and the well being cited from got a lot more uniform.

What the study actually covers

Pew published the analysis on August 20. About 490,000 English-language pages from Common Crawl, with publication dates spanning January 2021 to July 2026, scored with Open Pangram, an open-weight detection model.

Two limits deserve to be stated plainly rather than buried.

The sample is English only. Whatever the Spanish-language number is, this study does not contain it, and I would not assume the rates transfer. Our own audience reads in both languages and I have no data for half of it.

And Pew says its detectors misclassify individual documents in both directions, with the findings holding across large collections. That is an honest caveat and also a hard ceiling on what anyone can do with the result. Nobody gets to point at a competitor’s page and declare it generated. Anyone selling you that service is selling you a coin flip with a dashboard.

The four markers, and why they are worth greping for

Pew went past the percentage and measured what changed in the prose of the web since 2023.

Em dash usage doubled. Oxford commas rose 63%. Vocabulary associated with models, words like delve, interplay, testament, pivotal, tapestry, bolstered, showcase and additionally, more than doubled. And negative parallelism, the construction that denies one thing in order to assert another inside the same sentence, nearly tripled.

That last one is banned in everything we publish.

On August 16 we added scripts/check-ai-slop.mjs to this site’s repository. It blocks negative parallelism and it blocks em dashes, among other rules, runs in about half a second across 424 files, and exits with an error so a draft carrying either one cannot ship. Four days later Pew published that those two constructions are among the fastest-growing patterns on the web.

That timing is luck, not insight. We had found both by hand, reading batches of drafts and noticing they all opened the same way.

One caution before anyone turns this into policy. These markers are a moving target. The moment everyone greps for em dashes, the em dash stops carrying information, because people writing by hand will scrub them too. Pew’s list describes the web through July 2026 and it will age. The habit that survives is checking against something concrete instead of against a feeling.

Three things worth doing to your archive this week

Pull your last thirty published articles and count. Em dashes per thousand characters, and how many pieces open or close on the negative-parallelism move. An afternoon of work, and at the end you have a number instead of an impression. If the number surprises you, there was a template running and nobody had named it.

Then read those pieces in sequence, the way a subscriber sees them, and watch the shape rather than the quality. Any single article can be good while the set gives away the process: same paragraph length, same section count, same kind of ending. If you get bored on the third one, your reader got bored on the first.

Last, put the proprietary fact first. When a draft starts with the model and someone goes looking for data afterward, the data arrives as decoration and the structure has already set. When a real number from your operation exists first, the piece organizes itself around something no model had. We put numbers on that in the post on earned media and the 2% overlap in AI citations.

What I would skip: writing to beat a detector. That is optimizing against an exam that gets rewritten every quarter, and we all already ran that play with meta keywords in 2005. It did not end well then either!

This pairs with what we published Sunday about LinkedIn’s button for reporting AI slop. That one is about reach inside one platform. This one is about the whole web and who ends up quoted. Same editorial habit, two different problems.

Start with the thirty most recent pieces. Those are the ones an answer engine is looking at right now.

Let us audit what your content is actually being checked against

Frequently Asked Questions

In the study Pew Research Center published on August 20, 2026, 10% of all pages in its July 2026 sample showed signs of AI authorship. Narrowed to pages published after ChatGPT launched in November 2022, that figure rises to more than one third of pages.

Pew drew roughly 490,000 English-language pages from the Common Crawl archive, published between January 2021 and July 2026, and scored them with Open Pangram, an open-weight detection model. Pew notes that AI detectors misclassify individual documents and that its conclusions apply to large collections rather than single pages.

Pew tracked four signals that grew across the web since 2023. Em dash usage doubled, Oxford comma usage rose 63%, vocabulary associated with language models more than doubled, and negative parallelism, the construction that denies one thing to assert another in the same sentence, nearly tripled.

Answer engines assemble responses from web pages, and that pool now contains a high share of machine-written text. As those systems tune which sources they treat as trustworthy, content carrying the stylistic markers of generated text competes at a disadvantage for citations, regardless of how accurate it happens to be.

Pew Research Center AEO AI content Common Crawl B2B marketing answer engines AI detection

Related Articles

Your Shorts are deposits into the AI that cites you
AI in Marketing
· 6 min read

Your Shorts are deposits into the AI that cites you

Gary Vaynerchuk says YouTube Shorts became his number one platform. Not for the views, but because every video is a deposit into the AEO fight ahead.

AEO AI marketing YouTube Shorts
Perplexity at $30B on $750M, and Nvidia May Fund It
Business Strategy
· 4 min read

Perplexity at $30B on $750M, and Nvidia May Fund It

The Information reports Nvidia in talks to back Perplexity above $30B. In January it committed its entire annual revenue to three years of Azure compute.

Perplexity Nvidia answer engines
AI Writing Policies Are Brand Voice Now, Not IT Policy
AI in Marketing
· 4 min read

AI Writing Policies Are Brand Voice Now, Not IT Policy

Clay and Heidi Health posted their AI writing policies on LinkedIn instead of a handbook. What a marketing team's version needs to cover, and why it matters.

AI policy brand voice Clay
Nature Measured What AI Polishing Deletes
AI in Marketing
· 5 min read

Nature Measured What AI Polishing Deletes

A Nature Human Behaviour paper names marketing among the affected industries. Polishing with AI keeps what you said and flattens how you said it.

brand voice AI content Nature Human Behaviour
LinkedIn's AI Slop Button Puts a Price on Your Reach
AI in Marketing
· 6 min read

LinkedIn's AI Slop Button Puts a Price on Your Reach

Pangram found over 40% of LinkedIn long-form posts are fully AI-written. LinkedIn added a report button and classifiers that trim recommended reach.

LinkedIn AI content Pangram
Your AI Marketing Doesn't Need a Smarter Model
AI in Marketing
· 6 min read

Your AI Marketing Doesn't Need a Smarter Model

Your team ships AI content fast and it all looks great. The problem is not the model, it's that nobody verifies before it goes out. That is your edge.

AI marketing AI content verification