Pew Measured It: A Third of the New Web Is AI-Written
Ricardo Argüello, August 26, 2026
CEO & Founder
General summary
On August 20, 2026, Pew Research Center published an analysis of roughly 490,000 English-language pages pulled from Common Crawl. Among pages published after ChatGPT launched, more than a third show signs of machine authorship, and the concentration sits almost entirely on commercial domains. The part worth acting on is not the headline percentage but the four measurable style markers Pew tracked, because a marketing team can audit its own archive against those in an afternoon.
- Pew analyzed about 490,000 English-language pages from Common Crawl covering January 2021 through July 2026, using the Open Pangram detector
- In the July 2026 sample, 10% of all pages showed AI authorship signs, rising above one third for pages published after November 2022
- Commercial domains carry it: .com sits near 10%, .org at 4.6%, and .edu and .gov both around 1%
- Four markers grew since 2023: em dashes doubled, Oxford commas rose 63%, AI-associated vocabulary more than doubled, and negative parallelism nearly tripled
- Pew states detectors misclassify individual documents, so the findings hold in aggregate and cannot be used to accuse a single page
Imagine twenty years of handwritten mail, and then one third of the envelopes start arriving in the same printed typeface. You do not need to read them to notice. The stack looks different from across the room. That is what happened to the text of the web, and the systems deciding which sources to quote are learning to recognize that typeface before they read a word of what it says.
AI-generated summary
Here is the number from the Pew study that a marketing director should read first, and it is not the headline.
Government pages sit at about 1% AI authorship. University pages, also about 1%. Nonprofits, 4.6%. Commercial domains sit near 10%, and among everything published since ChatGPT launched, more than a third of pages show signs of machine authorship.
The machine text pooled exactly where the selling happens.
Your buyers are asking a system that reads that pool
Someone evaluating your company this quarter is going to ask a model before they ask you. The model composes an answer out of web pages, and Pew just told us what that set of pages is made of.
So the practical problem moved. For two years the argument inside marketing teams was whether to use AI to produce content, and the market settled that one without waiting for the argument to finish. What stays open is how an answer engine decides who to cite when a large share of the candidates read the same way.
We looked at the measurement side of this recently and found it messy: measuring AI visibility produces more sampling noise than signal at the volumes most companies can afford to run. Pew adds the other half of the picture. Tracking whether you get cited is hard, and the well being cited from got a lot more uniform.
What the study actually covers
Pew published the analysis on August 20. About 490,000 English-language pages from Common Crawl, with publication dates spanning January 2021 to July 2026, scored with Open Pangram, an open-weight detection model.
Two limits deserve to be stated plainly rather than buried.
The sample is English only. Whatever the Spanish-language number is, this study does not contain it, and I would not assume the rates transfer. Our own audience reads in both languages and I have no data for half of it.
And Pew says its detectors misclassify individual documents in both directions, with the findings holding across large collections. That is an honest caveat and also a hard ceiling on what anyone can do with the result. Nobody gets to point at a competitor’s page and declare it generated. Anyone selling you that service is selling you a coin flip with a dashboard.
The four markers, and why they are worth greping for
Pew went past the percentage and measured what changed in the prose of the web since 2023.
Em dash usage doubled. Oxford commas rose 63%. Vocabulary associated with models, words like delve, interplay, testament, pivotal, tapestry, bolstered, showcase and additionally, more than doubled. And negative parallelism, the construction that denies one thing in order to assert another inside the same sentence, nearly tripled.
That last one is banned in everything we publish.
On August 16 we added scripts/check-ai-slop.mjs to this site’s repository. It blocks negative parallelism and it blocks em dashes, among other rules, runs in about half a second across 424 files, and exits with an error so a draft carrying either one cannot ship. Four days later Pew published that those two constructions are among the fastest-growing patterns on the web.
That timing is luck, not insight. We had found both by hand, reading batches of drafts and noticing they all opened the same way.
One caution before anyone turns this into policy. These markers are a moving target. The moment everyone greps for em dashes, the em dash stops carrying information, because people writing by hand will scrub them too. Pew’s list describes the web through July 2026 and it will age. The habit that survives is checking against something concrete instead of against a feeling.
Three things worth doing to your archive this week
Pull your last thirty published articles and count. Em dashes per thousand characters, and how many pieces open or close on the negative-parallelism move. An afternoon of work, and at the end you have a number instead of an impression. If the number surprises you, there was a template running and nobody had named it.
Then read those pieces in sequence, the way a subscriber sees them, and watch the shape rather than the quality. Any single article can be good while the set gives away the process: same paragraph length, same section count, same kind of ending. If you get bored on the third one, your reader got bored on the first.
Last, put the proprietary fact first. When a draft starts with the model and someone goes looking for data afterward, the data arrives as decoration and the structure has already set. When a real number from your operation exists first, the piece organizes itself around something no model had. We put numbers on that in the post on earned media and the 2% overlap in AI citations.
What I would skip: writing to beat a detector. That is optimizing against an exam that gets rewritten every quarter, and we all already ran that play with meta keywords in 2005. It did not end well then either!
This pairs with what we published Sunday about LinkedIn’s button for reporting AI slop. That one is about reach inside one platform. This one is about the whole web and who ends up quoted. Same editorial habit, two different problems.
Start with the thirty most recent pieces. Those are the ones an answer engine is looking at right now.
Let us audit what your content is actually being checked againstFrequently Asked Questions
In the study Pew Research Center published on August 20, 2026, 10% of all pages in its July 2026 sample showed signs of AI authorship. Narrowed to pages published after ChatGPT launched in November 2022, that figure rises to more than one third of pages.
Pew drew roughly 490,000 English-language pages from the Common Crawl archive, published between January 2021 and July 2026, and scored them with Open Pangram, an open-weight detection model. Pew notes that AI detectors misclassify individual documents and that its conclusions apply to large collections rather than single pages.
Pew tracked four signals that grew across the web since 2023. Em dash usage doubled, Oxford comma usage rose 63%, vocabulary associated with language models more than doubled, and negative parallelism, the construction that denies one thing to assert another in the same sentence, nearly tripled.
Answer engines assemble responses from web pages, and that pool now contains a high share of machine-written text. As those systems tune which sources they treat as trustworthy, content carrying the stylistic markers of generated text competes at a disadvantage for citations, regardless of how accurate it happens to be.
Related Articles
Your Shorts are deposits into the AI that cites you
Gary Vaynerchuk says YouTube Shorts became his number one platform. Not for the views, but because every video is a deposit into the AEO fight ahead.
LinkedIn's AI Slop Button Puts a Price on Your Reach
Pangram found over 40% of LinkedIn long-form posts are fully AI-written. LinkedIn added a report button and classifiers that trim recommended reach.
Your AI Marketing Doesn't Need a Smarter Model
Your team ships AI content fast and it all looks great. The problem is not the model, it's that nobody verifies before it goes out. That is your edge.
85.5% Earned Media, 2% Overlap With What You Pitch
The stat every PR team is sharing and the 2% overlap between pitches and AI citations come from different measurements. Here is what each one is good for.
That AI Visibility Percentage Your Vendor Sold You May Be Noise
Profound and Evertune are publicly arguing over how to measure whether your brand shows up in ChatGPT. Most visibility numbers never say how they sampled.
Your marketing asks AI what to cut. Ask this instead.
Box created 13 new AI roles, one to market to industries it could not staff before. The question is not what AI lets you cut, but what it makes possible.