Nature Measured What AI Polishing Deletes
Ricardo Argüello, September 9, 2026
CEO & Founder
General summary
A study published in Nature Human Behaviour on August 24, 2026 by a University of Southern California team found that when a language model polishes a text, meaning survives almost intact while writing-complexity variance drops by 21 to 50 percent. The paper names marketing among the industries that stand to be affected by that loss of linguistic variability.
- Writing-complexity variance falls 21 to 50 percent after a model polishes the text
- Meaning survives, rated 2.97 out of 3.00 by human evaluators
- Rewrites drift toward an older, male, politically liberal profile with higher moral valence and lower empathy
- None of the twelve prompts the study tested asked the model to preserve the author's voice
- Roughly half of articles published on the web are primarily AI, but only 7 percent of those ranking first on Google
Record ten people telling the same story, then run all ten recordings through the same audio filter. The stories are still different. The voices are not. That is what a style pass does to your content, and in marketing the voice was the thing you were selling.
AI-generated summary
Half the articles published on the web are now primarily AI-generated. Seven percent of the articles ranking first on Google are.
Supply doubled. Visibility did not move.
A paper published in Nature Human Behaviour on August 24 explains a good part of why, and it names our department out loud.
The paper says “marketing”
The authors write that industries relying on nuanced text analysis for mass personalization, marketing among them, stand to be affected, and that by blurring linguistic variability the models may reduce the effectiveness of targeted advertising.
A University of Southern California team analyzed more than 880,000 texts across seven datasets.
One correction worth making before anything else, because most coverage got it wrong and so did my own notes until I opened the PDF. Those 880,000 texts were analyzed, not rewritten. Most of that count is an observational corpus of Reddit, local news and arXiv that no model ever touched. The rewriting experiment ran on about 9,400 source texts, each passed through three models and up to twelve prompts.
The core result holds anyway. When a model polishes a text, the content survives while writing-complexity variance falls between 21 and 50 percent. The paper sits behind a paywall, but the author’s version is open on arXiv.
Variance, not level. The writing does not get simpler. It gets more like the writing next to it.
Four raters scored meaning preservation at 2.97 out of 3.00. What you said is still there. How you said it is what leaves, and in marketing that was the asset.
The drift has a direction
Classifiers trained on the originals, then run against the polished versions, lost about 6 points of accuracy at identifying author traits.
The rewrites consistently align with authors who are older, male, politically liberal, higher in moral valence and lower in empathy.
Effect sizes matter here and almost no summary lists them in order. The largest by a wide margin is political, at 1.35. Age follows at 1.09. Empathy, the one every headline leads with, turns out to be the weakest of the five at 0.35.
Study 3 gets specific about what dies. The link between gender and negative-emotion words stops being significant after rewriting. So does the link between extraversion and pronoun use, and the one between age and future-focused words. Others survive untouched, like neuroticism and negative emotion.
Nothing gets erased evenly. Some markers go and some stay, and the ones that go are the ones that told you apart.
Read the twelve prompts
I went and read the twelve prompts. They are all in the methods section.
“Rewrite the following text using the best syntax and grammar.” “Rephrase the following text.” “Rewrite the following text to improve clarity and readability.” Twelve variations on the same instruction.
Not one of them asks the model to keep the author’s voice. The paper says so plainly, that it prompted without emphasizing any particular stylistic features, to capture neutral rewriting effects rather than targeted transformations.
So the study does not show that telling a model to write in your voice fails. It shows that twelve different ways of saying “improve this” all produce the same drift. That is a real difference and a lot of people are skipping past it this week.
One more limit. The models tested were GPT-3.5, Llama 3 70B and Gemini Pro. All 2023 and 2024 generation, none of them current frontier models.
Where this shows up before your brand notices
Back to the numbers I opened with, because they are the business case.
Graphite’s Common Crawl analysis puts primarily-AI articles at roughly half of everything published in the first quarter of 2026. Among articles ranking in Google’s top two positions, 14 percent. Among those ranking first, 7 percent.
LinkedIn already has a button for the same pattern. When the platform shipped its “seems like AI slop” report, a million people used it in the first two weeks, and LinkedIn says flagged content lost 40 percent of its views. We covered that in the LinkedIn slop button and what it costs your reach, and how much of the new web already reads that way in Pew’s Common Crawl study.
What I would do with your content team
Three concrete things, and none of them is stopping AI use.
Split the two jobs. Drafting and style-polishing are different tasks and should not run in the same step. The drift the paper measured lives in the polish, not in the generation.
Keep your unpolished originals. If everything you publish goes through a style pass, in six months you will have nothing left to compare against. The corpus of your own voice is an asset and it overwrites itself quietly.
Then run the paper’s test in your own building. Take ten pieces before and after the AI pass, hand both versions to someone on your team who did not write them, and ask which one sounds more like you. You do not need a classifier. You need half an hour.
Start with the ten pieces. The answer tells you whether you have a voice problem or you were worrying for free.
Let’s check whether your content still sounds like your brandFrequently Asked Questions
That polishing preserves content while homogenizing style, cutting writing-complexity variance by 21 to 50 percent depending on the dataset and model. The paper was published on August 24, 2026 by a team at the University of Southern California across seven datasets and more than 880,000 analyzed texts.
It flattens the markers that identify it. The study measured an average drop of about 6 points in the accuracy of classifiers trained to recognize author traits. Specific markers, such as the link between extraversion and pronoun use, stopped being statistically significant after rewriting.
The study never tested it. Its twelve prompts were all variants of improve this text, and the paper states it prompted without emphasizing any particular stylistic features in order to capture neutral rewriting effects. Voice instructions failing is a gap in the evidence, not a finding.
Roughly half of articles published on the web are primarily AI-generated according to Graphite's Common Crawl analysis, yet only 14 percent of articles ranking in Google's top two positions are AI-generated, and just 7 percent of those ranking first.
Related Articles
Your Shorts are deposits into the AI that cites you
Gary Vaynerchuk says YouTube Shorts became his number one platform. Not for the views, but because every video is a deposit into the AEO fight ahead.
Pew Measured It: A Third of the New Web Is AI-Written
Pew ran 490,000 pages through a detector. The useful part is the list of stylistic markers it published, because you can audit your own archive against it.
LinkedIn's AI Slop Button Puts a Price on Your Reach
Pangram found over 40% of LinkedIn long-form posts are fully AI-written. LinkedIn added a report button and classifiers that trim recommended reach.
Your AI Marketing Doesn't Need a Smarter Model
Your team ships AI content fast and it all looks great. The problem is not the model, it's that nobody verifies before it goes out. That is your edge.
A One-Dollar CRM That Fills Itself, From HubSpot's CTO
Dharmesh Shah launched YouSpot, a solo CRM that assembles relationship context from your inbox and calendar. Here is what that does to a marketing team's job.
85.5% Earned Media, 2% Overlap With What You Pitch
The stat every PR team is sharing and the 2% overlap between pitches and AI citations come from different measurements. Here is what each one is good for.