Study finds AI writing tools narrow linguistic diversity across 880,000 texts

A USC study in Nature Human Behaviour found LLM writing assistants are flattening linguistic diversity, analyzing 880,000 texts and showing AI rewriting cut writing complexity variance by up to 50%.

Categorized in: AI News Writers
Published on: Aug 28, 2026
Study finds AI writing tools narrow linguistic diversity across 880,000 texts

A University of Southern California research team published a study in Nature Human Behaviour this week showing that the widespread adoption of large language model (LLM) writing assistants is accelerating linguistic homogenization. Analyzing more than 880,000 texts, the researchers found that writing style differences across arXiv, Reddit, and local news narrowed significantly after ChatGPT's November 2022 launch.

In controlled experiments, AI rewriting reduced the variance in writing complexity by 21% to 50%. The study also found that prediction accuracy for identity traits such as age and personality based on text dropped by roughly 6 percentage points on average, with the shift systematically biased toward specific profiles.

Writing style convergence after AI intervention

The team examined three large longitudinal datasets: approximately 80,000 arXiv paper abstracts, roughly 379,600 U.S. local news articles, and about 318,000 Reddit creative stories, spanning 2018 to 2024. Using Binoculars, a high-precision AI text detection tool, they tracked changes in the variance of writing complexity metrics. After ChatGPT's public release, writing style differences in lexical diversity and sentence structure narrowed noticeably in all three datasets.

To establish causality, the researchers randomly sampled 1,000 human-written texts from Reddit and arXiv predating ChatGPT's release and had them rewritten by GPT-3.5, Llama 3 70B, and Gemini Pro using neutral prompts such as "improve grammar," "enhance readability," and "write more naturally." Even though AI maintained semantic similarity above 0.95 in 87% of texts, the variance in writing complexity after rewriting was still significantly compressed, with reductions ranging from 21% to 50%.

The effect appeared regardless of which LLM was used or how rewriting instructions were adjusted. The convergence trend in writing style repeated across models and prompt types.

Identity signals shift toward a "default persona"

The second sub-study tested whether personal traits could still be predicted from AI-rewritten text. The team trained classifiers to predict age, gender, Big Five personality traits, empathy, and moral values from textual features, then compared prediction accuracy between original and rewritten texts.

Prediction accuracy dropped by an average of approximately 6 percentage points (absolute F1). Age was the most affected trait, with F1 scores falling from 0.351 to 0.260. However, classifier performance remained above random chance, indicating that AI weakened identity signals rather than completely erasing them.

The change was not random noise but systematically biased. LLM-rewritten texts were more frequently classified as older, more morally grounded, less empathetic, and less extraverted, while higher in openness and agreeableness. The researchers said LLMs do not simply "erase" identity signals-they introduce new biases that push different authors' writing toward a similar "default persona."

Selective erosion of language-psychology associations

The third sub-study tested whether established associations between word categories and psychological traits still held after LLM rewriting. In original texts, the team replicated many classic findings, such as extraverts using more positive emotion words and openness correlating with complex vocabulary. After LLM rewriting, many of these associations were "washed out," while others-such as the link between neuroticism and negative emotion words-were preserved.

The team described this as "selective erosion" rather than wholesale distortion. For disciplines that rely on lexical cues for research, this may be more problematic than complete distortion: researchers must test which associations remain reliable and which have broken down, otherwise they may draw biased conclusions about identity, group differences, or cultural patterns.

Korean media reports noted a significant limitation: the paper did not directly prove that stylistic homogenization in actual online texts was caused by AI use. However, the controlled experiment results provide strong support for the association.

From clinical screening to cultural preservation

The research team enumerated several domains affected by linguistic homogenization. In clinical and mental health contexts, identifying linguistic signals of depression or suicidal ideation through text may become more difficult as writing styles converge, hindering early detection. In hiring, applicants whose writing has been AI-polished to conform to mainstream stylistic expectations may gain an advantage over those with distinctive but less "polished" writing, exacerbating unfairness. In cultural preservation, the cultural markers embedded in language are being smoothed away.

A more concerning mechanism is the feedback loop: LLM-influenced texts continuously enter online corpora, and those corpora will in turn be used to train the next generation of models. The loss of linguistic diversity may progressively self-reinforce.

The researchers also raised a broader question: given that language and thought are inseparable, linguistic homogenization may ultimately constrain cognitive flexibility and creativity. In extreme cases, as George Orwell warned in 1984, this could "narrow the range of thought."

Why this matters for writers

For working writers, the study's practical implication is direct: if you rely on LLM rewriting tools to polish drafts, your distinctive voice is being systematically averaged toward a statistical mean. The research shows this happens even with neutral prompts like "improve grammar"-not just with aggressive rewriting. Writers who want to preserve their individual style should consider using AI for discrete tasks like checking grammar or generating alternatives, rather than feeding entire passages through rewriting tools. The study also suggests that for AI for Writers professionals, understanding how Generative AI and LLM tools shape output is becoming a core competency-not just for quality control, but for maintaining the stylistic diversity that makes writing valuable in the first place.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)