Pew Research Center has found signs of AI writing or substantial editing on 35% of English-language webpages posted since ChatGPT launched in November 2022. The analysis, reported by TechRadar on Aug. 25, drew on about 490,000 webpages collected from the Common Crawl archive between 2021 and July 2026, with a randomly selected sample of 10,000 pages from July 2026 examined for AI authorship markers.
Across the full sample - which included pages dating back to the early days of the web - 9.6% showed signs of AI involvement. The jump to 35% when isolating post-ChatGPT pages suggests AI's role in web publishing is growing quickly.
Commercial sites lead AI adoption
The share of AI-influenced pages varied sharply by domain type. Commercial .com pages showed the highest concentration of AI writing markers in the 2026 sample. By contrast, .edu pages from U.S. educational institutions sat at about 1%, .gov pages at 0.8%, and .org pages at about 4.6%.
That gap points to commercial pressures driving AI adoption, while education and government sites remain more conservative in their publishing practices.
Linguistic markers of AI text
Pew also tracked common linguistic traits in AI-written text across the broader corpus. On pages published since 2023, em dash usage doubled and Oxford comma usage rose 63%. Certain words favored by AI models - including "delve," "interplay," and "testament" - more than doubled in frequency. Negative parallel constructions like "not only X but also Y" nearly tripled.
Pew stressed that these traits alone cannot prove a given text was written by AI. Em dashes and Oxford commas appear in human writing too. The researchers said accurate judgment requires combining statistical patterns in word choice and sentence structure across large-scale document sets rather than inspecting individual pages.
The findings also expose the limits of AI content detection. Punctuation and word frequency markers can't reliably separate human from machine writing in isolation. But at scale, they confirm a measurable shift in how web content is being produced.
Why this matters for writers
For writers, the takeaway is practical: AI-generated text now has identifiable statistical fingerprints, but those fingerprints aren't definitive proof. The same markers that flag AI writing - em dashes, Oxford commas, parallel structures - are common in skilled human prose. That ambiguity cuts both ways. Editors can't rely on style alone to police AI use, and writers who use ChatGPT as a drafting tool need to understand which patterns will make their work look machine-assisted.
The domain gap matters too. If commercial sites are leaning on AI far more heavily than institutional publishers, that's a signal about where editorial standards are heading. Writers who want to stand out - or who work for outlets that want to stay credible - have a reason to understand what AI text looks like and to develop editing habits that go beyond surface style. The tools for AI for Writers are evolving, but so is the scrutiny applied to the content writers produce.
Your membership also unlocks: