Article is online

Pew Research Finds One-Third of Post-ChatGPT Webpages Show AI Authorship Signals

Pew Research Finds One-Third of Post-ChatGPT Webpages Show AI Authorship Signals

Table of Contents




You might want to know


Has the introduction of ChatGPT and similar models dramatically changed the makeup of English-language web content? How reliably can detectors distinguish human-written text from AI-assisted or AI-generated text?



Main Topic


A new analysis from the Pew Research Center examined a large sample of English-language web pages and found that a substantial portion show patterns consistent with AI authorship, particularly for content published after the public release of ChatGPT in November 2022. Researchers used roughly 490,000 pages drawn from the Common Crawl archive covering January 2021 through July 2026. They passed the extracted text through an AI-detection model called Open Pangram to estimate where AI had likely played a role in composing the material.



The headline finding is straightforward: across that multi-year sample, about 10% of pages show significant signs of AI authorship. When the dataset is limited to pages published after ChatGPT's launch in November 2022, the estimated share rises sharply to about 35%. This suggests a meaningful acceleration of AI use in web publishing following the mainstream availability of large language models.



The distribution of detected AI authorship is not uniform across domain types. Commercial ".com" sites exhibit AI-authorship indicators at substantially higher rates than academic (".edu") or government (".gov") domains. Pew reports that ".com" domains show signs of AI involvement at roughly ten times the rate of ".edu" and ".gov" sites, each of which sits near the 1% level, and about twice the rate observed on ".org" domains (around 4.6% in the sample).



Part of the domain gap reflects who publishes what and how. Many academic and government pages undergo editorial review, institutional approval, and slower publication cycles, which reduces the incentive and opportunity for rapid, high-volume AI-generated output. Commercial web spaces, by contrast, include everything from newsroom content to search-traffic-driven affiliate and marketing pages that can be produced and published at high velocity—conditions where AI assistance or full automation is more attractive.



Pew’s approach to detecting AI does not rely on single ‘‘tells’’ but rather looks at statistical patterns across large text samples. Nonetheless, the researchers identified several stylistic and lexical markers that have grown more common since 2023. Em dashes, for example, appear about twice as often in more recent text; Oxford commas have increased by roughly 63%; and certain words favored by contemporary AI outputs (words such as "delve," "interplay," and "testament") have more than doubled in frequency. Another pattern dubbed "negative parallelism"—phrases like "it's not just X, it's Y"—has nearly tripled in incidence since 2023, though it remains uncommon overall.



Those linguistic shifts are consistent with other monitoring efforts and lexicographic observations: trackers and analysts have flagged similar signs in previous research, and dictionary publishers have noted shifts in language use that reflect larger trends. Even so, Pew emphasizes important caveats. Detection models can misclassify individual pages in either direction: they may mistakenly flag human-authored text as AI-generated or fail to identify AI-influenced material. Moreover, a label of "significant signs of AI authorship" does not imply that a page was entirely machine-generated—many pages likely contain a mix of human and AI contributions, such as AI-assisted drafts later edited by people.



Open Pangram, the detector used in Pew’s analysis, was developed by Pangram Labs and has appeared in other studies. For instance, researchers using similar detection tools have reported AI-generated content in a portion of U.S. newspaper articles, including opinion pieces at major outlets. The presence of detection signals across different datasets suggests that the imprint of AI on written content is measurable, even if individual classifications are sometimes uncertain.



Looking forward, AI detection itself could change as the technology evolves. Large AI developers are exploring model-level text fingerprinting—cryptographic or statistical markers embedded at generation time—that would make AI-produced text much easier to identify while aiming to limit false positives. If widely adopted, such approaches would alter the detection landscape by enabling more reliable attribution of machine-generated text.



At scale and over time, Pew’s data show a steep upward trend for AI-authorship markers on commercial sites: the rate for ".com" pages climbed from roughly 1% in January 2021 to about 9.35% in January 2026. That rise illustrates how quickly AI tools have been integrated into workflows that prioritize volume, search visibility, or rapid content turnover.



In summary, Pew’s large-scale web sweep finds that AI influence on English-language web content is widespread and rising, particularly for material published after ChatGPT’s debut. While detection has limitations and the line between AI-assisted and fully AI-generated content can be blurred, the patterns identified point to a significant shift in how online content is produced and highlight the need for continued monitoring, thoughtful editorial policies, and technological solutions that improve transparency and attribution.



Key Insights Table













AspectDescription
Sample size~490,000 English-language pages from Common Crawl (Jan 2021–Jul 2026)
Overall AI signal~10% of pages show significant signs of AI authorship
Post-ChatGPT pages~35% of pages published since Nov 2022 show AI signals
Domain differences.com domains show much higher AI rates (~10x .edu/.gov; ~2x .org)
Detection methodOpen Pangram (statistical pattern detection across large text batches)
Notable linguistic markersIncreased em dashes, Oxford commas, certain favored words, and "negative parallelism"
LimitationsPotential misclassification; "signs" do not confirm full AI generation


Afterwards...


The growing footprint of AI in web content raises practical and ethical questions about transparency, quality, and attribution. As detection tools and model-level fingerprinting evolve, stakeholders—publishers, platforms, researchers, and readers—will need to balance automation benefits with verification safeguards. Continued large-scale monitoring, clearer labeling practices, and advances in detection fidelity will shape how the web adapts to a future in which human and machine authorship increasingly intertwine.



Understanding these dynamics matters for search quality, academic integrity, media trust, and content governance. The data from Pew provide a benchmark: AI influence on the web is measurable, accelerating, and concentrated where publication speed and scale are prioritized. That combination suggests both opportunity and risk, and it calls for policies and tools that preserve information quality while enabling responsible use of AI.


Last edited at:2026/8/22

Claude AI

AI Smart Editor