Researchers measured the scale against a huge archive. They analyzed nearly 490,000 pages pulled from the Common Crawl archive between January 2021 and July 2026. Common Crawl is a massive public database that captures snapshots of the internet over time. The team processed the sampled text using Open Pangram, an AI detection model built to flag synthetic writing. The dataset offers one of the broadest looks yet at how fast generative tools have spread.
Linguistic fingerprints and detection limits
The analysis revealed clear patterns in how generative systems construct text. Language models overused specific vocabulary. Terms such as “compliant with”, “key”, “significant”, and “valuable” appeared more than twice as often in AI-written or AI-edited material. Em dashes also showed up twice as frequently in automated text as in human writing. These repetitive phrasing habits leave a distinct mark across synthetic content.
Pew also noticed more negative parallelism. This is a writing style where a sentence states that something is not only X, it is Y. AI models use this comparison three times more often than humans do. For newsroom leaders, the pattern cuts both ways. Machines can fill pages fast, but they also create a sea of sameness that buries human journalism under low-cost filler.

Pew emphasized that detection technology has clear limits. AI detectors can misclassify individual documents, and the results should not be treated as definitive proof of a page’s authorship. These scores are estimates based on probability. They do not settle the question of who wrote a single page. Instead, they show how the internet is changing as a whole.
A divide between commercial and institutional webs
The study found that machine-written text is concentrated heavily in commercial spaces. In 2026 samples, roughly one in ten web pages on .com domains showed machine-written markers. By comparison, synthetic markers appeared on 4.6 percent of .org domains, and on only about 1 percent each of .edu and .gov sites. Across all domains, about 9.6 percent of analyzed websites showed indicators of automation, regardless of when they were published.
If you run a local newsroom, this distribution matters. Most news sites live on .com domains. This is exactly where the flood of AI filler is highest. The noise around genuine reporting keeps growing. In a full random sample of 10,000 pages collected in July 2026, roughly 10 percent showed significant signs of AI authorship.
One methodological note deserves attention. The researchers found that only about 10 to 15 percent of web pages in each sample carried a usable publication date field in their HTML code. That narrow slice was the basis for the one-third finding. In other words, the 35 percent figure describes the pages whose dates could be verified, not the whole archive at once.
What a synthetic web means for trust
Synthetic content is becoming a standard feature of the open web, and readers are noticing. A majority of American adults under 30 now say they are more concerned than excited about the increased use of AI in daily life. When one in three new pages relies on automated text, a trusted brand name and verified human oversight become the clearest way to stand apart from the filler.
The challenge for publishers is simple but hard. Newsrooms must show readers that a person checked the facts and stood behind the words on the page. Transparent standards, not speed, will decide who keeps an audience as the synthetic web grows. The question for the next few years is whether readers will pay for that difference.
Written by Dominik Czarnota using the Tribune Desk AI platform. Every claim in this article was fact-checked against its sources, and an editor read, edited and approved it before publication.



