โ† Blog Statistics Library Aug 2026 13 min read

AI Generated Content Statistics 2026: How Much and Can We Tell

A data-led look at AI-generated content in 2026: how much of the web is machine-written, whether it ranks, how unreliable detectors are, the dead internet theory and EU labelling law, with every stat cited.

Hero graphic reading AI Content Statistics 2026, How much of the web is machine-written and can detectors tell, in Tom Riley's pink and cream brand style

The web has quietly crossed a line. On the best available measurement, machine-written articles now account for roughly half of everything newly published online, and on some platforms the human share is a small minority.

That raises two questions that this page tries to answer with verified data. First, how much of the internet is actually AI-generated, as opposed to how much people claim it is? Second, can anyone reliably tell the difference, given that the detection tools themselves have a documented accuracy problem?

The answers matter commercially. Google has rewritten its spam policies around scaled content, the EU's labelling rules took legal effect on 2 August 2026, and the ranking data suggests AI-heavy pages perform far worse than their volume implies. Every figure below was checked against the original source.

74.2%of 900,000 new webpages contained some AI content, but only 2.5% were purely AI [1]
50.9%of new online articles were primarily AI-generated in Q4 2025 [2]
86%of articles in top Google results are still human-written [2]
26%of AI text OpenAI's own withdrawn classifier actually caught [9]
61.22%average false-flag rate for non-native English writing across 7 detectors [12]
45%cut in low-quality content in results after Google's March 2024 update [14]

Key AI generated content statistics for 2026

How much of the web is now AI-generated?

Most new pages contain some AI content, but almost none are fully machine-written. Ahrefs found 74.2% of new pages included AI-generated text in April 2025, yet pure AI pages were just 2.5% of the sample. The web is hybrid, not synthetic.

The Ahrefs study ran 900,000 newly created English-language pages, one per domain, through its bot_or_not detector [1]. It found 25.8% of pages were purely human, 71.7% mixed, and 2.5% purely AI.

Within the mixed group, 25.86% showed moderate AI use (11-40% of the text), 20.50% substantial use (41-70%) and 15.51% dominant use (71-99%). The same study surveyed 879 content marketers and found 87% use AI to create or assist with content.

Statistic callout: 74.2% of 900,000 new webpages sampled contained some AI-generated content, but only 2.5% were purely AI [1]

The headline "three quarters of the web is AI" therefore needs care. What the data actually shows is that untouched human prose is now a minority category, whilst fully automated pages remain rare. Most of the shift is people drafting and editing with a model, which is a very different phenomenon from a robot-authored internet.

Has AI-written content overtaken human writing?

For full articles, yes, just. Graphite found primarily AI-generated articles reached parity with human-written ones in early 2025 and edged past them at 50.9% in Q4 2025, before settling at roughly half in Q1 2026.

The Graphite research team analysed 55,400 randomly selected English-language articles from Common Crawl, published between January 2020 and March 2026, and averaged the verdicts of three detectors, Pangram, Copyleaks and GPTZero, rather than trusting any single tool [2].

Statistic callout: 50.9% of new online articles were primarily AI-generated in Q4 2025, overtaking human writing for the first time [2]

The more interesting finding is the plateau. The AI share climbed steeply after ChatGPT's launch, then flattened near 50% through 2025 and early 2026. Graphite's hypothesis is that publishers noticed primarily AI-generated articles "do not perform well in search" and throttled back.

If that is right, the market is partly self-correcting. The incentive to pump out automated articles weakens when they fail to earn traffic, which is exactly what the ranking data below shows.

Far less than its share of the web. Originality.ai puts AI-generated content at 17.31% of top-20 Google results as of September 2025, whilst Graphite finds 86% of top results are human-written.

Originality.ai has tracked 500 informational keywords every two months since January 2019, scanning around 10,000 top-20 results per period with its own detector [3]. The AI share rose from 2.27% in February 2019 to a peak of 19.56% in July 2025, then dipped to 17.31% in September 2025.

The series also shows Google's enforcement leaving a mark. The AI share fell to 7.43% in March 2024, the month Google launched its scaled content crackdown, before resuming its climb to 19.10% by January 2025.

SurfaceAI-generated shareSource and date
New webpages (any AI content)74.2%Ahrefs, April 2025
New webpages (purely AI)2.5%Ahrefs, April 2025
New online articles (primarily AI)50.9%Graphite, Q4 2025
Top-20 Google results17.31%Originality.ai, September 2025
Top Google results (article level)14%Graphite, June 2025
ChatGPT-cited articles18%Graphite, June 2025
Perplexity-cited articles18%Graphite, June 2025
Long-form LinkedIn posts81.2%Originality.ai, July 2026
Bar chart: AI-generated share by surface. Long-form LinkedIn posts 81.2%, New webpages any AI content 74.2%, New online articles primarily AI 50.9%, Top-20 Google results 17.31%. Sources: Ahrefs, Graphite and Originality.ai, 2025 to 2026 [1][2][3][7].

The gap between production and visibility is the single most important pattern in this dataset. Roughly half of new articles are primarily AI-written, yet they hold under a fifth of search visibility. Someone is writing an enormous amount of content that almost nobody sees.

Does AI content actually rank?

It can, but it underperforms badly relative to its volume. Graphite found only 14% of articles in top Google results were AI-generated, falling to 7% among the very top-ranking articles, with ChatGPT and Perplexity citing AI articles just 18% of the time.

The Graphite companion study checked results across 31,493 keywords in ten categories in June 2025, using Surfer's AI detector, and ran parallel tests on ChatGPT and Perplexity citations [4].

Bar chart: AI content, production far outruns visibility. New articles primarily AI 50.9%, Top-20 Google results AI 17.31%, ChatGPT-cited articles AI 18%, Top-ranking articles AI 7%. Sources: Graphite and Originality.ai, June to September 2025 [3][4].

This is not evidence that Google demotes AI text as such. Ahrefs' data shows mixed human-AI pages are the majority of the web, including plenty that rank. What fails is the low-effort, purely automated kind; the AI content that does rank tends to be the edited, hybrid sort that detectors read as human.

The LLMs have their own reason to filter. Graphite's June 2026 follow-up found that in simulations where AI retrieves its own generations, 79.6% of runs end in response collapse [4].

Is the dead internet theory true?

Partly, on the traffic side. Automated traffic overtook human traffic for the first time in a decade in 2024, at 51% of all web traffic, whilst the most famous dead internet statistic, that 90% of content would be synthetic by 2026, cannot be verified in its supposed source.

Imperva's 2025 Bad Bot Report found bots generated 51% of web traffic in 2024, with malicious bots alone at 37% [5]. The report states this is "the first time in a decade" that automation has surpassed human activity.

The 90% claim deserves a burial. It is routinely attributed to a 2022 Europol deepfakes report, but the current version of that report, revised January 2024, contains no such figure at all [6]. A number this popular, resting on a citation that no longer exists, says more about appetite for the theory than about the internet.

So the honest 2026 position is that the dead internet is half true. Traffic is majority non-human and new articles are roughly half machine-written, yet what people actually find via search remains overwhelmingly human-made. The internet is not dead; its long tail is.

How much AI slop is on social media?

On LinkedIn, most of it. Originality.ai classified 81.2% of long-form LinkedIn posts as likely AI in July 2026, and Stanford-Georgetown research documented AI image spam reaching hundreds of millions of Facebook users.

The Originality.ai LinkedIn tracker analysed 5,000 public posts of at least 100 words in July 2026, finding likely-AI content has moved "from roughly half of sampled long-form posts in late 2024 to well over three-quarters today" [7]. LinkedIn is now rolling out a "Seems like AI slop" feedback option.

On Facebook, Harvard's Misinformation Review published a Stanford Internet Observatory and Georgetown CSET study of 125 spam and scam pages, each posting at least 50 AI-generated images [8]. The pages averaged 146,681 followers, their images drew hundreds of millions of exposures, and one AI image post hit 40 million views and 1.9 million interactions, a Facebook top-20 most viewed post in Q3 2023. The feed recommended these unlabelled images to users who did not follow the pages, and comments suggested many viewers had no idea they were synthetic.

Caveat the LinkedIn number properly, though. It relies on one commercial detector, samples only long-form posts, and LinkedIn is the platform where formulaic business prose most resembles AI output even when a human wrote it. The trend line is more trustworthy than the level.

How accurate are AI content detectors?

Not accurate enough to act on alone. OpenAI scrapped its own classifier at a 26% detection rate, and the most comprehensive academic test rated 14 tools "neither accurate nor reliable".

OpenAI's classifier correctly identified just 26% of AI-written text, whilst labelling human writing as AI 9% of the time; the company withdrew it on 20 July 2023 "due to its low rate of accuracy" [9].

Statistic callout: 26% of AI-written text OpenAI's own classifier caught before it was withdrawn for low accuracy [9]

The Weber-Wulff et al. study tested 12 public tools plus Turnitin and PlagiarismCheck, concluding they are "neither accurate nor reliable", skew towards calling text human, and degrade sharply when content is paraphrased or machine translated [10]. Turnitin itself claims a false positive rate below 1%, whilst conceding a residual risk remains [11].

Detection benchmarkResultSource
OpenAI classifier, true positive rate26%OpenAI, 2023
OpenAI classifier, false positive rate9%OpenAI, 2023
14 tools tested on obfuscated AI text"Neither accurate nor reliable"Weber-Wulff et al., 2023
Turnitin claimed false positive rateBelow 1%Turnitin, 2023
Average false flag rate, non-native essays, 7 detectors61.22%Liang et al., Stanford, 2023
Non-native essays flagged by at least one detector97.8%Liang et al., Stanford, 2023

The asymmetry of errors is the real problem. A missed AI text is a nuisance; a human falsely accused of using AI can lose a grade, a job or a byline. Even single-digit false positive rates generate a steady stream of wrongful accusations at institutional scale.

Current commercial detectors are better than the 2023 cohort, and ensembles like Graphite's three-detector average reduce noise. But every detector-derived percentage in this article is an estimate with error bars, not a count.

Are AI detectors biased against non-native English speakers?

Demonstrably, yes. Stanford researchers found seven widely used GPT detectors falsely flagged non-native English writing 61.22% of the time on average, whilst classifying native-written essays near-perfectly.

The Stanford study by Liang et al., published in Patterns, ran 91 human-written TOEFL essays and 88 US eighth-grade essays through seven detectors [12]. All seven unanimously flagged 18 of the 91 TOEFL essays (19.78%) as AI-authored, and 89 of 91 (97.8%) were flagged by at least one detector.

Bar chart: AI detector accuracy and bias. Non-native essays falsely flagged, 7 detectors 61.22%, OpenAI classifier true positive rate 26%, OpenAI classifier false positive rate 9%. Sources: OpenAI 2023 and Liang et al., Stanford, 2023 [9][12].

The mechanism is perplexity. Detectors treat predictable word choices as machine-like, and writers with constrained vocabulary in a second language produce exactly that. When the researchers prompted GPT to rewrite the essays in more varied language, the bias largely vanished, which also means the detectors are trivially easy to evade.

For any organisation running detectors on staff, students or freelancers, this is a discrimination risk, not just an accuracy one. The authors explicitly caution against evaluative use.

What is Google's official stance on AI content, and did the spam updates work?

Google rewards quality "however it is produced", but treats content generated primarily to manipulate rankings as spam, and its March 2024 crackdown cut low-quality content in results by 45%.

Google's February 2023 guidance states that "using automation, including AI, to generate content with the primary purpose of manipulating ranking in search results is a violation of our spam policies", whilst confirming that "appropriate use of AI or automation is not against our guidelines" [13].

The March 2024 core update added a scaled content abuse policy targeting content produced at scale "whether automation, humans or a combination are involved", and Google later reported the changes reduced low-quality, unoriginal content in results by 45%, ahead of its 40% forecast [14].

The policy is deliberately method-agnostic. Google polices whether words were made in bulk to game rankings, not how they were made, which squares with the ranking data: hybrid pages rank routinely, content-farm output does not.

The Originality.ai time series gives enforcement a visible signature, with the AI share of top-20 results dropping to 7.43% around the March 2024 update before recovering. Enforcement works, but it decays.

What are the rules on labelling AI-generated content?

In the EU, mandatory transparency arrived days ago. Article 50 of the EU AI Act applies from 2 August 2026 and requires AI-generated content to be marked in machine-readable form, with deepfakes and AI text on matters of public interest disclosed.

Under Article 50, providers of generative systems must ensure outputs are "marked in a machine-readable format and detectable as artificially generated or manipulated" [15]. Deployers must disclose deepfakes, and AI-generated text published to inform the public must be disclosed unless it underwent human review with editorial responsibility.

That editorial-control exemption is loophole and lifeline in one clause. A newsroom that genuinely reviews AI drafts need not label; a fully automated content farm targeting EU audiences now has a legal labelling duty, not just an SEO problem.

Given the detection figures above, enforcement will be the hard part. Regulators are mandating disclosure of something third parties cannot reliably measure, which is precisely why the Act pushes provenance marking at generation rather than detection after the fact.

How to read these numbers

Every "percentage of the web that is AI" figure is a detector output, not an audit. Ahrefs, Originality.ai, Graphite and Surfer all use proprietary classifiers with different thresholds, which is the main reason credible estimates range from 17% to 74%.

Sampling frames differ wildly. Ahrefs measured new pages, Graphite measured schema-marked articles in Common Crawl, Originality.ai measured 500 informational keywords. None is "the web"; each is a defensible slice of it.

Detector error compounds at scale, and the Stanford findings show it is not evenly distributed across writers. Figures for non-English content are close to non-existent, and heavily edited AI text sits in a grey zone every tool classifies differently.

Finally, some numbers are claims by interested parties. Turnitin's sub-1% false positive rate is a vendor statement, Google's 45% reduction is self-reported, and detection companies benefit from a scarier-sounding AI share.

What this means going into 2027

Volume is no longer the game. With half of new articles machine-written and 74% of pages AI-touched, production capacity is worthless as a differentiator. The scarce assets are original data, expertise and distribution.

Bet on the visibility gap persisting. Google and the LLMs both surface predominantly human or human-edited work, and Graphite's collapse research gives AI platforms a survival reason to keep filtering synthetic text.

Treat detectors as smoke alarms, not verdicts. Use them for portfolio-level triage, never for individual accusations, and assume meaningful error, especially for non-native writers.

Prepare for provenance, not detection. Article 50's machine-readable marking and platform slop flags point to content labelled from birth. Publishers with clear human editorial control hold the strongest position under both Google policy and EU law.

Expect the plateau to be tested. AI article share flatlined near 50% because pure AI content fails commercially. If models close the quality gap, or platforms stop filtering effectively, that equilibrium breaks, and the 2027 numbers will show which way it went. If you want help auditing where your own content sits on the human-to-AI spectrum, you can book a call with me.

Sources

  1. Ahrefs, "74.2% of New Webpages Include AI Content (Study of 900K Pages)", April 2025. https://ahrefs.com/blog/what-percentage-of-new-content-is-ai-generated/
  2. Graphite, "AI Now Writes as Many Online Articles as Humans Do", May 2026. https://graphite.io/five-percent/research/ai-now-writes-as-many-online-articles-as-humans-do
  3. Originality.ai, "AI Content in Google Search Results", September 2025 update. https://originality.ai/ai-content-in-google-search-results
  4. Graphite, "AI Content in Search and LLMs", June 2025. https://graphite.io/five-percent/ai-content-in-search-and-llms
  5. Imperva (a Thales company), "2025 Bad Bot Report", April 2025. https://www.imperva.com/resources/resource-library/reports/2025-bad-bot-report/
  6. Europol Innovation Lab, "Facing Reality? Law Enforcement and the Challenge of Deepfakes", 2022, revised January 2024. https://www.europol.europa.eu/cms/sites/default/files/documents/Europol_Innovation_Lab_Facing_Reality_Law_Enforcement_And_The_Challenge_Of_Deepfakes.pdf
  7. Originality.ai, "Study: AI Content Published on LinkedIn", July 2026. https://originality.ai/blog/ai-content-published-linkedin
  8. Harvard Kennedy School Misinformation Review, DiResta and Goldstein, "How Spammers and Scammers Leverage AI-Generated Images on Facebook for Audience Growth", August 2024. https://misinforeview.hks.harvard.edu/article/how-spammers-and-scammers-leverage-ai-generated-images-on-facebook-for-audience-growth/
  9. OpenAI, "New AI Classifier for Indicating AI-Written Text", January 2023, withdrawn July 2023. https://openai.com/index/new-ai-classifier-for-indicating-ai-written-text/
  10. Weber-Wulff et al., "Testing of Detection Tools for AI-Generated Text", International Journal for Educational Integrity, December 2023. https://edintegrity.biomedcentral.com/articles/10.1007/s40979-023-00146-z
  11. Turnitin, "Understanding False Positives Within Our AI Writing Detection Capabilities", March 2023. https://www.turnitin.com/blog/understanding-false-positives-within-our-ai-writing-detection-capabilities
  12. Liang, Yuksekgonul, Mao, Wu and Zou, "GPT Detectors Are Biased Against Non-Native English Writers", Patterns / arXiv, April 2023. https://arxiv.org/abs/2304.02819
  13. Google Search Central, "Google Search's Guidance About AI-Generated Content", February 2023. https://developers.google.com/search/blog/2023/02/google-search-and-ai-content
  14. Google, "New Ways We're Tackling Spammy, Low-Quality Content on Search", March 2024. https://blog.google/products/search/google-search-update-march-2024/
  15. EU Artificial Intelligence Act, "Article 50: Transparency Obligations for Providers and Deployers of Certain AI Systems", applicable from 2 August 2026. https://artificialintelligenceact.eu/article/50/

Key facts about this post

What this article is about A sourced overview of AI-generated content statistics for 2026, covering how much of the web is machine-written, whether it ranks, detector reliability, the dead internet theory and EU labelling law
Type Statistics / research roundup
Author Tom Riley, AI SEO consultant in London
Key stat 1 74.2% of 900,000 new webpages contained some AI content, but only 2.5% were purely AI (Ahrefs)
Key stat 2 50.9% of new online articles were primarily AI-generated in Q4 2025, overtaking human writing (Graphite)
Key stat 3 Only 17.31% of top-20 Google results were AI-generated, and 86% of top articles are human-written (Originality.ai, Graphite)
Key stat 4 Seven GPT detectors falsely flagged non-native English writing 61.22% of the time (Stanford, Liang et al.)
Key law EU AI Act Article 50 requires machine-readable marking of AI content from 2 August 2026
Sources cited 15 (Ahrefs, Graphite, Originality.ai, Imperva, Europol, Harvard Misinformation Review, OpenAI, Weber-Wulff et al., Turnitin, Stanford, Google, EU AI Act)
Why it matters Production capacity is no longer a differentiator; human editorial control wins visibility in search and AI answers and holds the strongest legal position

Using these stats? Please credit this page with a link back to AI Generated Content Statistics 2026. It keeps research like this free.

About the author: AI SEO consultant

Written by Tom Riley, an AI SEO and AI search consultant in London. He helps brands get recommended by ChatGPT, Google AI Overviews, Perplexity and Gemini, using the same AI SEO playbook he runs on his own site. Read his author profile or connect on LinkedIn.

Want results like this for your brand?

Book a call to discuss your growth and visibility. We'll go through your SEO and AI search together. No pitch deck, no hard sell. Just what I'd do if it were my site.

๐Ÿ‘‰ Book a call with Tom