โ† Blog Statistics Library Aug 2026 13 min read

Semantic SEO & Entity Statistics 2026: What the Data Says About Markup, Knowledge Graphs and AI Visibility

A data-led look at semantic SEO and entities in 2026: how much of the web uses structured data, which schema types are deployed, the size of Google's Knowledge Graph, Wikipedia and Wikidata, branded search, AI-visibility correlations and the Google leak, with every stat cited.

Hero graphic reading Semantic SEO & Entity Stats 2026, half the web carries markup, nearly half of Google searches are branded, every figure cited, in Tom Riley's pink and cream brand style

Semantic SEO used to be an argument about vocabulary. Whether you tagged your pages with schema.org, whether Google understood "things not strings", whether a Wikipedia page was worth chasing.

In 2026 it is an argument about supply. Large language models answer questions by assembling facts about entities, and the places they get those facts from are countable: structured data on the open web, Wikipedia and Wikidata, and whatever a system already believes about a brand from its wider footprint.

So the numbers below are worth more than the theory. Roughly half the web now carries some machine-readable markup, nearly half of Google searches contain a brand name, and the correlation studies that exist point at prominence rather than markup. Where the evidence is thin, I have said so, because a lot of what passes for entity SEO advice is community inference dressed up as fact.

50%of home pages carried structured data in the July 2025 crawl, per the HTTP Archive Web Almanac [1]
45.7%of Google searches are branded, across a study of ~150 million US keywords by Ahrefs [10]
73.99bntriples yielded by the October 2024 Web Data Commons extraction [3][4]
76.10%of AI Overview citations come from pages already ranking in Google's top 10 (Ahrefs) [11]
0.70Spearman correlation for AI Overviews between mentions on heavily linked pages and visibility, against 0.12 for ChatGPT (Ahrefs) [12]
122,786,550items in Wikidata, built from over 2.52 billion edits, as of 7 August 2026 [6]

Key semantic SEO and entity statistics for 2026

How much of the web actually uses structured data?

About half of home pages, and a little under half of all domains. Adoption is high but has stopped growing quickly, and the growth that remains is mostly JSON-LD replacing older formats rather than new sites adopting markup for the first time.

The HTTP Archive Web Almanac, based on a July 2025 crawl of 16,213,084 websites, found structured data on 50% of home pages on both desktop and mobile, against 48% and 49% in 2024 [1][2]. JSON-LD sits at 43% of home pages, up from 41% desktop and 40% mobile. Microdata fell by a percentage point to 17% desktop and 16% mobile. Microformats2 is a rounding error at 0.14%.

Statistic callout: 50% of home pages carried structured data in the July 2025 crawl (HTTP Archive Web Almanac) [1]

Web Data Commons, which extracts markup from Common Crawl, tells a compatible story at larger scale. Its October 2024 corpus found structured data on 16,525,070 of 37,447,141 pay-level domains and on 1,245,622,627 of 2,391,039,772 parsed URLs, producing 73,993,669,093 triples [3][4].

The interesting split is by format weight, not just by site count.

FormatDomains (WDC, Oct 2024)TriplesHome page use (HTTP Archive, 2025, desktop)
JSON-LD11,562,35947,979,634,59743%
Microdata7,599,79221,825,355,45017%
Microformats3,940,0873,740,945,8860.14% (Microformats2)
RDFa474,635458,151,723not separately reported
Bar chart: extracted triples by format (Web Data Commons, October 2024). JSON-LD 47,979,634,597, Microdata 21,825,355,450, Microformats 3,740,945,886, RDFa 458,151,723. Source: Web Data Commons, October 2024 corpus [3][4].

JSON-LD carries roughly 65% of all extracted triples despite being the newest format, so the markup machines actually read is concentrated in the format Google recommends. But the legacy tail is heavier than it looks: Microdata still sits on 7.6 million domains, mostly old CMS templates nobody has revisited.

Which schema types are actually deployed?

Mostly the boring ones. The web's most common markup describes the website itself rather than the products, people or organisations that AI systems are trying to reason about.

On home pages, the Web Almanac found schema.org/WebSite on 37% of desktop pages, SearchAction on 28%, Organization on 26.74%, WebPage on 25% and an unresolved "UnknownType" on 24% [1]. On inner pages, ListItem leads at 29% desktop, followed by BreadcrumbList at 28%, WebSite at 26% and Organization at 26%.

Schema typeDesktop home pagesDesktop inner pages
WebSite37%26%
SearchAction28%17%
Organization26.74%26%
WebPage25%25%
ListItem21%29%
BreadcrumbList21%28%

Note the top two. WebSite and SearchAction exist to power Google's sitelinks search box, which was visually retired in November 2024 and yet still sits on more than a quarter of the web because nobody removes markup once a plugin has written it [1].

That is the honest state of semantic SEO adoption. A large share of the world's structured data is inert boilerplate, not a considered description of an entity.

The Almanac reports that Microsoft has confirmed Bing uses schema.org markup to help its models distinguish expert articles, products, reviews and FAQs [1]. That is a statement of use, not a measured lift. I have found no credible 2025 or 2026 study isolating schema markup as a cause of higher AI citation rates, and anyone claiming to have measured it should be asked for the control group.

How big is Google's Knowledge Graph?

Google's last public figure is over 500 billion facts about 5 billion entities, from a Google blog post published in May 2020 [5]. Six years on, it is still the number everyone quotes.

Treat it with suspicion. It is a marketing figure rather than a released dataset, it predates the generative search era entirely, and Google has not refreshed it despite the Knowledge Graph now feeding AI Overviews and Gemini grounding. Anyone citing it in 2026 as current is repeating a six-year-old press line.

How much of AI's entity layer comes from Wikipedia and Wikidata?

Enough that both projects are now feeling the load. English Wikipedia holds 7,220,552 articles and Wikidata holds 122,786,550 items built from more than 2.52 billion edits, checked live on 7 August 2026 [6][7].

Statistic callout: 73.99 billion triples extracted from 16,525,070 of 37,447,141 domains carrying structured data (Web Data Commons) [3][4]

The Wikimedia Foundation reported in April 2025 that 65% of its most expensive traffic comes from bots, even though bots make up only around 35% of pageviews, and that bandwidth for multimedia downloads had grown 50% since January 2024 [8]. Scrapers hit obscure pages that miss the edge cache, so they cost disproportionately.

In October 2025 the Foundation reported that human pageviews for May and June 2025 were down roughly 8% year on year after reclassifying bot traffic, and attributed it to search engines answering directly with generative AI and to people encountering Wikipedia knowledge second-hand [9]. Its own phrasing is blunt: almost all large language models train on Wikipedia datasets.

This is the clearest structural fact in entity SEO. A licence-free, heavily edited corpus of entity descriptions sits underneath most AI answers, and its traffic is falling whilst its influence rises. Getting your organisation accurately represented in Wikidata is cheap and legitimate. Buying a Wikipedia article is neither.

How much of Google search is branded?

Nearly half. Ahrefs studied roughly 150 million US keywords and found that 36.9% of unique queries are branded, but 45.7% of actual searches by volume are, because branded terms are searched more often [10].

Statistic callout: 45.7% of Google searches are branded, across a study of roughly 150 million US keywords (Ahrefs) [10]

That gap between 36.9% and 45.7% is the whole point. Brand demand is concentrated in a smaller set of very high-volume terms, which is exactly why keyword-level averages understate it.

It also reframes what a "keyword strategy" is worth. If close to half of search volume is people asking for something they already know the name of, then the marketing that creates that recall sits upstream of anything an SEO does on a category page. Ahrefs also found that queries of three or more words make up the largest slice of branded search, so branded demand is not just people typing "amazon", it is people doing comparison research with a brand already in mind.

Do brand and entity signals correlate with AI visibility?

They correlate, unevenly, and the strength depends entirely on which system you look at. This is the most useful body of evidence available and the most commonly over-read.

Ahrefs ran two correlation studies in mid-2025 across roughly 76.7 million AI Overviews, 957,000 ChatGPT prompts and 953,500 Perplexity prompts, looking at the top 50 most-mentioned domains in each system [12][13].

SignalAI OverviewsChatGPTPerplexity
Mentions on highly linked pages (Spearman ฯ)0.700.120.40
Organic search traffic vs mention share (Spearman ฯ)0.470.330.66
Bar chart: mentions on highly linked pages vs AI visibility, Spearman rho. AI Overviews 0.70, Perplexity 0.40, ChatGPT 0.12. Source: Ahrefs correlation study, mid-2025 [12].

Google AI Overviews reward the same prominence traditional search rewards, which is unsurprising given they draw on the same index. Perplexity tracks organic traffic most closely at 0.66. ChatGPT is the outlier on both measures, and Ahrefs suggests licensing partnerships explain part of that weakness [12].

The sample is the catch. Fifty domains per system is a small n, the p-value for the ChatGPT traffic correlation was 0.0515, which is not significant at the conventional threshold, and rank correlations among the fifty biggest sites on the internet tell you little about a mid-market brand [13].

A separate Ahrefs study of 1.9 million citations from 1 million AI Overviews found 76.10% of cited pages rank in Google's top 10, 9.50% rank between 11 and 100, and 14.40% do not rank in the top 100 at all [11]. Classic ranking still buys most of the AI citation supply on Google's own surface.

Bar chart: where AI Overview citations rank in Google. Rank in top 10 76.10%, rank between 11 and 100 9.50%, do not rank in the top 100 14.40%. Source: Ahrefs, 1.9 million citations from 1 million AI Overviews [11].

And the systems disagree with each other. Only 7 of the top 50 most-mentioned domains, 14%, were shared across ChatGPT, Perplexity and AI Overviews [14].

What did the Google API leak really show about site authority?

It showed that a stored feature named siteAuthority exists. It did not show how it is calculated or weighted, and that distinction is where most of the commentary went wrong.

iPullRank documented the March 2024 Content Warehouse documentation as 2,596 modules containing 14,014 attributes, and identified siteAuthority inside the Compressed Quality Signals stored per document, alongside click-based features such as goodClicks, badClicks and lastLongestClicks [15]. SparkToro, which received the documents, described the same 14,014 attributes and Chrome-derived click data such as chrome_trans_clicks [16].

Both authors put the same caveat in writing: the documents contain no scoring functions, so nobody outside Google knows whether a given feature is live, deprecated or weighted at zero.

That matters because the leak became the evidentiary basis for a lot of 2025 and 2026 brand-signal advice. The defensible claim is that Google stores a sitewide authority value, which contradicts years of public denials. The indefensible claim is that you can reverse-engineer it. The documents are also now over two years old.

Is topical authority a measurable thing?

Not as a published, third-party-verifiable metric. Topical authority is a useful planning heuristic with almost no public measurement behind it, and it should be labelled that way.

What is measurable is prominence, and prominence correlates with AI visibility in the Ahrefs data above. What is not measurable from outside Google is whether a site holds a per-topic authority score: no leaked attribute has been shown to be one, and no vendor has published a replicable methodology.

The practical version survives the scepticism. Covering a subject completely, marking up the organisation behind it, and being mentioned in places other people link to are all things the evidence supports. The pseudo-precise version, scoring your "topical authority" out of 100 and reporting it to a board, is theory with a dashboard on top.

Are site owners blocking the crawlers that build entity understanding?

A growing minority are, and it is concentrated among exactly the sites AI systems most want to cite. Blocking is still rare across the web as a whole.

The Web Almanac's generative AI chapter found that 94.1% of the roughly 12.9 million sites analysed serve a robots.txt file with at least one directive, and that gptbot is now the second most common user-agent named in those files, rising from 2.6% of robots.txt files in 2024 to around 4.5% in 2025 [17]. Directives aimed at AI bots are far more common on popular websites, and nearly all of them disallow access.

Bar chart: share of robots.txt files naming gptbot. 2024 2.6%, 2025 around 4.5%. Source: HTTP Archive Web Almanac generative AI chapter [17].

The emerging alternative, llms.txt, is barely used. Valid files appear on 2.13% of desktop sites and 2.10% of mobile sites, and 39.6% of those were generated by All in One SEO rather than written deliberately [1].

That combination is awkward for anyone selling llms.txt as an entity strategy. The file that would describe your organisation to a model is mostly being written by a plugin, whilst the sites with the strongest entity data are quietly closing the door.

How to read these numbers

Structured data adoption figures come from crawls, not from surveys, so they are observations rather than estimates. But HTTP Archive tests one page per site from a datacentre with an empty cache, and Web Data Commons parses Common Crawl, so neither sees logged-in states, JavaScript-injected markup that fails to render, or the whole of a large site. The Almanac notes only about 2% of crawls had structured data added via JavaScript [1].

Correlation studies here are small-sample rank correlations over the top 50 domains per system. They describe which giant websites get mentioned, not what would happen if your site added markup or earned a mention. One of the reported p-values, 0.0515 for ChatGPT and organic traffic, sits just outside significance [13].

Live counts move. The Wikipedia and Wikidata figures were read on 7 August 2026 and will be higher by the time you read them. Google's 500 billion facts figure is the opposite problem: it has not moved since 2020 because Google has not restated it.

Finally, keep provenance straight. Ahrefs, Semrush and Profound-style vendors measure AI systems using their own prompt panels, which are not random samples of what real people ask.

What this means going into 2027

Structured data adoption is plateauing near 50%, so markup is now table stakes rather than an edge. The differentiator is whether your Organization and Product data is accurate and complete, not whether it exists.

Expect scrutiny of the Wikipedia and Wikidata layer to increase. A knowledge base with declining human traffic and rising machine consumption is a governance problem, and the Foundation has started saying so publicly.

Branded search at 45.7% of volume means brand marketing and SEO budgets are increasingly measuring the same demand. Reporting that separates them will keep understating both.

Treat the Google leak as history from here. It is a March 2024 snapshot, two AI Overview generations behind, and its value now is confirming that sitewide authority exists rather than telling you how to earn it.

And demand better evidence on schema and AI citations. It is the biggest open question in this dataset: plenty of assertion, no controlled study. If you want help working out how findable your business really is in search and AI answers, you can book a call with me.

Sources

  1. HTTP Archive, "SEO", Web Almanac 2025, November 2025. https://almanac.httparchive.org/en/2025/seo
  2. HTTP Archive, "Methodology", Web Almanac 2025, November 2025. https://almanac.httparchive.org/en/2025/methodology
  3. Web Data Commons, "Structured Data extraction statistics, October 2024 corpus", 2025. https://webdatacommons.org/structureddata/
  4. Web Data Commons, "Detailed statistics for the October 2024 corpus", 2025. https://webdatacommons.org/structureddata/2024-12/stats/stats.html
  5. Google, "About Knowledge Graph and knowledge panels", May 2020. https://blog.google/products/search/about-knowledge-graph-and-knowledge-panels/
  6. Wikidata, "Wikidata:Statistics", accessed August 2026. https://www.wikidata.org/wiki/Wikidata:Statistics
  7. Wikipedia, "Special:Statistics", English Wikipedia, accessed August 2026. https://en.wikipedia.org/wiki/Special:Statistics
  8. Wikimedia Foundation, "How crawlers impact the operations of the Wikimedia projects", April 2025. https://diff.wikimedia.org/2025/04/01/how-crawlers-impact-the-operations-of-the-wikimedia-projects/
  9. Wikimedia Foundation, "New user trends on Wikipedia", October 2025. https://diff.wikimedia.org/2025/10/17/new-user-trends-on-wikipedia/
  10. Ahrefs, "Almost Half of Google Searches Are Branded. Here's Why That Matters", May 2025. https://ahrefs.com/blog/almost-half-of-google-searches-are-branded-study/
  11. Ahrefs, "76% of AI Overview Citations Pull From the Top 10", July 2025. https://ahrefs.com/blog/search-rankings-ai-citations/
  12. Ahrefs, "Does Being Mentioned on Highly Linked Pages Influence AI Mentions?", July 2025. https://ahrefs.com/blog/does-being-mentioned-on-highly-linked-pages-influence-ai-mentions/
  13. Ahrefs, "Websites With More Organic Search Traffic Get Mentioned More in AI Search", June 2025. https://ahrefs.com/blog/websites-with-more-traffic-have-more-mentions/
  14. Ahrefs, "86% of Top Mentioned Sources Are Not Shared Across ChatGPT, Perplexity, and AI Overviews", June 2025. https://ahrefs.com/blog/top-mentioned-sources-are-not-shared-across-ai-assistants/
  15. iPullRank, "Secrets From the Algorithm: Google Search's Internal Engineering Documentation Has Leaked", May 2024. https://ipullrank.com/google-algo-leak
  16. SparkToro, "An Anonymous Source Shared Thousands of Leaked Google Search API Documents with Me", May 2024. https://sparktoro.com/blog/an-anonymous-source-shared-thousands-of-leaked-google-search-api-documents-with-me-everyone-in-seo-should-see-them/
  17. HTTP Archive, "Generative AI", Web Almanac 2025, November 2025. https://almanac.httparchive.org/en/2025/generative-ai

Key facts about this post

What this article is about A sourced overview of semantic SEO and entity statistics for 2026, covering structured data adoption, which schema types are deployed, the size of Google's Knowledge Graph, Wikipedia and Wikidata, branded search, AI-visibility correlations, the Google API leak, topical authority and AI crawler blocking
Type Statistics / research roundup
Author Tom Riley, AI SEO consultant in London
Key stat 1 50% of home pages carried structured data in the July 2025 crawl (HTTP Archive Web Almanac)
Key stat 2 45.7% of Google searches are branded, across a study of roughly 150 million US keywords (Ahrefs)
Key stat 3 76.10% of AI Overview citations come from pages already ranking in Google's top 10 (Ahrefs)
Key stat 4 Mentions on heavily linked pages correlate with AI visibility at Spearman 0.70 for AI Overviews but only 0.12 for ChatGPT (Ahrefs)
Knowledge base finding Wikidata contains 122,786,550 items and English Wikipedia holds 7,220,552 articles, checked live on 7 August 2026 (Wikidata, Wikipedia)
Sources cited 17 (HTTP Archive, Web Data Commons, Google, Wikidata, Wikipedia, Wikimedia Foundation, Ahrefs, iPullRank, SparkToro)
Why it matters Semantic SEO in 2026 is about supply: structured data is table stakes, prominence rather than markup correlates with AI visibility, and Wikipedia and Wikidata sit underneath most AI answers

Using these stats? Please credit this page with a link back to Semantic SEO & Entity Statistics 2026. It keeps research like this free.

About the author: AI SEO consultant

Written by Tom Riley, an AI SEO and AI search consultant in London. He helps brands get recommended by ChatGPT, Google AI Overviews, Perplexity and Gemini, using the same AI SEO playbook he runs on his own site. Read his author profile or connect on LinkedIn.

Want results like this for your brand?

Book a call to discuss your growth and visibility. We'll go through your SEO and AI search together. No pitch deck, no hard sell. Just what I'd do if it were my site.

๐Ÿ‘‰ Book a call with Tom