Scrape Info

Structured Data Markup for AI Search Visibility

Schema markup helps AI systems confidently identify and cite your content.

Reporter · · 10 min read
Cover illustration for “Structured Data Markup for AI Search Visibility”
AI Search & GEO · October 11, 2026 · 10 min read · 2,166 words

An AI system that lands on an unstructured page has to guess. A string of digits might be a phone number, a product ID, or a stray detail in a sentence, and the system has to decide which before using that page. That guessing introduces risk, and a system weighing risk tends to look elsewhere for its answer. Structured data markup removes the guesswork by labeling entities explicitly, so the AI reads a price as a price, a publish date as a publish date, a person as a person with stated credentials, and extracts it without having to interpret anything.

That's the real shift between traditional SEO and AI search. AI search asks a different question: can the system confidently parse a page, verify the claim on it, and attribute that claim to a specific source. Confidence, not keyword density, decides whether a page gets cited.

Schema intervenes at three points in the pipeline that gets a page from the open web to an AI-generated answer. The first is crawling and indexing. AI Overviews and similar systems build on the same crawled index traditional search already uses, so a page that isn't crawlable, or one flagged as low quality, never reaches the AI layer no matter how well it's marked up. The second stage is entity extraction. Once a page is crawled, AI systems break it into passages, convert those passages into embeddings, and hunt for entities, people, organizations, products, concepts, along with their relationships and attributes. Schema hands the system these labels directly. The third stage is relevance and authority scoring, where clarity and structure help decide which sources actually get selected. A page with complete schema presents a cleaner, lower-ambiguity profile than one without it.

None of this works without understanding what modern search actually does. Schema is the mechanism that connects a page's content to the recognizable entities already sitting in an AI system's knowledge graph, which is the foundation everything else in this piece builds on.

How AI search platforms use structured data

Different AI search platforms process structured data in genuinely different ways, and treating them as one audience wastes effort. That split is the single most important fact for anyone planning an implementation strategy.

The same guide lists "overfocusing on structured data" as a common misconception among people doing generative engine optimization. That's a striking admission from the platform where schema is usually assumed to matter most.

On other platforms, schema's influence runs through entity authority and indexing quality rather than through markup being parsed at the moment an answer gets generated.

What controlled experiments found when schema markup was isolated as a variable

Pages with complete schema do show up in AI citations more often than pages without it. But the causal version of that claim, that adding schema directly increases AI brand mentions, does not hold up once researchers isolate schema as a variable and test it directly.

OtterlyAI ran a controlled experiment from December 2025 through March 2026, testing schema markup across seven AI platforms specifically to see whether it moved brand mention counts. Google AI Overviews and AI Mode were the exception, showing significant increases. The Markdown-stripping mechanism flagged in the previous section explains why. A signal that never reaches the model cannot shape what the model says.

A second experiment points the same direction from a different angle. Mark Williams-Cook ran a "Duck Test" in February 2026, placing a fake company address exclusively inside invalid, made-up JSON-LD schema, the kind that shouldn't parse as valid structured data. Williams-Cook was careful to note that this one experiment doesn't conclusively prove large language models ignore schema altogether, but it lines up with the OtterlyAI finding closely enough to take seriously.

Put the two together and the correlation everyone points to starts to look like a confound. The content quality is doing the work of earning the citation. The markup was just riding along next to it.

Why schema still matters after the causal claim fails

Schema's genuine value was never about whispering instructions into an AI's ear at the moment it generates an answer. The value sits earlier, in building the entity authority and machine-readable trust signals that decide whether a page counts as credible source material.

Organization and Person schema carry a property called sameAs, which links an entity to authoritative outside records: Wikidata, LinkedIn, professional registries. That link gives an AI system a verifiable chain of identity. It can confirm who published a piece of content and what their credentials are without trying to extract that information from a bio paragraph written in conversational prose. That verification lowers the system's uncertainty about whether a source is authoritative, and reduced uncertainty is a prerequisite for citation on "your money or your life" queries and other high-stakes topics where getting it wrong carries real consequences.

Properties like knowsAbout, hasCredential, and alumniOf push this further by making expertise itself machine-readable, instead of leaving a system to find and interpret a paragraph that says someone has a medical degree or a decade in a field. A page that sits clearly inside that graph is simply easier to trust and attribute than one floating unconnected to anything the system already recognizes.

John Mueller of Google summarized the mechanism directly: "We do use structured data to better understand the entities on a page and to find out where that page is more relevant." Understanding and relevance mapping, not ranking, and not a guarantee of citation. For teams building AI pipelines that consume web content at scale, there's a quieter benefit too: entity-rich structured content is easier to process, embed, and retrieve accurately, because the same properties that help a search engine understand a page also cut down the noise a scraping or parsing pipeline has to sort through downstream.

Implementing schema types that Google has already retired wastes effort and sets teams up to expect visibility benefits that no longer exist, which happens more often than it should when guidance gets passed around secondhand.

FAQPage markup had already been limited to authoritative government and health sites back in August 2023, before Google pulled it. HowTo rich results were retired on September 13, 2023, on desktop, after mobile had already lost the feature the month before. Sitelinks Search Box was retired on November 21, 2024. And Practice Problem and Dataset rich results within Google Search were announced for retirement on November 5, 2025. None of this means the markup is harmful to leave in place. Google has stated there's no need to remove retired schema: unused markup causes no problems and simply produces no visible effect.

Several schema types remain fully supported and worth prioritizing. And FAQPage markup, while it no longer earns a rich result, still serves a purpose, since AI assistants often draw answers directly from question-and-answer pairs on a page.

Schema.org's vocabulary keeps changing. Schema.org version 30.1 shipped on September 16, 2026, and the standard continues to evolve from there, so treating any list of schema types as permanent is a mistake. Monitor the changelog rather than the memory of a guide written a year earlier.

Implementing schema as a precision layer on content that already earns trust

Schema applied to thin or weak content doesn't rescue it. The sequence that actually works is to build content an AI system would want to cite on its own merits first, then apply markup to cut the ambiguity cost of extracting and attributing it.

Google has said directly that AI Overviews run on the same ranking systems and the same index as traditional search, with a generative layer added on top. A page flagged as low quality in traditional ranking will almost never make it into that generative layer, regardless of how thorough its markup is. Crawlability, indexing, and genuine usefulness come first. Schema comes after.

Google also accepts Microdata and RDFa, for teams already working in those formats. The entity connection matters as much as the format: using sameAs properties to link Organization and Person entities to authoritative outside records is what actually connects a page to the knowledge graph, rather than leaving its claims isolated and unverifiable.

Multiple schema types can and should stack on a single page. A product page, for instance, can carry Product, Review, and BreadcrumbList schema all at once, each one signaling a different dimension of the page's content and its entity relationships. Required properties make a page eligible for rich results, but recommended properties are what build out the fuller entity profile, and teams that stop at the bare minimum leave that signal incomplete.

Click-through rate changes by page type are the measurable near-term signal to track. AI citation frequency is real but far harder to attribute to any single cause.

Pages with clean, complete schema are structurally easier to scrape, parse, and convert into AI-ready formats, because the same entity labels that help Google also cut down the ambiguity a downstream pipeline has to resolve when turning a page into structured JSON or Markdown for a language model. A web data source that delivers pre-cleaned, entity-rich content removes the need to write a custom parser for every schema variation a team runs into at scale.

What the llms.txt standard and agent-ready infrastructure add

The logic behind schema markup, reduce the ambiguity cost for a machine reader, is now showing up in a newer layer of infrastructure built specifically for AI agents. The most talked-about piece of that layer serves a narrower job than its boosters suggest.

llms.txt is a proposed standard that Jeremy Howard of Answer.AI introduced in September 2024. It's a Markdown file placed at the root of a domain, meant to give AI systems a curated index of a site's most important content. The adoption numbers, though, are blunt. An Ahrefs study of 137,000 domains found that among those carrying a valid llms.txt file, a large majority received zero requests for it during the study month. The standard's real use case looks like Business-to-Agent work, developer tooling, and documentation sites.

Cloudflare's "Agents Week" research from April 2026 found a related wrinkle. AI crawlers visiting Cloudflare's own developer documentation consumed deprecated content at the same rate as current content, and canonical tags, the standard signal built for exactly this problem, made no measurable difference. Standard web signals built for human-facing search don't reliably govern how AI crawlers behave. Cloudflare responded by launching a toggle on paid plans that lets sites enforce content hierarchy for crawlers directly, and it launched a checker site, isitagentready.com, that evaluates whether a domain has a robots.txt file, has an llms.txt file, can deliver content in Markdown on request, and meets other markers of agent-readiness.

The same principle underlies schema markup. Machine readers, whether they're search crawlers, AI agents, or coding assistants working inside an IDE, perform better when a site explicitly declares its own structure, hierarchy, and the relationships between its pieces. For documentation sites, developer tools, and API products specifically, pairing llms.txt with schema markup covers a use case schema.org markup alone doesn't reach: an agent navigating a codebase or a docs site needs a curated index to work from, not just entity labels scattered across individual pages.

Structured data as a precision layer on genuine content authority

The pages that benefit most from structured data markup are, almost without exception, the pages that would have earned a citation anyway. Schema is the mechanism that makes that existing worth legible to a machine, at the speed and confidence level AI systems need to act on it.

Lay the evidence from this piece side by side and a clear picture forms. Schema does not directly move AI citation counts on platforms that strip markup out before inference happens. Schema does build entity authority, cut extraction ambiguity, and support the knowledge graph connections that shape whether a source gets treated as credible. And schema that's technically valid but sits on top of thin, vague, or untrustworthy content signals nothing useful at all, because the AI system's quality filter operates upstream of the markup layer entirely, at the level of the content itself.

The practical move for content and engineering teams is to audit whether a given page would earn a citation on the strength of its content alone, before spending time on markup. If the honest answer is no, the priority is fixing the content, not adding schema on top of it.

For teams building AI products and retrieval pipelines, the pages a system pulls from and cites inherit the credibility of the web those pages come from. Feeding a language model or a retrieval system with content that's entity-rich, properly marked up, and genuinely authoritative lowers the risk of hallucination and improves how accurately the system attributes what it says. That's where content quality and structured markup meet and become retrieval quality. Structured data earns its place as a precision layer on top of content that already deserves to be found, read, and trusted. It was never built to manufacture authority that isn't there to begin with.

Filed underAI Search & GEO

More in AI Search & GEO