Scrape Info

Schema Design for Web Extracted Data

Define your extraction schema before scraping, not after.

Features Editor · · 12 min read
Cover illustration for “Schema Design for Web Extracted Data”
Data Extraction · September 28, 2026 · 12 min read · 2,794 words

Schema Design for Web Extracted Data

Why defining your output schema before extraction is an architectural decision, not a cleanup step

The default workflow most teams follow is to scrape first and figure out the shape of the data later, which produces brittle, inconsistent outputs that break downstream AI pipelines.

The failure mode is predictable. Field names drift ("price" on one run, "cost" on the next), types wobble between strings and numbers, and the entire burden of cleaning up the mess lands on whoever consumes the data, be it an LLM, a database, or a very tired analyst.

The correction, and it's become close to gospel among people building extraction pipelines as of August 2026, is to define a JSON schema before the extraction API is ever called. Not after the first ugly output. Before a single URL gets fetched. Schema-first means the schema functions as a contract: it tells the extractor what to look for, what to ignore, and exactly how to format whatever it finds.

Why does this matter more now than it used to? Because the consumer of this data has changed. A human analyst looking at a spreadsheet can shrug off an inconsistent date format or a stray null. An LLM or an autonomous agent downstream can't do that gracefully, it either fails silently or hallucinates a plausible-looking answer to paper over the gap. There's no shrugging in a JSON parser.

So the stakes aren't abstract. Schema-first design is the decision that determines whether web-extracted data is usable by AI systems at all. Get it wrong, and everything built on top of it inherits the wobble.

How AI extraction differs from CSS/XPath scraping

Traditional scraping worked by pointing at a spot on the page. A CSS selector like div.product-price > span.value finds a price because that price lives at a specific spot in the DOM, until the site redesigns and the selector returns nothing, or worse, returns the wrong thing without so much as a warning.

AI extraction flips the logic. A developer sends a URL and a plain-English objective, and the model reads the page the way a person would, by meaning and layout, not by DOM position, then returns structured output. Research cited in Scientific Reports found AI-driven extraction beating conventional rule-based crawlers by 35% in accuracy and 40% in processing efficiency across benchmark datasets Structured Data AI Search: Schema Markup Guide (2026). That's not a marginal upgrade, that's a different category of tool.

As of mid-2026, three patterns have settled out for how teams actually use this. Some go fully DIY, wiring together a browser automation tool, an LLM, and a validation library like Pydantic by hand.

What changes across all three is the role schema plays. In selector-based scraping, the selector basically was the schema, location implied structure. In AI extraction, the schema has to stand on its own as an explicit contract the model is instructed to satisfy. Because AI extraction runs on objectives rather than coordinates, a vague schema doesn't just produce a bad selector, it produces vague output across the board. The schema now carries all the precision that used to live implicitly in a CSS path.

Raw HTML is not a reasonable input at LLM scale parallel.ai. It averages 5 to 10 times more tokens than the same content rendered as clean markdown parallel.ai. Feeding a model raw HTML means paying, in real dollars, to process cookie banners and footer links that carry zero signal parallel.ai. Clean input isn't a nicety, it's a budget line. AI is offered as a parameter on a managed API, as with ScrapingBee's ai_extract_rules An AI-native OSS framework is exemplified by ScrapeGraphAI and Crawl4AI

The structural anatomy of a well-designed extraction schema

Diagram: Schema-First vs. Scrape-First: Where the Wobble Enters. Visualizes: Visualize the contrast between the default 'scrape-first' workflow and the 'schema-first' workflow as a before/after or two-path flow.

Picture a product schema: name, price, currency, availability, source URL. Simple, copyable, the kind of thing a developer can drop into a new project without reinventing it each time. That's the shape to build toward: not an abstract idea of "structured data" but five fields that actually do work parallel.ai.

Typing discipline is where most of the payoff hides. Dates typed as strings in ISO 8601 format let the model take "June 29, 2026" on one page and "2 days ago" on another (when given the fetch timestamp as a reference point) and normalize both into the same format automatically. Type prices as numbers and currency symbols get stripped in-flight, just keep currency itself in its own separate field. Type availability as a boolean, and you force a binary answer instead of a string like "In Stock" or "Available Now" that some poor script has to parse later.

Field descriptions deserve more respect than they usually get. The description property in a JSON Schema isn't a comment for a future engineer to read, it's an instruction the LLM follows to decide what counts as a valid value. Skyvern's SDK shows this well: describing invoice_date as "the date the invoice was issued (YYYY-MM-DD)" gets the model to format it correctly on the first pass, no separate normalization step required.

Nested objects and arrays earn their place when the data is naturally nested, line items on an invoice, multiple authors on a byline. But push nesting too deep and ambiguity about where one entity ends and another begins becomes visible in the output, a condition that invites hallucination.

A few things to actively avoid. Recursive schemas trip up some providers: Anthropic bans them outright, along with complex enum types, and requires additionalProperties: false on every object; OpenAI also requires that same setting but does support recursive schemas through $ref, as of February 2026. Vague field names like "info" or "details" produce vague output, name a field for exactly what it holds. And schemas that try to grab everything on a page recreate the same noise problem raw HTML causes, a schema should reflect what the downstream system actually consumes, nothing more.

Know that provider-side schema transformation exists rather than discovering it in production.

None of this replaces a clear objective. "Extract the company name, founding year, CEO name, total funding raised, and headquarters city from this page" will consistently beat "get info from this page," because the objective and the schema work together to define the contract. One without the other is half a plan. The Anthropic SDK behavior worth knowing is that it handles schema transformation automatically, stripping unsupported constraints and moving them into field descriptions, and what the provider silently drops can affect output

Why syntactic validity is not enough: semantic correctness and the hallucination problem

A JSON object can be perfectly well-formed and completely wrong. That gap between well-formed and wrong is the one most teams don't check for.

The ScrapeGraphAI-100k dataset, built from a large volume of telemetry, found a high rate of syntactic compliance: outputs that parse cleanly and match the requested schema. That sounds like good news. It measures the wrong thing, though, because parsing cleanly says nothing about whether the values inside are true. A model can produce a tidy, valid JSON object where every value is invented, because the page simply didn't state what was asked for clearly, and the model filled the gap rather than admitting it couldn't find an answer.

Research under the name PARSE gets at why this happens structurally. JSON schemas were originally built as contracts between human developers and static systems. Ambiguous descriptions and unclear entity boundaries are exactly the conditions that trigger hallucination, the model isn't being sloppy, it's doing what an underspecified instruction invites it to do.

The scale of the problem is not trivial. GPT-4 shows a meaningfully high invalid response rate on complex extraction tasks, and PARSE's reflection-based validation approach lifts valid JSON rates substantially over baseline extraction in reinforcement-learning experiments. That difference separates data worth trusting from data worth double-checking.

A handful of design habits close the gap. Write field descriptions as falsifiable instructions, "the price listed in the main product section, excluding shipping" beats "price" every time, because it gives the model a rule to check against rather than a label to guess at. Add a source_text field that stores the verbatim excerpt the model used, that alone lets someone verify an answer without re-fetching the page. Mark fields as required rather than leaving them implicitly optional, because unmarked optional fields tend to vanish silently instead of flagging that something's missing. And for numeric data, run an agentic validation pass after extraction, check that line items actually sum to the stated invoice total, and if they don't, have the system re-evaluate rather than shipping bad data downstream.

There's also a difference between enforcing schema at the prompt level and enforcing it at the API level. Plain JSON prompting (just asking nicely in the prompt) produces shape drift over time. Pydantic schemas combined with structured output enforcement guarantee the shape at the API level instead, and libraries like Instructor in Python handle this by automatically picking the best mode per provider, defaulting to tool calling for OpenAI, with structured output and JSON modes available when needed. Moving enforcement out of the prompt and into the infrastructure is the difference between hoping and knowing.

Designing schemas that hold up across multiple sites and inconsistent layouts

If three retailers are asked whether an item is in stock, the result is three different answers dressed up as one: "In Stock," "Available Now," "Ships in 3-5 days". A schema that doesn't plan for that produces data that looks structured but can't actually be joined in a database, because none of the values match.

This is where AI extraction earns its keep. Giving a field the right type and a clear description lets the model normalize values in-flight rather than requiring a separate cleanup pass afterward, an architectural advantage over selector-based scraping, where normalization was always a bolted-on post-processing step. The fix isn't complicated: list the acceptable normalized values right in the field description ("normalize to one of: in_stock, out_of_stock, preorder, discontinued") and the model effectively becomes a classifier instead of just an extractor.

Missing data needs the same care. There's a real difference between a field that simply isn't present on a page and a field that's present but empty, and downstream systems need to treat those two states differently. The fix is to mark fields nullable: true, keep them required, and add a description telling the model to return null rather than quietly omit the field when the data isn't there.

Then there's the question of one canonical schema across every source versus per-source schemas that map back to a shared model in a transformation layer. It's a real tradeoff: more maintenance overhead against better per-source accuracy, and the right call depends on how many sources are in play and how often they change.

Tables deserve a specific rule: never split one across chunks. Keep it intact and represent it in the schema as an array of objects, with the column names becoming field names.

Tools like Skyvern illustrate the underlying advantage well. Because the model reads pages by meaning and visual position rather than raw HTML structure, extraction logic keeps working even after the underlying DOM shifts around, and cross-site consistency becomes a question you can actually design for up front, rather than a per-site selector maintenance chore.

From single-page extraction to pipeline: how schema design propagates through ingestion, chunking, and storage

Schema decisions don't stay contained to the extraction step. A well-typed schema removes most of the ETL grunt work later on, dates arrive already in ISO 8601, prices arrive as floats, availability arrives as a boolean, and the systems downstream get values they can use immediately rather than values they have to interrogate first. Because extraction-time schema design is, in effect, storage schema design too, a mismatch between the two doesn't just cause friction, it creates ETL debt that compounds every time the pipeline runs.

Chunking for retrieval-augmented pipelines follows its own logic. Fixed-size splits are a fine default, but for technical documents, splitting along heading boundaries first, then sub-splitting long sections, tends to perform better. Tables, again, should never get split across chunks. Code documentation chunks best by function or class rather than by arbitrary length. And every chunk should carry its parent document ID and section heading in its metadata, since that alone allows accurate citation later without re-fetching the original source.

A short list of metadata fields belongs in essentially every extraction schema, regardless of the use case: source URL, fetch timestamp, page title, and schema version. These cost almost nothing to include and they're what make audit and reprocessing possible months later, when someone inevitably asks "where did this number come from?".

Versioning affects whether historical records can be reprocessed correctly after a schema changes. When a target site changes its layout, the extraction schema built for it may need to change too, and versioning that schema means historical records can be reprocessed against the correct version instead of silently getting corrupted by a schema that no longer matches what was originally captured.

The discipline that ties this together: define the exact JSON schema requirements for model inputs before writing any retrieval logic, build rate-limiting and background workers as isolated microservices rather than tangled into the extraction code, and test embedding quality against sample data streams before anything goes to production. Done in that order, schema updates don't derail the rest of the build. Skipping the order turns every schema change into a fire drill.

What schema-first extraction requires from the underlying web data infrastructure

Schema-first design assumes the extraction layer beneath it already strips navigation, ads, cookie banners, and boilerplate before the schema ever gets applied.

Consistency of output format matters just as much as cleanliness. A unified API that returns clean markdown, HTML, or JSON, rather than a patchwork of bespoke adapters for every individual site, is what lets one schema actually work across a pile of unrelated sources. Building ten adapters for ten sites produces ten schemas wearing a trench coat and calling itself one parallel.ai.

Reliability at scale is the third leg. Schema validation failures don't stay small, they compound, and a high per-URL failure rate means the schema-conformant slice of your output is only a fraction of what actually got processed. The token math backs this up too: since raw HTML runs 5 to 10 times heavier in tokens than clean markdown, pre-cleaning pages before schema-guided extraction cuts costs substantially, meaning schema-first design only pays for itself when the layer underneath already eliminates the noise parallel.ai.

Bot detection complicates all of this in a very specific way. If a meaningful share of fetches come back as CAPTCHA pages or bot-blocking screens instead of actual content, the extractor will dutifully try to apply the schema to an error page, and it'll fail in a way that looks like a schema problem when it's actually an access problem. JavaScript rendering, proxy handling, and bot mitigation aren't nice extras, they're prerequisites for getting schema-conformant output at any real volume.

Given all that, most teams face a build-versus-buy decision that isn't close. Standing up enterprise-grade scraping infrastructure from scratch carries real upfront cost and ongoing maintenance that rarely gets easier, so a managed web data API delivering pre-cleaned, structured content through one endpoint tends to be the faster route to reliable output, especially across sources that are both varied and constantly changing. A platform that handles search, scraping, crawling, and monitoring under a single API lines up naturally with schema-first thinking, because when discovery, fetching, and structuring all share one data contract, consistency gets enforced end to end instead of getting renegotiated at every handoff between separate tools.

The strongest pipelines now give the LLM a search_docs tool rather than pre-retrieving everything up front, letting the agent ask for what it needs while the schema still governs exactly what comes back, a pattern worth watching heading further into 2026. Retrieval becomes a conversation instead of a dump truck.

Security and operational risks that schema design must account for

A schema is also a security boundary, whether anyone designed it to be one or not. Every field the schema allows through is a field an LLM downstream will treat as trustworthy, so a loosely bounded schema doesn't just risk messy data, it risks pulling in content an attacker planted specifically to be scraped, a price manipulated to look legitimate, or text crafted to steer an agent's next action. Marking additionalProperties: false, keeping fields narrow, and refusing to let a schema balloon into "grab whatever's on the page" isn't just about clean data, it's about not handing a model an open door dressed up as a JSON object. The tighter the contract, the smaller the surface something malicious has to work with, and that's as much an operational safeguard as it is a data quality one.

Sources

  1. Schema-Based Data Extraction Tools (May 2026)
  2. AI Data Extraction at Scale: Search, Extract & Task APIs
  3. Structured Data AI Search: Schema Markup Guide (2026)
  4. scrapegraphai.com
Filed underData Extraction

More in Data Extraction