Scrape Info

Entity and Relationship Extraction From Unstructured Web Text

AI systems hallucinate on messy web data without proper extraction pipelines.

Features Editor · · 11 min read
Cover illustration for “Entity and Relationship Extraction From Unstructured Web Text”
Data Extraction · September 30, 2026 · 11 min read · 2,549 words

Most of the internet is not written for machines, and that single fact explains why so much AI output still feels half-baked. IDC's figures, cited in KlearStack's 2026 guide, put unstructured data at over 80% of global data, including emails, PDFs, contracts, social media posts, and business reports, none of it sitting in tidy rows and columns. Compare that to structured data, which any SQL query or BI dashboard can chew through in seconds. Unstructured text has to be extracted, cleaned, and rebuilt before any system can act on it at all.

The web makes this worse, not better. Pages get written for human eyeballs, complete with dynamic content, inconsistent layouts, navigation menus, and scripts that have nothing to do with the actual information a reader wants. None of that structure was ever meant for machines to parse, which is a bit like designing a restaurant menu for someone who can't read and expecting them to order correctly anyway.

AI agents and large language models have turned this from an annoyance into an urgent problem. They're now first-class consumers of raw web content, and when they're fed noisy, unstructured input, they hallucinate rather than merely performing worse. That's a substantial quality drop. That's a system inventing facts because the input handed to it was garbage.

The thesis this entire piece rests on is that the quality of a final structured output, whether that's a triple, a JSON record, or a node in a knowledge graph, gets determined entirely by the quality of every stage that came before it. Each section from here forward walks through one layer of that pipeline, and each layer's failure mode is exactly what sets up the next section's fix.

The Pipeline: From Raw Text to Typed Relational Triples

The whole point of this process is to produce a structured triple, a clean unit shaped as (Entity 1, Relationship, Entity 2), per LlamaIndex's glossary definition. Think (Elon Musk, CEO of, Tesla), or (Google, acquired, DeepMind). Simple on paper, and that simplicity is exactly why it scales.

Key-value extraction and relation extraction get confused constantly, but they are separate things. KV extraction grabs explicit fields, like an invoice total or a due date. Relation extraction is a different animal entirely: implicit semantic connections between named entities appear in the text itself, in stuff that isn't sitting in a labeled field anywhere on the page. Parsing and extraction get conflated too, and they shouldn't be. Parsing represents content faithfully. Extraction converts that content into discrete, machine-actionable facts, which is a much higher bar to clear.

Ingestion comes first: gathering raw data from file systems, APIs, email servers, or web scraping. Then preprocessing and cleaning, which covers tokenization, normalization, OCR, and stripping out noise like HTML tags, stray whitespace, and headers. Extraction follows, applying NLP, ML, or LLMs to identify entities, relationships, and sentiment. After that comes validation, checking the extracted data for accuracy, completeness, and whether it actually conforms to the expected schema. Integration closes the loop, moving validated data into wherever it's meant to live, a database, a graph store, an ERP system.

That's the structural argument the rest of this piece keeps coming back to. Musk" don't become two separate people), then information processing on top of that. Extraction is where this whole story actually begins, and it's a lot more nuanced than it looks from the outside.

Diagram: Five Stages of the Extraction Pipeline. Visualizes: Visualize the five sequential stages of the structured data extraction pipeline described in the article: (1) Ingestion — gathering raw data from file systems, APIs, email servers, or web…

Named Entity Recognition: the foundation layer where pipeline errors are born

Named Entity Recognition (NER) identifies and labels entities (such as a person's name, a company, or a drug).

If NER misses an entity, or slaps the wrong label on it, everything downstream inherits that mistake. Relation extraction either skips a real connection entirely or classifies it against the wrong boundary. Garbage in, garbage stacked on top of garbage out.

Rule-based systems, things like spaCy rule matchers, GATE, or hand-built regex pipelines, hit high precision in narrow, formulaic domains. Feed them anything with real linguistic variety and they snap like a dry twig. Supervised ML models are more flexible, but they need large labeled datasets and careful feature engineering, and they still miss a lot of contextual nuance. Deep learning and transformer-based models, BERT, RoBERTa, SpanBERT, plus relation-focused tools like REBEL and OpenNRE, handle long-range dependencies and context far better, at the cost of heavier compute and less interpretability.

Hybrid setups try to get the best of both worlds, using rules to constrain what a model is allowed to output, trimming down the error rate that pure ML approaches tend to produce. And then there's LLM-based NER, which tends to outperform older fine-tuned models specifically on messy, semi-structured, free-text data, because it can actually read contextual nuance instead of pattern-matching against it. That matters most when document formats vary wildly across sources, which, on the open web, is basically always.

Once the entities are pinned down and labeled, the next question is how the system figures out what connects them. That architecture choice carries consequences far bigger than it sounds. LlamaIndex outlines three methodological approaches to named entity recognition, each with honest trade-offs. The knowledge graph construction pipeline source lists spaCy and GLiNER as part of a typical open-source extraction stack for knowledge graph construction pipelines.

Pipeline vs. joint extraction: how the architecture choice shapes accuracy and throughput

Diagram: Pipeline vs. Joint Extraction: Accuracy Trade-off. Visualizes: Contrast two relation extraction architectures side by side.

A 2024 ACM survey by Zhao and colleagues lays out two canonical architectures for relation extraction, and the difference between them is not a minor implementation detail. Pipeline-based RE splits the job into two separate stages: first find the entities, then figure out how they relate. Take the sentence "ChatGPT is a chatbot launched by OpenAI." A pipeline model extracts "ChatGPT" and "OpenAI" first, then separately predicts that the relation between them is "product". Joint RE does both at once, letting entity detection and relation detection inform each other in real time, which turns out to matter a lot when sentences get messier.

Overlapping relations are where pipeline models fall apart. According to the Tongji and Tsinghua joint extraction paper published in Frontiers in Neurorobotics in 2022, when a single entity shows up in multiple relationships within the same sentence, sequential two-stage models frequently miss one of them or assign it to the wrong pair. A relation-oriented model with global context, proposed by Han and colleagues that same year, tackles this head-on: it frames the task relation-first and encodes global context, purpose-built for joint entity-relation extraction feeding directly into knowledge graph construction.

Relation extraction gets called "the most overlooked NLP technique" in practice, and that's a fair label. Teams pour resources into NER tooling, benchmark it, tune it, and then treat relation classification as an afterthought, even though relation classification decides whether the output is worth anything. Zero-shot extraction is a third path that needs nothing but a prompt, with no training data required. Useful for brand-new domains where labeled data doesn't exist yet, though it carries a real hallucination risk that only gets caught later at the validation stage.

Academic architecture comparisons are useful, but they only show what's possible in a lab setting. What a production system looks like at real scale is a different story.

A Production Pipeline at Scale: The Political Elite Network Case

Solovev and Lasser, working out of the University of Graz's IDea_Lab, published a system on arXiv (arXiv:2606.27347v3) on 24 July 2026: a multilingual joint entity-relation extraction pipeline built to map political elite networks across Europe, pulled straight out of massive unstructured news corpora. The scale is large. It's built to process multi-million-article corpora across several languages at once, doing the work that historically required large coding teams manually working through biographical sources one person at a time.

A few architecture choices here are worth any production team's attention. The pipeline combines span-based NER with a three-stage entity linking cascade that resolves mentions down to language-independent Wikidata QID identifiers, enabling cross-lingual joins with structured datasets. Relation extraction runs against a SKOS ontology covering 109 entity types and 99 relationship types, structuring the output as cross-national, multiplex, signed networks Mapping Political-Elite Networks in Europe with a Multilingual Joint…. That ontology sits separate from the pipeline code itself, so other researchers can swap in their own taxonomy without touching the extraction logic Mapping Political-Elite Networks in Europe with a Multilingual Joint….

The model choice is notable too: Qwen3.6-35B-A3B-FP8, a mixture-of-experts model served through vLLM, structured with DSPy and typed Pydantic signatures, running in non-thinking instruct mode because throughput, not reasoning depth, is the design constraint that matters most at this scale. Guided decoding enforces the ontology constraints directly on the model's output, so the system can't wander outside the 109 defined entity types even if it wanted to Mapping Political-Elite Networks in Europe with a Multilingual Joint….

Against a gold standard of 3,491 hand-verified relations, textual correctness landed between 68.2% under a strict match and 93.7% under a lenient one Mapping Political-Elite Networks in Europe with a Multilingual Joint…. Two case studies back the pipeline against independent public records.

One more detail matters as much as the accuracy numbers: the whole system is fully open-weight, with no proprietary API dependency. That makes it replicable across institutions. That directly answers one of the sharpest critiques of LLM-based relation extraction, which is that when a vendor quietly updates its API, results change and nobody can reproduce last year's study. This pipeline produces triples at scale, but a triple by itself is only half the story. What matters just as much is what format it takes and what metadata rides along with it.

The internal data unit that determines whether extracted triples are usable downstream

A bare (subject, predicate, object) triple tells a reader what got extracted. It says nothing about whether to trust it, or where it came from, which turns out to be most of the ballgame. A production-grade triple needs to carry more: subject, predicate, object, a confidence score, a source document ID with character offsets, and an extracted_at timestamp. Per the knowledge graph construction pipeline source, these fields are "not optional extras".

Each one earns its place for a specific reason. The confidence score lets a team filter out low-quality extractions before they quietly corrupt an entire graph. Source document ID and character offsets make the extraction auditable, meaning anyone can trace a triple straight back to the exact sentence that produced it. The timestamp matters most for temporal knowledge graphs, where knowing when a relationship was true is often as important as knowing that it was true.

Format matters enormously at the boundary where web content meets the model. Raw HTML gets described, accurately, as "toxic" to LLMs: token-heavy, stuffed with scripts and styling and navigation junk that distracts the model and drives up API costs for no benefit. Converting that mess into clean Markdown or JSON-LD needs to happen before a triple ever enters a model, not after. A 2025 NEXT-EVAL benchmark study, cited on Firecrawl's blog, found LLMs hitting F1 scores above 0.95 on structured web extraction, but only when the input was properly formatted going in. The extraction layer, not the model itself, is the actual bottleneck now.

Pydantic does the validation work here, checking an LLM's output against a required JSON schema before that output touches anything downstream, which blocks hallucinated fields from ever propagating into the graph. It's a small gatekeeper doing a large job. And the quality of everything that reaches this gate is shaped, almost entirely, by the web retrieval layer that fed it in the first place, which loops the whole conversation back to where any pipeline actually starts.

Why the web retrieval layer determines everything that follows

Web scraping gets called the load-bearing layer of any agent that touches live sites, and when that layer buckles, the rest of the agent goes down with it. There are four points where an agent actually touches the web: fetching, observation, sessions, and tool integration. Only the point of integration changes.

Crawling and scraping are distinct, and production pipelines lean on both. Scraping pulls specific data off pages a system already knows about. Crawling discovers new pages by following links across a site. Most large pipelines run both together: crawl first to find the relevant pages, scrape second to pull the actual data off each one.

AI-powered scraping has changed the game here by cutting loose the old dependency on fragile CSS selectors and XPath rules. Instead of targeting a specific DOM element that breaks the moment a site redesigns its layout, the system describes what data it wants in plain language, and the model finds it even as the page structure shifts because the description no longer depends on fixed layout markers. Dynamic content and inconsistent schemas are the two specific reasons that parsing raw HTML directly is often not good enough for preparing LLM training data.

Caching strategy quietly eats budgets if nobody thinks about it. Scrape endpoints typically reuse cached results younger than a configured maxAgeMs, defaulting to about one day. Setting that value to zero forces a live fetch every time, the right call for something like a pricing page that changes constantly. Raising it for stable reference content, on the other hand, cuts both latency and cost on every repeat read. It's a small dial, but it's the difference between a pipeline that's affordable at scale and one that isn't.

Given how much the retrieval layer shapes everything downstream, picking the right retrieval infrastructure is a major engineering decision. It's arguably the first decision that decides whether the rest of the pipeline is worth building. Preparing web content for LLMs requires cleaning, reformatting, and transforming into Markdown or JSON, which is a distinct engineering step rather than a side effect of scraping.

Web scraping APIs for extraction pipelines: what the June 2026 benchmarks show

Its free tier offered 1,000 credits a month, one scrape per credit at the base rate, with no credit card required to sign up (extras like JSON extraction or stealth mode cost additional credits per page). It cleared every single Cloudflare-protected site in the benchmark, five out of five, making it the strongest option specifically for Cloudflare-heavy targets.

A separate enterprise-grade suite in the same test, Bright Data, covered an Unlocker API, an Agent Browser, and a Web Scraper API. Its unlocker component returns clean HTML, JSON, or Markdown through a built-in HTML-to-Markdown converter, though boilerplate removal and content scoring on top of that still fall to the team building the pipeline. It also ships with more than 600 prebuilt scrapers tuned for well-known platforms like Amazon and LinkedIn, returning consistent JSON without custom selector work needed for each one.

The honest takeaway across both: no tool erases the underlying trade-off between coverage, cost, and how much custom cleanup work lands back on the engineering team. Picking the right one is less about finding a silver bullet and more about matching a tool's strengths (Cloudflare resilience, prebuilt platform coverage, raw proxy scale) to the actual shape of the target sites a pipeline needs to hit. Get that match wrong, and every downstream layer, from NER to the final knowledge graph, inherits the mess. The API handles JS rendering, full-site crawling, and autonomous data gathering via its /agent endpoint, plus interactive browser sessions via /interact, scoped to an existing scrape job as /v2/scrape/{jobId}/interact, all within a single API.

Sources

  1. Extracting Data from Unstructured Text: NLP & LLM Guide 2026
  2. Mapping Political-Elite Networks in Europe with a Multilingual Joint Entity-Relation Extraction Pipeline
  3. What is Relationship Extraction?
  4. A Relation-Oriented Model With Global Context Information for Joint Extraction of Overlapping Relations and Entities
  5. A Joint Extraction Model for Entity Relationships Based on Span and Cascaded Dual Decoding - PMC
  6. A Comprehensive Survey on Relation Extraction: Recent Advances and New Frontiers | ACM Computing Surveys
Filed underData Extraction

More in Data Extraction