Building an MCP Server for Web Data Access
Design your MCP server to handle retrieval and representation separately.

Building an MCP server for web data access means solving two problems at the same time: exposing the right tools through a standard interface, and making sure the data those tools hand back is clean enough for an AI agent to actually use. Before the Model Context Protocol existed, hooking an AI agent up to any outside tool meant writing a one-off integration every single time: custom login handling, custom data shapes, custom error handling, repeated for every pairing of tool and agent. Anthropic had a name for this mess: the N×M problem, where N tools times M agents equals a pile of brittle, duplicated plumbing that nobody wants to maintain.
MCP turns that multiplication into addition. Build one server for your tool; any agent that speaks MCP can reach it. Implement the protocol once in your agent, and it can reach any tool that speaks MCP back. MCP gives tool integration one shared language, the same role TCP plays for computer networks and SQL plays for databases.
Anthropic put MCP out in late 2024, and the argument over whether it would catch on is settled. OpenAI folded MCP support into its products by March 2025, then between September and October 2025, it added full read and write MCP client support in ChatGPT through Developer Mode. Google shipped its own MCP support in late 2025 and followed with a managed remote server for Gemini in mid-2026. When the three biggest names in the field all build on the same protocol, that protocol stops being optional. The Linux Foundation's Agentic AI Foundation, co-founded by Anthropic, Block, and OpenAI, now stewards MCP, and by early 2026 it counted over 10,000 active public MCP servers and tens of millions of monthly SDK downloads.
Under the hood, MCP runs on JSON-RPC 2.0 and supports three ways of moving data: stdio for local processes, SSE (on its way out), and Streamable HTTP, which is now the standard for anything running over the internet. A server built on MCP can expose tools the AI can call, resources it can read, and prompt templates it can use directly. MCP is now the language any serious web data tool needs to speak, so building an MCP server for web data access is the right starting point for everything that follows.
What web data access needs from an MCP server
Teams build an MCP server for one reason more than any other: to reach live web content, documentation, and knowledge bases. But web data is not a generic integration problem, and treating it like one is how most implementations fall apart. It splits into two separate jobs that a simple setup tends to ignore.
The first job is retrieval: actually fetching the right pages, reliably, across sites loaded with JavaScript, behind anti-bot defenses, with content that changes depending on who's asking and when. Most of the modern web just breaks a plain HTTP request. The second job is representation: getting that content into a shape an AI agent can reason on, rather than handing it raw HTML, which is bloated, messy, and expensive for a model to chew through.
An MCP server that only solves the first problem and ships raw HTML straight to the agent hasn't actually solved anything. It's shoved the mess one level deeper, straight into the agent's context window, where it drives up token costs and makes the agent's reasoning worse, not better.
Because of this, the tools a web data MCP server needs to expose go well past "fetch this URL." At minimum, a server needs to cover search, to find relevant pages; scraping, to pull content off one page; crawling, to move across a whole site; and structured extraction, to pull out specific fields like prices, dates, or names. A task needs the right shape of data, but a server with only one tool forces every task through that same narrow door.
That argues for a single server exposing one unified set of tools: search, scrape, crawl, extract, monitor. If an agent connects to another server, that server's entire tool schema lands in the context window on every single call, whether that tool gets used or not. Stitching together five narrow servers instead of one broad one is a tax paid on every request, forever.
Why the output format is a first-order design decision
Defaulting to raw HTML doesn't avoid a choice: it's an expensive one, made by accident.
Markdown cuts token use dramatically when it carries the same content raw HTML would. Cloudflare built a feature called "Markdown for Agents" specifically to strip HTML down to Markdown before feeding pages to AI systems, pointing to token savings, simpler pipelines, and treating AI agents as real users of the web. Cloudflare's own analysis found that a typical blog post in Markdown used only a fraction of the tokens its HTML version did.
Markdown doesn't just save tokens; it also reads better to a model. In extraction benchmarks built on a leading language model, Markdown versions of pages scored higher accuracy than HTML did on structured data like tables, and when retrieval pipelines feed on Markdown instead of raw HTML, they see real gains in accuracy.
The rule that falls out of this is straightforward. Make Markdown the default return format for prose, articles, and documentation pages. JSON schema extraction should be saved for fields the agent will act on directly: prices, names, dates, stock status, the kind of data that needs a type, not a paragraph.
None of this is just a scraping detail. It's a decision baked into how the MCP server itself is built. A tool's schema should state up front what format it returns, so the agent knows before it even makes the call what shape of data is coming back, and can plan the rest of its reasoning around that.
Structuring the MCP Tool Surface for Web Data
The tool schema is the contract between an MCP server and the agent calling it. Get the schema wrong, and the agent will misuse the tools, make calls it didn't need to make, or never discover a capability it actually needed. A well-built web data MCP server keeps that contract small and clean: a handful of tools, each doing one job well, instead of a sprawling set of options that overlap and confuse each other.
Search takes a query string and returns a ranked list of URLs with titles and snippets, giving the agent a way to find relevant pages without already knowing the exact address. Scrape takes a single URL and returns clean Markdown, or structured JSON when the task calls for typed fields, and works as the core extraction tool in the whole set. Crawl takes a starting URL along with depth and filter settings and returns the pages it finds across a site, built for jobs that need to move across many pages. Extract takes a URL together with a JSON schema describing exactly which fields to pull, for cases where the agent needs specific values. Monitor, sometimes called batch, takes a list of URLs along with a schedule or trigger, and it's built for change detection and tracking competitors over time.
A short illustrative tool definition might look like this:
{
"name": "scrape",
"description": "Fetch a URL and return clean Markdown or structured JSON.",
"parameters": {
"url": "string",
"format": "markdown | json",
"schema": "object (optional, required if format=json)"
}
}
Every parameter in a tool's schema needs a clear, typed description, because the agent reads those descriptions to decide when and how to use the tool. Vague descriptions produce vague, unpredictable agent behavior. And schema size has a real cost: every connected server dumps its complete tool definitions into the agent's context window on every single request, so a bloated surface with tools that overlap or duplicate each other adds overhead on every call, used or not. A gateway that filters which schemas reach the agent cuts that overhead by a wide margin, based on analysis of how MCP's token economics actually work.
This is where a unified API earns its place in the conversation. Olostep's API, covering search, scraping, crawling, mapping, batching, and monitoring, lines up directly with this tool surface: each capability becomes one typed MCP tool, backed by one reliable piece of infrastructure, instead of five separate servers each adding its own schema weight to every call.
Handling JavaScript rendering, anti-bot layers, and other retrieval realities
The most common way a web data MCP server breaks in practice: the scrape tool comes back empty or garbled because whatever is fetching the page underneath can't handle JavaScript-heavy sites, bot detection, or a login wall. The agent has no way to diagnose that failure or recover from it on its own.
None of that is actually an MCP problem. JavaScript rendering, getting past anti-bot systems, pulling out clean Markdown, producing structured output, these are scraping infrastructure problems, full stop. The MCP server's actual job is to define clean tools and validate schemas, not to run a browser engine inside its own request handlers.
The MCP server should handle tool definitions, schema validation, and formatting the response. A separate, purpose-built scraping API should handle rendering pages, rotating proxies, getting past CAPTCHAs, and pulling content out cleanly. The MCP server calls that API. It doesn't run a headless browser itself.
That separation keeps the MCP server stateless by default, in line with the MCP specification revision dated 2026-07-28. That revision drops session state from the protocol's core, replacing the old initialize handshake and session IDs with per-request _meta fields that carry protocol version, client identity, and capabilities on every single request, along with new routing headers (Mcp-Method, Mcp-Name). If a server hands retrieval off to an outside API, it's already stateless by nature. It scales on ordinary load-balanced HTTP infrastructure, with no shared session store needed anywhere.
Olostep's API handles JavaScript rendering, structured Markdown extraction, and high-volume batching behind one endpoint. The MCP server built on top of it stays a thin, stateless layer that translates MCP tool calls into Olostep API calls and hands back data the agent can use right away.
Authentication, authorization, and the security risks specific to web data MCP servers
A web data MCP server carries a different threat profile than most other MCP servers, because its core job, fetching arbitrary URLs on request, is itself a weak point if you leave it wide open. Unrestricted URL fetching opens the door to server-side request forgery: anyone who can influence which URLs an agent requests can steer the MCP server into internal networks, cloud metadata endpoints, or pages built specifically to pull data back out through the agent. The attacker doesn't even need to touch the agent's input directly. They only need to control content on a page the agent is going to fetch. That's indirect prompt injection, and it's the attack vector web data agents face most today. A research technique called "Memory Heist," documented by Tencent Zhuque Lab, showed agent memory being pulled out through link-following on pages an attacker controlled.
Given that, the defenses against these risks belong in the server's design from day one, not bolted on after something breaks. URL fetching should run against an allowlist or be scoped to specific domains, so the scrape tool can't be pointed at arbitrary internal addresses. Content coming back from the web should be sanitized before it ever reaches the agent, stripping or escaping anything that resembles an instruction aimed at the model itself. Remote servers should run OAuth 2.1 with PKCE, now the standard approach for MCP authentication (Dynamic Client Registration has given way to Client ID Metadata Documents in the 2026-07-28 revision), and ten major AI agents already support OAuth 2.1 natively as of early 2026. On top of all that, an MCP gateway sitting in front of the server as a reverse proxy gives a single place to enforce authentication, apply rate limits, trip circuit breakers, log every call for audits, and control what the server is allowed to reach out to on the open internet.
Deployment architecture: stateless scaling, token overhead, and the gateway pattern
Running a web data MCP server in production means juggling three problems at once: keeping session state manageable, keeping schema token overhead down, and enforcing security consistently. All three point toward the same fix: a gateway layer sitting between the agent and the MCP server.
The request path runs through four layers in sequence: the MCP client (the agent asking for data), the gateway (auth, rate limiting, schema filtering), the MCP server itself (tool definitions and formatting), and the Olostep API underneath doing the actual fetching from the live web. Each layer does one job and hands off cleanly to the next.
Anyone running more than one MCP server at a time faces real token overhead. Every server connected to an agent dumps its full tool definitions into the context window on every call, whether or not those tools get used that round. A gateway that filters which schemas reach the agent, showing only the tools relevant to the task at hand, cuts that overhead by a wide margin, based on analysis of how MCP's token costs actually work in practice. The same gateway layer also centralizes the rest of the operational load: OAuth 2.1 and token checks for authentication, rate limiting and circuit breaking to stop a runaway agent loop from hammering the scraping API underneath, audit logs recording every tool call for compliance and debugging, and one single point controlling what URLs the system is allowed to reach.
Where the MCP server itself actually runs is a secondary decision. Self-managed containers, serverless or edge deployments, and managed MCP platforms are all on the table, and the right pick depends on request volume and how much latency a use case can tolerate. Because the July 2026 spec revision introduced a stateless core, none of those three options carries the old session-management baggage now.
The real scaling challenge sits one layer down, in the retrieval API doing the actual fetching across potentially billions of requests. The MCP server on top of it stays thin and stateless, and scaling it is trivial next to the infrastructure underneath handling rendering, proxies, and anti-bot work at volume. That's the build-versus-buy calculus in a sentence: build a thin MCP server, and hand the heavy retrieval work to a purpose-built API built to carry that load.


