
Building MCP Servers for Agentic Web Browsing with Structured Data Access
Learn how to create Model Context Protocol servers that give LLMs live, structured web data using AlterLab's scraping API for reliable, agent-driven browsing.
AlterLab handles this automatically — scrape any URL with one API call. No infrastructure required.
Try it freeTL;DR
Build an MCP server that exposes a simple HTTP endpoint. The server calls AlterLab’s scraping API to fetch live web pages, converts the result to structured JSON, and returns it to an LLM agent. This gives the agent real‑time, reliable data without custom scraping code for each site.
Introduction
LLMs excel at reasoning but are limited by static training data. To enable agents that browse the web, you need a way to feed them fresh, structured information on demand. The Model Context Protocol (MCP) defines a lightweight interface for this purpose. By pairing an MCP server with a robust scraping API like AlterLab, you can give agents access to up‑to‑date e‑commerce listings, news articles, or any public page, all while avoiding the complexity of managing proxies, browsers, or anti‑bot systems yourself.
What is MCP?
MCP is a request‑response protocol where an LLM agent sends a structured query describing the data it needs, and the server returns that data in a predefined format. The protocol is agnostic to the underlying source—it could be a database, a file system, or a web scraper. For web browsing, the MCP server translates the agent’s request into a scrape job, waits for the result, and shapes it into JSON that the LLM can consume directly.
Why Use MCP for Agentic Browsing?
- Real‑time freshness: Each request triggers a live scrape, so the agent sees the current state of a page.
- Structured output: Instead of raw HTML, you can ask for JSON, markdown, or extracted fields, reducing parsing load on the LLM.
- Separation of concerns: The LLM focuses on reasoning; the MCP server handles retrieval, retries, and format conversion.
- Reusability: Multiple agents or tools can share the same MCP endpoint, promoting consistent data access.
Architecture Overview
A minimal MCP server for web browsing consists of three parts:
- API endpoint that receives MCP‑style JSON requests.
- Scraper client that calls AlterLab with the requested URL and desired output format.
- Response formatter that alters the AlterLab payload into the MCP‑expected structure.
The flow is synchronous for simplicity, but you can replace the scraper call with a queue or background job for higher throughput.
Setting Up the MCP Server
We’ll use Python and the FastAPI framework for its speed and automatic OpenAPI docs. First, install the dependencies:
pip install fastapi uvicorn alterlabCreate a file mcp_server.py and initialize the AlterLab client with your API key (obtainable from the AlterLab dashboard ).
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
import alterlab
import json
app = FastAPI()
client = alterlab.Client("YOUR_ALTERLAB_API_KEY") # Initialize once at startupDefine the request model. The MCP spec can be simple: a url field and an optional format field (json, markdown, text). Additional fields like extract_schema can be passed to AlterLab’s Cortex AI if you need predefined fields.
class MCPRequest(BaseModel):
url: str
format: str = "json" # json, markdown, or text
extract_schema: dict | None = None # Optional schema for Cortex AINow implement the endpoint. It validates the URL, calls AlterLab, and returns a JSON object with a data key containing the scraped content.
@app.post("/mcp/scrape")
async def scrape_endpoint(req: MCPRequest):
# Basic validation – reject empty URLs
if not req.url.strip():
raise HTTPException(status_code=400, detail="URL cannot be empty")
# Prepare AlterLab parameters
params = {
"url": req.url,
"formats": [req.format],
}
if req.extract_schema:
params["extract"] = req.extract_schema
try:
# Call AlterLab – this handles proxies, browsers, anti‑bot
alter_response = client.scrape(**params)
except Exception as exc:
raise HTTPException(status_code=502, detail=f"Scrape failed: {exc}")
# AlterLab returns an object with .text, .json, etc., depending on format
if req.format == "json":
data = alter_response.json
elif req.format == "markdown":
data = alter_response.markdown
else:
data = alter_response.text
# MCP‑style response
return {"data": data, "format": req.format, "url": req.url}Run the server locally:
uvicorn mcp_server:app --reloadThe server now listens on http://127.0.0.1:8000/mcp/scrape.
Integrating AlterLab for Real‑Time Data
AlterLab abstracts away the hardest parts of web scraping:
- Automatic retry with exponential backoff.
- Rotating residential proxies to avoid IP bans.
- Headless Chrome with stealth plugins for JavaScript‑heavy sites.
- Built‑in anti‑bot handling that solves challenges without you managing CAPTCHA services.
- Multiple output formats including raw HTML, cleaned markdown, and JSON.
By calling client.scrape() inside the MCP endpoint, you delegate all of this to AlterLab. The MCP server remains thin—its only job is to map the agent’s request to AlterLab parameters and return the result in a consistent shape.
Example: Requesting JSON from a Product Listing
An LLM agent might ask for the latest price and availability from a public product feed. The MCP request would look like:
{
"url": "https://example-shop.com/category/widgets",
"format": "json",
"extract_schema": {
"price": "css:.price",
"availability": "css:.stock"
}
}AlterLab returns JSON matching the schema, and the MCP server forwards it unchanged.
Code Example: Python SDK
Below is a standalone snippet showing how an external service (or the LLM agent itself) could call the MCP server directly using the AlterLab Python SDK. This demonstrates the same operation as the endpoint but bypasses the MCP layer for debugging or testing.
import alterlab
import requests
MCP_ENDPOINT = "http://127.0.0.1:8000/mcp/scrape"
ALTERLAB_KEY = "YOUR_ALTERLAB_API_KEY"
def get_structured_data(url: str, fmt: str = "json"):
# Call the MCP server
resp = requests.post(
MCP_ENDPOINT,
json={"url": url, "format": fmt}
)
resp.raise_for_status()
return resp.json()["data"]
# Usage
data = get_structured_data("https://example.com/news")
print(json.dumps(data, indent=2))Code Example: cURL
The same request can be made from a shell or any HTTP client. This is useful for quick verification or integrating with non‑Python services.
undefinedWas this article helpful?
Frequently Asked Questions
Related Articles

Aligning Protected Storage Capture and Restore Contracts in AlterLab
Learn how AlterLab repaired its protected-storage producer/consumer contract to enable safe promotion from staging to main while preserving database snapshots, redaction, financial, and usage guarantees.
Herald Blog Service

How to Scrape Etsy Data: Complete Guide for 2026
Learn how to scrape Etsy for e-commerce data using Python and Node.js. This guide covers anti-bot challenges, structured extraction, and scaling your pipeline.
Herald Blog Service

Weekly Product Roundup: Reliability Fixes for AlterLab's Scraping API
This week's AlterLab update includes key fixes for job ordering, proxy caching, and dashboard parameters to improve scraping reliability.
Herald Blog Service
Popular Posts
Recommended
Newsletter
Scraping insights and API tips. No spam.
Recommended Reading

How to Scrape AliExpress: Complete Guide for 2026

Why Your Headless Browser Gets Detected (and How to Fix It)

AlterLab vs Firecrawl: Which Scraping API Is Better in 2026?

How to Scrape Twitter/X Data: Complete Guide for 2026

How to Scrape Cloudflare-Protected Sites in 2026
Stay in the Loop
Get scraping insights, API tips, and platform updates. No spam — we only send when we have something worth reading.
Explore AlterLab
Web Scraping API Resources
Part of the Web Scraping API Documentation cluster
Complete API reference with 5-tier auto-escalation — Curl to challenge resolution.
Pillar pageConfigure Tier 4 browser rendering for SPAs and dynamic content.
Scrape pages behind login using session management.
Real success rates and cost data across all 5 tiers.
MCP Server, Python SDK, and Firecrawl-compatible API for AI agent workflows.