Building MCP Servers for Agentic Web Browsing with Structured Data Access
Tutorials

Building MCP Servers for Agentic Web Browsing with Structured Data Access

Learn how to create Model Context Protocol servers that give LLMs live, structured web data using AlterLab's scraping API for reliable, agent-driven browsing.

H
Herald Blog Service
5 min read
4 views

AlterLab handles this automatically — scrape any URL with one API call. No infrastructure required.

Try it free

TL;DR

Build an MCP server that exposes a simple HTTP endpoint. The server calls AlterLab’s scraping API to fetch live web pages, converts the result to structured JSON, and returns it to an LLM agent. This gives the agent real‑time, reliable data without custom scraping code for each site.

Introduction

LLMs excel at reasoning but are limited by static training data. To enable agents that browse the web, you need a way to feed them fresh, structured information on demand. The Model Context Protocol (MCP) defines a lightweight interface for this purpose. By pairing an MCP server with a robust scraping API like AlterLab, you can give agents access to up‑to‑date e‑commerce listings, news articles, or any public page, all while avoiding the complexity of managing proxies, browsers, or anti‑bot systems yourself.

What is MCP?

MCP is a request‑response protocol where an LLM agent sends a structured query describing the data it needs, and the server returns that data in a predefined format. The protocol is agnostic to the underlying source—it could be a database, a file system, or a web scraper. For web browsing, the MCP server translates the agent’s request into a scrape job, waits for the result, and shapes it into JSON that the LLM can consume directly.

Why Use MCP for Agentic Browsing?

  • Real‑time freshness: Each request triggers a live scrape, so the agent sees the current state of a page.
  • Structured output: Instead of raw HTML, you can ask for JSON, markdown, or extracted fields, reducing parsing load on the LLM.
  • Separation of concerns: The LLM focuses on reasoning; the MCP server handles retrieval, retries, and format conversion.
  • Reusability: Multiple agents or tools can share the same MCP endpoint, promoting consistent data access.

Architecture Overview

A minimal MCP server for web browsing consists of three parts:

  1. API endpoint that receives MCP‑style JSON requests.
  2. Scraper client that calls AlterLab with the requested URL and desired output format.
  3. Response formatter that alters the AlterLab payload into the MCP‑expected structure.

The flow is synchronous for simplicity, but you can replace the scraper call with a queue or background job for higher throughput.

Setting Up the MCP Server

We’ll use Python and the FastAPI framework for its speed and automatic OpenAPI docs. First, install the dependencies:

Bash
pip install fastapi uvicorn alterlab

Create a file mcp_server.py and initialize the AlterLab client with your API key (obtainable from the AlterLab dashboard ).

Python
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
import alterlab
import json

app = FastAPI()
client = alterlab.Client("YOUR_ALTERLAB_API_KEY")  # Initialize once at startup

Define the request model. The MCP spec can be simple: a url field and an optional format field (json, markdown, text). Additional fields like extract_schema can be passed to AlterLab’s Cortex AI if you need predefined fields.

Python
class MCPRequest(BaseModel):
    url: str
    format: str = "json"  # json, markdown, or text
    extract_schema: dict | None = None  # Optional schema for Cortex AI

Now implement the endpoint. It validates the URL, calls AlterLab, and returns a JSON object with a data key containing the scraped content.

Python
@app.post("/mcp/scrape")
async def scrape_endpoint(req: MCPRequest):
    # Basic validation – reject empty URLs
    if not req.url.strip():
        raise HTTPException(status_code=400, detail="URL cannot be empty")

    # Prepare AlterLab parameters
    params = {
        "url": req.url,
        "formats": [req.format],
    }
    if req.extract_schema:
        params["extract"] = req.extract_schema

    try:
        # Call AlterLab – this handles proxies, browsers, anti‑bot
        alter_response = client.scrape(**params)
    except Exception as exc:
        raise HTTPException(status_code=502, detail=f"Scrape failed: {exc}")

    # AlterLab returns an object with .text, .json, etc., depending on format
    if req.format == "json":
        data = alter_response.json
    elif req.format == "markdown":
        data = alter_response.markdown
    else:
        data = alter_response.text

    # MCP‑style response
    return {"data": data, "format": req.format, "url": req.url}

Run the server locally:

Bash
uvicorn mcp_server:app --reload

The server now listens on http://127.0.0.1:8000/mcp/scrape.

Integrating AlterLab for Real‑Time Data

AlterLab abstracts away the hardest parts of web scraping:

  • Automatic retry with exponential backoff.
  • Rotating residential proxies to avoid IP bans.
  • Headless Chrome with stealth plugins for JavaScript‑heavy sites.
  • Built‑in anti‑bot handling that solves challenges without you managing CAPTCHA services.
  • Multiple output formats including raw HTML, cleaned markdown, and JSON.

By calling client.scrape() inside the MCP endpoint, you delegate all of this to AlterLab. The MCP server remains thin—its only job is to map the agent’s request to AlterLab parameters and return the result in a consistent shape.

Example: Requesting JSON from a Product Listing

An LLM agent might ask for the latest price and availability from a public product feed. The MCP request would look like:

JSON
{
  "url": "https://example-shop.com/category/widgets",
  "format": "json",
  "extract_schema": {
    "price": "css:.price",
    "availability": "css:.stock"
  }
}

AlterLab returns JSON matching the schema, and the MCP server forwards it unchanged.

Code Example: Python SDK

Below is a standalone snippet showing how an external service (or the LLM agent itself) could call the MCP server directly using the AlterLab Python SDK. This demonstrates the same operation as the endpoint but bypasses the MCP layer for debugging or testing.

Python
import alterlab
import requests

MCP_ENDPOINT = "http://127.0.0.1:8000/mcp/scrape"
ALTERLAB_KEY = "YOUR_ALTERLAB_API_KEY"

def get_structured_data(url: str, fmt: str = "json"):
    # Call the MCP server
    resp = requests.post(
        MCP_ENDPOINT,
        json={"url": url, "format": fmt}
    )
    resp.raise_for_status()
    return resp.json()["data"]

# Usage
data = get_structured_data("https://example.com/news")
print(json.dumps(data, indent=2))

Code Example: cURL

The same request can be made from a shell or any HTTP client. This is useful for quick verification or integrating with non‑Python services.

Bash
undefined
Share

Was this article helpful?

Frequently Asked Questions

An MCP (Model Context Protocol) server acts as a bridge between an LLM and external data sources, providing real-time, structured context. It lets agents fetch up-to-date information without retraining the model.
AlterLab provides a scraping API that handles anti-bot measures, proxies, and headless browsers, returning clean JSON or markdown. This simplifies data retrieval for the MCP server so it can focus on serving the LLM.
No. With AlterLab you can request structured formats like JSON or use Cortex AI extraction to get ready‑made fields, reducing the need for site‑specific parsers.