Craigslist Data API: Extract Structured JSON in 2026
Tutorials

Craigslist Data API: Extract Structured JSON in 2026

Learn how to extract structured JSON from Craigslist listings using AlterLab's Craigslist Data API – fast, typed output for AI pipelines and data workflows.

H
Herald Blog Service
4 min read
4 views

AlterLab handles this automatically — scrape any URL with one API call. No infrastructure required.

Try it free

This guide covers extracting publicly accessible data. Always review a site's robots.txt and Terms of Service before scraping.

TL;DR

Use AlterLab's Extract API to get typed JSON from Craigslist pages. Define a JSON schema for the fields you need, POST the URL and schema, and receive validated data instantly. No HTML parsing required.

Why use Craigslist data?

Craigslist hosts a constant stream of classifieds across housing, jobs, goods, and services. Engineers use this data to:

  • Train price prediction models for local markets
  • Monitor inventory trends for competitive analysis
  • Feed real‑time alerts into internal dashboards or AI agents

The volume and freshness make it a valuable signal for hyperlocal insights.

What data can you extract?

Every public listing exposes a set of predictable fields. You can request:

  • title – the headline of the post
  • price – the listed cost, often with currency
  • location – city, neighborhood, or distance marker
  • posted_date – when the listing went live
  • category – the Craigslist section (e.g., apartments, for sale)

These fields are enough to build structured datasets for analytics or machine learning pipelines. Because you specify the schema, AlterLab returns each field with the correct JSON type, eliminating downstream cleaning.

The extraction approach

Fetching a Craigslist page with plain HTTP and parsing HTML with regex or brittle selectors fails frequently. The site updates its markup, serves different layouts per region, and employs anti‑bot measures that block simple scrapers. A data API abstracts away:

  • Automatic retry with rotating proxies
  • JavaScript rendering for dynamic content
  • CAPTCHA solving when needed
  • Schema‑driven validation so you always get typed JSON

By treating Craigslist as a data source rather than a page to scrape, you gain reliability and reduce maintenance.

Quick start with AlterLab Extract API

First, install the AlterLab SDK (or use cURL directly). The Extract API endpoint accepts a URL and a JSON schema, then returns the extracted data.

See the Getting started guide for installation steps.

Python example

Python
import alterlab

client = alterlab.Client("YOUR_API_KEY")

schema = {
  "type": "object",
  "properties": {
    "title": {
      "type": "string",
      "description": "The title field"
    },
    "price": {
      "type": "string",
      "description": "The price field"
    },
    "location": {
      "type": "string",
      "description": "The location field"
    },
    "posted_date": {
      "type": "string",
      "description": "The posted date field"
    },
    "category": {
      "type": "string",
      "description": "The category field"
    }
  }
}

result = client.extract(
    url="https://craigslist.org/example-page",
    schema=schema,
)
print(result.data)

Line 5‑12 shows the schema definition. The call returns a Python dict matching the schema, ready for further processing.

cURL example

Bash
curl -X POST https://api.alterlab.io/v1/extract \
  -H "X-API-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://craigslist.org/example-page",
    "schema": {"properties": {"title": {"type": "string"}, "price": {"type": "string"}, "location": {"type": "string"}}}
  }'

The response body is a JSON object with the requested fields.

Batch/async example

For high‑volume workflows, submit multiple URLs as separate jobs and poll for completion.

Python
import alterlab
import time

client = alterlab.Client("YOUR_API_KEY")

urls = [
    "https://craigslist.org/search/sss?query=laptop",
    "https://craigslist.org/search/apa?availabilityMode=0",
    "https://craigslist.org/search/jjj?is_fulltime=1"
]

jobs = []
for u in urls:
    jobs.append(client.extract_async(url=u, schema=schema))

# Poll until all are done
while any(not j.done() for j in jobs):
    time.sleep(0.5)

for j in jobs:
    print(j.result().data)

This pattern lets you scale to thousands of pages while respecting rate limits.

Define your schema

The schema parameter drives the extraction. Use JSON Schema Draft‑07 syntax. AlterLab validates the LLM output against it and coerces types when possible. Example schema for a typical Craigslist posting:

JSON
{
  "type": "object",
  "properties": {
    "title": {"type": "string"},
    "price": {"type": "string"},
    "location": {"type": "string"},
    "posted_date": {"type": "string", "format": "date-time"},
    "category": {"type": "string"},
    "has_image": {"type": "boolean"}
  },
  "required": ["title", "price", "location"]
}

If a field is missing or malformed, the API returns an error detailing the validation failure, so you can handle it programmatically rather than guessing.

Handle pagination and scale

Craigslist results are paginated via URLs like ?s=0, ?s=100, etc. To collect all listings:

  1. Determine the total count from the first page (often in a header or summary text).
  2. Calculate the number of pages (total // page_size + 1).
  3. Dispatch extract calls for each s offset, either synchronously for low volume or asynchronously for larger jobs.

AlterLab’s pricing scales with usage. See the pricing page for details. Costs are clamped between $0.001 and $0.50 per extraction, and you can preview each call’s expense with the estimate endpoint before committing.

When running many parallel jobs, respect Craigslist’s crawl delay by adding a short sleep between batches or using the API’s built‑in rate‑limit headers. The platform returns `

Share

Was this article helpful?

Frequently Asked Questions

Craigslist does not offer a public API for classifieds listings. AlterLab provides a compliant way to extract publicly available data as structured JSON, handling anti-bot measures and schema validation.
You can extract any publicly visible fields such as title, price, location, posted date, and category. Define a JSON schema to get typed, validated output without manual parsing.
AlterLab charges per extraction based on compute usage, with a pay‑as‑you-go model and no minimums. Costs are clamped between $0.001 and $0.50 per call, and you can preview pricing with the Extract API estimate endpoint.