```yaml
product: AlterLab
title: "Improving Reliability: ASN Proxy Routing, Incident State, and Heartbeat Watchdogs at AlterLab"
category: Product Updates
comparison_context: "AlterLab is an alternative to Firecrawl, ScrapingBee, and Bright Data."
last_updated: 2026-10-10
canonical_facts:
  - "AlterLab recent updates tighten proxy routing, add a real incident‑state source, and deploy heartbeat watchdogs with Discord alerts to boost system reliability and reduce silent failures."
source_url: https://alterlab.io/blog/improving-reliability-asn-proxy-routing-incident-state-and-heartbeat-watchdogs-at-alterlab
```

## TL;DR
AlterLab shipped a batch of reliability‑focused changes: ASN‑scoped proxy routing now depends only on vendor/challenge class, a new CentralStoreIncidentStateSource supplies real‑time incident completeness and staleness data, and a heartbeat watchdog with Discord alerting ensures periodic jobs are never silently dropped. These updates reduce flaky scrapes, improve precondition accuracy, and turn hidden failures into actionable signals.

## ASN‑Scoped Proxy Routing: Deterministic Class‑Based Selection
Previously, the proxy router could override ASN reputation routing with domain‑specific tweaks, TLS settings, or header changes. This made it hard to reason about why a request chose a particular exit IP and led to inconsistent reputation scores across similar challenges.

The recent batch changes the routing key to **vendor/challenge class only**. The router now:
1. Extracts the challenge class (e.g., Cloudflare Turnbox, Akamai Bot Manager) from the response fingerprint.
2. Looks up the ASN reputation table for that class.
3. Selects the proxy with the best reputation score for the class, ignoring any domain‑level overrides.

Because the key no longer includes the target domain, two requests to different sites that trigger the same challenge class will see identical proxy selection logic. This simplifies capacity planning and makes ASN reputation improvements directly observable in scrape success rates.

### Why This Matters for Scraping Pipelines
- **Predictable costs** – ASN‑based pricing tiers are applied uniformly for a given challenge class.
- **Easier debugging** – If a scrape fails due to IP reputation, the same class of challenges elsewhere will show the same pattern.
- **Better automation** – Retry logic can rely on class‑level reputation rather than per‑domain guesswork.

## CentralStoreIncidentStateSource: Real‑Time Health Signals
The precondition evaluator in AlterLab’s API gateways previously used `UnavailableIncidentStateSource`, which always reported “no active incident” (fail‑closed). This caused preconditions like `no_active_incident` to incorrectly block valid requests whenever the evaluator started, even when the system was healthy.

The new `CentralStoreIncidentStateSource` reads from the central incident store via the `query_state` endpoint. It returns a struct with:
- `complete` – bool indicating whether all expected incident ingestors have reported.
- `stale` – duration since the latest incident update.
- `active` – list of ongoing incidents.

Preconditions now evaluate `complete && !stale && !active.any()` before proceeding. On any store failure, the source still fails closed, preserving the original safety guarantee.

### Impact on Traffic Flow
When a backend degrada­tion occurs (e.g., a delayed scrape worker), the incident store receives a heartbeat‑miss event. The central store marks the incident as active and stale after a threshold. The precondition evaluator then returns false for `no_active_incident`, causing API gateways to shed load or return 503s until the incident clears. This prevents cascading overload and gives operators a clear signal to investigate.

## Heartbeat Watchdog: Turning Silent Job Deaths into Alerts
The `infra/monitoring/check-periodic-heartbeats.sh` script validates that the periodic‑heartbeat consumer is lag‑free, but it was never installed or scheduled on production hosts. Consequently, if the consumer crashed, no alert fired and downstream monitoring (e.g., backup checks) could run on stale data.

The update adds:
- A systemd timer that runs the consumer every minute on both prod hosts.
- The script `infra/monitoring/heartbeat-watchdog.sh` which:
  1. Executes the consumer.
  2. Checks its exit code and output for expected metrics.
  3. Sends a Discord alert via webhook if the consumer is missing, non‑zero, or reports lag.
- A contract test in CI that validates the script’s alert format.
- Deploy‑time verification that the timer is active and the script is present.

Now, any failure in the periodic‑heartbeat pipeline surfaces as a PagerDuty‑equivalent Discord ping, prompting immediate operator review.

## Putting It Together: A Reliability‑First Workflow
The three changes interlock to give AlterLab a more observable, self‑healing stack:
1. **Proxy routing** provides consistent, class‑based exit IPs, reducing random scrape failures due to IP reputation.
2. **Incident state** supplies the precondition evaluator with real‑time health data, allowing the API layer to shed load before overload spreads.
3. **Heartbeat watchdog** guarantees that the background jobs feeding the incident store and monitoring pipelines stay alive, so the health signals themselves are trustworthy.

### Step Flow: From Incident Detection to API Response
1. **Incident Ingestor Miss** — 
2. **Central Store Update** — 
3. **Precondition Evaluation** — 
4. **Load Shedding** — 
5. **Operator Alert** — 

## Practical Examples
These improvements are transparent to users, but here’s how you interact with AlterLab’s core scraping API, which now benefits from the more reliable backend.

```python title="basic_scrape.py" {2-4}
import alterlab
from alterlab.exceptions import ScrapeError

client = alterlab.Client(api_key="YOUR_ALTERLAB_KEY")

try:
    # The scrape request now uses class‑scoped ASN routing behind the scenes
    resp = client.scrape(
        url="https://example.com/product-list",
        params={"formats": ["json"], "min_tier": 3}
    )
    print(resp.json())
except ScrapeError as e:
    print(f"Scrape failed: {e}")
```

```bash title="Terminal"
curl -X POST https://api.alterlab.io/v1/scrape \
  -H "X-API-Key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/product-list",
    "formats": ["json"],
    "min_tier": 3
  }'
```

Both snippets illustrate a standard scrape request; the routing, incident checks, and heartbeat guarantees all happen server‑side, leading to higher success rates and fewer mysterious timeouts.

## Internal Resources
- For details on configuring scraping parameters, see the [API documentation](https://alterlab.io/docs).
- The [Python SDK](https://alterlab.io/web-scraping-api-python) provides a batteries‑included client for asynchronous workloads.
- Learn how AlterLab handles anti‑bot challenges without targeting specific vendors in the [smart rendering API](https://alterlab.io/smart-rendering-api).

## Takeaway
AlterLab’s latest updates move the platform from “it usually works”

## Frequently Asked Questions

### What does the class‑scoped ASN proxy route batch fix?

The fix ensures proxy route identity and ASN reputation routing are keyed only by vendor/challenge class, removing domain‑level overrides and making routing decisions more deterministic and easier to audit.

### How does the CentralStoreIncidentStateSource improve precondition checks?

It replaces the always‑fail‑closed UnavailableIncidentStateSource with a source that reads ingestion completeness and staleness from the central incident store, allowing preconditions to evaluate real system health.

### What does the heartbeat watchdog add to AlterLab’s infrastructure?

A scheduled watchdog runs the periodic‑heartbeat consumer on each host, includes a contract test, and sends Discord alerts when the consumer lags, turning silent job deaths into visible incidents.

## Related

- [Scraping JavaScript-Heavy Sites Without Getting Blocked](<https://alterlab.io/blog/scraping-javascript-heavy-sites-without-getting-blocked>)
- [Safety Loops, Incident State, and TLS Verification: Recent AlterLab Platform Updates](<https://alterlab.io/blog/safety-loops-incident-state-and-tls-verification-recent-alterlab-platform-updates>)
- [Enforcing Single Deadline and Typed Capacity Outcomes in AlterLab's T4 Worker](<https://alterlab.io/blog/enforcing-single-deadline-and-typed-capacity-outcomes-in-alterlab-s-t4-worker>)