Worker Reconciliation, Trusted Runtime, Netcup Relay Fixes
Product Updates

Worker Reconciliation, Trusted Runtime, Netcup Relay Fixes

Deep dive into AlterLab's latest infra and worker improvements: bounded reconciliation retries, trusted root runtime rollout, and Netcup relay candidate recovery fixes for reliable scraping pipelines.

H
Herald Blog Service
4 min read
3 views

AlterLab handles this automatically — scrape any URL with one API call. No infrastructure required.

Try it free

TL;DR

We bounded worker reconciliation retries to prevent infinite loops, rolled out a trusted root runtime for secure execution, and fixed Netcup relay candidate recovery to eliminate backup‑safety CI failures. These changes improve reliability and security of AlterLab's scraping platform.

Introduction

AlterLab’s platform runs thousands of scraping jobs per minute, relying on workers that reconcile API outcomes, a runtime that executes privileged tasks, and a relay system that backs up candidate state. Recent work focused on three distinct but critical areas: making worker reconciliation resilient under contention, hardening the execution environment for root‑level operations, and repairing a subtle bug in the Netcup relay that caused intermittent CI failures. The following sections detail each change, the problem it solved, and why it matters for developers building reliable data pipelines.

Worker Reconciliation: Bounded Retries Under Shared Billing Lock

The Problem

When a worker processes a scrape request, it must confirm that the API call and the internal worker commit both succeeded under a shared billing lock. Ambiguous outcomes—such as a timeout where the API call may have succeeded but the worker lost its lock—triggered a reconciliation loop. Previously, the worker would retry indefinitely, holding the lock and blocking other jobs, while also risking duplicate billing if the lock was reacquired after a partial commit.

The Solution

We introduced two bounds:

  1. Attempt limit – a configurable maximum number of reconciliation retries (default 5).
  2. Deadline – an absolute time window (default 30 seconds) after the first ambiguous outcome, after which the worker aborts reconciliation and marks the job for a later sweep.

Crucially, claims on the scrape job are retained throughout the retry window, so a subsequent sweep can reacquire the lock and finish the job without losing progress.

Why It Matters

  • Predictable latency – jobs no longer stall indefinitely; they either complete quickly or are handed off cleanly.
  • Lock fairness – the shared billing lock is released promptly, reducing queue back‑pressure.
  • Exactly‑once semantics – retained claims ensure that a later sweep can finish the job without double‑charging.

Trusted Root Runtime Rollout

The Problem

Privileged operations (e.g., updating system‑level certificates, cleaning stale mounts) were previously performed by ad‑hoc scripts executed via root-cron. These scripts:

  • Pulled code from mutable checkouts, making it hard to verify what actually ran.
  • Mixed trusted executables with live deployment data, increasing the attack surface.
  • Relied on environment variables like PYTHONPATH, RCLONE_CONFIG, or shell aliases that could be injected by compromised jobs.

The Solution

We replaced every root-cron checkout path with a verified root‑owned delivery helper that:

  1. Pulls an immutable, signed artifact from our internal artifact registry.
  2. Executes the helper in a clean, isolated namespace where only a whitelist of environment variables (e.g., PATH, HOME) is permitted.
  3. Separates trusted executable sources (the helper binary) from validated live deployment data (read‑only mounts), ensuring that no job‑supplied variables can influence the helper’s behavior.

The helper itself is a small, audited binary written in Go, with no external dependencies beyond the standard library. All privileged tasks now go through this single, traceable entry point.

Why It Matters

  • Supply‑chain integrity – the helper’s binary is signed and its hash verified before execution.
  • Environment hygiene – by stripping dangerous variables, we eliminate a class of injection attacks.
  • Operational clarity – auditors can inspect a single binary and its invocation logs rather than chasing scattered cron jobs.

Netcup Relay Candidate Recovery Fix

The Problem

The staging promotion backup‑safety CI suite began failing on Netcup relay tests. Investigation revealed three related issues:

  1. Bash dependent‑local expansion in backup-relay-scp where ${var:-default} was expanded incorrectly when var was unset, causing malformed SCP paths.
  2. Fixture mismatch – the installer, SCP, and PITR test fixtures used static WAL file names, while the relay protocol now generates transaction‑specific names (e.g., wal_000000010000000A000000B0.partial).
  3. Missing regression test – there was no automated check that orphan‑candidate recovery preserved HMAC‑based authentication, quota limits, ownership, and admission controls.

The Solution

  1. Fixed the Bash expansion by quoting the variable and using a explicit test:
    Bash
    if [ -z "$var" ]; then
        src_path="/default/path"
    else
        src_path="$var"
    fi
  2. Aligned fixtures to generate the same transaction‑specific WAL partial names used in production, ensuring the recovery logic sees the expected files.
  3. Added an authenticated recovery test that verifies:
    • HMAC signatures on candidate manifests are validated.
    • Quota and ownership checks are enforced before restoring a candidate.
    • Admission control (e.g., max concurrent recoveries) remains active.

Why It Matters

  • Backup safety – reliable candidate recovery prevents data loss during promotion failures.
  • CI confidence – the Netcup relay suite now passes consistently, catching regressions early.
  • Security preservation – the fix does not weaken any existing cryptographic or access‑control guarantees.

Practical Examples

Below are two code snippets that illustrate how developers interact with the updated systems.

Example 1: Scrape with Automatic Tier Escalation and Monitoring

Python
import alterlab
from alterlab.models import ScrapeRequest

client = alterlab.Client("
Share

Was this article helpful?

Frequently Asked Questions

It limits how many times the worker retries ambiguous API outcomes, preventing infinite loops while preserving claims until a sweep can finish the job.
It replaces ad‑hoc root‑cron checks with a verified, root‑owned delivery helper and isolates executables from live deployment data, blocking unsafe environment variables.
A Bash expansion bug and mismatched WAL fixture names caused backup‑safety CI failures; the fix aligns paths with transaction‑specific names and adds regression tests without weakening security checks.