Warm worker pool
The warm worker pool provisions, leases, and reaps the SSH-addressable machines your agent runs execute on. It is a long-lived singleton in @lorenz/worker-pool that survives workflow hot-reload, calls a swappable WorkerDriver for provision/probe/destroy/list, and owns every lifecycle decision itself: leasing, warm top-up, a reaper, spend caps, a write-ahead ledger, and crash recovery. This page is for operators tuning the pool under worker.worker_pool.
The pool is the single dispatch path for every workerHost. The legacy static list worker.ssh_hosts folds into a static-ssh pool (one machine per host), so it is not a separate model; you cannot name a driver twice for the same hosts, so the parser rejects worker.ssh_hosts alongside worker.worker_pool.driver or worker.kind. See static SSH workers for that path.
Turning it on
The pool is always on - it is the single dispatch path. With no worker.worker_pool block and no worker.ssh_hosts, the pool defaults to the local driver at max: 1, which runs agents locally on the daemon's own in-process endpoint (no SSH, no provisioning). To put runs on real machines, set a driver (and its sizing knobs); there is no enabled flag to set:
server:
# Remote per-run claims use this loopback listener while the dashboard keeps
# the default server.port (4040).
mcp_port: 4041
worker:
worker_pool:
driver: docker
min: 1
max: 4
warm: 2
SSH-addressable pool drivers require a distinct server.mcp_port while the dashboard is enabled. The remote agents reach it through Lorenz's reverse SSH tunnels; local-driver workflows can omit it and keep sharing the dashboard port.
When worker.worker_pool is present but driver is unspecified, the driver defaults to fake (in-memory, never touches SSH or disk). Set it to a real driver for real machines.
You can also keep driver options in a named profile and point worker.kind at it:
worker:
kind: ci-docker
workers:
ci-docker:
driver: docker
image: ghcr.io/acme/agent-box:latest
The keys under workers.<name> (minus driver) pass to the driver factory verbatim. worker.kind cannot be combined with worker.worker_pool.driver; pick one place to name the driver.
Lifecycle
Each worker moves through a small state machine. The pool stamps a worker LEASED when a run acquires it and returns it to WARM_IDLE on a healthy release; a poisoned release or a failed lease flags the worker for destruction, and the reaper or the last lease return recycles it.
A freshly provisioned worker is probed before it can be leased. The pool calls driver.probe up to 3 attempts (50 ms times the attempt number between attempts) and only then admits the worker to inventory. An unready worker is destroyed: a failed grow returns no_capacity with reason driver_error, and a failed warm top-up is skipped. This is the reachable-before-leased contract, so a run never receives a machine that cannot answer.
Two states in the WorkerState vocabulary, WARMING and DRAINING, are never assigned by the pool. They are reserved for future async warmup and per-worker drain. You will not see them in a current snapshot.
The acquire path
acquire() resolves synchronously where it can and parks the caller only when it must. The order is fixed:
- Short-circuit if the pool is disabled or draining, returning
no_capacitywith reasonpool_disabled. - Roll the UTC day key for daily spend accounting.
- Check spend caps. If a cap blocks the acquire, return
no_capacitywith reasonspend_cap. - Run
selectAndStampsynchronously: prefer a sticky-affinity worker, then any idle worker, then an under-capacity worker (co-residence). - If nothing is free and the pool can grow, grow under a reservation.
- If growth is blocked only by a spend cap, return
spend_cap. - Otherwise park on the FIFO waiter queue and wait for capacity.
A parked waiter that is not woken before acquire_timeout_ms returns no_capacity with reason acquire_timeout. The dispatcher maps that signal onto its existing worker_host_capacity backpressure, so a timed-out acquire reschedules the run rather than failing it. The four no_capacity reasons are a closed set: acquire_timeout, spend_cap, pool_disabled, driver_error.
Growth is reservation-based and single-flight. The pool increments a reservation counter synchronously before any provision await, so concurrent grows can never overshoot max (or max_workers_per_issue, when set). The reservation is released in a finally block.
A lease settles exactly once. release('healthy') keeps the worker warm; fail(reason) or release('poison') flags it for destruction and recycles it when the last lease returns. A release against a stale or already-destroyed worker is a no-op that never touches the in-flight count.
Config reference
All keys live under worker.worker_pool. Write them in snake_case.
| Key | Default | Meaning |
|---|---|---|
driver |
fake when a block is present, local when absent |
Registered driver kind, or an out-of-tree module specifier. There is no enabled key; the pool is always live. |
min |
0 |
Floor the reaper keeps warm. Never reaped below this. |
max |
1 |
Ceiling on total workers. Must be >= min. |
warm |
1 |
Target idle workers the top-up maintains. Must be <= max. |
max_in_flight |
1 |
Deprecated alias for slotsPerMachine (co-residence slots per machine). |
ttl_ms |
3600000 |
Max worker lifetime. A LEASED worker past TTL is flagged; an idle one is reaped above min. |
idle_reap_ms |
300000 |
Idle duration before a warm worker is eligible for reaping above min. |
acquire_timeout_ms |
30000 |
How long a parked waiter waits before no_capacity:acquire_timeout. |
reap_interval_ms |
15000 |
Reaper tick cadence. |
stale_heartbeat_ms |
600000 |
Heartbeat staleness threshold. |
drain_deadline_ms |
30000 |
How long drain waits for in-flight leases before force-destroying. |
max_workers_per_issue |
unset | Per-issue fairness cap on concurrent workers. |
co_residence |
unset | Opt-in required for slotsPerMachine > 1. |
max_concurrent_tunnels |
unset | Ceiling on concurrent reverse SSH tunnels, counted per distinct host (co-resident runs share one). |
Co-residence
max_in_flight is deprecated. It parses into slotsPerMachine, the number of runs that may share one machine. Co-residence (slotsPerMachine > 1) requires both a runtime per-run-claim-enforcement capability (the MCP gateway re-checks each request's per-run scoped claim server-side, so co-resident runs sharing one host and one reverse tunnel cannot authorize against each other) and an explicit co_residence: true opt-in. The CLI enforces this with a post-construction gate; the pool and domain layers do not. If you raise max_in_flight without setting co_residence, startup fails loud.
Spend caps
Keys under worker.worker_pool.spend cap paid machine usage.
| Key | Meaning |
|---|---|
max_concurrent_workers |
Hard ceiling on live workers. Blocks growth, surfaced as spend_cap. |
max_worker_seconds |
Lifetime worker-seconds budget. Gates acquire entirely (spend_cap). |
daily_worker_seconds |
Per-UTC-day worker-seconds budget. Gates acquire, persisted across restart. |
Worker-seconds are billed per lease from its own acquire timestamp. The daily accumulator rolls on UTC day change and persists to a spend.json sidecar next to the ledger file (dirname(ledgerPath)/spend.json). Daily spend is recorded fire-and-forget on the hot path; only a clean drain flushes the absolute total, so a crash can lose the last few unpersisted deltas.
The reaper
A single serial reaper runs every reap_interval_ms. A WeakSet in-progress guard prevents overlapping ticks. Each tick is five phases over the inventory, with every per-worker mutation inside a per-worker mutex:
- Reconcile with
driver.list(). Destroy pool-owned machines the driver reports but the pool does not know (gated on the pool being hydrated). Mark registered workers the driver no longer lists asDESTROYED. - Reap orphans. Destroy machines carrying the pool ownership label that no record claims.
- Reap TTL and idle. Flag
LEASEDworkers pastttl_ms. Reap idle workers abovemin, oldest-idle-first. - Probe and demote. Probe every warm idle worker; demote a failing one to
DEGRADED, then destroy it. - Top up. Provision toward
max(min, warm)within the spend budget.
The reaper never force-returns a LEASED worker. A long single-turn run with no heartbeat is never killed mid-flight. Cross-restart orphans are recovered only by hydrate (below), not by the reaper aborting active leases.
Ledger and crash recovery
For drivers that mark usesLedger (cloud and disposable backends), the pool keeps a write-ahead JSON ledger. It writes a provisional row before each provision and correlates the real worker after. Writes are atomic (temp file plus rename). The ledger is inert with zero filesystem I/O unless both the driver's usesLedger capability is true and a ledgerPath is supplied, so fake and static-ssh never touch disk.
On startup the pool calls hydrate(). It:
- Seeds daily spend from
spend.json. - Re-adopts every machine carrying the pool ownership label that
driver.list()reports, asWARM_IDLE. - Drops orphan ledger rows, keeping provisional rows younger than
ttl_ms.
driver.list() is retried up to 3 times with backoff. A paid driver (one with usesLedger or ephemeral) that still cannot list throws worker_pool_hydrate_failed and fails the daemon startup loud, so a crash never leaks paid machines behind a blind pool. A non-paid driver returns null and proceeds with hydrated=false.
Drain on reload and shutdown
Disabling the pool or shutting the daemon down calls drain(). Drain is idempotent, terminal, and awaitable. It stops the reaper, rejects new acquires and parked waiters, awaits the in-flight count reaching zero up to drain_deadline_ms, then force-destroys every worker inside the per-worker mutex so no paid machine leaks.
A workflow hot-reload that re-enables the pool bumps an internal drain epoch, so an orphaned drain from the old configuration bails out without destroying the re-enabled pool's workers. Driver hot-reload (swapDriver) is transactional: the pool resolves the new driver and builds a fresh ledger into locals first (a failure mutates nothing), captures the origin driver on every existing record, recycles idle workers on their origin backend, then commits. In-flight grows that captured a stale driver generation route their teardown to the origin driver so no paid worker is orphaned across the swap.
Built-in drivers
The pool resolves driver through a registry keyed on driver kind.
| Kind | Package | SSH-addressable | Ephemeral | Ledger | Notes |
|---|---|---|---|---|---|
fake |
@lorenz/worker-sdk |
no | no | no | In-memory. workerHost is fake://worker-<id>. For tests and dry runs. |
static-ssh |
@lorenz/static-worker |
yes | no | no | Round-robins a fixed ssh_hosts list. destroy forgets the address, never deletes a machine. |
docker |
extensions/docker-worker |
yes | yes | yes | Disposable containers via docker run -d. destroy is docker rm -f. |
The static-ssh driver requires an ssh_hosts (or ssh_hosts) option, else it throws static_ssh_hosts_required at construction. The docker driver requires an image, else docker_image_required. See Docker workers for that driver's full setup.
Drivers can also load out-of-tree by module specifier with an SDK-version handshake, so a custom backend ships without touching the pool engine. See out-of-tree drivers and the worker driver contract.
Audit events
The pool emits structured events for every lifecycle decision. Watch these in observability:
- Driver loading:
worker_pool_driver_loaded,worker_pool_driver_module_pinned. - Provision and probe:
worker_pool_provision_failed,worker_pool_warm_provision_failed,worker_pool_worker_unready,worker_pool_probe_failed,worker_pool_degraded. - Reaper and reconcile:
worker_pool_list_failed,worker_pool_reconcile_destroy_unknown,worker_pool_reconcile_missing,worker_pool_orphan_reaped,worker_pool_topup_budget_blocked,worker_pool_reaper_failed,worker_pool_destroy_failed. - Hydrate:
worker_pool_hydrate_failed,worker_pool_hydrate_list_failed,worker_pool_hydrate_orphan_dropped. - Ledger and callbacks:
worker_pool_ledger_write_failed,worker_pool_recycling_callback_failed,worker_pool_capacity_callback_failed,worker_pool_endpoint_release_failed.
Driver resolution and out-of-tree loading throw fail-loud errors at startup: worker_pool_driver_unavailable (unknown kind), worker_pool_driver_module_invalid, worker_pool_driver_sdk_mismatch, and worker_pool_driver_invalid_specifier. See the full list in the events reference.
See also
- Workers overview - how workers fit the run lifecycle
- Static SSH workers - the legacy
worker.ssh_hostspath and thestatic-sshdriver - Docker workers - the disposable-container driver
- Worker driver contract - building your own driver
- Out-of-tree drivers - loading a driver by module specifier
- Configuration reference - every
worker.worker_poolkey in one table