Self-hosted runner pools

Run your agents on machines you host. FerrFleet hands each run to one of your runners, and the code and the Claude credential stay on your side.


A runner pool is a set of machines you host that take the runs of the agents pointed at it. The agent keeps every trigger it has on FerrFleet: schedules, webhooks, tickets, a run started by hand. What changes is where the run executes. Instead of a Job in the FerrFleet cluster, it waits until one of your runners asks for work, and runs there.

Why a pool

  • Your code stays on your infrastructure. The runner clones the repository on your machine and the agent works on that checkout there.
  • So does your Claude credential. The runner starts Claude with an Anthropic API key from its own environment. FerrFleet never receives that key.
  • Outbound connections only. Every call goes from the runner to https://api.ferrfleet.com. A runner works behind a firewall that only allows outbound HTTPS, with no inbound port to open.
  • Your own capacity. Pool runs never take a slot in the FerrFleet cluster and never wait for one. How many run at once is how many runners you start.

A pool is different from setting an agent to run on your own pipeline (the GitHub Action). There, your CI creates the run. With a pool, FerrFleet creates the run as usual and your runners execute it. An agent is one or the other, not both.

Create a pool

In the app, open Runner pools under Operate, then New pool, and give it a name. Names are unique among your organization's pools.

The pool's registration token, starting with ffrp_, is shown once, right after the pool is created. Copy it into your secret store then: FerrFleet keeps only a hash of it and cannot show it again. If it is lost or leaks, Rotate token issues a new one and the old one stops working at once, without touching the agents on the pool.

The token only lets a runner take runs from that one pool and read how many are waiting. It opens nothing else in your organization.

Run runners

A runner is the open source FerrFleet Runner started as ferrfleet-runner agent. It needs:

  • FERRFLEET_API_URL: https://api.ferrfleet.com
  • FERRFLEET_POOL_TOKEN: the pool's ffrp_... token
  • ANTHROPIC_API_KEY: your Anthropic API key, see The Claude credential
  • FERRFLEET_RUNNER_NAME, optional: the label shown on the run page, the hostname by default

One runner executes one run at a time. To run several at once, start several runners.

Kubernetes with KEDA

The Helm chart scales a pool from zero with KEDA, which must already be installed in the cluster. It reads how many runs of the pool are waiting every 10 seconds and starts one Job per waiting run. Each Job takes one run and exits, so every run starts on a clean pod. With nothing waiting, nothing runs.

kubectl create namespace ferrfleet
kubectl -n ferrfleet create secret generic ferrfleet-pool \
  --from-literal=pool-token="$FERRFLEET_POOL_TOKEN" \
  --from-literal=anthropic-api-key="$ANTHROPIC_API_KEY"

helm install ferrfleet-runner oci://ghcr.io/ferrlabs/charts/ferrfleet-runner \
  --namespace ferrfleet \
  --set poolToken.existingSecret=ferrfleet-pool \
  --set claudeCredential.existingSecret=ferrfleet-pool

scaledJob.maxReplicaCount caps how many runs execute at once, 10 by default. The pods run as a non-root user with a read-only root filesystem and no capabilities, and their service account has no Kubernetes permissions. Every value of the chart is listed in the runner README.

Kubernetes without KEDA

The same chart runs a fixed number of long-lived runners instead:

helm install ferrfleet-runner oci://ghcr.io/ferrlabs/charts/ferrfleet-runner \
  --namespace ferrfleet \
  --set mode=deployment \
  --set deployment.replicas=3 \
  --set poolToken.existingSecret=ferrfleet-pool \
  --set claudeCredential.existingSecret=ferrfleet-pool

Docker Compose

On a host without Kubernetes, the compose example runs a fixed number of runners with the same hardening. Copy docker-compose.yml and .env.example from it, then:

cp .env.example .env
$EDITOR .env
docker compose up -d

.env holds the pool token and the API key, so keep it out of version control. RUNNER_REPLICAS sets how many runners start.

A single container

docker run -d --restart unless-stopped \
  -e FERRFLEET_API_URL=https://api.ferrfleet.com \
  -e FERRFLEET_POOL_TOKEN \
  -e FERRFLEET_RUNNER_NAME=build-farm-07 \
  -e ANTHROPIC_API_KEY \
  ghcr.io/ferrlabs/ferrfleet/runner:1 agent

Stopping a runner

On SIGTERM a runner stops asking for work and lets the run in progress finish. Give it a grace period as long as your agents' timeout: the chart allows 30 minutes, and so does the compose example. A runner killed before then is handled like one that goes quiet, see below.

Point an agent at a pool

On the agent page, Edit, then the Execution tab. With Runner on FerrFleet, set Run on to the pool. Setting it back to FerrFleet cluster returns the agent to our cluster.

The change applies to runs created from then on. Runs already waiting stay with the pool they were created for.

What happens to a run

It waits for a runner. A pool run is never started in the FerrFleet cluster, even if no runner is up. Runs go to runners in the order they arrived, oldest first, and two runners never get the same run. The run page shows which runner took it.

The runner checks in. From the moment a runner takes a run until it ends, it checks in with FerrFleet every 30 seconds. That is also how a cancellation or a newer commit reaches it: the next check-in tells it to stop, so it stops within 30 seconds.

A runner that goes quiet. When the check-ins stop for 90 seconds, FerrFleet settles the run within the next minute:

  • If the runner had not started the run yet, the run goes back to the pool and the next runner that asks gets it. Nothing ran, so nothing is lost.
  • If it had started, the run fails, with exit code 126 and a line in the transcript saying the runner stopped responding. It is not handed out again: it may already have pushed a branch or commented on a pull request, and running it twice would do that twice.

Pool runs are not retried automatically.

The 6 hour wait. A run no runner takes waits up to 6 hours from its creation, then ends as timed out, and the run page says no runner came for it. That leaves time for a pool scaled to zero to start a machine, and for a pool stopped for the night to find its runs in the morning. Once a runner starts the run, the agent's usual run timeout applies.

Revoking a pool. A pool that agents still point at cannot be revoked: move them to another pool or back to the FerrFleet cluster first. Nothing falls back to our cluster on its own. Once a pool is revoked, its token stops working and its waiting runs are cancelled. A run a runner already started finishes there.

The Claude credential

Runners use an Anthropic API key from their own environment, in ANTHROPIC_API_KEY. It is the default in the Helm chart, the compose example and the container image, and it is the option we recommend. Create the key in an Anthropic Console workspace you keep for these runs, so their spend and rate limits show on their own.

The key stays on your machines. The runner never sends it to FerrFleet, and FerrFleet never asks for it. To change it, update it in your secret store and restart the runners: nothing changes on the FerrFleet side.

Reference