Skip to content

Upgrading Hivemind

Hivemind ships signed releases via the managed-update path (the hivemind-updater sidecar). The normal upgrade is a one-click apply from the admin Updates page; a manual install.sh re-run is the fallback. For v1, rollback is manual — back up before you click Apply.

The managed-update path (normal)

The hivemind-updater sidecar polls updates.seglamater.app on a schedule for new signed releases on your channel (your bundle pins the channel; the default is stable for new installs — install.sh sets CHANNEL="stable" unless you pass --channel). When a new release is available, the admin Updates page in the web UI surfaces it with the version, whether the release manifest's signature verified, and the release notes. Hit Apply to update.

How the Updates page learns about a release

The Hivemind server checks your channel for the newest release by itself: once when it starts, then every 6 hours, each wait lengthened by up to 10% at random. This check only refreshes what the Updates page shows; it never installs anything. Installing is always the Apply (or Install here) you press.

The top of the Updates page says how that check is going:

  • Last checked time: when the server last read the channel, whether the background check did it or you pressed Check for updates.
  • Last check failed: reason: the newest check could not read the channel. The release an earlier check found stays on the page; the failure does not hide it.
  • Automatic checks are off: the background check is switched off (below). Check for updates still reads the channel when you press it.

To change how often it checks, set HIVEMIND_UPDATES_CHECK_INTERVAL_SECS in your .env (seconds; values under 900, i.e. 15 minutes, are raised to 900). To turn it off, for example on an install that cannot reach the update host, set it to off (or 0), then recreate the server container. The setting reaches the server through docker-compose.yml, so an install whose compose file predates this setting needs the current docker-compose.yml first (see "Replacing docker-compose.yml by hand" below). The same values are reported by GET /api/admin/updates/status as checked_at, last_error, background_check and check_interval_secs.

The sidecar then:

  1. Fetches the release manifest for your channel and verifies its signature against the pinned cosign public key (/srv/hivemind/secrets/cosign_pub, pinned at install time). A manifest whose signature does not verify hard-fails the apply before anything else happens.
  2. Runs a pre-flight check in a one-shot container from the new image.
  3. Broadcasts a maintenance notice to connected users.
  4. Takes a pg_dump snapshot of the database.
  5. Pulls the new hivemind-server image from registry.seglamater.app, pinned to the digest named in the signed manifest — then re-inspects the pulled image and refuses to continue unless its digest matches.
  6. Re-renders the channel pointer in your bundle (the pinned image digest advances; channel stays the same).
  7. hivemind-server sends a signed POST /v1/apply to the sidecar — not the other way around. The orchestrator that drives the rest of this sequence lives inside hivemind-updater, not in-process in hivemind-server; the sidecar's own /v1/apply route is HMAC-protected with updater_token and reached over the internal updater_internal Docker network. From there, the orchestrator:
  8. Drains the current hivemind-server (graceful shutdown).
  9. Runs the schema migrations as a one-shot against the live database.
  10. Recreates the container at the new image digest.
  11. Waits for the healthcheck (/api/v1/health) to return ok.
  12. Reports success / failure back to the Updates page.

Where the signature sits, precisely

The cosign signature that gets verified is the one on the release manifest — not a separate signature object stored next to the image in the registry. The updater does not fetch a registry .sig, and does not need one: the signed manifest names the exact image digest, the pull is pinned to that digest, and the pulled image's digest is checked against it afterwards.

The practical guarantee is the one you would want from image signing. An attacker who can serve you a different image cannot get it installed without also forging the manifest signature, and that needs the private key. What they cannot do is substitute the image alone.

If the healthcheck doesn't pass within the grace window, the orchestrator rolls back automatically: it restores the pg_dump snapshot taken before the apply, then recreates the container at the prior image digest, then waits for that container to pass its own healthcheck. Restoring the database first matters — a rollback that only swapped the container back would leave the old server talking to a database the new version had already migrated. If the snapshot restore itself fails, the apply does not silently succeed or retry; it freezes in a FrozenMaintenanceRequired state that needs admin intervention, rather than leaving the instance in an inconsistent state under a "Rollback" label. You'll see a Rollback (or, on that failure path, a frozen) status on the Updates page — see "Where the updater logs live" below for where the real detail lives; a sudo docker compose logs hivemind-updater grep alone will not find it.

Expected runtime: 30 s - 2 min. There is a brief window (a few seconds) where the API is unreachable as the container is replaced. The Postgres volume is untouched — schema migrations run as a one-shot against the live database before the new container starts, and a migration failure triggers the same automatic rollback described above.

What this page describes but has not watched happen

The rollback sequence above (snapshot restore → container recreate → healthcheck) is real code, with unit tests covering a failed migration triggering a successful rollback and a failed snapshot restore triggering the freeze state (crates/hivemind-updater/src/update/orchestrator.rs). What this documentation sweep did NOT do is drive an actual hivemind-updater against a running Docker daemon and a real signed release and watch a rollback happen end to end. The 30s-2min runtime estimate and the "brief window" characterization are similarly unverified against a live run. Treat this section as code-accurate, not field-observed, until someone runs that drill. Tracked as HIVE-1094.

Manual upgrade (fallback)

If you want to skip the in-app path — e.g. you've held back the channel on the sidecar and you have a newer bundle in hand — you can re-run the install script with the new bundle:

cd /srv/hivemind                                 # or your install dir
sudo ./scripts/install.sh --bundle /path/to/new-bundle.json

install.sh is idempotent: it reuses your existing credential files, re-renders the environment file only for keys it doesn't already see, pulls the new image digest from the bundle, and runs docker compose up -d to recreate the affected containers. The same migrations + healthcheck gates apply. Time is similar (30 s - 2 min on a warm cache).

If you do not have a new bundle file, the simplest path is to ask your support contact to mint a fresh invite — that bootstrap fetches the latest bundle for your customer + re-applies cleanly.

Re-laying docker-compose.yml after an upgrade

An upgrade never rewrites your docker-compose.yml. It replaces container images. So if a release adds something to compose — a new environment variable, a new mount, a new service — an install created before that release will never gain it, and nothing in the upgrade output says so. Your instance keeps working; the new capability is simply absent.

That is worth checking after any upgrade, and it is the reason a feature can be "in the release" and still missing on your install.

First, diff — do not overwrite

Your compose file may carry local edits: ports, volumes, resource limits, an extra service. Overwriting it discards them. Compare first:

# from your install directory
diff -u docker-compose.yml /path/to/new/docker-compose.yml

Read the diff as two separate things: additions from the release (take these) and your own edits (keep these). Apply the additions by hand, or take the new file and re-apply your edits on top.

What to look for

Recent releases added these, and an install predating them will not have them:

Setting Goes to What you lose without it
HIVEMIND_HEALTH_URLS the runner the runner cannot check your endpoints' health
TLS_DOMAINS the runner the runner cannot check certificate expiry
HIVEMIND_OPENAI_API_KEY_FILE the server a key you supply as a FILE is ignored; the server falls back to HIVEMIND_OPENAI_API_KEY

If grep -c HIVEMIND_HEALTH_URLS docker-compose.yml returns 0, your compose is behind.

HIVEMIND_OPENAI_API_KEY_FILE was added in 0.1.95 and is checked the same way:

grep -c HIVEMIND_OPENAI_API_KEY_FILE docker-compose.yml

0 means the server never receives the path, so a key file you have configured is being ignored. Nothing breaks if you are not using the file form — the plain variable is unaffected. Re-laying compose is the only way this reaches an existing install: an upgrade replaces images, never this file.

Then apply

sudo docker compose up -d

Compose recreates only the containers whose definition changed. If you are adding runner variables, that is the runner alone:

sudo docker compose up -d --force-recreate hivemind-runner

If you are replacing the whole file, read the next section first — an install created before the admin-key change needs a secret file to exist before the new compose will start.

Replacing docker-compose.yml by hand: create the secret files first

The current docker-compose.yml hands the server its admin key as a file, secrets/admin_api_key, instead of an environment variable. An install set up before that change has the key only in .env as ADMIN_API_KEY. If you lay the new docker-compose.yml over such an install and run sudo docker compose up -d, Docker refuses to recreate the server:

invalid mount config for type bind: bind source path does not exist: …/secrets/admin_api_key

Your old containers keep running; nothing is lost. Create the file from the key you already have, before you replace docker-compose.yml. Run these in your install directory, as root or with sudo.

With this release's scripts/install.sh (it creates only what is missing, never changes an existing secret, and never prints the key):

cd /srv/hivemind                                 # or your install dir
sudo ./scripts/install.sh --provision-secrets-only

By hand. The key goes from .env straight into the file through a pipe: it is never typed, never on a command line, and never shown. The file is created readable by its owner only, and nothing happens if it already exists:

cd /srv/hivemind                                 # or your install dir
[ -e secrets/admin_api_key ] || ( umask 077; grep -E '^ADMIN_API_KEY=' .env | tail -1 \
    | cut -d= -f2- | sed -E 's/^"(.*)"$/\1/; s/^'"'"'(.*)'"'"'$/\1/' | tr -d '\n' \
    > secrets/admin_api_key )
sudo chown 1000:1000 secrets/admin_api_key       # the server runs as uid 1000
sudo chmod 600 secrets/admin_api_key

It keeps the admin key you have: anything already using it keeps working.

If you run the agent runner (the runner profile), it also needs secrets/claude_api_key, your Anthropic key for the runner. This one cannot be taken from .env: the Anthropic key .env may hold belongs to the server, not to the runner. Create the file empty and readable by its owner only, then put the key in it with an editor, never on a command line:

sudo install -m 600 /dev/null secrets/claude_api_key

Which upgrade paths do this for you:

  • The managed update (the update service swapping the image) never replaces docker-compose.yml, so it never meets this.
  • Re-running scripts/install.sh creates secrets/admin_api_key from .env on its own.
  • Replacing docker-compose.yml yourself needs the step above.

For any non-canary instance:

cd /srv/hivemind

# 1. Note the running version (so you can compare after).
curl -fsS http://localhost:8585/api/v1/health
# {"status":"ok","version":"0.1.75","integrations":{"chatalot_bot_ws":{"status":"inactive"}}}

# 2. Back up the postgres volume. The simplest snapshot is pg_dump.
#    --clean --if-exists is REQUIRED, not optional: without it the dump
#    contains bare CREATE statements that collide with the schema still
#    present after a failed upgrade, and the restore in step 2 of Rollback
#    silently does nothing. See the warning under Rollback.
sudo docker compose exec -T postgres pg_dump -U hivemind hivemind \
  --clean --if-exists \
  > "/var/backups/hivemind-pre-upgrade-$(date +%Y%m%d-%H%M%S).sql"

# 2b. Confirm the dump is complete AND self-cleaning (both, not either).
grep -c 'PostgreSQL database dump complete' \
  "/var/backups/hivemind-pre-upgrade-<timestamp>.sql"   # expect 1
grep -c '^DROP TABLE IF EXISTS' \
  "/var/backups/hivemind-pre-upgrade-<timestamp>.sql"   # expect > 0

# 3. If you're running Borg / Restic against /var/lib/docker/volumes/,
#    run a manual snapshot now so you have a known-good restore point.

# 4. Skim the boot log so any pre-existing WARNs aren't blamed on the upgrade.
sudo docker compose logs --tail=50 hivemind-server | grep -E "WARN|ERROR"

The pg_dump file is your point-in-time restore target if the upgrade introduces a schema problem. Hivemind migrations are forward-only — the v1 rollback strategy is restore-from-dump, not down-migration. See the "Rollback" section below.

Is that dump actually restorable?

A backup nobody has restored is a hypothesis. The snapshot tooling verifies a dump by checking that it is non-empty, gunzips, and contains CREATE TABLE — which is a plausibility check, not evidence that it will load.

The round trip is proven end-to-end by:

scripts/prove-snapshot-restore.sh

It provisions a throwaway cluster, applies the full migration set, takes a snapshot with the real hivemind-snapshot binary, and then proves three things: a restore onto a database that still has its schema succeeds and the data matches; a restore onto an empty database still works; and a dump taken without --clean fails on a schema-bearing database. The third is what makes the first two mean anything — without it, a pass shows only that something worked, not that --clean is what makes it work.

The script refuses rather than skipping when no database is available, because a green run that measured nothing is the failure it exists to remove.

Verify post-upgrade

# 1. Server is healthy at the new version.
curl -fsS http://localhost:8585/api/v1/health
# {"status":"ok","version":"<new-version>","integrations":{"chatalot_bot_ws":{"status":"inactive"}}}

# 2. Migrations applied without error.
sudo docker compose logs --tail=30 hivemind-server | grep -E "migration|database"

# 3. Provider init still green (no new LLM warnings).
sudo docker compose logs --tail=30 hivemind-server | grep -E "LLM|provider|integration credential"

# 4. Background workers came up.
sudo docker compose logs --tail=30 hivemind-server | grep -E "watchdog|prekey-replenish|bot-ws supervisor"

# 5. Sanity-call a real endpoint with your admin API key.
curl -fsS -H "X-API-Key: $ADMIN_API_KEY" \
  http://localhost:8585/api/v1/me | jq .

If any of those don't return the expected shape, see troubleshooting.md — most upgrade-time failures are the same boot-time failures as fresh installs (missing key, port conflict, DB not ready, etc.).

Rollback (v1: manual)

Hivemind v1 has no automated down-migration. Migrations are forward-only and Postgres state is the source of truth.

Your dump must have been taken with --clean --if-exists. After a failed upgrade your tables still exist — sudo docker compose down does not drop the pgdata volume — so a dump without --clean consists of CREATE statements that collide with what is already there. Replayed through a bare psql, that produces a handful of already exists errors, exits 0, and restores nothing, leaving you with the damaged data and no indication anything went wrong. This is why the restore below sets ON_ERROR_STOP=1 — so a restore that cannot proceed fails loudly instead of reporting success. Verified by execution, 2026-08-05 (HIVE-240fdeb1).

If you already hold a dump taken without --clean, it is not lost — it just cannot be replayed in place. Restore it into a new empty database and repoint Hivemind at that, rather than replaying it over the damaged one.

To roll back to the prior version:

cd /srv/hivemind

# 1. Stop the new stack.
sudo docker compose down

# 2. Restore the postgres dump you took pre-upgrade.
#    (If you didn't take one, restore from your borg / restic snapshot of
#    /var/lib/docker/volumes/hivemind_pgdata.)
sudo docker compose up -d postgres
sleep 5
cat /var/backups/hivemind-pre-upgrade-<timestamp>.sql | \
  sudo docker compose exec -T postgres \
    psql -U hivemind hivemind -v ON_ERROR_STOP=1 --single-transaction

# 3. Re-pin the OLD image digest. Easiest: ask Seglamater for the prior
#    bundle (they archive every cut), copy it into ./scripts/bundle.json,
#    then re-run install.sh:
sudo ./scripts/install.sh --bundle ./scripts/bundle.json

# 4. Bring the stack back up.
sudo docker compose up -d

Confirm /api/v1/health returns the prior version + the integration credential key load line is present. Open the Updates page in the admin UI and pin the prior channel pointer (or temporarily switch the sidecar to a paused channel) so the next sidecar tick doesn't immediately re-apply the version you just rolled back from.

Where the updater's real record lives

sudo docker compose logs hivemind-updater is NOT the source of truth for what an apply actually did. The full, step-by-step record is written to a SQLite table (apply_events) inside the sidecar itself, one row per phase (apply_started, manifest_verify_started, health_check_started, health_check_passed / health_check_failed, rollback_started, apply_completed, …). The sidecar's own internal API exposes that table at GET /v1/apply/{id}/events — but that route lives on the updater_internal Docker network, HMAC-authenticated, and is not reachable from outside the docker host. From the outside, the practical way to check what happened is:

curl -fsS -H "X-API-Key: $ADMIN_API_KEY" \
  http://localhost:8585/api/v1/admin/updates/apply/$APPLY_ID
# {"apply_id":"…","state":…,"outcome":"success"|"delivery_incomplete"|"rolled_back"|"frozen"|…,"error":null|"…","started_at":"…","ended_at":"…"}
#
# `delivery_incomplete` means the image swap SUCCEEDED and is healthy, and part
# of the release did not reach this installation. It is not a failure and it is
# not fixed by `docker compose up -d` — see the section below.

GET /api/v1/admin/updates/applies lists apply attempts (this is where you'd find the apply id for the command above if you don't already have it from the Updates page).

sudo docker compose logs hivemind-updater still shows SOMETHING, but not the lines this section used to quote — those exact strings don't exist anywhere in the code. What's actually there is terser and structured-field-heavy. A representative excerpt, using the real message text (crates/hivemind-updater/src/update/orchestrator.rs):

INFO apply_id=… manifest_url=… target_version=0.2.1 apply started
INFO apply_id=… apply completed successfully

or, on a failure that triggers rollback:

WARN apply_id=… failed_phase=health_check error=… apply failed post-disruption, rolling back
INFO apply_id=… rollback recovered successfully

(failed_phase is one of migrate, start_new, or health_check — whichever step actually failed. The health-check retry loop itself logs at debug, which the updater doesn't emit by default — the default level is info — so you will not see individual failed health probes in a normal log tail, only the final outcome.)

If a health check fails, the new container booted but /api/v1/health didn't return ok before the deadline. Pull the server logs for that container:

sudo docker compose logs --tail=200 hivemind-server | tail -200

and check troubleshooting.md. The most common causes: missing HIVEMIND_INTEGRATION_ENCRYPTION_KEY_FILE mount (compose mis-edit), DB schema mismatch (a migration partially applied), port conflict on the host (some other process took 8585 between stop_old and start_new).

delivery_incomplete: the upgrade worked and part of the release did not arrive

This outcome is not a failure, and the remedy is not the obvious one.

What it means. docker-compose.yml is delivered by the invite bootstrap (the script served at your https://s.seglamater.app/i/<invite-id> link), not by scripts/install.sh — install.sh never writes it, at install time or later. No local re-run of install.sh rewrites it. The managed updater rewrites only docker-compose.override.yml, and only the image digest inside it — it deliberately preserves every other byte, because customers do edit their compose and silently overwriting those edits would be worse than the problem it solves.

The consequence: a fix that lands in docker-compose.yml or .env.example reaches your installation only when you re-run the invite bootstrap. The new image arrives; the configuration change that belongs with it does not. Until this release the updater could not tell you that, so an instance could run a release for months while still carrying a configuration defect that release had fixed. That is exactly what happened to one installation: its compose file was two releases old, still carrying a defect the shipped code had fixed twice over.

The updater now checks, after every apply, whether the release's compose/installer requirements are actually present on the running container. If any are not — or if it could not verify one — the apply completes and is recorded as delivery_incomplete instead of success:

curl -fsS -H "X-API-Key: $ADMIN_API_KEY" \
  http://localhost:8585/api/v1/admin/updates/apply/$APPLY_ID
# {"outcome":"delivery_incomplete","error":"THE IMAGE SWAP SUCCEEDED AND IS HEALTHY. …"}

The error field is not an error. It carries the requirement ids, what is missing, and the exact action for each. Read it.

THE TRAP: docker compose up -d does NOT remediate this

The instinct is to recreate the container so it "picks up the fix". It does the opposite. sudo docker compose up -d server re-reads the same docker-compose.yml you already have — the old one — and re-applies the old declaration. On the case above it re-binds the credential the release had removed, and it discards the file the last managed update wrote into the container layer.

It exits 0. It looks like it worked. Nothing tells you otherwise except the delivery check, which is why the sidecar re-runs it on a timer rather than only during an apply.

The remedy

scripts/install.sh cannot fix this — it never writes docker-compose.yml. Re-run the invite bootstrap instead. It re-fetches docker-compose.yml, .env.example, and LICENSE for your current version, then hands off to install.sh for you — so one command does both:

curl -fsSL https://s.seglamater.app/i/<invite-id> -o hivemind-install.sh \
  && sudo bash hivemind-install.sh
sudo docker compose --profile updater up -d

This overwrites docker-compose.yml, .env.example, and LICENSE in place. If you have made manual edits to docker-compose.yml — adding a variable the shipped compose didn't declare yet is the common case — the bootstrap backs up each existing file first, verified, to a <file>.before-<timestamp> sibling in the same directory, before writing the new one (HIVE-fb491461). Nothing is lost, but the edit itself is gone from the live file the moment the bootstrap finishes: diff the backup against the new file and re-apply what you added.

If you don't have a live invite link — most installs won't, months after install — contact Seglamater support for one; making this self-service is an open question (HIVE-fb491461).

Then confirm it is closed:

sudo docker compose logs --tail=50 hivemind-updater | grep -i delivery
# INFO verified=4 delivery sweep: every shipped compose/installer requirement is present …

Checking on demand

The sidecar answers the same question at any time on its internal API (GET /v1/delivery, HMAC-authenticated on updater_internal), and re-evaluates on its own every HIVEMIND_UPDATER_DELIVERY_CHECK_SECS seconds — hourly by default, 0 to switch off. The sweep is what catches a sudo docker compose up -d that quietly undid a previous remediation; without it, the next reader would be your next managed update.

Log lines to look for:

ERROR unsatisfied=[…] changed=true DELIVERY DRIFT — A SECURITY PROPERTY OF THE INSTALLED RELEASE IS ABSENT …
WARN  unsatisfied=[…] changed=true delivery drift: part of the installed release is absent from this installation …
WARN  delivery sweep could not read the server container — this is the CHECK failing, NOT a finding about the installation

The third line is deliberately distinct from the first two: "I could not look" and "I looked and it is missing" are different facts, and only the second is about your installation.

The agent runner is NOT upgraded by an apply

An apply replaces the hivemind-server container and no other container. The agent runner (hivemind-runner, behind the compose profile runner) keeps running whatever image it started with, for as long as you leave it alone.

That is deliberate, not an oversight: recreating the runner restarts it, and anything it is running at that moment stops. An upgrade does not take that decision for you.

What an apply does do is record the release's runner image, so your configuration usually already names the right one. What it never does is replace the running container — so the version your runner reports is unchanged until you recreate it, however current your configuration says it is.

The consequence to plan around: a fix that lives in the runner does not reach you by updating. You get it when you recreate the runner, and an apply that succeeded says nothing about which runner you are on.

Before you start: where these commands run

Every command below runs in a shell on the host, from your installation directory — the one holding docker-compose.yml — and uses Docker Compose v2 (docker compose, two words). Run them anywhere else and Compose resolves a different project, or none: your override and pin files are part of this project only, and an edit that Compose never reads changes nothing.

The server container cannot do this itself. It has no access to Docker at all, deliberately — it cannot inspect, recreate or verify containers, including its own.

Your instance may still be able to do it for you, through the updater service that performs your upgrades: see Moving your agent runner to a new version, which explains when it can and when it hands you this procedure instead. This page is the procedure, and it is always available.

Step 1 — see what your runner is running now

sudo docker compose --profile runner ps --format '{{.Service}}\t{{.Image}}\t{{.State}}'

If that lists no hivemind-runner, one of two things is true, and they have different remedies:

  • The profile was never started here. Enabling agent execution for the first time is a separate setup step, not this procedure.
  • Your docker-compose.yml predates the agent runner entirely — it does not define the service, so there is nothing to start. Check with:
sudo docker compose --profile runner config --services

If hivemind-runner is absent from that list, you need a newer docker-compose.yml first. An update does not rewrite it: compose is delivered by the invite bootstrap, as described under delivery_incomplete above.

Step 2 — find the release's runner image

Every release manifest is public and needs no credentials to READ (pulling the image still needs your registry credentials):

curl -fsS https://updates.seglamater.app/hivemind/releases/<version>/manifest.json

Use registry.runner_ref_with_digest — the image and its digest together. If the manifest carries no runner_* fields at all, that release published no runner image and there is nothing to move to; stay where you are.

Step 3 — set your pin, in whichever file holds it

Installations differ here, so look before you edit:

  • .env contains HIVEMIND_RUNNER_IMAGE=… — edit it there.
  • docker-compose.override.yml has an image: line under hivemind-runner: — edit it there.
  • Neither — compose falls back to hivemind-runner:dev, a tag that exists in no registry, so the pull fails with a "not found" that looks like a registry problem and is not one. Add HIVEMIND_RUNNER_IMAGE= to .env.

Always pin the ref with its @sha256:… digest. A bare tag can be re-pointed; a digest names exact bytes.

Step 3b — confirm the pin took effect BEFORE you recreate

This is the step that does not care which file your pin lives in. Compose resolves every file plus .env and tells you the image it would actually use:

sudo docker compose --profile runner config --images | grep hivemind-runner

config resolves and prints; it creates, starts and changes nothing.

If that shows the OLD image, your edit is in a file this project does not read, or in a key Compose is not using — fix it before recreating, or you will recreate onto the same image and read that as a failed upgrade.

If it prints NOTHING, the service is not in this project at all: either you left out --profile runner, or your docker-compose.yml predates the runner (step 1). Empty output here means "not present", never "up to date".

Step 4 — recreate, and NAME THE PROFILE

sudo docker compose --profile runner up -d hivemind-runner

Without --profile runner, compose does not consider the service part of your project and will not recreate it. The command exits successfully having done nothing to the runner, which is the failure mode worth knowing about: it looks like it worked.

Step 5 — confirm from the container, not from the page

sudo docker compose --profile runner ps --format '{{.Service}}\t{{.Image}}'

Compare that image and digest with the manifest value from step 2.

The admin Machines page cannot confirm this for you today. No runner reports its version yet, so every machine reads version unreported — including one you just recreated on the current image. Treat that as "not known", never as "behind", and confirm from the container as above.

What the updater does and does not tell you about the runner

During an apply the updater tries to advance the runner's pin as well as the server's, and if it moves one it logs RUNNER RECREATE PENDING naming the exact command. Two limits are worth stating plainly:

  • It can only edit the pin file it has mounted, docker-compose.override.yml. That mount is deliberately narrow so the updater never has access to .env or your secrets. If your runner pin lives in .env, the updater cannot move it, and you will see a pin-sync failure for the runner in its log instead.
  • The pending-recreate line appears only when a pin actually moved. Not seeing it does not mean your runner is current — it more often means there was no pin for the updater to move. Step 1 and step 5 are how you know; the absence of a log line is not evidence.

A runner that connects through a reverse proxy: update it, then rotate its key

Runner images before this release sent their API key twice when they connected: in the X-API-Key header, which is what the server checks, and in the connection URL (/ws?api_key=...), which the server never read. A reverse proxy (Caddy, Traefik, nginx) writes request URLs to its access log. So if a runner on another machine reached your server through a reverse proxy, that runner's key is in the proxy's access log, and in any copy or log shipment of it. The runner that ships in the bundle connects to the server inside the compose network and does not pass through your proxy.

To close it, in this order:

  1. Update the runner image on that machine, using the steps above. The updated runner sends its key only in the header.
  2. Then rotate that runner's key (see Rotating an API key), and give the runner the new key. Rotating first would leave the old runner sending the new key in its URL.

Until the runner is updated, the server logs one warning per restart naming the account that runner uses: a runner sent its API key in the /ws URL. The warning names the account, never the key. The server's own request log no longer records query values on any route.

Moving the image pin onto a directory mount

If your installation was created before 0.1.86, its update service may be able to lose track of the image pin without any error, and the next sudo docker compose up or reboot can then start an older server than the one you are running. If the database has moved on since, that older server refuses to start. This section checks for that and fixes it for good. Nothing about it changes when you upgrade: the fix is a one-time change to your installation.

Why. After every managed update, the update service rewrites the image pin (the @sha256: digest in your override file) so that a recreate or a reboot keeps the release it applied. Installs created before 0.1.86 mount that file into the update service as a single file. A single-file mount follows the file's identity on disk, not its name, so anything that saves the file by writing a new copy and swapping it into place (sed -i, most editors, cp or mv over it) cuts the update service off from it. From then on its writes go to a copy nothing reads, and the file on the host stops changing. Installs from 0.1.86 on mount the pin's directory instead, which does not have this problem.

The Updates page tells you where you stand. On a current update service it reports the pin checks directly. On an older one it says the pin cannot be checked, which means: follow this section.

Until you have done this

  • Never edit the pin with sed -i or an editor that saves by replacing the file. If you must change it, write it in place, for example printf '%s\n' "$new_contents" > docker-compose.override.yml.
  • After any edit to it, restart the update service with docker restart <prefix>-updater (hivemind-updater unless you changed the prefix). sudo docker compose up -d does not do this: for an unchanged service it does nothing.

Step 1 — check the shape

From your installation directory:

docker inspect hivemind-updater \
  --format '{{range .Config.Env}}{{println .}}{{end}}' | grep HIVEMIND_UPDATER_PIN_FILE
docker inspect hivemind-updater \
  --format '{{range .Mounts}}{{.Source}} -> {{.Destination}}{{println}}{{end}}'
  • HIVEMIND_UPDATER_PIN_FILE=/host/pin/... and a mount ending -> /host/pin: you already have the directory mount. Stop here.
  • HIVEMIND_UPDATER_PIN_FILE=/host/image-pin.yml and a mount of a file ending -> /host/image-pin.yml: single-file mount. Continue. Note which file on the host the mount's source is; the steps below call it the pin file and assume it is docker-compose.override.yml. Substitute its real name throughout.

Step 2 — make sure the pin names the server you are running

docker inspect hivemind-server --format '{{.Config.Image}}'
grep '@sha256:' docker-compose.override.yml

The digest after hivemind-server@ must be the same in both. If it is not, stop and fix that first: a pin that names a different server is the outage this section prevents. Rewrite the digest in place (never sed -i), then restart the update service as above.

Step 3 — move the pin into a directory, without cutting the update service off

mkdir -p pin
ln docker-compose.override.yml pin/docker-compose.override.yml
ln -s pin/docker-compose.override.yml docker-compose.override.yml.new
mv -T docker-compose.override.yml.new docker-compose.override.yml

ln without -s makes the same file reachable from inside pin/, so nothing is copied and the running update service keeps seeing it. The last two lines swap the old name for a link in one step, so docker compose never finds the file missing.

Step 4 — point the update service at the directory

Add this under services: in pin/docker-compose.override.yml, keeping the file's two-space indentation:

  hivemind-updater:
    volumes:
      - ./pin:/host/pin:rw
    environment:
      HIVEMIND_UPDATER_PIN_FILE: /host/pin/docker-compose.override.yml

To have the update service report on this itself from now on, also pin it to the current release: open your channel manifest, read registry.updater_ref_with_digest (follow the pointer to the release manifest if needed, as in the runner section above), pull it with docker pull, and add image: <that reference> to the same block.

Then check that Compose reads it:

sudo docker compose --profile updater config | grep -A30 '^  hivemind-updater:' | grep -E '/host/pin|PIN_FILE'

Both the /host/pin mount and HIVEMIND_UPDATER_PIN_FILE: /host/pin/... must appear. If not, undo the edit and check the indentation.

Step 5 — recreate the update service, and confirm

sudo docker compose --profile updater up -d --no-deps --force-recreate hivemind-updater
docker exec hivemind-updater stat -c %i /host/pin/docker-compose.override.yml
stat -c %i pin/docker-compose.override.yml

The two numbers must match: the update service is now writing the file the host reads. From now on, editing the pin any way you like no longer detaches it.

With the Hivemind source, scripts/migrate-pin-mount.sh performs Steps 1–5 and every check above, and refuses anything it does not recognise. It changes nothing unless given --apply; add --updater-image <ref>@sha256:<digest> to move the update service to a current release in the same run.

Pausing or changing your channel

Channels are pinned in ./scripts/bundle.json. To pause auto-applies, ask Seglamater to issue a bundle pinned to a closed channel. To switch between canary, beta, and stable, the same — a new bundle from your support contact.

Hivemind does not currently expose a one-click channel switcher in the admin UI; that is on the roadmap. Until then, the channel is part of your bundle and managed out-of-band.