Upgrading Hivemind¶
Hivemind ships signed releases via the managed-update path (the
hivemind-updater sidecar). The normal upgrade is a one-click apply from
the admin Updates page; a manual install.sh re-run is the fallback. For
v1, rollback is manual — back up before you click Apply.
The managed-update path (normal)¶
The hivemind-updater sidecar polls updates.seglamater.app on a
schedule for new signed releases on your channel (your bundle pins
the channel; the default is stable for new installs — install.sh
sets CHANNEL="stable" unless you pass --channel). When a new release
is available, the admin Updates page in the web UI surfaces it with the
version, whether the release manifest's signature verified, and the release
notes. Hit Apply to update.
How the Updates page learns about a release¶
The Hivemind server checks your channel for the newest release by itself: once when it starts, then every 6 hours, each wait lengthened by up to 10% at random. This check only refreshes what the Updates page shows; it never installs anything. Installing is always the Apply (or Install here) you press.
The top of the Updates page says how that check is going:
- Last checked time: when the server last read the channel, whether the background check did it or you pressed Check for updates.
- Last check failed: reason: the newest check could not read the channel. The release an earlier check found stays on the page; the failure does not hide it.
- Automatic checks are off: the background check is switched off (below). Check for updates still reads the channel when you press it.
To change how often it checks, set HIVEMIND_UPDATES_CHECK_INTERVAL_SECS in
your .env (seconds; values under 900, i.e. 15 minutes, are raised to 900).
To turn it off, for example on an install that cannot reach the update host,
set it to off (or 0), then recreate the server container. The setting
reaches the server through docker-compose.yml, so an install whose compose
file predates this setting needs the current docker-compose.yml first (see
"Replacing docker-compose.yml by hand" below). The same values are reported
by GET /api/admin/updates/status as checked_at, last_error,
background_check and check_interval_secs.
The sidecar then:
- Fetches the release manifest for your channel and verifies its signature
against the pinned cosign public key (
/srv/hivemind/secrets/cosign_pub, pinned at install time). A manifest whose signature does not verify hard-fails the apply before anything else happens. - Runs a pre-flight check in a one-shot container from the new image.
- Broadcasts a maintenance notice to connected users.
- Takes a
pg_dumpsnapshot of the database. - Pulls the new
hivemind-serverimage fromregistry.seglamater.app, pinned to the digest named in the signed manifest — then re-inspects the pulled image and refuses to continue unless its digest matches. - Re-renders the channel pointer in your bundle (the pinned image digest
advances;
channelstays the same). hivemind-serversends a signedPOST /v1/applyto the sidecar — not the other way around. The orchestrator that drives the rest of this sequence lives insidehivemind-updater, not in-process inhivemind-server; the sidecar's own/v1/applyroute is HMAC-protected withupdater_tokenand reached over the internalupdater_internalDocker network. From there, the orchestrator:- Drains the current
hivemind-server(graceful shutdown). - Runs the schema migrations as a one-shot against the live database.
- Recreates the container at the new image digest.
- Waits for the healthcheck (
/api/v1/health) to return ok. - Reports success / failure back to the Updates page.
Where the signature sits, precisely
The cosign signature that gets verified is the one on the release
manifest — not a separate signature object stored next to the image in the
registry. The updater does not fetch a registry .sig, and does not need
one: the signed manifest names the exact image digest, the pull is pinned to
that digest, and the pulled image's digest is checked against it afterwards.
The practical guarantee is the one you would want from image signing. An attacker who can serve you a different image cannot get it installed without also forging the manifest signature, and that needs the private key. What they cannot do is substitute the image alone.
If the healthcheck doesn't pass within the grace window, the orchestrator
rolls back automatically: it restores the pg_dump snapshot taken
before the apply, then recreates the container at the prior image
digest, then waits for that container to pass its own healthcheck.
Restoring the database first matters — a rollback that only swapped
the container back would leave the old server talking to a database
the new version had already migrated. If the snapshot restore itself
fails, the apply does not silently succeed or retry; it freezes in a
FrozenMaintenanceRequired state that needs admin intervention, rather
than leaving the instance in an inconsistent state under a "Rollback"
label. You'll see a Rollback (or, on that failure path, a frozen)
status on the Updates page — see "Where the updater logs live" below
for where the real detail lives; a sudo docker compose logs
hivemind-updater grep alone will not find it.
Expected runtime: 30 s - 2 min. There is a brief window (a few seconds) where the API is unreachable as the container is replaced. The Postgres volume is untouched — schema migrations run as a one-shot against the live database before the new container starts, and a migration failure triggers the same automatic rollback described above.
What this page describes but has not watched happen
The rollback sequence above (snapshot restore → container recreate →
healthcheck) is real code, with unit tests covering a failed
migration triggering a successful rollback and a failed snapshot
restore triggering the freeze state
(crates/hivemind-updater/src/update/orchestrator.rs). What this
documentation sweep did NOT do is drive an actual hivemind-updater
against a running Docker daemon and a real signed release and watch
a rollback happen end to end. The 30s-2min runtime estimate and the
"brief window" characterization are similarly unverified against a
live run. Treat this section as code-accurate, not field-observed,
until someone runs that drill. Tracked as HIVE-1094.
Manual upgrade (fallback)¶
If you want to skip the in-app path — e.g. you've held back the channel on the sidecar and you have a newer bundle in hand — you can re-run the install script with the new bundle:
install.sh is idempotent: it reuses your existing credential files,
re-renders the environment file only for keys it doesn't already see,
pulls the new image digest from the bundle, and runs docker compose up
-d to recreate the affected containers. The same migrations + healthcheck
gates apply. Time is similar (30 s - 2 min on a warm cache).
If you do not have a new bundle file, the simplest path is to ask your support contact to mint a fresh invite — that bootstrap fetches the latest bundle for your customer + re-applies cleanly.
Re-laying docker-compose.yml after an upgrade¶
An upgrade never rewrites your docker-compose.yml. It replaces container images. So if
a release adds something to compose — a new environment variable, a new mount, a new service
— an install created before that release will never gain it, and nothing in the upgrade
output says so. Your instance keeps working; the new capability is simply absent.
That is worth checking after any upgrade, and it is the reason a feature can be "in the release" and still missing on your install.
First, diff — do not overwrite¶
Your compose file may carry local edits: ports, volumes, resource limits, an extra service. Overwriting it discards them. Compare first:
Read the diff as two separate things: additions from the release (take these) and your own edits (keep these). Apply the additions by hand, or take the new file and re-apply your edits on top.
What to look for¶
Recent releases added these, and an install predating them will not have them:
| Setting | Goes to | What you lose without it |
|---|---|---|
HIVEMIND_HEALTH_URLS |
the runner | the runner cannot check your endpoints' health |
TLS_DOMAINS |
the runner | the runner cannot check certificate expiry |
HIVEMIND_OPENAI_API_KEY_FILE |
the server | a key you supply as a FILE is ignored; the server falls back to HIVEMIND_OPENAI_API_KEY |
If grep -c HIVEMIND_HEALTH_URLS docker-compose.yml returns 0, your compose is behind.
HIVEMIND_OPENAI_API_KEY_FILE was added in 0.1.95 and is checked the same way:
0 means the server never receives the path, so a key file you have configured is being
ignored. Nothing breaks if you are not using the file form — the plain variable is
unaffected. Re-laying compose is the only way this reaches an existing install: an
upgrade replaces images, never this file.
Then apply¶
Compose recreates only the containers whose definition changed. If you are adding runner variables, that is the runner alone:
If you are replacing the whole file, read the next section first — an install created before the admin-key change needs a secret file to exist before the new compose will start.
Replacing docker-compose.yml by hand: create the secret files first¶
The current docker-compose.yml hands the server its admin key as a file,
secrets/admin_api_key, instead of an environment variable. An install set up
before that change has the key only in .env as ADMIN_API_KEY. If you lay the
new docker-compose.yml over such an install and run sudo docker compose up -d,
Docker refuses to recreate the server:
Your old containers keep running; nothing is lost. Create the file from the key you
already have, before you replace docker-compose.yml. Run these in your install
directory, as root or with sudo.
With this release's scripts/install.sh (it creates only what is missing, never
changes an existing secret, and never prints the key):
By hand. The key goes from .env straight into the file through a pipe: it is never
typed, never on a command line, and never shown. The file is created readable by its
owner only, and nothing happens if it already exists:
cd /srv/hivemind # or your install dir
[ -e secrets/admin_api_key ] || ( umask 077; grep -E '^ADMIN_API_KEY=' .env | tail -1 \
| cut -d= -f2- | sed -E 's/^"(.*)"$/\1/; s/^'"'"'(.*)'"'"'$/\1/' | tr -d '\n' \
> secrets/admin_api_key )
sudo chown 1000:1000 secrets/admin_api_key # the server runs as uid 1000
sudo chmod 600 secrets/admin_api_key
It keeps the admin key you have: anything already using it keeps working.
If you run the agent runner (the runner profile), it also needs
secrets/claude_api_key, your Anthropic key for the runner. This one cannot be taken
from .env: the Anthropic key .env may hold belongs to the server, not to the
runner. Create the file empty and readable by its owner only, then put the key in it
with an editor, never on a command line:
Which upgrade paths do this for you:
- The managed update (the update service swapping the image) never replaces
docker-compose.yml, so it never meets this. - Re-running
scripts/install.shcreatessecrets/admin_api_keyfrom.envon its own. - Replacing
docker-compose.ymlyourself needs the step above.
Pre-upgrade checklist (recommended)¶
For any non-canary instance:
cd /srv/hivemind
# 1. Note the running version (so you can compare after).
curl -fsS http://localhost:8585/api/v1/health
# {"status":"ok","version":"0.1.75","integrations":{"chatalot_bot_ws":{"status":"inactive"}}}
# 2. Back up the postgres volume. The simplest snapshot is pg_dump.
# --clean --if-exists is REQUIRED, not optional: without it the dump
# contains bare CREATE statements that collide with the schema still
# present after a failed upgrade, and the restore in step 2 of Rollback
# silently does nothing. See the warning under Rollback.
sudo docker compose exec -T postgres pg_dump -U hivemind hivemind \
--clean --if-exists \
> "/var/backups/hivemind-pre-upgrade-$(date +%Y%m%d-%H%M%S).sql"
# 2b. Confirm the dump is complete AND self-cleaning (both, not either).
grep -c 'PostgreSQL database dump complete' \
"/var/backups/hivemind-pre-upgrade-<timestamp>.sql" # expect 1
grep -c '^DROP TABLE IF EXISTS' \
"/var/backups/hivemind-pre-upgrade-<timestamp>.sql" # expect > 0
# 3. If you're running Borg / Restic against /var/lib/docker/volumes/,
# run a manual snapshot now so you have a known-good restore point.
# 4. Skim the boot log so any pre-existing WARNs aren't blamed on the upgrade.
sudo docker compose logs --tail=50 hivemind-server | grep -E "WARN|ERROR"
The pg_dump file is your point-in-time restore target if the upgrade introduces a schema problem. Hivemind migrations are forward-only — the v1 rollback strategy is restore-from-dump, not down-migration. See the "Rollback" section below.
Is that dump actually restorable?¶
A backup nobody has restored is a hypothesis. The snapshot tooling verifies a
dump by checking that it is non-empty, gunzips, and contains CREATE TABLE —
which is a plausibility check, not evidence that it will load.
The round trip is proven end-to-end by:
It provisions a throwaway cluster, applies the full migration set, takes a
snapshot with the real hivemind-snapshot binary, and then proves three things:
a restore onto a database that still has its schema succeeds and the data
matches; a restore onto an empty database still works; and a dump taken
without --clean fails on a schema-bearing database. The third is what
makes the first two mean anything — without it, a pass shows only that something
worked, not that --clean is what makes it work.
The script refuses rather than skipping when no database is available, because a green run that measured nothing is the failure it exists to remove.
Verify post-upgrade¶
# 1. Server is healthy at the new version.
curl -fsS http://localhost:8585/api/v1/health
# {"status":"ok","version":"<new-version>","integrations":{"chatalot_bot_ws":{"status":"inactive"}}}
# 2. Migrations applied without error.
sudo docker compose logs --tail=30 hivemind-server | grep -E "migration|database"
# 3. Provider init still green (no new LLM warnings).
sudo docker compose logs --tail=30 hivemind-server | grep -E "LLM|provider|integration credential"
# 4. Background workers came up.
sudo docker compose logs --tail=30 hivemind-server | grep -E "watchdog|prekey-replenish|bot-ws supervisor"
# 5. Sanity-call a real endpoint with your admin API key.
curl -fsS -H "X-API-Key: $ADMIN_API_KEY" \
http://localhost:8585/api/v1/me | jq .
If any of those don't return the expected shape, see troubleshooting.md
— most upgrade-time failures are the same boot-time failures as fresh
installs (missing key, port conflict, DB not ready, etc.).
Rollback (v1: manual)¶
Hivemind v1 has no automated down-migration. Migrations are forward-only and Postgres state is the source of truth.
Your dump must have been taken with
--clean --if-exists. After a failed upgrade your tables still exist —sudo docker compose downdoes not drop thepgdatavolume — so a dump without--cleanconsists ofCREATEstatements that collide with what is already there. Replayed through a barepsql, that produces a handful ofalready existserrors, exits 0, and restores nothing, leaving you with the damaged data and no indication anything went wrong. This is why the restore below setsON_ERROR_STOP=1— so a restore that cannot proceed fails loudly instead of reporting success. Verified by execution, 2026-08-05 (HIVE-240fdeb1).If you already hold a dump taken without
--clean, it is not lost — it just cannot be replayed in place. Restore it into a new empty database and repoint Hivemind at that, rather than replaying it over the damaged one.
To roll back to the prior version:
cd /srv/hivemind
# 1. Stop the new stack.
sudo docker compose down
# 2. Restore the postgres dump you took pre-upgrade.
# (If you didn't take one, restore from your borg / restic snapshot of
# /var/lib/docker/volumes/hivemind_pgdata.)
sudo docker compose up -d postgres
sleep 5
cat /var/backups/hivemind-pre-upgrade-<timestamp>.sql | \
sudo docker compose exec -T postgres \
psql -U hivemind hivemind -v ON_ERROR_STOP=1 --single-transaction
# 3. Re-pin the OLD image digest. Easiest: ask Seglamater for the prior
# bundle (they archive every cut), copy it into ./scripts/bundle.json,
# then re-run install.sh:
sudo ./scripts/install.sh --bundle ./scripts/bundle.json
# 4. Bring the stack back up.
sudo docker compose up -d
Confirm /api/v1/health returns the prior version + the integration
credential key load line is present. Open the Updates page in the admin
UI and pin the prior channel pointer (or temporarily switch the
sidecar to a paused channel) so the next sidecar tick doesn't immediately
re-apply the version you just rolled back from.
Where the updater's real record lives¶
sudo docker compose logs hivemind-updater is NOT the source of truth for
what an apply actually did. The full, step-by-step record is written
to a SQLite table (apply_events) inside the sidecar itself, one row
per phase (apply_started, manifest_verify_started,
health_check_started, health_check_passed /
health_check_failed, rollback_started, apply_completed, …). The
sidecar's own internal API exposes that table at
GET /v1/apply/{id}/events — but that route lives on the
updater_internal Docker network, HMAC-authenticated, and is not
reachable from outside the docker host. From the outside, the
practical way to check what happened is:
curl -fsS -H "X-API-Key: $ADMIN_API_KEY" \
http://localhost:8585/api/v1/admin/updates/apply/$APPLY_ID
# {"apply_id":"…","state":…,"outcome":"success"|"delivery_incomplete"|"rolled_back"|"frozen"|…,"error":null|"…","started_at":"…","ended_at":"…"}
#
# `delivery_incomplete` means the image swap SUCCEEDED and is healthy, and part
# of the release did not reach this installation. It is not a failure and it is
# not fixed by `docker compose up -d` — see the section below.
GET /api/v1/admin/updates/applies lists apply attempts (this is
where you'd find the apply id for the command above if you don't
already have it from the Updates page).
sudo docker compose logs hivemind-updater still shows SOMETHING, but not
the lines this section used to quote — those exact strings don't
exist anywhere in the code. What's actually there is terser and
structured-field-heavy. A representative excerpt, using the real
message text (crates/hivemind-updater/src/update/orchestrator.rs):
INFO apply_id=… manifest_url=… target_version=0.2.1 apply started
INFO apply_id=… apply completed successfully
or, on a failure that triggers rollback:
WARN apply_id=… failed_phase=health_check error=… apply failed post-disruption, rolling back
INFO apply_id=… rollback recovered successfully
(failed_phase is one of migrate, start_new, or health_check —
whichever step actually failed. The health-check retry loop itself
logs at debug, which the updater doesn't emit by default — the
default level is info — so you will not see individual failed
health probes in a normal log tail, only the final outcome.)
If a health check fails, the new container booted but /api/v1/health
didn't return ok before the deadline. Pull the server logs for that
container:
and check troubleshooting.md. The most common causes: missing
HIVEMIND_INTEGRATION_ENCRYPTION_KEY_FILE mount (compose mis-edit), DB
schema mismatch (a migration partially applied), port conflict on the
host (some other process took 8585 between stop_old and start_new).
delivery_incomplete: the upgrade worked and part of the release did not arrive¶
This outcome is not a failure, and the remedy is not the obvious one.
What it means. docker-compose.yml is delivered by the invite bootstrap
(the script served at your https://s.seglamater.app/i/<invite-id> link), not by
scripts/install.sh — install.sh never writes it, at install time or later.
No local re-run of install.sh rewrites it. The managed updater rewrites
only docker-compose.override.yml, and only the image digest inside it — it
deliberately preserves every other byte, because customers do edit their compose
and silently overwriting those edits would be worse than the problem it solves.
The consequence: a fix that lands in docker-compose.yml or .env.example
reaches your installation only when you re-run the invite bootstrap. The new
image arrives; the configuration change that belongs with it does not. Until this
release the updater could not tell you that, so an instance could run a release
for months while still carrying a configuration defect that release had fixed.
That is exactly what happened to one installation: its compose file was two
releases old, still carrying a defect the shipped code had fixed twice over.
The updater now checks, after every apply, whether the release's
compose/installer requirements are actually present on the running container.
If any are not — or if it could not verify one — the apply completes and is
recorded as delivery_incomplete instead of success:
curl -fsS -H "X-API-Key: $ADMIN_API_KEY" \
http://localhost:8585/api/v1/admin/updates/apply/$APPLY_ID
# {"outcome":"delivery_incomplete","error":"THE IMAGE SWAP SUCCEEDED AND IS HEALTHY. …"}
The error field is not an error. It carries the requirement ids, what is
missing, and the exact action for each. Read it.
THE TRAP: docker compose up -d does NOT remediate this¶
The instinct is to recreate the container so it "picks up the fix". It does
the opposite. sudo docker compose up -d server re-reads the same
docker-compose.yml you already have — the old one — and re-applies the old
declaration. On the case above it re-binds the credential the release had
removed, and it discards the file the last managed update wrote into the
container layer.
It exits 0. It looks like it worked. Nothing tells you otherwise except the delivery check, which is why the sidecar re-runs it on a timer rather than only during an apply.
The remedy¶
scripts/install.sh cannot fix this — it never writes docker-compose.yml.
Re-run the invite bootstrap instead. It re-fetches docker-compose.yml,
.env.example, and LICENSE for your current version, then hands off to
install.sh for you — so one command does both:
curl -fsSL https://s.seglamater.app/i/<invite-id> -o hivemind-install.sh \
&& sudo bash hivemind-install.sh
sudo docker compose --profile updater up -d
This overwrites docker-compose.yml, .env.example, and LICENSE in
place. If you have made manual edits to docker-compose.yml — adding a
variable the shipped compose didn't declare yet is the common case — the
bootstrap backs up each existing file first, verified, to a
<file>.before-<timestamp> sibling in the same directory, before writing the
new one (HIVE-fb491461). Nothing is lost, but the edit itself is gone from the
live file the moment the bootstrap finishes: diff the backup against the new
file and re-apply what you added.
If you don't have a live invite link — most installs won't, months after install — contact Seglamater support for one; making this self-service is an open question (HIVE-fb491461).
Then confirm it is closed:
sudo docker compose logs --tail=50 hivemind-updater | grep -i delivery
# INFO verified=4 delivery sweep: every shipped compose/installer requirement is present …
Checking on demand¶
The sidecar answers the same question at any time on its internal API
(GET /v1/delivery, HMAC-authenticated on updater_internal), and re-evaluates
on its own every HIVEMIND_UPDATER_DELIVERY_CHECK_SECS seconds — hourly by
default, 0 to switch off. The sweep is what catches a sudo docker compose up -d
that quietly undid a previous remediation; without it, the next reader would be
your next managed update.
Log lines to look for:
ERROR unsatisfied=[…] changed=true DELIVERY DRIFT — A SECURITY PROPERTY OF THE INSTALLED RELEASE IS ABSENT …
WARN unsatisfied=[…] changed=true delivery drift: part of the installed release is absent from this installation …
WARN delivery sweep could not read the server container — this is the CHECK failing, NOT a finding about the installation
The third line is deliberately distinct from the first two: "I could not look" and "I looked and it is missing" are different facts, and only the second is about your installation.
The agent runner is NOT upgraded by an apply¶
An apply replaces the hivemind-server container and no other container. The
agent runner (hivemind-runner, behind the compose profile runner) keeps
running whatever image it started with, for as long as you leave it alone.
That is deliberate, not an oversight: recreating the runner restarts it, and anything it is running at that moment stops. An upgrade does not take that decision for you.
What an apply does do is record the release's runner image, so your configuration usually already names the right one. What it never does is replace the running container — so the version your runner reports is unchanged until you recreate it, however current your configuration says it is.
The consequence to plan around: a fix that lives in the runner does not reach you by updating. You get it when you recreate the runner, and an apply that succeeded says nothing about which runner you are on.
Before you start: where these commands run¶
Every command below runs in a shell on the host, from your installation
directory — the one holding docker-compose.yml — and uses Docker Compose v2
(docker compose, two words). Run them anywhere else and Compose resolves a
different project, or none: your override and pin files are part of this
project only, and an edit that Compose never reads changes nothing.
The server container cannot do this itself. It has no access to Docker at all, deliberately — it cannot inspect, recreate or verify containers, including its own.
Your instance may still be able to do it for you, through the updater service that performs your upgrades: see Moving your agent runner to a new version, which explains when it can and when it hands you this procedure instead. This page is the procedure, and it is always available.
Step 1 — see what your runner is running now¶
If that lists no hivemind-runner, one of two things is true, and they have
different remedies:
- The profile was never started here. Enabling agent execution for the first time is a separate setup step, not this procedure.
- Your
docker-compose.ymlpredates the agent runner entirely — it does not define the service, so there is nothing to start. Check with:
If hivemind-runner is absent from that list, you need a newer
docker-compose.yml first. An update does not rewrite it: compose is
delivered by the invite bootstrap, as described under delivery_incomplete
above.
Step 2 — find the release's runner image¶
Every release manifest is public and needs no credentials to READ (pulling the image still needs your registry credentials):
Use registry.runner_ref_with_digest — the image and its digest together. If the
manifest carries no runner_* fields at all, that release published no runner
image and there is nothing to move to; stay where you are.
Step 3 — set your pin, in whichever file holds it¶
Installations differ here, so look before you edit:
.envcontainsHIVEMIND_RUNNER_IMAGE=…— edit it there.docker-compose.override.ymlhas animage:line underhivemind-runner:— edit it there.- Neither — compose falls back to
hivemind-runner:dev, a tag that exists in no registry, so the pull fails with a "not found" that looks like a registry problem and is not one. AddHIVEMIND_RUNNER_IMAGE=to.env.
Always pin the ref with its @sha256:… digest. A bare tag can be
re-pointed; a digest names exact bytes.
Step 3b — confirm the pin took effect BEFORE you recreate¶
This is the step that does not care which file your pin lives in. Compose
resolves every file plus .env and tells you the image it would actually use:
config resolves and prints; it creates, starts and changes nothing.
If that shows the OLD image, your edit is in a file this project does not read, or in a key Compose is not using — fix it before recreating, or you will recreate onto the same image and read that as a failed upgrade.
If it prints NOTHING, the service is not in this project at all: either you left
out --profile runner, or your docker-compose.yml predates the runner (step 1).
Empty output here means "not present", never "up to date".
Step 4 — recreate, and NAME THE PROFILE¶
Without --profile runner, compose does not consider the service part of your
project and will not recreate it. The command exits successfully having done
nothing to the runner, which is the failure mode worth knowing about: it looks
like it worked.
Step 5 — confirm from the container, not from the page¶
Compare that image and digest with the manifest value from step 2.
The admin Machines page cannot confirm this for you today. No runner reports its version yet, so every machine reads version unreported — including one you just recreated on the current image. Treat that as "not known", never as "behind", and confirm from the container as above.
What the updater does and does not tell you about the runner¶
During an apply the updater tries to advance the runner's pin as well as the server's, and if it moves one it logs RUNNER RECREATE PENDING naming the exact command. Two limits are worth stating plainly:
- It can only edit the pin file it has mounted,
docker-compose.override.yml. That mount is deliberately narrow so the updater never has access to.envor your secrets. If your runner pin lives in.env, the updater cannot move it, and you will see a pin-sync failure for the runner in its log instead. - The pending-recreate line appears only when a pin actually moved. Not seeing it does not mean your runner is current — it more often means there was no pin for the updater to move. Step 1 and step 5 are how you know; the absence of a log line is not evidence.
A runner that connects through a reverse proxy: update it, then rotate its key¶
Runner images before this release sent their API key twice when they
connected: in the X-API-Key header, which is what the server checks, and in
the connection URL (/ws?api_key=...), which the server never read. A reverse
proxy (Caddy, Traefik, nginx) writes request URLs to its access log. So if a
runner on another machine reached your server through a reverse proxy, that
runner's key is in the proxy's access log, and in any copy or log shipment
of it. The runner that ships in the bundle connects to the server inside the
compose network and does not pass through your proxy.
To close it, in this order:
- Update the runner image on that machine, using the steps above. The updated runner sends its key only in the header.
- Then rotate that runner's key (see Rotating an API key), and give the runner the new key. Rotating first would leave the old runner sending the new key in its URL.
Until the runner is updated, the server logs one warning per restart naming
the account that runner uses: a runner sent its API key in the /ws URL. The
warning names the account, never the key. The server's own request log no
longer records query values on any route.
Moving the image pin onto a directory mount¶
If your installation was created before 0.1.86, its update service may be
able to lose track of the image pin without any error, and the next
sudo docker compose up or reboot can then start an older server than the one you
are running. If the database has moved on since, that older server refuses to
start. This section checks for that and fixes it for good. Nothing about it
changes when you upgrade: the fix is a one-time change to your installation.
Why. After every managed update, the update service rewrites the image pin
(the @sha256: digest in your override file) so that a recreate or a reboot
keeps the release it applied. Installs created before 0.1.86 mount that file
into the update service as a single file. A single-file mount follows the
file's identity on disk, not its name, so anything that saves the file by
writing a new copy and swapping it into place (sed -i, most editors, cp or
mv over it) cuts the update service off from it. From then on its writes go
to a copy nothing reads, and the file on the host stops changing. Installs from
0.1.86 on mount the pin's directory instead, which does not have this problem.
The Updates page tells you where you stand. On a current update service it reports the pin checks directly. On an older one it says the pin cannot be checked, which means: follow this section.
Until you have done this¶
- Never edit the pin with
sed -ior an editor that saves by replacing the file. If you must change it, write it in place, for exampleprintf '%s\n' "$new_contents" > docker-compose.override.yml. - After any edit to it, restart the update service with
docker restart <prefix>-updater(hivemind-updaterunless you changed the prefix).sudo docker compose up -ddoes not do this: for an unchanged service it does nothing.
Step 1 — check the shape¶
From your installation directory:
docker inspect hivemind-updater \
--format '{{range .Config.Env}}{{println .}}{{end}}' | grep HIVEMIND_UPDATER_PIN_FILE
docker inspect hivemind-updater \
--format '{{range .Mounts}}{{.Source}} -> {{.Destination}}{{println}}{{end}}'
HIVEMIND_UPDATER_PIN_FILE=/host/pin/...and a mount ending-> /host/pin: you already have the directory mount. Stop here.HIVEMIND_UPDATER_PIN_FILE=/host/image-pin.ymland a mount of a file ending-> /host/image-pin.yml: single-file mount. Continue. Note which file on the host the mount's source is; the steps below call it the pin file and assume it isdocker-compose.override.yml. Substitute its real name throughout.
Step 2 — make sure the pin names the server you are running¶
docker inspect hivemind-server --format '{{.Config.Image}}'
grep '@sha256:' docker-compose.override.yml
The digest after hivemind-server@ must be the same in both. If it is not,
stop and fix that first: a pin that names a different server is the outage
this section prevents. Rewrite the digest in place (never sed -i), then
restart the update service as above.
Step 3 — move the pin into a directory, without cutting the update service off¶
mkdir -p pin
ln docker-compose.override.yml pin/docker-compose.override.yml
ln -s pin/docker-compose.override.yml docker-compose.override.yml.new
mv -T docker-compose.override.yml.new docker-compose.override.yml
ln without -s makes the same file reachable from inside pin/, so nothing
is copied and the running update service keeps seeing it. The last two lines
swap the old name for a link in one step, so docker compose never finds the
file missing.
Step 4 — point the update service at the directory¶
Add this under services: in pin/docker-compose.override.yml, keeping the
file's two-space indentation:
hivemind-updater:
volumes:
- ./pin:/host/pin:rw
environment:
HIVEMIND_UPDATER_PIN_FILE: /host/pin/docker-compose.override.yml
To have the update service report on this itself from now on, also pin it to
the current release: open your channel manifest, read
registry.updater_ref_with_digest (follow the pointer to the release manifest
if needed, as in the runner section above), pull it with docker pull, and
add image: <that reference> to the same block.
Then check that Compose reads it:
sudo docker compose --profile updater config | grep -A30 '^ hivemind-updater:' | grep -E '/host/pin|PIN_FILE'
Both the /host/pin mount and HIVEMIND_UPDATER_PIN_FILE: /host/pin/... must
appear. If not, undo the edit and check the indentation.
Step 5 — recreate the update service, and confirm¶
sudo docker compose --profile updater up -d --no-deps --force-recreate hivemind-updater
docker exec hivemind-updater stat -c %i /host/pin/docker-compose.override.yml
stat -c %i pin/docker-compose.override.yml
The two numbers must match: the update service is now writing the file the host reads. From now on, editing the pin any way you like no longer detaches it.
With the Hivemind source, scripts/migrate-pin-mount.sh performs Steps 1–5
and every check above, and refuses anything it does not recognise. It changes
nothing unless given --apply; add --updater-image <ref>@sha256:<digest> to
move the update service to a current release in the same run.
Pausing or changing your channel¶
Channels are pinned in ./scripts/bundle.json. To pause auto-applies,
ask Seglamater to issue a bundle pinned to a closed channel. To switch
between canary, beta, and stable, the same — a new bundle from
your support contact.
Hivemind does not currently expose a one-click channel switcher in the admin UI; that is on the roadmap. Until then, the channel is part of your bundle and managed out-of-band.