Skip to content

Troubleshooting

Common failure modes for a self-hosted Hivemind install, drawn from real deploy + dogfood experience. Each entry names the symptom you'll see, the root cause, and the concrete fix. American English, copy-pasteable.

If you hit something not listed here, the first thing to check is the server boot log — most install-time problems announce themselves there:

sudo docker compose logs --tail=50 hivemind-server

Look for INFO lines that confirm each subsystem (database migrations applied, BYOK LLM provider initialised, integration credential key loaded, listening) and any WARN/ERROR lines explaining what didn't.


Install + boot

Integration routes return 503 integration_key_missing

You see this from POST /api/v1/agents/{id}/chatalot/connect (or any other /chatalot/* lifecycle route) with body {"error":"integration_key_missing", ...}.

Cause: HIVEMIND_INTEGRATION_ENCRYPTION_KEY (or _FILE) is not wired into hivemind-server. The server logs a WARN at boot with the message secret not configured — neither the file form nor the environment form is set, plus structured fields naming which secret (secret="integration encryption key") and which two env vars it checked (file_var="HIVEMIND_INTEGRATION_ENCRYPTION_KEY_FILE", env_var="HIVEMIND_INTEGRATION_ENCRYPTION_KEY") — grep for "secret not configured", not the sentence this section used to quote.

Only the file form is declared in the shipped docker-compose.yml. The server checks both names, but on a compose install HIVEMIND_INTEGRATION_ENCRYPTION_KEY_FILE is the one passed through to the container and the bare HIVEMIND_INTEGRATION_ENCRYPTION_KEY is not, so setting the environment form cannot resolve this. Use the file form below.

Fix (v0.1.9 and later): scripts/install.sh provisions secrets/integration_encryption_key automatically. Confirm the file is there:

ls -la secrets/integration_encryption_key
# Expect: -rw------- ... 64 bytes (32 bytes hex)

If it's missing, re-run scripts/install.sh or generate it by hand:

openssl rand -hex 32 > secrets/integration_encryption_key
chmod 0600 secrets/integration_encryption_key
sudo docker compose up -d --force-recreate server

After the recreate, the boot log should read integration credential key loaded — chatalot/etc. integrations enabled.

Fix (pre-v0.1.9): upgrade. v0.1.9 added the compose secret + the installer provisioning together; older releases require provisioning the file manually and editing docker-compose.yml to mount it as a docker secret.

A route returns 503 feature_disabled

The surface exists in this build and is switched off. The body names the feature and the flag that controls it:

{"error":{"status":503,"code":"feature_disabled",
          "message":"feature 'feedback' is disabled on this instance: set HIVEMIND_FEEDBACK_ENABLED to enable it (accepted: 1, true, yes, on)",
          "feature":"feedback","env_var":"HIVEMIND_FEEDBACK_ENABLED"}}

Set the named flag to 1, true, yes or on and restart. See Feature flags.

A route returns 401 and you are sure the feature is enabled

401 means "authenticate first" — it does not tell you whether the feature is on, deliberately, so that an anonymous caller cannot enumerate which features your instance runs. Retry with a valid X-API-Key; if the feature is actually off you will then get the 503 feature_disabled above.

Check the boot log too — a flag set to an unrecognized value logs a WARN at startup naming the value it saw:

WARN  feature disabled — HIVEMIND_FEEDBACK_ENABLED is set to 'enabled', which is
      NOT a recognized truthy value (accepted: 1, true, yes, on).

Boot log warns LLM startup ping FAILED — missing HIVEMIND_LLM_API_KEY

The server boots healthy and reachable, but LLM-dependent surfaces (public chat endpoint, agent runtime tool-call loop, content studio drafting) are degraded.

Cause: the BYOK LLM provider isn't configured in .env.

Fix: set the provider triplet in .env and restart server:

# In .env (Anthropic example; see byok-llm.md for OpenAI / Ollama / OpenAI-compat)
HIVEMIND_LLM_PROVIDER=anthropic
HIVEMIND_LLM_API_KEY=sk-ant-...
HIVEMIND_LLM_MODEL=claude-haiku-4-5-20251001

sudo docker compose up -d server

The next boot should log BYOK LLM provider initialised provider="anthropic" endpoint=... model=... without the LLM startup ping FAILED line.

/api/v1/health is ok but /api/v1/agents/{id}/chatalot/* returns 404

The server is up and routes for the rest of the API work, but every /chatalot/* path 404s.

Cause: on releases before v0.1.9, the chatalot lifecycle routes are only mounted when the integration key is set. Missing key → no mount → 404 (rather than the v0.1.9+ 503).

Fix: wire the integration key as above (integration_key_missing entry) and upgrade to v0.1.9 or later, where the same condition surfaces as a clearer 503 instead of a misleading 404.


Chatalot integration

POST /chatalot/connect returns chatalot_mint_bot_failed / opaque "HTTP transport error"

Connecting to a chatalot instance produces:

{"error":"chatalot_mint_bot_failed",
 "detail":"HTTP transport error: ..."}

Cause: the chatalot instance uses a self-signed or private-CA TLS certificate that the hivemind-server container's system CA store doesn't trust. The connector verifies TLS by default. On v0.1.9 and later, the enriched error reads TLS/certificate error (connection failed): self-signed certificate — if the chatalot instance uses a self-signed or private-CA cert, set verify_tls=false on the integration.

Fix: include "verify_tls": false in the connect body — a deliberate, per-integration, logged opt-out:

curl -fsS -X POST -H "X-API-Key: $API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{
    "instance_url": "https://chat.example.internal",
    "provision_token": "cb_<your bot:provision token>",
    "bot_username": "ops-bot",
    "display_name": "Ops Bot",
    "expires_in_hours": 720,
    "verify_tls": false
  }' \
  "$HIVE/api/v1/agents/$AGENT_ID/chatalot/connect"

Hivemind logs a WARN on every client build with verify_tls=false. Disabling TLS verification drops MITM protection on that integration — use only against an instance you control. A publicly-trusted cert (Let's Encrypt, a private CA pinned into the container) is the secure path when available.

send_message tool returns 403 / require_webhook_manager upstream

An agent's send_message call returns an UpstreamError{status:403,...} or chatalot_tool_failed with chatalot complaining about webhook-manager permissions.

Cause: chatalot has no plain REST message-send. The Hivemind send_message tool routes through a per-channel webhook (POST /api/webhooks/execute/{token}), which means the bot has to auto-create that webhook on the first send. Webhook creation in chatalot requires the bot to hold a community owner/admin role on the target channel's community. A bot that's only a member of the channel gets 403.

Fix: grant the bot community-admin via chatalot's bot administration API (on the chatalot instance):

curl -fsS -X POST -H "Authorization: Bearer $CHATALOT_ADMIN_JWT" \
  -H 'Content-Type: application/json' \
  -d '{"community_id":"<community-uuid>","role":"admin"}' \
  "$CHATALOT/api/admin/bots/$BOT_ID/communities"

The next send_message call mints the webhook (POST /api/channels/{id}/webhooks) cleanly and proceeds to execute it.

Bot intermittently replies "try again in a moment" (HTTP 503)

A support bot occasionally returns "I'm having trouble reaching my AI service right now — please try again in a moment." (chat surface) or a 503 Service Unavailable on /api/v1/agents/{id}/chat, then works on the next try. The server log shows anthropic chat: upstream overloaded, retries exhausted.

Cause: the upstream LLM provider returned transient 429/529/5xx on every attempt within the retry budget — usually a brief Anthropic overload spike. Hivemind already retried (default 3 attempts with jittered backoff up to 8 s, honoring Retry-After); the blip simply outlasted it. This is expected, bounded behavior, not a fault — the 503 is deliberate so the customer sees "try again", not a broken/empty bubble.

Fix / tuning: usually none — it self-recovers. If your traffic sees sustained overload, raise the budget with HIVEMIND_LLM_RETRY_MAX / HIVEMIND_LLM_RETRY_CEIL_MS (see BYOK LLM), or move to a higher-quota model/tier. Persistent 529s across a whole day point at an account-level quota issue — check the provider console.


Routing + networking

hivemind-server container times out reaching an on-prem chatalot

Symptom: POST /chatalot/connect (or /test) against https://chat.your-domain times out from inside the hivemind-server container, even though curl https://chat.your-domain/api/health works from the host shell on the same machine.

Cause: the hostname resolves (via DNS) to the host's LAN IP, but container → host LAN IP traffic to :443 is dropped by the host firewall (ufw-docker and friends route container-network → host-LAN asymmetrically). The TLS path itself is fine; only the route through the host IP fails.

Fix: route the hostname directly to the reverse-proxy's container-network IP instead, by adding an extra_hosts entry in docker-compose.yml on the server service:

services:
  server:
    extra_hosts:
      - "chat.your-domain:<proxy-container-ip>"

<proxy-container-ip> is the reverse proxy's address on a docker network the hivemind-server container is also attached to (a shared overlay or bridge). The full HTTPS path then runs end-to-end inside the container network, never traverses the host LAN, and TLS / SNI stay intact.

Long-term, replacing the extra_hosts entry with a docker network alias on the proxy service (so resolution happens by docker DNS) drops the hardcoded IP.

Symptom: asks confirmed in the dashboard never reach the Commander, and the waiting ask's card shows the link as OFFLINE or never checked in.

Cause: the instance sits behind SSO forward auth, and the proxy answers the workstation's signed requests with a redirect to the login page. The tick never follows a redirect.

Fix: route /api/v1/commander-link/* to Hivemind outside forward auth, with the identity headers stripped. Then an unsigned curl -i https://hivemind.example.com/api/v1/commander-link/asks returns 401 from Hivemind instead of 302. Snippets and checks: Commander link.


Agent runner

The runner is nowhere — docker compose ps does not list it at all

The most likely answer is that it was never started, and the second most likely is that it exists but is hidden. sudo docker compose ps does not show a container that was created and never started. Ask again with -a:

sudo docker compose --profile runner ps -a hivemind-runner

Three outcomes, and they mean different things:

what you see what it means
nothing, even with -a The runner profile was never started. install.sh says why in its closing Agent runner section — re-read the end of the install output.
Created, never Up Docker could not set the container up, so the entrypoint never ran. Its logs are empty and that is expected — it is not evidence that nothing is wrong.
Exited or Restarting The entrypoint did run and refused. sudo docker compose --profile runner logs hivemind-runner names the exact cause in one line.

Created and never started

Almost always the repository path. The runner works inside a git repository on your host, and one setting in .env names it:

HIVEMIND_RUNNER_REPO_PATH=/path/to/your/repo

docker-compose.yml mounts that path into the runner at the same path — you do not write a volumes: block. Left unset it defaults to /srv/hivemind-workspaces.

Check that the path exists on your host. If it does not, Docker creates it as an empty root-owned directory rather than failing, and the runner then starts into an empty workspace and refuses. The runner's own log names HIVEMIND_RUNNER_REPO_PATH and the path it resolved to.

Docker records its own reason when it has one:

docker inspect -f '{{.State.Error}}' hivemind-runner

Re-running install.sh also re-checks this before starting anything, and refuses with the specific reason rather than leaving you a container to find.

The runner is Up, but agents never do anything

Up is not "working". Two failures look like this from the outside:

  • The Anthropic key is wrong or expired. It passes every startup check — those are structural and make no network call — so the container starts cleanly and every spawn fails. Check sudo docker compose --profile runner logs hivemind-runner for spawn failures.
  • The runner's own credential is rejected by your Hub. The same logs will show the registration being refused (a 401). That credential is HIVEMIND_RUNNER_API_KEY in .env, not the Anthropic key — they are different keys and are easy to confuse. The server mints its runner row from that same variable at startup, so both must be restarted together:

    sudo docker compose up -d server && docker compose --profile runner up -d
    

Do not use /api/v1/health to answer "is my runner up" — it has no runner field and cannot tell you.

Never point HIVEMIND_RUNNER_API_KEY at your admin key

It looks like it would work, and it does — which is the problem. It is not a shortcut; it is a one-way door.

What it costs immediately. An admin-role credential is a Commander credential, so the runner is no longer scoped to one machine row and no longer subject to the secret-reachability interlock. Both of the controls that keep one runner out of another host's records are skipped, because both apply only to a runner-scoped credential.

What it costs afterwards, and this is the part that is not obvious. Registering under an admin key creates the machines row without the service-principal claim that normally accompanies it. A machine row that exists and carries no claim is permanently refused to a runner-scoped credential — deliberately, because an unclaimed row means "somebody else's", never "free to take". So switching back to the correct HIVEMIND_RUNNER_API_KEY afterwards does not recover: the correct credential is refused on the row the wrong one created, and there is no takeover path by design.

This compounds if the runner container has been recreated: it sets no hostname:, so its machine identity is the container hostname and each forced recreate mints a new one. Several recreates under an admin key leaves several unclaimed rows, and only you can say which is current.

If you have already done this, do not try to fix it by editing the database. Contact support, and include the result of this query run against your Hivemind database:

SELECT m.machine_id, m.hostname, m.last_heartbeat, c.principal_id
  FROM machines m
  LEFT JOIN machine_service_claim c USING (machine_id)
 ORDER BY m.last_heartbeat DESC;

Rows with a NULL principal_id are the affected ones. Clearing them is an operator action with a decision attached — which row is current, and whether the stale ones should be deleted or assigned — so it is deliberately not a self-service step.

The secret-reachability probe returns 200 instead of 403

The check in install.md (Verifying the runner cannot read your secrets) returned 200 with a tar body. This is a live credential disclosure, not a misconfiguration to schedule. Agent code can read the contents of any file in any container on this host — every /run/secrets/* of every service, including admin_api_key, which is full Hivemind admin on your own Hub.

The cause is always the same one value:

grep -A2 'hivemind-runner-socket-proxy:' -n docker-compose.yml | grep CONTAINERS

Set it to "0" and restart just the proxy:

  hivemind-runner-socket-proxy:
    environment:
      CONTAINERS: "0"
sudo docker compose --profile runner up -d hivemind-runner-socket-proxy

Then re-run the probe and confirm 403, plus the _ping control beside it. A restart that silently failed leaves the old container running with the old ACL, and up -d prints success either way.

What you lose by setting it to 0: agents can no longer run docker ps, docker logs or docker inspect, and the infra_docker tool's ps/logs/ inspect actions return an error. Nothing else. Agents run as processes inside the runner container — no container is created to dispatch one — so task execution, worktrees, and results are unaffected. There is no narrower setting that keeps the three read verbs and drops file reads; the proxy filters by path prefix, and they are all under /containers.

If you deliberately need container observability, do not re-enable this flag: it grants far more than the three verbs you want. Raise it with your Seglamater contact so it is solved with a mediating component instead.

If the credential was reachable, treat it as exposed. Anyone who could run an agent task on this Hub during that window could have read it. Rotate ADMIN_API_KEY (and any other value under secrets/) rather than assuming nobody did.

Runner registration refused with "this server still carries ADMIN_API_KEY as a value in its container environment"

This is the same interlock as the probe above, from its other side: rather than waiting for a probe to prove the read is possible, the server refuses to create a machines row for a runner at all while it can see ADMIN_API_KEY as a literal value in its own environment. Registering a runner in that state would arm the exact privilege escalation the probe above measures, so the server declines up front instead of shipping a runner that could later prove it.

The shipped docker-compose.yml never hits this — ADMIN_API_KEY_FILE is what it sets, not ADMIN_API_KEY. You will see this refusal only if something overrode that: a manual docker run/docker compose run that injects ADMIN_API_KEY directly, a hand-edited compose override, or a local rehearsal of the server outside the documented compose file.

Fix: pass the credential as a file, the same way docker-compose.yml already does, and restart:

services:
  hivemind-server:
    environment:
      - ADMIN_API_KEY_FILE=/run/secrets/admin_api_key
      # not: ADMIN_API_KEY=...

Confirm the interlock actually cleared before trying to register a runner again — the refusal names two separate conditions, and this fixes only the one it can see:

docker inspect -f '{{range .Config.Env}}{{println .}}{{end}}' hivemind-server \
  | grep -q '^ADMIN_API_KEY=' && echo "still exposed" || echo "clear"

Updater

hivemind-updater exits at startup with a configuration error

hivemind-updater validates three preconditions at process start, before it serves anything — including GET /health, which is unauthenticated and needs none of them on its own read path. All three are provisioned automatically by install.sh on every real install; you will only see these if you are running the image by hand outside the documented compose file (a local rehearsal, a manual docker run against a published image, or a hand-edited override).

error what it means what to do
api token file missing or unreadable at /run/secrets/updater_token: No such file or directory HIVEMIND_UPDATER_API_TOKEN_FILE points at a path with nothing there. install.sh normally creates this with openssl rand -hex 32. Provision the file the same way: openssl rand -hex 32 > secrets/updater_token, mounted at the path the container expects.
configuration error: DATABASE_URL must be set The variable is empty or absent. This check runs regardless of which route you actually want to exercise — a bogus-but-present URL satisfies it for a /health-only rehearsal, since nothing on that path reads it. Set any syntactically valid postgres:// URL if you only need /health to answer; set the real one if you need the updater to actually do anything.
configuration error: cosign pubkey at /run/secrets/cosign_pub is not a valid ECDSA P-256 SPKI key: ASN.1 error: PEM error: PEM preamble contains invalid data The file exists but does not parse as a real key — a placeholder string fails exactly like a missing file. Generate a real (throwaway, for a rehearsal) keypair rather than a placeholder: openssl ecparam -name prime256v1 -genkey -noout -out key.pem && openssl ec -in key.pem -pubout -out pub.pem, and mount pub.pem at the path the container expects.

If you are probing a container you expect to fail, do not run it with --rm — a container that exits immediately is removed before docker logs can read it, and the refusal is lost with it.


API keys

Rotating an API key

If a key may have leaked (into a log, a ticket, a chat), replace it. Run this on the host, in a terminal, inside the server container:

sudo docker compose exec server hivemind-server recover rotate-api-key --username <name>

It prints the new key and records the rotation in api_key_rotations. The old key is refused from that moment: no restart needed. The command refuses to run without a terminal, and it refuses agent credentials, which have their own rotation path.

Where the old key may still be held. Rotation changes the server's record, not the copies elsewhere, so update every copy:

  • admin. ./secrets/admin_api_key still holds the OLD key. The server reads that file only to create the admin on a fresh install, so it ignores the stale value. The opt-in legacy web profile does read it, so replace the file if you run that profile. Also update any client or script that sends the admin key.
  • The runner (--username runner on a standard install). The command reports ROTATION INCOMPLETE and exits 3. It is not done yet: the running runner still holds the old key and is now being refused. To finish:
  • Set HIVEMIND_RUNNER_API_KEY in .env to the new key.
  • Recreate the runner so it reads the new value: sudo docker compose up -d hivemind-runner
  • About a minute later, check instead of assuming it worked:

    sudo docker compose exec server hivemind-server recover rotation-status --username runner
    

    VERIFIED (exit 0) means a runner heartbeat landed after the rotation. Only the new key can produce one. NOT VERIFIED (exit 3) means no heartbeat has landed yet, so the runner is still on the old key or is not running. CANNOT VERIFY (exit 4) means the runner has never registered a machine, so no heartbeat can show which key it holds. In that case, check the runner's own log for refused requests.

rotation-status also reports refused rotation attempts from the last 24 hours. Rotation over HTTP refuses API-key callers, so a run of refused attempts means someone holding a key tried to use it to rotate one.

Health + verification

docker compose ps shows hivemind-server as (unhealthy) after install

The container is running but its healthcheck (curl /api/v1/health) fails repeatedly. /api/v1/health is what scripts/deploy.sh and the container HEALTHCHECK directive both poll.

Cause: usually one of:

  1. The server failed to apply its migrations (Postgres not ready yet, or wrong DATABASE_URL). The logs show a sqlx::Error near the top.
  2. .env is missing a required key (DB_PASSWORD, ADMIN_API_KEY, COOKIE_SECRET, JWT_SIGNING_KEY). The server panics at startup — the logs name the missing key.
  3. Host port ${API_PORT:-8585} is already in use by another process. The container starts but can't bind; sudo docker compose logs shows the bind error.

Fix: read the actual error in the logs and act on it:

sudo docker compose logs --tail=80 hivemind-server
sudo docker compose logs --tail=40 postgres

For missing .env keys, re-run scripts/install.sh (idempotent — it will fill in any blanks without regenerating existing secrets). For port conflicts, set API_PORT=<free-port> in .env and sudo docker compose up -d server.

If the healthcheck still fails after these, curl -fsS http://localhost:8585/api/v1/health from the host should return {"status":"ok","version":"…","integrations":{"chatalot_bot_ws":{"status":"…"}}}. If that works but sudo docker compose ps still says unhealthy, the issue is the in-container curl (e.g., the image is missing curl); confirm you're on the published registry.seglamater.app/seglamater/hivemind-server:<ver> image.

unauthorized / 401 pulling the image

Symptom: the install stops with No registry credentials — this install cannot pull its image, or an older installer runs further and fails at docker compose pull with Error response from daemon: unauthorized.

Cause: the image registry requires authentication. It answers anonymous callers with 401 and its token service will not issue an anonymous pull token, so there is no unauthenticated way to fetch the image. The credentials come from your per-customer bundle, and that bundle comes from your invite URL.

The single most common cause is installing from a generic release bundle instead of an invite. The bundle published on updates.seglamater.app carries no credentials by design and cannot install on its own.

Fix — in order of likelihood:

  1. Re-run from your invite URL. curl -fsSL https://s.seglamater.app/i/<invite-id> -o hivemind-install.sh && bash hivemind-install.sh. If the invite is expired or used up the command says so in as many words and exits 1 without touching the host — ask your support contact for a fresh one. (Piping the URL into bash instead reduces that message to command not found, which is why the documented form fetches to a file.)
  2. Check your bundle actually carries credentials. Both a username and a token must be present and non-empty:
python3 -c 'import json,sys; r=json.load(open(sys.argv[1])).get("registry",{}); \
  print("username:", bool(r.get("username") or r.get("user")), \
        "token:", bool(r.get("token") or r.get("password")))' bundle.json

Two True values means the bundle is credential-bearing. Do not hand-edit bundle.json to add them — its SHA-256 is verified against the sibling .sha256 and editing it will fail the integrity check. 3. Or authenticate the host yourself, if your contact gave you credentials out of band:

docker login registry.seglamater.app

install.sh detects an existing login and proceeds.

Verify the registry is reachable at all (this 401 is expected and healthy — it proves you reached the registry rather than a proxy or a DNS wildcard):

curl -sS -i https://registry.seglamater.app/v2/ | head -5
# HTTP/2 401
# docker-distribution-api-version: registry/2.0
# www-authenticate: Bearer realm="https://forgejo.seglamater.app/v2/token",…

If you get anything other than a 401 with those headers — a timeout, a 502, or an HTML page — the problem is network reachability or egress filtering, not credentials.