Troubleshooting¶
Common failure modes for a self-hosted Hivemind install, drawn from real deploy + dogfood experience. Each entry names the symptom you'll see, the root cause, and the concrete fix. American English, copy-pasteable.
If you hit something not listed here, the first thing to check is the server boot log — most install-time problems announce themselves there:
Look for INFO lines that confirm each subsystem (database migrations
applied, BYOK LLM provider initialised, integration credential key
loaded, listening) and any WARN/ERROR lines explaining what didn't.
Install + boot¶
Integration routes return 503 integration_key_missing¶
You see this from POST /api/v1/agents/{id}/chatalot/connect (or any
other /chatalot/* lifecycle route) with body
{"error":"integration_key_missing", ...}.
Cause: HIVEMIND_INTEGRATION_ENCRYPTION_KEY (or _FILE) is not wired
into hivemind-server. The server logs a WARN at boot with the message
secret not configured — neither the file form nor the environment
form is set, plus structured fields naming which secret
(secret="integration encryption key") and which two env vars it
checked (file_var="HIVEMIND_INTEGRATION_ENCRYPTION_KEY_FILE",
env_var="HIVEMIND_INTEGRATION_ENCRYPTION_KEY") — grep for "secret not
configured", not the sentence this section used to quote.
Only the file form is declared in the shipped docker-compose.yml. The
server checks both names, but on a compose install
HIVEMIND_INTEGRATION_ENCRYPTION_KEY_FILE is the one passed through to the
container and the bare HIVEMIND_INTEGRATION_ENCRYPTION_KEY is not, so setting
the environment form cannot resolve this. Use the file form below.
Fix (v0.1.9 and later): scripts/install.sh provisions
secrets/integration_encryption_key automatically. Confirm the file is
there:
If it's missing, re-run scripts/install.sh or generate it by hand:
openssl rand -hex 32 > secrets/integration_encryption_key
chmod 0600 secrets/integration_encryption_key
sudo docker compose up -d --force-recreate server
After the recreate, the boot log should read
integration credential key loaded — chatalot/etc. integrations enabled.
Fix (pre-v0.1.9): upgrade. v0.1.9 added the compose secret + the
installer provisioning together; older releases require provisioning the
file manually and editing docker-compose.yml to mount it as a docker
secret.
A route returns 503 feature_disabled¶
The surface exists in this build and is switched off. The body names the feature and the flag that controls it:
{"error":{"status":503,"code":"feature_disabled",
"message":"feature 'feedback' is disabled on this instance: set HIVEMIND_FEEDBACK_ENABLED to enable it (accepted: 1, true, yes, on)",
"feature":"feedback","env_var":"HIVEMIND_FEEDBACK_ENABLED"}}
Set the named flag to 1, true, yes or on and restart. See
Feature flags.
A route returns 401 and you are sure the feature is enabled¶
401 means "authenticate first" — it does not tell you whether the feature is
on, deliberately, so that an anonymous caller cannot enumerate which features
your instance runs. Retry with a valid X-API-Key; if the feature is actually
off you will then get the 503 feature_disabled above.
Check the boot log too — a flag set to an unrecognized value logs a WARN at
startup naming the value it saw:
WARN feature disabled — HIVEMIND_FEEDBACK_ENABLED is set to 'enabled', which is
NOT a recognized truthy value (accepted: 1, true, yes, on).
Boot log warns LLM startup ping FAILED — missing HIVEMIND_LLM_API_KEY¶
The server boots healthy and reachable, but LLM-dependent surfaces (public chat endpoint, agent runtime tool-call loop, content studio drafting) are degraded.
Cause: the BYOK LLM provider isn't configured in .env.
Fix: set the provider triplet in .env and restart server:
# In .env (Anthropic example; see byok-llm.md for OpenAI / Ollama / OpenAI-compat)
HIVEMIND_LLM_PROVIDER=anthropic
HIVEMIND_LLM_API_KEY=sk-ant-...
HIVEMIND_LLM_MODEL=claude-haiku-4-5-20251001
sudo docker compose up -d server
The next boot should log
BYOK LLM provider initialised provider="anthropic" endpoint=... model=...
without the LLM startup ping FAILED line.
/api/v1/health is ok but /api/v1/agents/{id}/chatalot/* returns 404¶
The server is up and routes for the rest of the API work, but every
/chatalot/* path 404s.
Cause: on releases before v0.1.9, the chatalot lifecycle routes are only mounted when the integration key is set. Missing key → no mount → 404 (rather than the v0.1.9+ 503).
Fix: wire the integration key as above (integration_key_missing
entry) and upgrade to v0.1.9 or later, where the same condition surfaces
as a clearer 503 instead of a misleading 404.
Chatalot integration¶
POST /chatalot/connect returns chatalot_mint_bot_failed / opaque "HTTP transport error"¶
Connecting to a chatalot instance produces:
Cause: the chatalot instance uses a self-signed or private-CA TLS
certificate that the hivemind-server container's system CA store
doesn't trust. The connector verifies TLS by default. On v0.1.9 and
later, the enriched error reads TLS/certificate error (connection
failed): self-signed certificate — if the chatalot instance uses a
self-signed or private-CA cert, set verify_tls=false on the integration.
Fix: include "verify_tls": false in the connect body — a
deliberate, per-integration, logged opt-out:
curl -fsS -X POST -H "X-API-Key: $API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"instance_url": "https://chat.example.internal",
"provision_token": "cb_<your bot:provision token>",
"bot_username": "ops-bot",
"display_name": "Ops Bot",
"expires_in_hours": 720,
"verify_tls": false
}' \
"$HIVE/api/v1/agents/$AGENT_ID/chatalot/connect"
Hivemind logs a WARN on every client build with verify_tls=false.
Disabling TLS verification drops MITM protection on that integration —
use only against an instance you control. A publicly-trusted cert (Let's
Encrypt, a private CA pinned into the container) is the secure path
when available.
send_message tool returns 403 / require_webhook_manager upstream¶
An agent's send_message call returns an UpstreamError{status:403,...}
or chatalot_tool_failed with chatalot complaining about webhook-manager
permissions.
Cause: chatalot has no plain REST message-send. The Hivemind
send_message tool routes through a per-channel webhook
(POST /api/webhooks/execute/{token}), which means the bot has to
auto-create that webhook on the first send. Webhook creation in chatalot
requires the bot to hold a community owner/admin role on the target
channel's community. A bot that's only a member of the channel gets 403.
Fix: grant the bot community-admin via chatalot's bot administration API (on the chatalot instance):
curl -fsS -X POST -H "Authorization: Bearer $CHATALOT_ADMIN_JWT" \
-H 'Content-Type: application/json' \
-d '{"community_id":"<community-uuid>","role":"admin"}' \
"$CHATALOT/api/admin/bots/$BOT_ID/communities"
The next send_message call mints the webhook (POST
/api/channels/{id}/webhooks) cleanly and proceeds to execute it.
Bot intermittently replies "try again in a moment" (HTTP 503)¶
A support bot occasionally returns "I'm having trouble reaching my AI
service right now — please try again in a moment." (chat surface) or a
503 Service Unavailable on /api/v1/agents/{id}/chat, then works on the
next try. The server log shows anthropic chat: upstream overloaded,
retries exhausted.
Cause: the upstream LLM provider returned transient 429/529/5xx
on every attempt within the retry budget — usually a brief Anthropic
overload spike. Hivemind already retried (default 3 attempts with jittered
backoff up to 8 s, honoring Retry-After); the blip simply outlasted it.
This is expected, bounded behavior, not a fault — the 503 is deliberate so
the customer sees "try again", not a broken/empty bubble.
Fix / tuning: usually none — it self-recovers. If your traffic sees
sustained overload, raise the budget with HIVEMIND_LLM_RETRY_MAX /
HIVEMIND_LLM_RETRY_CEIL_MS (see BYOK LLM),
or move to a higher-quota model/tier. Persistent 529s across a whole day
point at an account-level quota issue — check the provider console.
Routing + networking¶
hivemind-server container times out reaching an on-prem chatalot¶
Symptom: POST /chatalot/connect (or /test) against
https://chat.your-domain times out from inside the hivemind-server
container, even though curl https://chat.your-domain/api/health works
from the host shell on the same machine.
Cause: the hostname resolves (via DNS) to the host's LAN IP, but
container → host LAN IP traffic to :443 is dropped by the host firewall
(ufw-docker and friends route container-network → host-LAN
asymmetrically). The TLS path itself is fine; only the route through the
host IP fails.
Fix: route the hostname directly to the reverse-proxy's
container-network IP instead, by adding an extra_hosts entry in
docker-compose.yml on the server service:
<proxy-container-ip> is the reverse proxy's address on a docker network
the hivemind-server container is also attached to (a shared overlay or
bridge). The full HTTPS path then runs end-to-end inside the container
network, never traverses the host LAN, and TLS / SNI stay intact.
Long-term, replacing the extra_hosts entry with a docker network
alias on the proxy service (so resolution happens by docker DNS) drops
the hardcoded IP.
Commander link: the workstation log says "answered the ask list with 302 — refused, not followed"¶
Symptom: asks confirmed in the dashboard never reach the Commander, and the
waiting ask's card shows the link as OFFLINE or never checked in.
Cause: the instance sits behind SSO forward auth, and the proxy answers the workstation's signed requests with a redirect to the login page. The tick never follows a redirect.
Fix: route /api/v1/commander-link/* to Hivemind outside forward auth,
with the identity headers stripped. Then an unsigned
curl -i https://hivemind.example.com/api/v1/commander-link/asks returns
401 from Hivemind instead of 302. Snippets and checks:
Commander link.
Agent runner¶
The runner is nowhere — docker compose ps does not list it at all¶
The most likely answer is that it was never started, and the second most
likely is that it exists but is hidden. sudo docker compose ps does not show a
container that was created and never started. Ask again with -a:
Three outcomes, and they mean different things:
| what you see | what it means |
|---|---|
nothing, even with -a |
The runner profile was never started. install.sh says why in its closing Agent runner section — re-read the end of the install output. |
Created, never Up |
Docker could not set the container up, so the entrypoint never ran. Its logs are empty and that is expected — it is not evidence that nothing is wrong. |
Exited or Restarting |
The entrypoint did run and refused. sudo docker compose --profile runner logs hivemind-runner names the exact cause in one line. |
Created and never started¶
Almost always the repository path. The runner works inside a git repository on
your host, and one setting in .env names it:
docker-compose.yml mounts that path into the runner at the same path — you do
not write a volumes: block. Left unset it defaults to
/srv/hivemind-workspaces.
Check that the path exists on your host. If it does not, Docker creates it as an
empty root-owned directory rather than failing, and the runner then starts into
an empty workspace and refuses. The runner's own log names
HIVEMIND_RUNNER_REPO_PATH and the path it resolved to.
Docker records its own reason when it has one:
Re-running install.sh also re-checks this before starting anything, and
refuses with the specific reason rather than leaving you a container to find.
The runner is Up, but agents never do anything¶
Up is not "working". Two failures look like this from the outside:
- The Anthropic key is wrong or expired. It passes every startup check —
those are structural and make no network call — so the container starts
cleanly and every spawn fails. Check
sudo docker compose --profile runner logs hivemind-runnerfor spawn failures. -
The runner's own credential is rejected by your Hub. The same logs will show the registration being refused (a
401). That credential isHIVEMIND_RUNNER_API_KEYin.env, not the Anthropic key — they are different keys and are easy to confuse. The server mints its runner row from that same variable at startup, so both must be restarted together:
Do not use /api/v1/health to answer "is my runner up" — it has no runner
field and cannot tell you.
Never point HIVEMIND_RUNNER_API_KEY at your admin key¶
It looks like it would work, and it does — which is the problem. It is not a shortcut; it is a one-way door.
What it costs immediately. An admin-role credential is a Commander credential, so the runner is no longer scoped to one machine row and no longer subject to the secret-reachability interlock. Both of the controls that keep one runner out of another host's records are skipped, because both apply only to a runner-scoped credential.
What it costs afterwards, and this is the part that is not obvious.
Registering under an admin key creates the machines row without the
service-principal claim that normally accompanies it. A machine row that exists
and carries no claim is permanently refused to a runner-scoped credential —
deliberately, because an unclaimed row means "somebody else's", never "free to
take". So switching back to the correct HIVEMIND_RUNNER_API_KEY afterwards does
not recover: the correct credential is refused on the row the wrong one created,
and there is no takeover path by design.
This compounds if the runner container has been recreated: it sets no
hostname:, so its machine identity is the container hostname and each forced
recreate mints a new one. Several recreates under an admin key leaves several
unclaimed rows, and only you can say which is current.
If you have already done this, do not try to fix it by editing the database. Contact support, and include the result of this query run against your Hivemind database:
SELECT m.machine_id, m.hostname, m.last_heartbeat, c.principal_id
FROM machines m
LEFT JOIN machine_service_claim c USING (machine_id)
ORDER BY m.last_heartbeat DESC;
Rows with a NULL principal_id are the affected ones. Clearing them is an
operator action with a decision attached — which row is current, and whether the
stale ones should be deleted or assigned — so it is deliberately not a
self-service step.
The secret-reachability probe returns 200 instead of 403¶
The check in install.md
(Verifying the runner cannot read your secrets)
returned 200 with a tar body. This is a live credential disclosure, not a
misconfiguration to schedule. Agent code can read the contents of any file in
any container on this host — every /run/secrets/* of every service, including
admin_api_key, which is full Hivemind admin on your own Hub.
The cause is always the same one value:
Set it to "0" and restart just the proxy:
Then re-run the probe and confirm 403, plus the _ping control beside it.
A restart that silently failed leaves the old container running with the old
ACL, and up -d prints success either way.
What you lose by setting it to 0: agents can no longer run docker ps,
docker logs or docker inspect, and the infra_docker tool's ps/logs/
inspect actions return an error. Nothing else. Agents run as processes
inside the runner container — no container is created to dispatch one — so task
execution, worktrees, and results are unaffected. There is no narrower setting
that keeps the three read verbs and drops file reads; the proxy filters by path
prefix, and they are all under /containers.
If you deliberately need container observability, do not re-enable this flag: it grants far more than the three verbs you want. Raise it with your Seglamater contact so it is solved with a mediating component instead.
If the credential was reachable, treat it as exposed. Anyone who could run
an agent task on this Hub during that window could have read it. Rotate
ADMIN_API_KEY (and any other value under secrets/) rather than assuming
nobody did.
Runner registration refused with "this server still carries ADMIN_API_KEY as a value in its container environment"¶
This is the same interlock as the probe above, from its other side: rather
than waiting for a probe to prove the read is possible, the server refuses to
create a machines row for a runner at all while it can see ADMIN_API_KEY
as a literal value in its own environment. Registering a runner in that state
would arm the exact privilege escalation the probe above measures, so the
server declines up front instead of shipping a runner that could later prove
it.
The shipped docker-compose.yml never hits this — ADMIN_API_KEY_FILE is
what it sets, not ADMIN_API_KEY. You will see this refusal only if something
overrode that: a manual docker run/docker compose run that injects
ADMIN_API_KEY directly, a hand-edited compose override, or a local
rehearsal of the server outside the documented compose file.
Fix: pass the credential as a file, the same way docker-compose.yml already
does, and restart:
services:
hivemind-server:
environment:
- ADMIN_API_KEY_FILE=/run/secrets/admin_api_key
# not: ADMIN_API_KEY=...
Confirm the interlock actually cleared before trying to register a runner again — the refusal names two separate conditions, and this fixes only the one it can see:
docker inspect -f '{{range .Config.Env}}{{println .}}{{end}}' hivemind-server \
| grep -q '^ADMIN_API_KEY=' && echo "still exposed" || echo "clear"
Updater¶
hivemind-updater exits at startup with a configuration error¶
hivemind-updater validates three preconditions at process start, before it
serves anything — including GET /health, which is unauthenticated and needs
none of them on its own read path. All three are provisioned automatically by
install.sh on every real install; you will only see these if you are running
the image by hand outside the documented compose file (a local rehearsal, a
manual docker run against a published image, or a hand-edited override).
| error | what it means | what to do |
|---|---|---|
api token file missing or unreadable at /run/secrets/updater_token: No such file or directory |
HIVEMIND_UPDATER_API_TOKEN_FILE points at a path with nothing there. install.sh normally creates this with openssl rand -hex 32. |
Provision the file the same way: openssl rand -hex 32 > secrets/updater_token, mounted at the path the container expects. |
configuration error: DATABASE_URL must be set |
The variable is empty or absent. This check runs regardless of which route you actually want to exercise — a bogus-but-present URL satisfies it for a /health-only rehearsal, since nothing on that path reads it. |
Set any syntactically valid postgres:// URL if you only need /health to answer; set the real one if you need the updater to actually do anything. |
configuration error: cosign pubkey at /run/secrets/cosign_pub is not a valid ECDSA P-256 SPKI key: ASN.1 error: PEM error: PEM preamble contains invalid data |
The file exists but does not parse as a real key — a placeholder string fails exactly like a missing file. | Generate a real (throwaway, for a rehearsal) keypair rather than a placeholder: openssl ecparam -name prime256v1 -genkey -noout -out key.pem && openssl ec -in key.pem -pubout -out pub.pem, and mount pub.pem at the path the container expects. |
If you are probing a container you expect to fail, do not run it with
--rm — a container that exits immediately is removed before docker logs
can read it, and the refusal is lost with it.
API keys¶
Rotating an API key¶
If a key may have leaked (into a log, a ticket, a chat), replace it. Run this on the host, in a terminal, inside the server container:
It prints the new key and records the rotation in api_key_rotations. The old
key is refused from that moment: no restart needed. The command refuses to run
without a terminal, and it refuses agent credentials, which have their own
rotation path.
Where the old key may still be held. Rotation changes the server's record, not the copies elsewhere, so update every copy:
admin../secrets/admin_api_keystill holds the OLD key. The server reads that file only to create the admin on a fresh install, so it ignores the stale value. The opt-in legacywebprofile does read it, so replace the file if you run that profile. Also update any client or script that sends the admin key.- The runner (
--username runneron a standard install). The command reports ROTATION INCOMPLETE and exits3. It is not done yet: the running runner still holds the old key and is now being refused. To finish: - Set
HIVEMIND_RUNNER_API_KEYin.envto the new key. - Recreate the runner so it reads the new value:
sudo docker compose up -d hivemind-runner -
About a minute later, check instead of assuming it worked:
VERIFIED(exit0) means a runner heartbeat landed after the rotation. Only the new key can produce one.NOT VERIFIED(exit3) means no heartbeat has landed yet, so the runner is still on the old key or is not running.CANNOT VERIFY(exit4) means the runner has never registered a machine, so no heartbeat can show which key it holds. In that case, check the runner's own log for refused requests.
rotation-status also reports refused rotation attempts from the last 24
hours. Rotation over HTTP refuses API-key callers, so a run of refused attempts
means someone holding a key tried to use it to rotate one.
Health + verification¶
docker compose ps shows hivemind-server as (unhealthy) after install¶
The container is running but its healthcheck (curl /api/v1/health)
fails repeatedly. /api/v1/health is what scripts/deploy.sh and the
container HEALTHCHECK directive both poll.
Cause: usually one of:
- The server failed to apply its migrations (Postgres not ready yet, or
wrong
DATABASE_URL). The logs show asqlx::Errornear the top. .envis missing a required key (DB_PASSWORD,ADMIN_API_KEY,COOKIE_SECRET,JWT_SIGNING_KEY). The server panics at startup — the logs name the missing key.- Host port
${API_PORT:-8585}is already in use by another process. The container starts but can't bind;sudo docker compose logsshows the bind error.
Fix: read the actual error in the logs and act on it:
For missing .env keys, re-run scripts/install.sh (idempotent — it
will fill in any blanks without regenerating existing secrets). For port
conflicts, set API_PORT=<free-port> in .env and sudo docker compose up
-d server.
If the healthcheck still fails after these, curl -fsS
http://localhost:8585/api/v1/health from the host should return
{"status":"ok","version":"…","integrations":{"chatalot_bot_ws":{"status":"…"}}}.
If that works but sudo docker compose ps
still says unhealthy, the issue is the in-container curl (e.g., the
image is missing curl); confirm you're on the published
registry.seglamater.app/seglamater/hivemind-server:<ver> image.
unauthorized / 401 pulling the image¶
Symptom: the install stops with No registry credentials — this
install cannot pull its image, or an older installer runs further and
fails at docker compose pull with Error response from daemon:
unauthorized.
Cause: the image registry requires authentication. It answers anonymous callers with 401 and its token service will not issue an anonymous pull token, so there is no unauthenticated way to fetch the image. The credentials come from your per-customer bundle, and that bundle comes from your invite URL.
The single most common cause is installing from a generic release
bundle instead of an invite. The bundle published on
updates.seglamater.app carries no credentials by design and cannot
install on its own.
Fix — in order of likelihood:
- Re-run from your invite URL.
curl -fsSL https://s.seglamater.app/i/<invite-id> -o hivemind-install.sh && bash hivemind-install.sh. If the invite is expired or used up the command says so in as many words and exits 1 without touching the host — ask your support contact for a fresh one. (Piping the URL intobashinstead reduces that message tocommand not found, which is why the documented form fetches to a file.) - Check your bundle actually carries credentials. Both a username and a token must be present and non-empty:
python3 -c 'import json,sys; r=json.load(open(sys.argv[1])).get("registry",{}); \
print("username:", bool(r.get("username") or r.get("user")), \
"token:", bool(r.get("token") or r.get("password")))' bundle.json
Two True values means the bundle is credential-bearing. Do not
hand-edit bundle.json to add them — its SHA-256 is verified against
the sibling .sha256 and editing it will fail the integrity check.
3. Or authenticate the host yourself, if your contact gave you
credentials out of band:
install.sh detects an existing login and proceeds.
Verify the registry is reachable at all (this 401 is expected and healthy — it proves you reached the registry rather than a proxy or a DNS wildcard):
curl -sS -i https://registry.seglamater.app/v2/ | head -5
# HTTP/2 401
# docker-distribution-api-version: registry/2.0
# www-authenticate: Bearer realm="https://forgejo.seglamater.app/v2/token",…
If you get anything other than a 401 with those headers — a timeout, a 502, or an HTML page — the problem is network reachability or egress filtering, not credentials.