Prometheus metrics¶
Hivemind exposes a Prometheus exposition endpoint at GET /metrics so an
existing Prometheus can scrape per-instance operational signals: LLM spend,
agent/profile health, live bot connections, and chat-request outcomes
(including rate-limit, cost-cap, injection, and upstream-overload counts).
Tailnet-only.
/metricsexposes cost and customer-count data. It is auth-exempt at the application layer (Prometheus cannot do SSO), so the network is the trust boundary: never put/metricson a public/internet route. Scrape it over the tailnet/LAN only. The optional Bearer token below adds defense in depth.
What's exposed¶
| Metric | Type | Labels | Meaning |
|---|---|---|---|
hivemind_build_info |
gauge | version |
Always 1; carries the running version. |
hivemind_uptime_seconds |
gauge | — | Seconds since server start. |
hivemind_agent_chat_requests_total |
counter | result |
Chat requests by outcome: ok, rate_limited, cost_cap, injection, input_too_long, overloaded, not_found, error. |
hivemind_agents |
gauge | status |
Agent runs by status (running/pending/failed/completed/…). |
hivemind_agent_profiles |
gauge | tier, enabled |
Agent-profile counts. |
hivemind_cost_usd_total |
gauge | — | Lifetime LLM spend across all agents, USD. Only with the token (below). |
hivemind_agent_cost_usd_daily |
gauge | profile |
Today's spend per profile, USD. Only with the token. |
hivemind_agent_cost_cap_usd |
gauge | profile |
Configured daily cost cap per profile, USD. Only with the token. |
hivemind_bot_ws_connections |
gauge | — | Bot chatalot WebSocket connections currently established. |
hivemind_bot_ws_auth_rejected |
gauge | — | Bots whose chatalot WS auth was rejected (bad token; will not reconnect until fixed). |
hivemind_conformance_pass |
gauge | — | Latest guardrail conformance self-check: 1 = all passed, 0 = regression or not-yet-run. |
hivemind_conformance_checks_failed |
gauge | — | Guardrail checks failing in the latest conformance self-check. |
hivemind_conformance_last_run_timestamp |
gauge | — | Unix time of the latest conformance self-check (0 = never run). |
Labels are bounded (status/tier/result/profile) — there are no per-visitor or other unbounded-cardinality labels.
A failed internal query degrades that metric to its default rather than
failing the scrape: /metrics always returns 200 so the target never flaps
TargetDown.
Optional Bearer token¶
Set HIVEMIND_METRICS_TOKEN to require Authorization: Bearer <token> on
/metrics. Unset (default) = open, relying on the tailnet boundary. The boot
log warns when it is unset, as a reminder that the endpoint must stay
tailnet-only.
The spend gauges need the token. The installation's spend is visible only
to administrators in the web UI, so /metrics serves the three spend gauges
(hivemind_cost_usd_total, hivemind_agent_cost_usd_daily,
hivemind_agent_cost_cap_usd) only when HIVEMIND_METRICS_TOKEN is set. With
it unset, every other series is still served, so an existing scraper keeps
working; it just receives no spend.
Prometheus scrape config (operator step)¶
Wiring the scrape target is an operator step — it is not part of the app. Scrape the container over the tailnet:
scrape_configs:
- job_name: hivemind
metrics_path: /metrics
scheme: https
static_configs:
- targets: ["hivemind.internal"] # your private/tailnet hostname, not a public route
# only if HIVEMIND_METRICS_TOKEN is set:
authorization:
type: Bearer
credentials: "<HIVEMIND_METRICS_TOKEN value>"
If Hivemind is fronted by Caddy with forward_auth on the web UI, add a
dedicated /metrics route that bypasses forward_auth and is reachable
only on the tailnet (mirror the in-Caddy direct-proxy pattern already used
for the internal /api/loki and /api/alertmanager routes — never expose it
on the public site block):
# tailnet-only listener / matcher — NOT the public site block
handle /metrics {
reverse_proxy 127.0.0.1:8585
}
Prometheus then scrapes the new hivemind target; add Grafana panels / alert
rules (spend rate, overloaded rate, bot-connection drops) as a follow-on.