Skip to content

Prometheus metrics

Hivemind exposes a Prometheus exposition endpoint at GET /metrics so an existing Prometheus can scrape per-instance operational signals: LLM spend, agent/profile health, live bot connections, and chat-request outcomes (including rate-limit, cost-cap, injection, and upstream-overload counts).

Tailnet-only. /metrics exposes cost and customer-count data. It is auth-exempt at the application layer (Prometheus cannot do SSO), so the network is the trust boundary: never put /metrics on a public/internet route. Scrape it over the tailnet/LAN only. The optional Bearer token below adds defense in depth.

What's exposed

Metric Type Labels Meaning
hivemind_build_info gauge version Always 1; carries the running version.
hivemind_uptime_seconds gauge — Seconds since server start.
hivemind_agent_chat_requests_total counter result Chat requests by outcome: ok, rate_limited, cost_cap, injection, input_too_long, overloaded, not_found, error.
hivemind_agents gauge status Agent runs by status (running/pending/failed/completed/…).
hivemind_agent_profiles gauge tier, enabled Agent-profile counts.
hivemind_cost_usd_total gauge — Lifetime LLM spend across all agents, USD. Only with the token (below).
hivemind_agent_cost_usd_daily gauge profile Today's spend per profile, USD. Only with the token.
hivemind_agent_cost_cap_usd gauge profile Configured daily cost cap per profile, USD. Only with the token.
hivemind_bot_ws_connections gauge — Bot chatalot WebSocket connections currently established.
hivemind_bot_ws_auth_rejected gauge — Bots whose chatalot WS auth was rejected (bad token; will not reconnect until fixed).
hivemind_conformance_pass gauge — Latest guardrail conformance self-check: 1 = all passed, 0 = regression or not-yet-run.
hivemind_conformance_checks_failed gauge — Guardrail checks failing in the latest conformance self-check.
hivemind_conformance_last_run_timestamp gauge — Unix time of the latest conformance self-check (0 = never run).

Labels are bounded (status/tier/result/profile) — there are no per-visitor or other unbounded-cardinality labels.

A failed internal query degrades that metric to its default rather than failing the scrape: /metrics always returns 200 so the target never flaps TargetDown.

Optional Bearer token

Set HIVEMIND_METRICS_TOKEN to require Authorization: Bearer <token> on /metrics. Unset (default) = open, relying on the tailnet boundary. The boot log warns when it is unset, as a reminder that the endpoint must stay tailnet-only.

The spend gauges need the token. The installation's spend is visible only to administrators in the web UI, so /metrics serves the three spend gauges (hivemind_cost_usd_total, hivemind_agent_cost_usd_daily, hivemind_agent_cost_cap_usd) only when HIVEMIND_METRICS_TOKEN is set. With it unset, every other series is still served, so an existing scraper keeps working; it just receives no spend.

HIVEMIND_METRICS_TOKEN=$(openssl rand -hex 32)   # in the server's environment

Prometheus scrape config (operator step)

Wiring the scrape target is an operator step — it is not part of the app. Scrape the container over the tailnet:

scrape_configs:
  - job_name: hivemind
    metrics_path: /metrics
    scheme: https
    static_configs:
      - targets: ["hivemind.internal"]  # your private/tailnet hostname, not a public route
    # only if HIVEMIND_METRICS_TOKEN is set:
    authorization:
      type: Bearer
      credentials: "<HIVEMIND_METRICS_TOKEN value>"

If Hivemind is fronted by Caddy with forward_auth on the web UI, add a dedicated /metrics route that bypasses forward_auth and is reachable only on the tailnet (mirror the in-Caddy direct-proxy pattern already used for the internal /api/loki and /api/alertmanager routes — never expose it on the public site block):

# tailnet-only listener / matcher — NOT the public site block
handle /metrics {
    reverse_proxy 127.0.0.1:8585
}

Prometheus then scrapes the new hivemind target; add Grafana panels / alert rules (spend rate, overloaded rate, bot-connection drops) as a follow-on.