Skip to content

Security mechanisms

The Hivemind Hub gives an AI the ability to act on real machines. That is only defensible if the boundary between the model and the action is rigorous, inspectable, and layered so that no single failure opens the door. This page documents each mechanism that makes up that boundary.

The mechanisms are defense in depth: they compose, and several are structural — enforced by the shape of the code or the database, not by a policy that could be forgotten. None is the sole boundary.

Deny-by-default

Every gate in the system refuses in the absence of an explicit allow. A principal with no matching grant is denied. A verb that does not resolve to a known tier is denied. An unknown target or action is denied. A newly-registered box is seeded from a deny-by-default template — mutating verbs require approval, unlisted actions are denied — never a blank check. A workflow may call only the actions on its explicit allow-set, and each is independently gated at call time.

This is the property that makes the platform safe to extend: a new box, verb, principal, or action starts with zero capability and must be explicitly granted one. Adding a surface cannot silently open a hole, because the default state of any new surface is "refused."

The tighten-only composition guarantee

The single most important structural property. Dispositions are ordered by restrictiveness and composed by min:

denied (0)  <  needs_approval (1)  <  allowed (2)

floor = min(L0, L1)          # hard floor: irreversible/prod classifier + capability
final = min(floor, L2, L3)   # harness policy + operator rules can only pull DOWN

Because the composition is min, any layer can only lower the result. An operator rule of allowed is the maximum value — it cannot raise final above what the floor already permits, so it is dominated unless every layer already allows. A per-target operator rule can therefore tighten an action (dev-box drive → ask-me) but is structurally incapable of loosening a floor (a prod deploy stays at needs-approval no matter what rule is written).

This is a property proven of the composition, not a convention. "The operator cannot self-loosen prod" is true because of arithmetic, not because a reviewer remembered to check. The composition is asserted directly by the gate's own property tests, which check that the composed disposition equals min across all four layers for every combination of inputs (orchestration_authz/gate_eval.rs). See The unified per-target gate.

Never self-approve

The decider of an approval must be a different principal than the asker. The model that proposes a gated action can never be the one that approves it — only a human operator identity can decide. This is enforced at two layers:

  • In the application — on the decision path, the approval's guarded update carries a "decider is distinct from the asker" clause (principal_id IS DISTINCT FROM decider), plus a pre-check that refuses a self-decision before it is attempted. An approval cannot be self-issued.
  • In the database — the standing-grant table carries a CHECK constraint that a grant's issuer is never its own beneficiary (issued_by <> principal_id), so that guarantee holds at the storage layer even against a code path that forgot the check.1

Only a human operator can mint an elevation, and an elevation is minted only from an operator-decided approval row. The AI can propose an irreversible action; it can never authorize one.

Fail-closed

Every uncertain outcome resolves to "did not happen," never to "happened without approval":

  • Approval timeout — a pending escalation carries a 1-hour TTL. If the operator never answers, an always-on watchdog expires it and delivers a rejection. An unanswered request is a denied request.
  • Evaluator or DB fault — a gate or target-resolution fault denies rather than guesses.
  • Credential resolution — a revoked or expired credential simply stops resolving; a mint failure means the job does not run, with no fallback to a broader credential.

Doubt always resolves toward refusal.

Params-hash anti-forge binding

When an action is escalated for approval, the pending approval is bound to a params_hash — a hash over the verb, the target, and the params. This nonce nails the approval to this exact action. On approval:

  • the one-shot elevation is minted bound to that params_hash, and
  • match-before-burn re-checks the actual action's params_hash against the grant before consuming it.

A mismatched action — an attempt to reuse an approval for a different action than the one you saw and approved — is refused without spending the grant. You approve a specific act, and only that specific act can execute. Combined with single-consume (the approval can be acted on exactly once, via a guarded conditional update) and principal-scoped grants (knowing an id is not the same as being able to burn another principal's grant), the approval token is unforgeable and unreplayable, and it never leaves the server on the executed path.

Confused-deputy defense

The Hub holds credentials; the model never sees them. When a workflow action needs a secret, the model emits a structured tool call with no credential field, and the server loads, decrypts, and injects the secret into the outbound client. The model names the action; the server binds the credential.

Because no tool-call schema carries a credential field, the model cannot name, read, or exfiltrate a secret — there is no slot in which a secret could appear in anything the model touches. The credential is bound to the verified principal and job at inject time, so an agent cannot re-target another owner's or tenant's credential. This is the same discipline the claim seam enforces when a session claims a job: identity is resolved and bound server-side, not asserted by the actor.

This mechanism is why the Hub can safely be a conduit for actions that require real credentials: the deputy (the model) is never handed the authority it would need to be confused into misusing.

Least-privilege minted credentials

Status

Per-job credential minting is merged on main but not wired. Measured on main: nothing in the server calls the minting entry point, so no credential is minted at any setting of HIVEMIND_CRED_SCOPING_ENABLED, and a job today authenticates with a standing credential. What is missing is a call site at job start — unbuilt work rather than a configuration step. The seams this rests on are live: sealing at rest, server-side injection, and the fail-closed key resolver. See Credential scoping for the same status in detail.

Corrected 2026-09-17 (HIVE-1279, finding E06): this section previously said that where the Hub can mint a scoped credential, "it does" — present tense. The HIVE-1276 correction reached credential-scoping.md and missed this page.

Where the Hub can mint a scoped credential rather than reuse a standing one, it is designed to do so. A job then authenticates as a short-lived, scoped, expiring principal — not the operator's standing key. A minted credential is:

  • scoped to the job's action allow-set and tier,
  • short-lived — a TTL no longer than the run's budget window, clamped to a hard ceiling regardless of what is requested,
  • job-bound — recorded against the run/deployment job id,
  • sealed at rest — stored under authenticated encryption (a modern AEAD cipher with a per-record nonce) for server-side injection only, and
  • revoked at the terminal run marks, with a watchdog sweep as a backstop.

The key resolver refuses any key that is not active and unexpired, so a revoked or timed-out credential stops working immediately — fail-closed at the resolution layer. A credential the Hub cannot mint-scope (an external third-party key) stays sealed and is bounded instead by the workflow's action allow-set plus rate and budget caps.

The result: a leaked or misused credential is scoped to one job's actions and dies with the job. There is no standing all-powerful credential handed to a run. See Per-job credential scoping.

The tamper-evident audit chain

Every gate decision — allow, deny, escalate, approve, expire — lands on a hash-chained audit ledger. Each entry incorporates the hash of the previous entry, so the log is tamper-evident: an entry cannot be altered or removed without breaking the chain from that point forward, which a verification pass detects. The chain is seeded per tenant and can be verified end-to-end — verify_chain walks the chain and the server exposes it as an admin route — so a shortened or forked chain is visible against the previously recorded chain head.

The ledger is protected at the database layer as well: triggers reject updates, deletes, and truncation of the audit table through the application's own normal code paths, the rule store is upsert-only, and standing grants are append-only (soft-revoke only, with truncation blocked). These triggers stop an application bug or an ordinary mistaken write; they are not a barrier against a party who already holds the same database privileges the application itself connects with. On an install where the application's database role also owns these tables — which is how a fresh install is provisioned today — a party with that role's access can disable a trigger before writing, the same way any database owner can. Separating that role from table ownership so it cannot is tracked, in progress, and not yet true on every install; do not read this paragraph as saying it already is.

That is exactly why the hash-chain above is the guarantee this section leads with rather than a footnote: prevention at the trigger layer is a second line of defense against ordinary mistakes, not the load-bearing one. The load-bearing property is detection — any mutation, however it happened, breaks the chain, and verify_chain finds it. Nothing described here claims a mutation is impossible; it claims a mutation cannot happen invisibly.

Because both executors and the escalation engine share one gate path, the audit is complete: there is no side door whose actions do not appear on the chain.

The code-enforced estate floor

Underneath the Hub's own gate model sits an un-forgeable floor enforced on the box where a command actually runs. Where a Commander runs, code-level hooks on that box enforce the hard floor independently of the Hub: a prod/irreversible classifier, forge-resistant operator-OK markers (mintable only through an operator action, not by the AI), a design-before-execute planning gate, and a pre-deploy security-review clearance.

Whose boxes these are, precisely. This floor is Seglamater's own operating discipline, enforced by hooks in the workstation configuration of the machine a Commander runs on. It is documented here because it is the second enforcement plane the Hub's model is designed against — not because a Hivemind install places it. A self-hosted install ships none of these hooks. Its agent runner is an opt-in part of the install; see Install.

This is deliberately a separate enforcement plane from the Hub's gate. Even if the Hub-side evaluator were bypassed, the box re-checks the floor where the command executes. The Hub honors the same floor (via a mirrored classifier) so it never proposes an action the box would then refuse — but the box's enforcement is the backstop, not the Hub's cooperation. The operator-only actions (prod deploys, new credentials, new DNS, destructive prod ops, history rewrites) always require the un-forgeable marker or a human decide, and no rule can lower them.

The two planes are kept coherent by a customer-host list mirrored in code from the estate classifier, using the same containment match, so subdomains, host:port forms and trailing-dot FQDNs classify the same way on both sides. Over-matching is deliberate and fail-safe: more actions floored is tighter, never looser. Collapsing the two copies into a single shared list, rather than a mirrored one, is a known follow-up and is not yet done.


Verification note

The claims on this page are checked against the Hivemind source, and the checking is done in audit passes rather than assumed to hold forever. Where a mechanism is merged on main but ships dark behind a flag that defaults OFF — dev-safe and not yet activated in production — or is still in design, the relevant page says so explicitly rather than implying it is live in production.

The most recent audit pass was September 2026. It removed a mechanism from this page that described an attestation component the product does not contain, and corrected the description of an enforcement plane that belongs to Seglamater's own operating discipline rather than to a customer install. Both are noted here because a page about security mechanisms should say when one of its mechanisms was withdrawn.


  1. Both enforcements were checked against the source (origin/main) while writing this documentation. The decision-path application guard is active; the database CHECK constraint on grant issuance is part of the per-target standing-grant work, which — like several capabilities on this site — is merged on main but dark behind a default-OFF flag (HIVEMIND_ORCH_STANDING_GRANT_REQUIRED), not yet activated in production. See How it is built + verified. ↩