How it is built + verified¶
The security of the Hivemind Hub is not only in its runtime mechanisms. It is also in the discipline by which those mechanisms are built, reviewed, and turned on. A gate model is only as trustworthy as the process that produced it. This page documents that process.
Survey before design; design before build¶
Each load-bearing piece of the Hub follows the same order:
- Survey the code — before proposing anything, the existing codebase is surveyed to find what already exists and can be reused. The recurring finding across the gate work was "the engine already exists": the approval/escalation spine, the per-action policy engine, the model tool-loop, the credential-sealing and injection seams, and the audit chain were all present and already reviewed before the north-star wiring began.
- Design first, in writing — a design document lays out context, the reuse-vs-net-new split, the data model, the composition, the safety invariants, the failure modes, and the open questions. Nothing is built past the design without an explicit go-ahead.
- Build the smallest vertical that proves the thesis — each piece defines a single end-to-end slice that exercises every load-bearing seam at minimum surface, rather than building the whole surface at once.
Reusing a proven engine rather than building a parallel one is itself a security decision: there is one approval mechanism to audit, one gate to attest, one audit chain — not two that might disagree. The design work is largely about resisting the duplicate-engine trap.
Reduce, don't add, engines¶
A guiding constraint runs through the whole architecture: unify, don't multiply. The unified gate does not add another gate system — it defines one composition over the systems that exist. The escalation path does not build a second approval queue — it routes both executors' approvals through one. The credential subsystem does not fork the secret store — it adds a lifecycle over the existing sealing and injection seams. Fewer engines means fewer boundaries to get wrong, and a smaller surface for a reviewer to audit.
Two security reviews, with different jobs¶
Security review happens in two distinct passes. They are not interchangeable, and only one of them is a gate:
- A fast pre-merge review scans for secrets about to ship and for vulnerabilities, and writes a per-commit clearance that the deploy gate checks. This one is the gate: no deploy command runs until the current commit is cleared, and it applies to every change that reaches a deploy.
- A deep security audit examines one security-critical surface at maximum scrutiny — least-privilege scoping, the guarantee that a credential is never leaked to the model, confused-deputy safety, revoke reliability, the caller-auth composition of an exposed surface, and injection into the prompt path. It is pointed at a specific target rather than run across every change, and it produces findings and tickets rather than a deploy clearance.
The credential-minting work and any live network exposure are explicitly gated on the deep review, because minting credentials and exposing an AI-driving control surface are the highest-consequence surfaces in the system.
Dark launch: flags default OFF¶
These security-tightening capabilities land behind a flag that defaults OFF, and are proven before they are switched on in any live deployment:
- the per-target standing-grant requirement,
- the credential-scoping subsystem,
- the unified gate,
- the productization/config-store records.
Not-yet-live does not imply flag-gated, and that is the more useful thing to know. Other capabilities are held back by other means: a variable that is deliberately not delivered to the container, or simply the absence of a call site that would invoke the code. So if a page says a capability is not live and you cannot find a flag for it, the flag is most likely not missing from this documentation — there may be no flag to find, and the page for that capability should say which it is.
A capability can therefore be merged, reviewed, and tested against a live-shaped environment while remaining inert in production. Enabling it is a deliberate, separate, operator act — and for the highest-consequence flags, such as the unified gate, one taken only after the deep security review of that surface.
Corrected 2026-09-17 (HIVE-1279): this section previously stated that every security-tightening capability lands behind a default-OFF flag, and named a live control surface among those flags. The list above measures true; the universal did not, and the control surface is not flag-gated. Two capabilities described on this site are held dark by something other than a flag — the LLM-workflow runtime (corrected under HIVE-1273) and the remote-Commander control surface. The verification note further down is the one claim here that fails on its own terms rather than through drift: a flag that never existed cannot have been checked against the source when the sentence was written.
Property-first and RED-first testing¶
The tests are written to be an executable specification of the invariants, and are written before the implementation (RED-first — the test fails first, then the code makes it pass). The acceptance verticals assert the invariants directly, for example:
- a minted credential resolves while live and stops resolving after revoke (fail-closed);
- the model never sees the secret — asserted structurally, that no tool-call shape carries a credential field;
- an expired mint is refused;
- a job under one owner cannot resolve or see another owner's credential (cross-owner and cross-tenant denied);
- the served control surface exposes exactly the intended verb set and nothing beyond (verified at the served layer, not by reading a config allowlist);
- a caller without a grant is denied, and a request missing its transport proof or its verified principal is rejected.
The composition guarantee in particular is meant to be provable: the min-structure
is what makes "an operator rule cannot loosen a floor" true, and the gate's property
tests target that structure directly — asserting the composed disposition equals min
across all four layers for every combination of inputs.
Evidence, not narration¶
A claim of "done," "passed," or "verified" is backed by the artifact that proves it — a commit hash, a test exit code, a served-surface probe — not by a narrative.
This documentation is held to the same standard, and the way it is held to it is by audit rather than by assumption. The site is re-checked against the source in passes and corrected where it has drifted; the most recent pass, in September 2026, found and fixed published claims about the service count, an example payload many releases stale, and a verification step the updater does not perform. What that gives you is a site verified in passes — not a guarantee that every sentence was re-confirmed the moment you read it. Where a page states that a capability is dark behind a default-OFF flag, designed rather than built, or absent from the install, that specific statement was checked against the source when it was written.