Not approved as drafted -- four things came out of checking them against each other, and one was a bug that would have broken every upgrade. The bug: 0059 specified Restart=on-failure while 0057 has the host restart onto a new binary by exiting CLEANLY. on-failure does not restart a process that exited zero, so every upgraded node would have been left stopped, having successfully upgraded. Found by reading the two records against each other rather than by either alone. Now Restart=always in all three places that mention it. The host cannot run in a container, and the reason is decisive rather than stylistic: step 0 of the substrate bootstrap installs the container runtime, so a host inside a container would need the thing it exists to install. It would also break 0041 -- copy it onto a machine and run it stops being true when the machine must already have a runtime. Everything above tier 0 is a container; the host is not. That split is the tier boundary, not an inconsistency. systemd is named rather than abstracted. An init is not a dependency in 0041's sense: 0041 is about what must be installed before the host works, and an init is not installed, it is what the machine already is. The unit file is the only systemd-specific artefact and it belongs to the package, so a machine with a different supervisor ships a different package. The mesh is a watchdog, and my first draft was half an answer. Recovery must be local -- nothing dials a node, and a host that cannot start cannot report. But detection is the mesh's, and a local supervisor structurally cannot do it: it sees one process failing and cannot tell a broken machine from a broken release. Only something watching every node can, and that distinction decides whether the response is "fix this machine" or "stop shipping this version". So a host rollout is staged -- a few nodes, wait for heartbeats, continue or stop on silence. Local rollback still needed, because the canary nodes break and because a node offline during the rollout gets the declaration later with no batch around it. The first declaration is the overlay and nothing else. Forced, because a node's address and peers are assigned rather than chosen. But also the way back in: a node reachable over the overlay can be fixed by hand if a later declaration breaks it, and a large first declaration risks a node that is broken and unreachable at once. Also stated plainly, because it reads as a contradiction: nodes reach each other over the overlay and every node consumes from the broker; what 0039 forbids is an inbound CONTROL surface, not reachability. And in 06: no node holds a credential to any control-plane store, for reads or writes. Four ADRs already say this separately and none of them said it in one place. Nodes state over the broker; the owning context writes. With a note that most high-frequency writes are observability's, not the registry's -- routing logs into the registry would be the shared-schema mistake arriving through a door marked performance.
02-DECISIONS
Architecture decision records — the "why" trail behind the rules in
00-META and the specifications in 03-DESIGN.
Numbered 02 because a decision precedes the design it authorises. Research concludes,
the decision is recorded here, and only then is the design written. Following the folder
numbers walks the process in the order it happens.
One file per decision, numbered, never deleted. A superseded record has its status: changed
and gains a pointer to what replaced it — its text is never edited. The reasoning that was
rejected is the expensive half to rediscover.
The records run in the order the decisions were taken, oldest first.
Every decision is a record. There is no ledger and no index file — if a decision is worth
recording it is worth a record, and if it is not worth a record it is not recorded
(ADR 0026). A "decision" small enough to be one line is
almost always a rule, and a rule belongs in
00-META/how-we-build.md, where it is enforced and keeps the
incident that earned it.
Frontmatter
---
status: proposed | accepted | superseded
date: YYYY-MM-DD # when the decision was taken, not when it was written down
deciders: name
reconstructed: true|false # true when the record was written after the fact from evidence
superseded-by: # 02-DECISIONS/NNNN-....md, when status is superseded
extends: # 02-DECISIONS/NNNN-....md, when this record widens an earlier one
---
Body
# N. Title in plain language
## Context what was true, with evidence
## Considered Options numbered, each with why it was rejected
## Decision what was decided
## Consequences what follows, including what got harder
## References commits, pull requests, knowledge-base entries, prior art
State evidence, not assertion. "Zero of 124 modules declare brain as a dependency"
outranks "the dependency rule is not followed".
Reconstructed records
Records 0001–0014 were written on 2026-08-23, after the decisions they describe. Records 0015 onward were taken as records. Those
decisions were taken in implementation rather than in a document; the records state what was
decided and the evidence it was decided from, and each carries reconstructed: true and says
so in its first lines.
A reconstructed record is not a transcript. Where the deliberation is not recoverable, the options section states what the alternatives were and why the chosen one won on the evidence available — not a discussion that did not happen. Where a date is not establishable it says so rather than guessing.
Index
The index is generated, not maintained — run the hq-status skill, which reads the
frontmatter of every record. A hand-written index drifts from the folder it describes, and
this one had already done so after a single addition.