Files
hq/03-DESIGN/01-to-be
jschoubben 19997d56c3 Approve 0057-0059, with four corrections from review
Not approved as drafted -- four things came out of checking them against each
other, and one was a bug that would have broken every upgrade.

The bug: 0059 specified Restart=on-failure while 0057 has the host restart onto
a new binary by exiting CLEANLY. on-failure does not restart a process that
exited zero, so every upgraded node would have been left stopped, having
successfully upgraded. Found by reading the two records against each other
rather than by either alone. Now Restart=always in all three places that
mention it.

The host cannot run in a container, and the reason is decisive rather than
stylistic: step 0 of the substrate bootstrap installs the container runtime, so
a host inside a container would need the thing it exists to install. It would
also break 0041 -- copy it onto a machine and run it stops being true when the
machine must already have a runtime. Everything above tier 0 is a container;
the host is not. That split is the tier boundary, not an inconsistency.

systemd is named rather than abstracted. An init is not a dependency in 0041's
sense: 0041 is about what must be installed before the host works, and an init
is not installed, it is what the machine already is. The unit file is the only
systemd-specific artefact and it belongs to the package, so a machine with a
different supervisor ships a different package.

The mesh is a watchdog, and my first draft was half an answer. Recovery must be
local -- nothing dials a node, and a host that cannot start cannot report. But
detection is the mesh's, and a local supervisor structurally cannot do it: it
sees one process failing and cannot tell a broken machine from a broken
release. Only something watching every node can, and that distinction decides
whether the response is "fix this machine" or "stop shipping this version". So
a host rollout is staged -- a few nodes, wait for heartbeats, continue or stop
on silence. Local rollback still needed, because the canary nodes break and
because a node offline during the rollout gets the declaration later with no
batch around it.

The first declaration is the overlay and nothing else. Forced, because a node's
address and peers are assigned rather than chosen. But also the way back in: a
node reachable over the overlay can be fixed by hand if a later declaration
breaks it, and a large first declaration risks a node that is broken and
unreachable at once.

Also stated plainly, because it reads as a contradiction: nodes reach each
other over the overlay and every node consumes from the broker; what 0039
forbids is an inbound CONTROL surface, not reachability.

And in 06: no node holds a credential to any control-plane store, for reads or
writes. Four ADRs already say this separately and none of them said it in one
place. Nodes state over the broker; the owning context writes. With a note that
most high-frequency writes are observability's, not the registry's -- routing
logs into the registry would be the shared-schema mistake arriving through a
door marked performance.
2026-08-27 22:12:12 +02:00
..
2026-08-27 21:16:44 +02:00

03-DESIGN / 01-to-be

The mesh being built toward. Every statement here traces to a record in 02-DECISIONS/; nothing arrives by drafting.

A document here describes an intention. What currently runs is in 00-as-is/, and the two are never merged — when something ships, the as-is document is written and this one's status becomes implemented.

Document Covers Rests on
00-work-breakdown.md How the decomposition gets built, in what order, and where a human must look ADR 0015
01-end-to-end-testing.md The lab: a real mesh a change can be run against before it reaches nodes ADR 0016, 0029
02-scenario-declaration.md What a scenario declares — the underlay, and what to place on it ADR 0031
03-scenario-lifecycle.md What happens to a scenario — raise, snapshot, restore, move, destroy ADR 0032
04-lab-installation.md Getting the lab onto a clean machine, and why it verifies capability rather than installation ADR 0008
05-the-node-host.md Tier 0 — the one thing installed by hand, and the only thing that changes a machine ADR 0037
06-the-control-plane.md Tier 2 — what the term means, and the test for what belongs in it ADR 0037
07-the-substrate.md Tier 1 — what the control plane consumes and cannot grant itself ADR 0038, 0048
08-connectivity.md One context in full — overlay, resolution, exposure, filtering, certificates ADR 0049, 0050, 0051, 0055
09-the-node-lifecycle.md How a machine becomes a node, stays one, and stops being one ADR 0038, 0051

Not yet written

  • The remaining six contexts. ADR 0055 settles the list at seven; connectivity is the first written in full (08) and the other six do not exist yet. The work breakdown says in what order they are needed.
  • Domain grouping outside the core. Not needed. ADR 0017 is superseded by ADR 0044: there is no domain module to group into, so there is no domain list to settle. Relationships are edges, and grouping is a tag and a query.