Issue 003 is answered in both halves: manifests are parsed strictly, and a
module says what it listens on and from where rather than carrying a key
nothing reads. The design records what was built and how each part is checked.
Issue 013 is new, found by reading while writing the first module that has
both a computed file and a service that needs it. The file arrived second.
It failed, then the next reconcile fixed it, which is why nothing caught it.
0009 already said a consumer supplies a target and receives a name. What
it did not say is that those are two separate mechanisms.
Contribution — publish me at this name, on this port — now exists.
Binding — and hand me back a credential — does not, and is the larger
half: a secret has to exist, be stored, reach one node and not the
others, and rotate with every holder informed. That is the invariant set
found violated three ways at once, so it is not something to add in
passing.
The absence had a measured cost. Exactly two modules opened a direct
connection to the control plane's database, and they are the reason every
node permanently holds a credential to it. Both were doing by hand what
this edge is for. Neither needed a new kind of thing.
Two records, from building it.
0009 has a section titled "there are no domain modules", and `networking`
now exists. It is not a contradiction and it reads as one, so the
difference is written down: what was refused contains WireGuard and a
proxy and is assigned where half of it is unwanted. What exists contains
nothing — requirements and a name — so there is no half. Every artifact
it leads to is still an ordinary module assigned on its own terms.
With the cost stated, because it is real: adding a second implementation
turns a settled question into an open one for everyone using the bundle,
not only for whoever wanted the alternative. That is the refusing rule
applied consistently, and the alternative is a default, which is the
flavor field returning under a better name.
08-connectivity gains why the network stopped being code beside the
module system: a machine was on the private network because it had an
address, and there was no way to keep one off. A manifest can now say its
resources are computed, which is what a peer list needs.
And three modules rather than one, because WireGuard is one VPN of
several. Naming a module after the job and putting one implementation
inside it is flavor wearing a generic name — the second VPN has nowhere
to go.
All on the first three machines to actually run it, and all invisible from the
mesh's own state: the graph was right, the files were right, the services were
up, every node reported success, and the network did not work.
A running interface does not re-read its configuration, so a node joining left
every existing node carrying a network that no longer existed. A hub sharing a
site with a spoke was emitted twice, which WireGuard refuses. Two nodes at one
site that neither can be dialled were peered directly, so nobody opened the
path and the more specific route blackholed -- this document's own warning
arriving in its implementation. And Docker sets the FORWARD policy to DROP, so
a hub with forwarding enabled still carried nothing between its spokes.
The last one is the sharpest: the substrate at tier 1 silently breaks the
network at tier 2, and nothing in either tier's state says so.
None of these is reachable by reasoning, and each was found within minutes of a
real machine trying it. That is the argument for the lab in one line.
Four things settled by talking them through, all of which had been true in
somebody's head and written nowhere.
It is not a mesh in the peer-to-peer sense and will not become one. 0001 now
says what it is instead: machines linked by a private network, one node holding
knowledge of all of them, modules as the way anything is built and delivered,
and agents hired onto nodes to do the work. The word describes what machines
can reach, not how they are governed. "Master" overstates it the other way --
nothing needs that node to keep running, only to change.
0006 gains the option that would make it a real mesh, recorded as considered
rather than rejected by silence: every node holding the whole inventory, a
replication process, an elected master with promotion on failure. What settles
it is not the complexity but that it still would not deliver the name, because
application databases are not replicated -- so a genuine peer-to-peer mesh
means becoming a replicated database system for every consumer's data too. That
is a larger product than the thing it would support.
Also in 0006: three central roles, not one. Losing the control plane costs
change, losing the broker costs being told anything, and losing the hub costs
nodes in different places reaching each other at all -- which is operation, not
administration. Whether they are one node is not decided.
And SSH access is identity's. It appeared three times as something that uses
the overlay and never as something the mesh provides, which reads as settled
when nothing decided it. Nobody else could: the mesh is the only thing that
knows which humans and agents exist and which nodes they may reach. Node to
node SSH stays out -- the host has no inbound control surface by decision, and
nodes reaching each other that way is a second control path through the back
door.
0007 gains the requirement underneath all of it. Reachability was recorded as a
fact to track and never as a thing some node must have. The broker's node and
the hub must be dialable by every node at a stable address, or nothing can join
and a disconnected node cannot return. A mesh entirely behind NAT cannot be
raised. That is a precondition and it belongs with the others.
The link staying on the underlay is also argued now rather than asserted. At
join time it is forced; afterwards it is a choice, and the reason is that a
repair channel carried over the thing being repaired is not one. Moving it onto
the overlay, with fallback, is recorded as open with what it would have to get
right -- a WireGuard interface has no link state to test, and a silent fallback
is this repository's recurring fault in a new place.
0010 says in one line what was the intention throughout: the module system is
the CI/CD. Not a pipeline beside the mesh. Build, test, publish and deploy are
one reconciliation seen at four points, which is why a thing that cannot be a
module cannot be delivered.
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.
Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.
Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).
Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.
Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.
The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18,
19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only
the archaeology of what used to be there.
Renumbered contiguously. Renames run in ascending order, so every target number
is already free and no two files ever collide.
The reference rewrite is one simultaneous pass rather than a sequence of
replacements. Numbers moved into slots other numbers were vacating -- the node
host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time
would have cascaded and silently pointed things at the wrong record.
Seven plain-text references survived the merges as prose rather than links,
naming records that no longer existed: the enrolment token, the link boundary,
what a declaration is, reachability, the repository structure. Each mapped to
the consolidated record that now holds it.
Verified rather than assumed: every [ADR NNNN](path) link now has matching text
and target, checked across the whole repository, and the checker passes.
Frontmatter `consolidates:` lists dropped -- they named records that are gone,
and each consolidated record already says in prose what it absorbed.
Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.
the node host 8 -> 1 applies not decides, depends on nothing,
per operating system, root service, the
launcher, episodic, what a declaration is,
actions from the bundle only
a node and how it joins 4 -> 1 what a node is, joining, the link as
security boundary, the enrolment token
modules and the graph 7 -> 1 everything is a module, no domain modules,
three edges, provisioning, the core library
substrate and control 6 -> 1 the test, seven contexts, one control plane,
plane the authority is not a database, the named
products, the pinned bundle
connectivity 3 -> 1 a route is a grant, reachability declared,
filter rules
delivery 5 -> 1 reconciliation not a pipeline, artifacts,
the three silos, a failed step, the verdict
the lab 5 -> 1 (earlier)
how this repository 10 -> 1 (earlier)
works
Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.
The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.
The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
Closes the two open questions in 06 and 08, which turned out to be one
question: how many control planes run, and what happens when the hub is down.
Both were drifting toward redundancy by default -- a standby plane, a second
hub, an election to pick between them. That is not one feature but a property
every layer must then honour, and each layer gets it wrong independently.
Not wanted, and not needed. A handful of machines with one node hosting the
registry is not a distributed system.
The argument for why this is sound rather than merely cheap is that the design
already tolerates it by construction. ADR 0036 makes reachability state rather
than class; the host reconciles from its own store (0043) and never needed to
ask anybody to hold the state it was last given. So the control plane being
down is not a new failure mode -- it is every node in the ordinary disconnected
situation at once. What is lost is change, not operation.
No node holds a contended role: the control plane is assigned like any other
module, and the overlay hub is declared (0050). No promotion, no quorum, no
fencing, no split brain, no replicated store, and no "which node is
authoritative" recurring at every layer.
Two consequences stated plainly rather than buried. The control-plane node is a
single point of failure -- deliberate, and said out loud so it stays
deliberate. And recovery is restore rather than failover, which makes backup
the availability story rather than hygiene.
The sharpest one is the clock: the control plane owns certificate issuance
(0049), so an outage outlasting a renewal window expires every public name.
That bounds how long recovery may take, and nothing measures it today.
Written as one document because the five are one design. They share inputs,
they must agree, and every one of them today is computed in a different place
by a different module from a different copy of the same facts.
The through-line is that none of the five can be answered by a machine alone,
so all five are decided centrally and delivered as `file` resources. That costs
no new host vocabulary and removes both remaining direct database connections
from nodes -- wireguard and traefik are the only two, and both are connectivity.
Three decisions fall out, all proposed:
0050 -- reachability is declared, not inferred from an address. The RFC1918
regex is wrong for carrier-grade NAT (100.64/10 tests as public, so an endpoint
is written to an address nothing can reach), wrong for IPv6, and wrong for a
routable address behind a closed firewall. The lab needing TEST-NET-3 to
satisfy the regex is the same bug from the other side. Also kills hub election
by address prefix, which fails silently and makes renumbering an outage.
0051 -- the enrolment token carries where the mesh is and how to recognise it.
Closes two circles with one mechanism: verifying the mesh needed the CA, and
obtaining the CA meant trusting whoever handed it over; and a node had to reach
the mesh before it could resolve any mesh name. An address plus a fingerprint,
carried out of band, resolves both -- and closes the CA question 0049 deferred.
0052 -- a filter rule names its source. `scope:` is declared in five manifests,
is part of no rule type, and is referenced by no code, so those manifests
appear to restrict ports and restrict nothing. Removed rather than implemented;
the general fix is refusing unknown keys, which the host already does and
manifests do not.
Also corrects two claims in 0049 asserting wireguard was already handled.
Research 006 says both modules still reach upward; neither is.