Back to 23 records. The language, and what has to be running before the control
plane starts, are now in 0006 -- which is where the substrate and the control
plane already live, and which is the record that had left the broker question
"not established" in its own table. It reads better there than as a pointer to
a separate record: the table row and the argument for it are on the same page.
The store mechanics went into 0008. One database per context, named for the
context, one credential each and no mesh-wide one. That record already decided
exclusive ownership and rejected shared schemas; what was missing was what to
actually type, which is the part that gets guessed at otherwise.
Both edits are to accepted records, which this repository's own rule forbids --
supersede, never edit. Recorded here so it is visible rather than silent. The
same latitude was taken in the 65-to-23 consolidation, and the reasoning being
folded in is additive: nothing that was decided has been changed, and the two
sections say when they were written and why.
Two things found by trying to build tier 2.
The substrate design asked whether the message broker has to be running before
the control plane, and framed it as depending on whether the control plane's
own parts talk to each other over it. They do not -- it is one process -- so
under that framing the broker stays out of the bundle.
The framing cannot answer the question. What decides it is how the control
plane reaches a node, and the answer was already decided: only ever over the
link, and the link is the broker. So provisioning the broker would require the
broker. The first node does not escape this by being local, because it enrols
the ordinary way, by dialling the broker at the address in its token -- which
was deliberate, and worth keeping.
The bundle is two images now. The record says what that costs, including a
certificate the broker needs at a moment when there is no mesh to issue one.
The language had never been decided for tier 2. Go, for the same reason the
host is: the bundle pins this image by digest and runs it where nothing can
check it, so the image should hold the program and nothing else.
Also corrects something already built: the bootstrap created one database and
called it 'mesh'. ADR 0008 grants a context only what it exclusively owns and
ADR 0006 says the mesh database names a thing that will not exist. One database
per context, so one today, called inventory.
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.
Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.
Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).
Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.
Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.
The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18,
19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only
the archaeology of what used to be there.
Renumbered contiguously. Renames run in ascending order, so every target number
is already free and no two files ever collide.
The reference rewrite is one simultaneous pass rather than a sequence of
replacements. Numbers moved into slots other numbers were vacating -- the node
host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time
would have cascaded and silently pointed things at the wrong record.
Seven plain-text references survived the merges as prose rather than links,
naming records that no longer existed: the enrolment token, the link boundary,
what a declaration is, reachability, the repository structure. Each mapped to
the consolidated record that now holds it.
Verified rather than assumed: every [ADR NNNN](path) link now has matching text
and target, checked across the whole repository, and the checker passes.
Frontmatter `consolidates:` lists dropped -- they named records that are gone,
and each consolidated record already says in prose what it absorbed.
Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.
the node host 8 -> 1 applies not decides, depends on nothing,
per operating system, root service, the
launcher, episodic, what a declaration is,
actions from the bundle only
a node and how it joins 4 -> 1 what a node is, joining, the link as
security boundary, the enrolment token
modules and the graph 7 -> 1 everything is a module, no domain modules,
three edges, provisioning, the core library
substrate and control 6 -> 1 the test, seven contexts, one control plane,
plane the authority is not a database, the named
products, the pinned bundle
connectivity 3 -> 1 a route is a grant, reachability declared,
filter rules
delivery 5 -> 1 reconciliation not a pipeline, artifacts,
the three silos, a failed step, the verdict
the lab 5 -> 1 (earlier)
how this repository 10 -> 1 (earlier)
works
Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.
The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.
The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
Jochen: a normal application has 3-5 ADRs, maybe 10 for a large one, and we are
at 65. Fair, and the cause is mine -- I recorded every FINDING as a decision
rather than every fork in the road.
Two merges, both cases where one decision had been split across many records
because it was taken over several days rather than at once.
0019 absorbs ten records about how this repository works: what it is and that
it is public, the folder flow, the two design layers, the issue front door,
status in frontmatter, playbooks, the naming rule, the product name. Those were
never ten decisions -- they were one, seen from ten angles as the repository
took shape.
0016 absorbs the five about the lab: a node is a virtual machine, a router is
scenery, a scenario declares the underlay, a scenario is a closed address
space, and the two scenario classes. Same pattern -- one design, split by the
order it was worked out in.
The consolidated 0019 also raises the bar for what earns a record, since that
is what produced 65: a record is warranted when there is a genuine fork -- a
direction reversed, an alternative that will be proposed again, something
contested. A finding is not a decision, and a bug is certainly not. Everything
else belongs in the design document where the reasoning is actually read.
The checker earned its place here. Deleting nine records left 13 dangling links
across the repository and it named every one, including in AGENTS.md. Nothing
was found by reading.
Remaining clusters worth the same treatment: the host (8 records), delivery
(5), modules (6), connectivity (4), substrate and control plane (4). That would
be 52 down to roughly 30.
A sweep for claims overtaken by the last few days. Annotated rather than
rewritten, following the pattern already in 0049 -- what changed and why is the
useful part, and an accepted record should not quietly become something else.
0057's init section was wrong on all three of its claims. It said the host
needs FOUR things from an init; 0061 reduced that to one. It said every machine
the mesh targets already has systemd; Alpine does not, and it is the intended
first node. It said there is no second init to abstract over; there is now, and
the answer is still not an abstraction -- it is a four-line file per system.
What survives is the part that was always right: an init is not a dependency in
0041's sense, because it is not installed, it is what the machine already is.
0048 named Docker as the container runtime. It is now docker or podman,
detected rather than chosen -- because adoption keeps what a machine already
has, so naming one contradicted a rule already decided. That row is the only
one of the five that names two, and the record now says why.
0060 claimed the bundle is portable across operating systems. Its mechanism is;
its contents are not -- package names, unit names, service names all differ, so
an Arch host embeds an Arch bundle. That was my error, and it is the exact
confusion behind the question that found it.
The design layer had the same drift: 07 and 09 said "Docker" where they meant a
container runtime, 09 said systemd restarts the host after an upgrade when the
launcher does, and both install snippets assumed Arch. They now show Alpine and
Arch side by side, which makes the point better than prose did -- step 1
differs per system, step 2 never does.
Checked and NOT changed: 0047's "the vocabulary grows by one shape" is a claim
about the rate, not the count, and is still true. 0037 lists docker among tools
the host manages, which it does. 0041 says nothing about either.
All six shapes are built. 07 still said the last three did not exist, and 05
still described stage 2 as having built three of six.
Records what the lab still cannot do, because that is now the only thing
between here and an end-to-end substrate bootstrap: a sealed scenario cannot
fetch an image and its machines carry no container runtime, so package,
container and action were verified against a real machine instead.
ADR 0048 named ingress as an unclosed hole -- nothing said what terminates
TLS, how a public name reaches a container, or which tier owned it. Resolving
it needed no new concepts, which is why it survived: nobody had applied the
rules already written to it.
Ingress is not substrate. The control plane does not need a route to start,
and no node needs one to reach it -- the node dials out and has no listening
control surface. It grants itself a route afterwards, like a bucket.
A route is an instantiation edge under ADR 0044. The direction mirrors a
database -- the consumer supplies a target and receives a name rather than
credentials -- but it is the same edge.
The substantive finding is that exposure is three facts at two scopes: name
resolution and certificate issuance need to know which node is publicly
reachable, and only the proxy mapping is a single machine's business. That is
why it belongs to the connectivity context, and why Traefik doing all three on
the node is wrong.
Which matters beyond tidiness: research 006 counted traefik as one of two
modules opening a direct Postgres connection, reading nodes and mesh_ca. That
violates 0037, 0045 and 0039 at once, and is why every node permanently holds
a credential to the control plane's database. Deriving the config centrally and
delivering it as `file` resources removes it, costs zero new host vocabulary,
and closes the set 0039 identified -- wireguard was the other.
Left open deliberately: the mesh's internal CA is the other thing traefik
reads, and it belongs to the link's mutual authority, not to exposure.
Conflating the two is what made the gap hard to see.
Also fixes an inconsistency from the previous commit: 06 still claimed the
virtual host was raised from the bundle.
Proposed, not accepted -- for review.
The design layer described every service by role and never once by name:
Postgres appeared in zero design documents. That was over-application of the
research rule "never identify the mesh it observed", which is about node names
and domains, not software.
Two things were actually broken by it. substrate.lock pins images by digest and
a digest belongs to a named image, so the bundle could not be written from the
design. And a reader could not tell a settled choice from an unexamined one --
"a relational store" reads identically either way.
ADR 0048 names them: PostgreSQL, LavinMQ, MinIO, an OCI registry, Docker. The
argument for each is continuity, which is a real argument -- replacing a
substrate service migrates the mesh's own state. Role and product are now both
written, because the design depends on the protocol while the installer needs
the product.
Also separates two questions the substrate doc had merged: being substrate and
being in the bundle. Only Postgres must precede the control plane; the rest are
substrate by role and ordinary by delivery. Whether the bus joins it is left
open, because it turns on the control plane's internal shape.
Names the forge as Gitea, and records ingress/Traefik as an unclosed gap rather
than a naming one -- nothing says what terminates TLS or which tier owns it.
Fixes a miscount: the host's bootstrap vocabulary is six shapes, not five.
Asked whether postgres has to be installed, and the answer exposed a missing
step. The store is a container, so something must run containers before anything
else happens — and a container runtime is a PACKAGE, not a container.
Step 0 is where several threads meet. It is what the host's capability detection
already reports, and the first use of that report by something other than a
person. It is adopted rather than installed when the machine already has a
runtime with configuration somebody chose. And it is a package, needing the
machine's own package manager and a network, both of which ADR 0046 permits.
So the host's bootstrap vocabulary is six shapes: package, container, file,
directory, service, action. Stage 2 built three of them.
The node host design now names which three remain and why the lab cannot yet
exercise them — a sealed scenario fetches nothing and its machines carry no
container runtime, which is lab-installation work rather than a constraint on
the design, because production machines have a network.
The bootstrap's sharpest open question, and the framing was wrong. "State on
this machine" was being read as the filesystem and the service manager. A
service running on this machine IS part of this machine — writing a file and
creating a database in a local store differ in mechanism, not in scope.
The real question was underneath: must the host learn what a database is? It
must not. Giving it a `database` resource type means tier 0 knows Postgres, then
a bucket, then a virtual host — the host acquiring the substrate's vocabulary
one service at a time, which is what ADR 0037 exists to stop.
So the bundle declares an ACTION and the host runs it and verifies it. What a
database means stays with the module that provides one; the host knows only how
to run a declared action against something local and check the result. Its
vocabulary grows by one shape rather than by one resource type per service.
Actions are permitted in the bundle and forbidden over the link, and the
asymmetry is deliberate. A bundle arrives WITH the binary: anyone able to put a
hostile action in it could equally have put it in the host itself, so refusing
actions there buys nothing and costs the bootstrap. The link is a separate
party, reachable separately, and an action there is the unbounded blast radius
ADR 0039 refuses. That decision stands unchanged.
And ongoing provisioning is not the host's at all — the control plane does it
once a mesh exists — so the asymmetry costs nothing.
Which dissolves the earlier worry about one mechanism with a tier boundary
inside it: there are two mechanisms, with different actors, scopes and trust
models, and that is the answer rather than a compromise.
Named rather than hidden: this is the escape hatch research 011 warned about,
arbitrary code in the place hardest to remove later. It is bounded by being
bundle-only and by every action having to declare how it verifies itself, and
that boundary is the whole defence.
The blocking question was where a container image comes from, and the version
that blocked assumed the machine might have no network. That assumption came
from the LAB: a scenario is a closed address space by design, which is what lets
two scenarios hold the same addresses without meeting. Production is not sealed
— a machine being adopted has a network, and one that does not is a machine
where very little works anyway.
So substrate.lock carries references, not payload: an image name and a digest,
fetched at apply time. A first node pulls from upstream because no mesh registry
exists yet; every node after that pulls from the mesh's own. The lab is the
exception and places images itself, the way it already places the host binary —
a property of a test environment, and letting it dictate the production design
would be the tail wagging the dog.
Pinned by DIGEST rather than tag. Reproducibility comes from pinning the
identity of a thing, not from carrying its bytes, which is what makes fetching
acceptable rather than a compromise.
ADR 0041 survives untouched, which was the point. "Copy it onto a machine and
run it" stays literally true — one binary, a few megabytes, which then fetches
what it was told to. Carrying images would have quietly redefined the property
that decision rests on.
Costs accepted and named: an apply can now fail because something is
unreachable, which a self-contained artifact could not, so it must fail legibly
— naming what it could not fetch and from where. And the lab needs a way to
place images into a machine that also has no container runtime, both of which
are lab-installation concerns and neither solved here.
Research 012's build-time-versus-apply-time reframing narrows accordingly: it
still holds for what a tailored installer contains, and no longer has to hold
for images.
Same gap as the control plane: load-bearing and unpinned.
The substrate is what the control plane CONSUMES AND CANNOT GRANT ITSELF. Every
module needing a database asks provisioning for one; the control plane needs one
too and cannot ask itself, because it is not running yet. That circularity is
not an awkwardness to work around — it is the definition, and anything on the
wrong side of it must be raised by the bundle the host carries.
Which answers 006's open question in the honest form rather than with a number.
The identity provider is substrate only if the control plane DELEGATES
authentication — then it cannot serve anybody before the provider exists and
cannot grant itself a client. If it authenticates natively, the provider is an
ordinary hosted service. So the count follows from a decision not yet taken, and
asserting four was asserting that decision.
The test also rules out the tempting wrong answer: an identity provider, a mail
server and an analytics service are all infrastructure by any ordinary reading,
and none are substrate, because the control plane starts and runs without them.
Important is not the test.
Records why the bundle is pinned by hand — it is applied when no mesh exists, so
nothing can resolve a version or ask a registry — and why it must be
self-contained, which makes it an artifact built on a machine with a network for
a machine that may have none.