The connectivity design still said a hub cannot be filtered — a gap recorded in
the morning and closed in the afternoon, left standing as though it were
current. Worse than a stale date: it would send somebody away from something
that works.
`restart-on` was described nowhere, including the part added today that lets a
service reflect a file another module put on the machine. A rule the host
enforces and no document mentions is a rule nobody can rely on.
And nine of fifteen design documents claimed an `updated:` older than their last
change, some by a week. That field is what cross-cutting views are generated
from, so it is not decoration.
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.
Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.
Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).
Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.
Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.
The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18,
19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only
the archaeology of what used to be there.
Renumbered contiguously. Renames run in ascending order, so every target number
is already free and no two files ever collide.
The reference rewrite is one simultaneous pass rather than a sequence of
replacements. Numbers moved into slots other numbers were vacating -- the node
host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time
would have cascaded and silently pointed things at the wrong record.
Seven plain-text references survived the merges as prose rather than links,
naming records that no longer existed: the enrolment token, the link boundary,
what a declaration is, reachability, the repository structure. Each mapped to
the consolidated record that now holds it.
Verified rather than assumed: every [ADR NNNN](path) link now has matching text
and target, checked across the whole repository, and the checker passes.
Frontmatter `consolidates:` lists dropped -- they named records that are gone,
and each consolidated record already says in prose what it absorbed.
Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.
the node host 8 -> 1 applies not decides, depends on nothing,
per operating system, root service, the
launcher, episodic, what a declaration is,
actions from the bundle only
a node and how it joins 4 -> 1 what a node is, joining, the link as
security boundary, the enrolment token
modules and the graph 7 -> 1 everything is a module, no domain modules,
three edges, provisioning, the core library
substrate and control 6 -> 1 the test, seven contexts, one control plane,
plane the authority is not a database, the named
products, the pinned bundle
connectivity 3 -> 1 a route is a grant, reachability declared,
filter rules
delivery 5 -> 1 reconciliation not a pipeline, artifacts,
the three silos, a failed step, the verdict
the lab 5 -> 1 (earlier)
how this repository 10 -> 1 (earlier)
works
Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.
The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.
The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
Jochen: a normal application has 3-5 ADRs, maybe 10 for a large one, and we are
at 65. Fair, and the cause is mine -- I recorded every FINDING as a decision
rather than every fork in the road.
Two merges, both cases where one decision had been split across many records
because it was taken over several days rather than at once.
0019 absorbs ten records about how this repository works: what it is and that
it is public, the folder flow, the two design layers, the issue front door,
status in frontmatter, playbooks, the naming rule, the product name. Those were
never ten decisions -- they were one, seen from ten angles as the repository
took shape.
0016 absorbs the five about the lab: a node is a virtual machine, a router is
scenery, a scenario declares the underlay, a scenario is a closed address
space, and the two scenario classes. Same pattern -- one design, split by the
order it was worked out in.
The consolidated 0019 also raises the bar for what earns a record, since that
is what produced 65: a record is warranted when there is a genuine fork -- a
direction reversed, an alternative that will be proposed again, something
contested. A finding is not a decision, and a bug is certainly not. Everything
else belongs in the design document where the reasoning is actually read.
The checker earned its place here. Deleting nine records left 13 dangling links
across the repository and it named every one, including in AGENTS.md. Nothing
was found by reading.
Remaining clusters worth the same treatment: the host (8 records), delivery
(5), modules (6), connectivity (4), substrate and control plane (4). That would
be 52 down to roughly 30.
Playbook 02 and 04 were followed for the substance — decisions before design,
design before build — and skipped for the bookkeeping. This closes that.
004 graduates. Its one open item was "not yet stood up"; the lab is stood up,
and the substitution the effort turned on is now enforced by the validator
before anything is raised rather than left as a thing to remember. Its
certificate conclusion has a home in 01-end-to-end-testing and is designed but
not built — implementation is a third axis, and an effort graduates on its
conclusions.
One item leaves 004 without a home and is recorded rather than lost: the reverse
proxy does not set caServer, so it defaults to the production endpoint.
The two lab designs read `designed` while running in production of a sort, so
they become `in-progress`.
And the lab gets an as-is document, which it did not have. It records what runs
including the parts nobody would choose again: that `place:` is refused and the
lab therefore raises EMPTY MACHINES, that the drawing shipped with no design
document behind it, that a router is tagged as a machine for a reason found by a
bug, and that the integration suite raises two of five scenarios while both
faults found so far lived in the three it does not.
006 stays active, deliberately. Two of its open questions ARE the tier 0 design
— whether absorbing six concerns makes the host too large, and whether an
unprivileged node earns a place in the inventory. Playbook 04 is explicit that
an open question is a reason to research, not to build around.
The lifecycle design asked whether a scenario snapshot needs the machines
stopped. The integration test answered it on its first run: no, but they
must be flushed.
A snapshot captures disk and not memory, so a write still in the guest's
page cache is absent from it — not stale, absent. A file written seconds
before a snapshot did not survive the restore.
Flushing first buys write-durability. It does not buy
application-consistency: anything mid-transaction is still captured
mid-transaction, and that limit is now stated rather than left implied.
The lifecycle design closed on a question that was measurable rather than
arguable, so it was measured. One virtual machine on a workstation with
hardware virtualisation and NVMe.
Raising: the launch call returns in 3.4s, the machine is actually usable
after 14.3s. The gap is a design constraint — raise must wait for the
second number, because reporting the first would be transport reported as
effect, which is the mesh's own recurring failure.
Snapshot: 9.9s and 1.6 GB for a 1.5 GB instance. A dir snapshot is a full
copy; nothing is shared. Restore: 10.4s, usable again after 20.1s.
The second snapshot exceeded two minutes and never completed. That is the
more troubling number: snapshot cost here is not merely high, it is
unpredictable, and a loop with a variable multi-minute step is one nobody
trusts.
Projected to a four-machine scenario, a reset-and-rerun cycle is about a
minute and a half at best and unbounded at worst, before any of the mesh's
own work begins. That is too slow for an inner loop, and ADR 0029's whole
argument — that making the bootstrap path the inner loop turns the
least-exercised code into the most-exercised — holds only while resetting
is cheap.
The cause is not virtual machines. Hardware virtualisation is present and
machines boot in fourteen seconds. It is that the daemon offers exactly one
storage driver, dir, which has no copy-on-write and therefore no cheap
snapshot. The btrfs kernel module is available; btrfs-progs is simply not
installed, which is the entire reason the driver is absent.
The copy-on-write comparison was deliberately NOT run, because running it
would mean installing a package by hand — which the rules forbid and which
would have made the measurement unreproducible. So the honest statement is
that the current configuration is too slow and the likely fix is known but
unverified, rather than that btrfs fixes it.
ADR 0032: a scenario is a closed address space. Every segment materialises
as its own isolated link belonging to one instance, so two scenarios raised
from the same declaration hold the same addresses and never meet. The
declaration keeps its literal addresses and they mean what they say —
allocating from a pool would have made them a fiction, so a scenario
reproducing a specific topology would stop reproducing it.
The constraint that follows shapes everything: the lab never reaches into a
scenario over IP. It talks to machines through the virtualisation layer's
own channel. If it reached them by address, the workstation would need a
route into each scenario, and two carrying the same prefix would give it
two routes to one destination — failing not with an error but by one
scenario's traffic arriving in another.
That also makes reachability an honest question. Can this machine reach
that one is asked from INSIDE, by executing on the first, rather than
probed from a workstation that is not on the network and whose opinion
would be a different question with a misleadingly similar answer.
The lifecycle itself: six verbs, of which raise and destroy are enough to
be useful and the rest are what make repetition cheap. Raising is
convergent rather than incremental, because a lab behaving differently
from the thing it tests teaches the wrong habit.
A failed raise leaves the wreckage standing. Tearing down on failure
destroys the only evidence, which is backwards — a scenario that failed to
raise is more interesting than one that succeeded.
Snapshots are whole-scenario. Per-machine would be cheaper and wrong: the
mesh keeps state spanning nodes, so restoring one machine while its peers
move on produces a mesh that has never existed, and faults found there
would be artefacts of the lab.
Closes the declaration's open question about running several scenarios at
once.