Compare commits

...
Author SHA1 Message Date
jschoubben ba30286896 Merge pull request 'The pointers back from what yesterday's records changed, which I missed twice' (#197) from decision/the-pointers-back-from-what-these-narrow into main 2026-09-29 22:47:45 +00:00
jschoubben 53b94c51bb The pointers back from what yesterday's records changed, which I missed twice
Three records were left describing a mechanism a new record had moved,
and a reader arrives at them by following a citation: 0066 still said a
routed name is written into every container after 0148 replaced that with
resolution; 0016 still read as though the lab were the test bed after
0149; and issues 109 and 135 said nothing about 0148 ending the copying
that 135's own fix made comparable. Each was a citation leading to the
wrong answer in a record that was not wrong about anything it decided.

This is the second time in one session. The playbook rule I added last
round did not stop it, so the convention is now written where the record
conventions live, with the shape to use and three worked examples — and
with the honest note that it is NOT machine-checked and cannot be from
`extends:` alone: 102 records extend another, 87 have no back-reference,
and that is correct, because extending usually means building on a
context. Making it mechanical means a record declaring the relationship
in frontmatter, which is a schema change and is not mine to decide.

Also: designs 18 and 20 claimed `updated:` dates from before I edited
them, and 117's `fixed-by` gained the commit beside the record.
2026-09-30 00:47:19 +02:00
jschoubben 83791f0921 Merge pull request 'The four open design questions, answered: ADRs 0148, 0149, 0150, and 0114 accepted' (#196) from decision/0148-the-meshs-names-are-resolved-not-copied into main 2026-09-29 22:39:38 +00:00
jschoubben bf4d4e2e7b The three parked questions are answered: 0114 accepted, 0068 superseded, 117 settled
**0114 accepted.** The separation it draws — a consumer's resource and a
credential that reaches it are different things — is a data-loss rule,
and a record naming one should not sit unresolved while the code that
could hit it is written. Not built, and accepting it schedules nothing;
accepted-and-not-built is where 0141 and 0142 already are.

**0068 superseded by 0149, the live mesh is the test bed.** Never built,
and contradicted by practice that was written down nowhere but a handoff
note. The faults that cost the most are faults of a mesh that already
exists — bound consumers, containers made against an older roster, an
adopted machine — and a bed is by construction a mesh that does not.
Issue 156 settles it: that change was verified honestly, on the only path
where it cannot fail. The lab is not retired; raising a mesh from bare is
now its whole job.

**117 answered by 0150.** A module's own code runs as supervised
processes under the module's one account. 0047's argument was for a
runtime per module, not for a container — a unit satisfies it and runs as
an account. The invariant is the account, not the process count: its
worry was a second identity to scope and seal, and processes sharing one
account create none. Designs 18 and 20 now cite it and 0047 carries a
dated pointer, which closes the third disagreement — that neither design
knew the record existed.
2026-09-30 00:39:09 +02:00
jschoubben ec42ee0846 ADR 0148: the mesh's names are resolved, not copied into every container
Answers issue 151. Copying the roster into each container made the roster
part of each container's identity, so one name moving replaced every
container in the mesh — and it never stopped the staleness it was for,
since a copy taken at creation is stale the moment the roster moves (109,
135).

A container resolves through its machine's resolver instead, and nothing
is copied. Staleness stops being possible rather than detected, and a
name's blast radius becomes nothing.

Scoping each container to the names it binds was the close call and is
rejected: it contradicts anything-calls-anything, and leaves the roster
in the digest so the churn returns for a widely-bound name.

Gated on issue 110 — a container on the runtime's default network has no
DNS at all today. Removing the copy first reintroduces 109 and 135
silently on a live mesh. 151 stays open until the code lands; design 08's
file-not-resolver passage is narrowed to the machine's own roster.
2026-09-30 00:34:30 +02:00
jschoubben 5a3dee9e9e Merge pull request 'The records pointed at branches that no longer exist, and two fixes had no sequel' (#195) from issue/pointers-that-resolve-to-nothing into main 2026-09-29 22:29:02 +00:00
jschoubben 47909c6b71 The records pointed at branches that no longer exist, and two fixes had no sequel
Three issues named the branch that fixed them, and a branch is deleted
when it merges — so every `fixed-by:` was a pointer that resolved to
nothing by the time anyone followed it. They name commits and pull
requests now, and playbook 03 says to.

Two records were missing the thing a reader arrives for. 146 did not say
that one of its fixes crash-looped the control plane on a running mesh,
which is the whole reason the delivery subject carries the stream and the
raise path was the only one exercised. 151 did not say that 152 removed
the false reasons its roster moved, or that it stays open for the real
ones.

ADR 0080 enumerates what cycle.py enforces and named four things; it
enforces five. A progressive insight names the fifth — the decision
stands, the list had gone stale. The checks README and playbook 03 gained
the same rule, and 155 points at all three.
2026-09-30 00:28:35 +02:00
jschoubben 7e5edab8da Merge pull request 'Issue 156: the wider grant has landed, and the notes the fix printed' (#194) from issue/156-the-grant-has-since-landed into main 2026-09-29 22:19:35 +00:00
jschoubben 3649f82204 Issue 156: the wider grant has landed, and the notes it printed
Measured on the broker's own config after the fixed controller composed
it: every node is now allowed _DELIVER.<node>.> as well as the bare
subject. That grant was missing when this was diagnosed, which is why
re-making the consumers would have silenced every machine.

Also records the five consumers it kept and reported, including the build
machine's worker, which is bound the same way and was not anticipated.
2026-09-30 00:18:57 +02:00
jschoubben b400ce2c54 Merge pull request 'Issue 156: moving a consumer's delivery subject stops a running mesh' (#193) from issue/156-a-consumer-that-works-is-not-replaced into main 2026-09-29 21:57:00 +00:00
jschoubben 995c8cb266 Issue 156: moving a consumer's delivery subject stops a running mesh
146's change is correct on a mesh being raised and fatal on one that is
running: the server will not move a push consumer's delivery subject
while a subscriber is bound, and a node is bound to its declaration
consumer the whole time it is up. Reproduced before fixing; an earlier
version of the test unsubscribed first and passed against the broken
code.

The consumer that works is kept and the assertion says so. Not re-made:
the nodes are granted the bare subject only, so re-making would have
silenced every machine.
2026-09-29 23:56:31 +02:00
jschoubben f89aef992d Merge pull request 'Two records shared a number, twice, and every check passed' (#192) from issue/two-records-share-a-number-and-nothing-says-so into main 2026-09-29 21:39:09 +00:00
jschoubben b967ef7be3 Two records shared a number, twice, and every check passed
Two machines filing issues in the same hour both read main correctly and
both took "the next free number". main lags every open pull request —
seven that evening — so they collided twice. The second collision
reached main with records, cycle and index all reporting success.

cycle.py now refuses a tree where two issue folders share a leading
number, and names both. Proven by adding a duplicate and watching it
fail. The colliding records become 153 and 154, renumbered in the branch
that lands last, because renumbering a branch whose author is still
pushing only moves the race.

The check catches a collision; it does not prevent one. Taking a number
still means reading the open pull requests as well as main — issue 155
says so.
2026-09-29 23:38:48 +02:00
jschoubben e417906241 Merge pull request 'Issue 150: a machine's own network is not a reach' (#190) from issue/150-a-machines-own-network-is-not-a-reach into main 2026-09-29 21:37:41 +00:00
jschoubben b1bf895688 Merge pull request 'Issue 149: an adopted machine's data cannot be placed where it is' (#189) from issue/149-adopted-data-cannot-be-placed-where-it-is into main 2026-09-29 21:37:34 +00:00
jschoubben 0d64677c70 Merge pull request 'Issue 152: a node the mesh could not read withdrew its names from every machine' (#191) from fix/152-a-lookup-failure-is-not-an-absence into main 2026-09-29 21:37:16 +00:00
jschoubben 04205c1dc8 Merge pull request 'Issues 147 and 148: a route before its module is taken; a new name recreates every container' (#188) from issue/147-148-found-migrating-ace into main 2026-09-29 21:37:09 +00:00
jschoubben 9dc49cd831 Merge pull request 'Grooming: five issues were fixed and never closed, and one is not' (#187) from grooming/stale-issues into main 2026-09-29 21:37:04 +00:00
jschoubben 66b413a076 Merge pull request 'Issue 140 is resolved, and was resolved before it was read again' (#186) from issue/140-resolved into main 2026-09-29 21:36:57 +00:00
jschoubben fcba05fed9 Merge pull request 'Issue 146: a first node now enrols, and is enrolled twice' (#185) from issue/146-diagnosis into main 2026-09-29 21:36:42 +00:00
jschoubben 18f37c25b2 Issue 152: the loop is metastable, and it cleared at 23:11
It ran 22:46-23:11, five full replacements, and stopped when a pass
happened to read the store during a window it was up. The record said
nothing about it stops; that was wrong. Exiting by luck is the finding,
not a mitigation.
2026-09-29 23:27:51 +02:00
jschoubben 3c2b4fc6b6 Issue 152 is fixed: a lookup failure is no longer an absence
The three gatherers pass over a node whose set does not compose, and
raise anything else. 151 stays open: this removes the false reasons a
roster changes, not the fact that a real change still replaces every
container.
2026-09-29 23:26:20 +02:00
jschoubben 72eaf52867 Issue 152: a node whose plan will not compose withdraws its names from every machine
ace numbered its two records 147 and 148, which this repository already
uses; they become 150 and 151, as 127 became 149.

The new record is why the control node could not stop applying: the
roster alternates between two values because routeNamesInTheMesh
swallows a per-node plan failure, and the roster is part of every
container's identity. The loop closes through the control plane's own
store, which each pass replaces.
2026-09-29 23:13:54 +02:00
jschoubben 741625e725 Merge remote-tracking branch 'origin/issue/147-148-found-migrating-ace' into issue/the-roster-flicker-recreates-every-container 2026-09-29 23:12:57 +02:00
jschoubben c3730b9a23 Issue 150: a machine's own network is not a reach
ADR 0138's reach is internal, public or both. ace's LAN devices (an IoT
switch on mosquitto, unifi access points, plex clients) are none of them;
the only reach that admits them is public, which states the wrong thing and
opens the port to the internet on any machine with a public address.
2026-09-29 23:06:08 +02:00
jschoubben 3a59099c81 Issue 149: an adopted machine's data cannot be placed where it is
ADR 0112 decided that an assignment places a directory and says where an
access's data is; the control plane implements neither. ace's media stack
(a 40 TB ZFS library, plex state on a second disk) cannot be expressed
without writing ace's paths into manifests.
2026-09-29 23:03:11 +02:00
jschoubben 3bd6f34de3 Issues 147 and 148: a route before its module is taken; a new name recreates every container
Both found migrating the first module on ace (searxng), adopted with route-adapter:
147 — assign is not a no-op for a routed module: its route reaches the predecessor's
proxy while its container is held, and the name serves a dead backend until take.
148 — every push that moved the mesh's names replaced every container on novox, twice,
the control plane's own store included.
2026-09-29 23:01:51 +02:00
jschoubben 2b5119ecd2 Issue 114 is answered: the controller is a process, by ADR 0142
It asked a question rather than reporting a defect, and the question was taken
five days later — the mesh's own components are binaries on the machine, and
third-party software stays a container because an image is the right way to
carry somebody else's build. The delivery of them is issue 142.
2026-09-29 22:38:11 +02:00
jschoubben 96bdffa9bc Two records were numbered 127; the second becomes 149, and is resolved
A number identifies a record, and two were given 127 on 2026-09-27. The one
three documents and three source files cite by number keeps it; the other
becomes 149, says so in its own heading, and its two inbound references are
repointed.

It is also resolved: an empty declaration is sent carrying owns_nothing rather
than skipped, and the host refuses an empty body that does not carry it, so
emptiness cannot be read as a truncated declaration.
2026-09-29 22:33:40 +02:00
jschoubben 4eb16f1028 Three proposed records were already built; two are still yours to call
0037 (where a module lives), 0113 (the vault makes every secret) and 0115 (one
assignment of a module per node) described arrangements the mesh has — a
catalogue of other people's software plus a table filled by module add, a vault
answering a secret requirement six modules make, and a rule the assignment
table's primary key already enforces. Each is accepted against what was built,
and says so in its own words.

0037's other half is not built: a manifest outside this catalogue has no check,
which is issue 148.

0068 (the lab takes requests) and 0114 (a shared credential rotates over two
credentials) stay proposed. Neither is built, and both are decisions rather
than records of something that happened.
2026-09-29 22:31:55 +02:00
jschoubben d199de40db Merge the 146/147 records, which 006's note links to 2026-09-29 22:31:46 +02:00
jschoubben 9a1dc4665c Grooming: issue 006's knowledge base is the predecessor's, and is gone
The record catches README claiming this repository is indexed into a knowledge
base. Nothing indexed it, and since the cut-over there is nothing to index it
into — the surface that answered is on the transport the mesh removed (issue
147). Noted where the record is, so the next reader does not go looking for a
search that cannot exist.
2026-09-29 22:26:13 +02:00
jschoubben 14be8576f8 Grooming: five issues were fixed and never closed, and one is not
088, 089, 120, 128 and 130 each name a commit that is on main and cites them —
the forge's address following a moved port, a route naming its endpoint, a
provisioner asking the backend what is there, the hosts file written into a
marked block, and undeclaring giving a unit back the state it was found in.
Each says it was closed by reading commits rather than by a run, so nobody
reads a green that was never measured.

129 stays located on purpose: ca-trust is merged and no machine holds it, so
the symptom it opened on is still true everywhere.
2026-09-29 22:25:25 +02:00
jschoubben ec8676c225 Issue 140 is resolved, and was resolved before it was read again
Written 2026-09-28, answered the same week by mesh-controller bdf965d and
c68d3a7, and never closed — so the mesh's own account said reach was declared
nowhere while three readers were reading it: the filter, the proxy's names, and
each authority's host policy. The resolution names them and what checks each.

What is left is retiring the older per-port keys the block replaces, which is
not a gap in what reach can say.
2026-09-29 22:19:24 +02:00
jschoubben 3c0f7082e6 Issue 147: the tool surface is not the mesh's, it is the predecessor's
Diagnosed from the configuration: the tool server is HAL's brain, a local
process on the workstation with the predecessor's broker URL in the assistant's
own config. No manifest, no assignment, no seat, no account. Nothing regressed
— the mesh removed a transport this program still dials, and the program was
never part of the mesh. The mesh has a tool model and nothing publishes an
operator-facing surface onto it.
2026-09-29 21:50:50 +02:00
jschoubben 0dd00e88b6 Issue 147: the operator's tools still dial the bus that was removed
Every tool call fails with 'AMQP not connected', on every node including the
local one, because the tool surface still opens an AMQP connection and that
transport was deleted at the cut-over. The mesh reports healthy throughout —
what broke is the thing standing outside asking it questions, so nothing the
mesh checks is about it.
2026-09-29 21:39:55 +02:00
jschoubben eef54917ec Issue 146: the double enrolment was two consumers sharing a delivery subject
Not about enrolment. A push consumer delivers onto an ordinary subject and
everything subscribed to it gets a copy; the controller's two consumers were
both named after it, so both were given the same subject and the one process
acted on every message twice. Enrolment is where it drew blood because a second
enrolment mints a second credential.
2026-09-29 21:32:52 +02:00
jschoubben e9b1010bc0 Issue 146: what made it slow, and what was changed so it is not 2026-09-29 17:45:25 +02:00
jschoubben f6ed3545b7 Issue 146: a first node now enrols, and is enrolled twice
Two more faults behind the three already fixed. The account a token is the
password of was never recorded, and the comment above the issuing code said
it was; issuing now records it. Placing the composed list at genesis is the
other half — the control plane says what it composed and whoever raises the
machine writes it beside the bus, because no declaration can reach a machine
that has not enrolled.

With that a first node enrols. It is then enrolled twice from one attempt,
each minting a credential, and it keeps the answer to the first while the mesh
keeps the second. The trail for that one stops at a duplicate that survived
message-id deduplication.
2026-09-29 17:36:40 +02:00
jschoubben 0b08cdfce1 Merge pull request 'ADR 0147: a module anchors the mesh's authority, and issue 146: the foundation cannot be raised' (#184) from decision/0147-a-module-anchors-the-meshs-authority into main 2026-09-29 14:06:41 +00:00
jschoubben bffd2af40c Issue 146 diagnosed: four faults stacked, three fixed, the fourth is genesis
Raising a first node hits them in order: the bundle's bus image named for a
registry that is gone; the certificate made by openssl in an image that has
none; enrolment dialling TLS at a bus that speaks first, then refusing its own
token for the empty bus. Each was right until the bus changed and nothing has
raised a foundation since.

The fourth is not a patch: the installer carries the bus's first user list and
the controller composes the rest through the bus being a module, and at genesis
there is no module — so the first node cannot be let onto the bus it just
raised.
2026-09-29 15:42:32 +02:00
jschoubben 4a51ea4b3a Issue 146: the foundation cannot be raised on the bus the mesh runs on
Found trying to run ADR 0147's bed. The older bundle raises the previous
broker and a control plane that refuses to start without MESH_BUS_NATS; the
newer one stops a step earlier, asking for a certificate from openssl in an
image that has none. No bed can run while this holds, so 0147's check section
now says what actually stands behind it — the rendering, not a machine.
2026-09-29 15:25:58 +02:00
jschoubben 6c14d313b8 ADR 0147: a module anchors the mesh's authority on a machine, and takes it away again
Issue 129: every internal HTTPS name fails verification on every machine,
because nothing has ever written the mesh's root into a trust store. The
report proposed the controller inject it the way the private network writes
the registry's trust; this record rejects that — reachability and trust are
not the same fact, and where anchors live is the host's difference, not the
controller's. A module requiring internal-acme-ca does the whole of it, and
being unassigned undoes it.
2026-09-29 15:02:37 +02:00
mesh-admin ced547dae9 Merge pull request 'ADR 0146: connectivity is checked by name, per hosting form' (#183) from decision/0146-connectivity-by-name into main 2026-09-29 12:01:57 +00:00
jschoubben 38482435af ADR 0146: connectivity is checked by name, per hosting form, with a valid certificate
0145's core stands — a module checks what the mesh claims, from where the callers
are — and everything it said about what to dial was wrong.

A raw port is not how anything in this mesh is reached. A real caller resolves a
name, the proxy answers, the proxy reaches the service; dialling a port tests the
last hop of a four-hop path and skips the three where most connectivity lives.

So each hosting form gets its own endpoint, route and name:
connect-docker.<node>.internal for a container, connect-process for the mesh's own
code in a unit it writes, connect-unit for a unit a package ships — and the same
under each public domain. Every machine checks every machine, by name, over TLS,
verifying the certificate against the authority that should have issued it. The
names are the instrument: a failure reads as connect-docker.g14.internal did not
answer.

No name is written anywhere: the machines come from the roster fact, the labels are
the module's, and a machine that joins is checked by the others on the next push.

Five things this needs that the mesh does not have, named rather than assumed —
a module every machine has (ScopeNode is exclusivity, not obligation), a container
running a bundle, a unit a package ships, the public domain in the roster fact, and
issue 129, until which every internal name will correctly fail verification.
2026-09-29 14:01:55 +02:00
mesh-admin cf8a8d78c9 Merge pull request 'ADR 0145: a module checks what the mesh claims is reachable' (#182) from decision/0145-a-module-checks-what-the-mesh-claims into main 2026-09-29 11:44:58 +00:00
jschoubben 9de25994e9 ADR 0145: a module checks what the mesh claims is reachable, and it checks itself
The mesh asserts three things are callable and has never checked any of them. A
module on every machine serves an endpoint of its own and dials every other
machine's, from the position the callers are in — not the host and not the control
plane, both of which reach those addresses by paths no ordinary caller uses and
would have passed throughout the outage.

Its probe is its own endpoint, declared reachable over the private network, so it
is admitted by exactly the rule that governs every internally-exposed service. The
tempting target is a service every machine has, and those are the ones never closed
— ssh above all — which would have passed for all eleven hours.

Resolution and connection reported separately, because they have different owners.
One failure is not a fault, and the count travels with the result. It reports and
repairs nothing.

Built on a scheduled container, a rendered roster fact and a bundle: no new
vocabulary. Adopted on its own merits, which is what 0143 was not.
2026-09-29 13:44:56 +02:00
mesh-admin 64ea47b11d Merge pull request 'ADR 0144: anything on a machine may call anything on it' (#181) from decision/0144-local-is-not-a-boundary into main 2026-09-29 11:32:03 +00:00
jschoubben eba24a72af ADR 0144: anything on a machine may call anything on it, superseding 0143
Everything should be able to call what runs on the same machine, another machine's
service exposed to the private network, and another machine's service exposed
publicly. Three cases; the filter had two.

The first was expressed as the machines' own addresses on the private network. A
caller on the machine carries such an address; a caller in one of its containers
carries a bridge address and matched nothing — measured, same destination and same
machine: src 10.10.0.1 against src 172.17.0.8. The second case worked by accident,
because the tunnel rewrites a caller's address to the sending machine's. Two of
three working is why it read as correct.

0143 answered the wrong question. It proposed verifying each grant from the
consumer's own network position and went to length about which position, because
whether a caller sat in a container changed the answer — and that difference was the
bug. Observing a configuration error is not its remedy. Superseded, and nothing
replaces it; whether the mesh should check a grant is still open in issue 145 and
must stand on its own.

And a module is not a container: 61 of 72 happen to use one, 11 do not, and a rule
reasoning about containers describes most of the mesh rather than the mesh.
2026-09-29 13:32:01 +02:00
mesh-admin 78d4873f4f Merge pull request 'ADR 0143: a consumer verifies the grant it is given' (#180) from decision/0143-a-consumer-verifies-its-grant into main 2026-09-29 11:07:20 +00:00
jschoubben ddfd62edf6 ADR 0143: a consumer verifies the grant it is given
A grant is four facts and a credential — the provision, the machine, the port, and
who the consumer is when it connects — and it is the whole mechanism by which
anything in the mesh reaches anything else. The mesh asserted it and never found
out whether it was true. Issue 145 is what that cost: eleven hours of 'every module
current with its source' while a module could not reach its database.

The consumer verifies it, from its own network position. Not the control plane,
which reaches the address by a path no consumer uses. Not the machine, whose own
packets carried a source address the filter admitted while every container's was
refused — so a host-side check would have passed throughout the outage it exists to
catch. That is inference from the rule that was loaded, stated as such.

A connection and nothing more; speaking each provision's protocol would be a second
implementation of every provision. One failure is not news, only consecutive ones,
and the count is reported rather than the last attempt. A consumer that is not
running reads unchecked, which is a different sentence from broken.

What it costs to be wrong is the constraint on all of it: a check that calls a
working provision broken trains a reader to ignore the report.
2026-09-29 13:07:18 +02:00
mesh-admin f65664640a Merge pull request 'Issue 145: a machine reads healthy while its modules cannot reach each other' (#179) from issue/145-healthy-while-broken into main 2026-09-29 10:37:10 +00:00
54 changed files with 2476 additions and 41 deletions
+3 -1
View File
@@ -54,4 +54,6 @@ whose failure has never been observed is a guess about its own correctness.
The development cycle, checked ([ADR 0080](../../02-DECISIONS/0080-the-development-cycle-is-checked.md)): The development cycle, checked ([ADR 0080](../../02-DECISIONS/0080-the-development-cycle-is-checked.md)):
a to-be design names a decision, an in-progress/implemented design names its owning code, a a to-be design names a decision, an in-progress/implemented design names its owning code, a
located/fixed issue names its owner, a fixed/resolved issue says what fixed it, a graduated located/fixed issue names its owner, a fixed/resolved issue says what fixed it, a graduated
research overview says what it became. `python3 00-META/checks/cycle.py` research overview says what it became, and no two issue records share a number (issue 155 — the
number is how a record is cited, and `main` lags every open pull request, so two people reading it
allocate the same one). `python3 00-META/checks/cycle.py`
+23 -1
View File
@@ -14,7 +14,8 @@ What is enforced:
its owning code (`code:`) -- no development without a design that says where. its owning code (`code:`) -- no development without a design that says where.
issues a known `status:`; once `located`, `located-in:` names the owner; issues a known `status:`; once `located`, `located-in:` names the owner;
once `resolved`, `fixed-by:` says what fixed it (prose counts -- once `resolved`, `fixed-by:` says what fixed it (prose counts --
"nothing, the capability existed" is an answer). "nothing, the capability existed" is an answer). And no two records share a
number -- the number is how a record is cited.
research a known `status:`; a `graduated` overview says what it `became:`, and every research a known `status:`; a `graduated` overview says what it `became:`, and every
target it names exists. target it names exists.
decisions every accepted record is REACHABLE from the cycle: cited by a design doc's decisions every accepted record is REACHABLE from the cycle: cited by a design doc's
@@ -109,6 +110,27 @@ def main():
"without a design that says where" % status) "without a design that says where" % status)
# ---- issues ------------------------------------------------------------------------ # ---- issues ------------------------------------------------------------------------
# Two records may not share a number. Numbers are taken as "next free after main", and work
# sits on unmerged branches for days -- so two people reading the same main allocate the same
# number, and nothing said so. It happened twice in one evening between two machines, and the
# second collision landed on main with all three checks passing (issue 155). An issue number is
# how every other record cites this one; two records answering to it means a pointer that
# resolves to whichever the reader happened to open.
seen = {}
for folder in sorted(glob.glob(os.path.join(ROOT, "04-ISSUES", "*", ""))):
name = os.path.basename(os.path.normpath(folder))
number = name.split("-", 1)[0]
if not number.isdigit():
continue
if number in seen:
bad(os.path.join("04-ISSUES", name),
"is numbered %s, and so is %s -- an issue number is how it is cited, and two "
"records answering to one means a citation that resolves to whichever the reader "
"opened. Take the next free number across main AND every open pull request"
% (number, seen[number]))
else:
seen[number] = name
for path in sorted(glob.glob(os.path.join(ROOT, "04-ISSUES", "*", "00-report.md"))): for path in sorted(glob.glob(os.path.join(ROOT, "04-ISSUES", "*", "00-report.md"))):
front = frontmatter(path) front = frontmatter(path)
if front is None: if front is None:
+14 -1
View File
@@ -21,7 +21,12 @@ incident someone must **clear**.
## Steps ## Steps
1. Take the next free number. Create `04-ISSUES/NNN-short-name/00-report.md`: 1. Take the next free number — **across `main` and every open pull request**, not `main` alone.
Work sits on unmerged branches for days, so two people both reading `main` allocate the same
number; it happened twice in one hour between two machines, and the second collision reached
`main` with every check passing (issue 155). `cycle.py` now refuses two records sharing a number,
which catches a collision but does not prevent one. Create
`04-ISSUES/NNN-short-name/00-report.md`:
```yaml ```yaml
--- ---
@@ -42,6 +47,14 @@ incident someone must **clear**.
## Rules ## Rules
- Closed issues are never deleted — they are the mesh's symptom-to-component memory. - Closed issues are never deleted — they are the mesh's symptom-to-component memory.
- `fixed-by:` names something that will still exist: a commit or a pull request, never a branch. A
branch is deleted when it merges, so a branch name there is a pointer that resolves to nothing by
the time anybody follows it.
- A fix that turns out to have broken something else is written back into the record that asked for
it, pointing at the new issue. Somebody arriving at a record to learn why the code is the way it
is must not have to already know there was a sequel.
- Renumbering a collision happens once, in the branch that lands last. Renumbering a branch whose
author is still pushing only moves the race.
- An issue whose answer is a general lesson should also be written to the knowledge base, so - An issue whose answer is a general lesson should also be written to the knowledge base, so
the next person searching a symptom finds it. Both, not either. the next person searching a symptom finds it. Both, not either.
- `status: wontfix` is legitimate and requires a sentence saying why. - `status: wontfix` is legitimate and requires a sentence saying why.
+8
View File
@@ -13,6 +13,14 @@ decisions taken over three days; the reasoning is kept, the fragmentation is not
The environment a change is run against before it reaches real machines. The environment a change is run against before it reaches real machines.
> **Still the lab, no longer the test bed — 2026-09-30, by [ADR 0149](0149-the-live-mesh-is-the-test-bed.md).**
> Everything here stands. What changed is what the lab is *for*: a change is verified against the mesh
> that is running, because the faults that cost the most are faults of a mesh that already exists —
> bound consumers, containers made against an older roster, an adopted machine — and a bed is by
> construction a mesh that does not. Raising a mesh from bare is now the lab's whole job, which is the
> one thing the live mesh cannot be asked to do. 0149 also supersedes
> [ADR 0068](0068-the-lab-takes-requests.md), which extended this one and was never built.
## A node in the lab is a virtual machine ## A node in the lab is a virtual machine
It boots a stock Linux image, runs the real install, and becomes a node. **It is not a model of It boots a stock Linux image, runs the real install, and becomes a node. **It is not a model of
+17 -1
View File
@@ -1,6 +1,6 @@
--- ---
topic: building it topic: building it
status: proposed status: accepted
date: 2026-09-01 date: 2026-09-01
deciders: jochen deciders: jochen
reconstructed: false reconstructed: false
@@ -101,3 +101,19 @@ the digest down after building.
**Whether kind 4 deserves a module at all.** Thirty-five descriptions that say *install this and **Whether kind 4 deserves a module at all.** Thirty-five descriptions that say *install this and
write these files* may be better as one module with settings than as thirty-five modules. Left write these files* may be better as one module with settings than as thirty-five modules. Left
open deliberately; it is a question about the shape of the catalogue, not about whether to have one. open deliberately; it is a question about the shape of the catalogue, not about whether to have one.
## Accepted, 2026-09-29, against what was built
*Marked in a grooming pass: the mesh was built to this record and the record still said `proposed`.*
The proposal is the arrangement that exists. `mesh-catalog` holds descriptions of software we did
not write and the programs that provision it, and holds neither the mesh's own components nor an
application's own module. The mesh's list of modules is a table in the control plane, filled by
`module add`, and every module records the source it came from with the commit it was read at.
**One half is not built: `module check` as a command on the control plane's binary.** A manifest is
still validated by a test that reaches into the control plane's internals — which works for this
catalogue and gives nothing at all to somebody describing their own application in their own
repository, which this record says is the case that matters most. That is
[issue 148](../04-ISSUES/148-a-manifest-outside-this-catalogue-has-no-check/00-report.md).
@@ -26,6 +26,14 @@ runtime is per-module, not per-node, and treating the audit-logger as special le
modules' code with nothing to run it: the conversion produced tools and events that, as it stands, modules' code with nothing to run it: the conversion produced tools and events that, as it stands,
never execute. never execute.
> **The hosting form is settled elsewhere — 2026-09-30.** Where this record says "a container", read
> [ADR 0150](0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md): a module's own
> code runs as supervised processes under this record's one account. Nothing else here changes — the
> per-module runtime, the per-tool key and the single scoped account are the argument this record made
> and they are why 0150 goes the way it does. The note is here because two design documents chose the
> other form without knowing this record existed
> ([issue 117](../04-ISSUES/117-a-modules-own-code-is-a-container-and-a-process/00-report.md)).
## Decision ## Decision
### A module with tools or events runs a process of its own ### A module with tools or events runs a process of its own
@@ -80,6 +80,15 @@ reaching the routed name, which the clause above has just made resolvable inside
three are one decision: **compose the name, propagate it, certify it** — each is meaningless without three are one decision: **compose the name, propagate it, certify it** — each is meaningless without
the one before it. the one before it.
> **The mechanism changed — 2026-09-30, by [ADR 0148](0148-the-meshs-names-are-resolved-not-copied-into-containers.md).**
> A routed name still reaches every asker in the mesh, which is what this record decided and it stands.
> It no longer reaches them by being written into each declared container: copying the roster in made the
> roster part of every container's identity, so one name moving replaced every container in the mesh
> ([issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). A
> container resolves through its machine's resolver instead. The consequence below — that an internal
> issuer's challenge needs the routed name resolvable inside the mesh — holds unchanged, by the means the
> machine itself already uses.
## Consequences ## Consequences
- **Lab-versus-production is one node-level `public-domain` setting**, not an override on every - **Lab-versus-production is one node-level `public-domain` setting**, not an override on every
+8 -1
View File
@@ -1,14 +1,21 @@
--- ---
topic: building it topic: building it
status: proposed status: superseded
date: 2026-09-12 date: 2026-09-12
deciders: jochen deciders: jochen
reconstructed: false reconstructed: false
extends: 0016-the-lab.md extends: 0016-the-lab.md
superseded-by: 02-DECISIONS/0149-the-live-mesh-is-the-test-bed.md
--- ---
# 68. The lab takes requests, one at a time, and runs each from its own copy # 68. The lab takes requests, one at a time, and runs each from its own copy
> **Superseded — 2026-09-30, by [ADR 0149](0149-the-live-mesh-is-the-test-bed.md).** Never built. The
> live mesh became the test bed, because the faults that cost the most are faults of a mesh that
> already exists — bound consumers, containers made against an older roster, an adopted machine — and a
> bed is by construction a mesh that does not. The rule worth keeping from below is that a run reads a
> copy that is not anybody's working tree.
## Context ## Context
**The lab is exclusive hardware, and today a person holds it.** Raising a scenario takes over **The lab is exclusive hardware, and today a person holds it.** Raising a scenario takes over
@@ -35,6 +35,16 @@ The development cycle is enforced mechanically, to the extent frontmatter can ca
capability existed" is an answer). capability existed" is an answer).
- **No silent graduation** — a `graduated` research overview says what it `became:`, and the - **No silent graduation** — a `graduated` research overview says what it `became:`, and the
targets exist. targets exist.
- **No two records answering to one number** — added 2026-09-30; see the insight below.
> **Progressive insight — 2026-09-30.** The list above named four things `cycle.py` enforces, and
> now names five. Nothing enforced that two issue records hold different numbers: two machines
> filing issues within one hour both read `main`, both took "the next free number", and collided
> twice — the second collision reaching `main` with `records.py`, `cycle.py` and `index.py` all
> reporting success (04-ISSUES/155). A number is how every other record cites one, so two records
> answering to it is a citation that resolves to whichever folder the reader opened. `cycle.py`
> refuses it now. The decision here stands exactly as written: this is one more thing frontmatter
> and file names can carry, found by its absence rather than by reasoning.
[`00-META/checks/cycle.py`](../00-META/checks/cycle.py) refuses violations, beside `records.py` [`00-META/checks/cycle.py`](../00-META/checks/cycle.py) refuses violations, beside `records.py`
and `index.py`; all three run before any merge here. What frontmatter cannot see — that code work and `index.py`; all three run before any merge here. What frontmatter cannot see — that code work
@@ -1,6 +1,6 @@
--- ---
topic: what runs on it topic: what runs on it
status: proposed status: accepted
date: 2026-09-25 date: 2026-09-25
deciders: jochen deciders: jochen
reconstructed: false reconstructed: false
@@ -275,3 +275,11 @@ On acceptance, each of these is superseded or amended by this record, not edited
everything a module needs is a requirement everything a module needs is a requirement
- [Issue 095](../04-ISSUES/095-a-module-assigned-after-genesis-has-no-broker-account/00-report.md), - [Issue 095](../04-ISSUES/095-a-module-assigned-after-genesis-has-no-broker-account/00-report.md),
[issue 103](../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md): what fails today [issue 103](../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md): what fails today
## Accepted, 2026-09-29, against what was built
*Marked in a grooming pass.* The vault is a module providing `secret` at mesh scope, and six
modules in the catalogue require it — so a shared secret is a requirement answered by the vault,
which is what this record asks for. Private keys are still made where they are used and never
travel, which is the other half and was never in question.
@@ -1,6 +1,6 @@
--- ---
topic: what runs on it topic: what runs on it
status: proposed status: accepted
date: 2026-09-26 date: 2026-09-26
deciders: jochen deciders: jochen
reconstructed: false reconstructed: false
@@ -189,6 +189,23 @@ On acceptance, each of these is amended by this record, not edited:
credential is no longer all-or-nothing with a window. It overlaps, with each step confirmed. credential is no longer all-or-nothing with a window. It overlaps, with each step confirmed.
A single-party credential is staged, not replaced. A single-party credential is staged, not replaced.
## Accepted, 2026-09-30, and not scheduled
Accepted as written. The separation it draws — a consumer's *resource* and a *credential that reaches
it* are different things with different lifecycles — is the part that had to be settled, because the
alternative is what the record was written against: retiring a credential taking the data it reached
with it. That is a data-loss shape, and a record that names it should not sit unresolved while the
code that could hit it is being written.
**It is not built, and accepting it does not schedule it.** The SDK's provisioner adapter is still
`create` / `remove` / `holds` rather than the four operations above, and no provider implements the
two-credential rotation. Accepted-and-not-built is an ordinary state here — 0141 and 0142 are both in
it — and it is the honest one: leaving this `proposed` made it invisible to anyone reading what the
mesh has decided, while changing nothing about what runs.
The work it implies belongs with the provisioner contract, beside
[issue 124](../04-ISSUES/124-a-consumer-cannot-be-told-what-its-provider-derived/00-report.md).
## Consequences ## Consequences
- **Every credential provider's adapter changes**, in two steps. The first separates *retire a - **Every credential provider's adapter changes**, in two steps. The first separates *retire a
@@ -1,6 +1,6 @@
--- ---
topic: what runs on it topic: what runs on it
status: proposed status: accepted
date: 2026-09-26 date: 2026-09-26
deciders: jochen deciders: jochen
extends: 0112-a-module-definition-names-no-node-mesh-or-path.md extends: 0112-a-module-definition-names-no-node-mesh-or-path.md
@@ -47,3 +47,10 @@ other boundary already is: the module name.
node runs one of each (ADR 0115)" — instead of failing on whichever name collides first. node runs one of each (ADR 0115)" — instead of failing on whichever name collides first.
- Multi-tenant asks are answered in the catalogue (a second module definition), not in the - Multi-tenant asks are answered in the catalogue (a second module definition), not in the
control plane. control plane.
## Accepted, 2026-09-29, against what was built
*Marked in a grooming pass.* The rule is enforced where it cannot be forgotten: `assignment`'s
primary key is `(node, module)`, so a second assignment of one module to one machine is not a thing
the mesh can hold. The record read `proposed` while the schema had already settled it.
@@ -12,7 +12,7 @@ extends: 0102-the-mesh-writes-into-a-shared-file-never-over-it.md
## Context ## Context
When a resource stops being declared — its module unassigned, the node sent a When a resource stops being declared — its module unassigned, the node sent a
deliberately-empty declaration ([issue 127](../04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)), deliberately-empty declaration ([issue 149](../04-ISSUES/149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)),
or a new catalogue version renaming its id — the host undoes it. The host's own code states or a new catalogue version renaming its id — the host undoes it. The host's own code states
the rule it means to follow: **it removes what it made and leaves what it merely configured.** the rule it means to follow: **it removes what it made and leaves what it merely configured.**
For almost every resource it does exactly that: For almost every resource it does exactly that:
@@ -0,0 +1,135 @@
---
topic: what runs on it
status: superseded
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0010-delivery.md
superseded-by: 02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md
---
# 143. A consumer verifies the grant it is given
## Context
A **grant** is what the mesh writes on a consumer's machine so it can reach a provider. The real one
the forge receives for its database, as it arrives:
```
provision postgres-database
at <the provider's machine, by name>
port the machine port the provider is published on
as the role the provider created for this consumer
```
with the credential sealed in a separate file. Four facts and a password, and they are the whole
mechanism by which anything in the mesh reaches anything else.
**The mesh asserts that claim and never finds out whether it is true.**
[Issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md):
converging a machine dropped the path from a container to a port on its own machine, and for eleven
hours the mesh answered *all doing what they were told, all heard from, every module current with its
source* while a web application logged, six thousand times:
```
connection to server at "<the machine>" (10.10.0.1), port 6852 failed: timeout expired
```
Every check the mesh makes passed, because every check it makes is about the relationship between the
mesh and a machine: the declaration was applied, the digest matched, every container named was running.
None of them asks whether a consumer can reach what it requires — though the mesh composed the grant
and therefore knows the consumer, the machine, the address, the port and the credential.
**And where the check runs decides whether it catches anything.** The rule in force admitted the
machines' own addresses on the private network. A dial from the *machine* to its own address carries
exactly such a source address, so a check run by the host on its own behalf would have matched that rule
and passed — while every container on the machine was refused. This is inference from the rule that was
loaded, not a measurement: the fault was found and fixed before anyone thought to dial from the host.
It is enough to decide the question, because a check whose position differs from the consumer's is
testing something nobody asked about.
## Considered Options
1. **The control plane dials each provision.** Rejected, and it is the tempting one because the control
plane holds every fact. It sits on the provider's machine for most provisions here and reaches the
address by a path no consumer uses; in the measured outage it would have passed throughout.
2. **The host dials on the consumer's behalf, from the machine.** Rejected for the reason above: the
machine's network position is not the consumer's, and the one outage this exists to catch is exactly
a difference between them.
3. **Ask the module.** Rejected: a module is arbitrary software that the mesh does not write. Some could
report on their provisions and most cannot, and a check that covers the modules that opted in tells
nobody anything about the rest.
4. **Read the module's logs.** Rejected: the failure was in a log the whole time, and reading a module's
logs makes the mesh depend on the wording of software it does not control.
5. **The consumer verifies it, from its own network position.** Adopted.
## Decision
**A consumer verifies each grant it is given, from its own network position.** After a reconcile has
applied a grant, the machine opens a connection to the address and port that grant names, from inside
the consumer's own network namespace — the same position the consumer's software dials from, which is
the only position that answers the question the grant asks.
**It is a connection, not a conversation.** Whether the port accepts a connection is what a grant
claims; whether the credential is right, the role exists or the schema is current is the provider's to
answer and the consumer's to discover. A check that spoke each provision's protocol would be a second
implementation of every provision, and would fail for reasons that are not the mesh's.
**One failure is not news.** A provider restarting is ordinary, and so is a consumer between containers.
A grant is reported unreachable only after it has failed on **consecutive** reconciles, and the count is
what the machine reports rather than the last attempt — so a reader can tell "it was briefly away" from
"it has never worked".
**A grant that cannot be checked is said to be unchecked, never assumed good.** A consumer that is not
running has no network position to dial from; that is not a broken grant and must not read as one. It is
also not a verified grant, and the two are different sentences.
**What it costs to be wrong is the constraint on all of it.** A check that reports a working provision
broken trains a reader to ignore the report, which is worse than having none — the fault this
repository keeps finding, one level up. So the threshold is consecutive failures, the check is the
cheapest thing that answers the question, and an unknown is reported as unknown.
**The mesh says it where it says everything else.** A machine's report carries its unreachable grants,
and `status` names them beside what is out of date — so "every module current with its source" stops
being the whole of what the mesh will tell you about a machine whose modules cannot reach each other.
## Consequences
- **The mesh can be wrong out loud.** It has been able to assert a grant and not check it; now a grant
that does not work is a thing the mesh says, and the eleven hours of issue 145 become minutes.
- **The host gains the ability to act from a container's network position**, which it has not needed
before. That is a real capability and the only one this needs.
- **A machine reports something that is not about the declaration.** Everything it reports today is
what it applied and what it holds; this is the first thing it says about whether what it applied
works.
- **A provision with no port is not checked**, because there is nothing to dial. Several are files and
secrets, and saying "checked" about those would be the appearance of verification that this record
exists to remove.
- **What got harder:** a reconcile does more than apply. Every grant adds a connection attempt on a
cadence, which is cheap individually and worth naming: a machine with many consumers dials once per
grant per reconcile.
## How it is checked
- **The outage is caught.** A bed drops the path from a consumer's network position to a provider's
port while leaving the machine's own path to it open — the exact shape of issue 145 — and the grant
reads unreachable. This fails against the previous behaviour, where nothing reported anything, and
against a check run from the machine, which passes while the consumer cannot reach it.
- **A restarting provider is not an outage.** One failed reconcile reports nothing; the count rises and
falls, and the grant reads reachable again without anybody acting.
- **A consumer that is not running reads unchecked, not broken**, asserted separately from the
unreachable case because they are different sentences.
- **A provision with no port is not claimed to be checked.**
- **The report carries the count, not the last attempt**, so "briefly away" and "never worked" are
distinguishable by a reader who sees only the report.
- **`status` names an unreachable grant**, asserted on the output, since a check nothing surfaces is
the same as no check.
## References
- [ADR 0010](0010-delivery.md) — the declaration is owned resources; a grant is one of them
- [ADR 0009](0009-modules-and-the-graph.md) — what a provision and a consumer are
- [issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
— the eleven hours
- [issue 136](../04-ISSUES/136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md) — the
same distance between a declaration and a machine, one level down
@@ -0,0 +1,121 @@
---
topic: what runs on it
status: accepted
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md
supersedes: 02-DECISIONS/0143-a-consumer-verifies-the-grant-it-is-given.md
---
# 144. Anything on a machine may call anything on it, and that is the whole of "local"
## Context
Everything in the mesh should be able to call:
- what runs on the same machine;
- another machine's service over the private network, if that service is exposed there;
- another machine's service over the public network, if it is exposed there.
Three cases. The filter had two of them.
**The first was broken and the break was invisible.** A service exposed to the private network rendered
as the machines' own addresses on it. A caller on the machine carries such an address; a caller inside
one of that machine's containers carries a bridge address and matched nothing. Measured:
```
the machine: local 10.10.0.1 dev lo src 10.10.0.1
a container: 10.10.0.1 via 172.17.0.1 dev eth0 src 172.17.0.8
```
Same destination, same machine, two source addresses. The rule named the first and silently refused the
second, so a module reaching its database on its own machine's name timed out for eleven hours
([issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)).
**The second case works, and by accident.** A caller on another machine reaches the private network over
the tunnel, and arrives carrying that machine's own address — so the rule matches. It would not have
matched the caller's own address either; the tunnel rewrites it. That two of three cases worked is why
this looked correct.
**[ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) answered the wrong question.** Written
hours earlier, it proposed that a consumer verify each grant it is given by opening a connection from
its own network position — and it went to some length about *which* position, because whether a caller
sat in a container changed the answer. That difference was the bug. A verification mechanism would have
reported this outage sooner and would not have prevented it, and the machinery it needed existed only
because the rule was wrong. The remedy for a configuration error is the correct configuration.
**And a module is not a container.** A module is software that delivers one or more services, and it may
do that as a container, an installed package with a unit, a binary, or files something else reads. Of 72
modules in the catalogue, 61 happen to use a container and 11 do not — among them the resolver, the ssh
daemon and the intrusion-prevention module. A rule that reasons about containers describes most of the
mesh and not the mesh.
## Considered Options
1. **A line per service admitting the machine's own callers.** Rejected: it is what was written first,
and it only ever covers the services somebody remembered to think about. It also states, service by
service, a thing that is true of the machine.
2. **Verify each grant from the consumer's position** ([ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md)).
Rejected as a remedy: it observes the fault rather than removing it, and the question it agonised over
— which network position — exists only while the fault does.
3. **Enumerate the addresses a machine's callers may have.** Rejected for the reason no address is named
anywhere in this filter any more ([ADR 0140](0140-the-filter-constrains-what-arrives-from-outside.md)):
a range describes one machine and goes stale in silence.
4. **Local is not filtered, stated once.** Adopted.
## Decision
**Anything on a machine may call anything on that machine, and the filter says so once.** Not per
service, not per port, and not by naming who the callers are: traffic that did not arrive from outside
the machine and did not arrive over the private network is the machine's own, and is admitted. It is
asked by the link the traffic arrived on, because that is a fact about the machine rather than a list
that describes one.
**Local is not a boundary this mesh draws.** Whether a caller is a container, a unit, or the operator's
shell changes nothing, because the thing being decided is "is this the same machine" and the answer does
not depend on the form the caller takes.
**The other two cases are unchanged and are now legible beside it.** A service exposed to the private
network admits the machines on it; a service exposed publicly admits anything. Three cases, three lines,
and a reader can see all three at once.
**[ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) is superseded and nothing replaces it.**
Whether the mesh should check that a grant works is a real question — it reported this machine healthy
for eleven hours — but it is a question about what the mesh can say, not about what it should do, and it
must stand on its own rather than as the remedy for a rule that was wrong. It is not built.
## Consequences
- **The three things everything should be able to call are three lines**, and the first is one line
rather than one per service, so a service added tomorrow is reachable locally without anybody
remembering to say so.
- **A form of module stops mattering to the filter.** The 11 modules that are not containers were never
affected by this bug and were never the reason it was hard to see; they are the reason the rule should
never have mentioned containers.
- **The mesh still cannot say when a grant stops working.** That is the live gap, recorded in issue 145
and no longer pretending to have an answer.
- **What got harder:** nothing. This removes a line per service and replaces it with one.
## How it is checked
- **A caller on the machine reaches a service on it, in the input chain**, asserted on that chain's own
body — because the forward chain carries the same line in the same words, and an assertion on the
whole rendered file passed with the input chain's copy deleted. That is what
[ADR 0137](0137-a-machine-says-which-networks-it-routes.md)'s tests already say to do.
- **It is one rule, not one per service.** Asserted by rendering two services of different reach and
refusing a per-port local line.
- **The three reaches render as three lines**, asserted together, so the whole of what the filter says
about who may call what is one test.
- **The measured case:** from a container on the machine, a service exposed to the private network on
that machine answers. This is the outage, and it fails against the rule this replaces.
## References
- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the filter is the sum
of what its modules listen on
- [ADR 0140](0140-the-filter-constrains-what-arrives-from-outside.md) — why no address is named
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — internal and public,
the other two cases
- [ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) — superseded here
- [issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
@@ -0,0 +1,120 @@
---
topic: what runs on it
status: superseded
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md
superseded-by: 02-DECISIONS/0146-connectivity-is-checked-by-name-per-hosting-form.md
---
# 145. A module checks what the mesh claims is reachable, and it checks itself
## Context
The mesh asserts three things are callable ([ADR 0144](0144-anything-on-a-machine-may-call-anything-on-it.md)):
what runs on the same machine, another machine's service exposed to the private network, and another
machine's service exposed publicly. It has never checked any of them.
[Issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md):
the first of the three was broken for eleven hours and the mesh answered *all heard from, every module
current with its source* throughout. Every check it makes is about the relationship between the mesh and
a machine — applied, current, containers running — and none about whether anything can reach anything.
**A first answer was drafted and withdrawn.** [ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md)
put the check inside the host, verifying each grant from the consumer's network position. It was
superseded because the difference it worked so hard to reproduce — whether a caller sat in a container —
was the bug itself. What survives from it is the part that was right: a check run from the wrong place
proves nothing, and the mesh's own reports are not evidence about the network.
**The mesh already has the shape for this and it is a module.** A module can declare a container that
runs on a cadence ([ADR 0053](0053-a-step-that-runs-on-a-schedule.md), and three modules already use
`*/5 * * * *`), can be given the mesh's roster as a rendered fact — every machine's name, address and
this node's own identity, the same mechanism the resolver and the operator's ssh configuration use — and
can emit what it found on the bus. Nothing new is needed to build this except the module.
**What it must not check is the trap.** The obvious probe target is ssh: present on every machine, never
closed by design. Dialling it would have passed throughout the outage, because ssh is admitted
unconditionally and the thing that broke was a service exposed to the private network. A checker whose
probe is unconditionally open measures the one path that cannot fail, which is the failure this whole
sequence keeps producing — a check that reads as verification and verifies nothing.
## Considered Options
1. **The host verifies each grant** ([ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md)).
Superseded. It needed the host to act from another network position, which is machinery that exists
only while local calls are filtered wrongly.
2. **The control plane dials every node.** Rejected: it sits on one machine and reaches the others by a
path no ordinary caller uses. It would have passed throughout the outage.
3. **Probe an existing service.** Rejected for the target problem above: the services guaranteed on every
machine are the ones that are never closed, so they cannot fail the way the mesh fails.
4. **A module on every machine that serves its own probe and dials the others'.** Adopted.
## Decision
**A module runs on every machine, serves an endpoint of its own, and dials every other machine's.** The
probe is the module's own endpoint, declared reachable over the private network — so the thing being
dialled is admitted by exactly the rule that governs every other internally-exposed service, and fails
when that rule is wrong. A second endpoint, declared public, does the same for the public path where a
machine has one.
**It checks the three cases the mesh claims, by name:**
- its **own machine**, by dialling its own machine's address — the case that broke, and the only one that
distinguishes a caller on the machine from a caller in one of its containers;
- **each other machine over the private network**;
- **each machine's public path**, where one is recorded.
**It resolves before it dials, and says which failed.** A name that does not resolve and a port that does
not answer are different faults with different owners, and a checker that reports one sentence for both
sends a reader to the wrong place.
**It runs where the callers run.** The module's own code in its own container, on the cadence the mesh
already has, from the same position as every other module on that machine. It is not the host and not the
control plane, and that is the whole point.
**It says what it found and nothing else.** It emits results; it repairs nothing, opens nothing and holds
no credential beyond its own. A checker that fixes things is a second control plane.
**One failure is not a fault.** A machine rebooting is ordinary. A path is reported broken after it has
failed on consecutive runs, and the count travels with the result so a reader can tell "briefly away"
from "never worked" — the one thing [ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) got
right and worth keeping.
## Consequences
- **The mesh gains the ability to be wrong out loud about the network.** Eleven hours becomes two runs.
- **It is a module, so it is assigned, built, pushed and reported on like everything else** — no new host
capability, no new vocabulary, nothing in the control plane that has to know about checking.
- **Its own endpoint is the instrument.** That is what makes it able to fail; it also means the checker
must be assigned to a machine before that machine can be checked, and a machine without it is
unchecked rather than healthy.
- **It cannot check what it cannot be told.** The roster gives it machines; it does not give it every
module's endpoints, so this checks the paths the mesh claims and not every grant in the mesh. That is
the honest scope of a first one, and the difference is worth saying rather than growing quietly.
- **What got harder:** one more module on every machine, and a module whose whole purpose is to fail
visibly when something else is wrong. Its own failures will be read as the mesh's, which is the cost of
an instrument.
## How it is checked
- **It catches the measured outage.** A bed closes the path from a container to a service exposed to the
private network on its own machine — issue 145's shape — and the checker reports its own machine
unreachable while every other path still reads reachable. This fails against a probe on a port that is
never closed, which is the wrong target this record exists to name.
- **A machine rebooting is not a fault**: one failed run reports nothing, the count rises and falls.
- **A name that does not resolve is reported as that**, not as a port that did not answer.
- **It reports and does not act**: asserted by giving it a broken path and checking nothing on the machine
changed.
- **A machine without the module reads unchecked**, never healthy — asserted on what the mesh says about
a machine it is not assigned to.
## References
- [ADR 0144](0144-anything-on-a-machine-may-call-anything-on-it.md) — the three things that must be callable
- [ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) — superseded; what survives is that the
position matters
- [ADR 0053](0053-a-step-that-runs-on-a-schedule.md) — the cadence
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — internal and public,
which the probe endpoints declare
- [issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
@@ -0,0 +1,125 @@
---
topic: what runs on it
status: accepted
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0145-a-module-checks-what-the-mesh-claims-is-reachable.md
supersedes: 02-DECISIONS/0145-a-module-checks-what-the-mesh-claims-is-reachable.md
---
# 146. Connectivity is checked by name, per hosting form, with a valid certificate
## Context
[ADR 0145](0145-a-module-checks-what-the-mesh-claims-is-reachable.md) decided that a module checks what
the mesh claims is reachable, from where the callers are, because the mesh reported four machines healthy
for eleven hours while a module could not reach its database
([issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)).
That decision stands. What it got wrong is everything about *what* is dialled.
It dialled a raw port on each machine's address. Three things are wrong with that:
- **A raw port is not how anything in this mesh is reached.** A real caller resolves a name, the proxy
answers it, and the proxy reaches the service. A check that dials a port tests the last hop of a path
with four hops in it, and the three it skips — resolution, the proxy, the certificate — are where most
of the mesh's connectivity actually lives.
- **It tested one hosting form.** A module is software that delivers services, and it may deliver them
from a container, from a unit the mesh writes for its own code, or from a unit a package ships. Those
are three different paths to the same machine, and the outage that produced this was two of them
disagreeing. A probe served one way measures one way.
- **It said nothing about certificates.** An internal name that resolves, routes and answers over TLS
that nothing can verify is not a working path; it is a working path for whoever holds the proxy's
trust and nobody else.
## Decision
**Each hosting form gets its own endpoint, its own route and therefore its own name.** On every machine:
| name | what serves it |
|---|---|
| `connect-docker.<node>.internal` | a container |
| `connect-process.<node>.internal` | the mesh's own code, in a unit the mesh writes |
| `connect-unit.<node>.internal` | a unit a package ships |
and the same set under each machine's public domain where it has one — `connect-docker.<domain>` and its
siblings. The names are the instrument: a failure reads as *`connect-docker.g14.internal` did not answer*,
which says which machine and which hosting form without anybody interpreting anything.
**Every machine checks every machine, by name, over TLS, verifying the certificate.** Not a port, not an
address: resolve the name, connect, complete the handshake, check the certificate against the authority
that should have issued it — the mesh's own for an internal name, a public one for a public name. That is
the whole path a real caller takes, and each step failing is reported as itself.
**No name is written anywhere.** The machines come from the roster the mesh already renders as a fact, and
the labels are the module's. A machine that joins appears in every other machine's roster on the next
push, and they begin checking it without an edit.
**And the module arrives on a machine because the machine exists, not because somebody assigned it.** A
machine that joins and does not have it is worse than unchecked: every other machine is already dialling
its names, so it reads as broken everywhere until someone notices. This is the part the mesh cannot
currently express — see below — and it is the part that makes the rest safe.
**What survives from 0145**, unchanged: it reports and repairs nothing; one failure is not a fault and a
path is broken after consecutive runs with the count travelling with the result; findings are said on the
bus, because a finding in a file on the machine is what this exists to end; and the bus is the one path
that cannot report its own failure, so an emit that does not land is written locally and nowhere else.
## What this needs that the mesh does not have
Named here rather than assumed, because each is a decision of its own and this record is not the place to
make them:
1. **A module that every machine has.** `ScopeNode` means *at most one holder per node* — an exclusivity
rule, not an obligation — and nothing assigns a module at enrolment. Today the resolver, the packet
filter, ssh and intrusion prevention are each assigned per machine by hand, which is the same gap
wearing different clothes.
2. **A container running a module's own bundle.** A `process` runs the mesh's own compiled code with no
image; a `container` needs an image of the module's own, which means a Dockerfile — the thing the
`bundle` artifact exists to abolish. Nothing in the catalogue runs a bundle in a container, so
`connect-docker` has no shape yet.
3. **A unit a package ships, for `connect-unit`.** The `service` resource puts an existing unit into a
state and deliberately installs none, so this form needs a package that serves a port — and naming a
program the machine may not have is
[issue 136](../04-ISSUES/136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md).
4. **A machine's public domain in the roster fact.** The fact carries each machine's name, mesh name,
address and operator account. The public names cannot be composed without the domain.
5. **Something that installs the mesh's own root on a machine.** This is
[issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md), open
since before any of this. Until it is closed, every internal name will fail certificate verification
from every machine — correctly, because nothing can verify it. That is the checker working, and it is
worth saying in advance so the first run is not read as the checker being broken.
## Consequences
- **A failure names the machine and the hosting form.** That is the whole gain over a port: eleven hours
became two runs under 0145, and under this it also becomes one line that says where to look.
- **The checker surfaces issue 129 immediately**, and will report every internal name unverifiable until
it is fixed. A reader must be told that before the first run rather than after.
- **Five things must be built before this is what it says it is**, and until they are, what exists is a
port dial from one position — useful, and not this.
- **What got harder:** a module with three hosting forms of the same trivial service is a strange thing to
read. It is justified only because those three forms are how the mesh actually runs software, and a
checker that tested one of them would keep the class of outage it exists to catch.
## How it is checked
- **A name per hosting form answers from every machine**, asserted by name and not by port.
- **A certificate that does not verify is reported as that**, distinctly from a name that does not resolve
and a port that does not answer — three faults, three owners.
- **A machine that joins is checked by every other machine without an edit**, asserted by adding one to a
bed and looking at what the others dial on their next run.
- **A machine that joins has the module**, which is gap 1 above and is the assertion that cannot be
written yet.
- **The measured outage is still caught**: the path from a container to a service on its own machine is
closed and `connect-docker.<that node>.internal` fails from that machine while the others still pass.
## References
- [ADR 0145](0145-a-module-checks-what-the-mesh-claims-is-reachable.md) — superseded; its core stands
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — the two reaches these
names come from
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — a label plus a domain, which is why no name is written
- [issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md) — what the
internal names will fail on until it is closed
- [issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
@@ -0,0 +1,139 @@
---
topic: what runs on it
status: accepted
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md
---
# 147. A module anchors the mesh's authority on a machine, and takes it away again
## Context
The mesh runs its own certificate authority and every internal name is served with a certificate
from it. No machine trusts it. On an enrolled, adopted workstation — on the private network,
resolving through the mesh's resolver — every internal HTTPS name fails verification with
*unable to get local issuer certificate*
([issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md)).
The certificates are genuine; nothing on the machine has ever been told what issued them.
The authority's only consumer today is a proxy, which fetches the root into a directory of its own
and hands it to one program ([ADR 0098](0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md)).
That is enough for the proxy and for nothing else: a browser, `git` over HTTPS, `curl`, a package
manager and every module that calls another module by an internal name read the machine's trust
store, which holds the predecessor's authority and a developer tool's local root, and nothing of
the mesh's.
The predecessor wrote its root into every machine it set up. Removing it was deliberate — an
honest failure beats a name that verifies for the wrong reason — and it leaves the mesh with no
answer at all until this one lands. It is also what keeps the predecessor alive on the machines
that still speak TLS to a mesh name.
**What makes this a decision rather than a patch** is where the knowledge goes. Two mechanisms in
the mesh already write things onto a machine because it is on the private network: `/etc/hosts`
and the registry's plaintext trust ([ADR 0082](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)).
Following that precedent, the controller would inject an anchor into every such machine's
declaration, and issue 129 proposed exactly that. It would work. It would also put *where this
operating system keeps trust anchors* and *which command refreshes its bundles* into the control
plane, for a fact the control plane does not have (the root does not exist until the authority has
run) and a machine that may have no reason to verify a mesh name at all.
## Considered Options
1. **The controller injects the anchor into every machine on the private network**, the
`/etc/hosts` and insecure-registry shape. Rejected: being on the network is what makes the
registry reachable, and that is why network presence is the right trigger *there* — the trust
and the reachability are the same fact. Trusting an authority is not the same fact as being
able to reach it, and the anchor's path and the bundle refresh are a property of the machine's
operating system, which is the host's half of the mesh, not the controller's.
2. **A new host primitive — a `trust-anchor` resource type.** Rejected for now, not on principle.
The host's vocabulary should grow when a shape cannot be said with what exists, and this one
can: a file and a service already express it, as the packet filter proves
([ADR 0140](0140-the-filter-constrains-what-arrives-from-outside.md), whose module writes a
unit file and a service and nothing else). The primitive becomes right the moment a second
operating system is in play, because the anchor directory and the refresh command are exactly
the difference `internal/system` exists to hold. Until then it would be a vocabulary word with
one speaker.
3. **A module that requires the authority, fetches its root, installs it as an anchor and
refreshes the machine's bundles — and removes both when it is no longer assigned.** Adopted.
## Decision
**A machine trusts the mesh's authority because a module put its root there, and stops trusting it
when that module is taken away.**
1. **The module requires `internal-acme-ca`** and reads the provider's bound address and the path
it serves its root at. It requires nothing else and provides nothing: it is a consumer of the
authority like any other.
2. **It fetches the root over the mesh's own network, without prior trust**, because there is no
prior trust to have — this is the module that establishes it — and the network is what
authenticates the fetch ([ADR 0098](0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md),
the same reasoning that lets the proxy fetch it). What it accepts is checked: a body that is
not a certificate fails, and the failure is the module's, not a later handshake's.
3. **It installs the root where this machine's TLS clients look, and refreshes the extracted
bundles** — the command that does the refresh is an ordinary part of the unit that places the
anchor, not a new thing the mesh can be asked to do.
4. **Removal is symmetric and is the same unit's business.** Undeclared, the host stops the unit;
stopping it removes the anchor and refreshes the bundles again. A machine that leaves the mesh
stops trusting the mesh, without anybody remembering to go and look.
5. **It is an ordinary assignment.** No machine is given it automatically. A machine that verifies
a mesh name is assigned it, and a machine that does not is not — which is the same statement
the mesh already makes about every other module, and is why this is not the controller's
business.
**One operating system, said out loud.** The anchor directory and the refresh command in the
module today are Arch's. On a machine that is not Arch the unit fails, visibly, rather than
writing a file nothing reads. That is the accurate failure, and it is the signal that option 2
above has become right.
## How this is checked
- **The verification that could not succeed before.** On a machine holding the module, a plain
client fetches an internal HTTPS name with no `-k` and no bundle argument and verifies. On a
machine without it, the same fetch fails with *unable to get local issuer certificate*. Both
halves, because only the pair distinguishes "the anchor works" from "something else already
trusted it".
- **The removal half, in the same bed:** unassign the module, refetch, and the failure returns.
Checking only the arrival is how a trust store fills up with authorities nobody can account for.
- **What is deliberately not checked here:** that the authority issues, that a name resolves, that
the proxy serves. Those have their own beds, and this module's bed passing for those reasons is
the failure mode this record is most exposed to — which is why the negative half is not optional.
**What this bed is dialled at, and why it is the authority itself.** The authority serves its own
API with a certificate it issued, so the handshake under test needs nothing else in the mesh to be
right. A trust bed that reached for a routed name through the proxy would be passing or failing for
the proxy's reasons and the resolver's.
**Written, and not yet run** *(2026-09-29)*. The bed is `trust-anchor` in the lab, and it cannot
execute: raising a foundation fails before any module is reached, in both bundles that exist
([issue 146](../04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md)).
So what stands behind this record today is the rendering — the script the machine would run names
the authority it was bound to, checked in the control plane's own test suite — and **not** a machine
that verified anything. That is a weaker thing than the paragraph above describes, and it stays
written this way until the bed runs.
## Consequences
The predecessor's authority can be retired from a machine once this module is assigned to it,
which is the first time that has been true. `git` over HTTPS to the mesh's forge starts working,
so the ssh-only clone URL stops being a rule. A module on any machine can call another module's
internal name and verify it.
What got harder: one more module to assign to a machine that needs it, and the machine's trust
store now changes when an assignment changes — which is the point, and is also a thing an operator
can be surprised by. The fetch without prior trust is the same exposure ADR 0098 accepted, now on
every machine that holds the module rather than only where a proxy runs: anything that can stand
in the middle of the mesh's own network at the moment of the fetch can be believed. The mesh
already treats that network as the thing it authenticates.
## References
- [issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md) —
the symptom and the evidence.
- [ADR 0098](0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md) — a fact made at
first start is fetched from its provider; this extends it from one program to the machine.
- [ADR 0082](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md) — the precedent
this deliberately does not follow, and why it is right where it is.
- [ADR 0005](0005-the-node-host.md) — the host is where one operating system's difference lives.
- [`03-DESIGN/01-to-be/08-connectivity.md`](../03-DESIGN/01-to-be/08-connectivity.md).
@@ -0,0 +1,159 @@
---
topic: the tiers
status: accepted
date: 2026-09-30
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0066-public-routing-is-name-agnostic.md
---
# 148. The mesh's names are resolved, not copied into every container
## Context
The mesh gives every container it declares the whole roster of mesh names as entries written into
the container's own hosts file at creation
([design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md),
[ADR 0066](0066-public-routing-is-name-agnostic.md)). A container takes those entries once and never
looks again.
Three issues are the same fact arriving three times.
**A container keeps the address it was made with.** Adopting the predecessor's tunnel moved the hub's
private address; the declaration followed it within one push and nothing on the machine did. The
forge's container held the old address, lost its database, reported healthy while its existing
connections lasted, and then the public name went down
([issue 109](../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md)).
**The same fault, four days later, undetected for five days.** One container had restarted 2286 times
against a database it could no longer find, while the mesh reported the machine as doing what it was
told. Inside it, `novox.internal` was an address that had not existed for five days
([issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md)). Forty-eight
other containers were current, none of them corrected — each had been recreated for some other
reason and picked up the roster on the way.
135 was fixed by putting the roster into the digest the host compares a container against, so a
container whose names moved is recreated like one whose image moved. **That made the roster part of
every container's identity**, which is the third arrival:
**One name moving replaces every container in the mesh.** Migrating one small module on one machine
took four routine actions; each changed the roster, and each replaced every container on the control
node — its own store, the registry, the edge proxy, the forge, the directory, mail. The control plane
was unreachable twice while its own store came back through crash recovery
([issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). None of
the replaced containers had anything to do with the module being migrated, or with its machine.
The blast radius of a name is now every container that carries the list, which is all of them. The
node-by-node migration ahead adds names one module at a time — on one machine alone that is around
twenty-five — and each would be a full restart of every service on the hub.
## Considered Options
**1. Keep the roster in every container and accept the churn.** Rejected. It is not a cost that can
be paid down: the mesh gets more names as it grows, and every name costs a restart of everything.
A rollback costs another.
**2. Scope each container's entries to the names it actually binds.** A container is given the names
of the things it declared a requirement on, so a name's blast radius is its consumers. Tidy, needs no
new mechanism, and keeps 135's guarantee exactly.
Rejected, and this is the close one. It contradicts the standing intent that **anything on the mesh
can call anything on it** — three cases, same machine, the private network, the public network, and
no fourth. Scoping resolution to declared couplings makes a name reachable only where the mesh was
told in advance that it would be wanted, and a person debugging inside a container would find names
missing that exist everywhere else on the machine. It also leaves the roster in the digest, so the
churn returns the moment a widely-bound name moves — smaller, not gone.
**3. Resolve at lookup time through the machine's resolver, and copy nothing.** Chosen.
## Decision
**A container resolves the mesh's names through its machine's resolver, at the moment it asks. No
mesh name and no mesh address is written into a container, and none is part of a container's
identity.**
The three consequences that make this worth doing:
- **Staleness stops being possible**, rather than being detected. 109 and 135 are not bugs that were
fixed; they are a shape that no longer exists. A name that moves is answered differently by the next
lookup, in every container, with nothing recreated and nothing restarted.
- **A name's blast radius becomes nothing.** Assigning a module on one machine does not touch a
container on another.
- **Anything can still call anything**, which option 2 gave up. The resolver answers every mesh name to
every asker on the machine, exactly as it answers the machine itself.
**The resolver is a machine-level process, not a container** — one of the modules that is not a
container at all — so a container depending on it is not the circularity it would be if the mesh's
own store had to resolve a name through something the store's own runtime had to start first.
**What a module declares for itself is untouched.** Entries a manifest asks for are the module's own,
stay in the container, and stay in its identity: they are part of what the module *is*, they do not
move when the mesh's roster does, and the mesh does not know what they mean.
**The machine's own roster file is untouched.** It is a file, rewritten in place, read by processes and
people; nothing restarts when it changes. It is only the *copy into each container* that this ends.
### The order this lands in, which is not a preference
**Nothing may stop copying names until resolution works from a container.** Removing the copy first
reintroduces 109 and 135 — silently, and on a live mesh, which is exactly how both were found.
1. **A container on any network can reach the resolver.** Today a container on the runtime's default
network asks from an address the converged filter drops, so it has no DNS at all
([issue 110](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md));
and on two of four machines the resolver binds loopback only, so the runtime hands containers a
public resolver instead. Both are prerequisites, not related work.
2. **The runtime is told which resolver to use, per machine, as a file** — not per container as a
creation-time argument, or the resolver's address is back in every container's identity and the
problem has only got smaller.
3. **Then, and only then, the roster leaves the declaration and the digest.**
Until step 3 the mesh keeps copying, and keeps comparing. 151 stays open until step 3 lands; it is not
closed by this record, only answered by it.
## How this is checked
- **A container resolves a name that moved, without being recreated.** Move a name the mesh serves;
from a container that was running before the move and has not been touched since, the name answers
with the new address. This is the one 109 and 135 would both have failed.
- **A name's blast radius is nothing.** Add a routed name on one machine; no container on any other
machine is recreated. The apply report on each machine says nothing changed. This is 151.
- **Anything calls anything.** From a container on any machine, every `<node>.internal` name and every
routed name the mesh serves resolves — including names the module never declared a requirement on,
which is the guarantee option 2 would have given up.
- **On every network the runtime offers.** The first three hold for a container on the runtime's
default network as well as one on a declared network, because the default network is the case that
has no DNS today.
- **No mesh name is in a container's spec.** A test asserts the digest a host computes for a container
does not move when the mesh's roster does, and does move when the module's own declared entries do.
## Consequences
- **The resolver becomes load-bearing for every container**, where before it was load-bearing for the
machine. This is a real cost and is accepted: a resolver that is down is a machine that cannot
resolve, which is already true of the machine itself, and is a smaller event than a roster change
destroying and recreating every container on the machine.
- **Design 08's "a file rather than a resolver" no longer describes containers.** It was written when
the mesh had no resolver and it gave the right answer then. The reasoning it rested on — every Linux
has a hosts file, no package needed — was already overtaken by names a hosts file cannot express:
service names and wildcards under `<node>.internal`, which is why the resolver was built.
- **ADR 0066's mesh-wide propagation is kept and its mechanism changes.** A routed name still reaches
every asker in the mesh; it reaches them through the resolver rather than by being written into each
container. The consequence 0066 records — that an internal issuer's challenge needs the routed name
resolvable inside the mesh — holds unchanged and by the same means the machine already uses.
- **Issue 110 stops being a container-DNS inconvenience and becomes a prerequisite** for the mesh not
restarting itself whenever it learns a name.
- **A container started by hand gets the mesh's names too**, where before only declared containers did.
Design 08 drew that boundary deliberately, on the grounds that reaching into every container is what
a nameserver would be for. This record accepts that consequence rather than working around it: a
person debugging in a hand-started container resolving the same names as everything else is the
behaviour worth having, and it is what "anything can call anything" means.
## References
- [issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md) — one name replaces every container; the question this answers
- [issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md) — the roster put into the digest
- [issue 109](../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md) — the first arrival
- [issue 110](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md) — the prerequisite
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — routed names propagate mesh-wide; extended here
- [design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md) — the file-not-resolver reasoning this narrows
@@ -0,0 +1,106 @@
---
topic: building it
status: accepted
date: 2026-09-30
deciders: jochen
reconstructed: false
supersedes: 02-DECISIONS/0068-the-lab-takes-requests.md
---
# 149. The live mesh is the test bed
## Context
[ADR 0068](0068-the-lab-takes-requests.md) proposed that the lab accept queued requests
— a bed and a commit — answer them one at a time from a copy it owns, and expose that through tools so
an agent could start a run and come back to it. It has been `proposed` since 2026-09-12 and nothing was
built.
What happened instead is that the mesh became the thing under test. It runs on four machines; every
fault worth finding in the last month was found on them, and none was found in a bed:
- a container holding an address that had not existed for five days, on the control node
([issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md));
- a machine reading healthy for eleven hours while no module could reach another
([issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md));
- one name replacing every container on the hub
([issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md));
- a consumer assertion that is correct on a mesh being raised and fatal on one that is running
([issue 156](../04-ISSUES/156-moving-a-consumers-delivery-subject-stops-the-control-plane/00-report.md)).
The last is the one that settles it. That change was exercised on the raise path — which is what a bed
*is* — and the raise path is the only path on which the fault cannot appear. A bed raises a mesh; it
does not have a mesh that has been running for weeks, with consumers already bound, containers created
against an older roster, and an adopted machine carrying a predecessor's configuration. **The faults
that cost the most were all faults of a mesh that already exists**, and a bed is by construction a mesh
that does not.
Lab runs are also expensive in a way that changed the behaviour around them: each costs a build and
several minutes, so they were batched, and a batched test is one whose result arrives after the next
three changes were already written.
## Considered Options
**1. Build 0068 as proposed.** Rejected. It answers a question nobody is asking: the bottleneck was
never that a person had to sit at the lab, it was that a bed cannot hold the state the faults live in.
Queueing and tooling a mechanism that finds the wrong class of fault faster is not an improvement.
**2. Leave 0068 `proposed`.** Rejected, and it is why this record exists rather than nothing. A record
that contradicts current practice and sits unresolved is worse than either answer: it reads as intent
to anyone who finds it, and the practice it contradicts is written down nowhere but a handoff note.
**3. Record that the live mesh is the test bed, and supersede 0068.** Chosen.
## Decision
**A change is verified against the mesh that is running.** Not because a bed would be unwelcome, but
because the state that breaks things is state a bed does not have: containers made against an older
roster, consumers already bound, an adopted machine, a store with weeks of history.
**A change that can only be exercised on the raise path is not verified.** If the only test available
raises a fresh mesh, the record says so, and says which case was therefore not covered. The words
"exercised on a fresh mesh" are a statement about coverage, not a pass.
**The lab is not retired**, and [ADR 0016](0016-the-lab.md) stands. It remains the place to raise a
mesh from bare, which is the one thing the live mesh cannot be asked to do and the one thing a bed does
better than anything else. What this record removes is the lab as the *default* answer to "is this
change good", and with it 0068's queue, tools and request protocol.
**Accuracy over a green run.** An honest failure on the live mesh beats a pass in a bed that could not
have failed — and a change that is risky on the running mesh is a reason to make the change smaller,
not a reason to test it somewhere it cannot break.
**What 0068 got right is kept as a rule, not a mechanism:** a run reads a copy that is not anybody's
working tree. Every run of the lab that mattered was pinned to a checkout rather than a worktree, and
the ones that were not produced results about code nobody had written down.
## How this is checked
- **A record that says a change was verified says on what.** Where it was a fresh mesh, it says which
case is uncovered. This is the clause that would have caught 156: its change was verified, honestly,
on the only path where it works.
- **The lab is not in the path of a merge.** No check, playbook or handoff requires a bed to have run.
- **0068 is unreachable as intent.** Its status is `superseded` and it names this record, so a reader
arriving at the queue design finds out immediately that it was not built and why.
## Consequences
- **A fault can be introduced on the machines that serve.** This is the cost, it is real, and it was
paid twice in one evening — a control plane crash-looping for half an hour, and every container on the
hub recreated five times. Both were found in minutes because they were live, and both would have
passed a bed.
- **There is no pre-merge gate beyond the repositories' own suites.** `make check` and the three hq
checks are what stands between a change and the machines, which raises what those suites are worth
and makes a test that cannot fail a genuine defect rather than an untidiness.
- **Raising a mesh from bare is now the lab's whole job**, and is exercised deliberately rather than
as a side effect of testing something else. The foundation work
([issue 146](../04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md))
is that job, and it is also the proof that the mesh can make another of itself.
- **An agent cannot hand a run to a queue and come back**, which 0068 would have given. In practice it
watches a push and reads the machines, which is what happened anyway.
## References
- [ADR 0068](0068-the-lab-takes-requests.md) — superseded by this
- [ADR 0016](0016-the-lab.md) — the lab, which stands
- [issue 156](../04-ISSUES/156-moving-a-consumers-delivery-subject-stops-the-control-plane/00-report.md) — correct on the raise path, fatal on a running mesh
@@ -0,0 +1,113 @@
---
topic: what runs on it
status: accepted
date: 2026-09-30
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md
---
# 150. A module's own code runs as supervised processes under the module's one account
## Context
The repository answers "what runs a module's own code" two ways and reconciles them nowhere
([issue 117](../04-ISSUES/117-a-modules-own-code-is-a-container-and-a-process/00-report.md)).
[ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) is accepted and
says **a container** — "the tool runtime carrying that module's compiled code" — and "one module, one
process, one account". Two `proposed` design documents say a **`process`** resource running an argv,
supervised by the machine, and one of them declares *four* of them for a single module and presents
four as the point. Neither design document names 0047 in its `decisions:`, and the string `process`
as a resource type appears in no decision record at all. The thing as built is the container.
Two things have happened since 0047 was written that bear on it directly.
[**ADR 0142**](0142-the-mesh-delivers-its-own-components-as-binaries.md) decided that the mesh's own
components are binaries on the machine rather than container images, and
[issue 114](../04-ISSUES/114-should-the-controller-be-a-container-or-a-process/00-report.md) was closed
by it. That settled the mesh's components and deliberately said nothing about a module's.
And the standing definition of a module hardened: **a module is software that delivers one or more
services, and a module is not a container.** It may deliver them as a container, an installed package
with a unit, a binary, or configuration files; 61 of 73 happen to use a container and 11 do not,
including the resolver, sshd and fail2ban. A rule that a module's *own code* must be a container makes
the one kind of module the mesh writes itself the only kind that has no choice.
## Considered Options
**1. Hold 0047 as written: a container.** Rejected. Its own reasoning does not require one. What 0047
argued for was a runtime **per module** rather than one for the whole node, because a node-wide runtime
could not hold a per-module broker account and per-module runtimes competing on one tool key would each
be handed calls for tools they do not have. A supervised unit per module satisfies that argument
exactly — it is per module, and a unit runs as an account. The container was the mechanism to hand, not
the conclusion.
**2. Let each design document choose.** Rejected; that is the present state and it is what issue 117
reports. A module author reading the guide writes four processes; a module author reading the record
writes a container; nothing tells either that the other exists.
**3. Settle the hosting form as a supervised process, and settle the count separately.** Chosen.
## Decision
**A module's own code runs as one or more supervised processes on the machine, under the module's single
account.** Where [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) says
"a container, the tool runtime carrying that module's compiled code", read this record. Everything else
0047 decided stands untouched: a tool is served on its own key, only the module that serves it answers,
and the module's account is scoped to exactly its tool keys.
**The invariant is the account, not the process count.** 0047's "one module, one process, one account"
carried its weight in the last clause. Its stated worry about a second process was "not a second one to
scope and seal" — a second *identity* to grant, seal a secret to, and scope on the bus. Several
processes sharing the module's one account create no second identity, so nothing further is scoped or
sealed, and a module may therefore declare as many as its work has shapes: events, tools, a
provisioner, a scheduled ingest. **What a module may not have is two accounts.**
**A module that delivers its service as a container still does.** This record is about the code the
module itself carries — its tools, its events, its provisioner — and not about the software it delivers.
A module wrapping a third-party image wraps a third-party image.
**Why supervised by the machine rather than by the mesh:** it is the same answer ADR 0142 gave for the
mesh's own components, for the same reason. A unit the machine restarts needs no image, no registry
pull and no runtime to be up before the mesh's own code can run — which matters most for exactly the
modules whose code the mesh cannot start any other way.
## How this is checked
- **No design document describes a hosting form for a module's own code without citing this record.**
Designs 18 and 20 name it in `decisions:`; this is the gap issue 117's third point reports, and
`cycle.py` already enforces that a to-be design names its decisions.
- **A module declaring several processes resolves to one account.** A test composes a module with more
than one process resource and asserts the mesh mints exactly one broker account for it, scoped to that
module's tool keys and nothing else — which is 0047's invariant stated as an assertion rather than a
sentence.
- **A module's own code does not require the container runtime.** A machine with no container runtime
can still run a module whose code is its own, which is the claim that separates this from option 1 and
is checkable on a machine that has one by asserting the declaration names no image for it.
## Consequences
- **The sidecar port stops being needed.** [ADR 0029](0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md)
records that "anything that is a service plus a sidecar currently has to publish a port to talk to
itself", and the host's `network` shape exists partly for it. A process beside the service on the same
machine reaches it without publishing anything, so that pressure goes.
- **Something must supervise, and it is the machine.** This adds a unit per module's code to what the
host writes and owns. The mesh already writes and owns units — `nftables` proves a module can write one
and run it — so the mechanism exists; the count grows.
- **A module's code is delivered, not pulled**, which puts it behind the same gap as the host's own
delivery ([ADR 0141](0141-the-host-delivers-its-own-successor.md), not built): nothing yet delivers a
version of a module's binary to a machine. A container's code arrives by `docker pull`, and this does
not. **This is the cost of the decision and it is not paid**; until delivery exists, a module whose code
is its own is a module somebody places by hand.
- **Issue 117 is answered and its three disagreements close differently:** container-or-unit is decided
here; one-process-or-several is decided here as several under one account; and whether the record was
consulted is fixed by designs 18 and 20 naming this one.
## References
- [issue 117](../04-ISSUES/117-a-modules-own-code-is-a-container-and-a-process/00-report.md) — the contradiction this answers
- [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) — extended; its "a container" clause is settled here
- [ADR 0142](0142-the-mesh-delivers-its-own-components-as-binaries.md) — the same answer for the mesh's own components
- [ADR 0141](0141-the-host-delivers-its-own-successor.md) — the delivery this depends on and which is not built
- [ADR 0029](0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md) — the sidecar port this relieves
+38 -5
View File
@@ -48,6 +48,31 @@ form above and dated no earlier than the record's own `date:` — an unmarked ed
violation the reviewer looks for in the diff, and a marked one is legible in the record itself. violation the reviewer looks for in the diff, and a marked one is legible in the record itself.
The git history is the backstop, not the record of intent; the note is the record of intent. The git history is the backstop, not the record of intent; the note is the record of intent.
## A pointer back from what a record changes
A new record naming an old one is not enough. **Where a record changes a mechanism an older record
states — without reversing the decision, so no supersession — the older record gets a dated note
saying where its mechanism now lives.** A reader arrives at the old record by following a citation,
and finds text that is still the decision and no longer the method; nothing in it says a later record
moved the method, and the new record is not in their hands.
> **The mechanism changed — YYYY-MM-DD, by ADR NNNN.** What still stands, what moved,
> and why.
Three examples of the shape, all found by being missed: ADR 0066 still described a routed name being
written into every container after 0148 replaced that with resolution; ADR 0047 still said a module's
code runs in a container after 0150 made it a supervised process; and ADR 0016 still read as though the
lab were the test bed after 0149 said the live mesh is. Each was a citation leading to the wrong
answer, in a record that was not wrong about anything it decided.
**This is not machine-checked, and it cannot be from `extends:` alone.** 102 records extend another and
87 name a parent that does not mention them, which is correct: extending usually means building on a
context, and a one-directional pointer is the right shape for that. What needs a note is the narrower
case where the parent's own text has gone stale, and which case that is, is a judgement — so it is a
rule for the author and the reviewer, and the diff is where it is caught. Making it mechanical would
mean a record declaring the relationship in its frontmatter, which is a change to the record schema and
has not been decided.
The records run in the order the decisions were taken, oldest first. The records run in the order the decisions were taken, oldest first.
**Every decision is a record.** There is no ledger and no index file — if a decision is worth **Every decision is a record.** There is no ledger and no index file — if a decision is worth
@@ -174,6 +199,7 @@ python3 00-META/checks/index.py fail if stale
- **0108** — [A route carries the policy applied to a request, and names a secret rather than holding one](0108-a-route-carries-the-policy-applied-to-a-request.md) - **0108** — [A route carries the policy applied to a request, and names a secret rather than holding one](0108-a-route-carries-the-policy-applied-to-a-request.md)
- **0109** — [A package registry seat is one per ecosystem, not one for all of them](0109-a-package-registry-seat-is-one-per-ecosystem.md) - **0109** — [A package registry seat is one per ecosystem, not one for all of them](0109-a-package-registry-seat-is-one-per-ecosystem.md)
- **0126** — [A module declares its own seats; the mesh reserves its own](0126-a-module-declares-its-own-seats.md) - **0126** — [A module declares its own seats; the mesh reserves its own](0126-a-module-declares-its-own-seats.md)
- **0148** — [The mesh's names are resolved, not copied into every container](0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
### What runs on them, and how it gets there ### What runs on them, and how it gets there
@@ -207,9 +233,9 @@ python3 00-META/checks/index.py fail if stale
- **0099** — [A step that runs once names what it reads, and runs again when it changed](0099-a-step-that-runs-once-names-what-it-reads.md) - **0099** — [A step that runs once names what it reads, and runs again when it changed](0099-a-step-that-runs-once-names-what-it-reads.md)
- **0110** — [A seat is held by one assignment, from a closed set, and it may deliver a provision](0110-a-seat-is-a-module-assignment-from-a-closed-set.md) - **0110** — [A seat is held by one assignment, from a closed set, and it may deliver a provision](0110-a-seat-is-a-module-assignment-from-a-closed-set.md)
- **0112** — [A module definition names no node, no mesh and no path: everything it needs is a requirement the mesh resolves](0112-a-module-definition-names-no-node-mesh-or-path.md) - **0112** — [A module definition names no node, no mesh and no path: everything it needs is a requirement the mesh resolves](0112-a-module-definition-names-no-node-mesh-or-path.md)
- **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md) *(proposed)* - **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md)
- **0114** — [A credential two parties hold rotates over two credentials; one a single party holds rotates in place, staged; and retiring a credential never removes what it reached](0114-a-shared-credential-rotates-over-two-credentials.md) *(proposed)* - **0114** — [A credential two parties hold rotates over two credentials; one a single party holds rotates in place, staged; and retiring a credential never removes what it reached](0114-a-shared-credential-rotates-over-two-credentials.md)
- **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md) *(proposed)* - **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md)
- **0117** — [A machine's uplink is a seat: the mesh configures the manager, never the link](0117-a-machines-uplink-is-a-seat.md) - **0117** — [A machine's uplink is a seat: the mesh configures the manager, never the link](0117-a-machines-uplink-is-a-seat.md)
- **0118** — [Undeclaring removes what the mesh made, and gives a unit back the state it was found in](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md) - **0118** — [Undeclaring removes what the mesh made, and gives a unit back the state it was found in](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md)
- **0120** — [A roster fact carries its format as a template: the mesh owns the data, the module owns the format](0120-a-roster-fact-carries-its-format-as-a-template.md) - **0120** — [A roster fact carries its format as a template: the mesh owns the data, the module owns the format](0120-a-roster-fact-carries-its-format-as-a-template.md)
@@ -223,6 +249,12 @@ python3 00-META/checks/index.py fail if stale
- **0139** — [A network is forwarded because a module declared it](0139-a-network-is-forwarded-because-a-module-declared-it.md) *(superseded)* - **0139** — [A network is forwarded because a module declared it](0139-a-network-is-forwarded-because-a-module-declared-it.md) *(superseded)*
- **0140** — [The filter constrains what arrives from outside, and says nothing about a machine's own guests](0140-the-filter-constrains-what-arrives-from-outside.md) - **0140** — [The filter constrains what arrives from outside, and says nothing about a machine's own guests](0140-the-filter-constrains-what-arrives-from-outside.md)
- **0141** — [The host delivers its own successor, and versions live side by side](0141-the-host-delivers-its-own-successor.md) - **0141** — [The host delivers its own successor, and versions live side by side](0141-the-host-delivers-its-own-successor.md)
- **0143** — [A consumer verifies the grant it is given](0143-a-consumer-verifies-the-grant-it-is-given.md) *(superseded)*
- **0144** — [Anything on a machine may call anything on it, and that is the whole of "local"](0144-anything-on-a-machine-may-call-anything-on-it.md)
- **0145** — [A module checks what the mesh claims is reachable, and it checks itself](0145-a-module-checks-what-the-mesh-claims-is-reachable.md) *(superseded)*
- **0146** — [Connectivity is checked by name, per hosting form, with a valid certificate](0146-connectivity-is-checked-by-name-per-hosting-form.md)
- **0147** — [A module anchors the mesh's authority on a machine, and takes it away again](0147-a-module-anchors-the-meshs-authority.md)
- **0150** — [A module's own code runs as supervised processes under the module's one account](0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md)
### How it is built ### How it is built
@@ -232,9 +264,9 @@ python3 00-META/checks/index.py fail if stale
- **0014** — [No workspace — each module is a standalone package consuming published dependencies](0014-no-npm-workspace.md) - **0014** — [No workspace — each module is a standalone package consuming published dependencies](0014-no-npm-workspace.md)
- **0015** — [Applications live in their own repository; the monorepo is for the mesh](0015-applications-live-in-their-own-repository.md) - **0015** — [Applications live in their own repository; the monorepo is for the mesh](0015-applications-live-in-their-own-repository.md)
- **0016** — [The lab](0016-the-lab.md) - **0016** — [The lab](0016-the-lab.md)
- **0037** — [Where a module lives](0037-where-a-module-lives.md) *(proposed)* - **0037** — [Where a module lives](0037-where-a-module-lives.md)
- **0039** — [What the SDK holds, and what it refuses](0039-what-the-sdk-holds-and-refuses.md) - **0039** — [What the SDK holds, and what it refuses](0039-what-the-sdk-holds-and-refuses.md)
- **0068** — [The lab takes requests, one at a time, and runs each from its own copy](0068-the-lab-takes-requests.md) *(proposed)* - **0068** — [The lab takes requests, one at a time, and runs each from its own copy](0068-the-lab-takes-requests.md) *(superseded)*
- **0069** — [A module is a repository and a path within it](0069-a-module-is-a-repository-and-a-path.md) - **0069** — [A module is a repository and a path within it](0069-a-module-is-a-repository-and-a-path.md)
- **0076** — [The SDK is a published package, and the toolchain resolves it by version](0076-the-sdk-is-a-published-package.md) - **0076** — [The SDK is a published package, and the toolchain resolves it by version](0076-the-sdk-is-a-published-package.md)
- **0082** — [The registry is reached by name, and the overlay is its security](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md) - **0082** — [The registry is reached by name, and the overlay is its security](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)
@@ -243,6 +275,7 @@ python3 00-META/checks/index.py fail if stale
- **0097** — [A vendor image is a declared build input, and a recipe fetches nothing undeclared](0097-a-vendor-image-is-a-declared-build-input.md) - **0097** — [A vendor image is a declared build input, and a recipe fetches nothing undeclared](0097-a-vendor-image-is-a-declared-build-input.md)
- **0107** — [Persistent data is a directory bind, never a named volume](0107-persistent-data-is-a-directory-bind-never-a-named-volume.md) - **0107** — [Persistent data is a directory bind, never a named volume](0107-persistent-data-is-a-directory-bind-never-a-named-volume.md)
- **0111** — [A build source is on the mesh's git seat, or it is an external repository](0111-a-build-source-is-on-the-git-seat-or-external.md) - **0111** — [A build source is on the mesh's git seat, or it is an external repository](0111-a-build-source-is-on-the-git-seat-or-external.md)
- **0149** — [The live mesh is the test bed](0149-the-live-mesh-is-the-test-bed.md)
### How it is checked ### How it is checked
+43 -1
View File
@@ -7,8 +7,10 @@ code:
- mesh-controller internal/identity/authority.go - mesh-controller internal/identity/authority.go
- mesh-host internal/identity/serving.go - mesh-host internal/identity/serving.go
- mesh-host internal/apply (the service that reflects a rule set) - mesh-host internal/apply (the service that reflects a rule set)
updated: 2026-09-28 updated: 2026-09-30
decisions: decisions:
- 02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md
- 02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md
- 02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md - 02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md
- 02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md - 02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md
- 02-DECISIONS/0104-a-provision-may-be-answered-by-an-adapter-to-the-predecessor.md - 02-DECISIONS/0104-a-provision-may-be-answered-by-an-adapter-to-the-predecessor.md
@@ -288,6 +290,30 @@ hosts file by the runtime. That extends the file decision rather than overturnin
mesh and not chosen by a module: a module that listed the machines would go stale the day one mesh and not chosen by a module: a module that listed the machines would go stale the day one
joins, and a module that did not would be one whose containers cannot reach anything by name. joins, and a module that did not would be one whose containers cannot reach anything by name.
**Superseded for containers by [ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
(2026-09-30).** Everything above still describes the machine's own roster file, which is how it works
and how it will keep working. It no longer describes containers.
Copying the roster into each container made the roster part of each container's identity, so one name
moving replaced every container in the mesh — a module assigned on one machine restarted the store, the
registry, the edge and mail on another
([issue 151](../../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). And it
did not stop the staleness it was meant to: a copy taken at creation is stale the moment the roster
moves, twice found as a container holding an address that had not existed for days
([issues 109](../../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md)
and [135](../../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md)).
**A container resolves the mesh's names through its machine's resolver, at the moment it asks, and
nothing is copied.** The resolver is a machine-level process rather than a container, so nothing
circular is being asked for. This is gated on a container being able to reach the resolver from any of
the runtime's networks, which it cannot today
([issue 110](../../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)) —
until that lands the mesh keeps copying and keeps comparing, and the order is stated in the record.
The paragraph below states the old boundary, and 0148 deliberately gives it up: a container somebody
started by hand resolves the same names as everything else, because the resolver answers the machine,
not a list of containers.
**The boundary, which is deliberate and worth stating:** *declared* containers. A container **The boundary, which is deliberate and worth stating:** *declared* containers. A container
somebody starts by hand is not the mesh's to configure, and reaching into every container on a somebody starts by hand is not the mesh's to configure, and reaching into every container on a
machine — declared or not — is what a nameserver in `resolv.conf` would be for. machine — declared or not — is what a nameserver in `resolv.conf` would be for.
@@ -710,6 +736,22 @@ step, so when the authority moves the root is fetched again and the proxy is rec
*How it is checked:* the route-forwarding bed installs the authority, the proxy and a consumer *How it is checked:* the route-forwarding bed installs the authority, the proxy and a consumer
from the catalogue and asserts the routed name is served. from the catalogue and asserts the routed name is served.
**And a machine trusts that authority because a module put its root in its trust store**
([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md)). The proxy's fetch
answers for the proxy and for nothing else: a browser, `git` over HTTPS, a package manager and
every module calling another by an internal name read the machine's own trust store, and the mesh
had never written anything there. A module requiring the authority does the whole of it — fetch
the root over the mesh network, place it where this machine's TLS clients look, refresh the
extracted bundles — and stopping it, which is what being unassigned does, takes the anchor away
and refreshes them again. Not the controller's business, because being on the private network is
what makes the authority *reachable* and is not the same fact as having a reason to *verify* a
mesh name; and because where anchors live and which command refreshes them is one operating
system's difference, which is the host's half of the mesh
([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)).
*How it is checked:* on a machine holding the module a plain client verifies an internal HTTPS
name with no bundle argument, and on one without it the same fetch fails to find an issuer — both
halves, because only the pair tells the anchor apart from something that already trusted it.
### What was built ### What was built
*2026-08-31.* *2026-08-31.*
+32 -1
View File
@@ -5,7 +5,7 @@ code:
- mesh-controller internal/builder - mesh-controller internal/builder
- mesh-controller cmd/mesh-controller (build, build --behind, push, status) - mesh-controller cmd/mesh-controller (build, build --behind, push, status)
- mesh-controller internal/inventory/builds.go - mesh-controller internal/inventory/builds.go
updated: 2026-09-21 updated: 2026-09-29
decisions: decisions:
- 02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md - 02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md
- 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md - 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md
@@ -226,3 +226,34 @@ when the current failure began and how many reports in a row have said it — th
id, whatever the words; three make the machine stuck, and `status` says so beside the failure. The host keeps trying — stuck is what the mesh id, whatever the words; three make the machine stuck, and `status` says so beside the failure. The host keeps trying — stuck is what the mesh
knows, not what the machine is told. *How it is checked:* an inventory test counts three identical knows, not what the machine is told. *How it is checked:* an inventory test counts three identical
reports, a different one, and a clean apply; the status test asserts the word appears. reports, a different one, and a clean apply; the status test asserts the word appears.
## Everything may call what is exposed to it, and local is not a boundary
*2026-09-29, from an outage that ran eleven hours —
[issue 145](../../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
settled by [ADR 0144](../../02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md).*
A grant is four facts and a credential: the provision, the machine, the port, and who the consumer is
when it connects. It is the whole mechanism by which anything in the mesh reaches anything else, and it
rests on three things being callable — what runs on the same machine, another machine's service over the
private network where it is exposed there, and another machine's service over the public network where it
is exposed there.
The filter had two of those. A service exposed to the private network admitted the machines' own addresses
on it; a caller on the machine carries such an address, and a caller inside one of that machine's
containers carries a bridge address and matched nothing. Measured, same destination and same machine:
`src 10.10.0.1` from the machine, `src 172.17.0.8` from a container on it. So a module reaching its
database on its own machine's name timed out for eleven hours while the mesh called the machine healthy.
The second case worked by accident: a caller on another machine arrives over the tunnel carrying that
machine's address, which the rule matched. Two of three working is why this read as correct.
**So local is not a boundary this mesh draws, and the filter says so once.** Traffic that did not arrive
from outside the machine and did not arrive over the private network is the machine's own, and is
admitted — for every service there, not per service. Whether the caller is a container, a unit or a shell
decides nothing, because the question is "is this the same machine".
A verification mechanism was drafted for this and withdrawn. It would have reported the outage sooner and
would not have prevented it, and the part of it that was hard — deciding which network position to check
from — existed only because the rule was wrong. Whether the mesh should check that a grant works is still
open, in issue 145; it is not the remedy for a configuration error.
+2 -1
View File
@@ -5,8 +5,9 @@ code:
- mesh-controller cmd/mesh-builder - mesh-controller cmd/mesh-builder
- mesh-controller internal/builder - mesh-controller internal/builder
- mesh-catalog modules/builder - mesh-catalog modules/builder
updated: 2026-09-29 updated: 2026-09-30
decisions: decisions:
- 02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md
- 02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md - 02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md
- 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md - 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md
- 02-DECISIONS/0097-a-vendor-image-is-a-declared-build-input.md - 02-DECISIONS/0097-a-vendor-image-is-a-declared-build-input.md
+2 -1
View File
@@ -5,8 +5,9 @@ code:
- mesh-catalog modules/showcase - mesh-catalog modules/showcase
- mesh-controller internal/builder - mesh-controller internal/builder
- mesh-sdk src - mesh-sdk src
updated: 2026-09-21 updated: 2026-09-30
decisions: decisions:
- 02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md
- 02-DECISIONS/0099-a-step-that-runs-once-names-what-it-reads.md - 02-DECISIONS/0099-a-step-that-runs-once-names-what-it-reads.md
- 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md - 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md
- 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md - 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md
@@ -155,3 +155,17 @@ design document here, and get it back. That check fails today by design.
**What stands until then** is the signpost, and the honest description of it: reachable, not **What stands until then** is the signpost, and the honest description of it: reachable, not
surfacing. surfacing.
## Where this stands, 2026-09-29
*Added in a grooming pass.* The knowledge base this record is about is the **predecessor's**, and it
is no longer reachable from anything: the surface that answered `recall_search` speaks the transport
the mesh removed at the cut-over
([issue 147](../147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md)).
So the sentence in `README.md` that this record catches — *these documents are still indexed into
the knowledge base* — is now wrong twice over: nothing indexed them, and there is nothing to index
them into. The record stays open, and its answer is no longer "index this repository somewhere"; it
is whatever the mesh grows as its own knowledge surface, if it grows one. Until then the honest fix
is the README, which should stop claiming a property nothing provides.
@@ -1,8 +1,8 @@
--- ---
status: open status: resolved
opened: 2026-09-22 opened: 2026-09-22
located-in: [] located-in: [mesh-catalog modules/gitea, mesh-controller internal/catalogue/declaration.go]
fixed-by: fixed-by: mesh-controller 7352c84, merged in #46 — a module is told its port in a container's environment too, as a file already was
amended-design: amended-design:
--- ---
@@ -39,3 +39,9 @@ the assignment happens to differ.
- Should composition refuse an environment value that names a port the module does not fix, the - Should composition refuse an environment value that names a port the module does not fix, the
way it refuses other claims a module cannot make? way it refuses other claims a module cannot make?
- Which other modules write their own address, with a port, into their environment? - Which other modules write their own address, with a port, into their environment?
## Closed
*2026-09-29, in a grooming pass rather than by whoever fixed it.* The forge's address follows a moved port the same way every other reader does. Found by
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
@@ -1,8 +1,8 @@
--- ---
status: open status: resolved
opened: 2026-09-22 opened: 2026-09-22
located-in: [] located-in: [mesh-controller internal/catalogue/declaration.go, mesh-catalog (every routed module)]
fixed-by: fixed-by: mesh-controller bdf965d (a route names the endpoint it serves) with `portOfEndpoint` and `AtPublishedPort` — the contribution carries the endpoint's declared port and the machine-side redirection is applied to it like any other
amended-design: amended-design:
--- ---
@@ -45,3 +45,9 @@ precisely because the predecessor holds the usual one.
keeping the mapping out of rendered configuration? keeping the mapping out of rendered configuration?
- What should refuse a declaration whose contributed route names a port nothing on that node - What should refuse a declaration whose contributed route names a port nothing on that node
listens on? listens on?
## Closed
*2026-09-29, in a grooming pass rather than by whoever fixed it.* A route names an endpoint rather than a port, and the redirection that turns a declared port into the published one is applied to contributions too. Found by
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
@@ -59,3 +59,17 @@ container is made with are the same kind of input, read once at creation, and ar
and is the stronger statement; it is also what the mesh's own resolver exists for. and is the stronger statement; it is also what the mesh's own resolver exists for.
- Either way: what tells an operator that a container is running with an address the node no longer - Either way: what tells an operator that a container is running with an address the node no longer
has? Nothing did. has? Nothing did.
## Answered at the cause (2026-09-30)
This was the first of three arrivals of one fact: a container is given the mesh's names when it is
created and never looks again, so a name that moves afterwards is wrong inside it for as long as it
runs. It arrived again as [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md),
whose fix made the names comparable — and that fix made the roster part of every container's identity,
which arrived as [issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md).
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) ends the
copying: a container resolves through its machine's resolver at the moment it asks. The shape this
record reports then has nowhere to occur. It is gated on
[issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md), so
until that lands the mesh still copies and still compares.
@@ -51,3 +51,19 @@ knows that is what the rule means.
the runtime's default one? That is a stronger rule and would have prevented 109 as well. the runtime's default one? That is a stronger rule and would have prevented 109 as well.
- What checks it? A converged bed with a container on the default network resolving a mesh name is - What checks it? A converged bed with a container on the default network resolving a mesh name is
the missing assertion; nothing in the resolver's own beds covers the filter. the missing assertion; nothing in the resolver's own beds covers the filter.
## What now depends on this (2026-09-30)
This stopped being a container-DNS inconvenience.
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) decides
that a container resolves the mesh's names rather than being given a copy of them, which is what stops
one name moving from replacing every container in the mesh
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)) and what makes a
stale address impossible rather than merely noticed
([issues 109](../109-a-container-keeps-the-address-it-was-made-with/00-report.md)
and [135](../135-a-containers-mesh-names-are-not-compared/00-report.md)).
**That decision cannot land until this one does**, and not partly: a container on the runtime's default
network is the case with no DNS at all, and it is the case the mesh's own forge runs in. Two of four
machines also bind the resolver to loopback only, so the runtime hands their containers a public
resolver. Both halves are this issue.
@@ -1,8 +1,8 @@
--- ---
status: open status: resolved
opened: 2026-09-24 opened: 2026-09-24
located-in: [mesh-controller module.json, mesh-host internal/apply] located-in: [mesh-controller module.json, mesh-host internal/apply]
fixed-by: fixed-by: ADR 0142 — the mesh's own components are delivered as binaries on the machine, so the controller is a process; the delivery itself is issue 142
amended-design: amended-design:
--- ---
@@ -84,3 +84,21 @@ restart and run-to-completion semantics — so this would not need host-side wor
The two do not collapse into one. The controller is not a code-carrying sidecar, and `network: host` The two do not collapse into one. The controller is not a code-carrying sidecar, and `network: host`
is what makes the asymmetry visible here and nowhere else. is what makes the asymmetry visible here and nowhere else.
## Answered
*2026-09-29, in a grooming pass.* This asked a question rather than reporting a defect, and the
question was taken: [ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md)
decides that the mesh's own components — the host, the controller, the catalogue, the builder, the
vault — are **binaries on the machine**, delivered by the mechanism
[ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md) built, and that
third-party software (the store, the registry, the broker) stays a container because an image is the
right way to carry somebody else's build.
So the operating experience this record was written from — every mutating command reached through
`docker exec mesh-controller` — is answered, and answered against the container.
**The delivery is a separate matter and is not this record's.** Step 1 of it is built and no
component travels yet; that is
[issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md).
@@ -1,8 +1,8 @@
--- ---
status: located status: resolved
opened: 2026-09-25 opened: 2026-09-25
located-in: [hq, mesh-catalog modules/showcase, mesh-sdk src/tools/index.ts, mesh-tools] located-in: [hq, mesh-catalog modules/showcase, mesh-sdk src/tools/index.ts, mesh-tools]
fixed-by: fixed-by: hq 83791f0 (PR 196) — ADR 0150: a module's own code runs as supervised processes under the module's one account; designs 18 and 20 now cite it, and ADR 0047 carries a dated note pointing at it
amended-design: amended-design:
--- ---
@@ -99,3 +99,27 @@ what the sidecar is has nowhere in the design layer to look, which is how this w
is mechanically checkable: the resource types a design doc names are a closed set, and every is mechanically checkable: the resource types a design doc names are a closed set, and every
member of it either appears in a decision or does not. Whether that check is worth writing is member of it either appears in a decision or does not. Whether that check is worth writing is
part of this issue, not settled by it. part of this issue, not settled by it.
## Answered (2026-09-30)
[ADR 0150](../../02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md)
settles all three disagreements, and the design documents win two of them:
1. **Container or unit — a supervised process.** ADR 0047's argument never required a container. It
argued for a runtime *per module*, because a node-wide one could not hold a per-module account and
per-module runtimes on one tool key would be handed calls for tools they do not have. A unit per
module satisfies that exactly, and a unit runs as an account. The container was the mechanism to
hand. [ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md) had
already gone the same way for the mesh's own components, and a module is not a container.
2. **One process or several — several, under one account.** 0047's "one module, one process, one
account" carried its weight in the last clause; its stated worry was "not a second one to scope and
seal", which is about a second *identity*. Processes sharing the module's one account create none.
What a module may not have is two accounts.
3. **Whether the record was consulted — fixed rather than answered.** Designs 18 and 20 now name 0150
in `decisions:`, and 0047 carries a dated note saying where its hosting form was settled, so neither
door leads to the wrong answer any more.
**What this does not fix.** A module's code delivered as a binary is behind the same gap as the host's
own ([ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md), accepted and not
built): a container's code arrives by `docker pull` and this does not, so until delivery exists such a
module is one somebody places by hand. 0150 records that as the cost of the decision, unpaid.
@@ -1,8 +1,8 @@
--- ---
status: located status: resolved
opened: 2026-09-26 opened: 2026-09-26
located-in: [mesh-sdk src/provisioner, mesh-catalog modules/redis] located-in: [mesh-sdk src/provisioner, mesh-catalog modules/redis]
fixed-by: fixed-by: mesh-catalog bbda88c, merged in #84 — every credential provider says whether it still holds a consumer
amended-design: amended-design:
--- ---
@@ -60,3 +60,9 @@ checks it after the first pass.
instance and leaves the gap for the others. instance and leaves the gap for the others.
- Where does the record of what was applied live, if not in memory? ADR 0114, still - Where does the record of what was applied live, if not in memory? ADR 0114, still
proposed, puts rotation state with the vault. The same place may answer this. proposed, puts rotation state with the vault. The same place may answer this.
## Closed
*2026-09-29, in a grooming pass rather than by whoever fixed it.* A provisioner asks the backend what is there rather than trusting what it remembers doing. Found by
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
@@ -1,5 +1,6 @@
--- ---
status: located status: resolved
fixed-by: mesh-host 1cb8953 and fdc768c — the mesh writes into a marked block of a text file instead of over it
opened: 2026-09-26 opened: 2026-09-26
located-in: [mesh-controller internal/catalogue/facts.go, mesh-host internal/apply] located-in: [mesh-controller internal/catalogue/facts.go, mesh-host internal/apply]
--- ---
@@ -71,3 +72,9 @@ private network loses that name too.
- The host's file resource supports `into: "json"` only; anything else is a whole write. - The host's file resource supports `into: "json"` only; anything else is a whole write.
- `node show <node>` on the adopted workstation: `holds file /etc/hosts - `node show <node>` on the adopted workstation: `holds file /etc/hosts
mesh-wireguard.fact-node-names`, original kept. mesh-wireguard.fact-node-names`, original kept.
## Closed
*2026-09-29, in a grooming pass rather than by whoever fixed it.* A shared hosts file keeps every line that is not the mesh's. Found by
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
@@ -1,7 +1,7 @@
--- ---
status: open status: located
opened: 2026-09-26 opened: 2026-09-26
located-in: [mesh-controller, mesh-catalog step-ca] located-in: [mesh-catalog ca-trust]
--- ---
# 129 — nothing makes a machine trust the mesh's own certificate authority # 129 — nothing makes a machine trust the mesh's own certificate authority
@@ -0,0 +1,45 @@
# Diagnosis
*2026-09-29.*
## What was ruled out
**That something already carries the root and it is only misplaced.** It does not. The authority
serves its root at a path beside its ACME directory, and the one thing that fetches it — the route
proxy — puts it in a directory of its own and hands it to one program. Nothing has ever written
into a machine's trust store. Measured on three converged machines: the anchors present are the
predecessor's authority and a developer tool's local root, and on the machines where the
predecessor's was deliberately removed, every internal name fails verification.
**That the private network could carry it, the way it carries the registry's trust.** That is what
the report proposed, and it was rejected on consideration rather than on difficulty
([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md), option 1): being on
the network is what makes the registry *reachable* and is therefore the right trigger there, while
trusting an authority is a separate fact from being able to reach it. The anchor's directory and
the command that refreshes the extracted bundles are also one operating system's difference, which
is the host's half of the mesh and not the controller's.
**That it needs a new host resource type.** It does not, today. A file and a service say the whole
of it, which the packet filter already proves. The primitive becomes the right answer when a second
operating system is in play, and not before.
## Where it belongs
A module in the catalogue: it requires `internal-acme-ca`, fetches the root over the mesh's own
network, installs it as a trust anchor, refreshes the machine's bundles, and — because being
unassigned stops its unit, and stopping the unit is what undoes it — takes both away again.
The owner is therefore `mesh-catalog`, module `ca-trust`, and nothing in the control plane.
## The module exists, and this stays open until a machine holds it
*2026-09-29.* `ca-trust` is in the catalogue and merged
([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md)), and what it renders
is checked in the control plane's own suite: the script fetches from the authority it was bound to,
and the unit runs it both ways.
**No machine has been assigned it, and nothing has verified a name because of it.** The bed written
for that cannot run ([issue 146](../146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md)),
and the live mesh has not been given the module. So the symptom this record opened on — every
internal name failing verification on every machine — is still true everywhere, and the record stays
`located` until it is not. Closing it on a module that exists would be closing it on an intention.
@@ -1,5 +1,6 @@
--- ---
status: located status: resolved
fixed-by: mesh-host 3112c88 — undeclaring gives a unit back the state it was found in, and removes only a process the mesh made
opened: 2026-09-27 opened: 2026-09-27
located-in: [mesh-host internal/apply/apply.go (remove)] located-in: [mesh-host internal/apply/apply.go (remove)]
amended-design: 02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md amended-design: 02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md
@@ -15,7 +16,7 @@ found that the host's `remove` path stops every `service` resource that is no lo
delete". `store.Orphans` matches by id alone. So any of these stops the unit: delete". `store.Orphans` matches by id alone. So any of these stops the unit:
- the module is unassigned — by mistake, or to switch it for another; - the module is unassigned — by mistake, or to switch it for another;
- the node is sent a deliberately-empty declaration ([issue 127](../127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)); - the node is sent a deliberately-empty declaration ([issue 149](../149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md));
- a later catalogue version renames the resource's `id`. - a later catalogue version renames the resource's `id`.
That is right for a service the mesh brought into being. It is wrong for a unit the mesh That is right for a service the mesh brought into being. It is wrong for a unit the mesh
@@ -64,3 +65,9 @@ something to settle in passing.
The unassign preview is partly answered — the host's plan names each unit it will stop — and the The unassign preview is partly answered — the host's plan names each unit it will stop — and the
controller's side is left open. controller's side is left open.
## Closed
*2026-09-29, in a grooming pass rather than by whoever fixed it.* Undeclaring no longer stops a unit the mesh only reloaded or only kept running. Found by
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
@@ -68,3 +68,20 @@ container runtime's shape, not a choice; the answer is to recreate, which is wha
that would rather re-read a roster from a file can already ask for one as a fact that would rather re-read a roster from a file can already ask for one as a fact
([ADR 0120](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md)) and restart on ([ADR 0120](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md)) and restart on
it. it.
## What replaced this fix (2026-09-30)
The fix here — putting the mesh's names into the digest the host compares, so a container whose names
moved is recreated like one whose image moved — worked, and cost more than it was worth. It made the
roster part of every container's identity, so one name moving replaced every container in the mesh: a
module assigned on one machine restarted the store, the registry, the edge and mail on another
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)).
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) goes at
the cause this record only described: **a copy taken at creation is stale the moment the roster moves,
and detecting that is not as good as not copying.** A container resolves the mesh's names through its
machine's resolver, at the moment it asks, so the fault this record reports cannot occur rather than
being noticed a restart later.
Recorded here because this is where somebody arrives to find out why the digest carries names, and the
answer is that it did, for two days short of a month, and stopped.
@@ -1,5 +1,5 @@
--- ---
status: located status: resolved
opened: 2026-09-28 opened: 2026-09-28
located-in: located-in:
- mesh-controller internal/catalogue/manifest.go - mesh-controller internal/catalogue/manifest.go
@@ -7,7 +7,7 @@ located-in:
- mesh-controller internal/catalogue/declaration.go - mesh-controller internal/catalogue/declaration.go
- mesh-controller examples/route-proxy - mesh-controller examples/route-proxy
- mesh-catalog (every routed module manifest) - mesh-catalog (every routed module manifest)
fixed-by: fixed-by: mesh-controller bdf965d (a module names its endpoints) and c68d3a7 (an assignment configures an endpoint as one thing) — the filter, the proxy's names and both authorities now read one statement
amended-design: 03-DESIGN/01-to-be/08-connectivity.md amended-design: 03-DESIGN/01-to-be/08-connectivity.md
--- ---
@@ -0,0 +1,47 @@
# Resolution
*2026-09-29.*
**Built, and this record did not say so.** The issue was written on 2026-09-28 and answered the same
week by two commits in `mesh-controller`; nothing came back to close it, so the mesh's own account of
itself said for a day that reach was declared nowhere while the code read it in three places.
- `bdf965d` — *a module names its endpoints, and a route names the one it serves*. `listens[].name`
is the endpoint; a route contribution names the endpoint rather than repeating a port.
- `c68d3a7` — *an assignment configures an endpoint as one thing*. The `endpoints` settings key, per
node, by endpoint name: `{"endpoints": {"ssh": {"port": 20134, "reach": "public"}}}` — port, label
and reach in one block, which is what [ADR 0138](../../02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md)
asked for and what [ADR 0046](../../02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md)
said configuration is.
## The three readers, which is what the issue was about
The complaint was that the per-node source override had exactly one caller. It now has three, and
they are the three mechanisms reach was decided to settle at once:
| reader | what it does with it |
|---|---|
| the filter | `Reaches` turns each endpoint's reach into the rule for its machine port |
| the proxy's names | `composeName` composes the public name, the internal name, or both — and a name nobody asked for is not composed |
| the authorities | the proxy certifies only names it was actually given, each from its own authority, through two host policies rather than one |
**A routed endpoint keeps the manifest's port**, which is ADR 0138's own insight and older than it
([ADR 0045](../../02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)): the
proxy is how it is reached, so `public` there asks for a public *name*, not an open port.
**An endpoint that is not routed is reached and never named.** Git over ssh is that case — the one
the issue said the model could not express — and it is now the ordinary one.
## How it is checked
`internal/catalogue/endpoints_setting_test.go`: a block says port, label and reach; a block may say
only a reach; a name the module does not declare is refused; a reach outside the four values is
refused; and saying the same thing twice — once in the block, once through the older per-port keys —
is refused rather than resolved by whichever is read last. The proxy's half is `policy_test.go` and
`authority_test.go`: a name the mesh did not send is not certified, by either authority.
## What is left, and it is not this
The older keys (`ports`, `expose`, and reach keyed by port) still work beside the block. They are
what the block replaces, and retiring them is its own small change — not a gap in what reach can
say.
@@ -5,7 +5,7 @@ located-in:
- mesh-controller internal/catalogue/filtering.go (fixed for this instance) - mesh-controller internal/catalogue/filtering.go (fixed for this instance)
- mesh-controller (what status reports, and what it does not ask) - mesh-controller (what status reports, and what it does not ask)
fixed-by: fixed-by:
amended-design: amended-design: 03-DESIGN/01-to-be/10-delivery.md
--- ---
# 145 — A machine reads healthy while its modules cannot reach each other # 145 — A machine reads healthy while its modules cannot reach each other
@@ -72,6 +72,36 @@ distance between a declaration and the machine, and in both cases the report was
**And the eleven hours are the measurement, not the bug.** The filter fault was one line and is fixed. **And the eleven hours are the measurement, not the bug.** The filter fault was one line and is fixed.
What is not fixed is that nothing in the mesh would have told anybody. What is not fixed is that nothing in the mesh would have told anybody.
## What was decided
*2026-09-29, the same day, in two steps and the first was wrong.*
The first answer was [ADR 0143](../../02-DECISIONS/0143-a-consumer-verifies-the-grant-it-is-given.md):
the consumer verifies each grant from its own network position, because whether a caller sat in a
container changed whether it could reach the provider. **That difference was the fault**, and the record
is superseded. A verification mechanism would have reported this sooner and would not have prevented it,
and the part of it that was difficult — deciding which network position to check from — existed only
while the rule was wrong.
The remedy is [ADR 0144](../../02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md):
anything on a machine may call anything on it, said once rather than per service, and asked by the link
traffic arrives on rather than the address it carries. Everything should be able to call what runs on the
same machine, another machine's service exposed to the private network, and another machine's service
exposed publicly. The filter had the second and third and expressed the first as a list of addresses that
no container could match.
**And then the question this issue is actually about was answered on its own terms.**
[ADR 0145](../../02-DECISIONS/0145-a-module-checks-what-the-mesh-claims-is-reachable.md): a module on
every machine serves an endpoint of its own and dials every other machine's, from the position the
callers are in. Its probe is its own endpoint declared reachable over the private network, so it is
admitted by exactly the rule that governs every internally-exposed service and fails when that rule is
wrong — where a probe on a service every machine has would have passed for all eleven hours, because the
services every machine has are the ones never closed.
Adopted on its merits rather than as the remedy for a configuration error, which is what 0143 was and
why it went. The module is written and merged; it is not yet assigned, so every machine currently reads
unchecked.
## Open questions ## Open questions
- Should a grant be checked? The mesh knows the consumer, the provider, the address and the port, so a - Should a grant be checked? The mesh knows the consumer, the provider, the address and the port, so a
@@ -0,0 +1,85 @@
---
status: located
opened: 2026-09-29
located-in: [mesh-host examples + internal/link, mesh-controller internal/broker]
---
# 146 — the foundation cannot be raised on the bus the mesh runs on
## What was observed
Raising a first node in the lab, to check a module against a real mesh, fails before any module is
reached. Two separate faults, in the two bundles that exist:
**The older bundle raises a control plane that cannot start.** It brings up the previous broker,
and the control plane it then starts says, once every few seconds, for ever:
```
mesh-controller: this control plane has no MESH_BUS_NATS, so it cannot reach the mesh's bus
```
That is the control plane being right. The mesh moved to one bus
([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)) and the
bundle did not. Every bed that raises a foundation raises this one, so every bed is in this state.
**The newer bundle, written for the new bus, stops one step earlier.** Its certificate step asks a
container to make the broker's certificate:
```
docker run --rm --entrypoint sh -v <the broker's tls volume>:/tls <the bus image> \
-c "test -f /tls/tls.crt || (openssl req -x509 ... )"
...
failed bus-certificate: running the action: docker exited 127
```
127 is *command not found*. The bus's image has a shell and no `openssl`; the previous broker's
image had both, which is why the step worked when it was written against that one. Substituting the
store's image — the only other image the bundle carries — does not help: it has no `openssl`
either. So the step as written cannot succeed with anything the bundle names, and the fault is not
one image's: **the bundle asks for a certificate to be made by a tool it never says must be there.**
Measured 2026-09-29 on a fresh lab machine, both bundles, from bare.
## Why this is here and not a note in the knowledge base
The mesh's own foundation is the one thing it cannot raise. Nothing reports that: the bundles are
files in a repository, nothing applies them but a person raising a node, and the last thing that
did was the hand-driven cut-over
([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md), whose
work was done on the machines rather than from a bundle). So the state where the mesh cannot make
another one of itself is reachable, and was reached, without anything saying so.
It is also load-bearing for everything else: a lab bed proves a claim by raising a mesh, so while
this holds, **no bed can run**, and every "checked in the lab" written from now on is a promise
against a suite nobody can execute.
## What would have prevented it
- **Something raising the foundation on a schedule, from the bundle, as it is written** — the
bundle is the mesh's own installer and nothing installs from it. A bed that raises a first node
is exactly that check, and it is the bed that cannot run.
- **A step naming what it needs.** The certificate step names an image and assumes a program inside
it. An action that said which tool it requires would have failed at the declaration rather than
at 127 on a machine.
## Evidence to carry into diagnosis
- `mesh-host examples/foundation-first-node.lock` — the previous broker, no `MESH_BUS_NATS`.
- `mesh-host examples/foundation-first-node-nats.lock` — the new bus; `bus-certificate` and its
`verify` both run `openssl` in the bus's image.
- The bus image the bundle pins has `sh` and no `openssl`; the store's image likewise.
- The lab rewrites a bundle's registry-prefixed third-party references to upstream ones for a
machine with an uplink (`test/integration/harness.ts`); the new bundle's bus reference needed
that rule added, which is done and is not this issue.
## What one of its fixes then did to a running mesh (2026-09-30)
The change that stopped the doubling — putting the stream into a push consumer's delivery subject —
is correct on a foundation being raised and fatal on a mesh that is already running: the server will
not move that subject while a subscriber is bound, and a node is bound to its declaration consumer
the whole time it is up. The control plane crash-looped on the first build that carried it.
Recorded and fixed as [issue 156](../156-moving-a-consumers-delivery-subject-stops-the-control-plane/00-report.md).
Noted here because this record is where somebody will arrive when reading why the subject carries the
stream at all, and the answer is incomplete without it: **the raise path was the only one exercised,
and it is the one path on which nothing is bound.**
@@ -0,0 +1,184 @@
# Diagnosis
*2026-09-29, by raising a first node in the lab over and over and writing down each thing it hit.*
Not one fault. **Four, stacked**, each hidden behind the one before it, and every one of them the
same shape: a step that was right while the mesh ran on the previous broker and was never asked a
question again after the bus changed. Nothing had raised a foundation since, so nothing said so.
## 1 — the bundle's bus image is named for a registry that is gone *(fixed)*
The newer bundle pins `<a lab registry>/nats@…`, which resolves nowhere outside the lab that
raised that registry. The lab already rewrites the store's and the previous broker's references to
upstream ones for a machine with an uplink; the bus had no such rule because no bed had ever tried
to raise this bundle. Added (`mesh-lab test/integration/harness.ts`). The digest is the bundle's
own — what the registry served was a copy, so the same digest resolves upstream, and this is a
prefix being removed rather than a reference being replaced.
## 2 — the bus's certificate was made by a tool the bus does not have *(fixed)*
```
failed bus-certificate … docker exited 127
```
The step ran `openssl` inside the broker's image. The previous broker's image carried it; the bus's
does not — it is Alpine with a shell and no `openssl` — and neither does any other image the bundle
names, so there was nothing to substitute. **The program that needs the certificate now makes it**:
`mesh-controller broker certificate --into <dir>`, with `--check` as the step's verify. The
controller is already on the machine at that point (the schema step ran it) and needs nothing from
the image it writes into. Self-signed, as before and on purpose — a host pins this server's exact
certificate ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)) and at that moment
there is no authority to ask. Idempotent, because a second certificate is one every host that
pinned the first no longer believes. It runs `--user 0:0`: the volume is root's, and the control
plane's image runs as nobody, which is right for the long-lived server and wrong for a one-shot
writing into a fresh volume.
## 3 — enrolment dialled TLS at a bus that speaks first *(fixed)*
```
mesh-host: cannot reach the broker at …:5671: tls: first record does not look like a TLS handshake
```
Enrolment opened a raw TLS connection to check the pinned certificate before saying anything. NATS
speaks its own protocol and upgrades afterwards, so the handshake met a plaintext greeting. The pin
was never the problem: the same pinned configuration is handed to the client that presents the
token, and the verification runs inside *that* handshake — so what
[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) requires still holds, and holds
better, because the one-time secret is sent only after the certificate has been checked. The raw
dial is gone from the enrolment path and kept only as what its tests always proved: that a wrong
certificate is refused before a byte of application data is sent.
**Then, immediately behind it:**
```
mesh-host: this token is for the "" bus, and the mesh's bus is nats
```
The enrolment left the transport empty and meant *whatever the mesh runs today*, which was true
while two buses existed and became a refusal the moment one did. The host knows which bus the mesh
runs; it says so now.
## 4 — a first node cannot be let onto its own bus *(open, and this is the real one)*
```
mesh-host: cannot reach the bus at …:5671 as anchor: nats: Authorization Violation
```
The bus's user list is a file beside its configuration. The installer carries the first one — the
controller's own account at a bootstrap password — and **the controller composes every user after
that** (design 25 §6; the controller's own test asserts the carried list matches what it would
derive). On the running mesh that composition reaches the bus because the bus is a *module*, with
the list delivered to it the way anything is delivered to a module.
At genesis there is no module. The foundation's bus is raised by the installer, the control plane is
given no way to write beside its configuration — it mounts the certificate and nothing else — and so
the account a joining node needs cannot come into existence. **The first node cannot join the mesh
it just raised.**
That is not a line to fix in a bundle. It is the open half of the mesh delivering its own components
([ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md)) and of the
bus becoming a module: either the installer's bus is raised as the module the mesh will go on
managing, or genesis carries a user list that includes the first node's enrolment and the controller
takes over from there. Both are decisions, not patches, and both belong to the genesis step that was
deliberately left until last.
## 5 — the composed user list has to be placed by hand at genesis *(fixed)*
The account a token is the password of is **not recorded at all**: the composer names an enrolment
user for every machine with a live token, nothing minted a credential for it, and the composition
left it out as a user with no password. The comment above the issuing code already claimed
otherwise — *"the account is created before the token is handed over"* — which is how it went
unnoticed. Issuing a token now records that account, with the token's own secret as its password,
because that is the string the machine will present.
Placing it is the other half. The list reaches the machine running the bus in that machine's
declaration, which a machine that has not enrolled does not get, so at genesis it cannot arrive
that way. **The control plane composes and says what it composed** — `broker accounts`, to standard
output — and whoever is raising the machine writes it beside the bus's configuration and makes the
server re-read it. Twice, because two accounts come into existence at different moments: the
enrolment when the token is issued, and the machine's own when it enrols. A control plane that
wrote the file itself would have to know where the bus keeps its configuration and how to make it
reload, which is the module's knowledge and is what the module takes over on the first push.
With that, **a first node enrols against the bus it just raised** — measured, from bare, in the
lab.
## 6 — and is enrolled twice, keeping a credential the mesh has replaced *(fixed)*
```
mesh-controller: enrolled anchor
mesh-controller: enrolled anchor (the same second)
```
One `enrol` on the machine, two enrolments in the control plane. Each mints the node a fresh bus
password and returns it; the machine keeps the answer to the first, and the mesh keeps the hash of
the second. The machine then reconnects for ever as a user whose password the mesh rotated out from
under it — *authentication error - User "node.anchor"* on the bus, `Authorization Violation` in the
host's log, and a node that never reports.
What is ruled out: the host asking twice — it asks again only when the mesh says *try again*, and
a refused attempt is not logged as an enrolment. Redelivery by the consumer — there is one
consumer, its acknowledgement window is thirty seconds, and the handler is quick.
What is left: the client re-publishing when an acknowledgement is slow, which is what its defaults
do. That was addressed by giving the publish a message id derived from its own bytes, so the stream
discards the copy — **and the duplicate survived it**, so either the id is not reaching the stream
or the second copy is not a copy. This is where the trail stops.
Worth saying plainly: **the mint is the fragile part, not the delivery.** An enrolment answered
twice is survivable if the answer is the same both times, and it cannot be — the mesh keeps only
the hash, so a second answer is necessarily a different credential.
**Found, and it is not about enrolment at all** *(fixed)*. The bus's own counters settled it: one
message published, one held in the stream, one delivery, nothing redelivered — and the controller
enrolled the machine twice. So the handler ran twice on one delivery.
A push consumer delivers onto an ordinary subject, and **everything subscribed to that subject gets
a copy**. The controller holds a consumer called `controller` on CONTROL and another called
`controller` on EVENTS, and the delivery subject was derived from the consumer's name alone — so
both were `_DELIVER.controller`, the one process held both subscriptions, and every message from
either stream was acted on twice.
Enrolment is where it drew blood, because enrolling twice mints twice and the second credential
replaces the first. But it applied to **every report and every event the controller follows**, and
it is the kind of fault that leaves no trace: nothing is redelivered, no counter is wrong, the work
simply happens twice. The comment in the receiving code about a merge that ran the whole catalogue
five times over on 2026-09-28 is the same shape seen from the other end.
The stream is in the delivery subject now, because the pair is what identifies a consumer — the
server scopes a durable's name to its stream, and this subject was the one place that scoping was
dropped. A subscriber's permission gains the same shape, keeping the bare name so an existing
consumer keeps working until the controller's next assertion moves it.
*How it is checked:* the consumers the mesh asks for are asserted to deliver onto distinct subjects,
in the controller's own suite. Against a server it would be invisible, which is the point.
## Where it belongs
`mesh-host` (the bundle and the enrolment path) and `mesh-controller` (the certificate command, and
the composition that cannot reach the bus at genesis). Three of the four are fixed on branches; the
fourth is the genesis work.
## What made it slow, and what was changed so it is not
Six faults behind one another, each found by raising a machine and reading what it said. What cost
the most was not the faults:
- **Every bed's own instructions named the bundle that cannot work**, so the first three attempts
ended in a control plane crash-looping on a missing bus. They name the working one now.
- **A host binary built without its system** refuses everything it is given with *this host was
built for ""*, which reads like a broken bundle. The lab's README says so.
- **`make image` in the control plane had been broken for as long as its base was pinned**: the
Dockerfile's fallback is a Go older than the module asks for, and the pipeline never saw it
because the pipeline passes the declared base in. It reads the base from the manifest now.
- **Leaving the machine standing is what answers the question.** Every finding above came from
shelling in afterwards — the host's log, the bus's log, the file the bus was actually handed —
and none from the test's own output, which says only that nothing converged. The bed takes
`MESH_LAB_KEEP`, and the README says to reach for it first.
## What it cost, for the next person
Every lab bed still names `foundation-first-node.lock` in its own instructions, and that bundle
raises the previous broker with a control plane that refuses to start without `MESH_BUS_NATS`. Until
the fourth fault is answered and the two bundles become one, a bed runs with `MESH_LAB_BUNDLE`
pointing at the NATS bundle by hand, and stops at the enrolment.
@@ -0,0 +1,75 @@
---
status: located
opened: 2026-09-29
located-in: [nothing in the mesh — the surface in use is the predecessor's, installed on the workstation]
---
# 147 — the operator's tools still dial the bus that was removed
## What was observed
Every tool call an operator makes against the mesh fails, on every machine, with the same answer:
```
AMQP not connected — cannot reach hal/mesh@novox
AMQP not connected — cannot reach hal/mesh@shanks
```
The mesh moved to one bus and the previous transport was deleted
([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md), cut over
2026-09-28). The surface an operator drives the mesh through — the tool bridge their assistant
speaks to — still opens an AMQP connection, so it cannot reach anything. Not one node answers,
including the machine the operator is sitting at.
**What that leaves.** The mesh itself is healthy: nodes are current, modules run, the bus carries
the mesh's own traffic. What is gone is the way a person asks it anything. Every question — what a
node runs, what is assigned, what a module's settings are — has to be asked by opening a shell on
the machine and running the control plane's binary inside its container, which is precisely the
path the tool surface exists to remove, and which nothing checks, records or permits.
It also silently changes how work gets done: an assistant told to use the mesh's tools finds them
dead, falls back to `ssh` and `docker exec`, and the operator discovers the fallback rather than
the fault.
## Why this is here and not a note in the knowledge base
The rule is that everything happens over the mesh's bus. A surface that cannot reach that bus is
that rule enforced by nothing — and the mesh reports itself healthy throughout, because what broke
is not a module, a node or a provision but the thing standing outside asking them questions.
The same cut-over removed the assistant's long-term memory (`recall`, `memorize`) for the same
reason, which is why lessons from the last two days were written into this repository by hand.
## What would have prevented it
- **The tool surface as a consumer of the bus like any other**, so moving the bus moves it — rather
than a separate bridge with its own connection settings that nothing resolves.
- **A check that a tool call reaches a node**, run where the mesh's other checks run. Every check
the mesh makes today is about the relationship between the mesh and a machine; none asks whether
a person can ask it anything ([issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
the same shape one level out).
## Diagnosed at once, because the answer was in the configuration
**The surface is not the mesh's.** The tool server the operator's assistant speaks to is the
predecessor's brain, installed on the workstation and started as a local process, with the
predecessor's broker URL — `amqp://…@<the control node>` — written into the assistant's own
configuration. Nothing about it is a module: it has no manifest, no assignment, no seat, no account
on the mesh's bus, and the mesh has never known it exists.
So nothing regressed. The mesh removed a transport that this program still dials, and the program
was never part of the mesh to be moved. **The mesh has a tool model** — `tools` in a manifest, the
calls a seat accepts, the subjects a module answers on, and a control-plane command that asks one —
and **nothing publishes an operator-facing surface onto it.** What an operator uses is the thing
that came before, kept alive by a URL in a file.
That is the issue, and it is larger than a broken connection: the way a person drives this mesh is
outside the mesh.
## Evidence to carry into diagnosis
- The failure text names `hal/mesh@<node>`, so the bridge is resolving a node and then dialling the
old transport.
- It fails identically for the local machine, which rules out reachability and points at the
transport alone.
- The mesh's own traffic over the new bus is unaffected: nodes report, declarations apply.
@@ -0,0 +1,37 @@
---
status: open
opened: 2026-09-29
located-in: [mesh-controller cmd/mesh-controller]
---
# 148 — a manifest outside this catalogue has no check
## What was observed
A module's manifest is validated by a **test** — `internal/catalogue`'s suite parses every manifest
in the catalogue checkout beside it and fails on one it cannot resolve. That works, and it is how
several real faults were caught before a machine saw them.
It is available to exactly one repository: this one. Somebody describing their own application in
their own repository — the case
[ADR 0037](../../02-DECISIONS/0037-where-a-module-lives.md) calls *the case that matters most* —
has no check at all. They write a manifest, register it with a running mesh, and find out whether
it is valid when the mesh refuses it, or later, when a machine applies something that resolved and
should not have.
The same record asks for the answer: **a `module check` command on the control plane's binary**, so
a manifest is checked by the tool rather than by a test that imports the tool's internals.
## What would have prevented it
Nothing prevents this; it was noticed and left. ADR 0037 named it on 2026-09-01 and the record sat
`proposed` until 2026-09-29, so the missing half was never anybody's task.
## Evidence to carry into diagnosis
- `mesh-controller/internal/catalogue` — `ParseManifest` and `CatalogueProblems` are the check, and
both are internal.
- The catalogue-wide test is `TestEveryCatalogueManifestDeclaresWhatItMounts` and its siblings; they
take a path from `MESH_CATALOG`, so the mechanism is already path-driven and not repository-bound.
- `mesh-controller module add` refuses a bad manifest at registration, which is the same check far
too late: by then it is in a running mesh's records.
@@ -1,10 +1,15 @@
--- ---
status: located status: resolved
opened: 2026-09-27 opened: 2026-09-27
located-in: [mesh-controller cmd/mesh-controller/push.go] located-in: [mesh-controller cmd/mesh-controller/push.go, mesh-controller cmd/mesh-controller/sendable.go]
fixed-by: mesh-controller sendable.go and push.go — an empty declaration is sent carrying `owns_nothing`, and the host refuses an empty body that does not carry it
--- ---
# A declaration that shrinks to empty is skipped, so the node keeps what it should drop # 149 — a declaration that shrinks to empty is skipped, so the node keeps what it should drop
*Opened as 127 and renumbered on 2026-09-29: two records were given that number on the same day.*
*The other kept it, because three documents and three source files cite it by number and nothing
cited this one but a decision and a sibling issue, both corrected with this move.*
## What was observed ## What was observed
@@ -37,3 +42,17 @@ mean "own nothing", which the host already applies correctly when it receives on
On ace, one command drops it permanently (the corrected controller never re-composes it): On ace, one command drops it permanently (the corrected controller never re-composes it):
`sudo ufw delete allow 5671`. At ace's converge it would clear on its own. `sudo ufw delete allow 5671`. At ace's converge it would clear on its own.
## Closed
*2026-09-29, in a grooming pass.* Both halves are on `main` and both name this issue.
- The control plane **sends** it: a declaration that composes to no resources goes out with
`owns_nothing`, and `push` says *sent, not skipped*.
- The host **refuses an empty body that does not carry it**, so a truncated or mis-composed
declaration can never be read as "own nothing" — which is the failure the fix had to avoid while
making the empty case expressible.
Closed by reading the code rather than by watching a machine let go of a stray resource; the record
says so rather than implying a run.
@@ -0,0 +1,52 @@
---
status: open
opened: 2026-09-29
located-in:
- mesh-controller internal/catalogue/declaration.go (contributions are emitted for an assigned module, taken or not)
fixed-by:
---
# 150 — A route is contributed before its module is taken
## What was observed
ace is adopted and runs `route-adapter` beside the predecessor's traefik (ADR 0104). The first web
module migrated there was searxng. `assign ace searxng` + `push ace` — the step that is supposed to
change nothing on an adopted machine (ADR 0100, hq 125: assign holds, take replaces) — took
`searxng.zurag.be` down:
```
https://searxng.zurag.be/ → 502
```
for about five minutes, until the module was unassigned again.
## Why
Assign held everything it found on the machine — the predecessor's `searxng` container, its
directories — exactly as designed. But the module's **route contribution** is not a resource on the
machine, so nothing held it: it reached `route-adapter` at once, which wrote
`mesh-searxng.zurag.be.yml` into traefik's file provider pointing at the mesh's assigned machine port
(`http://ace.internal:20000`) — where nothing listened, because the container that would was held.
traefik's file router for the name then won over the predecessor's docker-label router for the same
name, and the name served a dead backend.
On this occasion the window was lengthened by an unrelated failure (the machine's docker address
pools were exhausted, so the module's network could not be created), but the fault does not depend on
it: **between assign and take, every routed module's public name points at a backend the mesh has
deliberately not started.** The runbook's §6 step 2 ("check the preparation — HAL's service still
serves") is false for every routed module on a node running route-adapter.
## What the operator did
Rolled back per runbook (unassign before take), then re-ran as assign → push → wait for the node to
report its holds → take → push, keeping the window to the ~10 s between the two pushes. A workaround,
not a fix: a module whose data must move between assign and take (runbook §6 step 3) cannot shrink
its window this way.
## What would be right
A contribution from a module that is assigned but **not taken** on an adopted node should be held like
the module's resources are — withheld from the provider until take — or the provider should be told
the contributor is held and leave the predecessor's route alone. Either keeps "assign changes nothing"
true for routed modules.
@@ -0,0 +1,94 @@
---
status: open
opened: 2026-09-29
located-in:
- mesh-host internal/apply/apply.go (containerSpecReading hashes every `host` entry)
- mesh-controller internal/catalogue/declaration.go (withMeshNames gives every container the mesh's names)
fixed-by:
---
# 151 — A new name recreates every container in the mesh
## What was observed
Migrating one small module on ace (searxng) took four routine controller actions: `node public-domain
ace zurag.be`, `assign ace searxng` + push, `unassign ace searxng` + push (a rollback, see
[issue 150](../150-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take
+ push again. Each push to ace also pushed novox ("this push left g14, novox, shanks behind … sending it
too"). novox's host then **replaced every container it runs, twice**:
```
22:46 … updated postgres.server (mesh-store): replaced; a container's configuration is fixed when it is created
22:46 … updated distribution.store (mesh-registry): replaced; …
22:46 … updated route-proxy.server (route-proxy): recreated …
22:51 … updated postgres.server (mesh-store): replaced; …
22:52 … updated gitea.server (gitea): replaced; …
22:52–22:55 mesh-vault, builder, all of mailu, keycloak, minio, mongodb, invoicing, photos, umami, …
```
Replacements per minute on novox, 22:46–22:55: 12, 8, 13, 5, 6, 8, 8, 8, 8, 4. The control plane was
unreachable twice while its own store came back through crash recovery
(`FATAL: the database system is starting up`), and dependants (umami) crash-looped until it did. None
of the replaced containers belonged to the module being migrated, or to ace.
## Why (confirmed part)
Every container the mesh runs is given the mesh's names as `--add-host` entries
(`withMeshNames`), and the host puts every entry into the container's spec digest — deliberately,
since [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md): a container left alone when the roster moved kept a five-day-old
address and restarted 2286 times. A running container cannot have its hosts changed, so a changed
digest means a replace.
The consequence is that **the roster is part of every container everywhere**: anything that adds,
removes or moves one name — a node's public domain, a routed name, an unassign — replaces every
container on every machine that carries the list. On the hub that includes the control plane's store,
the registry, the edge and mail.
## Not yet established
Which name moved in each pass. gitea's hosts after the fact list the `.internal` names and novox's
internally routed `*.novox.be` names, and **not** `searxng.zurag.be` — so it is not simply "a routed
name was added". Two passes suggest the set changed at assign and changed back at unassign; the
declarations before and after would say, and nothing on the machine records the previous one.
## Why it matters now
The node-by-node migration adds names one module at a time. On ace alone that is ~25 web modules; if
each assignment moves the roster, each is a full restart of every hub service, and a rollback is
another. The migration is paused on this.
## What has since been ruled out as a cause (2026-09-30)
The roster was also moving for a reason that was not a name changing at all: a node whose plan would
not compose had its routed names silently dropped from the roster handed to every machine, so a
briefly unreachable store withdrew and restored a name on alternating passes. That is
[issue 152](../152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md), and it is
fixed — it is what made the churn on the control node repeat every six minutes for twenty-five
minutes rather than once per operator action.
**It does not close this record.** A roster that changes for a real reason — a machine joining, a
module assigned, a public domain set — still replaces every container on every machine that carries
the list. 152 removed the false reasons; the question below is still open.
## What would be right (for diagnosis)
Keep 135's guarantee (no container runs with a stale address) without making the roster part of every
container's identity — e.g. resolve mesh names through a resolver the container asks at lookup time
rather than baked entries, or scope each container's entries to the names it actually binds.
## Answered (2026-09-30): the first of those two
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) takes
the resolver, not the scoping. Scoping was the close call and was rejected for contradicting the
standing intent that anything on the mesh can call anything on it: it would make a name reachable only
where the mesh was told in advance it would be wanted, and it would leave the roster in the digest, so
the churn returns whenever a widely-bound name moves.
Resolving at lookup time makes staleness impossible rather than detected — 109 and 135 stop being bugs
that were fixed and become a shape that does not exist — and makes a name's blast radius nothing.
**This record stays open**, because the record answers it and the code does not. Nothing may stop
copying names until a container can reach the resolver from any of the runtime's networks
([issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)),
which it cannot on two of four machines today. Removing the copy first reintroduces 109 and 135
silently, on a live mesh, which is how both were found. The order is in the record.
@@ -0,0 +1,124 @@
---
status: resolved
opened: 2026-09-29
located-in:
- mesh-controller cmd/mesh-controller/plan.go (routeNamesInTheMesh skips a node whose plan will not compose)
fixed-by: mesh-controller 6c5dfd0 (PR 147)
amended-design:
---
# 152 — A node whose plan will not compose silently removes its names from every machine
## What was observed
For twenty-five minutes, and for seventeen of them after the last operator action, the control-node's
host applied all 327 of its resources every ~6.5 minutes without pause, replacing every container on
the machine each time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail,
and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine
sat near 8. Nothing was converging: each pass ended and the next began four seconds later.
The two machines carrying no containers were not churning. They were only knocked off the bus each
time the control node re-created it, reconnected, re-heard the same declaration and applied it again
as a no-op.
## Why: the roster alternates between two values, and it is part of every container
Two consecutive declarations were compared by reading the `--add-host` entries of four containers
the moment each pass created them:
```
23:02 ace.internal drive.novox.be g14.internal keycloak.novox.be novox.internal
office.novox.be portainer.novox.be shanks.internal umami.novox.be (9 names)
23:07 … the same nine, and searxng.zurag.be (10 names)
```
One routed name — belonging to a module on another machine entirely — leaves the roster and comes
back. Because the roster is part of every container's spec digest ([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)),
each flip is a different identity for every container on the machine, and a running container cannot
have its hosts changed. So every flip replaces all of them.
**What makes it flip is a swallowed error.** `routeNamesInTheMesh` composes every node's plan to
find the names it serves, and when one will not compose it moves on:
```
plan, settings, err := planFor(ctx, open, n.Name)
if err != nil {
continue
}
```
A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does
not know", but "the mesh states these names do not exist", to every machine at once.
## Why it sustains itself
The loop closes through the control plane's own database:
1. An apply replaces `mesh-store` — the store the control plane reads — by removing the container,
so postgres comes back through crash recovery.
2. While it recovers it refuses connections: `FATAL: the database system is not yet accepting
connections / Consistent recovery state has not been yet reached` (observed, 21:08:39 UTC, every
pass).
3. `planFor` for the other machine fails against that store. `routeNamesInTheMesh` swallows it and
drops its routed name.
4. The roster changed, so all 327 resources differ, so all are replaced — including `mesh-store`,
and including the bus, which is why the host also cannot report: `applied, and could not tell the
mesh: reporting: nats: connection closed`.
5. Back to 1.
Every pass destroys the evidence the next pass needs to decide it has nothing to do, and nothing
outside the machine has to be wrong for it to continue.
**It is metastable, not permanent.** It ran from 22:46 to 23:11 — five full replacements of every
container on the machine — and then stopped on its own, when one pass happened to read the store
during a window it was up, composed the same roster twice running, and found nothing to do. Load fell
from 7.8 to 1.7 and the machine returned to its five-minute idle tick.
That it ends by luck is the point, not a mitigation. The exit condition is a race the mesh does not
control, the operator cannot see, and nothing reports; the same four actions on a slower machine, or
a larger store, would not have found it. An outage that clears itself after twenty-five minutes and
five restarts of the forge, the directory and mail is not a smaller fault than one that does not — it
is the same fault, harder to catch.
## What it is not
- Not the operator's four actions on the other machine. Those explain the first passes
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)); they were
finished seventeen minutes and three full passes before these measurements.
- Not a file that keeps changing. `/etc/hosts`, the bus's account list and the vault's export were
hashed across passes and are byte-identical, and the directory reported `mode 755 to 700` every
pass while already being `700`. Those resources are **misreported as changed** and are worth their
own question, but they are not what moves a container's identity.
- Not the lost report alone. A report that cannot be delivered explains a re-apply; it does not
explain a re-apply that finds 327 differences.
## Why it matters beyond this outage
The same `continue` makes every routed name in the mesh conditional on every node's plan composing at
the moment any machine is pushed to. One unreachable or half-migrated machine is enough to withdraw
its names from everywhere — and the withdrawal is indistinguishable, on the receiving machine, from
the operator having removed them.
The codebase already states the rule this breaks, forty lines away, about the same kind of lookup:
> A lookup failure is an error, never "not found": collapsing the two composed a declaration without
> the trust whenever the inventory hiccuped, delivered by a push that reported success.
## How it was fixed, and how the fix is checked
`planFor` now marks the two failures that really are a statement about the node — its set not
composing, and a setting that reaches nothing — and the three gatherers pass over those and only
those. Every other failure is raised, naming the machine and the read.
Five tests hold it: a set that cannot compose is marked as the node's own; a store that cannot be
read is *not*; one incoherent node still does not cost the rest their names; a roster is never
returned beside an error; and the raised failure names what could not be read.
The two sibling gatherers were audited and fixed the same way — the grant composer, which would have
withheld a consumer's credential, and the private-network membership, which would have taken a
machine off the overlay. Three other `planFor` callers were audited and left alone: they refuse or
report rather than silently withdraw, which is the safe direction.
**This does not close [issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md).**
A roster that changes for a real reason still replaces every container in the mesh. This removes the
false reasons; whether the roster belongs in a container's identity at all is that record's question.
@@ -0,0 +1,57 @@
---
status: open
opened: 2026-09-29
located-in:
- mesh-controller internal/catalogue/dir_into.go (dirsFor: a stated path or <data root>/<module>/<id>, nothing else)
- mesh-controller (accesses: the path is the manifest's literal)
fixed-by:
amended-design:
---
# 153 — An adopted machine's data cannot be placed where it is
## What was observed
Preparing ace's media modules (plex, sonarr, radarr, lidarr, bazarr, nzbget, qbittorrent, bookshelf)
for migration. ace is adopted; its data is where the predecessor put it and **must stay there**:
- the library and download spool: `/storage/media/*`, `/storage/downloads` — a separate ZFS pool,
~40 TB, the operator's shared data (ADR 0051);
- plex's own state: `/mnt/plex/{config,data,temp}` — 133 GB on a second disk;
- large configuration directories held in place: lidarr 46 GB, radarr 17 GB, sonarr 3.2 GB.
The catalogue's manifests name `/services/media/*` (as `accesses`) and `/services/<m>/config` (as
owned directories), which is novox's layout, not ace's, and not a value a definition may carry.
## What was decided, and what exists
[ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md) (accepted) says
exactly what is needed:
> *Where* it is on the machine is the assignment's. A node has a default layout, and an assignment may
> place a directory elsewhere: on a second disk, or where an adopted machine's data already is.
> [ADR 0051]: an access keeps its shape and its semantics; its path moves from the definition to the
> assignment.
What the control plane implements (`dirsFor`): a directory is either a path the manifest states, or
`<data root>/<module>/<id>` under the node's one data root. There is **no per-assignment placement of
one directory**, and an `access` path is the manifest's literal — no setting reaches either.
## Consequence
Every module whose data an adopted machine already holds somewhere other than the default layout can
only be migrated by (a) writing the machine's path into the manifest — which 0112 forbids and which
is wrong on the next machine — or (b) moving the data into the placed layout in a window. (b) is
acceptable for a 40 MB configuration and impossible for a 40 TB library the operator has ruled must
never be moved, copied or re-owned.
The same gap covers ownership: the predecessor runs ace's media stack as `1001:2000`; a manifest's
`owner` is one value for every machine.
## What would be right
The two assignment halves 0112 decided: a setting that places a declared directory (by id) at a given
path on this node, and a setting that says where an access's data is — both validated like
`endpoints` (unknown ids refused), and an access placed by the assignment still never created,
chowned or removed.
@@ -0,0 +1,42 @@
---
status: open
opened: 2026-09-29
located-in:
- hq 02-DECISIONS/0138 (reach: internal | public | both)
- mesh-controller internal/catalogue/filtering.go (Reaches)
fixed-by:
amended-design:
---
# 154 — A machine's own network is not a reach
## What was observed
Preparing ace's modules. ace sits on a home network (192.168.1.0/24) behind a router, and several of
its services are reached **from that network by devices that will never be mesh machines**:
- mosquitto `1883` — an IoT light switch (`sonoff-office-light-switch`) and home-assistant;
- unifi `8080`/`3478 udp`/`10001 udp` — the access points' inform, STUN and discovery;
- plex `32400` — LAN streaming clients (three connected at survey time);
- home-assistant `8123`, and the resolver on the LAN address.
ADR 0138 gives an endpoint's reach as `internal` (the private overlay), `public` (anywhere) or `both`.
None of them says *this machine's own network*. The predecessor could: its unifi manifest opened
inform/STUN/discovery `from: 192.168.0.0/16, 10.0.0.0/8, 172.16.0.0/12`.
## Consequence
The only reach that includes a LAN device is `public`. While ace is adopted that is harmless — its
own firewall stays and admits the LAN — and behind NAT "anywhere" happens to mean the LAN. But:
- it states the wrong thing: an operator reading `reach: public` on an IoT broker believes it is on
the internet, and a router port-forward added later for something else makes it so;
- at `converge ace`, the mesh's filter is the sum of what it listens on (ADR 0045). An endpoint left
`internal` cuts every LAN device off at the flip; one set `public` opens it to the internet on any
machine with a public address.
## What would be right (for diagnosis)
A reach — or a source — that means the networks the machine is directly attached to (its uplink's
subnets, as the machine reports them), so a LAN-only service is declared as exactly that and the
filter can admit it without admitting the internet.
@@ -0,0 +1,58 @@
---
status: resolved
opened: 2026-09-29
located-in: [hq 00-META/checks/cycle.py]
fixed-by: hq f89aef9 (PR 192)
amended-design:
---
# 155 — Two records may share a number, and every check passes
## What was observed
On 2026-09-29 two machines opened issues against this repository within the same hour. Both read
`main` correctly and both took "the next free number", and they collided twice:
| | one machine opened | the other had already used |
|---|---|---|
| first | 147, 148 | 147, 148 on an unmerged branch |
| second | 149, 150 | 149 on an unmerged branch, 150 from renumbering the first collision |
The first collision was reconciled by hand before merging. The second was **merged into `main`**, and
`records.py`, `cycle.py` and `index.py` all reported success over a tree holding
`149-a-declaration-that-shrinks-to-empty` beside `149-an-adopted-machines-data-cannot-be-placed-where-it-is`,
and two folders numbered 150.
## Why
The number is allocated as `max(main) + 1`, and `main` lags every open pull request — seven of them
that evening. Two readers of the same `main` therefore compute the same next number, and neither is
doing anything wrong. The existing reconciliation precedent (a second record numbered 127 became 149)
assumed a single writer, which stopped being true when a second machine began filing its own findings.
## Why it matters
An issue number is how every other record cites this one — `fixed-by:`, `located-in:`, a decision
record's consequence, a commit message. Two records answering to one number is a citation that
resolves to whichever folder the reader happened to open, and the failure is silent on both sides:
the citer is not wrong, and the cited record exists.
It is also exactly the class this repository says it does not permit — a rule (`00-META/process/03-issues.md`:
"take the next free number") enforced by nothing.
## How it was fixed, and how the fix is checked
`cycle.py` now refuses a tree in which two issue folders share a leading number, and names both.
Proven by adding a duplicate and watching it fail, then removing it and watching it pass.
The colliding records were renumbered 153 and 154, in the branch that landed last — renumbering a
branch whose author is still pushing only moves the race.
**The check catches the collision; it does not prevent it.** Allocating a number still needs the open
pull requests read as well as `main`. That is a habit the check now backstops rather than one it
replaces, and playbook [03](../../00-META/process/03-issues.md) now says so at the step where the
number is taken.
[ADR 0080](../../02-DECISIONS/0080-the-development-cycle-is-checked.md) enumerates what `cycle.py`
enforces and named four things; it carries a progressive insight naming the fifth. The decision
stands — this is one more thing frontmatter and file names can carry, found by its absence.
@@ -0,0 +1,98 @@
---
status: resolved
opened: 2026-09-29
located-in:
- mesh-controller internal/broker/jetstream.go (EnsureConsumer)
fixed-by: mesh-controller e7da39d (PR 148)
amended-design:
---
# 156 — Moving a consumer's delivery subject stops the control plane, and only on a mesh that is running
## What was observed
The control plane crash-looped, every restart ending the same way:
```
mesh-controller: asserting how novox hears its declaration:
bringing consumer novox on NODES to match: nats: consumer name already in use
```
It came up on the first build of the controller in eight hours. The machines kept running what they
already held — this stops the mesh being *changed*, not the services it placed — and nothing could be
pushed, no report was consumed and no enrolment answered, for as long as it lasted.
## Why
[Issue 146](../146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md) put the
stream into a push consumer's delivery subject, because one process holding two consumers of the same
name on two streams was given one subject and acted on every message twice.
**The server will not move a push consumer's delivery subject while a subscriber is bound to it.** It
refuses with `consumer name already in use` — a message about the name, for a conflict about the
subject, which is why the trail starts in the wrong place.
A node is bound to its declaration consumer the whole time it is up. That *is* a node listening for
what it should be. So every node consumer in a mesh that is running is one the assertion cannot bring
to match — and the assertion happens before the controller serves, so it never serves.
The controller's own two consumers moved without trouble, and are on the new subject in the live mesh.
It asserts them before it subscribes, so nothing was bound.
## Why nothing caught it
The change was exercised on a mesh being raised, where every consumer is created rather than updated
and nothing is bound to any of them. On that path the code is correct. The test that would have caught
it needs a mesh that is already running: an existing consumer, a subscriber still attached, and then
the assertion.
Reproduced exactly that way before the fix — same server version, same stream shape, same consumer —
and it fails with the same words as the machine did. An earlier version of the same test unsubscribed
first and passed against the code that was crash-looping on the control node.
## What it is not
- Not the change it shipped beside. The merge that triggered this build carried
[issue 152](../152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md)'s fix and three
other commits that had never been deployed; this one is 146's.
- Not a version difference. The test server and the mesh's broker are both nats-server v2.10.29.
## How it was fixed
The consumer that works is kept, and the assertion says so instead of failing.
**Not deleted and re-made.** Re-making moves the subject, and a holder may not yet be allowed to
subscribe to the new one: the wider grant travels in the bus's user list, which the control plane
composes and a machine applies minutes later. On the live mesh the nodes are granted `_DELIVER.<node>`
and not `_DELIVER.<node>.>` — re-making would have silenced every machine in the mesh, which is worse
than the collision it was fixing and far harder to undo. That was the first fix written here, and the
permission is the reason it was not shipped.
**Not fatal**, which is what 146's change intended and did not do: the bare subject still delivers, and
collides only where one holder has two consumers of one name. A node has one.
## How the fix is checked
Two tests against a real server: a consumer with a subscriber bound keeps its subject, is reported,
and still delivers to that subscriber; a consumer with nothing bound moves, so 146's fix still applies
where the collision actually was.
## What is left
The node consumers stay on the bare subject, which is correct and not tidy. Nothing is wrong while
they do not move: one consumer per name per stream cannot collide with itself.
The controller reports each one it kept, and did, on the start that fixed this — four node consumers
and the build machine's worker, which is bound the same way and was not anticipated here:
```
consumer novox on NODES still delivers to "_DELIVER.novox" and not "_DELIVER.novox.NODES":
nats: consumer name already in use. It keeps working; the subject moves on an assertion
made while nothing is bound to it
```
**The wider grant has since landed** (2026-09-30, measured on the mesh's own broker config): every
node is now allowed `_DELIVER.<node>.>` as well as the bare subject. That was the thing missing when
this was diagnosed, and it is why re-making the consumers then would have silenced every machine.
What remains is only the second half — an assertion made while each node is detached from its
consumer — and that is its own piece of work, not a side effect of a restart.