Merge main: renumber this branch's records around the trunk's

Both lines of work numbered from the same point, so four decision records and one design
document existed twice with different content. The trunk keeps its numbers and this branch
yields — the only rule that scales, because the trunk's are already cited by what merged
before them.

  0117 the bus is the only broker        -> 0125
  0118 a module declares its own seats   -> 0126
  0119 amqp is a provision, not the bus  -> 0127
  0120 the mesh bus is required          -> 0128
  0123 a seat carries its role's protocol -> 0129
  0124 the predecessor is ending          -> 0130
  design 29, what a module declares       -> design 32

Applied to the code repositories too, because a stale reference is worse when numbers
collide than when they dangle: the reader lands on a real record that decided something
else.

Two reconciliations the merge forced, both real:

**0110 was marked wholly superseded and was not.** Its successor says in as many words that
everything 0110 decided about what a seat *is* stands untouched — and two records that
landed on the trunk rest on exactly that part. So it is accepted again, extended rather than
replaced, with a note saying which of its claims moved and where.

**A seat's protocol becomes columns, not fields.** The trunk moved the seat set out of
compiled code into a table the controller owns. This branch had added what a role accepts,
emits and serves to the Go slice. The decision is unaffected and the mechanism is better for
it: giving a role a protocol is now a write rather than a rebuild, which is the trunk's own
argument applied to what this branch added.

One check still fails and it fails on main too: a record resting on ADR 0112 while that is
still 'proposed'. Left alone — it is not this merge's to answer.
This commit is contained in:
2026-09-27 18:23:41 +02:00
36 changed files with 1489 additions and 104 deletions
@@ -1,8 +1,8 @@
---
status: located
status: fixed
opened: 2026-09-24
located-in: [mesh-controller internal/inventory, mesh-controller internal/catalogue, mesh-controller cmd/mesh-controller, mesh-catalog modules/dnsmasq]
fixed-by:
fixed-by: [mesh-controller#73 overlay-name + namesInTheMesh, mesh-catalog#106 dnsmasq daemon.json merge]
amended-design:
---
@@ -0,0 +1,39 @@
---
status: located
opened: 2026-09-27
located-in: [mesh-controller cmd/mesh-controller/push.go]
---
# A declaration that shrinks to empty is skipped, so the node keeps what it should drop
## What was observed
Fixing the broker-opening leak (the foundation port scoped to the broker's host) made ace's
declaration compose to **zero resources** — ace is adopted with nothing assigned, and the
stray opening was its only resource. `push ace` then printed `ace is assigned nothing —
skipped` and sent nothing. ace goes on holding `adoption.opening-tcp-5671-incoming` in its
ufw, because it was never told the resource is gone.
`composeEach` (push.go) skips any node whose composed declaration has no resources. That is
right for a node that never had anything. It is wrong for a node that **had** resources and
now composes to none: the empty declaration is the correction, and skipping it leaves the last
non-empty one in force forever.
## Why it matters
Any adopted node whose openings (or other baseline resources) are all removed keeps the stale
ones until something else pushes a non-empty declaration to it. Converge is unaffected — it
composes the full ruleset fresh — so this is an incremental-push gap, not a firewall-safety
one. But "the mesh cannot tell a node to drop its last resource" is a real hole in reconcile.
## The fix, roughly
Send the empty declaration when the node's last-sent declaration was non-empty — i.e. skip
only when empty-and-was-already-empty. Requires push to know (or the host to be told) that the
node held something. Simplest: always send to a placed, enrolled node; let an empty declaration
mean "own nothing", which the host already applies correctly when it receives one.
## Workaround used
On ace, one command drops it permanently (the corrected controller never re-composes it):
`sudo ufw delete allow 5671`. At ace's converge it would clear on its own.
@@ -3,14 +3,14 @@ status: resolved
opened: 2026-09-27
located-in: [mesh-catalog modules, mesh-control internal/catalogue, mesh-tools src]
fixed-by: mesh-catalog 7b06a7a, mesh-tools fbeb373, mesh-control 05ff606
amended-design: 03-DESIGN/01-to-be/29-what-a-module-declares.md
amended-design: 03-DESIGN/01-to-be/32-what-a-module-declares.md
---
# 127 — A module's event derives a subject nothing publishes
## What was observed
[Design 29](../../03-DESIGN/01-to-be/29-what-a-module-declares.md) §1 says a module names an event
[Design 29](../../03-DESIGN/01-to-be/32-what-a-module-declares.md) §1 says a module names an event
locally and the mesh derives the subject: `emits: order.placed` becomes
`mesh.mod.<module>.event.order.placed`, and a consumer declaring `consumes: shop.order.placed`
subscribes the emitter's own subject. That derivation is built and tested.
@@ -37,7 +37,7 @@ Two further consequences of the same cause, found in the same check:
and the module cannot be assigned.
- One module emits under a name that is not its own — it declares `module.<other>.image.pushed`
while being a differently named module — which the derivation puts inside *its* namespace. Whether
that is legitimate is a design question: design 29 §2 makes an event's source a fact the server
that is legitimate is a design question: design 32 §2 makes an event's source a fact the server
enforces, and this is a module claiming another's name in its own event.
None of it fails on the bus the mesh runs on today, where a routing key is matched literally and
@@ -47,7 +47,7 @@ the first mesh raised on the new bus, and not before.
Evidence: run against the controller's own `PermissionsFor` on the current feature branch, with the
declarations read from the catalogue's manifests. Found while wiring the controller's consume side
(design 28 step 3.4), when the controller's own subscription had to be written and the subject it
would have to name turned out not to be the one design 29 specifies.
would have to name turned out not to be the one design 32 specifies.
## Why it matters beyond this instance
@@ -6,7 +6,7 @@ Three places, and only one of them is a bug in code.
**The manifests, in the module catalogue.** Thirty-seven declare events, and every one of them
spells an event the way a routing key on the bus the mesh runs on today is spelled —
`module.<module>.<verb>`. [Design 29](../../03-DESIGN/01-to-be/29-what-a-module-declares.md) §1 says
`module.<module>.<verb>`. [Design 29](../../03-DESIGN/01-to-be/32-what-a-module-declares.md) §1 says
a module names an event **locally and bare** (`emits: order.placed`) and a consumer names
`<emitter>.<event>` (`consumes: billing.order.placed`). So the manifests are stale against a rule
that was already decided, not wrong against an undecided one. **This is the whole of the reported
@@ -27,7 +27,7 @@ first thing that notices is a subscription that never fires.
## What was ruled out
**The derivation is not wrong.** Asked directly, with the module names and declarations the
catalogue holds, `PermissionsFor` produces exactly what design 29 §1 specifies for the input it is
catalogue holds, `PermissionsFor` produces exactly what design 32 §1 specifies for the input it is
given: it reads a consumer's `<emitter>.<event>` and builds the emitter's subject. Given
`module.builder.built` it reads the emitter as `module`, which is a correct reading of an incorrect
declaration.
@@ -0,0 +1,73 @@
---
status: located
opened: 2026-09-26
located-in: [mesh-controller internal/catalogue/facts.go, mesh-host internal/apply]
---
# 128 — the machine's hosts file is written whole, and on a workstation it is shared
## What was observed
The private network asks for the `node-names` fact, and the mesh delivers it as
`/etc/hosts`. `nodeNames` composes a **complete** file — its own header, `localhost`, the
machine's own name, and every name in the mesh — and the host writes it over whatever is there.
On an adopted workstation the file the mesh holds contains, besides the predecessor's block of
mesh names:
- the distribution's own lines (`localhost`, the machine's `.localdomain` name);
- two marked blocks (`# BEGIN … # END …`) maintained by a local-development tool, pointing a
dozen development hostnames at `127.0.0.1` — rewritten by that tool whenever its project
list changes;
- hand-added entries of the operator's.
Today the file is only **held** ([ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)):
the private network was assigned, not yet taken, so nothing was lost. Taking it — or
converging the node, which takes everything — replaces the file. The development tool's
entries disappear, its projects stop resolving, and every later write it makes is overwritten
at the next change to the mesh's names (a machine joins, a route is contributed), silently and
without a failure anywhere: the development tool thinks it wrote its block, and the mesh thinks
it owns the file.
This is [ADR 0102](../../02-DECISIONS/0102-the-mesh-writes-into-a-shared-file-never-over-it.md)'s
failure exactly — a file the mesh shares with software it did not install, written over — in a
file 0102 did not name, because its merge verb is structured (`into: json`) and a hosts file is
not JSON.
A second, smaller finding from the same reading: the fact's contents depend on which machines
hold the private network. A machine that is enrolled but not yet assigned the private network
is in neither `node-names` nor `node-zones`; its name resolves on the others only for as long
as a predecessor's hosts block survives. Taking the hosts file before every machine is on the
private network loses that name too.
## What would have prevented it
- A **marked-region** merge in the host's vocabulary: `into: "block"` (or similar) — the host
owns only the lines between its own begin and end markers, keeps everything outside them
byte for byte, records what the region held before, and on undeclare removes the region and
nothing else. The shape local tools already use for this very file.
- The `node-names` fact written as that region — no header of its own, no `localhost`, no
machine name — so the distribution's lines and every other tool's stay where they are.
- A converge preview that names a held file the take would replace *whole*, with its line
count before and after, so a person sees "hosts: 31 lines → 12" before the flip.
## The fix, as built (in review)
- **Host:** a file resource may say `"into": "block"`. The host owns only the lines between
`# BEGIN mesh <id>` and `# END mesh <id>` and keeps everything outside them byte for byte. A
new region goes at the `end` by default, or at the `start` (`"at": "start"`) for files where a
line's meaning depends on what stands above it; a region already present is never moved.
Undeclared, what the region held before is put back, or the region is removed and nothing
else. Replacing nothing, it is written on an adopted node without being held — so a machine
gets the mesh's names before its private network is taken.
- **Controller:** `node-names` is a fact written into a shared file, emitted as that region: the
mesh's names only, no header, no `localhost`, no `127.0.1.1` line.
- **Order:** a host older than the block mode refuses the whole declaration on an unknown
`into`, so hosts are upgraded before the controller that emits it.
## Evidence to carry into diagnosis
- `internal/catalogue/facts.go`, `nodeNames`: the complete file is built here.
- The host's file resource supports `into: "json"` only; anything else is a whole write.
- `node show <node>` on the adopted workstation: `holds file /etc/hosts
mesh-wireguard.fact-node-names`, original kept.
@@ -0,0 +1,62 @@
---
status: open
opened: 2026-09-26
located-in: [mesh-controller, mesh-catalog step-ca]
---
# 129 — nothing makes a machine trust the mesh's own certificate authority
## What was observed
On an enrolled, adopted workstation — on the private network, resolving the mesh's names
through the mesh's resolver — every HTTPS name the mesh serves internally fails verification:
```
curl https://<a name the mesh routes internally>/
curl: (60) SSL certificate OpenSSL verify result: unable to get local issuer certificate (20)
```
The route proxy presents a certificate issued by the mesh's internal authority (step-ca, the
`internal-acme-ca` provision). The machine's trust store holds the **predecessor's** authority
and a developer tool's local root, and nothing of the mesh's. No module installs the mesh's
root, and no fact carries it: step-ca's only consumers are proxies, which obtain certificates
over ACME and never need the root on the machine they run on.
[Issue 048](../048-nothing-makes-a-machine-trust-the-mesh-registry/00-report.md) found the same
shape for the mesh's registry and resolved it by treating the private network as the transport
security ([ADR 0082](../../02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)):
the runtime pulls in the clear, over the tunnel. That answer does not carry over. A browser, git
over HTTPS, a package manager and every TLS client a person or a module uses verify the
certificate chain, and there is no "insecure registries" for them — nor should there be.
Consequences today, all silent until someone tries:
- a person on a workstation cannot open any internal HTTPS name without a warning;
- git over HTTPS to the mesh's forge fails, so the working clone URL is ssh-only;
- a module on a non-hub machine that calls another module's internal HTTPS name fails
verification unless its image happens to carry the root;
- the predecessor's authority cannot be retired from any machine while anything there still
speaks TLS to a mesh name, because it is the only authority those machines trust.
## What would have prevented it
- A **mesh fact carrying the internal authority's root** (public material; the controller or
the step-ca module is its source), written onto every machine on the private network — the
same reasoning that has the private network write the registry trust and the names: being on
the network is what makes a machine one that speaks to the mesh's names.
- A resource that puts it where the machine's TLS clients look — on Arch,
`/etc/ca-certificates/trust-source/anchors/` — and **refreshes the extracted bundles**
(`update-ca-trust`). The refresh is the open design question: it is a command, and the link
may not carry an action ([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)). A declared
one-shot unit, or a host primitive for "trust this anchor", are the obvious candidates.
- Removal symmetric to arrival: undeclared, the anchor goes and the bundles are refreshed again,
so a machine leaving the mesh stops trusting it.
## Evidence to carry into diagnosis
- `step-ca` module: provides `acme-ca` / `internal-acme-ca`, listens on 9000 for proxies; no
resource writes its root anywhere but its own state directory.
- The private network's generator writes `/etc/hosts` and the registry trust, and nothing
about certificates.
- On the workstation, the trust anchors present are the predecessor's authority and a local
development root; `trust list` shows no entry for the mesh.
@@ -0,0 +1,66 @@
---
status: located
opened: 2026-09-27
located-in: [mesh-host internal/apply/apply.go (remove)]
amended-design: 02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md
---
# 130 — undeclaring a service stops it, even one the mesh only reloads or only keeps running
## What was observed
Reviewing the uplink modules ([ADR 0125](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md))
found that the host's `remove` path stops every `service` resource that is no longer declared:
`SetServiceState(..., "stopped")`, reported as "stopped; the unit file is not the host's to
delete". `store.Orphans` matches by id alone. So any of these stops the unit:
- the module is unassigned — by mistake, or to switch it for another;
- the node is sent a deliberately-empty declaration ([issue 127](../127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md));
- a later catalogue version renames the resource's `id`.
That is right for a service the mesh brought into being. It is wrong for a unit the mesh
declares only to act on — and the catalogue already has two:
- **The private network declares `docker.service`** (`registry-trust-reload`, state `running`)
so that a change to the registry trust reloads the runtime ([ADR 0102](../../02-DECISIONS/0102-the-mesh-writes-into-a-shared-file-never-over-it.md)).
Unassigning the private network stops the container runtime, and every container on the
machine with it — including ones the mesh does not manage.
- **The sshd module declares `sshd.service`.** Unassigning it stops the machine's ssh daemon:
the lockout the same module's `listens` rule says a firewall must never arrange.
The uplink modules would have added a third and a fourth: unassigning the network manager's
module would have stopped the network manager, taking the machine off the only link the mesh
reaches it by.
## What would have prevented it
- A service resource that says the unit's **lifecycle is the machine's**: declared with no
`state`, the mesh never starts, stops, enables or disables it; it only reloads or restarts a
*running* unit when a trigger changes; undeclared, it is left exactly as it is. (Being built
on mesh-host `feat/a-file-written-into-a-marked-block` for the uplink modules.)
- Then: `registry-trust-reload` declared that way (the runtime is the machine's), and the sshd
module's service too — a machine's ssh daemon outlives any module that configures it.
- A plan or unassign preview that names every unit an undeclare will stop, so the consequence
is read before it happens.
## Resolution
[ADR 0126](../../02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md):
undeclaring removes what the mesh made and gives back what it changed. The host records the state
it first found a unit in, and undeclaring returns the unit to it — a unit found running (the
container runtime, sshd, a network manager) is left running; one the mesh started (the packet
filter a converge loaded) is stopped again; nothing is started on the way out; a record from
before the host kept what it found leaves the unit alone. That covers the runtime, sshd and the
uplink modules at once, without each module opting out; the private network and the sshd module
need no change.
A first draft — never stop a unit the mesh did not create — was rejected while implementing it:
returning a converged node to adopted unloads the mesh's filter by exactly this path.
Found on the way: an undeclared `process` failed every apply on its node (`remove` had no case
for it). Now removed with its unit, timer and bundle — the mesh's own code. `user` and `archive`
have the same gap and are left for their own decisions: removing a login or unpacked files is not
something to settle in passing.
The unassign preview is partly answered — the host's plan names each unit it will stop — and the
controller's side is left open.