Merge remote-tracking branch 'origin/main' into decision/docker-module

This commit is contained in:
2026-10-02 12:23:27 +02:00
33 changed files with 1103 additions and 50 deletions
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-09-22
located-in: [mesh-controller internal/overlay, mesh-host internal/apply]
fixed-by:
fixed-by: ADR 0102 (mesh-controller internal/overlay: the runtime file written into, reloaded), issue 128 (the hosts file as a region)
amended-design: 03-DESIGN/01-to-be/05-the-node-host.md
---
@@ -56,3 +56,12 @@ and nothing checks for it today.
- Should an adopted node that cannot trust the registry be refused a module that needs to pull?
Or should the refusal come earlier, when the node joins?
## Resolved, 2026-10-02
The runtime's file is written into and the runtime reloaded, never restarted
([ADR 0102](../../02-DECISIONS/0102-the-mesh-writes-into-a-shared-file-never-over-it.md), the
diagnosis above); the hosts file is a marked region the mesh owns alone
([issue 128](../128-the-hosts-file-is-written-whole/00-report.md)). Neither whole file remains. Read
into [ADR 0168](../../02-DECISIONS/0168-a-converged-machine-is-filtered-by-the-mesh-alone.md), rule 5,
which names the one whole machine-wide file the mesh still writes — its own filter at the
distribution's path — as a difference a take shows, not a fault.
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-09-22
located-in: [mesh-controller cmd/mesh-controller/adoption.go (take previews nothing), mesh-host internal/apply (the comparison and the record)]
fixed-by:
fixed-by: mesh-controller 201 (the preview names the narrowing and the port's reach), 206 (`take --yes <digest>` acts on the preview read)
amended-design:
---
@@ -40,3 +40,13 @@ it changes before it changes it, and for taking a module this one does not.
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rules 1 and 2: the preview names a narrowing. Building follows,
host first, then the controller's `take`.
## Built, 2026-10-02
mesh-controller 201 and the pull request after it: the preview names it, and `take --yes <digest>`
acts on the preview that was read. Stays located until a take is read on an adopted machine — every
machine of this mesh is converged today, so the record's live row has not been run.
## Resolved, 2026-10-02
Closed on the operator's decision of 2026-10-02 with the built and tested code live on every machine (mesh-controller 206, mesh-host 64), not on a take read on an adopted machine: every machine of this mesh is converged, so none holds a found thing to compare, and the record's live row — ADR 0163's last — will be read at the next real adoption rather than staged. Said here so nobody later believes that row was run.
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-09-22
located-in: [mesh-controller cmd/mesh-controller/adoption.go (take previews nothing), mesh-host internal/apply (the comparison and the record)]
fixed-by:
fixed-by: mesh-host 64 (genesis raises the forge as `gitea`, on the module's image digest, with the module's data directory at /data; a test holds the installer to the module's manifest)
amended-design:
---
@@ -53,3 +53,17 @@ network, or it is not a takeover.
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rule 7: genesis raises as the module declares. Building follows,
host first, then the controller's `take`.
## Built in part, 2026-10-02
mesh-host 64: genesis raises the forge under the module's container name (`gitea`), pinned to the
module's image digest, with the module's data directory mounted at `/data` — so the module finds it,
holds it, and a take compares equal images and the same data. A test holds the installer's constants
to the module's manifest where the catalogue is checked out beside it. The network is the difference
left: the bootstrap forge runs on the machine's network to reach the store on its loopback, the module
runs bridged and publishes its ports, and the take says so. Closing waits for group 9's genesis test —
a mesh raised, the module assigned, and the module found holding rather than raising a second forge.
## Resolved, 2026-10-02
Closed on the operator's decision of 2026-10-02. Name, image and data directory align; the network does not — the bootstrap forge runs on the machine's network to reach the store on its loopback, the module runs bridged — and a take says so rather than hides it. Whether genesis should move the forge onto a bridge, and the test that raises a mesh and finds the module holding rather than raising a second forge, belong to group 9's genesis work and are not owed by this record any more.
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-09-23
located-in: [mesh-controller cmd/mesh-controller/adoption.go (take previews nothing), mesh-host internal/apply (the comparison and the record)]
fixed-by:
fixed-by: mesh-controller (the pull request after 201: JudgeSettings, LeftOut), mesh-host 64 (left_out kept)
amended-design:
---
@@ -63,3 +63,12 @@ knowing the code.
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rule 6: judged where stored; an impossible statement costs a module. Building follows,
host first, then the controller's `take`.
## Resolved, 2026-10-02
One judgement, in the catalogue, run where a setting is stored and where a machine is composed. Stored,
a setting that cannot compose with the module's current definition is refused naming the node, the
module, the layer and the key; a key that reaches nothing is refused there too. Composed, a definition
that moved under a stored setting leaves that module out of the machine's declaration — the envelope
names it, the host keeps what it holds and wrote for it, `plan` and `push` say it — and the machine is
told everything else. A stray setting no longer refuses the whole machine where it is read.
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-09-23
located-in: [mesh-controller cmd/mesh-controller/adoption.go (take previews nothing), mesh-host internal/apply (the comparison and the record)]
fixed-by:
fixed-by: mesh-host 63 (former targets removed, strays reported), mesh-controller 201/202 (strays shown)
amended-design:
---
@@ -80,3 +80,11 @@ found, and so would be kept for ever on purpose.
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rule 5: former targets are removed and strays reported. Building follows,
host first, then the controller's `take`.
## Resolved, 2026-10-02
mesh-host 63: the host's record keeps a resource's former targets, removes a container or file it
wrote under a name the declaration no longer names, never what was found, and reports strays — what
runs on the machine that the mesh neither wrote nor holds. mesh-controller 201 and 202 show strays
on `node show` for an adopted and a converged machine alike; the live mesh reported four on the
control node the evening it rolled.
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-09-23
located-in: [mesh-controller cmd/mesh-controller/adoption.go (take previews nothing), mesh-host internal/apply (the comparison and the record)]
fixed-by:
fixed-by: mesh-host 63 (the kept original's difference), mesh-controller 201 (shown; a differing file refuses unless `--replace <path>`)
amended-design:
---
@@ -68,3 +68,13 @@ written.
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rules 1 and 2: the difference is shown and a differing file refuses. Building follows,
host first, then the controller's `take`.
## Built, 2026-10-02
mesh-host 63 reports the difference between the kept original and the declared content; mesh-controller
201 shows it in the preview and refuses a differing file unless `--replace <path>` names it, or the
module declares the file partially. Stays located until a take is read on an adopted machine.
## Resolved, 2026-10-02
Closed on the operator's decision of 2026-10-02 with the built and tested code live on every machine (mesh-controller 206, mesh-host 64), not on a take read on an adopted machine: every machine of this mesh is converged, so none holds a found thing to compare, and the record's live row — ADR 0163's last — will be read at the next real adoption rather than staged. Said here so nobody later believes that row was run.
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-09-23
located-in: [mesh-controller cmd/mesh-controller/adoption.go (take previews nothing), mesh-host internal/apply (the comparison and the record)]
fixed-by:
fixed-by: mesh-host 63 (both images' creation dates), mesh-controller 201 (DOWNGRADE said; refused unless `--downgrade`)
amended-design:
---
@@ -65,3 +65,13 @@ expected rate.
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rules 1 and 2: the images are compared by age and a downgrade refuses. Building follows,
host first, then the controller's `take`.
## Built, 2026-10-02
mesh-host 63 reports the found image and both images' creation dates; mesh-controller 201 says
DOWNGRADE and refuses unless `--downgrade` is said. Stays located until a take is read on an adopted
machine.
## Resolved, 2026-10-02
Closed on the operator's decision of 2026-10-02 with the built and tested code live on every machine (mesh-controller 206, mesh-host 64), not on a take read on an adopted machine: every machine of this mesh is converged, so none holds a found thing to compare, and the record's live row — ADR 0163's last — will be read at the next real adoption rather than staged. Said here so nobody later believes that row was run.
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-09-23
located-in: [mesh-controller cmd/mesh-controller/adoption.go (take previews nothing), mesh-host internal/apply (the comparison and the record)]
fixed-by:
fixed-by: mesh-controller 201 (`secret accept --provider` reaches a required secret), 206 (a module's secrets listed with origin; a minted one for found data refuses unless `--mint <name>`)
amended-design:
---
@@ -70,3 +70,15 @@ the module can only be installed fresh.
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rules 2 and 3: a minted secret for found data refuses; secret accept reaches required secrets. Building follows,
host first, then the controller's `take`.
## Built, 2026-10-02
`secret accept <node> <module> <name> --provider <node>` reaches a required secret (mesh-controller 201).
The pull request after it reads every secret a module holds on a machine with its origin, and a take
of a module whose data was found refuses a minted, unaccepted one — naming the accept that carries
the existing value in, or `--mint <name>` to let the service take the new one. Stays located until
a take is read on an adopted machine.
## Resolved, 2026-10-02
Closed on the operator's decision of 2026-10-02 with the built and tested code live on every machine (mesh-controller 206, mesh-host 64), not on a take read on an adopted machine: every machine of this mesh is converged, so none holds a found thing to compare, and the record's live row — ADR 0163's last — will be read at the next real adoption rather than staged. Said here so nobody later believes that row was run.
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-09-23
located-in: [mesh-controller cmd/mesh-controller/adoption.go (take previews nothing), mesh-host internal/apply (the comparison and the record)]
fixed-by:
fixed-by: mesh-controller 201 (the neighbours on a found network are named), 206 (the per-machine `networks` setting), mesh-host 64 (the taken container joins the kept network)
amended-design:
---
@@ -66,3 +66,14 @@ exercise.
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rules 1 and 4: the neighbours are named; a found network may be kept by a setting. Building follows,
host first, then the controller's `take`.
## Built, 2026-10-02
The preview names every neighbour on a found network (mesh-controller 201). The pull request after it
adds the per-machine setting `networks` — a container id to the found networks it keeps — judged for an
adopted machine only, and mesh-host 64 has the taken container join each once it runs. Stays located
until a take is read on an adopted machine.
## Resolved, 2026-10-02
Closed on the operator's decision of 2026-10-02 with the built and tested code live on every machine (mesh-controller 206, mesh-host 64), not on a take read on an adopted machine: every machine of this mesh is converged, so none holds a found thing to compare, and the record's live row — ADR 0163's last — will be read at the next real adoption rather than staged. Said here so nobody later believes that row was run.
@@ -1,7 +1,9 @@
---
status: located
status: resolved
opened: 2026-09-26
located-in: [mesh-host internal/apply]
fixed-by: mesh-host 63 (every written field compared), mesh-controller 201 (build says the policy)
amended-design:
---
# 126 — a volume path is not in the spec comparison, and a roll-out raced a data move
@@ -50,3 +52,9 @@ the install-page junk was discarded twice.
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rules 5 and 7: every field compared; build says the policy. Building follows,
host first, then the controller's `take`.
## Resolved, 2026-10-02
mesh-host 63: every field the host writes is compared before a container is called current, volumes
and paths included. mesh-controller 201: `build` and the take-in say when a module's policy rolls a
result out at once; under ADR 0162 the plan says it too.
@@ -1,10 +1,10 @@
---
status: located
status: resolved
opened: 2026-09-28
located-in:
- mesh-controller internal/catalogue/filtering.go
- mesh-host internal/apply
fixed-by:
fixed-by: ADR 0140 — mesh-controller (the filter around outward links; no network ranges anywhere)
amended-design: 03-DESIGN/01-to-be/08-connectivity.md
---
@@ -86,3 +86,11 @@ supersedes both 0137 and the first attempt at answering this.
runtime, or left as the one constant?
- Should the preview say which of a machine's networks are the mesh's and which are not, so a range
that exists to protect a leftover is visible as such?
## Resolved, 2026-10-02
By [ADR 0140](../../02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md), built and
live since 2026-09-29: the forward chain constrains what arrives on the machine's outward links and
says nothing about networks, so there is no list to derive and nothing for a preview to tell apart.
Read into [ADR 0168](../../02-DECISIONS/0168-a-converged-machine-is-filtered-by-the-mesh-alone.md),
rule 5, which closes it.
@@ -1,10 +1,10 @@
---
status: located
status: resolved
opened: 2026-09-29
located-in:
- mesh-host internal/apply/opening.go (retireFirewall)
- mesh-host internal/apply/apply.go (the condition it is called under)
fixed-by:
fixed-by: mesh-host 67 (retire on every converged apply; found-inactive apart from disabled-by-mesh; a skipped step said), mesh-controller 211 (the found firewall's state on node show)
amended-design:
---
@@ -100,3 +100,21 @@ harmless, but the mesh's belief about which firewall is in force has been wrong
so". Should it?
- Why do the host's own detail lines not reach the journal? Everything it decided during the flip is
unrecoverable, which is why this account has candidates instead of a cause.
## Decided, 2026-10-02
[ADR 0168](../../02-DECISIONS/0168-a-converged-machine-is-filtered-by-the-mesh-alone.md), rule 1:
convergence is a state the host keeps — the found firewall active again is retired again and said, a
reconcile that finds it inactive records *found so* and never *done by the mesh*, and a step skipped
after a failed apply is said. Built in mesh-host on `feat/one-thing-filters-a-converged-machine`; the
record of both machines of this mesh is corrected by the first report under it.
## Resolved, 2026-10-02
mesh-host 67 and mesh-controller 211, live on every machine at 10:10Z. The step now runs on every
converged apply and says what it did; a found firewall enabled again is retired again. The record's
one inherited lie stands as history: on the control node the machine's own record already said the
mesh had disabled the firewall, and the host trusts its record, so `node show` says "retired by the
mesh" there. From this build on, a reconcile that finds the firewall inactive records *found inactive*
and never the other thing. Whether the flip's step took on 2026-09-29 is not recoverable and is not
owed by this record any more.
@@ -1,10 +1,10 @@
---
status: located
status: resolved
opened: 2026-09-29
located-in:
- mesh-host internal/apply/opening.go
- mesh-controller cmd/mesh-controller (the converge preview)
fixed-by:
fixed-by: mesh-host 67 (every refusing table and legacy chain classified with an owner; the runtime's user chain is other), mesh-controller 211 (kept, shown on node show, named by status, previewed with fates)
amended-design:
---
@@ -82,3 +82,24 @@ everything reached from within.
- Is the bus and the registry being reachable from anywhere still what the mesh wants on a machine that
faces the internet? The design says yes, for enrolment. It deserves asking on its own rather than
being answered by a leftover.
## Decided, 2026-10-02
[ADR 0168](../../02-DECISIONS/0168-a-converged-machine-is-filtered-by-the-mesh-alone.md), rules 2 and
3: the host reports every table and legacy chain that refuses, with an owner, and the runtime's user
chain's refusals as *other*; `node show`, `status` and the converge preview say it. Built on
`feat/one-thing-filters-a-converged-machine` in mesh-host and mesh-controller. On 2026-10-02 the home
server still carries the predecessor's chain in its legacy filter; the record's live row is reading it
there.
## Resolved, 2026-10-02
mesh-host 67 and mesh-controller 211, live at 10:10Z. The live row of
[ADR 0168](../../02-DECISIONS/0168-a-converged-machine-is-filtered-by-the-mesh-alone.md) was read the
same hour: the home server's record names the predecessor's chain in the legacy filter's user chain
as *other*, with what it refuses, beside two chains a retired front end left in the IPv6 legacy filter;
the control node's record names the same two leftovers; the laptop and the workstation read *the mesh
alone*. `status` names both machines and is not well until the operator removes what the mesh did not
write. The allowance the predecessor's chain carried is
[issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)'s,
and that record is not closed by this one.
@@ -0,0 +1,80 @@
---
status: resolved
opened: 2026-10-01
located-in: [mesh-controller examples/route-proxy/main.go (routesFrom requires a route's public `name` and treats `internal-name` only as an alias of it; the handler serves every routed name to any source), mesh-controller internal/broker/membership.go (a membership says nothing of what its module receives or who the mesh is)]
fixed-by: mesh-controller PR 207 (the membership carries what a module receives and who the mesh is; the proxy follows it and serves internal names to the mesh only), mesh-catalog PR 211 (the proxy's bus account), mesh-controller PR 208 (the issue verb that delivers it), live 2026-10-02
amended-design: [03-DESIGN/01-to-be/08-connectivity.md, 03-DESIGN/01-to-be/25-the-bus-on-nats.md]
---
# 191 — A route with only an internal name is dropped as naming nothing
## What was observed
A module whose endpoint reaches only the private network could not be reached by its internal name.
The module ran and answered on its own port. Its route's internal name resolved to the serving node.
The request failed during the TLS handshake:
```
http: TLS handshake error from …: no public route for "unifi.home-server.internal" in this mesh,
so no certificate is asked for
```
The proxy's own log said why, every time it re-read its routes:
```
unifi on home-server asked for a route and named nothing; skipped
```
The route it skipped was not empty. The mesh had given it an endpoint, a port, a scheme and an internal
name, and no public name:
| route | `name` | `internal-name` | served |
|---|---|---|---|
| home-assistant | a public name | `home-assistant.home-server.internal` | under both |
| unifi | — | `unifi.home-server.internal` | under neither |
Three other modules on the same node were skipped with the same line on the same pass.
## Why it matters
**Reach is decided in one place, and the proxy reads the old shape of the decision.**
[ADR 0138](../../02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md)
made an endpoint's reach decide which names exist. The controller composes the public name, the
internal name, or both, and composes no name that nobody asked for
([issue 140](../140-an-endpoints-reach-is-not-declared/01-resolution.md)). An endpoint that reaches
only the private network is the ordinary case for anything that should not face the internet. It is
exactly the case the proxy drops.
**The failure is quiet and points the wrong way.** Nothing marks the module unhealthy. The handshake
error says *no public route*, which reads as a certificate fault on the reader's side. The line that
gives the real cause is one of four identical lines repeated every few seconds in a log nobody reads
until they already suspect the proxy.
**The opposite move is not a workaround.** Giving the endpoint public reach makes the proxy serve it.
It also publishes an administration interface to the internet to get a name on the private network.
## Open questions
- Should the proxy refuse a route it cannot serve in a way the controller or an operator sees, not
only in its own log? The same silent skip covers a route with no usable port or an unknown scheme.
- What checks that what the controller composes and what the proxy serves stay the same shape? ADR
0138 changed one side and nothing failed on the other.
## Resolved (2026-10-02)
Built as [ADR 0167](../../02-DECISIONS/0167-a-membership-carries-what-its-module-receives-and-who-the-mesh-is.md)
decided, and live on both machines that run the proxy. Each logs that its routes now come from its
membership, and serves internal names to the four machines the mesh names. Checked by hand:
- the internal-only route answers through the proxy from the serving machine and from two other
machines of the mesh, over a certificate from the mesh's own authority that each verifies;
- the same name asked from an address outside the mesh is answered as a name never routed, over plain
HTTP, and refused in the TLS handshake; the list of served names it is shown leaves out every internal
name.
Two things the rollout found are their own records: the proxy's bus account could be issued only from
the controller's command line, until mesh-controller PR 208 added the `issue` verb, and the status line
counting every module as a bus user without a credential is
[issue 195](../195-every-assigned-module-is-counted-as-a-bus-user-without-a-credential/00-report.md).
The serving machine also lacked the certificate-trust module, so it could not verify the mesh's own
certificates until it was assigned there.
@@ -0,0 +1,75 @@
# Diagnosis
*2026-10-01.*
**Ruled out first: the module itself.** Its container was up and had not restarted. The controller's
status endpoint answered on its own port with `"up": true`. The tool wrapper beside it was serving
all its tools.
**Ruled out: name resolution.** The internal name resolved to the serving node's private-network
address, which is where the proxy listens. Plain HTTP to the name reached the proxy and got a 404.
HTTPS failed in the handshake, and the proxy logged that it had no route for the name.
**The route as the proxy received it.** The mesh-written route file held a complete contribution
for the module: endpoint `web`, port, scheme `https`, `insecure`, a label, and `internal-name`. It had
no `name`. That is what the controller composes for an endpoint whose reach stops at the private
network (ADR 0138, `composeName`). The contribution was correct.
**Located: `routesFrom` in the proxy.** It reads `name` first and skips the contribution if `name`
is empty. It reads `internal-name` only at the end, as a second host for a rule that already has a
public one. So the proxy can serve an internal name only next to a public one. That matched the
mesh before ADR 0138, when both names were always composed. It has been wrong since then.
The other half of the proxy already handles the case. Certificates for a host are split by whether it
is in the public set: hosts outside it go to the internal authority, and only hosts inside it are
eligible for ACME. A host that is only ever an internal name falls on the correct side of both checks
without change. For certificates, only reading the route was wrong; who may reach the route is the next section.
**The fix.** `routesFrom` takes a route that names either host, serves each name it carries, and
marks only the public one as public. It still skips a route that names neither, with the same log line.
A test proves an internal-only route is served, certified by the internal authority, and refused by
the public one. That test fails against the code before the change.
## The first fix would have made the name public — 2026-10-02, from review
Serving the dropped route was not enough. The proxy picks a route from the name a request carries
and never from where the request came from, and it answers public and internal names on the same
listeners. Its public names resolve to an address the internet reaches. So once the internal-only
route was served, any request from the internet carrying `unifi.home-server.internal` — a name of a
fixed, guessable shape — would have reached an administration interface that reach `internal` was
chosen to keep private. Before the fix the route was unreachable from everywhere. After it, it would
have been reachable from everywhere. Two more leaks came with it: the proxy's answer for an unrouted
name listed every name it serves, internal ones included, and the handshake handed a certificate
naming the internal host to any client.
Nothing showed this while every routed endpoint also had a public name: its internal name exposed
nothing the public one did not. It is a gap in the decision's wording, not only in the proxy — ADR
0138 says the proxy *serves* the internal name without saying to whom — so it is recorded there as a
progressive insight and in the to-be connectivity design.
**Where "inside" is decided: told, not worked out.** The first correction had the proxy work it out
for itself — the mesh's range from an environment variable the catalogue wrote, and the machine's
container bridges from its own interfaces. That was a second definition of "the mesh", kept by one
module beside the one the controller already has: it resolves "from the mesh" to every machine's
address on the private network, and the packet filter is rendered from that list. Reviewed, it was
replaced: [ADR 0167](../../02-DECISIONS/0167-a-membership-carries-what-its-module-receives-and-who-the-mesh-is.md)
has every membership on the bus carry what its module receives and that list, and the proxy follows
its membership. One composition, read by the filter and by the proxy.
The proxy reads the source address, where the guard reads the interface, because it cannot see the
interface a request arrived on. A claimed source does not carry here: a connection needs its replies,
and replies to a mesh address leave by the tunnel.
**What changed with it.** The internal name of a route that also has a public one is now served to the
mesh only, like any other internal name. Outsiders have the public name, so nothing they could reach is
lost. A container calling its own machine's internal name arrives from its container network and is
refused; whether the mesh should issue those networks too is left open in ADR 0167.
**Order of release.**
1. The catalogue change, which gives the proxy a bus account. A machine running the proxy is not
composed until its account is issued, so the account is issued straight after
(`module issue route-proxy --node <machine>`), and then the machine is pushed.
2. The controller and proxy change. The push after it publishes memberships that carry the routes and
the mesh, and each proxy takes them. Until then, a proxy serves its file, and internal names to its
own machine alone.
@@ -0,0 +1,71 @@
---
status: resolved
opened: 2026-10-02
located-in: [mesh-host internal/apply (removeOrphan: a former target of a kind with no removal was fatal), mesh-host internal/store (Record keeps a former target for every kind, the host's own archive included)]
fixed-by: mesh-host 65 — a former target of a kind the host cannot remove is left in place, said and forgotten; a dropped archive still refuses
amended-design:
---
# 194 — The host's own former archive stops every machine applying anything
## What was observed
2026-10-02, 00:34Z, on all four machines of this mesh, the first time a host carrying former
targets ([ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rule 5,
built in mesh-host 63) replaced itself with a newer host (mesh-host 64).
The host delivers its own successor as an archive whose target is a versioned directory
([ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md)): every new version
is the same resource with a new target. Since mesh-host 63 the record keeps a resource's former
target so the next apply removes what the host wrote under it
([issue 097](../097-a-resource-that-changes-target-leaves-the-old-one-behind/00-report.md)). So
the new host's first apply found the previous version's directory as a former target of its own
archive, and asked the removal for an archive — which does not exist
([issue 162](../162-an-archive-cannot-be-undeclared/00-report.md)):
```
applying "mesh-host.next@former:/usr/lib/nox-mesh-host/versions/3c906749ad27": no way to remove a "archive"
0 resource(s) were applied and remain
```
Orphans are removed before any resource is applied on a converged machine, so the refusal ended
every apply at its first step. Every machine reported `failed`, applied nothing, and would have
gone on doing so: a host fix is itself an archive the same apply would have to write, and the apply
never reached it. The machines kept running what they had; nothing new from the mesh could land.
## Why it matters beyond this instance
Two rules that are each right met in the one resource the host cannot afford to stop on. Rule 5
says a former target is removed and said; issue 162 says an archive has no removal, deliberately,
so an unassignment nothing can undo is never reported as done. Neither rule was wrong; their
meeting was never tested, because the bed that would have found it is a host replacing itself
under the new rule, and the first such replacement was the live one. The fix is narrow: a former
target of a kind the host cannot remove is left in place, said, and forgotten — never fatal,
because nobody dropped it. An archive the declaration dropped still refuses, as 162 has it.
## What it took to recover
The broken host cannot apply its own fix: the fix is delivered as an archive, and the apply fails
before writing anything. On each machine the host's record (`/var/lib/mesh-host/state.json`) had to
lose the one `@former:` entry by hand, once, so that the next push could write the fixed archive and
stand aside for it. A manual edit of the host's record is otherwise never done; it is written here
because the alternative was four machines that could apply nothing.
## Open questions
- Should `Record` keep a former target for a kind the host cannot remove at all? The trace is
useful; the removal it implies is not. Keeping it and letting the apply forget it is what the fix
does; not recording it would be quieter.
- Should the host's own versions directory be cleaned by the launcher rather than by the apply —
the one archive whose former targets are genuinely removable, by the thing that knows which one
runs?
- Is there a bed that replaces a host under the current rules before the live mesh does
(the proof row of [ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md) was
a single crossover, before former targets existed)?
## Resolved, 2026-10-02
mesh-host 65, merged 07:35Z. Recovered as the record above says: the operator dropped the one
`@former:` entry from each machine's host record and pushed; the fixed host then ran on all four and
its first apply said `forgotten mesh-host.next@former:… a former target left in place` and applied the
rest. The open questions stand as questions for the host's own versions, not as faults.
@@ -0,0 +1,62 @@
---
status: open
opened: 2026-10-02
located-in: []
fixed-by:
amended-design:
---
# 195 — Every assigned module is counted as a bus user without a credential, and the real gaps are lost in the count
## What was observed
`status`, and `plan` for any machine, open with one line before anything else:
```
the bus's user list leaves out 49 user(s) the mesh has minted no credential for: <node>.<module>, …
Each is a user that cannot connect until one is issued
```
The 49 are spread over four machines and name 26 distinct modules. Checked against the catalogue on
2026-10-02:
| what the module's definition says | modules |
|---|---|
| declares an own secret named `broker` | 1 — the route proxy, which needed a bus account for issue 191 |
| declares no `broker` secret, and emits, consumes and serves nothing on the bus | 17 — the packet filter, the intrusion filter, the ssh daemon, the resolver configuration, the certificate authority, the broker itself and others |
| declares no `broker` secret, and **emits events** | 1 |
| not in this catalogue, so not checked | 7 |
So the line counts every module assigned anywhere as a bus user. For almost all of them that is not a
missing credential. A module with no `broker` secret has nowhere to receive one, and the mesh already
says an account nothing reads is an orphan ([issue 078](../078-a-delivered-secret-is-accepted-under-any-name/00-report.md)).
Two real gaps sit inside the count and cannot be told from the noise:
- **A declared `broker` secret was filled with a value that is not an account.** Before its account was
issued, the route proxy's plan on both machines already carried a sealed `broker` file, while the same
status line said no credential had been minted for it. A push had made the declared secret the way it
makes any own secret. The module would have started with a credential the bus does not know, and
nothing would have said why. It was found only because the account was being issued by hand.
- **A module that emits events declares no way to reach the bus.** Its events can go nowhere, and no
check refuses that.
## Why it matters
**A warning that is always on is read as never on.** The line names 49 users on every `status` and every
`plan`. An operator, or an agent, learns to scroll past it. The one entry that was a real fault looked
exactly like the 48 that were not.
**The fault that was real is the silent kind.** A module whose broker credential is a generated value
starts, fails to authenticate, and reports that three layers away from the cause. That is the failure
the composition already refuses for a secret that was never made at all ("declared and not made"). Here
a value was made, so the refusal never fired.
## Open questions
- Should a bus user be composed for a module that declares no `broker` secret at all? If not, the line
shrinks to the modules that can actually use an account.
- Is a `broker` secret ever correctly made by the generic generator? If not, should composition refuse
a declared `broker` until it is issued, or should the mesh issue it as part of placing the module?
- Should a module that emits, consumes or serves on the bus be refused when it declares no `broker`
secret?
@@ -0,0 +1,54 @@
---
status: resolved
opened: 2026-10-02
located-in: [mesh-controller internal/catalogue/filtering.go (AsNftables: the forward chain has no rule for the mesh passing through, so a relayed packet is judged by this machine's own published ports)]
fixed-by: mesh-controller PR 209 (the forward chain relays what comes in and goes out on the tunnel), live 2026-10-02
amended-design: []
---
# 196 — The hub relays the mesh only on the ports it publishes for itself
## What was observed
A sweep of every listening port on every machine, from every other machine, on 2026-10-02. Two home
machines, neither of which can be dialled, reach a third home machine through the hub, as
[ADR 0007](../../02-DECISIONS/0007-connectivity.md) says every path between machines that are not
co-located does.
From either of the two, the third answered on **17 of its 55** listening ports over the mesh. The hub
itself, probing the same machine directly, reached all the ports that machine's rules open to the mesh.
The result was the same at 40 probes in parallel and at 4, so it was not load.
The 17 were not a property of the target. They were exactly the ports **the hub** publishes for its own
containers: ssh, the proxy's two, and the hub's own block of published ports. A capture on the target
during one probe to a port that answered and one that did not:
- the answering one: the SYN arrives on the tunnel, reaches the container, and the reply leaves by the
tunnel;
- the other: nothing arrives at all, on any interface.
## Why it matters
**ADR 0007's hub carries every path between machines that are not co-located, and the filter breaks
that path without saying so.** Whether one home machine can reach a service on another depends on
whether the hub happens to publish the same port number for something of its own. Adding or removing
a module on the hub silently opens or closes paths between two other machines that it has nothing to
do with.
It also hid behind another fault. A missing placement made the same pair look disconnected earlier the
same day, and that explanation fit well enough that the per-port pattern was not looked for.
## Open questions
- The relaying rule accepts what comes in on the tunnel and leaves on it, and leaves judging to the
machine it is for. Should the hub also restrict relayed traffic to what that machine opens to the
mesh? That would duplicate the target's rules on the hub.
- No test raises two machines behind a hub and checks a port between them that the hub does not
publish. The lab's beds have one machine per site.
## Resolved (2026-10-02)
Live on all four machines after one push each. The same sweep, from both home machines to the third
over the mesh: 45 of 55 ports answer, the same 45 the hub reaches directly. The 9 that do not are
ports the target opens to nobody on the mesh, and one is refused because it listens only on a LAN
address. Nothing answers that the target's rules do not open.
@@ -0,0 +1,27 @@
# Diagnosis
*2026-10-02.*
**Not the tunnel.** The route from either home machine to the target is the tunnel, and traffic to the
17 ports travels it in both directions. A placement fault would have stopped every port.
**Not the target's filter.** The target opens the failing ports to every address of the mesh in its
input chain and its forward chain, the hub reaches them directly, and the SYN for a failing port never
arrived at the target to be judged.
**The hub's forward chain.** A relayed packet comes in on the tunnel and leaves on it, so the hub's
forward hook judges it. The chain the controller renders (`AsNftables`) has a default of drop, accepts
established traffic, and accepts what did not arrive on an outward link or the tunnel. That last rule is
for the machine's own containers reaching outward. After that come the rules for this machine's own
published ports, each matching the **original destination port** of the connection. None of them names
an outgoing interface or a destination. So a relayed packet to another machine's port 20000 matched the
hub's own rule for its own port 20000 and passed. One to port 8080, which the hub does not publish,
matched nothing and was dropped.
**The fix.** One rule: in on the tunnel **and** out on the tunnel is accepted. That is the mesh passing
through to another of its machines, which filters it against its own rules. It does not widen anything
on the hub. A packet for the hub itself is the input chain's, and one for the hub's own containers
leaves by a bridge, not the tunnel. Both still meet their rules. WireGuard only accepts a packet from a
peer whose address that peer is allowed to use, so in-on-the-tunnel means from a machine of the mesh.
A controller test asserts the rule in the forward chain only, never in the input chain, and absent on a
machine with no tunnel. It fails without the fix, and the rendered set loads with `nft -c`.
@@ -0,0 +1,43 @@
---
status: resolved
opened: 2026-10-02
located-in: [mesh-host internal/outward (Links reported only the links carrying a default route)]
fixed-by: mesh-host PR 66 (a link backed by a physical device is named outward, up or down), live 2026-10-02
amended-design: []
---
# 197 — A physical link that is down is not filtered when it comes up
## What was observed
A sweep of every machine's filter on 2026-10-02. A laptop-class machine connected by its radio has a
wired port that was unplugged. Its filter guarded the radio and the tunnel, and accepted everything
arriving on any other link:
```
iifname != { "mesh0", "<radio>" } accept
```
The wired port was not in the list. Plugged in, everything arriving on it would have been accepted,
every port of the machine open to whatever network the cable reached. That would last until the
machine reported again and was pushed a new filter.
## Why it matters
**The filter's one rule about links fails open.** [ADR 0140](../../02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md)
has the filter constrain what arrives from outside, and has the machine say which links face outside.
Everything not named is treated as the machine's own, its containers and bridges. So a link the machine
fails to name is not filtered at all. The host named only the links carrying a default route at the
moment it reported. A cable plugged in later is the ordinary case for a laptop. A second wired network
that never carries the default route, such as a direct link to a storage box, is never named at all.
## Open questions
- A virtual link that faces outside (a VPN client's interface, a USB tether that appears as a virtual
device) has no physical device behind it. It is named only while it carries the default route. Is
that enough?
## Resolved (2026-10-02)
Live on the affected machine after the host was delivered and one more push: its filter now guards the
radio, the tunnel and the unplugged wired port, before anything is plugged into it.
@@ -0,0 +1,14 @@
# Diagnosis
*2026-10-02.*
**Located in `mesh-host` `internal/outward`.** `Links` read the kernel's routing tables and returned the
interfaces carrying a default route. An unplugged port carries none, so it was never reported, and the
controller rendered the filter around the links it was given.
**The fix.** A link faces outside if it carries a default route **or** has a physical device behind it.
The kernel lists every interface under `/sys/class/net`, with a `device` entry for one backed by
hardware. A bridge, a veth, the tunnel and the loopback have none, so they stay the machine's own. The
wired port is now reported up or down, and the filter guards it before anything is plugged in. Tested
with a radio carrying the default route and an unplugged wired port beside a bridge, a veth, the docker
bridge, the tunnel and the loopback: the two physical links are reported, nothing else.
@@ -0,0 +1,71 @@
---
status: resolved
opened: 2026-10-02
located-in: [mesh-catalog modules/dnsmasq (listens on loopback and the machine's mesh address only), the home-server's DNS (a predecessor's dnsmasq configuration the mesh did not own), the home network's DHCP (hands out the home-server as every device's DNS)]
fixed-by: mesh-catalog PR 214 (dnsmasq listens on addresses from a setting; docker's file takes no settings), mesh-controller PR 210 (the settings verb), mesh-catalog PR 215 (unifi network DNS tools), live 2026-10-02
amended-design: []
---
# 198 — The home network's DNS server ran outside the mesh, and the mesh's filter closed it
## What was observed
Every phone on the home Wi-Fi had no internet, while a laptop on the same Wi-Fi did. The router's
DHCP hands every device the home-server's LAN address as its DNS server. The home-server's DNS daemon
was listening on that address, and every query to it timed out. The router itself answered the same
query at once. The laptop worked because it resolves through its own local resolver, not through the
server DHCP names.
## Why it happened
The DNS daemon on the home-server was not the mesh's. It ran under a configuration file a predecessor
generated, listening on loopback, the mesh address and the LAN address. The mesh's `dnsmasq` module was
assigned to the other three machines and not to this one, so no module on the home-server declared
port 53. Its filter opens only what a module declares, so DNS from the LAN was dropped. It started when
the home-server applied the filter this morning, after nine hours of applying nothing
([issue 194](../194-the-hosts-own-former-archive-stops-every-apply/00-report.md)).
Nothing said so. The daemon reported running, the filter applied cleanly, and the mesh had no record
that the home network depended on a service it did not know.
## Why it matters
**A service the mesh does not know is closed by the mesh's filter, by design, and nothing asks whether
something depends on it.** That is the right default for an unknown port. It is the wrong outcome for
the one service a whole network was told to use. The gap is that a machine can run something
important outside the mesh with nothing to show it.
**The mesh's `dnsmasq` could not have served the LAN either.** It listened on loopback and the mesh
address only. The reach of its DNS endpoints opens the filter, but the daemon would not have been
listening on the LAN address anyway.
## Open questions
- Should a machine report the listening services the mesh does not own, the way it reports the links
that face outside? This one would have been visible before the filter closed it.
- The LAN address the home-server answers on is now a setting, beside the reach that opens the filter.
Two statements that must agree. Should reach `public` on a DNS endpoint imply listening beyond the
mesh?
## Resolved (2026-10-02)
The home network was pointed at the gateway for DNS while the fix was built, which got the phones back
within minutes. Then:
- the mesh's `dnsmasq` takes the addresses it listens on beside the machine's from a setting, with
loopback as the mesh-wide default, so no other machine changed;
- the home-server's layer adds its LAN address, and its DNS endpoints' reach is `public`. The router
forwards no DNS, so that means the LAN;
- the module and its sibling `resolv-conf` were assigned to the home-server, replacing the
predecessor's daemon and configuration, which were kept aside;
- the home network was pointed back at the home-server, through a new `unifi` tool.
Checked live: from another machine on the LAN, public names and mesh names both resolve through the
home-server's LAN address, and the mesh and the machine itself resolve as before.
**One fault found on the way, and caught before it reached any machine.** A module's settings are
merged into every mergeable file the module owns. The first attempt therefore put the new setting into
docker's `daemon.json` as well as into dnsmasq's config, and dockerd refuses keys it does not know. The
plan showed it before any push. The change was reverted and redone with docker's file declared to take
no settings. The general fault, a module's settings reaching files they were not meant for, is still
there for any module with more than one file.