Compare commits

..
Author SHA1 Message Date
jschoubben af170e3a67 Issue 171: a module that names its own resolver knows no mesh name
Found and fixed the afternoon ADR 0148 landed: mailu-admin lost its
database behind Mailu's own resolver. Two catalogue PRs; an insight on
0148 that a container's dns is a decision, not a preference.
2026-09-30 15:20:26 +02:00
jschoubben 6c2d5f5913 Merge pull request 'Group 2 is resolved: containers resolve, nothing is copied, a route's name says where it arrives' (#219) from issue/110-resolved into main 2026-09-30 13:07:14 +00:00
jschoubben 04c9500b5b Group 2 is resolved: containers resolve, nothing is copied, a route's name says where it arrives
Issue 110's cause was not the filter: the runtime had never been told,
and the resolver dropped a query arriving on a bridge. ADR 0148 step 3
landed once it did (109, 151 resolved). ADR 0151 composes a route's
internal name under the serving node and drops the suffixed alias
(139, 157 resolved). Design 08 amended; a fact in 0148 corrected.
2026-09-30 14:56:43 +02:00
jschoubben 846c1f85f2 Merge pull request 'Issue 107 is resolved: a declaration carries its order' (#217) from issue/107-resolved into main 2026-09-30 12:13:57 +00:00
jschoubben 9eef0bd525 Issue 107 is resolved: a declaration carries its order
Hosts first, then the controller — a build and a push each, now that the
mesh delivers the host. The host refuses a lower sequence than it kept
and drains a batch by sequence rather than arrival; the controller
numbers each send under the node's hold, inside the signed bytes.

Measured: two pushes, sequence 2 in the kept declaration, counters in
the store agree, no machine reads as behind. That last one is the
subtlety: the mesh compares the digest of what it would send against
what it did, and a number changes the bytes, so the read-only comparison
composes with the last number sent rather than a fresh one.
2026-09-30 14:13:50 +02:00
jschoubben 6e08cdf3d6 Merge pull request 'Every machine self-updates, verified, and 107's gate has opened' (#216) from issue/142-self-update-on-every-machine into main 2026-09-30 11:51:53 +00:00
jschoubben 02f291a129 Every machine self-updates, verified, and 107's gate has opened
All four machines run a host the mesh built, published and delivered, the
last delivery unattended: each stood aside once for a genuinely newer
version and the delivered launcher started it. A following push that
delivered nothing new was applied and reported by every machine and stood
nobody aside.

The crossover needs one restart of the unit per machine, once, because
the running launcher executes from its own inode. Measured timing: three
seconds on the machine, 17-20 as the operator sees it, the difference
being the control plane composing before it sends.

107 is unblocked: a declaration field is now a build and a push.
2026-09-30 13:51:46 +02:00
jschoubben 9b14430d3f Merge pull request 'Issue 163: a delivered host stood aside on every push and reported nothing' (#214) from issue/163-a-delivered-host-stands-aside-on-every-push into main 2026-09-30 11:47:35 +00:00
jschoubben a4384f13d3 Issue 163: a delivered host stood aside on every push and reported nothing
Asked whether a newer host was delivered using the link-time stamp, which
every delivered host carries as 'development build' now that the version
comes from where the binary sits. Never matched, so it stood aside on
every push for ever; standing aside cancels the report, so the mesh never
heard from it. Read as healthy throughout.

The three-minute push wait made it invisible: a wait long enough to
absorb a whole apply is long enough to hide that the machine never
answered.
2026-09-30 13:47:28 +02:00
jschoubben a841e2c173 Merge pull request 'The host self-updates, and an archive cannot be undeclared' (#212) from issue/161-resolved-and-162-an-archive-cannot-be-removed into main 2026-09-30 11:16:44 +00:00
jschoubben 3c535ead31 The host self-updates, and an archive cannot be undeclared
161 resolved and verified on a machine: the workstation runs a host the
mesh compiled, published, delivered and started, applying declarations
and reporting the version it was delivered as.

The system it was built for comes from the artifact — the one thing a
toolchain takes from a module, which 0142 already allowed because the
target is a property of the artifact. The version comes from where the
binary sits, which 0142 decided and nothing had implemented.

Two mistakes on the way, both caught by reading the output rather than
the line that claimed success. A second -ldflags does not merge with the
first: the binary gained its system and lost -s -w, 12.2MB against 8.5MB.
And the delivered binary was named after its package, so the first
delivery was correct, reported success and was invisible to the launcher.

A delivered host that cannot apply is a machine the mesh cannot repair,
because the declaration that would fix it is the one it cannot apply. The
launcher's fallback is what made that an inconvenience instead of an
expedition.

162 is new and not about the host: an archive has no removal, so a module
using one can never be unassigned, and the attempt takes the whole apply
with it — the machine applies nothing else either. It is how undoing the
first delivery froze the workstation.
2026-09-30 13:16:37 +02:00
jschoubben 4cf941d858 Merge pull request 'Self-update works, and a delivered host is one fact short of usable' (#211) from issue/161-a-delivered-host-has-no-link-time-facts into main 2026-09-30 10:28:08 +00:00
jschoubben 5042ffd8d3 Self-update works, and a delivered host is one fact short of usable
The loop closed on the workstation: the version landed, the launcher was
replaced, the running host stood aside, and after one restart the launcher
started a binary the mesh had compiled, published and delivered.

The launcher goes as a file resource rather than inside the archive, and
that is the safety rather than a preference. A file is written atomically,
so the running launcher keeps the inode it started from; an archive writes
in place with truncate and would cut a script a shell is reading. The
manifest carries a second copy and a test refuses any drift from the one
in packaging.

Then it would have refused the first declaration it was asked to apply.
The Makefile links in two facts the mesh's toolchain does not, on purpose,
and one of them is the system the host was built for — read before
anything is applied, so the failure is safe and total. Nothing reports it:
the unit is active, the bus link is up, and the log says it is hearing
what the node should be.

Worse, the declaration that would fix it is the declaration it cannot
apply, so the mesh cannot repair such a machine. Restored by moving the
delivered versions aside and letting the launcher fall back, which is the
fallback working as designed.

0142 already settles the version — it comes from where the component sits,
not from its linker — and that is unimplemented. The system pin has no
answer, and the candidates are a decision rather than a fix: put it in the
path too, carry it in a file beside the binary, or stop pinning at link
time at all, which is 0005's to change.
2026-09-30 12:28:01 +02:00
19 changed files with 759 additions and 21 deletions
@@ -103,6 +103,13 @@ reintroduces 109 and 135 — silently, and on a live mesh, which is exactly how
([issue 110](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md));
and on two of four machines the resolver binds loopback only, so the runtime hands containers a
public resolver instead. Both are prerequisites, not related work.
> **Progressive insight — 2026-09-30, later the same day. The loopback claim was wrong.** The
> resolver bound the private address on all four machines; on two the runtime had never been told
> to use it, and on all four the resolver discarded a query that arrived on the runtime's bridge.
> The step stands; the facts under it were those. Both fixed the same day
> ([issue 110's resolution](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/01-resolution.md)),
> and step 3 landed after them.
2. **The runtime is told which resolver to use, per machine, as a file** — not per container as a
creation-time argument, or the resolver's address is back in every container's identity and the
problem has only got smaller.
@@ -143,6 +150,16 @@ closed by this record, only answered by it.
resolvable inside the mesh — holds unchanged and by the same means the machine already uses.
- **Issue 110 stops being a container-DNS inconvenience and becomes a prerequisite** for the mesh not
restarting itself whenever it learns a name.
- **A container that names a resolver of its own has opted out of the machine's**, and the copy this
record removes was the only reason such a container could reach anything by a mesh name.
> **Progressive insight — 2026-09-30, the afternoon this landed. Found the hard way.** The mail
> system's admin, behind Mailu's own resolver, lost its database the moment the copy went
> ([issue 171](../04-ISSUES/171-a-modules-own-resolver-knows-no-mesh-name/00-report.md)). A `dns` on
> a container is a decision about whether mesh names exist inside it, not a preference; the module
> was corrected, and whether the controller should refuse the contradiction is that issue's open
> question.
- **A container started by hand gets the mesh's names too**, where before only declared containers did.
Design 08 drew that boundary deliberately, on the grounds that reaching into every container is what
a nameserver would be for. This record accepts that consequence rather than working around it: a
@@ -0,0 +1,99 @@
---
topic: the tiers
status: accepted
date: 2026-09-30
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0066-public-routing-is-name-agnostic.md
---
# 151. A route's internal name is composed under the node that serves it
## Context
A module that requires a route is given two names from one label: a public one, `<label>.<public
domain>`, and an internal one, `<label>.<node>.internal`
([ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md)). Both were composed
from the node the module runs on.
The two are answered differently. The public name is published into every machine's roster at the
address of the node whose proxy serves it ([ADR 0066](0066-public-routing-is-name-agnostic.md)), so
it reaches the proxy from anywhere in the mesh. The internal name is answered by every machine's
resolver as *anything under a node's name goes to that node*
([design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md)) — the node it was composed from, which is
the consumer's. Where the proxy runs on another machine, that name sends a client to a machine with
nothing listening, while the public name works
([issue 139](../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md)).
Every route on this mesh today is served beside its module, so it has not been seen; `route` is
provided mesh-wide precisely so that stops being true.
Beside it, the roster gave every routed name a second entry with the mesh's suffix appended —
`<name>.<public domain>.internal` — because it composed a full name for every entry as it does for a
machine. That name resolved on every machine, was served by nothing, and was refused by the proxy at
the handshake; the first three names tried while reproducing an unrelated issue were those, and the
evidence pointed at a regression that had not happened
([issue 157](../04-ISSUES/157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md)).
## Considered Options
**1. Keep the consumer's name and publish it at the serving node's address**, as the public name is.
The name stays `<label>.<consumer>.internal` and an exact roster entry overrides the wildcard.
Rejected: it makes `<x>.<node>.internal` mean *goes to that node* except when it does not, which is
the one rule the resolver design states; it needs an entry per route where the wildcard needed none;
and which of an exact entry and a wildcard a resolver answers first is the resolver's business, which
the mesh deliberately does not know.
**2. A proxy on every machine, so the serving node is always the consumer's.** Rejected for this
question: it is a different decision about what `route` is — a node-scoped seat with a mesh-wide
fallback — and this mesh runs one proxy on the hub today. Whatever is decided there, a route served
from another machine must have a name that reaches it.
**3. Compose the internal name under the node that serves the route.** Chosen.
## Decision
**A route's internal name is `<label>.<serving node>.internal` — composed under the node whose proxy
answers the route, which is the machine the request arrives at.** The public name is unchanged:
`<label>.<public domain>` of the node the module runs on, which is where the operator put it.
Where the proxy runs beside the module — every route on this mesh today — the two nodes are one and
nothing changes. Where it does not, the name says where the request goes, which is what a name under
a node's name has always meant.
**A routed name has no mesh form.** The roster publishes it as itself, once, at the serving node's
address. Only a machine has a bare name beside its full one.
What certifies the internal name is unchanged by this: the proxy that terminates it obtains a
certificate from the mesh's authority for the names it is given, and it is given this one.
Taken on the operator's standing instruction to answer the open design questions in the work order.
## How this is checked
- **Composition.** A controller test contributes a route from a module on one node to a proxy offered
from another, gathered the way the controller gathers a consumer's contribution for a provider on
another machine, and asserts the internal name carries the serving node.
- **Publication.** A controller test renders a roster with a machine and a routed name and asserts
the routed name appears as itself, once, and never with the suffix appended.
- **On the mesh.** After the change no machine's roster carries a `<domain>.internal` entry, and a
route's internal name still answers from a container with a certificate from the mesh's authority.
## Consequences
- **A route served from another machine now has a usable internal name.** The first module assigned
that way will resolve, where before it would have resolved to the wrong machine with no error.
- **The internal name of a route can change when its proxy moves.** A route re-homed from one proxy
to another gets a new internal name, as the design's rule implies; clients that dialled the old one
reach the old machine. The public name does not move with the proxy and is the stable one.
- **The roster is one line shorter per routed name**, and a person reading a hosts file no longer
finds names that resolve to a refusal.
- **Issue 139's second question — a per-node route holder — is left open**, and is a decision about
what a seat is rather than about a name.
## References
- [issue 139](../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md) — the question
- [issue 157](../04-ISSUES/157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md) — the alias
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — routed names propagate mesh-wide; extended here
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — how the two names are composed and how far each reaches
- [design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md) — anything under a node's name goes to that node
+1
View File
@@ -200,6 +200,7 @@ python3 00-META/checks/index.py fail if stale
- **0109** — [A package registry seat is one per ecosystem, not one for all of them](0109-a-package-registry-seat-is-one-per-ecosystem.md)
- **0126** — [A module declares its own seats; the mesh reserves its own](0126-a-module-declares-its-own-seats.md)
- **0148** — [The mesh's names are resolved, not copied into every container](0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
- **0151** — [A route's internal name is composed under the node that serves it](0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md)
### What runs on them, and how it gets there
+17 -5
View File
@@ -10,6 +10,7 @@ code:
updated: 2026-09-30
decisions:
- 02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md
- 02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md
- 02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md
- 02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md
- 02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md
@@ -305,10 +306,11 @@ and [135](../../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report
**A container resolves the mesh's names through its machine's resolver, at the moment it asks, and
nothing is copied.** The resolver is a machine-level process rather than a container, so nothing
circular is being asked for. This is gated on a container being able to reach the resolver from any of
the runtime's networks, which it cannot today
([issue 110](../../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)) —
until that lands the mesh keeps copying and keeps comparing, and the order is stated in the record.
circular is being asked for. It was gated on a container being able to reach the resolver from any of
the runtime's networks
([issue 110](../../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)),
and landed the day that did, 2026-09-30: the controller writes no mesh name into a container and the
host's digest carries only what the module declared for itself.
The paragraph below states the old boundary, and 0148 deliberately gives it up: a container somebody
started by hand resolves the same names as everything else, because the resolver answers the machine,
@@ -324,6 +326,14 @@ machine — declared or not — is what a nameserver in `resolv.conf` would be f
service, the rest is the node — so what resolves is *anything under a node's name*, going to that
node. What routes it once it arrives is a proxy's, and stays separate.
*2026-09-30.* **So the node in a route's internal name is the one whose proxy answers it**
([ADR 0151](../../02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md)).
Composed from the node the module ran on, the name sent a client to a machine with nothing listening
whenever the proxy ran elsewhere
([issue 139](../../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md));
composed from the serving node, the rule above holds without exception. The public name stays the
module's node's, which is where the operator put it.
**The mesh writes the data and runs no daemon.** One wildcard per machine, from the same set that
writes the hosts file. A resolver is third-party software and runs *on* the mesh rather than being
*of* it: the mesh has no business shipping one, choosing which one, or knowing its configuration
@@ -399,7 +409,9 @@ can reach from the outside but cannot resolve from the inside is a name it canno
authority of its own.
**So a granted route is published into internal resolution as well** — the routed name to the node
that serves it, mesh-wide, by the same mechanism that writes the node names. It is *given by the
that serves it, mesh-wide, by the same mechanism that writes the node names — and as itself: a routed
name has no mesh form, and the suffixed alias the roster once added beside it resolved to a refusal
([issue 157](../../04-ISSUES/157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md)). It is *given by the
mesh, not chosen by a module*, for the same reason the node names are: a module listing the routes
would go stale the day one changes. The mesh propagates the names it was told to serve and still
knows nothing about what they mean
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-09-23
located-in: [mesh-controller internal/link, mesh-host internal/link]
fixed-by:
fixed-by: mesh-host PR 59 (the host refuses an older sequence and drains by it), mesh-controller PR 160 (each send is numbered under the node's hold) — measured 2026-09-30, 02-resolution.md
amended-design:
---
@@ -66,3 +66,11 @@ Left `located`. The owner is unchanged, the shape of the fix is agreed, and the
[issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md) rather than anything
in this record. **This is a judgement about order, not a refusal** — it is cheap to overrule, and the
code is a day's work once a host can be delivered.
## The gate has opened (2026-09-30, evening)
The mesh delivers the host now — built by its own toolchain, published to its own registry, delivered
over the bus and started by the launcher, on all four machines
([issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/01-progress.md)). A declaration
field is a build and a push, not an expedition. The order this record asked for — hosts first, then
the controller — is now two commands and a status line that says when the first has finished.
@@ -0,0 +1,60 @@
# 107 — resolved: a declaration carries its order
*2026-09-30. Measured on the mesh.*
## What was done
**Hosts first, then the controller** — the order [issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md)
says a new declaration field needs, and now a build and a push rather than an expedition
([issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/01-progress.md)).
The host understands a `sequence` on a declaration and tolerates its absence: absent reads as "no
order claimed", not "first", so a controller that sends none is still understood and a host that kept
a declaration before it understood the field compares nothing. It refuses a declaration with a lower
sequence than the one it kept, whole, and says why; and the drain that picks one declaration from a
batch keeps the highest sequence rather than the last to arrive — which is the case the report
constructed, a backlog drained out of order.
The controller numbers each send: the next number for that node, taken under the node's hold, before
the body exists, so the number is inside what the mesh signs and a replayed older declaration cannot
borrow a newer one's.
## Measured
```
push shanks; push shanks
sequence in kept declaration: 2
node sequence
novox 2
shanks 2
ace (none — not sent since numbering)
g14 (none)
status: nobody "not running what the mesh would send them"
```
Both applies went through; neither was refused; the machine holding the earlier one accepted the later.
## The subtlety, which would have read every machine as behind for ever
The mesh decides a machine is behind by comparing the digest of what it **would** send against what it
**did** send. A number changes the bytes. So the read-only comparison composes with the number the
machine was *last* sent — not a fresh one — and is byte for byte what was sent when nothing else
changed. Without that, numbering would have made `status` name all four machines as out of date on
every reading, permanently.
## The open questions
- *A per-node `sequence` under the controller's node hold?* Yes, as described. **`supersedes` — the
previous digest — is not added.** A strictly-greater sequence gives the ordering; a chain of digests
would give continuity, which nothing here needs yet and which every re-composition would break.
- *Genesis signing its bundle as sequence zero?* Zero is "no order claimed", which is what the bundle
carries by carrying nothing. Same rule, no genesis branch.
- *A marker for a mode change?* Not needed for the incident it guards: a replayed converged declaration
reaching a node returned to adopted is already refused **by mode**, before this check runs.
## How it is checked
Host: an older sequence is refused, a newer or equal one is not, and no order claimed on either side
compares nothing; the drain keeps the highest sequence, and falls back to arrival when none is claimed.
Controller: a send carries its number inside the signed bytes, an unnumbered send is byte for byte what
it was before, and each node's counter is one higher per send and readable for the comparison.
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-09-24
located-in: [mesh-host internal/apply]
fixed-by:
located-in: [mesh-controller internal/catalogue/declaration.go (every container was given the roster at creation)]
fixed-by: mesh-controller PR 161 — no container is given a mesh name; it resolves through its machine's resolver (ADR 0148, landed 2026-09-30 once issue 110 did)
amended-design:
---
@@ -73,3 +73,13 @@ copying: a container resolves through its machine's resolver at the moment it as
record reports then has nowhere to occur. It is gated on
[issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md), so
until that lands the mesh still copies and still compares.
## Resolved (2026-09-30)
110 landed the same day ([its resolution](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/01-resolution.md)),
and mesh-controller PR 161 then removed the copy: no container is given a mesh name or a mesh address,
and a module's own declared entries are the only `host` lines it carries. Verified on the control-node
after its containers were recreated once — the last time a name will do that: the forge's container
carries no extra hosts and resolves another machine and a routed name through the machine's resolver,
so the shape this record describes has nowhere to occur. Checked in the controller's tests: a
container's declaration is byte-for-byte the same under a roster of one machine and a roster of three.
@@ -1,8 +1,9 @@
---
status: open
status: resolved
opened: 2026-09-24
located-in: []
fixed-by:
located-in:
- mesh-catalog modules/dnsmasq (the runtime was never told; the resolver answered by interface)
fixed-by: mesh-catalog PR 175 (the runtime is reloaded and keeps its containers over a restart) and PR 176 (the resolver answers by address, so a query from a bridge is admitted) — measured 2026-09-30, 01-resolution.md
amended-design:
---
@@ -67,3 +68,6 @@ and [135](../135-a-containers-mesh-names-are-not-compared/00-report.md)).
network is the case with no DNS at all, and it is the case the mesh's own forge runs in. Two of four
machines also bind the resolver to loopback only, so the runtime hands their containers a public
resolver. Both halves are this issue.
*Later the same day: the second half was wrong, and the first had a different cause than the one above.
[01-resolution.md](01-resolution.md) has what was actually found.*
@@ -0,0 +1,64 @@
# 110 — resolved: a container on any network reaches the resolver, and is answered
*2026-09-30. Measured on the three converged machines; the adopted one holds its resolver module until it
is taken and is not covered.*
## What was actually wrong
Not what the report predicted. The report named the filter: a container on the runtime's default
network asks from a bridge address, and the converged filter admitted queries by source address only.
That was true when it was written and was fixed before this issue was ever tested — the filter admits
by the link a packet arrives on ([ADR 0144](../../02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md)),
and a container's bridge is admitted whole. Tested on every machine: the query arrives, the filter
passes it.
Three other things were wrong, each hiding the next.
**The runtime had never been told.** The resolver module writes the runtime's `dns` key into the
runtime's own configuration file. The runtime reads that key when it starts and not on a reload, and on
two machines the runtime predated the file — so every container they started got a public resolver, and
`novox.internal` came back as not existing. Nothing reported this: the file was present and current,
the resolver ran, and a name not existing is a valid answer. Fixed in mesh-catalog PR 175: the module
also sets `live-restore` and reloads the runtime when its file changes, so the one restart the `dns` key
needs no longer stops every container. The restart is then the operator's, once per machine; done on
both today, with every running container kept.
**The resolver dropped the query.** With the runtime corrected, a container's query reached the resolver
— and got no answer, on every machine, including the one whose runtime had been right all along. The
socket was bound to the private address; the filter admitted the packet; dnsmasq received it and
discarded it without a line of log. Its configuration said `interface=mesh0`, and dnsmasq admits a
query by the interface it arrives on when told an interface: a container's query is addressed to the
private address but arrives on the runtime's bridge, and the bridge is not `mesh0`. Fixed in mesh-catalog
PR 176: the resolver is told the address to answer on, not the interface that carries it, and a query to
that address is admitted whatever bridge brings it. The bridges are the runtime's to name.
**The report's second half was wrong.** "Two of four machines bind the resolver to loopback only" was
an inference from the containers' behaviour, and the behaviour had the cause above. The resolver bound
the private address on all four; nothing had asked it there.
## What is verified
From a container on the runtime's default network, started by hand and given nothing, on each of the
three converged machines: `novox.internal` answers with the hub's private address, through the machine's
own resolver. That is the fourth check of
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) — "on every
network the runtime offers" — and its first step; the record's step 2 (the runtime told per machine, as
a file) was already how the module works. Step 3 may now begin.
## What checks it
By hand, today. Nothing in the mesh asserts that a container can resolve a mesh name: the resolver's
own tests cover what it answers, not who can ask. The check that would have caught all three faults is
the one the report asked for and 0148 lists — a container on the default network resolving a mesh name
— and it is not built. It belongs with the reachability check of
[issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
which is parked; until then this is a thing a person verifies after touching the resolver, the filter,
or the runtime's configuration.
## What this cost to find
The three faults produced one symptom — a container that cannot resolve — and each fix revealed the
next. The first was found by reading the runtime's own view of its configuration rather than the file;
the second by capturing the query on the bridge and finding it arrive and go unanswered; the third only
by admitting the first belief was wrong. A machine that had been believed to work all day had never
worked either.
@@ -1,9 +1,9 @@
---
status: open
status: resolved
opened: 2026-09-28
located-in: [mesh-controller internal/catalogue]
fixed-by:
amended-design:
located-in: [mesh-controller internal/catalogue/declaration.go (composeName took the consumer's own name as the internal domain)]
fixed-by: mesh-controller PR 163 — the internal name composes under the node whose proxy serves the route (ADR 0151, 2026-09-30)
amended-design: 03-DESIGN/01-to-be/08-connectivity.md
---
# 139 — An internal route name resolves to the consumer's node, not the one that serves it
@@ -49,3 +49,15 @@ resolve whether or not anything answers.
and if so, is `route` still one mesh-wide provision or a node-scoped seat with a mesh-wide fallback?
- What certifies the name in either case? The certificate is obtained by whoever terminates TLS, and
that is the question above in another form.
## Answered (2026-09-30)
[ADR 0151](../../02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md):
the internal name is composed under the node that serves the route — the machine the request arrives
at — because `<x>.<node>.internal` means *goes to that node* and nothing else. The public name stays
the consumer node's, which is where the operator put it. The first question is answered that way; the
second, a per-node route holder, is a decision about seats and is left where it is; the third is
unchanged, since the proxy that terminates the name is given it and certifies it.
mesh-controller PR 163 carries it. On this mesh every route is served beside its module, so no name
changed; the controller's tests hold the case where it would.
@@ -95,3 +95,91 @@ and then read by nothing: it does not reach the compiler, no machine is matched
chooses between two artifacts by it. The host built here is x86-64 because the build machine is, not
because anything in the declaration said so — correct for this mesh by coincidence. That is
[issue 159](../159-an-artifacts-system-is-checked-and-then-ignored/00-report.md).
## Delivered, started, and one fact short (2026-09-30, later)
The loop closed. The launcher is delivered as a **file** resource rather than inside the archive, and
that difference is the safety: a file is written atomically — temp file, then rename — so the running
launcher keeps the inode it was started from, where an archive writes in place with truncate and would
cut the script a running shell is reading. The manifest carries a second copy of the launcher and a
test refuses any difference from `packaging/nox-mesh-host-launch`.
On the workstation, in order: the version landed, the launcher was replaced, the running host saw a
delivered version and stood aside, and after one restart of the unit the launcher started
`/usr/lib/nox-mesh-host/versions/637f65559d16/nox-mesh-host`. **A host the mesh compiled, published,
delivered and started.**
It would then have refused the first declaration it was asked to apply. The host's Makefile links in
two facts the mesh's toolchain does not, and one of them — the system it was built for — is read before
anything is applied. That is
[issue 161](../161-a-delivered-host-carries-none-of-its-link-time-facts/00-report.md), and the machine
is back on its hand-placed binary until it is answered.
**The fallback is what made that safe**, and it was not luck: the launcher runs the pinned version, or
the newest delivered one, or the one placed by hand — so moving the delivered versions aside restored
the machine in one step.
## Self-update works (2026-09-30, end of the day)
```
running: /usr/lib/nox-mesh-host/versions/093231796eb0/nox-mesh-host
mesh-controller node show shanks
host 093231796eb0
```
One machine runs a host the mesh compiled, published, delivered and started, applying declarations and
reporting the version it was delivered as. The two facts a delivered binary was missing are
[issue 161](../161-a-delivered-host-carries-none-of-its-link-time-facts/01-resolution.md) and resolved:
the system comes from the artifact, the version from where the binary sits.
**Three machines still run a hand-placed host**, and rolling each forward is one assignment and one
push. The control node is worth last.
One thing this found on the way out: an archive cannot be undeclared, and the attempt stops the machine
applying anything at all — [issue 162](../162-an-archive-cannot-be-undeclared/00-report.md). It is how
undoing the first delivery froze the workstation, and it is not specific to the host.
## Every machine self-updates (2026-09-30, evening)
```
shanks 76f4566bef3d/nox-mesh-host active
g14 76f4566bef3d/nox-mesh-host active
novox 76f4566bef3d/nox-mesh-host active
ace 76f4566bef3d/nox-mesh-host active
mesh-controller status: (no host split)
```
The last delivery was unattended on all four: the fixed host was built, pushed, each machine stood
aside exactly once for the genuinely newer version, and the delivered launcher started it — no
restart by hand. A following push that delivered nothing new was applied and reported by every
machine and stood nobody aside, which is the check
[issue 163](../163-a-delivered-host-stood-aside-on-every-push-and-reported-nothing/00-report.md) asks
for.
**Two more faults on the way, both mine, both found by reading the machine rather than the success
line.** A delivered host compared the newest delivered version against its link-time stamp rather
than the version it was running, so it stood aside on every push and — because standing aside cancels
the report — never reported again (163). And the adopted machine kept its found launcher as the
adoption rule says, so the delivery there needed a `take` before the launcher moved.
**The crossover needs one restart of the unit per machine, once.** The launcher process that was
running on each machine was the old script, executing from its own inode; a new file beside it
changes nothing until the unit restarts. Every subsequent delivery is unattended.
**Timing, measured:** on a machine, hearing a declaration to reporting it applied is about three
seconds. A push as the operator sees it takes 17–20 seconds, and the difference is the control plane
composing the declaration before it sends. A `--wait` shorter than that reads as "did not report" for
a machine that did; the three-minute default read as slowness for a machine that never would. Neither
number is a defect being chased here, and both are worth knowing before reading a push's answer.
## What this leaves
- [Issue 162](../162-an-archive-cannot-be-undeclared/00-report.md): an archive cannot be undeclared, so
the host module — and any module with an archive — cannot be unassigned, and trying stops the machine
applying anything.
- [Issue 107](../107-a-declaration-carries-no-order/00-report.md) is unblocked: a declaration field is
now a build and a push rather than an expedition.
- Three stale version directories on the workstation from the first attempts, moved aside under
`/var/lib/mesh-host/versions-held-back/`, and a backup of the adopted machine's hand-placed binary
beside its state. Both are safe to delete and are not the mesh's to delete.
@@ -1,10 +1,10 @@
---
status: open
status: resolved
opened: 2026-09-29
located-in:
- mesh-host internal/apply/apply.go (containerSpecReading hashes every `host` entry)
- mesh-controller internal/catalogue/declaration.go (withMeshNames gives every container the mesh's names)
fixed-by:
fixed-by: mesh-controller PR 161 — the roster left every container's declaration and so its digest (ADR 0148 step 3, 2026-09-30)
---
# 151 — A new name recreates every container in the mesh
@@ -92,3 +92,15 @@ copying names until a container can reach the resolver from any of the runtime's
([issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)),
which it cannot on two of four machines today. Removing the copy first reintroduces 109 and 135
silently, on a live mesh, which is how both were found. The order is in the record.
## Resolved (2026-09-30)
Step 3 landed the day 110 did. mesh-controller PR 161 stops writing the roster into any container, so
a container's digest no longer carries a name that is not its own. The controller's tests hold the
record's check — a container's declaration does not move when the mesh's roster does, and does move
when the module's own declared entries do.
The first push after the change recreated every container once, because every digest lost its host
entries at the same moment. That was the last such event: from here a name added or moved on one
machine changes no container anywhere, and the record's second check — add a routed name, watch every
other machine's apply report say nothing changed — is what the next module assignment will show.
@@ -1,9 +1,9 @@
---
status: located
status: resolved
opened: 2026-09-30
located-in:
- mesh-controller internal/catalogue/roster.go (the roster's entries for a routed name)
fixed-by:
fixed-by: mesh-controller PR 163 — a routed name is published as itself, once, with no suffixed alias (ADR 0151, 2026-09-30)
amended-design:
---
@@ -61,3 +61,10 @@ composition produces `keycloak.novox.be.internal`, which is not a name anything
The fix is a judgement about what a routed name's internal form is, and 139 is the record that asks it;
this one is the evidence that the current answer publishes a third thing that is neither.
## Resolved (2026-09-30)
The judgement 139 asked for is [ADR 0151](../../02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md):
a routed name has no mesh form. The roster now publishes it as itself, once, at the serving node's
address; the `<domain>.internal` line is gone from every machine's hosts file, and a controller test
refuses it coming back.
@@ -0,0 +1,95 @@
---
status: resolved
opened: 2026-09-30
located-in:
- mesh-host cmd/mesh-host/main.go (version and builtFor, both set at link time)
- mesh-controller internal/builder (the toolchain, which deliberately takes nothing from the module)
fixed-by: mesh-controller (the system stamp, and one linker flag rather than two), mesh-host (the version read from the path) — verified on a machine 2026-09-30, 01-resolution.md
amended-design:
---
# 161 — A host the mesh built carries none of the facts its Makefile stamps in
## What was observed
*2026-09-30, on the workstation, having just made the host self-updating.*
The mesh compiled the host, published it, delivered it and the launcher started it. It ran, read the
machine correctly, and **would have refused the first declaration it was asked to apply.**
The host's own Makefile links in two facts:
```
LDFLAGS := -s -w -X main.builtFor=$(SYSTEM) -X main.version=$(VERSION)
```
The mesh's Go toolchain links in neither, on purpose: a toolchain accepts nothing from the module,
because anything a module could override there it would be writing a Dockerfile to override
([ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md)). So a
delivered host has `builtFor = ""` and `version = "development build"`.
**`builtFor` empty is the one that bites.** Before applying anything, the host asks which system it
was built for:
```go
sys, err := system.For(builtFor)
```
and that answers, for an empty name:
```
this host was built for "", which is not a system it knows. Built hosts are: …
```
It is called before any resource is applied, so the failure is in the safe direction — the machine is
not half-configured. It is still a host that cannot do its job, and nothing about it looks wrong: the
unit is active, the link to the bus is up, and the log says it is hearing what the node should be.
Measured: after the crossover the machine logged nothing further, where the previous host had written
a reconcile line every five minutes.
## Why this was found rather than reported
Nothing reports it. The host does not check its own stamps at start, the mesh does not ask, and the
declaration that would fail is the same declaration that would deliver a fix — so **a machine in this
state cannot be repaired by the mesh.** It was restored by moving the delivered versions aside and
letting the launcher fall back to the hand-placed binary, which is the fallback working exactly as
designed.
## What the records already say about half of it
[ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md) settles the
version and its answer is not implemented:
> A component's version comes from where it sits, not from its linker. It is unpacked into a directory
> named for its version, so it can read its own version from its path. The stamp goes, and with it the
> need for a build to know what it will be called.
That is exactly right and would also fix what the mesh reports: a delivered host would say
`637f65559d16` rather than `development build`, and
[issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md)'s host comparison would
mean something for delivered hosts.
**The system pin has no answer yet**, and it needs one before any mesh-built host can apply anything.
The tension is real: the target is a property of the artifact and 0142 says so, but a toolchain that
passed it would be linking a value into a variable whose name belongs to the module — which is the
coupling the toolchain exists to avoid. Candidates, none decided:
- the path carries it as well as the version, so the host reads both from where it sits, as 0142 does
for the version;
- the bundle carries a small file beside the binary saying what it was built for, written by the
builder from the artifact's declaration;
- the host stops being pinned at link time and refuses on a fact it reads from the machine instead —
which changes what [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) decided and is the biggest of
the three.
## What is true in the meantime
Self-update works end to end and is one fact short of usable: the mesh builds the host, publishes it,
delivers it to a machine, the running host stands aside, and the launcher starts the delivered one. The
machine is left on its hand-placed binary until this is answered, which is one command to undo.
## How a fix is checked
A host the mesh built and delivered applies a declaration on a machine, shown by the machine's own
reconcile line; and it reports a version that names the build it came from rather than a placeholder.
@@ -0,0 +1,59 @@
# 161 — resolved: a host the mesh built runs a machine
*2026-09-30. Measured on the workstation.*
```
running: /usr/lib/nox-mesh-host/versions/093231796eb0/nox-mesh-host
agent: active
reconciles in the last six minutes: 1
mesh-controller node show shanks
host 093231796eb0
```
A binary the mesh compiled, published to its own registry, delivered over the bus, started by the
launcher, applying declarations, and reporting a version that names the build it came from.
## The two facts, and where each now comes from
**The system it was built for comes from the artifact.** ADR 0142 already made the target a property
of the artifact rather than of the recipe, so the toolchain names the variable it fills and the
artifact supplies the value. It is the one thing a toolchain takes from a module, and it is stated
rather than inferred.
**The version comes from where the binary sits**, which is what
[ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md) decided and
nothing had implemented: a delivered host reads the directory it was unpacked into. A host placed by
hand keeps its link-time stamp, which is the honest answer for one the mesh did not deliver — and is
every other machine today.
## Two mistakes on the way, both found by reading the output
**A repeated flag is not a merged one.** The stamp was appended as a second `-ldflags`, and the Go
command takes the last and drops the first. The binary gained its system and lost `-s -w`: 12.2MB
against 8.5MB, with its debug info. The comment I had written said the linker "accepts and merges"
them. It does not. Linker flags are the toolchain's own list now, composed into one flag, and a test
refuses a compile line that carries `-ldflags` itself.
**The delivered binary was named after its package.** `cmd/mesh-host` builds `mesh-host`; every
machine runs `nox-mesh-host`, which is what the launcher looks for inside a version. The first
delivery landed, reported `created … 1 file(s)`, and was invisible. An artifact says what its
executable is called now.
Both were caught by listing the directory and reading the binary rather than believing the line that
said it worked.
## What this cost while it was wrong, and what saved it
A delivered host that cannot apply is a machine the mesh cannot repair, because the declaration that
would fix it is the declaration it cannot apply. The workstation was restored by moving the delivered
versions aside so the launcher fell back to the hand-placed binary — **the fallback in the launcher,
working exactly as designed**, and the reason this was an inconvenience rather than an expedition.
It also loops if you are not careful: the working binary applies, delivers a version, stands aside,
and the broken one starts. Stopping the unit while the fix was built was the way through.
## What is left
**Three machines still run a hand-placed host.** Rolling them forward is one assignment and one push
each, and the control node is worth doing last and watching.
@@ -0,0 +1,56 @@
---
status: located
opened: 2026-09-30
located-in: [mesh-host internal/apply (no removal for an archive)]
fixed-by:
amended-design:
---
# 162 — An archive cannot be undeclared, and trying stops the machine applying anything
## What was observed
*2026-09-30, unassigning the host module from the workstation to undo a delivery.*
```
holding this machine: 0 applied, and map[apply:applying "mesh-host.next": no way to remove a "archive"
0 resource(s) were applied and remain; everything was attempted, so what is not listed as failed was done.
```
**Nothing was applied at all** — not the archive, not the other forty resources that had nothing to do
with it. The machine stopped reconciling and stayed that way until the module was assigned again.
## Why it matters
Every other resource kind can be taken away. A file is removed and what was found under it is put
back; a container is stopped and removed; a unit is given back the state it was found in
([ADR 0118](../../02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md)). An
archive has no removal at all, so:
- **a module with an archive can never be unassigned** — the attempt fails for ever;
- **the failure takes the whole apply with it**, so the machine applies nothing else either, and one
unassignable resource is a machine frozen against every other change;
- it is silent in the mesh's terms: the push reported sent, and only the machine's own journal says
what happened.
The host module is the obvious case and not the only one. An archive is for what inlining cannot
serve — a theme, an icon set, a tree of configuration — and any module using one is in the same
position.
## What the right answer probably is, and the question in it
The other kinds answer this by remembering what they found. An archive unpacks many files into a
directory the mesh did not necessarily create, so removal has a real question in it: **remove what the
archive put there, or remove the directory?** The first needs the applier to have recorded the file
list; the second would delete whatever else lives there — and for the host's own versions directory,
that is every other delivered version.
Recording what was unpacked is the answer that matches how the rest of the host behaves, and it is
what [issue 126](../126-a-volume-path-is-not-in-the-spec-comparison/00-report.md) and ADR 0118 already
argue for elsewhere: the mesh gives back what it found.
## How a fix is checked
A module with an archive is assigned, pushed, unassigned and pushed again; what the archive put on the
machine is gone, anything that was in the directory beforehand is still there, and the apply that
removed it applied everything else in the same declaration.
@@ -0,0 +1,56 @@
---
status: resolved
opened: 2026-09-30
located-in: [mesh-host cmd/mesh-host/main.go (the successor check after an apply)]
fixed-by: mesh-host PR 58 — the check asks with the running version, read from the binary's path, not the link-time stamp
amended-design:
---
# 163 — A delivered host stood aside on every push, and reported nothing
## What was observed
*2026-09-30, rolling the mesh-built host onto the last two machines.*
Every push to a machine running a delivered host produced, in order:
```
host 093231796eb0 is delivered; standing aside so the launcher runs it
applied 333 resource(s)
applied, and could not tell the mesh: reporting: context canceled
nox-mesh-host-launch: the host exited cleanly; starting it again
nox-mesh-host-launch: running /usr/lib/nox-mesh-host/versions/093231796eb0/nox-mesh-host
```
— for the version it was **already running**. It restarted itself on every push, for ever, and the mesh
never received a single report from it: `node show` kept the version from before the crossover, and
the operator's push waited its full three minutes for an answer that was never coming.
Read as healthy throughout: unit active, bus link up, "hearing what this node should be".
## Why
After an apply the host asks whether a newer host has been delivered than the one running, and the
question was asked with the **link-time version stamp**. Since
[issue 161](../161-a-delivered-host-carries-none-of-its-link-time-facts/01-resolution.md) a delivered
host's version comes from where it sits and its stamp is `development build` — so the comparison never
matched the newest delivered version, and "a newer host is waiting" was always true.
Standing aside cancels the context the report is published with, so the report was lost on every one
of those applies. Two faults from one wrong argument.
The change that moved the version to the path was applied to the report and to the known-good record,
and not here. Half a change, and the half left behind was the one that decides whether to exit.
## Why the three-minute wait made it invisible
The push's `--wait` timing out read as *slow*. It was not slow: **the report was never going to arrive.**
The operator put it exactly: *if you don't get a response in five seconds, something is wrong.* A wait
long enough to absorb a machine's whole apply is a wait long enough to hide that the machine never
answered.
## How it is checked
A machine running a delivered host is pushed a declaration that delivers nothing new; it applies,
reports, and does not stand aside. A machine running a delivered host is pushed a genuinely newer
version; it stands aside once, and the next push it does not.
@@ -0,0 +1,78 @@
---
status: resolved
opened: 2026-09-30
located-in:
- mesh-catalog modules/mailu (eight containers name Mailu's own resolver, and one of them binds a mesh name)
- mesh-catalog modules/dnsmasq (dropped the DNSSEC bit its upstreams set)
fixed-by: mesh-catalog PR 178 (mailu-admin uses the machine's resolver) and PR 179 (the machine's resolver passes the DNSSEC bit down) — 2026-09-30, the same afternoon
amended-design:
---
# 171 — A module that names its own resolver knows no mesh name
## What was observed
The afternoon [ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
landed — no container is given the mesh's names any more; it asks the machine's resolver — the mail
system's admin container began logging, 523 times in three minutes:
```
psycopg2.OperationalError: could not translate host name "novox.internal" to address: Name does not resolve
```
Mail was accepted on every port and the web front answered; the admin and the spam filter beside it
were unhealthy, and anything that needed the database — a mailbox change through the API, the spam
filter's domain list — failed. Found by the operator asking whether mail was back, forty minutes in.
Mailu ships its own resolver, an unbound in a container, and every other Mailu container is told to
use it — the module carries `dns: [192.168.203.254]` on eight containers. That resolver recurses from the
root and knows nothing under `.internal`. Until that afternoon the admin container had the database's
name anyway, because the mesh wrote every name into every container at creation; the copy was the only
reason a container behind its own resolver could reach anything by a mesh name, and nobody knew it was
load-bearing.
**Removing the override was not enough.** Given the machine's resolver instead, the admin refused to
start: `Your DNS resolver at 127.0.0.11 isn't doing DNSSEC validation`. Mailu checks, at start, that
its resolver returns the Authenticated Data bit for a signed name. The mesh's resolver forwards to two
upstreams that validate and set the bit, and dropped it on the way down — dnsmasq does unless told
otherwise. Mailu's own unbound has no hook to forward a zone elsewhere, so it could not be taught the
mesh's names either.
## Why it matters beyond this instance
**A container with a resolver of its own has opted out of the machine's, and nothing says so.** 0148
made the machine's resolver load-bearing for every container; a `dns` on a container is a quiet
exception to that, and the exception used to be papered over by the copy the record removed. The
manifest field reads like a preference and is a decision about whether mesh names exist inside the
container.
**A resolver that forwards to validating upstreams and hides the fact is less useful than it could
be**, and the first program to check found out.
**The mesh reported nothing.** Every container ran; the failing one accepted connections; the report
was about bytes. It is [issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
again, and the check that would have caught it is the same unbuilt one.
## What was done
- The one Mailu container that binds a mesh name — the admin, through the database it is granted —
no longer names Mailu's resolver and uses the machine's, like every container without a `dns` of
its own (mesh-catalog PR 178). The other seven keep unbound: the spam filter needs a validating
resolver for its blocklist lookups, and none of them asks for a mesh name.
- The machine's resolver passes the DNSSEC bit down from its upstreams, `proxy-dnssec` (PR 179). It
does not validate itself; the trust is the upstream's and the path to it, as a forwarding resolver's
always was, and the configuration says so.
## What checks it
The admin container's own start-up check, which is what failed, and the mesh's status once it reads
healthy. A container-level check that a mesh name resolves from inside every declared container is
the one 110 and 145 both ask for and is not built.
## Open questions
- Should a container's `dns` be refused, or made to say what it gives up? A module that names its
own resolver and binds a mesh name is a contradiction the controller can see at composition — the
grant hands it a name its resolver will not answer.
- Should the machine's resolver validate rather than proxy? It would cost a trust anchor on every
machine and make the resolver slower to start; proxying was enough for the one program that asked.