Compare commits
1
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
49b1136ded |
@@ -103,13 +103,6 @@ reintroduces 109 and 135 — silently, and on a live mesh, which is exactly how
|
||||
([issue 110](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md));
|
||||
and on two of four machines the resolver binds loopback only, so the runtime hands containers a
|
||||
public resolver instead. Both are prerequisites, not related work.
|
||||
|
||||
> **Progressive insight — 2026-09-30, later the same day. The loopback claim was wrong.** The
|
||||
> resolver bound the private address on all four machines; on two the runtime had never been told
|
||||
> to use it, and on all four the resolver discarded a query that arrived on the runtime's bridge.
|
||||
> The step stands; the facts under it were those. Both fixed the same day
|
||||
> ([issue 110's resolution](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/01-resolution.md)),
|
||||
> and step 3 landed after them.
|
||||
2. **The runtime is told which resolver to use, per machine, as a file** — not per container as a
|
||||
creation-time argument, or the resolver's address is back in every container's identity and the
|
||||
problem has only got smaller.
|
||||
@@ -150,16 +143,6 @@ closed by this record, only answered by it.
|
||||
resolvable inside the mesh — holds unchanged and by the same means the machine already uses.
|
||||
- **Issue 110 stops being a container-DNS inconvenience and becomes a prerequisite** for the mesh not
|
||||
restarting itself whenever it learns a name.
|
||||
- **A container that names a resolver of its own has opted out of the machine's**, and the copy this
|
||||
record removes was the only reason such a container could reach anything by a mesh name.
|
||||
|
||||
> **Progressive insight — 2026-09-30, the afternoon this landed. Found the hard way.** The mail
|
||||
> system's admin, behind Mailu's own resolver, lost its database the moment the copy went
|
||||
> ([issue 171](../04-ISSUES/171-a-modules-own-resolver-knows-no-mesh-name/00-report.md)). A `dns` on
|
||||
> a container is a decision about whether mesh names exist inside it, not a preference; the module
|
||||
> was corrected, and whether the controller should refuse the contradiction is that issue's open
|
||||
> question.
|
||||
|
||||
- **A container started by hand gets the mesh's names too**, where before only declared containers did.
|
||||
Design 08 drew that boundary deliberately, on the grounds that reaching into every container is what
|
||||
a nameserver would be for. This record accepts that consequence rather than working around it: a
|
||||
|
||||
@@ -1,99 +0,0 @@
|
||||
---
|
||||
topic: the tiers
|
||||
status: accepted
|
||||
date: 2026-09-30
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 02-DECISIONS/0066-public-routing-is-name-agnostic.md
|
||||
---
|
||||
|
||||
# 151. A route's internal name is composed under the node that serves it
|
||||
|
||||
## Context
|
||||
|
||||
A module that requires a route is given two names from one label: a public one, `<label>.<public
|
||||
domain>`, and an internal one, `<label>.<node>.internal`
|
||||
([ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md)). Both were composed
|
||||
from the node the module runs on.
|
||||
|
||||
The two are answered differently. The public name is published into every machine's roster at the
|
||||
address of the node whose proxy serves it ([ADR 0066](0066-public-routing-is-name-agnostic.md)), so
|
||||
it reaches the proxy from anywhere in the mesh. The internal name is answered by every machine's
|
||||
resolver as *anything under a node's name goes to that node*
|
||||
([design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md)) — the node it was composed from, which is
|
||||
the consumer's. Where the proxy runs on another machine, that name sends a client to a machine with
|
||||
nothing listening, while the public name works
|
||||
([issue 139](../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md)).
|
||||
Every route on this mesh today is served beside its module, so it has not been seen; `route` is
|
||||
provided mesh-wide precisely so that stops being true.
|
||||
|
||||
Beside it, the roster gave every routed name a second entry with the mesh's suffix appended —
|
||||
`<name>.<public domain>.internal` — because it composed a full name for every entry as it does for a
|
||||
machine. That name resolved on every machine, was served by nothing, and was refused by the proxy at
|
||||
the handshake; the first three names tried while reproducing an unrelated issue were those, and the
|
||||
evidence pointed at a regression that had not happened
|
||||
([issue 157](../04-ISSUES/157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md)).
|
||||
|
||||
## Considered Options
|
||||
|
||||
**1. Keep the consumer's name and publish it at the serving node's address**, as the public name is.
|
||||
The name stays `<label>.<consumer>.internal` and an exact roster entry overrides the wildcard.
|
||||
Rejected: it makes `<x>.<node>.internal` mean *goes to that node* except when it does not, which is
|
||||
the one rule the resolver design states; it needs an entry per route where the wildcard needed none;
|
||||
and which of an exact entry and a wildcard a resolver answers first is the resolver's business, which
|
||||
the mesh deliberately does not know.
|
||||
|
||||
**2. A proxy on every machine, so the serving node is always the consumer's.** Rejected for this
|
||||
question: it is a different decision about what `route` is — a node-scoped seat with a mesh-wide
|
||||
fallback — and this mesh runs one proxy on the hub today. Whatever is decided there, a route served
|
||||
from another machine must have a name that reaches it.
|
||||
|
||||
**3. Compose the internal name under the node that serves the route.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**A route's internal name is `<label>.<serving node>.internal` — composed under the node whose proxy
|
||||
answers the route, which is the machine the request arrives at.** The public name is unchanged:
|
||||
`<label>.<public domain>` of the node the module runs on, which is where the operator put it.
|
||||
|
||||
Where the proxy runs beside the module — every route on this mesh today — the two nodes are one and
|
||||
nothing changes. Where it does not, the name says where the request goes, which is what a name under
|
||||
a node's name has always meant.
|
||||
|
||||
**A routed name has no mesh form.** The roster publishes it as itself, once, at the serving node's
|
||||
address. Only a machine has a bare name beside its full one.
|
||||
|
||||
What certifies the internal name is unchanged by this: the proxy that terminates it obtains a
|
||||
certificate from the mesh's authority for the names it is given, and it is given this one.
|
||||
|
||||
Taken on the operator's standing instruction to answer the open design questions in the work order.
|
||||
|
||||
## How this is checked
|
||||
|
||||
- **Composition.** A controller test contributes a route from a module on one node to a proxy offered
|
||||
from another, gathered the way the controller gathers a consumer's contribution for a provider on
|
||||
another machine, and asserts the internal name carries the serving node.
|
||||
- **Publication.** A controller test renders a roster with a machine and a routed name and asserts
|
||||
the routed name appears as itself, once, and never with the suffix appended.
|
||||
- **On the mesh.** After the change no machine's roster carries a `<domain>.internal` entry, and a
|
||||
route's internal name still answers from a container with a certificate from the mesh's authority.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **A route served from another machine now has a usable internal name.** The first module assigned
|
||||
that way will resolve, where before it would have resolved to the wrong machine with no error.
|
||||
- **The internal name of a route can change when its proxy moves.** A route re-homed from one proxy
|
||||
to another gets a new internal name, as the design's rule implies; clients that dialled the old one
|
||||
reach the old machine. The public name does not move with the proxy and is the stable one.
|
||||
- **The roster is one line shorter per routed name**, and a person reading a hosts file no longer
|
||||
finds names that resolve to a refusal.
|
||||
- **Issue 139's second question — a per-node route holder — is left open**, and is a decision about
|
||||
what a seat is rather than about a name.
|
||||
|
||||
## References
|
||||
|
||||
- [issue 139](../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md) — the question
|
||||
- [issue 157](../04-ISSUES/157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md) — the alias
|
||||
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — routed names propagate mesh-wide; extended here
|
||||
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — how the two names are composed and how far each reaches
|
||||
- [design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md) — anything under a node's name goes to that node
|
||||
@@ -200,7 +200,6 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0109** — [A package registry seat is one per ecosystem, not one for all of them](0109-a-package-registry-seat-is-one-per-ecosystem.md)
|
||||
- **0126** — [A module declares its own seats; the mesh reserves its own](0126-a-module-declares-its-own-seats.md)
|
||||
- **0148** — [The mesh's names are resolved, not copied into every container](0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
|
||||
- **0151** — [A route's internal name is composed under the node that serves it](0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md)
|
||||
|
||||
### What runs on them, and how it gets there
|
||||
|
||||
|
||||
@@ -10,7 +10,6 @@ code:
|
||||
updated: 2026-09-30
|
||||
decisions:
|
||||
- 02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md
|
||||
- 02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md
|
||||
- 02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md
|
||||
- 02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md
|
||||
- 02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md
|
||||
@@ -306,11 +305,10 @@ and [135](../../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report
|
||||
|
||||
**A container resolves the mesh's names through its machine's resolver, at the moment it asks, and
|
||||
nothing is copied.** The resolver is a machine-level process rather than a container, so nothing
|
||||
circular is being asked for. It was gated on a container being able to reach the resolver from any of
|
||||
the runtime's networks
|
||||
([issue 110](../../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)),
|
||||
and landed the day that did, 2026-09-30: the controller writes no mesh name into a container and the
|
||||
host's digest carries only what the module declared for itself.
|
||||
circular is being asked for. This is gated on a container being able to reach the resolver from any of
|
||||
the runtime's networks, which it cannot today
|
||||
([issue 110](../../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)) —
|
||||
until that lands the mesh keeps copying and keeps comparing, and the order is stated in the record.
|
||||
|
||||
The paragraph below states the old boundary, and 0148 deliberately gives it up: a container somebody
|
||||
started by hand resolves the same names as everything else, because the resolver answers the machine,
|
||||
@@ -326,14 +324,6 @@ machine — declared or not — is what a nameserver in `resolv.conf` would be f
|
||||
service, the rest is the node — so what resolves is *anything under a node's name*, going to that
|
||||
node. What routes it once it arrives is a proxy's, and stays separate.
|
||||
|
||||
*2026-09-30.* **So the node in a route's internal name is the one whose proxy answers it**
|
||||
([ADR 0151](../../02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md)).
|
||||
Composed from the node the module ran on, the name sent a client to a machine with nothing listening
|
||||
whenever the proxy ran elsewhere
|
||||
([issue 139](../../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md));
|
||||
composed from the serving node, the rule above holds without exception. The public name stays the
|
||||
module's node's, which is where the operator put it.
|
||||
|
||||
**The mesh writes the data and runs no daemon.** One wildcard per machine, from the same set that
|
||||
writes the hosts file. A resolver is third-party software and runs *on* the mesh rather than being
|
||||
*of* it: the mesh has no business shipping one, choosing which one, or knowing its configuration
|
||||
@@ -409,9 +399,7 @@ can reach from the outside but cannot resolve from the inside is a name it canno
|
||||
authority of its own.
|
||||
|
||||
**So a granted route is published into internal resolution as well** — the routed name to the node
|
||||
that serves it, mesh-wide, by the same mechanism that writes the node names — and as itself: a routed
|
||||
name has no mesh form, and the suffixed alias the roster once added beside it resolved to a refusal
|
||||
([issue 157](../../04-ISSUES/157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md)). It is *given by the
|
||||
that serves it, mesh-wide, by the same mechanism that writes the node names. It is *given by the
|
||||
mesh, not chosen by a module*, for the same reason the node names are: a module listing the routes
|
||||
would go stale the day one changes. The mesh propagates the names it was told to serve and still
|
||||
knows nothing about what they mean
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: resolved
|
||||
status: located
|
||||
opened: 2026-09-23
|
||||
located-in: [mesh-controller internal/link, mesh-host internal/link]
|
||||
fixed-by: mesh-host PR 59 (the host refuses an older sequence and drains by it), mesh-controller PR 160 (each send is numbered under the node's hold) — measured 2026-09-30, 02-resolution.md
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
|
||||
@@ -66,11 +66,3 @@ Left `located`. The owner is unchanged, the shape of the fix is agreed, and the
|
||||
[issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md) rather than anything
|
||||
in this record. **This is a judgement about order, not a refusal** — it is cheap to overrule, and the
|
||||
code is a day's work once a host can be delivered.
|
||||
|
||||
## The gate has opened (2026-09-30, evening)
|
||||
|
||||
The mesh delivers the host now — built by its own toolchain, published to its own registry, delivered
|
||||
over the bus and started by the launcher, on all four machines
|
||||
([issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/01-progress.md)). A declaration
|
||||
field is a build and a push, not an expedition. The order this record asked for — hosts first, then
|
||||
the controller — is now two commands and a status line that says when the first has finished.
|
||||
|
||||
@@ -1,60 +0,0 @@
|
||||
# 107 — resolved: a declaration carries its order
|
||||
|
||||
*2026-09-30. Measured on the mesh.*
|
||||
|
||||
## What was done
|
||||
|
||||
**Hosts first, then the controller** — the order [issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md)
|
||||
says a new declaration field needs, and now a build and a push rather than an expedition
|
||||
([issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/01-progress.md)).
|
||||
|
||||
The host understands a `sequence` on a declaration and tolerates its absence: absent reads as "no
|
||||
order claimed", not "first", so a controller that sends none is still understood and a host that kept
|
||||
a declaration before it understood the field compares nothing. It refuses a declaration with a lower
|
||||
sequence than the one it kept, whole, and says why; and the drain that picks one declaration from a
|
||||
batch keeps the highest sequence rather than the last to arrive — which is the case the report
|
||||
constructed, a backlog drained out of order.
|
||||
|
||||
The controller numbers each send: the next number for that node, taken under the node's hold, before
|
||||
the body exists, so the number is inside what the mesh signs and a replayed older declaration cannot
|
||||
borrow a newer one's.
|
||||
|
||||
## Measured
|
||||
|
||||
```
|
||||
push shanks; push shanks
|
||||
sequence in kept declaration: 2
|
||||
node sequence
|
||||
novox 2
|
||||
shanks 2
|
||||
ace (none — not sent since numbering)
|
||||
g14 (none)
|
||||
status: nobody "not running what the mesh would send them"
|
||||
```
|
||||
|
||||
Both applies went through; neither was refused; the machine holding the earlier one accepted the later.
|
||||
|
||||
## The subtlety, which would have read every machine as behind for ever
|
||||
|
||||
The mesh decides a machine is behind by comparing the digest of what it **would** send against what it
|
||||
**did** send. A number changes the bytes. So the read-only comparison composes with the number the
|
||||
machine was *last* sent — not a fresh one — and is byte for byte what was sent when nothing else
|
||||
changed. Without that, numbering would have made `status` name all four machines as out of date on
|
||||
every reading, permanently.
|
||||
|
||||
## The open questions
|
||||
|
||||
- *A per-node `sequence` under the controller's node hold?* Yes, as described. **`supersedes` — the
|
||||
previous digest — is not added.** A strictly-greater sequence gives the ordering; a chain of digests
|
||||
would give continuity, which nothing here needs yet and which every re-composition would break.
|
||||
- *Genesis signing its bundle as sequence zero?* Zero is "no order claimed", which is what the bundle
|
||||
carries by carrying nothing. Same rule, no genesis branch.
|
||||
- *A marker for a mode change?* Not needed for the incident it guards: a replayed converged declaration
|
||||
reaching a node returned to adopted is already refused **by mode**, before this check runs.
|
||||
|
||||
## How it is checked
|
||||
|
||||
Host: an older sequence is refused, a newer or equal one is not, and no order claimed on either side
|
||||
compares nothing; the drain keeps the highest sequence, and falls back to arrival when none is claimed.
|
||||
Controller: a send carries its number inside the signed bytes, an unnumbered send is byte for byte what
|
||||
it was before, and each node's counter is one higher per send and readable for the comparison.
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: resolved
|
||||
status: located
|
||||
opened: 2026-09-24
|
||||
located-in: [mesh-controller internal/catalogue/declaration.go (every container was given the roster at creation)]
|
||||
fixed-by: mesh-controller PR 161 — no container is given a mesh name; it resolves through its machine's resolver (ADR 0148, landed 2026-09-30 once issue 110 did)
|
||||
located-in: [mesh-host internal/apply]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -73,13 +73,3 @@ copying: a container resolves through its machine's resolver at the moment it as
|
||||
record reports then has nowhere to occur. It is gated on
|
||||
[issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md), so
|
||||
until that lands the mesh still copies and still compares.
|
||||
|
||||
## Resolved (2026-09-30)
|
||||
|
||||
110 landed the same day ([its resolution](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/01-resolution.md)),
|
||||
and mesh-controller PR 161 then removed the copy: no container is given a mesh name or a mesh address,
|
||||
and a module's own declared entries are the only `host` lines it carries. Verified on the control-node
|
||||
after its containers were recreated once — the last time a name will do that: the forge's container
|
||||
carries no extra hosts and resolves another machine and a routed name through the machine's resolver,
|
||||
so the shape this record describes has nowhere to occur. Checked in the controller's tests: a
|
||||
container's declaration is byte-for-byte the same under a roster of one machine and a roster of three.
|
||||
|
||||
+3
-7
@@ -1,9 +1,8 @@
|
||||
---
|
||||
status: resolved
|
||||
status: open
|
||||
opened: 2026-09-24
|
||||
located-in:
|
||||
- mesh-catalog modules/dnsmasq (the runtime was never told; the resolver answered by interface)
|
||||
fixed-by: mesh-catalog PR 175 (the runtime is reloaded and keeps its containers over a restart) and PR 176 (the resolver answers by address, so a query from a bridge is admitted) — measured 2026-09-30, 01-resolution.md
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -68,6 +67,3 @@ and [135](../135-a-containers-mesh-names-are-not-compared/00-report.md)).
|
||||
network is the case with no DNS at all, and it is the case the mesh's own forge runs in. Two of four
|
||||
machines also bind the resolver to loopback only, so the runtime hands their containers a public
|
||||
resolver. Both halves are this issue.
|
||||
|
||||
*Later the same day: the second half was wrong, and the first had a different cause than the one above.
|
||||
[01-resolution.md](01-resolution.md) has what was actually found.*
|
||||
|
||||
-64
@@ -1,64 +0,0 @@
|
||||
# 110 — resolved: a container on any network reaches the resolver, and is answered
|
||||
|
||||
*2026-09-30. Measured on the three converged machines; the adopted one holds its resolver module until it
|
||||
is taken and is not covered.*
|
||||
|
||||
## What was actually wrong
|
||||
|
||||
Not what the report predicted. The report named the filter: a container on the runtime's default
|
||||
network asks from a bridge address, and the converged filter admitted queries by source address only.
|
||||
That was true when it was written and was fixed before this issue was ever tested — the filter admits
|
||||
by the link a packet arrives on ([ADR 0144](../../02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md)),
|
||||
and a container's bridge is admitted whole. Tested on every machine: the query arrives, the filter
|
||||
passes it.
|
||||
|
||||
Three other things were wrong, each hiding the next.
|
||||
|
||||
**The runtime had never been told.** The resolver module writes the runtime's `dns` key into the
|
||||
runtime's own configuration file. The runtime reads that key when it starts and not on a reload, and on
|
||||
two machines the runtime predated the file — so every container they started got a public resolver, and
|
||||
`novox.internal` came back as not existing. Nothing reported this: the file was present and current,
|
||||
the resolver ran, and a name not existing is a valid answer. Fixed in mesh-catalog PR 175: the module
|
||||
also sets `live-restore` and reloads the runtime when its file changes, so the one restart the `dns` key
|
||||
needs no longer stops every container. The restart is then the operator's, once per machine; done on
|
||||
both today, with every running container kept.
|
||||
|
||||
**The resolver dropped the query.** With the runtime corrected, a container's query reached the resolver
|
||||
— and got no answer, on every machine, including the one whose runtime had been right all along. The
|
||||
socket was bound to the private address; the filter admitted the packet; dnsmasq received it and
|
||||
discarded it without a line of log. Its configuration said `interface=mesh0`, and dnsmasq admits a
|
||||
query by the interface it arrives on when told an interface: a container's query is addressed to the
|
||||
private address but arrives on the runtime's bridge, and the bridge is not `mesh0`. Fixed in mesh-catalog
|
||||
PR 176: the resolver is told the address to answer on, not the interface that carries it, and a query to
|
||||
that address is admitted whatever bridge brings it. The bridges are the runtime's to name.
|
||||
|
||||
**The report's second half was wrong.** "Two of four machines bind the resolver to loopback only" was
|
||||
an inference from the containers' behaviour, and the behaviour had the cause above. The resolver bound
|
||||
the private address on all four; nothing had asked it there.
|
||||
|
||||
## What is verified
|
||||
|
||||
From a container on the runtime's default network, started by hand and given nothing, on each of the
|
||||
three converged machines: `novox.internal` answers with the hub's private address, through the machine's
|
||||
own resolver. That is the fourth check of
|
||||
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) — "on every
|
||||
network the runtime offers" — and its first step; the record's step 2 (the runtime told per machine, as
|
||||
a file) was already how the module works. Step 3 may now begin.
|
||||
|
||||
## What checks it
|
||||
|
||||
By hand, today. Nothing in the mesh asserts that a container can resolve a mesh name: the resolver's
|
||||
own tests cover what it answers, not who can ask. The check that would have caught all three faults is
|
||||
the one the report asked for and 0148 lists — a container on the default network resolving a mesh name
|
||||
— and it is not built. It belongs with the reachability check of
|
||||
[issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
|
||||
which is parked; until then this is a thing a person verifies after touching the resolver, the filter,
|
||||
or the runtime's configuration.
|
||||
|
||||
## What this cost to find
|
||||
|
||||
The three faults produced one symptom — a container that cannot resolve — and each fix revealed the
|
||||
next. The first was found by reading the runtime's own view of its configuration rather than the file;
|
||||
the second by capturing the query on the bridge and finding it arrive and go unanswered; the third only
|
||||
by admitting the first belief was wrong. A machine that had been believed to work all day had never
|
||||
worked either.
|
||||
@@ -1,9 +1,9 @@
|
||||
---
|
||||
status: resolved
|
||||
status: open
|
||||
opened: 2026-09-28
|
||||
located-in: [mesh-controller internal/catalogue/declaration.go (composeName took the consumer's own name as the internal domain)]
|
||||
fixed-by: mesh-controller PR 163 — the internal name composes under the node whose proxy serves the route (ADR 0151, 2026-09-30)
|
||||
amended-design: 03-DESIGN/01-to-be/08-connectivity.md
|
||||
located-in: [mesh-controller internal/catalogue]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 139 — An internal route name resolves to the consumer's node, not the one that serves it
|
||||
@@ -49,15 +49,3 @@ resolve whether or not anything answers.
|
||||
and if so, is `route` still one mesh-wide provision or a node-scoped seat with a mesh-wide fallback?
|
||||
- What certifies the name in either case? The certificate is obtained by whoever terminates TLS, and
|
||||
that is the question above in another form.
|
||||
|
||||
## Answered (2026-09-30)
|
||||
|
||||
[ADR 0151](../../02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md):
|
||||
the internal name is composed under the node that serves the route — the machine the request arrives
|
||||
at — because `<x>.<node>.internal` means *goes to that node* and nothing else. The public name stays
|
||||
the consumer node's, which is where the operator put it. The first question is answered that way; the
|
||||
second, a per-node route holder, is a decision about seats and is left where it is; the third is
|
||||
unchanged, since the proxy that terminates the name is given it and certifies it.
|
||||
|
||||
mesh-controller PR 163 carries it. On this mesh every route is served beside its module, so no name
|
||||
changed; the controller's tests hold the case where it would.
|
||||
|
||||
@@ -138,48 +138,3 @@ push. The control node is worth last.
|
||||
One thing this found on the way out: an archive cannot be undeclared, and the attempt stops the machine
|
||||
applying anything at all — [issue 162](../162-an-archive-cannot-be-undeclared/00-report.md). It is how
|
||||
undoing the first delivery froze the workstation, and it is not specific to the host.
|
||||
|
||||
## Every machine self-updates (2026-09-30, evening)
|
||||
|
||||
```
|
||||
shanks 76f4566bef3d/nox-mesh-host active
|
||||
g14 76f4566bef3d/nox-mesh-host active
|
||||
novox 76f4566bef3d/nox-mesh-host active
|
||||
ace 76f4566bef3d/nox-mesh-host active
|
||||
|
||||
mesh-controller status: (no host split)
|
||||
```
|
||||
|
||||
The last delivery was unattended on all four: the fixed host was built, pushed, each machine stood
|
||||
aside exactly once for the genuinely newer version, and the delivered launcher started it — no
|
||||
restart by hand. A following push that delivered nothing new was applied and reported by every
|
||||
machine and stood nobody aside, which is the check
|
||||
[issue 163](../163-a-delivered-host-stood-aside-on-every-push-and-reported-nothing/00-report.md) asks
|
||||
for.
|
||||
|
||||
**Two more faults on the way, both mine, both found by reading the machine rather than the success
|
||||
line.** A delivered host compared the newest delivered version against its link-time stamp rather
|
||||
than the version it was running, so it stood aside on every push and — because standing aside cancels
|
||||
the report — never reported again (163). And the adopted machine kept its found launcher as the
|
||||
adoption rule says, so the delivery there needed a `take` before the launcher moved.
|
||||
|
||||
**The crossover needs one restart of the unit per machine, once.** The launcher process that was
|
||||
running on each machine was the old script, executing from its own inode; a new file beside it
|
||||
changes nothing until the unit restarts. Every subsequent delivery is unattended.
|
||||
|
||||
**Timing, measured:** on a machine, hearing a declaration to reporting it applied is about three
|
||||
seconds. A push as the operator sees it takes 17–20 seconds, and the difference is the control plane
|
||||
composing the declaration before it sends. A `--wait` shorter than that reads as "did not report" for
|
||||
a machine that did; the three-minute default read as slowness for a machine that never would. Neither
|
||||
number is a defect being chased here, and both are worth knowing before reading a push's answer.
|
||||
|
||||
## What this leaves
|
||||
|
||||
- [Issue 162](../162-an-archive-cannot-be-undeclared/00-report.md): an archive cannot be undeclared, so
|
||||
the host module — and any module with an archive — cannot be unassigned, and trying stops the machine
|
||||
applying anything.
|
||||
- [Issue 107](../107-a-declaration-carries-no-order/00-report.md) is unblocked: a declaration field is
|
||||
now a build and a push rather than an expedition.
|
||||
- Three stale version directories on the workstation from the first attempts, moved aside under
|
||||
`/var/lib/mesh-host/versions-held-back/`, and a backup of the adopted machine's hand-placed binary
|
||||
beside its state. Both are safe to delete and are not the mesh's to delete.
|
||||
|
||||
@@ -1,10 +1,10 @@
|
||||
---
|
||||
status: resolved
|
||||
status: open
|
||||
opened: 2026-09-29
|
||||
located-in:
|
||||
- mesh-host internal/apply/apply.go (containerSpecReading hashes every `host` entry)
|
||||
- mesh-controller internal/catalogue/declaration.go (withMeshNames gives every container the mesh's names)
|
||||
fixed-by: mesh-controller PR 161 — the roster left every container's declaration and so its digest (ADR 0148 step 3, 2026-09-30)
|
||||
fixed-by:
|
||||
---
|
||||
|
||||
# 151 — A new name recreates every container in the mesh
|
||||
@@ -92,15 +92,3 @@ copying names until a container can reach the resolver from any of the runtime's
|
||||
([issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)),
|
||||
which it cannot on two of four machines today. Removing the copy first reintroduces 109 and 135
|
||||
silently, on a live mesh, which is how both were found. The order is in the record.
|
||||
|
||||
## Resolved (2026-09-30)
|
||||
|
||||
Step 3 landed the day 110 did. mesh-controller PR 161 stops writing the roster into any container, so
|
||||
a container's digest no longer carries a name that is not its own. The controller's tests hold the
|
||||
record's check — a container's declaration does not move when the mesh's roster does, and does move
|
||||
when the module's own declared entries do.
|
||||
|
||||
The first push after the change recreated every container once, because every digest lost its host
|
||||
entries at the same moment. That was the last such event: from here a name added or moved on one
|
||||
machine changes no container anywhere, and the record's second check — add a routed name, watch every
|
||||
other machine's apply report say nothing changed — is what the next module assignment will show.
|
||||
|
||||
@@ -1,9 +1,9 @@
|
||||
---
|
||||
status: resolved
|
||||
status: located
|
||||
opened: 2026-09-30
|
||||
located-in:
|
||||
- mesh-controller internal/catalogue/roster.go (the roster's entries for a routed name)
|
||||
fixed-by: mesh-controller PR 163 — a routed name is published as itself, once, with no suffixed alias (ADR 0151, 2026-09-30)
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -61,10 +61,3 @@ composition produces `keycloak.novox.be.internal`, which is not a name anything
|
||||
|
||||
The fix is a judgement about what a routed name's internal form is, and 139 is the record that asks it;
|
||||
this one is the evidence that the current answer publishes a third thing that is neither.
|
||||
|
||||
## Resolved (2026-09-30)
|
||||
|
||||
The judgement 139 asked for is [ADR 0151](../../02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md):
|
||||
a routed name has no mesh form. The roster now publishes it as itself, once, at the serving node's
|
||||
address; the `<domain>.internal` line is gone from every machine's hosts file, and a controller test
|
||||
refuses it coming back.
|
||||
|
||||
-56
@@ -1,56 +0,0 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-30
|
||||
located-in: [mesh-host cmd/mesh-host/main.go (the successor check after an apply)]
|
||||
fixed-by: mesh-host PR 58 — the check asks with the running version, read from the binary's path, not the link-time stamp
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 163 — A delivered host stood aside on every push, and reported nothing
|
||||
|
||||
## What was observed
|
||||
|
||||
*2026-09-30, rolling the mesh-built host onto the last two machines.*
|
||||
|
||||
Every push to a machine running a delivered host produced, in order:
|
||||
|
||||
```
|
||||
host 093231796eb0 is delivered; standing aside so the launcher runs it
|
||||
applied 333 resource(s)
|
||||
applied, and could not tell the mesh: reporting: context canceled
|
||||
nox-mesh-host-launch: the host exited cleanly; starting it again
|
||||
nox-mesh-host-launch: running /usr/lib/nox-mesh-host/versions/093231796eb0/nox-mesh-host
|
||||
```
|
||||
|
||||
— for the version it was **already running**. It restarted itself on every push, for ever, and the mesh
|
||||
never received a single report from it: `node show` kept the version from before the crossover, and
|
||||
the operator's push waited its full three minutes for an answer that was never coming.
|
||||
|
||||
Read as healthy throughout: unit active, bus link up, "hearing what this node should be".
|
||||
|
||||
## Why
|
||||
|
||||
After an apply the host asks whether a newer host has been delivered than the one running, and the
|
||||
question was asked with the **link-time version stamp**. Since
|
||||
[issue 161](../161-a-delivered-host-carries-none-of-its-link-time-facts/01-resolution.md) a delivered
|
||||
host's version comes from where it sits and its stamp is `development build` — so the comparison never
|
||||
matched the newest delivered version, and "a newer host is waiting" was always true.
|
||||
|
||||
Standing aside cancels the context the report is published with, so the report was lost on every one
|
||||
of those applies. Two faults from one wrong argument.
|
||||
|
||||
The change that moved the version to the path was applied to the report and to the known-good record,
|
||||
and not here. Half a change, and the half left behind was the one that decides whether to exit.
|
||||
|
||||
## Why the three-minute wait made it invisible
|
||||
|
||||
The push's `--wait` timing out read as *slow*. It was not slow: **the report was never going to arrive.**
|
||||
The operator put it exactly: *if you don't get a response in five seconds, something is wrong.* A wait
|
||||
long enough to absorb a machine's whole apply is a wait long enough to hide that the machine never
|
||||
answered.
|
||||
|
||||
## How it is checked
|
||||
|
||||
A machine running a delivered host is pushed a declaration that delivers nothing new; it applies,
|
||||
reports, and does not stand aside. A machine running a delivered host is pushed a genuinely newer
|
||||
version; it stands aside once, and the next push it does not.
|
||||
@@ -0,0 +1,29 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-30
|
||||
located-in:
|
||||
- mesh-controller internal/inventory/secrets.go (SecretFor mints a pair credential nobody accepted)
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 164 — A credential that must be accepted is minted anyway
|
||||
|
||||
## What was observed
|
||||
|
||||
Provisioning ace's modules. Several providers hold exactly one credential they did not get from the
|
||||
mesh and cannot take one from it: a Servarr app's API key (sonarr, radarr, lidarr), jackett's API key,
|
||||
plex's X-Plex-Token, nzbget's ControlPassword, qBittorrent's WebUI password. Their consumers' pair
|
||||
credential must be **accepted** by the operator (ADR 0092). Until it is, `SecretFor` mints a random
|
||||
value, seals it to both ends, and reports nothing: the value can never work.
|
||||
|
||||
Every consumer therefore had to learn to detect it — try the credential against the provider first,
|
||||
refuse a value the provider rejects, print the `secret accept` command — six write-in steps, one probe
|
||||
each (ombi, home-assistant, and the four download-stack consumers). qBittorrent bans an address after
|
||||
five failed logins, so a consumer retrying a minted value locks itself out.
|
||||
|
||||
## What would be right
|
||||
|
||||
A provision (or a provider's `serves`) can declare its pair credential **accepted-only**. The plan then
|
||||
refuses the pair — naming the accept command — instead of minting, and a consumer is never handed a
|
||||
value the mesh knows cannot work.
|
||||
@@ -0,0 +1,24 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-30
|
||||
located-in:
|
||||
- mesh-controller internal/inventory/secrets.go (AcceptSecretForPair is per consumer)
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 165 — One accepted value must be accepted once per consumer
|
||||
|
||||
## What was observed
|
||||
|
||||
On ace, jackett's API key is the pair credential for sonarr, radarr, lidarr and bookshelf; sonarr's is
|
||||
the credential for ombi, bazarr and home-assistant. It is **one value**, owned by the provider — yet
|
||||
`secret accept` is per pair, so ace's download stack alone needs 12 accepts of 3 values, and rotating
|
||||
a provider's key means finding and re-accepting every pair. Missing one leaves that consumer on a
|
||||
stale (or minted, 164) value.
|
||||
|
||||
## What would be right
|
||||
|
||||
A provider-level accept: "this provider's credential for `<provision>` is X" — delivered to every
|
||||
consumer pair, current and future, and rotated in one place. Pairs whose credential is genuinely per
|
||||
consumer (postgres, keycloak, mosquitto, influxdb — minted and created by a provisioner) are unaffected.
|
||||
@@ -0,0 +1,24 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-30
|
||||
located-in:
|
||||
- mesh-controller internal/catalogue (requires is a list of hard requirements)
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 166 — A requirement cannot be optional
|
||||
|
||||
## What was observed
|
||||
|
||||
Making every dependency on ace a provision turned soft dependencies into hard ones. grafana now
|
||||
requires `influxdb-api` (a data source), ombi requires `sonarr-api`, `radarr-api` and `lidarr-api`,
|
||||
home-assistant requires the Servarr APIs and `mqtt-topic`. Each is optional to the software — grafana
|
||||
runs without a data source, ombi without lidarr — but a mesh without influxdb cannot assign grafana at
|
||||
all, and a mesh without lidarr cannot run ombi.
|
||||
|
||||
## What would be right
|
||||
|
||||
A requirement a module can run without: resolved and bound when a provider exists, absent (with its
|
||||
`${bound:…}` placeholders refused or defaulted explicitly, never rendered empty) when none does — so
|
||||
the module description stays true on every mesh.
|
||||
@@ -0,0 +1,26 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-30
|
||||
located-in:
|
||||
- mesh-catalog (each module builds from its own directory, ADR 0069)
|
||||
- mesh-sdk
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 167 — Code several modules share has no home
|
||||
|
||||
## What was observed
|
||||
|
||||
The download-stack write-in step (register download clients and torznab indexers through the Servarr
|
||||
API) is identical for sonarr, radarr, lidarr and bookshelf. Because a module builds from its own
|
||||
directory, it now exists as four byte-identical copies under `modules/<m>/downloads/`, kept honest by a
|
||||
test that fails when one differs. The same shape repeats: an MQTT probe copied into two modules, and a
|
||||
"write the provider into the app through its API, idempotently, refuse a minted value" step in ombi,
|
||||
home-assistant, nodered, tautulli and the four downloaders.
|
||||
|
||||
## What would be right
|
||||
|
||||
A home for shared module code the builder can use — an sdk helper (a write-in step harness: read
|
||||
bindings and pair credentials, probe the provider, diff, write, report) or a shared package the
|
||||
catalogue builds once — so a fix lands in one place.
|
||||
@@ -0,0 +1,32 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-30
|
||||
located-in:
|
||||
- mesh-controller internal/catalogue/settings.go (settle: every key but `ports` merges into every mergeable file and every contribution)
|
||||
- mesh-controller internal/catalogue/declaration.go (a provider's settings are laid over what it serves)
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 168 — A setting reaches every file and every contribution
|
||||
|
||||
## What was observed
|
||||
|
||||
Settings merge key by key into **every** `"merge": "json"` file of a module **and** every contribution
|
||||
it makes; a provider's settings are also laid over what it serves. Seen on ace:
|
||||
|
||||
- searxng's `endpoints` and a route `label` land in searxng's own `settings.yml`; nodered's
|
||||
`timeZone` and `mqtt` keys land in mosquitto's grants file; keycloak's `issuer` lands in its
|
||||
`postgres-database` and `route` contributions.
|
||||
- every consumer's `plex-api` binding carries plex's `endpoints` and `expose` settings — and a provider
|
||||
setting named `port` would silently redirect every consumer.
|
||||
- a module cannot have two configurable files: searxng's sidecar config had to stop being mergeable
|
||||
so searxng's keys would not reach it.
|
||||
|
||||
Harmless today only because every receiver happens to ignore unknown keys.
|
||||
|
||||
## What would be right
|
||||
|
||||
A setting is aimed: at a file (by resource id), at a contribution (by requirement), or at what the
|
||||
module serves — declared settable by the module (ADR 0046 already says settings drive "the fields the
|
||||
manifest marks") — and an unaimed key is refused like any unknown setting.
|
||||
@@ -1,78 +0,0 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-30
|
||||
located-in:
|
||||
- mesh-catalog modules/mailu (eight containers name Mailu's own resolver, and one of them binds a mesh name)
|
||||
- mesh-catalog modules/dnsmasq (dropped the DNSSEC bit its upstreams set)
|
||||
fixed-by: mesh-catalog PR 178 (mailu-admin uses the machine's resolver) and PR 179 (the machine's resolver passes the DNSSEC bit down) — 2026-09-30, the same afternoon
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 171 — A module that names its own resolver knows no mesh name
|
||||
|
||||
## What was observed
|
||||
|
||||
The afternoon [ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
|
||||
landed — no container is given the mesh's names any more; it asks the machine's resolver — the mail
|
||||
system's admin container began logging, 523 times in three minutes:
|
||||
|
||||
```
|
||||
psycopg2.OperationalError: could not translate host name "novox.internal" to address: Name does not resolve
|
||||
```
|
||||
|
||||
Mail was accepted on every port and the web front answered; the admin and the spam filter beside it
|
||||
were unhealthy, and anything that needed the database — a mailbox change through the API, the spam
|
||||
filter's domain list — failed. Found by the operator asking whether mail was back, forty minutes in.
|
||||
|
||||
Mailu ships its own resolver, an unbound in a container, and every other Mailu container is told to
|
||||
use it — the module carries `dns: [192.168.203.254]` on eight containers. That resolver recurses from the
|
||||
root and knows nothing under `.internal`. Until that afternoon the admin container had the database's
|
||||
name anyway, because the mesh wrote every name into every container at creation; the copy was the only
|
||||
reason a container behind its own resolver could reach anything by a mesh name, and nobody knew it was
|
||||
load-bearing.
|
||||
|
||||
**Removing the override was not enough.** Given the machine's resolver instead, the admin refused to
|
||||
start: `Your DNS resolver at 127.0.0.11 isn't doing DNSSEC validation`. Mailu checks, at start, that
|
||||
its resolver returns the Authenticated Data bit for a signed name. The mesh's resolver forwards to two
|
||||
upstreams that validate and set the bit, and dropped it on the way down — dnsmasq does unless told
|
||||
otherwise. Mailu's own unbound has no hook to forward a zone elsewhere, so it could not be taught the
|
||||
mesh's names either.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
**A container with a resolver of its own has opted out of the machine's, and nothing says so.** 0148
|
||||
made the machine's resolver load-bearing for every container; a `dns` on a container is a quiet
|
||||
exception to that, and the exception used to be papered over by the copy the record removed. The
|
||||
manifest field reads like a preference and is a decision about whether mesh names exist inside the
|
||||
container.
|
||||
|
||||
**A resolver that forwards to validating upstreams and hides the fact is less useful than it could
|
||||
be**, and the first program to check found out.
|
||||
|
||||
**The mesh reported nothing.** Every container ran; the failing one accepted connections; the report
|
||||
was about bytes. It is [issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
|
||||
again, and the check that would have caught it is the same unbuilt one.
|
||||
|
||||
## What was done
|
||||
|
||||
- The one Mailu container that binds a mesh name — the admin, through the database it is granted —
|
||||
no longer names Mailu's resolver and uses the machine's, like every container without a `dns` of
|
||||
its own (mesh-catalog PR 178). The other seven keep unbound: the spam filter needs a validating
|
||||
resolver for its blocklist lookups, and none of them asks for a mesh name.
|
||||
- The machine's resolver passes the DNSSEC bit down from its upstreams, `proxy-dnssec` (PR 179). It
|
||||
does not validate itself; the trust is the upstream's and the path to it, as a forwarding resolver's
|
||||
always was, and the configuration says so.
|
||||
|
||||
## What checks it
|
||||
|
||||
The admin container's own start-up check, which is what failed, and the mesh's status once it reads
|
||||
healthy. A container-level check that a mesh name resolves from inside every declared container is
|
||||
the one 110 and 145 both ask for and is not built.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should a container's `dns` be refused, or made to say what it gives up? A module that names its
|
||||
own resolver and binds a mesh name is a contradiction the controller can see at composition — the
|
||||
grant hands it a name its resolver will not answer.
|
||||
- Should the machine's resolver validate rather than proxy? It would cost a trust anchor on every
|
||||
machine and make the resolver slower to start; proxying was enough for the one program that asked.
|
||||
Reference in New Issue
Block a user