Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
105ae9a56a | ||
|
|
ee801a6441 | ||
|
|
90b89aa1c9 | ||
|
|
e33191161d | ||
|
|
a0a930b1cd | ||
|
|
8d9c9ab6b5 |
@@ -103,13 +103,6 @@ reintroduces 109 and 135 — silently, and on a live mesh, which is exactly how
|
||||
([issue 110](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md));
|
||||
and on two of four machines the resolver binds loopback only, so the runtime hands containers a
|
||||
public resolver instead. Both are prerequisites, not related work.
|
||||
|
||||
> **Progressive insight — 2026-09-30, later the same day. The loopback claim was wrong.** The
|
||||
> resolver bound the private address on all four machines; on two the runtime had never been told
|
||||
> to use it, and on all four the resolver discarded a query that arrived on the runtime's bridge.
|
||||
> The step stands; the facts under it were those. Both fixed the same day
|
||||
> ([issue 110's resolution](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/01-resolution.md)),
|
||||
> and step 3 landed after them.
|
||||
2. **The runtime is told which resolver to use, per machine, as a file** — not per container as a
|
||||
creation-time argument, or the resolver's address is back in every container's identity and the
|
||||
problem has only got smaller.
|
||||
|
||||
@@ -1,99 +0,0 @@
|
||||
---
|
||||
topic: the tiers
|
||||
status: accepted
|
||||
date: 2026-09-30
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 02-DECISIONS/0066-public-routing-is-name-agnostic.md
|
||||
---
|
||||
|
||||
# 151. A route's internal name is composed under the node that serves it
|
||||
|
||||
## Context
|
||||
|
||||
A module that requires a route is given two names from one label: a public one, `<label>.<public
|
||||
domain>`, and an internal one, `<label>.<node>.internal`
|
||||
([ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md)). Both were composed
|
||||
from the node the module runs on.
|
||||
|
||||
The two are answered differently. The public name is published into every machine's roster at the
|
||||
address of the node whose proxy serves it ([ADR 0066](0066-public-routing-is-name-agnostic.md)), so
|
||||
it reaches the proxy from anywhere in the mesh. The internal name is answered by every machine's
|
||||
resolver as *anything under a node's name goes to that node*
|
||||
([design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md)) — the node it was composed from, which is
|
||||
the consumer's. Where the proxy runs on another machine, that name sends a client to a machine with
|
||||
nothing listening, while the public name works
|
||||
([issue 139](../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md)).
|
||||
Every route on this mesh today is served beside its module, so it has not been seen; `route` is
|
||||
provided mesh-wide precisely so that stops being true.
|
||||
|
||||
Beside it, the roster gave every routed name a second entry with the mesh's suffix appended —
|
||||
`<name>.<public domain>.internal` — because it composed a full name for every entry as it does for a
|
||||
machine. That name resolved on every machine, was served by nothing, and was refused by the proxy at
|
||||
the handshake; the first three names tried while reproducing an unrelated issue were those, and the
|
||||
evidence pointed at a regression that had not happened
|
||||
([issue 157](../04-ISSUES/157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md)).
|
||||
|
||||
## Considered Options
|
||||
|
||||
**1. Keep the consumer's name and publish it at the serving node's address**, as the public name is.
|
||||
The name stays `<label>.<consumer>.internal` and an exact roster entry overrides the wildcard.
|
||||
Rejected: it makes `<x>.<node>.internal` mean *goes to that node* except when it does not, which is
|
||||
the one rule the resolver design states; it needs an entry per route where the wildcard needed none;
|
||||
and which of an exact entry and a wildcard a resolver answers first is the resolver's business, which
|
||||
the mesh deliberately does not know.
|
||||
|
||||
**2. A proxy on every machine, so the serving node is always the consumer's.** Rejected for this
|
||||
question: it is a different decision about what `route` is — a node-scoped seat with a mesh-wide
|
||||
fallback — and this mesh runs one proxy on the hub today. Whatever is decided there, a route served
|
||||
from another machine must have a name that reaches it.
|
||||
|
||||
**3. Compose the internal name under the node that serves the route.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**A route's internal name is `<label>.<serving node>.internal` — composed under the node whose proxy
|
||||
answers the route, which is the machine the request arrives at.** The public name is unchanged:
|
||||
`<label>.<public domain>` of the node the module runs on, which is where the operator put it.
|
||||
|
||||
Where the proxy runs beside the module — every route on this mesh today — the two nodes are one and
|
||||
nothing changes. Where it does not, the name says where the request goes, which is what a name under
|
||||
a node's name has always meant.
|
||||
|
||||
**A routed name has no mesh form.** The roster publishes it as itself, once, at the serving node's
|
||||
address. Only a machine has a bare name beside its full one.
|
||||
|
||||
What certifies the internal name is unchanged by this: the proxy that terminates it obtains a
|
||||
certificate from the mesh's authority for the names it is given, and it is given this one.
|
||||
|
||||
Taken on the operator's standing instruction to answer the open design questions in the work order.
|
||||
|
||||
## How this is checked
|
||||
|
||||
- **Composition.** A controller test contributes a route from a module on one node to a proxy offered
|
||||
from another, gathered the way the controller gathers a consumer's contribution for a provider on
|
||||
another machine, and asserts the internal name carries the serving node.
|
||||
- **Publication.** A controller test renders a roster with a machine and a routed name and asserts
|
||||
the routed name appears as itself, once, and never with the suffix appended.
|
||||
- **On the mesh.** After the change no machine's roster carries a `<domain>.internal` entry, and a
|
||||
route's internal name still answers from a container with a certificate from the mesh's authority.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **A route served from another machine now has a usable internal name.** The first module assigned
|
||||
that way will resolve, where before it would have resolved to the wrong machine with no error.
|
||||
- **The internal name of a route can change when its proxy moves.** A route re-homed from one proxy
|
||||
to another gets a new internal name, as the design's rule implies; clients that dialled the old one
|
||||
reach the old machine. The public name does not move with the proxy and is the stable one.
|
||||
- **The roster is one line shorter per routed name**, and a person reading a hosts file no longer
|
||||
finds names that resolve to a refusal.
|
||||
- **Issue 139's second question — a per-node route holder — is left open**, and is a decision about
|
||||
what a seat is rather than about a name.
|
||||
|
||||
## References
|
||||
|
||||
- [issue 139](../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md) — the question
|
||||
- [issue 157](../04-ISSUES/157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md) — the alias
|
||||
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — routed names propagate mesh-wide; extended here
|
||||
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — how the two names are composed and how far each reaches
|
||||
- [design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md) — anything under a node's name goes to that node
|
||||
@@ -200,7 +200,6 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0109** — [A package registry seat is one per ecosystem, not one for all of them](0109-a-package-registry-seat-is-one-per-ecosystem.md)
|
||||
- **0126** — [A module declares its own seats; the mesh reserves its own](0126-a-module-declares-its-own-seats.md)
|
||||
- **0148** — [The mesh's names are resolved, not copied into every container](0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
|
||||
- **0151** — [A route's internal name is composed under the node that serves it](0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md)
|
||||
|
||||
### What runs on them, and how it gets there
|
||||
|
||||
|
||||
@@ -10,7 +10,6 @@ code:
|
||||
updated: 2026-09-30
|
||||
decisions:
|
||||
- 02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md
|
||||
- 02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md
|
||||
- 02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md
|
||||
- 02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md
|
||||
- 02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md
|
||||
@@ -306,11 +305,10 @@ and [135](../../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report
|
||||
|
||||
**A container resolves the mesh's names through its machine's resolver, at the moment it asks, and
|
||||
nothing is copied.** The resolver is a machine-level process rather than a container, so nothing
|
||||
circular is being asked for. It was gated on a container being able to reach the resolver from any of
|
||||
the runtime's networks
|
||||
([issue 110](../../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)),
|
||||
and landed the day that did, 2026-09-30: the controller writes no mesh name into a container and the
|
||||
host's digest carries only what the module declared for itself.
|
||||
circular is being asked for. This is gated on a container being able to reach the resolver from any of
|
||||
the runtime's networks, which it cannot today
|
||||
([issue 110](../../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)) —
|
||||
until that lands the mesh keeps copying and keeps comparing, and the order is stated in the record.
|
||||
|
||||
The paragraph below states the old boundary, and 0148 deliberately gives it up: a container somebody
|
||||
started by hand resolves the same names as everything else, because the resolver answers the machine,
|
||||
@@ -326,14 +324,6 @@ machine — declared or not — is what a nameserver in `resolv.conf` would be f
|
||||
service, the rest is the node — so what resolves is *anything under a node's name*, going to that
|
||||
node. What routes it once it arrives is a proxy's, and stays separate.
|
||||
|
||||
*2026-09-30.* **So the node in a route's internal name is the one whose proxy answers it**
|
||||
([ADR 0151](../../02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md)).
|
||||
Composed from the node the module ran on, the name sent a client to a machine with nothing listening
|
||||
whenever the proxy ran elsewhere
|
||||
([issue 139](../../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md));
|
||||
composed from the serving node, the rule above holds without exception. The public name stays the
|
||||
module's node's, which is where the operator put it.
|
||||
|
||||
**The mesh writes the data and runs no daemon.** One wildcard per machine, from the same set that
|
||||
writes the hosts file. A resolver is third-party software and runs *on* the mesh rather than being
|
||||
*of* it: the mesh has no business shipping one, choosing which one, or knowing its configuration
|
||||
@@ -409,9 +399,7 @@ can reach from the outside but cannot resolve from the inside is a name it canno
|
||||
authority of its own.
|
||||
|
||||
**So a granted route is published into internal resolution as well** — the routed name to the node
|
||||
that serves it, mesh-wide, by the same mechanism that writes the node names — and as itself: a routed
|
||||
name has no mesh form, and the suffixed alias the roster once added beside it resolved to a refusal
|
||||
([issue 157](../../04-ISSUES/157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md)). It is *given by the
|
||||
that serves it, mesh-wide, by the same mechanism that writes the node names. It is *given by the
|
||||
mesh, not chosen by a module*, for the same reason the node names are: a module listing the routes
|
||||
would go stale the day one changes. The mesh propagates the names it was told to serve and still
|
||||
knows nothing about what they mean
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: resolved
|
||||
status: located
|
||||
opened: 2026-09-23
|
||||
located-in: [mesh-controller internal/link, mesh-host internal/link]
|
||||
fixed-by: mesh-host PR 59 (the host refuses an older sequence and drains by it), mesh-controller PR 160 (each send is numbered under the node's hold) — measured 2026-09-30, 02-resolution.md
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
|
||||
@@ -66,11 +66,3 @@ Left `located`. The owner is unchanged, the shape of the fix is agreed, and the
|
||||
[issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md) rather than anything
|
||||
in this record. **This is a judgement about order, not a refusal** — it is cheap to overrule, and the
|
||||
code is a day's work once a host can be delivered.
|
||||
|
||||
## The gate has opened (2026-09-30, evening)
|
||||
|
||||
The mesh delivers the host now — built by its own toolchain, published to its own registry, delivered
|
||||
over the bus and started by the launcher, on all four machines
|
||||
([issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/01-progress.md)). A declaration
|
||||
field is a build and a push, not an expedition. The order this record asked for — hosts first, then
|
||||
the controller — is now two commands and a status line that says when the first has finished.
|
||||
|
||||
@@ -1,60 +0,0 @@
|
||||
# 107 — resolved: a declaration carries its order
|
||||
|
||||
*2026-09-30. Measured on the mesh.*
|
||||
|
||||
## What was done
|
||||
|
||||
**Hosts first, then the controller** — the order [issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md)
|
||||
says a new declaration field needs, and now a build and a push rather than an expedition
|
||||
([issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/01-progress.md)).
|
||||
|
||||
The host understands a `sequence` on a declaration and tolerates its absence: absent reads as "no
|
||||
order claimed", not "first", so a controller that sends none is still understood and a host that kept
|
||||
a declaration before it understood the field compares nothing. It refuses a declaration with a lower
|
||||
sequence than the one it kept, whole, and says why; and the drain that picks one declaration from a
|
||||
batch keeps the highest sequence rather than the last to arrive — which is the case the report
|
||||
constructed, a backlog drained out of order.
|
||||
|
||||
The controller numbers each send: the next number for that node, taken under the node's hold, before
|
||||
the body exists, so the number is inside what the mesh signs and a replayed older declaration cannot
|
||||
borrow a newer one's.
|
||||
|
||||
## Measured
|
||||
|
||||
```
|
||||
push shanks; push shanks
|
||||
sequence in kept declaration: 2
|
||||
node sequence
|
||||
novox 2
|
||||
shanks 2
|
||||
ace (none — not sent since numbering)
|
||||
g14 (none)
|
||||
status: nobody "not running what the mesh would send them"
|
||||
```
|
||||
|
||||
Both applies went through; neither was refused; the machine holding the earlier one accepted the later.
|
||||
|
||||
## The subtlety, which would have read every machine as behind for ever
|
||||
|
||||
The mesh decides a machine is behind by comparing the digest of what it **would** send against what it
|
||||
**did** send. A number changes the bytes. So the read-only comparison composes with the number the
|
||||
machine was *last* sent — not a fresh one — and is byte for byte what was sent when nothing else
|
||||
changed. Without that, numbering would have made `status` name all four machines as out of date on
|
||||
every reading, permanently.
|
||||
|
||||
## The open questions
|
||||
|
||||
- *A per-node `sequence` under the controller's node hold?* Yes, as described. **`supersedes` — the
|
||||
previous digest — is not added.** A strictly-greater sequence gives the ordering; a chain of digests
|
||||
would give continuity, which nothing here needs yet and which every re-composition would break.
|
||||
- *Genesis signing its bundle as sequence zero?* Zero is "no order claimed", which is what the bundle
|
||||
carries by carrying nothing. Same rule, no genesis branch.
|
||||
- *A marker for a mode change?* Not needed for the incident it guards: a replayed converged declaration
|
||||
reaching a node returned to adopted is already refused **by mode**, before this check runs.
|
||||
|
||||
## How it is checked
|
||||
|
||||
Host: an older sequence is refused, a newer or equal one is not, and no order claimed on either side
|
||||
compares nothing; the drain keeps the highest sequence, and falls back to arrival when none is claimed.
|
||||
Controller: a send carries its number inside the signed bytes, an unnumbered send is byte for byte what
|
||||
it was before, and each node's counter is one higher per send and readable for the comparison.
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: resolved
|
||||
status: located
|
||||
opened: 2026-09-24
|
||||
located-in: [mesh-controller internal/catalogue/declaration.go (every container was given the roster at creation)]
|
||||
fixed-by: mesh-controller PR 161 — no container is given a mesh name; it resolves through its machine's resolver (ADR 0148, landed 2026-09-30 once issue 110 did)
|
||||
located-in: [mesh-host internal/apply]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -73,13 +73,3 @@ copying: a container resolves through its machine's resolver at the moment it as
|
||||
record reports then has nowhere to occur. It is gated on
|
||||
[issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md), so
|
||||
until that lands the mesh still copies and still compares.
|
||||
|
||||
## Resolved (2026-09-30)
|
||||
|
||||
110 landed the same day ([its resolution](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/01-resolution.md)),
|
||||
and mesh-controller PR 161 then removed the copy: no container is given a mesh name or a mesh address,
|
||||
and a module's own declared entries are the only `host` lines it carries. Verified on the control-node
|
||||
after its containers were recreated once — the last time a name will do that: the forge's container
|
||||
carries no extra hosts and resolves another machine and a routed name through the machine's resolver,
|
||||
so the shape this record describes has nowhere to occur. Checked in the controller's tests: a
|
||||
container's declaration is byte-for-byte the same under a roster of one machine and a roster of three.
|
||||
|
||||
+3
-7
@@ -1,9 +1,8 @@
|
||||
---
|
||||
status: resolved
|
||||
status: open
|
||||
opened: 2026-09-24
|
||||
located-in:
|
||||
- mesh-catalog modules/dnsmasq (the runtime was never told; the resolver answered by interface)
|
||||
fixed-by: mesh-catalog PR 175 (the runtime is reloaded and keeps its containers over a restart) and PR 176 (the resolver answers by address, so a query from a bridge is admitted) — measured 2026-09-30, 01-resolution.md
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -68,6 +67,3 @@ and [135](../135-a-containers-mesh-names-are-not-compared/00-report.md)).
|
||||
network is the case with no DNS at all, and it is the case the mesh's own forge runs in. Two of four
|
||||
machines also bind the resolver to loopback only, so the runtime hands their containers a public
|
||||
resolver. Both halves are this issue.
|
||||
|
||||
*Later the same day: the second half was wrong, and the first had a different cause than the one above.
|
||||
[01-resolution.md](01-resolution.md) has what was actually found.*
|
||||
|
||||
-64
@@ -1,64 +0,0 @@
|
||||
# 110 — resolved: a container on any network reaches the resolver, and is answered
|
||||
|
||||
*2026-09-30. Measured on the three converged machines; the adopted one holds its resolver module until it
|
||||
is taken and is not covered.*
|
||||
|
||||
## What was actually wrong
|
||||
|
||||
Not what the report predicted. The report named the filter: a container on the runtime's default
|
||||
network asks from a bridge address, and the converged filter admitted queries by source address only.
|
||||
That was true when it was written and was fixed before this issue was ever tested — the filter admits
|
||||
by the link a packet arrives on ([ADR 0144](../../02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md)),
|
||||
and a container's bridge is admitted whole. Tested on every machine: the query arrives, the filter
|
||||
passes it.
|
||||
|
||||
Three other things were wrong, each hiding the next.
|
||||
|
||||
**The runtime had never been told.** The resolver module writes the runtime's `dns` key into the
|
||||
runtime's own configuration file. The runtime reads that key when it starts and not on a reload, and on
|
||||
two machines the runtime predated the file — so every container they started got a public resolver, and
|
||||
`novox.internal` came back as not existing. Nothing reported this: the file was present and current,
|
||||
the resolver ran, and a name not existing is a valid answer. Fixed in mesh-catalog PR 175: the module
|
||||
also sets `live-restore` and reloads the runtime when its file changes, so the one restart the `dns` key
|
||||
needs no longer stops every container. The restart is then the operator's, once per machine; done on
|
||||
both today, with every running container kept.
|
||||
|
||||
**The resolver dropped the query.** With the runtime corrected, a container's query reached the resolver
|
||||
— and got no answer, on every machine, including the one whose runtime had been right all along. The
|
||||
socket was bound to the private address; the filter admitted the packet; dnsmasq received it and
|
||||
discarded it without a line of log. Its configuration said `interface=mesh0`, and dnsmasq admits a
|
||||
query by the interface it arrives on when told an interface: a container's query is addressed to the
|
||||
private address but arrives on the runtime's bridge, and the bridge is not `mesh0`. Fixed in mesh-catalog
|
||||
PR 176: the resolver is told the address to answer on, not the interface that carries it, and a query to
|
||||
that address is admitted whatever bridge brings it. The bridges are the runtime's to name.
|
||||
|
||||
**The report's second half was wrong.** "Two of four machines bind the resolver to loopback only" was
|
||||
an inference from the containers' behaviour, and the behaviour had the cause above. The resolver bound
|
||||
the private address on all four; nothing had asked it there.
|
||||
|
||||
## What is verified
|
||||
|
||||
From a container on the runtime's default network, started by hand and given nothing, on each of the
|
||||
three converged machines: `novox.internal` answers with the hub's private address, through the machine's
|
||||
own resolver. That is the fourth check of
|
||||
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) — "on every
|
||||
network the runtime offers" — and its first step; the record's step 2 (the runtime told per machine, as
|
||||
a file) was already how the module works. Step 3 may now begin.
|
||||
|
||||
## What checks it
|
||||
|
||||
By hand, today. Nothing in the mesh asserts that a container can resolve a mesh name: the resolver's
|
||||
own tests cover what it answers, not who can ask. The check that would have caught all three faults is
|
||||
the one the report asked for and 0148 lists — a container on the default network resolving a mesh name
|
||||
— and it is not built. It belongs with the reachability check of
|
||||
[issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
|
||||
which is parked; until then this is a thing a person verifies after touching the resolver, the filter,
|
||||
or the runtime's configuration.
|
||||
|
||||
## What this cost to find
|
||||
|
||||
The three faults produced one symptom — a container that cannot resolve — and each fix revealed the
|
||||
next. The first was found by reading the runtime's own view of its configuration rather than the file;
|
||||
the second by capturing the query on the bridge and finding it arrive and go unanswered; the third only
|
||||
by admitting the first belief was wrong. A machine that had been believed to work all day had never
|
||||
worked either.
|
||||
@@ -1,9 +1,9 @@
|
||||
---
|
||||
status: resolved
|
||||
status: open
|
||||
opened: 2026-09-28
|
||||
located-in: [mesh-controller internal/catalogue/declaration.go (composeName took the consumer's own name as the internal domain)]
|
||||
fixed-by: mesh-controller PR 163 — the internal name composes under the node whose proxy serves the route (ADR 0151, 2026-09-30)
|
||||
amended-design: 03-DESIGN/01-to-be/08-connectivity.md
|
||||
located-in: [mesh-controller internal/catalogue]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 139 — An internal route name resolves to the consumer's node, not the one that serves it
|
||||
@@ -49,15 +49,3 @@ resolve whether or not anything answers.
|
||||
and if so, is `route` still one mesh-wide provision or a node-scoped seat with a mesh-wide fallback?
|
||||
- What certifies the name in either case? The certificate is obtained by whoever terminates TLS, and
|
||||
that is the question above in another form.
|
||||
|
||||
## Answered (2026-09-30)
|
||||
|
||||
[ADR 0151](../../02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md):
|
||||
the internal name is composed under the node that serves the route — the machine the request arrives
|
||||
at — because `<x>.<node>.internal` means *goes to that node* and nothing else. The public name stays
|
||||
the consumer node's, which is where the operator put it. The first question is answered that way; the
|
||||
second, a per-node route holder, is a decision about seats and is left where it is; the third is
|
||||
unchanged, since the proxy that terminates the name is given it and certifies it.
|
||||
|
||||
mesh-controller PR 163 carries it. On this mesh every route is served beside its module, so no name
|
||||
changed; the controller's tests hold the case where it would.
|
||||
|
||||
@@ -138,48 +138,3 @@ push. The control node is worth last.
|
||||
One thing this found on the way out: an archive cannot be undeclared, and the attempt stops the machine
|
||||
applying anything at all — [issue 162](../162-an-archive-cannot-be-undeclared/00-report.md). It is how
|
||||
undoing the first delivery froze the workstation, and it is not specific to the host.
|
||||
|
||||
## Every machine self-updates (2026-09-30, evening)
|
||||
|
||||
```
|
||||
shanks 76f4566bef3d/nox-mesh-host active
|
||||
g14 76f4566bef3d/nox-mesh-host active
|
||||
novox 76f4566bef3d/nox-mesh-host active
|
||||
ace 76f4566bef3d/nox-mesh-host active
|
||||
|
||||
mesh-controller status: (no host split)
|
||||
```
|
||||
|
||||
The last delivery was unattended on all four: the fixed host was built, pushed, each machine stood
|
||||
aside exactly once for the genuinely newer version, and the delivered launcher started it — no
|
||||
restart by hand. A following push that delivered nothing new was applied and reported by every
|
||||
machine and stood nobody aside, which is the check
|
||||
[issue 163](../163-a-delivered-host-stood-aside-on-every-push-and-reported-nothing/00-report.md) asks
|
||||
for.
|
||||
|
||||
**Two more faults on the way, both mine, both found by reading the machine rather than the success
|
||||
line.** A delivered host compared the newest delivered version against its link-time stamp rather
|
||||
than the version it was running, so it stood aside on every push and — because standing aside cancels
|
||||
the report — never reported again (163). And the adopted machine kept its found launcher as the
|
||||
adoption rule says, so the delivery there needed a `take` before the launcher moved.
|
||||
|
||||
**The crossover needs one restart of the unit per machine, once.** The launcher process that was
|
||||
running on each machine was the old script, executing from its own inode; a new file beside it
|
||||
changes nothing until the unit restarts. Every subsequent delivery is unattended.
|
||||
|
||||
**Timing, measured:** on a machine, hearing a declaration to reporting it applied is about three
|
||||
seconds. A push as the operator sees it takes 17–20 seconds, and the difference is the control plane
|
||||
composing the declaration before it sends. A `--wait` shorter than that reads as "did not report" for
|
||||
a machine that did; the three-minute default read as slowness for a machine that never would. Neither
|
||||
number is a defect being chased here, and both are worth knowing before reading a push's answer.
|
||||
|
||||
## What this leaves
|
||||
|
||||
- [Issue 162](../162-an-archive-cannot-be-undeclared/00-report.md): an archive cannot be undeclared, so
|
||||
the host module — and any module with an archive — cannot be unassigned, and trying stops the machine
|
||||
applying anything.
|
||||
- [Issue 107](../107-a-declaration-carries-no-order/00-report.md) is unblocked: a declaration field is
|
||||
now a build and a push rather than an expedition.
|
||||
- Three stale version directories on the workstation from the first attempts, moved aside under
|
||||
`/var/lib/mesh-host/versions-held-back/`, and a backup of the adopted machine's hand-placed binary
|
||||
beside its state. Both are safe to delete and are not the mesh's to delete.
|
||||
|
||||
@@ -1,10 +1,10 @@
|
||||
---
|
||||
status: resolved
|
||||
status: open
|
||||
opened: 2026-09-29
|
||||
located-in:
|
||||
- mesh-host internal/apply/apply.go (containerSpecReading hashes every `host` entry)
|
||||
- mesh-controller internal/catalogue/declaration.go (withMeshNames gives every container the mesh's names)
|
||||
fixed-by: mesh-controller PR 161 — the roster left every container's declaration and so its digest (ADR 0148 step 3, 2026-09-30)
|
||||
fixed-by:
|
||||
---
|
||||
|
||||
# 151 — A new name recreates every container in the mesh
|
||||
@@ -92,15 +92,3 @@ copying names until a container can reach the resolver from any of the runtime's
|
||||
([issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)),
|
||||
which it cannot on two of four machines today. Removing the copy first reintroduces 109 and 135
|
||||
silently, on a live mesh, which is how both were found. The order is in the record.
|
||||
|
||||
## Resolved (2026-09-30)
|
||||
|
||||
Step 3 landed the day 110 did. mesh-controller PR 161 stops writing the roster into any container, so
|
||||
a container's digest no longer carries a name that is not its own. The controller's tests hold the
|
||||
record's check — a container's declaration does not move when the mesh's roster does, and does move
|
||||
when the module's own declared entries do.
|
||||
|
||||
The first push after the change recreated every container once, because every digest lost its host
|
||||
entries at the same moment. That was the last such event: from here a name added or moved on one
|
||||
machine changes no container anywhere, and the record's second check — add a routed name, watch every
|
||||
other machine's apply report say nothing changed — is what the next module assignment will show.
|
||||
|
||||
@@ -1,9 +1,9 @@
|
||||
---
|
||||
status: resolved
|
||||
status: located
|
||||
opened: 2026-09-30
|
||||
located-in:
|
||||
- mesh-controller internal/catalogue/roster.go (the roster's entries for a routed name)
|
||||
fixed-by: mesh-controller PR 163 — a routed name is published as itself, once, with no suffixed alias (ADR 0151, 2026-09-30)
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -61,10 +61,3 @@ composition produces `keycloak.novox.be.internal`, which is not a name anything
|
||||
|
||||
The fix is a judgement about what a routed name's internal form is, and 139 is the record that asks it;
|
||||
this one is the evidence that the current answer publishes a third thing that is neither.
|
||||
|
||||
## Resolved (2026-09-30)
|
||||
|
||||
The judgement 139 asked for is [ADR 0151](../../02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md):
|
||||
a routed name has no mesh form. The roster now publishes it as itself, once, at the serving node's
|
||||
address; the `<domain>.internal` line is gone from every machine's hosts file, and a controller test
|
||||
refuses it coming back.
|
||||
|
||||
@@ -0,0 +1,128 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-30
|
||||
located-in:
|
||||
- mesh-catalog (no module shares a path over the network)
|
||||
- hq 02-DECISIONS (a file-share seat, per ADR 0126, is a module's own to define)
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 169 — A machine shares its files, and the mesh does not know
|
||||
|
||||
## What was observed
|
||||
|
||||
ace serves the operator's media library to the home network with two host services no module
|
||||
declares and HAL never managed either:
|
||||
|
||||
```
|
||||
/etc/exports: /storage/media 192.168.1.0/24(rw,sync,root_squash,…) nfs-server active, :2049
|
||||
/etc/samba/smb.conf: [media] path = /storage/media/ valid users = media smb active, :139/:445
|
||||
```
|
||||
|
||||
Two LAN clients were connected at survey (2026-09-30). The library itself is operator data
|
||||
(ADR 0051: ~40 TB on ZFS, the mesh owns nothing about it — [issue 153](../153-an-adopted-machines-data-cannot-be-placed-where-it-is/00-report.md)
|
||||
is about modules reaching it in place).
|
||||
|
||||
Under the mesh as it stands, this arrangement has no expression and one failure mode:
|
||||
|
||||
- **Nothing declares the listens.** At `converge ace` the filter is the sum of what modules listen
|
||||
on (ADR 0045); 2049 and 445 are nobody's, so the shares close — silently, for the two clients
|
||||
that mount them.
|
||||
- **Nothing owns the configuration.** `/etc/exports` and `smb.conf` are hand-written files on one
|
||||
machine; a second machine sharing a directory would be written by hand again.
|
||||
- **Nothing can consume it.** A module on another node that wanted the library (a player, an
|
||||
indexer, a backup) has no `requires` to state and no binding to read; it would mount by a
|
||||
hand-typed host and path.
|
||||
- The clients are LAN devices, so this also meets [issue 154](../154-a-machines-own-network-is-not-a-reach/00-report.md)
|
||||
(no reach for the machine's own network).
|
||||
|
||||
## The proposal (the operator's, 2026-09-30, settled after two rounds)
|
||||
|
||||
**Two module-defined seats, one per protocol, because NFS and SMB share an intent and not a
|
||||
contract.** A seat in the mesh's sense is a contract — what it accepts, emits and serves, and the
|
||||
tools its holder must answer (ADR 0126, 0132) — and lined up, the two share almost none of it:
|
||||
|
||||
| | `nfs-share` | `smb-share` |
|
||||
|---|---|---|
|
||||
| serves | export path(s); the client ranges allowed (`sec=sys` authorises by address) | share name(s), path |
|
||||
| pair credential | none | a user and password per consumer |
|
||||
| consumer's mount | `at:/path` | `//at/share` with credentials |
|
||||
| holder's tools | export / unexport a path for a range | add / remove a share, create a user |
|
||||
|
||||
One `file-share` seat would be the union with every field optional — a consumer could bind it and
|
||||
still not know how to mount what it got (the emptiness ADR 0129 warns against). "Export a path to
|
||||
the network" is a category, and the mesh needs no seat category: a consumer requires the one it
|
||||
can mount. If "give me the library, however" is ever needed, it is a provision an umbrella module
|
||||
serves, not a seat.
|
||||
|
||||
Both are node-scoped, one holder per node (ADR 0110), so ace holds both. `nfs` and `samba` are the
|
||||
first implementations; a second (Ganesha for `nfs-share`, ksmbd for `smb-share`) is what proves
|
||||
0126's promise that "replacing the implementation changes nothing for any caller".
|
||||
|
||||
The holder module:
|
||||
|
||||
- declares the exported paths as `accesses` (ADR 0051: it owns nothing about them — never creates,
|
||||
chowns or removes), and *which* paths as the assignment's settings (ADR 0046/0112);
|
||||
- writes the share configuration (`/etc/exports`, `smb.conf`) as mesh-managed files and drives the
|
||||
units, like `dnsmasq`/`sshd` do for theirs;
|
||||
- declares its endpoints (`nfs` 2049/tcp; `smb` 445/tcp, …) so the reach — internal, or the LAN
|
||||
once 154 has an answer — is the assignment's, and converge keeps them open;
|
||||
- **provides** the seat's provision, so a consumer on another node `requires nfs-share` (or
|
||||
`smb-share`) and reads `${bound:nfs-share:at}` and the path from its binding instead of a
|
||||
hand-typed mount.
|
||||
|
||||
## The design gap this exposes
|
||||
|
||||
**A seat definition has no home outside the module that first declared it.** Today a seat is
|
||||
declared inside a manifest (`showcase` declares `the-showcase`, `ca-trust` its own). If `nfs`
|
||||
declared `nfs-share`, Ganesha could hold it only by depending on nfs's manifest — the coupling
|
||||
0126 removed for callers, reintroduced for implementations. The protocol needs a neutral place in
|
||||
the catalogue beside the modules (a seat definition registered like a manifest), with a module
|
||||
saying which seats it implements. This is the first role with an obvious second implementation,
|
||||
which is what makes it the exemplar for that mechanism.
|
||||
|
||||
## The consumer's half: a module mounts it (2026-09-30, third and fourth round)
|
||||
|
||||
A binding tells a consumer *where* the share is; it does not put the files on its machine. Mounting
|
||||
is something done on a machine, and something done on a machine is a module's work — not the host's
|
||||
(the vocabulary stays closed; no `mount` resource kind).
|
||||
|
||||
**A consumer-side module, `network-share` — the module responsible for setting up the network
|
||||
shares a node uses** (the operator's framing). A node role, like `node-uplink` or
|
||||
`node-dns-resolver`: each machine has it at most once, which is a reason for it to hold a
|
||||
node-scoped seat, so two modules can never both be writing mount units on one machine. Assigned on
|
||||
the node that wants the files:
|
||||
|
||||
- `requires nfs-share` (or `smb-share`); several shares on one node are several local names of
|
||||
the requirement (ADR 0094);
|
||||
- its manifest is a `package` (nfs-utils), a `file` writing a systemd `.mount` unit filled from
|
||||
the binding — `What=${bound:nfs-share:at}:${bound:nfs-share:path}` — and a `service` enabling it
|
||||
after the overlay is up: the same shape as `resolv-conf` or `sshd`, files and a unit;
|
||||
- *where* it mounts is the assignment's setting (`/srv/media` on one machine, elsewhere on
|
||||
another); which machine mounts what is an operator decision made at assignment, exactly as which
|
||||
paths a machine shares is.
|
||||
|
||||
**The modules that use the files never learn about NFS.** A player, an indexer, a backup declares
|
||||
the mounted path as an `access` — an operator-chosen, pre-existing path the mesh never owns
|
||||
(ADR 0051), exactly as `/storage/media` is on ace. The same app manifest then runs on ace against
|
||||
the local library and on another node against the mounted one, with only its assignment differing.
|
||||
|
||||
**The one check to add, because it is the data-loss case.** An `access` is confirmed today by the
|
||||
path being present. For a mountpoint that is not enough: a writer whose container starts before the
|
||||
mount is up writes into the empty directory underneath it, and the files vanish when the mount
|
||||
lands. The access check must confirm the path is *a mountpoint* when the module says so (or the
|
||||
module's unit is ordered before the consumer's container — which crosses modules and is exactly
|
||||
what the mesh does not order). Which of the two is the decision's.
|
||||
|
||||
**Identity crosses the wire.** `sec=sys` NFS trusts the client's uid, so a consumer must run as the
|
||||
library's owner on the server (ace: `media`, 1001:2000) — hq 153's `${access:<id>:uid}`, read from
|
||||
the mounted tree, answers it on the consumer's side too.
|
||||
|
||||
## Open questions for the decision
|
||||
|
||||
- Whether an NFS export over the overlay is an `internal` reach of the same endpoint or a second
|
||||
export line — NFS authorises by client address, so the mesh range and the LAN range are two
|
||||
entries in one file.
|
||||
- How a consumer's binding expresses a *path* to mount (today bindings carry `at`, `port`, `as` and
|
||||
whatever the provider `serves`), and whether one share can serve several paths.
|
||||
Reference in New Issue
Block a user