Compare commits
8
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
105ae9a56a | ||
|
|
ee801a6441 | ||
|
|
90b89aa1c9 | ||
|
|
e33191161d | ||
|
|
a0a930b1cd | ||
|
|
8d9c9ab6b5 | ||
|
|
9b14430d3f | ||
|
|
a4384f13d3 |
+56
@@ -0,0 +1,56 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-30
|
||||
located-in: [mesh-host cmd/mesh-host/main.go (the successor check after an apply)]
|
||||
fixed-by: mesh-host PR 58 — the check asks with the running version, read from the binary's path, not the link-time stamp
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 163 — A delivered host stood aside on every push, and reported nothing
|
||||
|
||||
## What was observed
|
||||
|
||||
*2026-09-30, rolling the mesh-built host onto the last two machines.*
|
||||
|
||||
Every push to a machine running a delivered host produced, in order:
|
||||
|
||||
```
|
||||
host 093231796eb0 is delivered; standing aside so the launcher runs it
|
||||
applied 333 resource(s)
|
||||
applied, and could not tell the mesh: reporting: context canceled
|
||||
nox-mesh-host-launch: the host exited cleanly; starting it again
|
||||
nox-mesh-host-launch: running /usr/lib/nox-mesh-host/versions/093231796eb0/nox-mesh-host
|
||||
```
|
||||
|
||||
— for the version it was **already running**. It restarted itself on every push, for ever, and the mesh
|
||||
never received a single report from it: `node show` kept the version from before the crossover, and
|
||||
the operator's push waited its full three minutes for an answer that was never coming.
|
||||
|
||||
Read as healthy throughout: unit active, bus link up, "hearing what this node should be".
|
||||
|
||||
## Why
|
||||
|
||||
After an apply the host asks whether a newer host has been delivered than the one running, and the
|
||||
question was asked with the **link-time version stamp**. Since
|
||||
[issue 161](../161-a-delivered-host-carries-none-of-its-link-time-facts/01-resolution.md) a delivered
|
||||
host's version comes from where it sits and its stamp is `development build` — so the comparison never
|
||||
matched the newest delivered version, and "a newer host is waiting" was always true.
|
||||
|
||||
Standing aside cancels the context the report is published with, so the report was lost on every one
|
||||
of those applies. Two faults from one wrong argument.
|
||||
|
||||
The change that moved the version to the path was applied to the report and to the known-good record,
|
||||
and not here. Half a change, and the half left behind was the one that decides whether to exit.
|
||||
|
||||
## Why the three-minute wait made it invisible
|
||||
|
||||
The push's `--wait` timing out read as *slow*. It was not slow: **the report was never going to arrive.**
|
||||
The operator put it exactly: *if you don't get a response in five seconds, something is wrong.* A wait
|
||||
long enough to absorb a machine's whole apply is a wait long enough to hide that the machine never
|
||||
answered.
|
||||
|
||||
## How it is checked
|
||||
|
||||
A machine running a delivered host is pushed a declaration that delivers nothing new; it applies,
|
||||
reports, and does not stand aside. A machine running a delivered host is pushed a genuinely newer
|
||||
version; it stands aside once, and the next push it does not.
|
||||
@@ -0,0 +1,128 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-30
|
||||
located-in:
|
||||
- mesh-catalog (no module shares a path over the network)
|
||||
- hq 02-DECISIONS (a file-share seat, per ADR 0126, is a module's own to define)
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 169 — A machine shares its files, and the mesh does not know
|
||||
|
||||
## What was observed
|
||||
|
||||
ace serves the operator's media library to the home network with two host services no module
|
||||
declares and HAL never managed either:
|
||||
|
||||
```
|
||||
/etc/exports: /storage/media 192.168.1.0/24(rw,sync,root_squash,…) nfs-server active, :2049
|
||||
/etc/samba/smb.conf: [media] path = /storage/media/ valid users = media smb active, :139/:445
|
||||
```
|
||||
|
||||
Two LAN clients were connected at survey (2026-09-30). The library itself is operator data
|
||||
(ADR 0051: ~40 TB on ZFS, the mesh owns nothing about it — [issue 153](../153-an-adopted-machines-data-cannot-be-placed-where-it-is/00-report.md)
|
||||
is about modules reaching it in place).
|
||||
|
||||
Under the mesh as it stands, this arrangement has no expression and one failure mode:
|
||||
|
||||
- **Nothing declares the listens.** At `converge ace` the filter is the sum of what modules listen
|
||||
on (ADR 0045); 2049 and 445 are nobody's, so the shares close — silently, for the two clients
|
||||
that mount them.
|
||||
- **Nothing owns the configuration.** `/etc/exports` and `smb.conf` are hand-written files on one
|
||||
machine; a second machine sharing a directory would be written by hand again.
|
||||
- **Nothing can consume it.** A module on another node that wanted the library (a player, an
|
||||
indexer, a backup) has no `requires` to state and no binding to read; it would mount by a
|
||||
hand-typed host and path.
|
||||
- The clients are LAN devices, so this also meets [issue 154](../154-a-machines-own-network-is-not-a-reach/00-report.md)
|
||||
(no reach for the machine's own network).
|
||||
|
||||
## The proposal (the operator's, 2026-09-30, settled after two rounds)
|
||||
|
||||
**Two module-defined seats, one per protocol, because NFS and SMB share an intent and not a
|
||||
contract.** A seat in the mesh's sense is a contract — what it accepts, emits and serves, and the
|
||||
tools its holder must answer (ADR 0126, 0132) — and lined up, the two share almost none of it:
|
||||
|
||||
| | `nfs-share` | `smb-share` |
|
||||
|---|---|---|
|
||||
| serves | export path(s); the client ranges allowed (`sec=sys` authorises by address) | share name(s), path |
|
||||
| pair credential | none | a user and password per consumer |
|
||||
| consumer's mount | `at:/path` | `//at/share` with credentials |
|
||||
| holder's tools | export / unexport a path for a range | add / remove a share, create a user |
|
||||
|
||||
One `file-share` seat would be the union with every field optional — a consumer could bind it and
|
||||
still not know how to mount what it got (the emptiness ADR 0129 warns against). "Export a path to
|
||||
the network" is a category, and the mesh needs no seat category: a consumer requires the one it
|
||||
can mount. If "give me the library, however" is ever needed, it is a provision an umbrella module
|
||||
serves, not a seat.
|
||||
|
||||
Both are node-scoped, one holder per node (ADR 0110), so ace holds both. `nfs` and `samba` are the
|
||||
first implementations; a second (Ganesha for `nfs-share`, ksmbd for `smb-share`) is what proves
|
||||
0126's promise that "replacing the implementation changes nothing for any caller".
|
||||
|
||||
The holder module:
|
||||
|
||||
- declares the exported paths as `accesses` (ADR 0051: it owns nothing about them — never creates,
|
||||
chowns or removes), and *which* paths as the assignment's settings (ADR 0046/0112);
|
||||
- writes the share configuration (`/etc/exports`, `smb.conf`) as mesh-managed files and drives the
|
||||
units, like `dnsmasq`/`sshd` do for theirs;
|
||||
- declares its endpoints (`nfs` 2049/tcp; `smb` 445/tcp, …) so the reach — internal, or the LAN
|
||||
once 154 has an answer — is the assignment's, and converge keeps them open;
|
||||
- **provides** the seat's provision, so a consumer on another node `requires nfs-share` (or
|
||||
`smb-share`) and reads `${bound:nfs-share:at}` and the path from its binding instead of a
|
||||
hand-typed mount.
|
||||
|
||||
## The design gap this exposes
|
||||
|
||||
**A seat definition has no home outside the module that first declared it.** Today a seat is
|
||||
declared inside a manifest (`showcase` declares `the-showcase`, `ca-trust` its own). If `nfs`
|
||||
declared `nfs-share`, Ganesha could hold it only by depending on nfs's manifest — the coupling
|
||||
0126 removed for callers, reintroduced for implementations. The protocol needs a neutral place in
|
||||
the catalogue beside the modules (a seat definition registered like a manifest), with a module
|
||||
saying which seats it implements. This is the first role with an obvious second implementation,
|
||||
which is what makes it the exemplar for that mechanism.
|
||||
|
||||
## The consumer's half: a module mounts it (2026-09-30, third and fourth round)
|
||||
|
||||
A binding tells a consumer *where* the share is; it does not put the files on its machine. Mounting
|
||||
is something done on a machine, and something done on a machine is a module's work — not the host's
|
||||
(the vocabulary stays closed; no `mount` resource kind).
|
||||
|
||||
**A consumer-side module, `network-share` — the module responsible for setting up the network
|
||||
shares a node uses** (the operator's framing). A node role, like `node-uplink` or
|
||||
`node-dns-resolver`: each machine has it at most once, which is a reason for it to hold a
|
||||
node-scoped seat, so two modules can never both be writing mount units on one machine. Assigned on
|
||||
the node that wants the files:
|
||||
|
||||
- `requires nfs-share` (or `smb-share`); several shares on one node are several local names of
|
||||
the requirement (ADR 0094);
|
||||
- its manifest is a `package` (nfs-utils), a `file` writing a systemd `.mount` unit filled from
|
||||
the binding — `What=${bound:nfs-share:at}:${bound:nfs-share:path}` — and a `service` enabling it
|
||||
after the overlay is up: the same shape as `resolv-conf` or `sshd`, files and a unit;
|
||||
- *where* it mounts is the assignment's setting (`/srv/media` on one machine, elsewhere on
|
||||
another); which machine mounts what is an operator decision made at assignment, exactly as which
|
||||
paths a machine shares is.
|
||||
|
||||
**The modules that use the files never learn about NFS.** A player, an indexer, a backup declares
|
||||
the mounted path as an `access` — an operator-chosen, pre-existing path the mesh never owns
|
||||
(ADR 0051), exactly as `/storage/media` is on ace. The same app manifest then runs on ace against
|
||||
the local library and on another node against the mounted one, with only its assignment differing.
|
||||
|
||||
**The one check to add, because it is the data-loss case.** An `access` is confirmed today by the
|
||||
path being present. For a mountpoint that is not enough: a writer whose container starts before the
|
||||
mount is up writes into the empty directory underneath it, and the files vanish when the mount
|
||||
lands. The access check must confirm the path is *a mountpoint* when the module says so (or the
|
||||
module's unit is ordered before the consumer's container — which crosses modules and is exactly
|
||||
what the mesh does not order). Which of the two is the decision's.
|
||||
|
||||
**Identity crosses the wire.** `sec=sys` NFS trusts the client's uid, so a consumer must run as the
|
||||
library's owner on the server (ace: `media`, 1001:2000) — hq 153's `${access:<id>:uid}`, read from
|
||||
the mounted tree, answers it on the consumer's side too.
|
||||
|
||||
## Open questions for the decision
|
||||
|
||||
- Whether an NFS export over the overlay is an `internal` reach of the same endpoint or a second
|
||||
export line — NFS authorises by client address, so the mesh range and the LAN range are two
|
||||
entries in one file.
|
||||
- How a consumer's binding expresses a *path* to mount (today bindings carry `at`, `port`, `as` and
|
||||
whatever the provider `serves`), and whether one share can serve several paths.
|
||||
Reference in New Issue
Block a user