Compare commits

..
Author SHA1 Message Date
jschoubben 846c1f85f2 Merge pull request 'Issue 107 is resolved: a declaration carries its order' (#217) from issue/107-resolved into main 2026-09-30 12:13:57 +00:00
jschoubben 9eef0bd525 Issue 107 is resolved: a declaration carries its order
Hosts first, then the controller — a build and a push each, now that the
mesh delivers the host. The host refuses a lower sequence than it kept
and drains a batch by sequence rather than arrival; the controller
numbers each send under the node's hold, inside the signed bytes.

Measured: two pushes, sequence 2 in the kept declaration, counters in
the store agree, no machine reads as behind. That last one is the
subtlety: the mesh compares the digest of what it would send against
what it did, and a number changes the bytes, so the read-only comparison
composes with the last number sent rather than a fresh one.
2026-09-30 14:13:50 +02:00
jschoubben 6e08cdf3d6 Merge pull request 'Every machine self-updates, verified, and 107's gate has opened' (#216) from issue/142-self-update-on-every-machine into main 2026-09-30 11:51:53 +00:00
jschoubben 02f291a129 Every machine self-updates, verified, and 107's gate has opened
All four machines run a host the mesh built, published and delivered, the
last delivery unattended: each stood aside once for a genuinely newer
version and the delivered launcher started it. A following push that
delivered nothing new was applied and reported by every machine and stood
nobody aside.

The crossover needs one restart of the unit per machine, once, because
the running launcher executes from its own inode. Measured timing: three
seconds on the machine, 17-20 as the operator sees it, the difference
being the control plane composing before it sends.

107 is unblocked: a declaration field is now a build and a push.
2026-09-30 13:51:46 +02:00
5 changed files with 115 additions and 130 deletions
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-09-23
located-in: [mesh-controller internal/link, mesh-host internal/link]
fixed-by:
fixed-by: mesh-host PR 59 (the host refuses an older sequence and drains by it), mesh-controller PR 160 (each send is numbered under the node's hold) — measured 2026-09-30, 02-resolution.md
amended-design:
---
@@ -66,3 +66,11 @@ Left `located`. The owner is unchanged, the shape of the fix is agreed, and the
[issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md) rather than anything
in this record. **This is a judgement about order, not a refusal** — it is cheap to overrule, and the
code is a day's work once a host can be delivered.
## The gate has opened (2026-09-30, evening)
The mesh delivers the host now — built by its own toolchain, published to its own registry, delivered
over the bus and started by the launcher, on all four machines
([issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/01-progress.md)). A declaration
field is a build and a push, not an expedition. The order this record asked for — hosts first, then
the controller — is now two commands and a status line that says when the first has finished.
@@ -0,0 +1,60 @@
# 107 — resolved: a declaration carries its order
*2026-09-30. Measured on the mesh.*
## What was done
**Hosts first, then the controller** — the order [issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md)
says a new declaration field needs, and now a build and a push rather than an expedition
([issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/01-progress.md)).
The host understands a `sequence` on a declaration and tolerates its absence: absent reads as "no
order claimed", not "first", so a controller that sends none is still understood and a host that kept
a declaration before it understood the field compares nothing. It refuses a declaration with a lower
sequence than the one it kept, whole, and says why; and the drain that picks one declaration from a
batch keeps the highest sequence rather than the last to arrive — which is the case the report
constructed, a backlog drained out of order.
The controller numbers each send: the next number for that node, taken under the node's hold, before
the body exists, so the number is inside what the mesh signs and a replayed older declaration cannot
borrow a newer one's.
## Measured
```
push shanks; push shanks
sequence in kept declaration: 2
node sequence
novox 2
shanks 2
ace (none — not sent since numbering)
g14 (none)
status: nobody "not running what the mesh would send them"
```
Both applies went through; neither was refused; the machine holding the earlier one accepted the later.
## The subtlety, which would have read every machine as behind for ever
The mesh decides a machine is behind by comparing the digest of what it **would** send against what it
**did** send. A number changes the bytes. So the read-only comparison composes with the number the
machine was *last* sent — not a fresh one — and is byte for byte what was sent when nothing else
changed. Without that, numbering would have made `status` name all four machines as out of date on
every reading, permanently.
## The open questions
- *A per-node `sequence` under the controller's node hold?* Yes, as described. **`supersedes` — the
previous digest — is not added.** A strictly-greater sequence gives the ordering; a chain of digests
would give continuity, which nothing here needs yet and which every re-composition would break.
- *Genesis signing its bundle as sequence zero?* Zero is "no order claimed", which is what the bundle
carries by carrying nothing. Same rule, no genesis branch.
- *A marker for a mode change?* Not needed for the incident it guards: a replayed converged declaration
reaching a node returned to adopted is already refused **by mode**, before this check runs.
## How it is checked
Host: an older sequence is refused, a newer or equal one is not, and no order claimed on either side
compares nothing; the drain keeps the highest sequence, and falls back to arrival when none is claimed.
Controller: a send carries its number inside the signed bytes, an unnumbered send is byte for byte what
it was before, and each node's counter is one higher per send and readable for the comparison.
@@ -138,3 +138,48 @@ push. The control node is worth last.
One thing this found on the way out: an archive cannot be undeclared, and the attempt stops the machine
applying anything at all — [issue 162](../162-an-archive-cannot-be-undeclared/00-report.md). It is how
undoing the first delivery froze the workstation, and it is not specific to the host.
## Every machine self-updates (2026-09-30, evening)
```
shanks 76f4566bef3d/nox-mesh-host active
g14 76f4566bef3d/nox-mesh-host active
novox 76f4566bef3d/nox-mesh-host active
ace 76f4566bef3d/nox-mesh-host active
mesh-controller status: (no host split)
```
The last delivery was unattended on all four: the fixed host was built, pushed, each machine stood
aside exactly once for the genuinely newer version, and the delivered launcher started it — no
restart by hand. A following push that delivered nothing new was applied and reported by every
machine and stood nobody aside, which is the check
[issue 163](../163-a-delivered-host-stood-aside-on-every-push-and-reported-nothing/00-report.md) asks
for.
**Two more faults on the way, both mine, both found by reading the machine rather than the success
line.** A delivered host compared the newest delivered version against its link-time stamp rather
than the version it was running, so it stood aside on every push and — because standing aside cancels
the report — never reported again (163). And the adopted machine kept its found launcher as the
adoption rule says, so the delivery there needed a `take` before the launcher moved.
**The crossover needs one restart of the unit per machine, once.** The launcher process that was
running on each machine was the old script, executing from its own inode; a new file beside it
changes nothing until the unit restarts. Every subsequent delivery is unattended.
**Timing, measured:** on a machine, hearing a declaration to reporting it applied is about three
seconds. A push as the operator sees it takes 17–20 seconds, and the difference is the control plane
composing the declaration before it sends. A `--wait` shorter than that reads as "did not report" for
a machine that did; the three-minute default read as slowness for a machine that never would. Neither
number is a defect being chased here, and both are worth knowing before reading a push's answer.
## What this leaves
- [Issue 162](../162-an-archive-cannot-be-undeclared/00-report.md): an archive cannot be undeclared, so
the host module — and any module with an archive — cannot be unassigned, and trying stops the machine
applying anything.
- [Issue 107](../107-a-declaration-carries-no-order/00-report.md) is unblocked: a declaration field is
now a build and a push rather than an expedition.
- Three stale version directories on the workstation from the first attempts, moved aside under
`/var/lib/mesh-host/versions-held-back/`, and a backup of the adopted machine's hand-placed binary
beside its state. Both are safe to delete and are not the mesh's to delete.
@@ -1,128 +0,0 @@
---
status: open
opened: 2026-09-30
located-in:
- mesh-catalog (no module shares a path over the network)
- hq 02-DECISIONS (a file-share seat, per ADR 0126, is a module's own to define)
fixed-by:
amended-design:
---
# 169 — A machine shares its files, and the mesh does not know
## What was observed
ace serves the operator's media library to the home network with two host services no module
declares and HAL never managed either:
```
/etc/exports: /storage/media 192.168.1.0/24(rw,sync,root_squash,…) nfs-server active, :2049
/etc/samba/smb.conf: [media] path = /storage/media/ valid users = media smb active, :139/:445
```
Two LAN clients were connected at survey (2026-09-30). The library itself is operator data
(ADR 0051: ~40 TB on ZFS, the mesh owns nothing about it — [issue 153](../153-an-adopted-machines-data-cannot-be-placed-where-it-is/00-report.md)
is about modules reaching it in place).
Under the mesh as it stands, this arrangement has no expression and one failure mode:
- **Nothing declares the listens.** At `converge ace` the filter is the sum of what modules listen
on (ADR 0045); 2049 and 445 are nobody's, so the shares close — silently, for the two clients
that mount them.
- **Nothing owns the configuration.** `/etc/exports` and `smb.conf` are hand-written files on one
machine; a second machine sharing a directory would be written by hand again.
- **Nothing can consume it.** A module on another node that wanted the library (a player, an
indexer, a backup) has no `requires` to state and no binding to read; it would mount by a
hand-typed host and path.
- The clients are LAN devices, so this also meets [issue 154](../154-a-machines-own-network-is-not-a-reach/00-report.md)
(no reach for the machine's own network).
## The proposal (the operator's, 2026-09-30, settled after two rounds)
**Two module-defined seats, one per protocol, because NFS and SMB share an intent and not a
contract.** A seat in the mesh's sense is a contract — what it accepts, emits and serves, and the
tools its holder must answer (ADR 0126, 0132) — and lined up, the two share almost none of it:
| | `nfs-share` | `smb-share` |
|---|---|---|
| serves | export path(s); the client ranges allowed (`sec=sys` authorises by address) | share name(s), path |
| pair credential | none | a user and password per consumer |
| consumer's mount | `at:/path` | `//at/share` with credentials |
| holder's tools | export / unexport a path for a range | add / remove a share, create a user |
One `file-share` seat would be the union with every field optional — a consumer could bind it and
still not know how to mount what it got (the emptiness ADR 0129 warns against). "Export a path to
the network" is a category, and the mesh needs no seat category: a consumer requires the one it
can mount. If "give me the library, however" is ever needed, it is a provision an umbrella module
serves, not a seat.
Both are node-scoped, one holder per node (ADR 0110), so ace holds both. `nfs` and `samba` are the
first implementations; a second (Ganesha for `nfs-share`, ksmbd for `smb-share`) is what proves
0126's promise that "replacing the implementation changes nothing for any caller".
The holder module:
- declares the exported paths as `accesses` (ADR 0051: it owns nothing about them — never creates,
chowns or removes), and *which* paths as the assignment's settings (ADR 0046/0112);
- writes the share configuration (`/etc/exports`, `smb.conf`) as mesh-managed files and drives the
units, like `dnsmasq`/`sshd` do for theirs;
- declares its endpoints (`nfs` 2049/tcp; `smb` 445/tcp, …) so the reach — internal, or the LAN
once 154 has an answer — is the assignment's, and converge keeps them open;
- **provides** the seat's provision, so a consumer on another node `requires nfs-share` (or
`smb-share`) and reads `${bound:nfs-share:at}` and the path from its binding instead of a
hand-typed mount.
## The design gap this exposes
**A seat definition has no home outside the module that first declared it.** Today a seat is
declared inside a manifest (`showcase` declares `the-showcase`, `ca-trust` its own). If `nfs`
declared `nfs-share`, Ganesha could hold it only by depending on nfs's manifest — the coupling
0126 removed for callers, reintroduced for implementations. The protocol needs a neutral place in
the catalogue beside the modules (a seat definition registered like a manifest), with a module
saying which seats it implements. This is the first role with an obvious second implementation,
which is what makes it the exemplar for that mechanism.
## The consumer's half: a module mounts it (2026-09-30, third and fourth round)
A binding tells a consumer *where* the share is; it does not put the files on its machine. Mounting
is something done on a machine, and something done on a machine is a module's work — not the host's
(the vocabulary stays closed; no `mount` resource kind).
**A consumer-side module, `network-share` — the module responsible for setting up the network
shares a node uses** (the operator's framing). A node role, like `node-uplink` or
`node-dns-resolver`: each machine has it at most once, which is a reason for it to hold a
node-scoped seat, so two modules can never both be writing mount units on one machine. Assigned on
the node that wants the files:
- `requires nfs-share` (or `smb-share`); several shares on one node are several local names of
the requirement (ADR 0094);
- its manifest is a `package` (nfs-utils), a `file` writing a systemd `.mount` unit filled from
the binding — `What=${bound:nfs-share:at}:${bound:nfs-share:path}` — and a `service` enabling it
after the overlay is up: the same shape as `resolv-conf` or `sshd`, files and a unit;
- *where* it mounts is the assignment's setting (`/srv/media` on one machine, elsewhere on
another); which machine mounts what is an operator decision made at assignment, exactly as which
paths a machine shares is.
**The modules that use the files never learn about NFS.** A player, an indexer, a backup declares
the mounted path as an `access` — an operator-chosen, pre-existing path the mesh never owns
(ADR 0051), exactly as `/storage/media` is on ace. The same app manifest then runs on ace against
the local library and on another node against the mounted one, with only its assignment differing.
**The one check to add, because it is the data-loss case.** An `access` is confirmed today by the
path being present. For a mountpoint that is not enough: a writer whose container starts before the
mount is up writes into the empty directory underneath it, and the files vanish when the mount
lands. The access check must confirm the path is *a mountpoint* when the module says so (or the
module's unit is ordered before the consumer's container — which crosses modules and is exactly
what the mesh does not order). Which of the two is the decision's.
**Identity crosses the wire.** `sec=sys` NFS trusts the client's uid, so a consumer must run as the
library's owner on the server (ace: `media`, 1001:2000) — hq 153's `${access:<id>:uid}`, read from
the mounted tree, answers it on the consumer's side too.
## Open questions for the decision
- Whether an NFS export over the overlay is an `internal` reach of the same endpoint or a second
export line — NFS authorises by client address, so the mesh range and the LAN range are two
entries in one file.
- How a consumer's binding expresses a *path* to mount (today bindings carry `at`, `port`, `as` and
whatever the provider `serves`), and whether one share can serve several paths.