Compare commits

..
Author SHA1 Message Date
jschoubben 743051efe7 Issue 163 (was 161): another record took 161 on main first 2026-09-30 13:28:22 +02:00
jschoubben e8470057aa Merge remote-tracking branch 'origin/main' into issue/161-an-assignment-does-not-record-its-provider 2026-09-30 13:28:22 +02:00
jschoubben 39340fcd76 Issue 161: an assignment does not record which provider answers it
ADR 0110 decided each assignment records where its requirements are answered
from; the control plane keeps only a per-machine pin (none recorded) and
resolves every requirement implicitly. Harmless with one provider; a second
one silently moves consumers' data.
2026-09-30 11:54:31 +02:00
7 changed files with 65 additions and 234 deletions
@@ -1,8 +1,8 @@
---
status: resolved
status: located
opened: 2026-09-23
located-in: [mesh-controller internal/link, mesh-host internal/link]
fixed-by: mesh-host PR 59 (the host refuses an older sequence and drains by it), mesh-controller PR 160 (each send is numbered under the node's hold) — measured 2026-09-30, 02-resolution.md
fixed-by:
amended-design:
---
@@ -66,11 +66,3 @@ Left `located`. The owner is unchanged, the shape of the fix is agreed, and the
[issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md) rather than anything
in this record. **This is a judgement about order, not a refusal** — it is cheap to overrule, and the
code is a day's work once a host can be delivered.
## The gate has opened (2026-09-30, evening)
The mesh delivers the host now — built by its own toolchain, published to its own registry, delivered
over the bus and started by the launcher, on all four machines
([issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/01-progress.md)). A declaration
field is a build and a push, not an expedition. The order this record asked for — hosts first, then
the controller — is now two commands and a status line that says when the first has finished.
@@ -1,60 +0,0 @@
# 107 — resolved: a declaration carries its order
*2026-09-30. Measured on the mesh.*
## What was done
**Hosts first, then the controller** — the order [issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md)
says a new declaration field needs, and now a build and a push rather than an expedition
([issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/01-progress.md)).
The host understands a `sequence` on a declaration and tolerates its absence: absent reads as "no
order claimed", not "first", so a controller that sends none is still understood and a host that kept
a declaration before it understood the field compares nothing. It refuses a declaration with a lower
sequence than the one it kept, whole, and says why; and the drain that picks one declaration from a
batch keeps the highest sequence rather than the last to arrive — which is the case the report
constructed, a backlog drained out of order.
The controller numbers each send: the next number for that node, taken under the node's hold, before
the body exists, so the number is inside what the mesh signs and a replayed older declaration cannot
borrow a newer one's.
## Measured
```
push shanks; push shanks
sequence in kept declaration: 2
node sequence
novox 2
shanks 2
ace (none — not sent since numbering)
g14 (none)
status: nobody "not running what the mesh would send them"
```
Both applies went through; neither was refused; the machine holding the earlier one accepted the later.
## The subtlety, which would have read every machine as behind for ever
The mesh decides a machine is behind by comparing the digest of what it **would** send against what it
**did** send. A number changes the bytes. So the read-only comparison composes with the number the
machine was *last* sent — not a fresh one — and is byte for byte what was sent when nothing else
changed. Without that, numbering would have made `status` name all four machines as out of date on
every reading, permanently.
## The open questions
- *A per-node `sequence` under the controller's node hold?* Yes, as described. **`supersedes` — the
previous digest — is not added.** A strictly-greater sequence gives the ordering; a chain of digests
would give continuity, which nothing here needs yet and which every re-composition would break.
- *Genesis signing its bundle as sequence zero?* Zero is "no order claimed", which is what the bundle
carries by carrying nothing. Same rule, no genesis branch.
- *A marker for a mode change?* Not needed for the incident it guards: a replayed converged declaration
reaching a node returned to adopted is already refused **by mode**, before this check runs.
## How it is checked
Host: an older sequence is refused, a newer or equal one is not, and no order claimed on either side
compares nothing; the drain keeps the highest sequence, and falls back to arrival when none is claimed.
Controller: a send carries its number inside the signed bytes, an unnumbered send is byte for byte what
it was before, and each node's counter is one higher per send and readable for the comparison.
@@ -138,48 +138,3 @@ push. The control node is worth last.
One thing this found on the way out: an archive cannot be undeclared, and the attempt stops the machine
applying anything at all — [issue 162](../162-an-archive-cannot-be-undeclared/00-report.md). It is how
undoing the first delivery froze the workstation, and it is not specific to the host.
## Every machine self-updates (2026-09-30, evening)
```
shanks 76f4566bef3d/nox-mesh-host active
g14 76f4566bef3d/nox-mesh-host active
novox 76f4566bef3d/nox-mesh-host active
ace 76f4566bef3d/nox-mesh-host active
mesh-controller status: (no host split)
```
The last delivery was unattended on all four: the fixed host was built, pushed, each machine stood
aside exactly once for the genuinely newer version, and the delivered launcher started it — no
restart by hand. A following push that delivered nothing new was applied and reported by every
machine and stood nobody aside, which is the check
[issue 163](../163-a-delivered-host-stood-aside-on-every-push-and-reported-nothing/00-report.md) asks
for.
**Two more faults on the way, both mine, both found by reading the machine rather than the success
line.** A delivered host compared the newest delivered version against its link-time stamp rather
than the version it was running, so it stood aside on every push and — because standing aside cancels
the report — never reported again (163). And the adopted machine kept its found launcher as the
adoption rule says, so the delivery there needed a `take` before the launcher moved.
**The crossover needs one restart of the unit per machine, once.** The launcher process that was
running on each machine was the old script, executing from its own inode; a new file beside it
changes nothing until the unit restarts. Every subsequent delivery is unattended.
**Timing, measured:** on a machine, hearing a declaration to reporting it applied is about three
seconds. A push as the operator sees it takes 17–20 seconds, and the difference is the control plane
composing the declaration before it sends. A `--wait` shorter than that reads as "did not report" for
a machine that did; the three-minute default read as slowness for a machine that never would. Neither
number is a defect being chased here, and both are worth knowing before reading a push's answer.
## What this leaves
- [Issue 162](../162-an-archive-cannot-be-undeclared/00-report.md): an archive cannot be undeclared, so
the host module — and any module with an archive — cannot be unassigned, and trying stops the machine
applying anything.
- [Issue 107](../107-a-declaration-carries-no-order/00-report.md) is unblocked: a declaration field is
now a build and a push rather than an expedition.
- Three stale version directories on the workstation from the first attempts, moved aside under
`/var/lib/mesh-host/versions-held-back/`, and a backup of the adopted machine's hand-placed binary
beside its state. Both are safe to delete and are not the mesh's to delete.
@@ -1,56 +0,0 @@
---
status: resolved
opened: 2026-09-30
located-in: [mesh-host cmd/mesh-host/main.go (the successor check after an apply)]
fixed-by: mesh-host PR 58 — the check asks with the running version, read from the binary's path, not the link-time stamp
amended-design:
---
# 163 — A delivered host stood aside on every push, and reported nothing
## What was observed
*2026-09-30, rolling the mesh-built host onto the last two machines.*
Every push to a machine running a delivered host produced, in order:
```
host 093231796eb0 is delivered; standing aside so the launcher runs it
applied 333 resource(s)
applied, and could not tell the mesh: reporting: context canceled
nox-mesh-host-launch: the host exited cleanly; starting it again
nox-mesh-host-launch: running /usr/lib/nox-mesh-host/versions/093231796eb0/nox-mesh-host
```
— for the version it was **already running**. It restarted itself on every push, for ever, and the mesh
never received a single report from it: `node show` kept the version from before the crossover, and
the operator's push waited its full three minutes for an answer that was never coming.
Read as healthy throughout: unit active, bus link up, "hearing what this node should be".
## Why
After an apply the host asks whether a newer host has been delivered than the one running, and the
question was asked with the **link-time version stamp**. Since
[issue 161](../161-a-delivered-host-carries-none-of-its-link-time-facts/01-resolution.md) a delivered
host's version comes from where it sits and its stamp is `development build` — so the comparison never
matched the newest delivered version, and "a newer host is waiting" was always true.
Standing aside cancels the context the report is published with, so the report was lost on every one
of those applies. Two faults from one wrong argument.
The change that moved the version to the path was applied to the report and to the known-good record,
and not here. Half a change, and the half left behind was the one that decides whether to exit.
## Why the three-minute wait made it invisible
The push's `--wait` timing out read as *slow*. It was not slow: **the report was never going to arrive.**
The operator put it exactly: *if you don't get a response in five seconds, something is wrong.* A wait
long enough to absorb a machine's whole apply is a wait long enough to hide that the machine never
answered.
## How it is checked
A machine running a delivered host is pushed a declaration that delivers nothing new; it applies,
reports, and does not stand aside. A machine running a delivered host is pushed a genuinely newer
version; it stands aside once, and the next push it does not.
@@ -0,0 +1,63 @@
---
status: open
opened: 2026-09-30
located-in:
- mesh-controller cmd/mesh-controller/modules.go (assign takes no provider; pin is a separate, per-machine command)
- mesh-controller internal/inventory (provision_pin keyed by (node, name))
fixed-by:
amended-design:
---
# 163 — An assignment does not record which provider answers it
## What was observed
Planning ace's modules that need a database (baserow, letta, n8n, and the apps using ace's
predecessor postgres). The operator's model — and ADR 0110's — is that **an assignment states where
each of its requirements is answered from**: gitea's assignment on novox says its `postgres-database`
comes from novox; an app assigned to ace says whether its database comes from ace or from novox.
The mesh holds no such statement for any assignment. Read on novox (2026-09-30):
```
select … from provision_pin; -- 0 rows
```
Every requirement in the mesh resolves implicitly, each time, by ADR 0084's order (a pin, then the
provider on the consumer's own node, then the only provider).
## What was decided, and what exists
[ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md):
> Where several remain and none is local, **a person chooses when the module is assigned**.
> Assignment lists the candidates, with the holder of a seat that delivers the provision suggested
> first, and records the answer on the assignment as its pin. Without an answer the module is not
> assigned.
What the control plane implements:
| decided | implemented |
|---|---|
| the answer is recorded **on the assignment** | `provision_pin` is keyed `(node, name)` — one answer per machine per provision, shared by every module on it |
| chosen **at assignment** | `assign <node> <module>` takes no provider; `pin <node> <provision> <from-node>` is a separate command |
| an assignment may be answered from its own machine (gitea ← novox) | `pin` refuses a machine pinning to itself ("does not need saying") |
| every assignment has an answer | none recorded; resolution guesses the same answer every time |
## Consequence
Nothing is wrong *today* — with one postgres provider, every guess is the intended answer. But the
answer is not a fact anyone stated, so:
- **it changes silently** the day a second provider appears (e.g. a postgres assigned on ace): every
unpinned consumer re-resolves — a consumer on ace moves from novox's database to an empty one on ace
at the next push, which is data a module stops seeing without anything saying so;
- two modules on one machine cannot take one provision from different providers;
- a person reading an assignment cannot see where its data lives.
## What would be right
ADR 0110 as written: `assign` records, per requirement, the node that answers it (its own node
included), offering the candidates and refusing an assignment without an answer where several exist;
the per-machine `provision_pin` becomes a per-assignment record, with existing assignments backfilled
from what they resolve to now so nothing moves.
@@ -1,63 +0,0 @@
---
status: open
opened: 2026-09-30
located-in:
- mesh-controller internal/catalogue/resolve.go (holdings are derived from every resolved assignment's manifest `claims`)
- mesh-controller cmd/mesh-controller/seats.go (the deliberate act exists — HoldSeat, "recording … as its standing holder" — beside it)
fixed-by:
amended-design:
---
# 170 — Assigning a module claims every seat it could hold
## What was observed
ace's migration needs a postgres of its own: the operator's decision is that a `postgres` module
assigned on ace provides `postgres-database` to ace's modules and has **nothing to do with the
`mesh-store` seat**, which novox's assignment holds by a deliberate act already taken ("make
novox's postgres the mesh-store").
`assign ace postgres` (2026-09-30):
```
ace is assigned postgres
AND 1 other machine(s) cannot be worked out as things stand, so nothing will be sent to them:
novox
- postgres on novox claims "mesh-store", which postgres on ace already holds — one per mesh
mesh-controller: these assignments cannot be applied:
- postgres on ace claims "mesh-store", which postgres on novox already holds — one per mesh
```
The second assignment did not merely fail: it made **the control plane's own store's
assignment unresolvable** until unassigned. Nothing was pushed; the state is restored.
## Why
`resolve.go` derives what a node holds from the manifest's `claims` of every module resolved on
it, so a claim in a definition is a claim by every assignment of that module. The deliberate
act ADR 0110 describes exists beside it — `seat …` records "X on Y as its standing holder"
(`HoldSeat`) — but resolution does not consult that record; it consults the manifests.
## What was decided
[ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md):
> **A definition says which seats a module *can* hold. An assignment says which it *does*
> hold.** The store module can hold `mesh-store`, and it may be assigned to every node. Exactly
> one of those assignments holds the seat, because that assignment said so.
The manifest's `claims` is being read as *does hold*.
## What would be right
Resolution takes the holder of a seat from the recorded holding (the seat's standing holder),
not from the manifests: a module whose definition can hold a seat is assignable anywhere, and only
the assignment recorded as holder claims it — with the refusal reserved for a second *recorded*
holder at the seat's scope. Assigning postgres to ace is then exactly what the operator said it
is: a database provider on ace, and no more.
## Until then
`postgres` cannot be assigned on any second node; ace's database windows (baserow, letta, n8n,
car-hunter, txt-game) wait on this.