Files
hq/04-ISSUES/010-the-first-declaration-destroys-the-substrate/00-report.md
T
jschoubben 36d9b38a0d An issue is open, diagnosing, located, resolved or wontfix — nothing else
The playbook, the README and the status skill knew five statuses; the cycle check
knew a sixth, 'fixed', and not 'wontfix'. Eleven issues sat in the sixth for weeks
with their fixes shipped, one step short of closed. They are resolved; the check
refuses the word from now on and accepts the one the playbook allows.
2026-09-21 17:38:17 +02:00

5.6 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
resolved 2026-08-29
mesh-host
mesh-control
mesh-host: the store records where each resource came from — carried or declared — and each origin removes only its own. Verified in the lab on the exact scenario that caused this: the substrate survived, and a later declaration still removed what it had itself declared.

010 — The first declaration a node receives destroys the substrate it raised

Symptom

A first node was raised from its carried bundle: container runtime, store, two context databases, their schemas, broker, control plane. Eleven resources, all running. It then enrolled against the control plane on its own machine, held its link open, and was sent a declaration naming two resources — a directory and a file.

Both were applied correctly. And every container on the machine was removed: the store, the broker, and the control plane that had sent the declaration. The link died mid-sentence with the link closed: Exception (501) Reason: "EOF", because the broker carrying it had just been torn down by the message it carried.

Afterwards mesh-host owned listed two resources. The mesh had deleted itself.

What is actually wrong

Nothing in the code is behaving incorrectly. apply removes what the store holds and the incoming declaration does not name, which is what reconciliation means — the declaration is the desired state, not a patch, and anything else would make it impossible to remove a resource by omission.

The fault is that the carried bundle and mesh declarations share one store. The host cannot tell "this machine raised this for itself before there was a mesh" from "the mesh told this machine to have this", so the second overwrites the first completely.

That is invisible until the two meet, which happens exactly once per mesh: on the first node, after enrolment, at the moment the control plane first speaks.

Why it matters more than a footgun

The first node is the only node where the substrate is not the mesh's doing. Every other node receives everything it runs from the control plane, so a complete declaration is complete by construction. The first node raised its own substrate from a file it carried, and the control plane has never been told about it — so the control plane cannot include it in a declaration even if it wanted to.

So the first node is left in a state no other node is in, and the ordinary path destroys it.

What is not the answer

  • Making the control plane send the substrate back. It does not know what the bundle contained and should not: the bundle exists precisely because there was no control plane yet.
  • Making apply stop removing orphans. Removal by omission is how a declaration says stop running this, and losing it costs the property that a node converges on what it was told rather than accumulating.
  • Special-casing the first node. ADR 0004 is explicit that its specialness lasts two commands, and this would extend it for ever.

The fix

The store records where each resource came from — carried or declared — and each origin removes only its own. A declaration removes what the mesh previously declared and never what the bundle raised; reconciling the bundle removes what the bundle previously raised and never what the mesh assigned.

State written before the field existed reads as carried, because everything a host had applied at that point came from its bundle — there was no other way to tell it anything. Guessing the other way would have the first upgrade remove the substrate, which is this fault arriving through the change that fixes it.

Verified on the scenario that caused it. A first node raised eleven resources, enrolled, and was sent the same two-resource declaration. Both applied; the store, the broker and the control plane were still running afterwards. A second declaration dropping one resource removed that resource and nothing else, so removal by omission still works — which is the property that had to survive the fix.

What remains open is what happens when the mesh eventually declares the substrate, which it must, or the substrate can never be upgraded. Two sources claiming one container is the ambiguity this issue is made of, narrowed rather than removed: it can no longer happen by accident, and nothing yet says what it means when it happens on purpose.

How it was found

In the lab, on a sealed machine, by doing the ordinary thing: raise a first node, enrol it, and tell it something. It was not a test of this — it was the first end-to-end run of the link, and this fell out of it.

The declaration was two lines and destroyed a working mesh in under a second, which is worth holding on to: this is not an edge case reached by trying, it is the first thing that happens.

Two more faults found while fixing it

Both of the same shape, and worth recording because the shape is the point.

A report published to an unbound routing key vanishes. The control plane bound enrol and not report, so nodes announced what they had applied into a void — the broker accepted each message, found no queue for it, and dropped it. The publisher was told nothing. Reports are now published mandatory, so anything unroutable comes back and is said out loud, and the binding covers every key a node may publish.

A publish failure was being swallowed. publishReport discarded its error, so a node that could not tell the mesh what it had done looked identical to one that had. That is the fault this repository keeps cataloguing, written by hand into the newest code in it.