4.1 KiB
status, opened, located-in, fixed-by, amended-design
| status | opened | located-in | fixed-by | amended-design | ||||
|---|---|---|---|---|---|---|---|---|
| resolved | 2026-10-01 |
|
mesh-controller PR 199 (each machine resolved with its own pins; a dropped machine said in both passes), rolled 2026-10-01 evening; PR 200 keeps the per-machine view quiet |
188 — A refusal inside "who is on the network" drops a machine silently, and every symptom points elsewhere
What was observed
At 15:28Z on 2026-10-01 the controller rolled to a build that refuses a machine with two modules
answering one provision and no pin naming which (mesh-controller 195). The control node had two
issuers of acme-ca. From that moment every plan of the control node failed with step-ca has a
content that says ${machine:at}, and this machine says mesh-range or name; seats listed every
seat as unheld; the build machine refused the builder's and the proxy's builds with no clone base
for that seat — nothing holds it; and the roll-out of the next controller was refused with the
${machine:at} words. Not one of those names the cause. It was found by running the previous image
as a one-shot beside the current one and reading the difference, forty minutes later.
Why this is here
onTheNetwork decides which machines have an address by resolving each one, unchecked, and
skipping any whose resolution errs. A machine skipped there has no at, so its own plan fails on
the first placeholder that needs one, in another module's words; everything held on it reads as
unheld; everything built from it cannot be built. The design lets one refusal become four unrelated
symptoms and no sentence about the refusal itself. It is the same shape as
issue 187: a fault that is swallowed
where it happens and discovered where it hurts.
What a fix needs
- A machine whose resolution refuses is said, by
onTheNetwork's caller or instatus: the control node does not resolve: more than one module provides acme-ca; pin one — the resolver's own words, which exist and were dropped. - A refusal that a release introduces for a machine already converged — a new rule the stored state does not meet — must not be silent at the roll either; the controller's prepare or first resolution after a roll should name every machine it now refuses.
How this would be checked: a controller test where one machine's unchecked resolution refuses:
status names the machine and the refusal, and the other machines keep their addresses.
Resolved in the live mesh, 2026-10-01
Two faults, one on top of the other. The refusal was mesh-controller 195's new rule — two modules
answering one provision on one machine need a pin — which its author hotfixed for the first pass
(196). The second pass of "the rest of the mesh" resolves every machine without its pins, so the
control node, pinned or not, was refused there and vanished: every seat it holds read as unheld,
the builder's and the proxy's builds were refused for want of the git seat's clone base, the
roll-out of the next controller was refused, and the first tiered plan failed at its first tier.
Found by a diagnostic build counting what each machine yielded. Fixed by mesh-controller PR
fix/a-machine-not-on-the-network-is-said: each machine is resolved with its own pins, and a machine
left out is named with the resolver's words in both places. The pin itself (step-ca, the issuer
the proxy already had) was made by hand and stands.
Resolved, 2026-10-01 evening
The controller holding the fix was rolled onto the control node by the operator and a colleague (the running one could not roll itself); after it every seat read as held again, a build asked through the git seat worked, all four machines resolved and pushed. The line naming a dropped machine spoke once too often — in the per-machine view, where the others are resolved without the planned machine's offers and may fail by design — and is quiet there since mesh-controller PR 200.