3.2 KiB
status, opened, located-in, fixed-by, amended-design
| status | opened | located-in | fixed-by | amended-design | ||
|---|---|---|---|---|---|---|
| located | 2026-10-01 |
|
188 — A refusal inside "who is on the network" drops a machine silently, and every symptom points elsewhere
What was observed
At 15:28Z on 2026-10-01 the controller rolled to a build that refuses a machine with two modules
answering one provision and no pin naming which (mesh-controller 195). The control node had two
issuers of acme-ca. From that moment every plan of the control node failed with step-ca has a
content that says ${machine:at}, and this machine says mesh-range or name; seats listed every
seat as unheld; the build machine refused the builder's and the proxy's builds with no clone base
for that seat — nothing holds it; and the roll-out of the next controller was refused with the
${machine:at} words. Not one of those names the cause. It was found by running the previous image
as a one-shot beside the current one and reading the difference, forty minutes later.
Why this is here
onTheNetwork decides which machines have an address by resolving each one, unchecked, and
skipping any whose resolution errs. A machine skipped there has no at, so its own plan fails on
the first placeholder that needs one, in another module's words; everything held on it reads as
unheld; everything built from it cannot be built. The design lets one refusal become four unrelated
symptoms and no sentence about the refusal itself. It is the same shape as
issue 187: a fault that is swallowed
where it happens and discovered where it hurts.
What a fix needs
- A machine whose resolution refuses is said, by
onTheNetwork's caller or instatus: the control node does not resolve: more than one module provides acme-ca; pin one — the resolver's own words, which exist and were dropped. - A refusal that a release introduces for a machine already converged — a new rule the stored state does not meet — must not be silent at the roll either; the controller's prepare or first resolution after a roll should name every machine it now refuses.
How this would be checked: a controller test where one machine's unchecked resolution refuses:
status names the machine and the refusal, and the other machines keep their addresses.
Resolved in the live mesh, 2026-10-01
The refusal itself was a fault of mesh-controller 195, already corrected on main by its author's
hotfix (196) when the control node was found refusing; the running controller was the one build in
between. Found by running the previous image and main's image as one-shots beside the running one
and reading which resolved. Resolved by pushing the control node from a one-shot of main's image, as
the recipe for a controller that cannot roll itself says. A pin naming the proxy's issuer
(step-ca, the one the previous plan had bound) was made first and kept; it changes nothing. The
design fault above stands and is fixed by mesh-controller PR fix/a-machine-not-on-the-network-is-said:
the dropped machine and the resolver's words are said where the drop happens.