--- status: located opened: 2026-10-01 located-in: [mesh-controller cmd/mesh-controller/network.go (onTheNetwork resolves every machine unchecked and skips one that refuses, saying nothing)] fixed-by: amended-design: [] --- # 188 — A refusal inside "who is on the network" drops a machine silently, and every symptom points elsewhere ## What was observed At 15:28Z on 2026-10-01 the controller rolled to a build that refuses a machine with two modules answering one provision and no pin naming which (mesh-controller 195). The control node had two issuers of `acme-ca`. From that moment every plan of the control node failed with *step-ca has a content that says `${machine:at}`, and this machine says mesh-range or name*; `seats` listed every seat as unheld; the build machine refused the builder's and the proxy's builds with *no clone base for that seat — nothing holds it*; and the roll-out of the next controller was refused with the `${machine:at}` words. Not one of those names the cause. It was found by running the previous image as a one-shot beside the current one and reading the difference, forty minutes later. ## Why this is here `onTheNetwork` decides which machines have an address by resolving each one, unchecked, and *skipping* any whose resolution errs. A machine skipped there has no `at`, so its own plan fails on the first placeholder that needs one, in another module's words; everything held on it reads as unheld; everything built from it cannot be built. The design lets one refusal become four unrelated symptoms and no sentence about the refusal itself. It is the same shape as [issue 187](../187-the-mesh-tells-nobody-when-it-stops-working/00-report.md): a fault that is swallowed where it happens and discovered where it hurts. ## What a fix needs - A machine whose resolution refuses is said, by `onTheNetwork`'s caller or in `status`: *the control node does not resolve: more than one module provides acme-ca; pin one* — the resolver's own words, which exist and were dropped. - A refusal that a release introduces for a machine already converged — a new rule the stored state does not meet — must not be silent at the roll either; the controller's prepare or first resolution after a roll should name every machine it now refuses. *How this would be checked:* a controller test where one machine's unchecked resolution refuses: `status` names the machine and the refusal, and the other machines keep their addresses. ## Resolved in the live mesh, 2026-10-01 By hand: `pin novox acme-ca novox step-ca`, the provider the previous controller had in fact bound the proxy to (read from its plan), after which the control node resolved, every seat read as held and the roll-out proceeded. The design fault above stands.