Files
hq/04-ISSUES/079-every-machine-is-named-twice-over/01-diagnosis.md

2.2 KiB

Diagnosis — 2026-09-22

  1. The control plane collects every placed machine's name and address for a resolution keyed by the internal name (homer.internal), because that is the map every container is given as its hosts. The facts that render the hosts file and the resolver's zones took the same map as bare names and appended .internal to each — their unit tests fed them bare names and passed.
  2. Fixed in the facts: a name is read as either form, and each machine is written once as <machine>.internal with its bare name beside it. A test feeds both keyings and holds the files equal, and holds the zones free of a doubled suffix.

Located in: the facts. Not a decision. Proven by the unit test and by the large mesh bed's name tests, once its controller image is rebuilt from the fix.

On review. The first fix wrote the suffix a second time, in the facts, which is the shape that produced the defect: two places composing one name. The control plane now hands the suffix it composed the names with down to the facts, so an operator who chose another gets that one and nothing appended. What remains unproven by a bed is a mesh with a suffix other than the default.

Later. With the suffix right, the large bed's name test asked the second half of its question: a machine that leaves the private network must not be answered for. The names were every placed machine with an address, while the resolver's "on the private network" is a machine that also runs the module that puts it there. The names follow the resolver's rule now — one predicate, used by both — held by a controller test that unassigns one machine's networking and reads the names back.

Two consequences of the one predicate, on review. The filter's "from the mesh" set is that same set now, so a rule admits exactly the machines the mesh names. And a machine whose assigned set fails to resolve is no longer on the network in the mesh's eyes: its name leaves every other machine's hosts file and every resolver's zones on the next push, and the mesh's report of that machine's failure is what says why — the stated preference for a name that fails at once over a connection that hangs, at the cost of one machine's failure being visible mesh-wide.