Issue 112: diagnose — the predecessor's own DNS config already names the carried peers #101

Merged
jschoubben merged 3 commits from issue/112-diagnosis into main 2026-09-24 12:49:51 +00:00
Owner

Diagnosis pass on #112 before touching the resolver (HANDOFF §3.4).

Two findings:

  1. ADR 0104's forward-to-predecessor shape doesn't transfer to DNS the way it did the proxy. The proxy's predecessor kept running as a separate process throughout its migration. DNS is one process on one port — assigning the mesh's dnsmasq module replaces /etc/dnsmasq.conf whole and restarts the same dnsmasq.service unit the predecessor's own resolver is running as right now. There's no predecessor process left standing to forward to after cutover.

  2. The report's first open question — should a carried peer be nameable — doesn't need a guess. /etc/dnsmasq.d/hal-dns.conf, still live on the control-node, already states ace.internal/shanks.internal/g14.internal against 10.10.0.2/.3/.4. Checked against the mesh's own overlay show output for the carried tunnel peers: exact match, all three. It's a transcription of a record the predecessor already has and has been correctly serving for six days, not an assertion the mesh can't verify.

located-in updated: mesh-controller internal/inventory (where CarriedPeer/TunnelPeer records live today — public key and address, no name field) and cmd/mesh-controller (nothing today lets an operator attach a name to a carried-peer record). Separately, mesh-controller internal/catalogue (facts.go's nodeZones) would need to emit an address= wildcard for a named carried peer the same way it does for a node. mesh-catalog modules/dnsmasq itself needs no change — it already restarts on the node-zones fact and would pick up the new wildcard automatically.

status moved to located (playbook 03: diagnosing, then located once the owner is known — it is).

Not implemented — this is diagnosis only, per playbook 03. Checks pass (records.py, cycle.py, index.py).

Diagnosis pass on #112 before touching the resolver (HANDOFF §3.4). Two findings: 1. **ADR 0104's forward-to-predecessor shape doesn't transfer to DNS the way it did the proxy.** The proxy's predecessor kept running as a separate process throughout its migration. DNS is one process on one port — assigning the mesh's `dnsmasq` module replaces `/etc/dnsmasq.conf` whole and restarts the same `dnsmasq.service` unit the predecessor's own resolver is running as right now. There's no predecessor process left standing to forward to after cutover. 2. **The report's first open question — should a carried peer be nameable — doesn't need a guess.** `/etc/dnsmasq.d/hal-dns.conf`, still live on the control-node, already states `ace.internal`/`shanks.internal`/`g14.internal` against `10.10.0.2`/`.3`/`.4`. Checked against the mesh's own `overlay show` output for the carried tunnel peers: exact match, all three. It's a transcription of a record the predecessor already has and has been correctly serving for six days, not an assertion the mesh can't verify. `located-in` updated: `mesh-controller internal/inventory` (where `CarriedPeer`/`TunnelPeer` records live today — public key and address, no name field) and `cmd/mesh-controller` (nothing today lets an operator attach a name to a carried-peer record). Separately, `mesh-controller internal/catalogue` (`facts.go`'s `nodeZones`) would need to emit an `address=` wildcard for a named carried peer the same way it does for a node. `mesh-catalog modules/dnsmasq` itself needs no change — it already restarts on the `node-zones` fact and would pick up the new wildcard automatically. `status` moved to `located` (playbook 03: diagnosing, then located once the owner is known — it is). Not implemented — this is diagnosis only, per playbook 03. Checks pass (`records.py`, `cycle.py`, `index.py`).
jschoubben added 1 commit 2026-09-24 11:25:56 +00:00
Checked why ADR 0104's forward-to-predecessor shape doesn't transfer to the
resolver the way it did the proxy: DNS is one process on one port, and
assigning the mesh's dnsmasq module replaces it in place, so there is no
predecessor process left standing to forward to.

But /etc/dnsmasq.d/hal-dns.conf's static address= lines for ace/shanks/g14
match the mesh's own carried-peer addresses from overlay show exactly. The
name a carried peer needs isn't a guess the operator has to make under
pressure — it's a transcription of a record the predecessor already has and
has been correctly serving for six days. Located in mesh-controller (no way
to attach a name to a carried peer today) and the dnsmasq module (doesn't
emit a wildcard for a named-but-uncarried peer). Not implemented.
jschoubben added 1 commit 2026-09-24 12:01:49 +00:00
The carried-peer record (CarriedPeer/TunnelPeer) lives in mesh-controller
internal/inventory, not internal/catalogue. internal/catalogue is the
right package for the zone-generation side of the fix (facts.go's
nodeZones), but a different concern from where the name field itself
would go. Split the two so a decision record doesn't get pointed at the
wrong package.
jschoubben added 1 commit 2026-09-24 12:27:27 +00:00
Playbook 03 step 2: move status to diagnosing, then located once the
owner is known. located-in is filled with four confirmed packages —
the owner is known.
jschoubben merged commit 183b22997c into main 2026-09-24 12:49:51 +00:00
jschoubben deleted branch issue/112-diagnosis 2026-09-24 12:49:51 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: novox/hq#101