Issues 258–260: what the move to one resolver met

This commit is contained in:
jochen
2026-10-05 21:15:09 +02:00
parent 55e776a1cd
commit a0de734ed6
3 changed files with 128 additions and 0 deletions
@@ -0,0 +1,47 @@
---
status: located
opened: 2026-10-05
located-in: [mesh-controller]
fixed-by: novox/mesh-controller#59
amended-design:
---
# 258. Every machine bound the mesh's resolver to itself
## Symptom
After the mesh moved to one resolver (ADR 0194, 0196), the seat `mesh-dns-resolver` was recorded as
held on the anchor, and each other machine was pinned to it for `wildcard-resolution`. Each machine
was then pushed. The anchor's resolver configuration named its own private address first. Every
other machine's did too: each named **its own** private address, not the anchor's.
Nothing broke, because each machine still ran a resolver of its own while the move was under way. But
the mesh had decided on one resolver, recorded who held it and pinned every machine to it, and no
machine used it.
## Cause
The controller binds a requirement in one of two branches:
- **Answered here**, when a module on the same machine provides it.
- **Answered elsewhere**, when one on another machine does.
Only the second branch read the seat's holder (ADR 0110) and a pin. The first took the local provider,
and read a pin only to choose between two local ones. Every machine still had its own resolver
assigned, so every machine took the first branch.
The seat's own definition says the requirement "resolves to the holder wherever it is placed". The
code did not.
## Fix
For a mesh-wide provision, a pin naming another machine, or the seat's holder on another machine,
wins over a provider on this one. With neither, the local provider answers as before.
This is checked by three resolve tests in mesh-controller:
- the holder elsewhere answers;
- a pin elsewhere answers;
- a holder here still answers here.
The tests fail without the fix.
@@ -0,0 +1,41 @@
---
status: open
opened: 2026-10-05
located-in: [mesh-controller]
fixed-by:
amended-design:
---
# 259. A push to one machine sent every machine
## Symptom
A change to the resolver modules was merged with the upgrade policy `record`, so that each machine
would get it only when pushed. The plan was to go one machine at a time: the anchor first, then each
of the others, each checked before the next.
`push <anchor>` sent all four machines, each a new declaration carrying the change. A later
`push <laptop>` did the same. A further fault in the change (issue 260) was therefore met on every
machine at once, not on one.
## Cause
A named push ends by flushing every other machine whose declaration differs from what it was last
sent (issue 057, ADR 0083). This was meant for consequences of the push, such as a provider's grant
list after a consumer was assigned. It cannot tell a consequence from a change the upgrade policy is
holding back. Under `record`, every machine running the module differs, so every machine is flushed.
Issue 249 met the same confusion for the machine holding the bus, and narrowed that check to the user
list alone. The flush has no such narrowing.
## What it costs
`record` is the policy for a change that must be walked through the mesh by hand. A named push cannot
do that, so the policy does not hold the change back from anything but the merge. The command's name
says one machine, and the command's output is the only place that says otherwise.
## Open
- Should the flush send a machine only what changed as a consequence: grants, user lists and bound
facts, not module versions held by a policy?
- Or should it list such machines as behind and leave them, as `push --behind` would find them?
@@ -0,0 +1,40 @@
---
status: open
opened: 2026-10-05
located-in: [mesh-host, mesh-catalog]
fixed-by:
amended-design:
---
# 260. The resolver was restarted before the file it reads existed
## Symptom
The first apply of the resolver's new configuration failed on every machine. Each host's journal:
> failed dnsmasq.service: restarting dnsmasq.service: starting it again: systemctl exited 1
The resolver's own journal:
> dnsmasq: cannot read /etc/mesh-resolver/zones.conf: No such file or directory
The service manager's automatic restart started it again in the same second, and it answered. The
host still reported the apply as failed, on every machine.
## Cause
The new configuration names a second file, the mesh's zones. The host wrote the configuration and
restarted the service on it **before** it created the zones file. The zones file was created next in
the same apply. The service's `restart-on` names both files, but nothing orders the restart after
every file it reads has been written.
## What it costs
A configuration that adds a file it reads fails its first start everywhere. Here the service manager's
restart policy hid it within a second. A service without one would have stayed down, holding the
machine's name resolution with it.
## Open
- Should the host run every restart and reload after all files of the apply are written?
- Or should it order each restart after every resource the service names in `restart-on`?