Issues 258–260: what the move to one resolver met
This commit is contained in:
@@ -0,0 +1,47 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-controller]
|
||||
fixed-by: novox/mesh-controller#59
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 258. Every machine bound the mesh's resolver to itself
|
||||
|
||||
## Symptom
|
||||
|
||||
After the mesh moved to one resolver (ADR 0194, 0196), the seat `mesh-dns-resolver` was recorded as
|
||||
held on the anchor, and each other machine was pinned to it for `wildcard-resolution`. Each machine
|
||||
was then pushed. The anchor's resolver configuration named its own private address first. Every
|
||||
other machine's did too: each named **its own** private address, not the anchor's.
|
||||
|
||||
Nothing broke, because each machine still ran a resolver of its own while the move was under way. But
|
||||
the mesh had decided on one resolver, recorded who held it and pinned every machine to it, and no
|
||||
machine used it.
|
||||
|
||||
## Cause
|
||||
|
||||
The controller binds a requirement in one of two branches:
|
||||
|
||||
- **Answered here**, when a module on the same machine provides it.
|
||||
- **Answered elsewhere**, when one on another machine does.
|
||||
|
||||
Only the second branch read the seat's holder (ADR 0110) and a pin. The first took the local provider,
|
||||
and read a pin only to choose between two local ones. Every machine still had its own resolver
|
||||
assigned, so every machine took the first branch.
|
||||
|
||||
The seat's own definition says the requirement "resolves to the holder wherever it is placed". The
|
||||
code did not.
|
||||
|
||||
## Fix
|
||||
|
||||
For a mesh-wide provision, a pin naming another machine, or the seat's holder on another machine,
|
||||
wins over a provider on this one. With neither, the local provider answers as before.
|
||||
|
||||
This is checked by three resolve tests in mesh-controller:
|
||||
|
||||
- the holder elsewhere answers;
|
||||
- a pin elsewhere answers;
|
||||
- a holder here still answers here.
|
||||
|
||||
The tests fail without the fix.
|
||||
@@ -0,0 +1,41 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-controller]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 259. A push to one machine sent every machine
|
||||
|
||||
## Symptom
|
||||
|
||||
A change to the resolver modules was merged with the upgrade policy `record`, so that each machine
|
||||
would get it only when pushed. The plan was to go one machine at a time: the anchor first, then each
|
||||
of the others, each checked before the next.
|
||||
|
||||
`push <anchor>` sent all four machines, each a new declaration carrying the change. A later
|
||||
`push <laptop>` did the same. A further fault in the change (issue 260) was therefore met on every
|
||||
machine at once, not on one.
|
||||
|
||||
## Cause
|
||||
|
||||
A named push ends by flushing every other machine whose declaration differs from what it was last
|
||||
sent (issue 057, ADR 0083). This was meant for consequences of the push, such as a provider's grant
|
||||
list after a consumer was assigned. It cannot tell a consequence from a change the upgrade policy is
|
||||
holding back. Under `record`, every machine running the module differs, so every machine is flushed.
|
||||
|
||||
Issue 249 met the same confusion for the machine holding the bus, and narrowed that check to the user
|
||||
list alone. The flush has no such narrowing.
|
||||
|
||||
## What it costs
|
||||
|
||||
`record` is the policy for a change that must be walked through the mesh by hand. A named push cannot
|
||||
do that, so the policy does not hold the change back from anything but the merge. The command's name
|
||||
says one machine, and the command's output is the only place that says otherwise.
|
||||
|
||||
## Open
|
||||
|
||||
- Should the flush send a machine only what changed as a consequence: grants, user lists and bound
|
||||
facts, not module versions held by a policy?
|
||||
- Or should it list such machines as behind and leave them, as `push --behind` would find them?
|
||||
@@ -0,0 +1,40 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-host, mesh-catalog]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 260. The resolver was restarted before the file it reads existed
|
||||
|
||||
## Symptom
|
||||
|
||||
The first apply of the resolver's new configuration failed on every machine. Each host's journal:
|
||||
|
||||
> failed dnsmasq.service: restarting dnsmasq.service: starting it again: systemctl exited 1
|
||||
|
||||
The resolver's own journal:
|
||||
|
||||
> dnsmasq: cannot read /etc/mesh-resolver/zones.conf: No such file or directory
|
||||
|
||||
The service manager's automatic restart started it again in the same second, and it answered. The
|
||||
host still reported the apply as failed, on every machine.
|
||||
|
||||
## Cause
|
||||
|
||||
The new configuration names a second file, the mesh's zones. The host wrote the configuration and
|
||||
restarted the service on it **before** it created the zones file. The zones file was created next in
|
||||
the same apply. The service's `restart-on` names both files, but nothing orders the restart after
|
||||
every file it reads has been written.
|
||||
|
||||
## What it costs
|
||||
|
||||
A configuration that adds a file it reads fails its first start everywhere. Here the service manager's
|
||||
restart policy hid it within a second. A service without one would have stayed down, holding the
|
||||
machine's name resolution with it.
|
||||
|
||||
## Open
|
||||
|
||||
- Should the host run every restart and reload after all files of the apply are written?
|
||||
- Or should it order each restart after every resource the service names in `restart-on`?
|
||||
Reference in New Issue
Block a user