From a0de734ed68534b9540b8a6ca3fa8d6425d62326 Mon Sep 17 00:00:00 2001 From: jochen Date: Mon, 5 Oct 2026 21:15:09 +0200 Subject: [PATCH] =?UTF-8?q?Issues=20258=E2=80=93260:=20what=20the=20move?= =?UTF-8?q?=20to=20one=20resolver=20met?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .../00-report.md | 47 +++++++++++++++++++ .../00-report.md | 41 ++++++++++++++++ .../00-report.md | 40 ++++++++++++++++ 3 files changed, 128 insertions(+) create mode 100644 04-ISSUES/258-every-machine-bound-the-resolver-to-itself/00-report.md create mode 100644 04-ISSUES/259-a-named-push-sent-every-machine/00-report.md create mode 100644 04-ISSUES/260-the-resolver-started-before-its-zones-file-existed/00-report.md diff --git a/04-ISSUES/258-every-machine-bound-the-resolver-to-itself/00-report.md b/04-ISSUES/258-every-machine-bound-the-resolver-to-itself/00-report.md new file mode 100644 index 0000000..d22c268 --- /dev/null +++ b/04-ISSUES/258-every-machine-bound-the-resolver-to-itself/00-report.md @@ -0,0 +1,47 @@ +--- +status: located +opened: 2026-10-05 +located-in: [mesh-controller] +fixed-by: novox/mesh-controller#59 +amended-design: +--- + +# 258. Every machine bound the mesh's resolver to itself + +## Symptom + +After the mesh moved to one resolver (ADR 0194, 0196), the seat `mesh-dns-resolver` was recorded as +held on the anchor, and each other machine was pinned to it for `wildcard-resolution`. Each machine +was then pushed. The anchor's resolver configuration named its own private address first. Every +other machine's did too: each named **its own** private address, not the anchor's. + +Nothing broke, because each machine still ran a resolver of its own while the move was under way. But +the mesh had decided on one resolver, recorded who held it and pinned every machine to it, and no +machine used it. + +## Cause + +The controller binds a requirement in one of two branches: + +- **Answered here**, when a module on the same machine provides it. +- **Answered elsewhere**, when one on another machine does. + +Only the second branch read the seat's holder (ADR 0110) and a pin. The first took the local provider, +and read a pin only to choose between two local ones. Every machine still had its own resolver +assigned, so every machine took the first branch. + +The seat's own definition says the requirement "resolves to the holder wherever it is placed". The +code did not. + +## Fix + +For a mesh-wide provision, a pin naming another machine, or the seat's holder on another machine, +wins over a provider on this one. With neither, the local provider answers as before. + +This is checked by three resolve tests in mesh-controller: + +- the holder elsewhere answers; +- a pin elsewhere answers; +- a holder here still answers here. + +The tests fail without the fix. diff --git a/04-ISSUES/259-a-named-push-sent-every-machine/00-report.md b/04-ISSUES/259-a-named-push-sent-every-machine/00-report.md new file mode 100644 index 0000000..3bbe9d9 --- /dev/null +++ b/04-ISSUES/259-a-named-push-sent-every-machine/00-report.md @@ -0,0 +1,41 @@ +--- +status: open +opened: 2026-10-05 +located-in: [mesh-controller] +fixed-by: +amended-design: +--- + +# 259. A push to one machine sent every machine + +## Symptom + +A change to the resolver modules was merged with the upgrade policy `record`, so that each machine +would get it only when pushed. The plan was to go one machine at a time: the anchor first, then each +of the others, each checked before the next. + +`push ` sent all four machines, each a new declaration carrying the change. A later +`push ` did the same. A further fault in the change (issue 260) was therefore met on every +machine at once, not on one. + +## Cause + +A named push ends by flushing every other machine whose declaration differs from what it was last +sent (issue 057, ADR 0083). This was meant for consequences of the push, such as a provider's grant +list after a consumer was assigned. It cannot tell a consequence from a change the upgrade policy is +holding back. Under `record`, every machine running the module differs, so every machine is flushed. + +Issue 249 met the same confusion for the machine holding the bus, and narrowed that check to the user +list alone. The flush has no such narrowing. + +## What it costs + +`record` is the policy for a change that must be walked through the mesh by hand. A named push cannot +do that, so the policy does not hold the change back from anything but the merge. The command's name +says one machine, and the command's output is the only place that says otherwise. + +## Open + +- Should the flush send a machine only what changed as a consequence: grants, user lists and bound + facts, not module versions held by a policy? +- Or should it list such machines as behind and leave them, as `push --behind` would find them? diff --git a/04-ISSUES/260-the-resolver-started-before-its-zones-file-existed/00-report.md b/04-ISSUES/260-the-resolver-started-before-its-zones-file-existed/00-report.md new file mode 100644 index 0000000..f9b7888 --- /dev/null +++ b/04-ISSUES/260-the-resolver-started-before-its-zones-file-existed/00-report.md @@ -0,0 +1,40 @@ +--- +status: open +opened: 2026-10-05 +located-in: [mesh-host, mesh-catalog] +fixed-by: +amended-design: +--- + +# 260. The resolver was restarted before the file it reads existed + +## Symptom + +The first apply of the resolver's new configuration failed on every machine. Each host's journal: + +> failed dnsmasq.service: restarting dnsmasq.service: starting it again: systemctl exited 1 + +The resolver's own journal: + +> dnsmasq: cannot read /etc/mesh-resolver/zones.conf: No such file or directory + +The service manager's automatic restart started it again in the same second, and it answered. The +host still reported the apply as failed, on every machine. + +## Cause + +The new configuration names a second file, the mesh's zones. The host wrote the configuration and +restarted the service on it **before** it created the zones file. The zones file was created next in +the same apply. The service's `restart-on` names both files, but nothing orders the restart after +every file it reads has been written. + +## What it costs + +A configuration that adds a file it reads fails its first start everywhere. Here the service manager's +restart policy hid it within a second. A service without one would have stayed down, holding the +machine's name resolution with it. + +## Open + +- Should the host run every restart and reload after all files of the apply are written? +- Or should it order each restart after every resource the service names in `restart-on`?