Merge pull request 'Issues 260, 262: a service waits for what it reads; a mesh name has no IPv6 address rather than no name' (#111) from issues/260-262 into main
This commit was merged in pull request #111.
This commit is contained in:
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: open
|
||||
status: located
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-host, mesh-catalog]
|
||||
fixed-by:
|
||||
fixed-by: novox/mesh-host#25 (a566add)
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -38,3 +38,13 @@ machine's name resolution with it.
|
||||
|
||||
- Should the host run every restart and reload after all files of the apply are written?
|
||||
- Or should it order each restart after every resource the service names in `restart-on`?
|
||||
|
||||
## Fix
|
||||
|
||||
The host now orders the apply so that a service comes after every resource it names under
|
||||
`restart-on` or `reload-on`. Only services move, and only as far as the last of what they name; every
|
||||
other resource keeps its declared place. The first open question above is answered that way, per
|
||||
service, rather than by moving every restart to the end. This is checked by mesh-host's
|
||||
`restart_order_test.go`: a service restarting on two files, one declared after it, is restarted once
|
||||
both exist, and a later change to the second file still restarts it. The test fails without the
|
||||
ordering.
|
||||
|
||||
+46
@@ -0,0 +1,46 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-catalog]
|
||||
fixed-by: novox/mesh-catalog#71 (20603b6)
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 262. An Alpine container could not find a machine by its mesh name
|
||||
|
||||
## Symptom
|
||||
|
||||
After the mesh moved to one resolver (ADR 0194, 0196), a workflow container on the home server was
|
||||
restarted so that it would ask the mesh's resolver. It then crash-looped every fourteen seconds:
|
||||
|
||||
> getaddrinfo ENOTFOUND <anchor>.internal
|
||||
|
||||
From the same container, `getent hosts <anchor>.internal` answered correctly. So did the machine
|
||||
itself, every time, in about forty-five milliseconds.
|
||||
|
||||
## Cause
|
||||
|
||||
The mesh's resolver answered a machine's name only through a wildcard rule:
|
||||
|
||||
- asked for the IPv4 address, it gave the address;
|
||||
- asked for the IPv6 address, it answered NXDOMAIN, "no such name", where the correct answer is
|
||||
NODATA, "the name exists and has no such record".
|
||||
|
||||
glibc ignores that. musl, the C library of every Alpine image, asks for both records and takes the
|
||||
NXDOMAIN as final, so the whole lookup failed. The per-machine resolvers this one replaced answered
|
||||
each machine's name from `/etc/hosts`, which gives NODATA, so nothing had depended on the difference
|
||||
before.
|
||||
|
||||
Reproduced with a throwaway resolver of the same version, given only the wildcard rule.
|
||||
|
||||
## Fix
|
||||
|
||||
Each machine's name is now also a host record in the machine list the resolver reads. Asked for an
|
||||
IPv6 address, the resolver then answers that the name exists and has none. A name under a machine,
|
||||
answered only by the wildcard, still answers NXDOMAIN for IPv6, exactly as it did before the move.
|
||||
|
||||
## How it is checked
|
||||
|
||||
Ask the mesh's resolver for the IPv6 address of a machine's name. The answer must be NOERROR with no
|
||||
records. Today that check is done by hand. A controller test that renders the machine list and
|
||||
requires one host record per machine should be added.
|
||||
Reference in New Issue
Block a user