Issues 260, 262: a service waits for what it reads; a mesh name has no IPv6 address rather than no name

This commit is contained in:
jochen
2026-10-05 22:08:12 +02:00
parent 553d0c8893
commit eb4b4256a1
2 changed files with 58 additions and 2 deletions
@@ -1,8 +1,8 @@
---
status: open
status: located
opened: 2026-10-05
located-in: [mesh-host, mesh-catalog]
fixed-by:
fixed-by: novox/mesh-host#25 (a566add)
amended-design:
---
@@ -38,3 +38,13 @@ machine's name resolution with it.
- Should the host run every restart and reload after all files of the apply are written?
- Or should it order each restart after every resource the service names in `restart-on`?
## Fix
The host now orders the apply so that a service comes after every resource it names under
`restart-on` or `reload-on`. Only services move, and only as far as the last of what they name; every
other resource keeps its declared place. The first open question above is answered that way, per
service, rather than by moving every restart to the end. This is checked by mesh-host's
`restart_order_test.go`: a service restarting on two files, one declared after it, is restarted once
both exist, and a later change to the second file still restarts it. The test fails without the
ordering.
@@ -0,0 +1,46 @@
---
status: located
opened: 2026-10-05
located-in: [mesh-catalog]
fixed-by: novox/mesh-catalog#71 (20603b6)
amended-design:
---
# 262. An Alpine container could not find a machine by its mesh name
## Symptom
After the mesh moved to one resolver (ADR 0194, 0196), a workflow container on the home server was
restarted so that it would ask the mesh's resolver. It then crash-looped every fourteen seconds:
> getaddrinfo ENOTFOUND <anchor>.internal
From the same container, `getent hosts <anchor>.internal` answered correctly. So did the machine
itself, every time, in about forty-five milliseconds.
## Cause
The mesh's resolver answered a machine's name only through a wildcard rule:
- asked for the IPv4 address, it gave the address;
- asked for the IPv6 address, it answered NXDOMAIN, "no such name", where the correct answer is
NODATA, "the name exists and has no such record".
glibc ignores that. musl, the C library of every Alpine image, asks for both records and takes the
NXDOMAIN as final, so the whole lookup failed. The per-machine resolvers this one replaced answered
each machine's name from `/etc/hosts`, which gives NODATA, so nothing had depended on the difference
before.
Reproduced with a throwaway resolver of the same version, given only the wildcard rule.
## Fix
Each machine's name is now also a host record in the machine list the resolver reads. Asked for an
IPv6 address, the resolver then answers that the name exists and has none. A name under a machine,
answered only by the wildcard, still answers NXDOMAIN for IPv6, exactly as it did before the move.
## How it is checked
Ask the mesh's resolver for the IPv6 address of a machine's name. The answer must be NOERROR with no
records. Today that check is done by hand. A controller test that renders the machine list and
requires one host record per machine should be added.