One dropped lookup should not cost a two-hour run #24

Merged
jschoubben merged 1 commits from fix/one-dropped-lookup-should-not-cost-a-run into main 2026-09-14 16:23:49 +00:00
Owner

Three runs have now died at a DNS timeout, each one well after the egress check had passed. The path was never broken — a query went unanswered while four machines were pulling images at once, and the uplink's resolver is a single server under exactly that load.

The check itself is unchanged and still uses the uplink alone, because proving that path is the point of it. What is added happens after it passes: the uplink stays first and keeps being asked, public resolvers sit behind it and answer only when it does not, and the retry is tightened so a silent server costs seconds rather than a step.

A genuinely broken uplink still fails the check, before any of this applies.

Worth the change because the loop is two hours wide: a lookup that drops at minute forty is not visible until minute ninety, and it has cost three full cycles.

Three runs have now died at a DNS timeout, each one well after the egress check had passed. The path was never broken — a query went unanswered while four machines were pulling images at once, and the uplink's resolver is a single server under exactly that load. The check itself is unchanged and still uses the uplink alone, because proving that path is the point of it. What is added happens after it passes: the uplink stays first and keeps being asked, public resolvers sit behind it and answer only when it does not, and the retry is tightened so a silent server costs seconds rather than a step. A genuinely broken uplink still fails the check, before any of this applies. Worth the change because the loop is two hours wide: a lookup that drops at minute forty is not visible until minute ninety, and it has cost three full cycles.
jschoubben added 1 commit 2026-09-14 16:23:42 +00:00
The egress check proves the uplink resolves and reaches the internet, and that
is the right thing to check. What it does not cover is the hours afterwards: a
bed pulls images on four machines at once, the uplink's resolver is one server
under exactly that load, and three runs have now died at a lookup timeout long
after the check passed — the path was never broken, a query just went
unanswered.

The uplink stays first and keeps proving the path. Public resolvers sit behind
it and answer only when it does not, and the retry is tightened so a silent
server costs seconds. A genuinely broken uplink still fails the check, before
any of this applies.
jschoubben merged commit eed5647146 into main 2026-09-14 16:23:49 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: novox/mesh-lab#24