A network-checker module: dial what the mesh claims, from where the callers are #140

Merged
mesh-admin merged 1 commits from feat/a-network-checker-module into main 2026-09-29 11:44:37 +00:00
Contributor

novox/hq ADR 0145, answering the question issue 145 left open: the mesh asserts three things are
callable and has never checked any of them.

It runs on every machine, every five minutes, in its own container — the same position every other
module on that machine calls from. Not the host and not the control plane: both reach these addresses
by paths no ordinary caller uses, and both would have passed throughout the outage that produced this.

The probe is this module's own endpoint, and that is the whole design. Declared reachable over the
private network like any other service, so it is admitted by exactly the rule that governs every
internally-exposed service — and fails when that rule is wrong. The tempting target is a service every
machine has, and the services every machine has are the ones that are never closed, ssh above all.
Dialling ssh would have passed for all eleven hours, because ssh is admitted unconditionally and what
broke was a service exposed to the private network. A probe on a port that cannot fail measures nothing.

What it checks, named so a failure says which claim broke:

  • its own machine, by address and by name — the case that distinguishes a caller on the machine from
    a caller in one of its containers, and the one that broke
  • every other machine over the private network
  • each machine's public path, only where a public name is known, because a machine with no public
    face has nothing to fail

It resolves before it dials and reports which step failed, since a name that does not resolve and a port
that does not answer have different owners and one sentence for both sends a reader to the wrong place.

One failure is not a fault — a machine rebooting is ordinary — so a path is reported broken after
consecutive runs, and the count travels with the result so "briefly away" and "never worked" are
distinguishable. It reports and repairs nothing; a checker that fixed things would be a second control
plane. A broken path exits zero, because a scheduled step that fails is retried rather than believed.

Built on what already exists: a container with schedule: "*/5 * * * *" (three modules use it), the
roster rendered as a fact (the resolver and the operator's ssh config use the same mechanism), and a
bundle artifact. No new vocabulary anywhere.

Eight tests. The consecutive-failure logic proved by reverting it once: the "one failure is not a fault"
case keeps passing and the two that depend on carrying a count fail.

Not registered and not assigned — that is a deliberate next step, since a machine without it reads
unchecked rather than healthy and I would rather that be a decision than a side effect of a merge.

novox/hq ADR 0145, answering the question issue 145 left open: the mesh asserts three things are callable and has never checked any of them. It runs on every machine, every five minutes, in its own container — the same position every other module on that machine calls from. Not the host and not the control plane: both reach these addresses by paths no ordinary caller uses, and both would have passed throughout the outage that produced this. **The probe is this module's own endpoint, and that is the whole design.** Declared reachable over the private network like any other service, so it is admitted by exactly the rule that governs every internally-exposed service — and fails when that rule is wrong. The tempting target is a service every machine has, and the services every machine has are the ones that are never closed, ssh above all. Dialling ssh would have passed for all eleven hours, because ssh is admitted unconditionally and what broke was a service exposed to the private network. A probe on a port that cannot fail measures nothing. What it checks, named so a failure says which claim broke: - **its own machine**, by address and by name — the case that distinguishes a caller on the machine from a caller in one of its containers, and the one that broke - **every other machine over the private network** - **each machine's public path**, only where a public name is known, because a machine with no public face has nothing to fail It resolves before it dials and reports which step failed, since a name that does not resolve and a port that does not answer have different owners and one sentence for both sends a reader to the wrong place. One failure is not a fault — a machine rebooting is ordinary — so a path is reported broken after consecutive runs, and the count travels with the result so "briefly away" and "never worked" are distinguishable. It reports and repairs nothing; a checker that fixed things would be a second control plane. A broken path exits zero, because a scheduled step that fails is retried rather than believed. Built on what already exists: a container with `schedule: "*/5 * * * *"` (three modules use it), the roster rendered as a `fact` (the resolver and the operator's ssh config use the same mechanism), and a `bundle` artifact. No new vocabulary anywhere. Eight tests. The consecutive-failure logic proved by reverting it once: the "one failure is not a fault" case keeps passing and the two that depend on carrying a count fail. Not registered and not assigned — that is a deliberate next step, since a machine without it reads unchecked rather than healthy and I would rather that be a decision than a side effect of a merge.
mesh-admin added 1 commit 2026-09-29 11:44:36 +00:00
The mesh asserts three things are callable (ADR 0144) — what runs on the same
machine, another machine's service exposed to the private network, and another
machine's service exposed publicly — and has never checked any of them. The first
was broken for eleven hours while the mesh reported every machine healthy.

This runs on every machine, on the cadence the mesh already has, in its own
container: the same position every other module calls from. Not the host and not
the control plane, both of which reach these addresses by paths no ordinary caller
uses and would have passed throughout that outage.

**Its probe is its own endpoint, and that is the point.** Declared reachable over
the private network like any other service, so it is admitted by exactly the rule
that governs every internally-exposed service and fails when that rule is wrong.
The tempting target is a service every machine has, and those are the ones never
closed — ssh above all — which would have passed while the thing that actually
broke was a service exposed to the private network.

It resolves before it dials and says which failed, because a name that does not
resolve and a port that does not answer have different owners. One failure is not
a fault: a machine rebooting is ordinary, so a path is broken after consecutive
runs and the count travels with the result. It reports and repairs nothing.

novox/hq ADR 0145. Eight tests; the consecutive-failure logic proved by reverting
it once. Not yet registered or assigned.
mesh-admin merged commit bbac08a7d2 into main 2026-09-29 11:44:37 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: novox/mesh-catalog#140