The mesh asserts three things are callable (ADR 0144) — what runs on the same machine, another machine's service exposed to the private network, and another machine's service exposed publicly — and has never checked any of them. The first was broken for eleven hours while the mesh reported every machine healthy. This runs on every machine, on the cadence the mesh already has, in its own container: the same position every other module calls from. Not the host and not the control plane, both of which reach these addresses by paths no ordinary caller uses and would have passed throughout that outage. **Its probe is its own endpoint, and that is the point.** Declared reachable over the private network like any other service, so it is admitted by exactly the rule that governs every internally-exposed service and fails when that rule is wrong. The tempting target is a service every machine has, and those are the ones never closed — ssh above all — which would have passed while the thing that actually broke was a service exposed to the private network. It resolves before it dials and says which failed, because a name that does not resolve and a port that does not answer have different owners. One failure is not a fault: a machine rebooting is ordinary, so a path is broken after consecutive runs and the count travels with the result. It reports and repairs nothing. novox/hq ADR 0145. Eight tests; the consecutive-failure logic proved by reverting it once. Not yet registered or assigned.
9 lines
315 B
JSON
9 lines
315 B
JSON
{
|
|
"name": "@novox/module-network-checker",
|
|
"version": "0.1.0",
|
|
"description": "network-checker — dials what the mesh claims is reachable, from where the callers are, and says what it found.",
|
|
"type": "module",
|
|
"private": true,
|
|
"devDependencies": { "@types/node": "^22.0.0", "typescript": "^5.6.0" }
|
|
}
|