Files
jschoubben 492ac7be18 Issue 145: a machine reads healthy while its modules cannot reach each other
Converging the control node closed every path by which a module reached another by
the machine's own name, and it ran for eleven hours while the mesh answered 'all
doing what they were told, all heard from, running what the mesh would send them,
and every module current with its source'.

6,154 database failures in one module's log, beginning at the minute of the flip.
The service accepted TCP and never answered HTTP. Every check the mesh makes
passed, because every check the mesh makes is about the relationship between the
mesh and a machine — applied, current, containers running. None asks whether a
module can reach what it requires, though the mesh composes every grant and so
knows exactly who requires what.

The filter fault is fixed. The eleven hours are the measurement, not the bug.

Also records on 144 that 'more closed than the mesh believes' was harmless only
for what is reached from outside, and an outage for what is reached from within.
2026-09-29 12:37:08 +02:00

4.8 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
located 2026-09-29
mesh-controller internal/catalogue/filtering.go (fixed for this instance)
mesh-controller (what status reports, and what it does not ask)

145 — A machine reads healthy while its modules cannot reach each other

What was observed

Converging the control-node closed every path by which a module on that machine reached another module by the machine's own name. It ran for eleven hours. Throughout, the mesh answered:

4 machine(s), all doing what they were told, all heard from,
running what the mesh would send them, and every module current with its source

What was actually happening, from one affected module's own log:

Doctrine\DBAL\Exception: Failed to connect to the database:
SQLSTATE[08006] connection to server at "novox.internal" (10.10.0.1), port 6852 failed: timeout expired

6,154 of them, beginning at the minute of the flip. The web application accepted TCP connections and never answered an HTTP request; a client waited 35 seconds and gave up. Confirmed from a throwaway container on the machine: neither the store nor the forge was reachable on the machine's own address.

The cause is issue 144's sibling and is fixed: a port declared reachable from the private network admitted the machines' own overlay addresses, and a container on the machine comes from a bridge address, matching none of them. What this issue is about is the eleven hours.

Nothing the mesh reports would have shown it. Every check the mesh makes passed, because every check the mesh makes is about the relationship between the mesh and a machine:

  • the machine applied what it was sent, and said so;
  • its declaration digest matches what the mesh would send;
  • every module's source commit matches what the mesh holds;
  • every container the declaration names is running.

None of those asks whether a module can reach what it requires. The mesh knows precisely who requires what — it composes the grants — and never checks that the grant works.

Nor would an operator's usual look. The ports were probed from outside and behaved correctly; the routed services answered; a container's egress to the internet worked. Those are the paths a person checks after changing a firewall, and all three were fine. The broken path was module-to-module over the machine's own name, which nothing routine exercises.

Why it matters beyond this instance

A mesh that composes a dependency and never tests it can only report on itself. Every provision the mesh grants is a claim that a consumer can reach a provider. The mesh asserts that claim, delivers credentials for it, and has no mechanism that ever finds out. "Every module current with its source" is a statement about bytes, not about whether anything works.

The failure was silent in the direction that hides longest. A service that will not start is noticed. A service that starts, accepts connections and then cannot reach its database serves errors under a healthy-looking process, and the machine's own report says the container is running — which it is.

It is the same shape as issue 136, one level up. There, a module named a program the machine lacked and everything reported success. Here, the mesh granted a provision the filter refused and everything reported success. Both are the distance between a declaration and the machine, and in both cases the report was about the declaration.

And the eleven hours are the measurement, not the bug. The filter fault was one line and is fixed. What is not fixed is that nothing in the mesh would have told anybody.

Open questions

  • Should a grant be checked? The mesh knows the consumer, the provider, the address and the port, so a reachability check is expressible — but from where: the consumer's machine, as part of a reconcile, or the provider's?
  • What would it cost to be wrong in the other direction? A check that reports a provision broken while it works is worse than none, because it trains a reader to ignore the report. A provider restarting is ordinary; a consumer between containers is ordinary.
  • What should status say about a machine whose modules cannot reach each other? It currently has one vocabulary for "heard from and current", and that sentence was true the whole time.
  • Is there a cheaper signal than a probe? The affected module was logging the failure 6,154 times. The mesh reads no module's logs and arguably should not — but something a module could say about its own provisions would have surfaced this in minutes.
  • Does the same blindness apply to the other direction — a provider that lost a consumer's grant and is refusing it? Nothing checks that either.