--- status: located opened: 2026-09-29 located-in: - mesh-controller internal/catalogue/filtering.go (fixed for this instance) - mesh-controller (what status reports, and what it does not ask) fixed-by: amended-design: --- # 145 — A machine reads healthy while its modules cannot reach each other ## What was observed Converging the control-node closed every path by which a module on that machine reached another module by the machine's own name. It ran for **eleven hours**. Throughout, the mesh answered: ``` 4 machine(s), all doing what they were told, all heard from, running what the mesh would send them, and every module current with its source ``` What was actually happening, from one affected module's own log: ``` Doctrine\DBAL\Exception: Failed to connect to the database: SQLSTATE[08006] connection to server at "novox.internal" (10.10.0.1), port 6852 failed: timeout expired ``` 6,154 of them, beginning at the minute of the flip. The web application accepted TCP connections and never answered an HTTP request; a client waited 35 seconds and gave up. Confirmed from a throwaway container on the machine: neither the store nor the forge was reachable on the machine's own address. The cause is [issue 144](../144-the-predecessors-rules-outlive-the-firewall-it-was-found-as/00-report.md)'s sibling and is fixed: a port declared reachable from the private network admitted the machines' own overlay addresses, and a container on the machine comes from a bridge address, matching none of them. What this issue is about is the eleven hours. **Nothing the mesh reports would have shown it.** Every check the mesh makes passed, because every check the mesh makes is about the relationship between the mesh and a machine: - the machine applied what it was sent, and said so; - its declaration digest matches what the mesh would send; - every module's source commit matches what the mesh holds; - every container the declaration names is running. None of those asks whether a module can reach what it requires. The mesh knows precisely who requires what — it composes the grants — and never checks that the grant works. **Nor would an operator's usual look.** The ports were probed from outside and behaved correctly; the routed services answered; a container's egress to the internet worked. Those are the paths a person checks after changing a firewall, and all three were fine. The broken path was module-to-module over the machine's own name, which nothing routine exercises. ## Why it matters beyond this instance **A mesh that composes a dependency and never tests it can only report on itself.** Every provision the mesh grants is a claim that a consumer can reach a provider. The mesh asserts that claim, delivers credentials for it, and has no mechanism that ever finds out. "Every module current with its source" is a statement about bytes, not about whether anything works. **The failure was silent in the direction that hides longest.** A service that will not start is noticed. A service that starts, accepts connections and then cannot reach its database serves errors under a healthy-looking process, and the machine's own report says the container is running — which it is. **It is the same shape as [issue 136](../136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md), one level up.** There, a module named a program the machine lacked and everything reported success. Here, the mesh granted a provision the filter refused and everything reported success. Both are the distance between a declaration and the machine, and in both cases the report was about the declaration. **And the eleven hours are the measurement, not the bug.** The filter fault was one line and is fixed. What is not fixed is that nothing in the mesh would have told anybody. ## Open questions - Should a grant be checked? The mesh knows the consumer, the provider, the address and the port, so a reachability check is expressible — but from where: the consumer's machine, as part of a reconcile, or the provider's? - What would it cost to be wrong in the other direction? A check that reports a provision broken while it works is worse than none, because it trains a reader to ignore the report. A provider restarting is ordinary; a consumer between containers is ordinary. - What should `status` say about a machine whose modules cannot reach each other? It currently has one vocabulary for "heard from and current", and that sentence was true the whole time. - Is there a cheaper signal than a probe? The affected module was logging the failure 6,154 times. The mesh reads no module's logs and arguably should not — but something a module could *say* about its own provisions would have surfaced this in minutes. - Does the same blindness apply to the other direction — a provider that lost a consumer's grant and is refusing it? Nothing checks that either.