diff --git a/04-ISSUES/144-the-predecessors-rules-outlive-the-firewall-it-was-found-as/00-report.md b/04-ISSUES/144-the-predecessors-rules-outlive-the-firewall-it-was-found-as/00-report.md index 63ac152..3a9dd56 100644 --- a/04-ISSUES/144-the-predecessors-rules-outlive-the-firewall-it-was-found-as/00-report.md +++ b/04-ISSUES/144-the-predecessors-rules-outlive-the-firewall-it-was-found-as/00-report.md @@ -57,6 +57,20 @@ enrolment path the design guarantees. **The safe direction is not the same as the correct one.** Being more closed than intended broke nothing visible, which is exactly why it went unnoticed for as long as the mesh has been on this machine. +## What it cost, measured later the same day + +*2026-09-29.* The predecessor's chain was removed, and something it had been carrying went with it. It +admitted the private ranges wholesale, which is how a container on the machine reached a port declared +for the private network — the mesh's own filter admits the machines' overlay addresses, and a container +comes from a bridge. Every module that reached another by the machine's own name had been relying on the +predecessor's rule without anybody knowing. + +That is [issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md), +and it ran for eleven hours while the mesh reported the machine healthy. The filter is fixed. What this +adds to the account here is that "the machine is more closed than the mesh believes" was not the +harmless direction after all — it was harmless for everything reached from outside, and an outage for +everything reached from within. + ## Open questions - Should the host report every place the machine filters from, rather than one kind — the front-end, diff --git a/04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md b/04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md new file mode 100644 index 0000000..65afda0 --- /dev/null +++ b/04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md @@ -0,0 +1,89 @@ +--- +status: located +opened: 2026-09-29 +located-in: + - mesh-controller internal/catalogue/filtering.go (fixed for this instance) + - mesh-controller (what status reports, and what it does not ask) +fixed-by: +amended-design: +--- + +# 145 — A machine reads healthy while its modules cannot reach each other + +## What was observed + +Converging the control-node closed every path by which a module on that machine reached another module +by the machine's own name. It ran for **eleven hours**. Throughout, the mesh answered: + +``` +4 machine(s), all doing what they were told, all heard from, +running what the mesh would send them, and every module current with its source +``` + +What was actually happening, from one affected module's own log: + +``` +Doctrine\DBAL\Exception: Failed to connect to the database: +SQLSTATE[08006] connection to server at "novox.internal" (10.10.0.1), port 6852 failed: timeout expired +``` + +6,154 of them, beginning at the minute of the flip. The web application accepted TCP connections and +never answered an HTTP request; a client waited 35 seconds and gave up. Confirmed from a throwaway +container on the machine: neither the store nor the forge was reachable on the machine's own address. + +The cause is [issue 144](../144-the-predecessors-rules-outlive-the-firewall-it-was-found-as/00-report.md)'s +sibling and is fixed: a port declared reachable from the private network admitted the machines' own +overlay addresses, and a container on the machine comes from a bridge address, matching none of them. +What this issue is about is the eleven hours. + +**Nothing the mesh reports would have shown it.** Every check the mesh makes passed, because every +check the mesh makes is about the relationship between the mesh and a machine: + +- the machine applied what it was sent, and said so; +- its declaration digest matches what the mesh would send; +- every module's source commit matches what the mesh holds; +- every container the declaration names is running. + +None of those asks whether a module can reach what it requires. The mesh knows precisely who requires +what — it composes the grants — and never checks that the grant works. + +**Nor would an operator's usual look.** The ports were probed from outside and behaved correctly; the +routed services answered; a container's egress to the internet worked. Those are the paths a person +checks after changing a firewall, and all three were fine. The broken path was module-to-module over +the machine's own name, which nothing routine exercises. + +## Why it matters beyond this instance + +**A mesh that composes a dependency and never tests it can only report on itself.** Every provision the +mesh grants is a claim that a consumer can reach a provider. The mesh asserts that claim, delivers +credentials for it, and has no mechanism that ever finds out. "Every module current with its source" +is a statement about bytes, not about whether anything works. + +**The failure was silent in the direction that hides longest.** A service that will not start is +noticed. A service that starts, accepts connections and then cannot reach its database serves errors +under a healthy-looking process, and the machine's own report says the container is running — which it +is. + +**It is the same shape as [issue 136](../136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md), +one level up.** There, a module named a program the machine lacked and everything reported success. +Here, the mesh granted a provision the filter refused and everything reported success. Both are the +distance between a declaration and the machine, and in both cases the report was about the declaration. + +**And the eleven hours are the measurement, not the bug.** The filter fault was one line and is fixed. +What is not fixed is that nothing in the mesh would have told anybody. + +## Open questions + +- Should a grant be checked? The mesh knows the consumer, the provider, the address and the port, so a + reachability check is expressible — but from where: the consumer's machine, as part of a reconcile, + or the provider's? +- What would it cost to be wrong in the other direction? A check that reports a provision broken while + it works is worse than none, because it trains a reader to ignore the report. A provider restarting is + ordinary; a consumer between containers is ordinary. +- What should `status` say about a machine whose modules cannot reach each other? It currently has one + vocabulary for "heard from and current", and that sentence was true the whole time. +- Is there a cheaper signal than a probe? The affected module was logging the failure 6,154 times. The + mesh reads no module's logs and arguably should not — but something a module could *say* about its + own provisions would have surfaced this in minutes. +- Does the same blindness apply to the other direction — a provider that lost a consumer's grant and + is refusing it? Nothing checks that either.