Issue 145: a machine reads healthy while its modules cannot reach each other #179
+14
@@ -57,6 +57,20 @@ enrolment path the design guarantees.
|
||||
**The safe direction is not the same as the correct one.** Being more closed than intended broke nothing
|
||||
visible, which is exactly why it went unnoticed for as long as the mesh has been on this machine.
|
||||
|
||||
## What it cost, measured later the same day
|
||||
|
||||
*2026-09-29.* The predecessor's chain was removed, and something it had been carrying went with it. It
|
||||
admitted the private ranges wholesale, which is how a container on the machine reached a port declared
|
||||
for the private network — the mesh's own filter admits the machines' overlay addresses, and a container
|
||||
comes from a bridge. Every module that reached another by the machine's own name had been relying on the
|
||||
predecessor's rule without anybody knowing.
|
||||
|
||||
That is [issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
|
||||
and it ran for eleven hours while the mesh reported the machine healthy. The filter is fixed. What this
|
||||
adds to the account here is that "the machine is more closed than the mesh believes" was not the
|
||||
harmless direction after all — it was harmless for everything reached from outside, and an outage for
|
||||
everything reached from within.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should the host report every place the machine filters from, rather than one kind — the front-end,
|
||||
|
||||
+89
@@ -0,0 +1,89 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-09-29
|
||||
located-in:
|
||||
- mesh-controller internal/catalogue/filtering.go (fixed for this instance)
|
||||
- mesh-controller (what status reports, and what it does not ask)
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 145 — A machine reads healthy while its modules cannot reach each other
|
||||
|
||||
## What was observed
|
||||
|
||||
Converging the control-node closed every path by which a module on that machine reached another module
|
||||
by the machine's own name. It ran for **eleven hours**. Throughout, the mesh answered:
|
||||
|
||||
```
|
||||
4 machine(s), all doing what they were told, all heard from,
|
||||
running what the mesh would send them, and every module current with its source
|
||||
```
|
||||
|
||||
What was actually happening, from one affected module's own log:
|
||||
|
||||
```
|
||||
Doctrine\DBAL\Exception: Failed to connect to the database:
|
||||
SQLSTATE[08006] connection to server at "novox.internal" (10.10.0.1), port 6852 failed: timeout expired
|
||||
```
|
||||
|
||||
6,154 of them, beginning at the minute of the flip. The web application accepted TCP connections and
|
||||
never answered an HTTP request; a client waited 35 seconds and gave up. Confirmed from a throwaway
|
||||
container on the machine: neither the store nor the forge was reachable on the machine's own address.
|
||||
|
||||
The cause is [issue 144](../144-the-predecessors-rules-outlive-the-firewall-it-was-found-as/00-report.md)'s
|
||||
sibling and is fixed: a port declared reachable from the private network admitted the machines' own
|
||||
overlay addresses, and a container on the machine comes from a bridge address, matching none of them.
|
||||
What this issue is about is the eleven hours.
|
||||
|
||||
**Nothing the mesh reports would have shown it.** Every check the mesh makes passed, because every
|
||||
check the mesh makes is about the relationship between the mesh and a machine:
|
||||
|
||||
- the machine applied what it was sent, and said so;
|
||||
- its declaration digest matches what the mesh would send;
|
||||
- every module's source commit matches what the mesh holds;
|
||||
- every container the declaration names is running.
|
||||
|
||||
None of those asks whether a module can reach what it requires. The mesh knows precisely who requires
|
||||
what — it composes the grants — and never checks that the grant works.
|
||||
|
||||
**Nor would an operator's usual look.** The ports were probed from outside and behaved correctly; the
|
||||
routed services answered; a container's egress to the internet worked. Those are the paths a person
|
||||
checks after changing a firewall, and all three were fine. The broken path was module-to-module over
|
||||
the machine's own name, which nothing routine exercises.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
**A mesh that composes a dependency and never tests it can only report on itself.** Every provision the
|
||||
mesh grants is a claim that a consumer can reach a provider. The mesh asserts that claim, delivers
|
||||
credentials for it, and has no mechanism that ever finds out. "Every module current with its source"
|
||||
is a statement about bytes, not about whether anything works.
|
||||
|
||||
**The failure was silent in the direction that hides longest.** A service that will not start is
|
||||
noticed. A service that starts, accepts connections and then cannot reach its database serves errors
|
||||
under a healthy-looking process, and the machine's own report says the container is running — which it
|
||||
is.
|
||||
|
||||
**It is the same shape as [issue 136](../136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md),
|
||||
one level up.** There, a module named a program the machine lacked and everything reported success.
|
||||
Here, the mesh granted a provision the filter refused and everything reported success. Both are the
|
||||
distance between a declaration and the machine, and in both cases the report was about the declaration.
|
||||
|
||||
**And the eleven hours are the measurement, not the bug.** The filter fault was one line and is fixed.
|
||||
What is not fixed is that nothing in the mesh would have told anybody.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should a grant be checked? The mesh knows the consumer, the provider, the address and the port, so a
|
||||
reachability check is expressible — but from where: the consumer's machine, as part of a reconcile,
|
||||
or the provider's?
|
||||
- What would it cost to be wrong in the other direction? A check that reports a provision broken while
|
||||
it works is worse than none, because it trains a reader to ignore the report. A provider restarting is
|
||||
ordinary; a consumer between containers is ordinary.
|
||||
- What should `status` say about a machine whose modules cannot reach each other? It currently has one
|
||||
vocabulary for "heard from and current", and that sentence was true the whole time.
|
||||
- Is there a cheaper signal than a probe? The affected module was logging the failure 6,154 times. The
|
||||
mesh reads no module's logs and arguably should not — but something a module could *say* about its
|
||||
own provisions would have surfaced this in minutes.
|
||||
- Does the same blindness apply to the other direction — a provider that lost a consumer's grant and
|
||||
is refusing it? Nothing checks that either.
|
||||
Reference in New Issue
Block a user