145, partly resolved. The sentence that was true for eleven hours of a mesh in which no module could reach another now says what it is not a claim about: that is the mesh and the machines agreeing, and nothing here dials a provision. It checks nothing and does not pretend to — ADR 0146 decides the check and is deliberately not built. What changed is that the report no longer implies otherwise. Stays open for that reason. Carried forward: 0146's check needs an internal name fetched over TLS with the certificate verified, and until today no machine trusted the mesh's authority. Three of four do now, so whoever builds it does not have to solve that first. 107, diagnosed and deliberately not built. The premise is confirmed in the host's own words — unknown fields are refused because "a field the host does not know is a thing the control plane believes it asked for" — so the fix is a flag day, not an addition. 087 now makes the cost measurable, and the measurement is why it waits: one machine of four runs an older host, it is parked, and nothing delivers a host at all (142). Shipping the field means hand-placing binaries and unparking a machine, and one missed in that sequence is unreachable, not degraded. The fault it prevents has never been observed. 142 gains the note that it is 107's gate, and that it is what makes a declaration field cost a rollout instead of an expedition. A judgement about order, not a refusal, and cheap to overrule.
7.0 KiB
status, opened, located-in, fixed-by, amended-design
| status | opened | located-in | fixed-by | amended-design | ||
|---|---|---|---|---|---|---|
| located | 2026-09-29 |
|
partly — mesh-controller 1da96e8 makes the report state its own scope; nothing dials a provision yet, which is ADR 0146 and is not built | 03-DESIGN/01-to-be/10-delivery.md |
145 — A machine reads healthy while its modules cannot reach each other
What was observed
Converging the control-node closed every path by which a module on that machine reached another module by the machine's own name. It ran for eleven hours. Throughout, the mesh answered:
4 machine(s), all doing what they were told, all heard from,
running what the mesh would send them, and every module current with its source
What was actually happening, from one affected module's own log:
Doctrine\DBAL\Exception: Failed to connect to the database:
SQLSTATE[08006] connection to server at "novox.internal" (10.10.0.1), port 6852 failed: timeout expired
6,154 of them, beginning at the minute of the flip. The web application accepted TCP connections and never answered an HTTP request; a client waited 35 seconds and gave up. Confirmed from a throwaway container on the machine: neither the store nor the forge was reachable on the machine's own address.
The cause is issue 144's sibling and is fixed: a port declared reachable from the private network admitted the machines' own overlay addresses, and a container on the machine comes from a bridge address, matching none of them. What this issue is about is the eleven hours.
Nothing the mesh reports would have shown it. Every check the mesh makes passed, because every check the mesh makes is about the relationship between the mesh and a machine:
- the machine applied what it was sent, and said so;
- its declaration digest matches what the mesh would send;
- every module's source commit matches what the mesh holds;
- every container the declaration names is running.
None of those asks whether a module can reach what it requires. The mesh knows precisely who requires what — it composes the grants — and never checks that the grant works.
Nor would an operator's usual look. The ports were probed from outside and behaved correctly; the routed services answered; a container's egress to the internet worked. Those are the paths a person checks after changing a firewall, and all three were fine. The broken path was module-to-module over the machine's own name, which nothing routine exercises.
Why it matters beyond this instance
A mesh that composes a dependency and never tests it can only report on itself. Every provision the mesh grants is a claim that a consumer can reach a provider. The mesh asserts that claim, delivers credentials for it, and has no mechanism that ever finds out. "Every module current with its source" is a statement about bytes, not about whether anything works.
The failure was silent in the direction that hides longest. A service that will not start is noticed. A service that starts, accepts connections and then cannot reach its database serves errors under a healthy-looking process, and the machine's own report says the container is running — which it is.
It is the same shape as issue 136, one level up. There, a module named a program the machine lacked and everything reported success. Here, the mesh granted a provision the filter refused and everything reported success. Both are the distance between a declaration and the machine, and in both cases the report was about the declaration.
And the eleven hours are the measurement, not the bug. The filter fault was one line and is fixed. What is not fixed is that nothing in the mesh would have told anybody.
What was decided
2026-09-29, the same day, in two steps and the first was wrong.
The first answer was ADR 0143: the consumer verifies each grant from its own network position, because whether a caller sat in a container changed whether it could reach the provider. That difference was the fault, and the record is superseded. A verification mechanism would have reported this sooner and would not have prevented it, and the part of it that was difficult — deciding which network position to check from — existed only while the rule was wrong.
The remedy is ADR 0144: anything on a machine may call anything on it, said once rather than per service, and asked by the link traffic arrives on rather than the address it carries. Everything should be able to call what runs on the same machine, another machine's service exposed to the private network, and another machine's service exposed publicly. The filter had the second and third and expressed the first as a list of addresses that no container could match.
And then the question this issue is actually about was answered on its own terms. ADR 0145: a module on every machine serves an endpoint of its own and dials every other machine's, from the position the callers are in. Its probe is its own endpoint declared reachable over the private network, so it is admitted by exactly the rule that governs every internally-exposed service and fails when that rule is wrong — where a probe on a service every machine has would have passed for all eleven hours, because the services every machine has are the ones never closed.
Adopted on its merits rather than as the remedy for a configuration error, which is what 0143 was and why it went. The module is written and merged; it is not yet assigned, so every machine currently reads unchecked.
Open questions
- Should a grant be checked? The mesh knows the consumer, the provider, the address and the port, so a reachability check is expressible — but from where: the consumer's machine, as part of a reconcile, or the provider's?
- What would it cost to be wrong in the other direction? A check that reports a provision broken while it works is worse than none, because it trains a reader to ignore the report. A provider restarting is ordinary; a consumer between containers is ordinary.
- What should
statussay about a machine whose modules cannot reach each other? It currently has one vocabulary for "heard from and current", and that sentence was true the whole time. - Is there a cheaper signal than a probe? The affected module was logging the failure 6,154 times. The mesh reads no module's logs and arguably should not — but something a module could say about its own provisions would have surfaced this in minutes.
- Does the same blindness apply to the other direction — a provider that lost a consumer's grant and is refusing it? Nothing checks that either.