Compare commits

..
Author SHA1 Message Date
jschoubben ddfd62edf6 ADR 0143: a consumer verifies the grant it is given
A grant is four facts and a credential — the provision, the machine, the port, and
who the consumer is when it connects — and it is the whole mechanism by which
anything in the mesh reaches anything else. The mesh asserted it and never found
out whether it was true. Issue 145 is what that cost: eleven hours of 'every module
current with its source' while a module could not reach its database.

The consumer verifies it, from its own network position. Not the control plane,
which reaches the address by a path no consumer uses. Not the machine, whose own
packets carried a source address the filter admitted while every container's was
refused — so a host-side check would have passed throughout the outage it exists to
catch. That is inference from the rule that was loaded, stated as such.

A connection and nothing more; speaking each provision's protocol would be a second
implementation of every provision. One failure is not news, only consecutive ones,
and the count is reported rather than the last attempt. A consumer that is not
running reads unchecked, which is a different sentence from broken.

What it costs to be wrong is the constraint on all of it: a check that calls a
working provision broken trains a reader to ignore the report.
2026-09-29 13:07:18 +02:00
mesh-admin f65664640a Merge pull request 'Issue 145: a machine reads healthy while its modules cannot reach each other' (#179) from issue/145-healthy-while-broken into main 2026-09-29 10:37:10 +00:00
jschoubben 492ac7be18 Issue 145: a machine reads healthy while its modules cannot reach each other
Converging the control node closed every path by which a module reached another by
the machine's own name, and it ran for eleven hours while the mesh answered 'all
doing what they were told, all heard from, running what the mesh would send them,
and every module current with its source'.

6,154 database failures in one module's log, beginning at the minute of the flip.
The service accepted TCP and never answered HTTP. Every check the mesh makes
passed, because every check the mesh makes is about the relationship between the
mesh and a machine — applied, current, containers running. None asks whether a
module can reach what it requires, though the mesh composes every grant and so
knows exactly who requires what.

The filter fault is fixed. The eleven hours are the measurement, not the bug.

Also records on 144 that 'more closed than the mesh believes' was harmless only
for what is reached from outside, and an outage for what is reached from within.
2026-09-29 12:37:08 +02:00
mesh-admin 5ac77e3cef Merge pull request 'ADR 0138: reach asks for names on a routed endpoint' (#178) from decision/0138-insight-reach-and-the-proxy into main 2026-09-29 00:50:28 +00:00
jschoubben a619022c35 ADR 0138: a progressive insight — reach asks for names on a routed endpoint
The record says internal and public each mean something to the filter. For an
endpoint the proxy serves, the second half is wrong, and ADR 0045 said so first:
a public service is exposed through the proxy, listening from the mesh, not by
opening its own port.

Found by trying to express one real module, not by review — routed name public
because browsers post to it, machine port private because it serves a dashboard in
cleartext. Under one value for both, saying public would have reopened a port an
operator had just closed. Measured the same evening: the routed name answered from
the internet over TLS while the port was refused from the same place.

Corrects a fact. One statement per endpoint with three things derived from it
stands; the filter column applies to an unrouted endpoint.
2026-09-29 02:50:25 +02:00
mesh-admin 49c065c204 Merge pull request 'Issue 143: correct the diagnosis' (#177) from issue/143-corrected-diagnosis into main 2026-09-29 00:33:13 +00:00
6 changed files with 312 additions and 1 deletions
@@ -99,6 +99,31 @@ composed, so it is not certified.
binding. The per-node source override becomes its reach, widened from the filter alone to the names
and the certificate as well.
## Progressive insight — 2026-09-29, from building it
**Reach does not mean the same thing to the filter for an endpoint the proxy serves.** The decision
above says `internal` means "the filter opens the machine port to the private network" and `public`
means "the filter opens it to anywhere". For a routed endpoint the second half is wrong, and
[ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) already said so before this
record was written: *a public service is exposed through the proxy, not by opening its own port* — it
listens `from: mesh`, only the proxy reaches it, and it is exposed by name.
Found by trying to express one real module, not by review. Its routed name must be public, because
browsers post to it; its machine-side port must not be, because that port serves the dashboard in
cleartext. Under one value driving both, saying "public" would have reopened a port an operator had
just closed. Measured the same evening: that module's routed name answered from the internet over TLS
while its machine-side port was refused from the same place. The port is not the path.
So the reach of a **routed** endpoint asks for names, and its port keeps what the manifest said. The
reach of an **unrouted** endpoint — git over ssh, a mail port, the bus — governs the port, because
there is no name and the port is the only way in. That is the same split this record already draws in
*an endpoint that is not routed is reached but never named*; what it got wrong was carrying the filter
across it.
This corrects a fact, not the decision: one statement per endpoint, three things derived from it and
none of them deciding on its own, all stand. The table in the decision should be read with the filter
column applying to an unrouted endpoint.
## Consequences
- **A manifest gains endpoint names, and a route contribution names an endpoint instead of a port.**
@@ -0,0 +1,134 @@
---
topic: what runs on it
status: accepted
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0010-delivery.md
---
# 143. A consumer verifies the grant it is given
## Context
A **grant** is what the mesh writes on a consumer's machine so it can reach a provider. The real one
the forge receives for its database, as it arrives:
```
provision postgres-database
at <the provider's machine, by name>
port the machine port the provider is published on
as the role the provider created for this consumer
```
with the credential sealed in a separate file. Four facts and a password, and they are the whole
mechanism by which anything in the mesh reaches anything else.
**The mesh asserts that claim and never finds out whether it is true.**
[Issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md):
converging a machine dropped the path from a container to a port on its own machine, and for eleven
hours the mesh answered *all doing what they were told, all heard from, every module current with its
source* while a web application logged, six thousand times:
```
connection to server at "<the machine>" (10.10.0.1), port 6852 failed: timeout expired
```
Every check the mesh makes passed, because every check it makes is about the relationship between the
mesh and a machine: the declaration was applied, the digest matched, every container named was running.
None of them asks whether a consumer can reach what it requires — though the mesh composed the grant
and therefore knows the consumer, the machine, the address, the port and the credential.
**And where the check runs decides whether it catches anything.** The rule in force admitted the
machines' own addresses on the private network. A dial from the *machine* to its own address carries
exactly such a source address, so a check run by the host on its own behalf would have matched that rule
and passed — while every container on the machine was refused. This is inference from the rule that was
loaded, not a measurement: the fault was found and fixed before anyone thought to dial from the host.
It is enough to decide the question, because a check whose position differs from the consumer's is
testing something nobody asked about.
## Considered Options
1. **The control plane dials each provision.** Rejected, and it is the tempting one because the control
plane holds every fact. It sits on the provider's machine for most provisions here and reaches the
address by a path no consumer uses; in the measured outage it would have passed throughout.
2. **The host dials on the consumer's behalf, from the machine.** Rejected for the reason above: the
machine's network position is not the consumer's, and the one outage this exists to catch is exactly
a difference between them.
3. **Ask the module.** Rejected: a module is arbitrary software that the mesh does not write. Some could
report on their provisions and most cannot, and a check that covers the modules that opted in tells
nobody anything about the rest.
4. **Read the module's logs.** Rejected: the failure was in a log the whole time, and reading a module's
logs makes the mesh depend on the wording of software it does not control.
5. **The consumer verifies it, from its own network position.** Adopted.
## Decision
**A consumer verifies each grant it is given, from its own network position.** After a reconcile has
applied a grant, the machine opens a connection to the address and port that grant names, from inside
the consumer's own network namespace — the same position the consumer's software dials from, which is
the only position that answers the question the grant asks.
**It is a connection, not a conversation.** Whether the port accepts a connection is what a grant
claims; whether the credential is right, the role exists or the schema is current is the provider's to
answer and the consumer's to discover. A check that spoke each provision's protocol would be a second
implementation of every provision, and would fail for reasons that are not the mesh's.
**One failure is not news.** A provider restarting is ordinary, and so is a consumer between containers.
A grant is reported unreachable only after it has failed on **consecutive** reconciles, and the count is
what the machine reports rather than the last attempt — so a reader can tell "it was briefly away" from
"it has never worked".
**A grant that cannot be checked is said to be unchecked, never assumed good.** A consumer that is not
running has no network position to dial from; that is not a broken grant and must not read as one. It is
also not a verified grant, and the two are different sentences.
**What it costs to be wrong is the constraint on all of it.** A check that reports a working provision
broken trains a reader to ignore the report, which is worse than having none — the fault this
repository keeps finding, one level up. So the threshold is consecutive failures, the check is the
cheapest thing that answers the question, and an unknown is reported as unknown.
**The mesh says it where it says everything else.** A machine's report carries its unreachable grants,
and `status` names them beside what is out of date — so "every module current with its source" stops
being the whole of what the mesh will tell you about a machine whose modules cannot reach each other.
## Consequences
- **The mesh can be wrong out loud.** It has been able to assert a grant and not check it; now a grant
that does not work is a thing the mesh says, and the eleven hours of issue 145 become minutes.
- **The host gains the ability to act from a container's network position**, which it has not needed
before. That is a real capability and the only one this needs.
- **A machine reports something that is not about the declaration.** Everything it reports today is
what it applied and what it holds; this is the first thing it says about whether what it applied
works.
- **A provision with no port is not checked**, because there is nothing to dial. Several are files and
secrets, and saying "checked" about those would be the appearance of verification that this record
exists to remove.
- **What got harder:** a reconcile does more than apply. Every grant adds a connection attempt on a
cadence, which is cheap individually and worth naming: a machine with many consumers dials once per
grant per reconcile.
## How it is checked
- **The outage is caught.** A bed drops the path from a consumer's network position to a provider's
port while leaving the machine's own path to it open — the exact shape of issue 145 — and the grant
reads unreachable. This fails against the previous behaviour, where nothing reported anything, and
against a check run from the machine, which passes while the consumer cannot reach it.
- **A restarting provider is not an outage.** One failed reconcile reports nothing; the count rises and
falls, and the grant reads reachable again without anybody acting.
- **A consumer that is not running reads unchecked, not broken**, asserted separately from the
unreachable case because they are different sentences.
- **A provision with no port is not claimed to be checked.**
- **The report carries the count, not the last attempt**, so "briefly away" and "never worked" are
distinguishable by a reader who sees only the report.
- **`status` names an unreachable grant**, asserted on the output, since a check nothing surfaces is
the same as no check.
## References
- [ADR 0010](0010-delivery.md) — the declaration is owned resources; a grant is one of them
- [ADR 0009](0009-modules-and-the-graph.md) — what a provision and a consumer are
- [issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
— the eleven hours
- [issue 136](../04-ISSUES/136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md) — the
same distance between a declaration and a machine, one level down
+1
View File
@@ -223,6 +223,7 @@ python3 00-META/checks/index.py fail if stale
- **0139** — [A network is forwarded because a module declared it](0139-a-network-is-forwarded-because-a-module-declared-it.md) *(superseded)*
- **0140** — [The filter constrains what arrives from outside, and says nothing about a machine's own guests](0140-the-filter-constrains-what-arrives-from-outside.md)
- **0141** — [The host delivers its own successor, and versions live side by side](0141-the-host-delivers-its-own-successor.md)
- **0143** — [A consumer verifies the grant it is given](0143-a-consumer-verifies-the-grant-it-is-given.md)
### How it is built
+37 -1
View File
@@ -5,8 +5,9 @@ code:
- mesh-controller internal/builder
- mesh-controller cmd/mesh-controller (build, build --behind, push, status)
- mesh-controller internal/inventory/builds.go
updated: 2026-09-21
updated: 2026-09-29
decisions:
- 02-DECISIONS/0143-a-consumer-verifies-the-grant-it-is-given.md
- 02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md
- 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md
- 02-DECISIONS/0010-delivery.md
@@ -226,3 +227,38 @@ when the current failure began and how many reports in a row have said it — th
id, whatever the words; three make the machine stuck, and `status` says so beside the failure. The host keeps trying — stuck is what the mesh
knows, not what the machine is told. *How it is checked:* an inventory test counts three identical
reports, a different one, and a clean apply; the status test asserts the word appears.
## A consumer verifies the grant it is given
*2026-09-29, from an outage that ran eleven hours —
[issue 145](../../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
settled by [ADR 0143](../../02-DECISIONS/0143-a-consumer-verifies-the-grant-it-is-given.md).*
A grant is four facts and a credential: the provision, the machine, the port, and who the consumer is
when it connects. It is the whole mechanism by which anything in the mesh reaches anything else, and the
mesh asserted it without ever finding out whether it was true.
What that cost: a machine was converged, the path from a container to a port on its own machine closed,
and for eleven hours the mesh answered *all heard from, every module current with its source* while a
module logged a connection timeout to its database six thousand times. Every check the mesh makes passed,
because every one is about the relationship between the mesh and a machine — applied, current, running.
None asks whether a consumer can reach what it requires.
**So a consumer verifies its own grants, from its own network position.** Not the control plane, which
reaches the address by a path no consumer uses; not the machine, whose own packets carried a source
address the filter admitted while every container's was refused. The position is the point: a check
somewhere else is testing something nobody asked about.
It opens a connection and nothing more. Whether the credential is right or the schema current is the
provider's to answer and the consumer's to discover; a check that spoke each provision's protocol would
be a second implementation of every provision.
One failure is not news — a provider restarting is ordinary — so a grant reads unreachable only after
consecutive reconciles, and the machine reports the count rather than the last attempt, which is what
separates "briefly away" from "never worked". A consumer that is not running has no position to dial
from: that grant reads unchecked, which is a different sentence from broken and must not be written as
one.
*How it is checked* is stated with the decision, and the first of them is the outage itself: a bed drops
the path from a consumer's position while leaving the machine's own open, and the grant must read
unreachable — which fails both against reporting nothing and against a check run from the machine.
@@ -57,6 +57,20 @@ enrolment path the design guarantees.
**The safe direction is not the same as the correct one.** Being more closed than intended broke nothing
visible, which is exactly why it went unnoticed for as long as the mesh has been on this machine.
## What it cost, measured later the same day
*2026-09-29.* The predecessor's chain was removed, and something it had been carrying went with it. It
admitted the private ranges wholesale, which is how a container on the machine reached a port declared
for the private network — the mesh's own filter admits the machines' overlay addresses, and a container
comes from a bridge. Every module that reached another by the machine's own name had been relying on the
predecessor's rule without anybody knowing.
That is [issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
and it ran for eleven hours while the mesh reported the machine healthy. The filter is fixed. What this
adds to the account here is that "the machine is more closed than the mesh believes" was not the
harmless direction after all — it was harmless for everything reached from outside, and an outage for
everything reached from within.
## Open questions
- Should the host report every place the machine filters from, rather than one kind — the front-end,
@@ -0,0 +1,101 @@
---
status: located
opened: 2026-09-29
located-in:
- mesh-controller internal/catalogue/filtering.go (fixed for this instance)
- mesh-controller (what status reports, and what it does not ask)
fixed-by:
amended-design: 03-DESIGN/01-to-be/10-delivery.md
---
# 145 — A machine reads healthy while its modules cannot reach each other
## What was observed
Converging the control-node closed every path by which a module on that machine reached another module
by the machine's own name. It ran for **eleven hours**. Throughout, the mesh answered:
```
4 machine(s), all doing what they were told, all heard from,
running what the mesh would send them, and every module current with its source
```
What was actually happening, from one affected module's own log:
```
Doctrine\DBAL\Exception: Failed to connect to the database:
SQLSTATE[08006] connection to server at "novox.internal" (10.10.0.1), port 6852 failed: timeout expired
```
6,154 of them, beginning at the minute of the flip. The web application accepted TCP connections and
never answered an HTTP request; a client waited 35 seconds and gave up. Confirmed from a throwaway
container on the machine: neither the store nor the forge was reachable on the machine's own address.
The cause is [issue 144](../144-the-predecessors-rules-outlive-the-firewall-it-was-found-as/00-report.md)'s
sibling and is fixed: a port declared reachable from the private network admitted the machines' own
overlay addresses, and a container on the machine comes from a bridge address, matching none of them.
What this issue is about is the eleven hours.
**Nothing the mesh reports would have shown it.** Every check the mesh makes passed, because every
check the mesh makes is about the relationship between the mesh and a machine:
- the machine applied what it was sent, and said so;
- its declaration digest matches what the mesh would send;
- every module's source commit matches what the mesh holds;
- every container the declaration names is running.
None of those asks whether a module can reach what it requires. The mesh knows precisely who requires
what — it composes the grants — and never checks that the grant works.
**Nor would an operator's usual look.** The ports were probed from outside and behaved correctly; the
routed services answered; a container's egress to the internet worked. Those are the paths a person
checks after changing a firewall, and all three were fine. The broken path was module-to-module over
the machine's own name, which nothing routine exercises.
## Why it matters beyond this instance
**A mesh that composes a dependency and never tests it can only report on itself.** Every provision the
mesh grants is a claim that a consumer can reach a provider. The mesh asserts that claim, delivers
credentials for it, and has no mechanism that ever finds out. "Every module current with its source"
is a statement about bytes, not about whether anything works.
**The failure was silent in the direction that hides longest.** A service that will not start is
noticed. A service that starts, accepts connections and then cannot reach its database serves errors
under a healthy-looking process, and the machine's own report says the container is running — which it
is.
**It is the same shape as [issue 136](../136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md),
one level up.** There, a module named a program the machine lacked and everything reported success.
Here, the mesh granted a provision the filter refused and everything reported success. Both are the
distance between a declaration and the machine, and in both cases the report was about the declaration.
**And the eleven hours are the measurement, not the bug.** The filter fault was one line and is fixed.
What is not fixed is that nothing in the mesh would have told anybody.
## What was decided
*2026-09-29, the same day.* The first open question below — should a grant be checked, and from where —
is answered by [ADR 0143](../../02-DECISIONS/0143-a-consumer-verifies-the-grant-it-is-given.md): the
consumer verifies it, from its own network position, because that is the only position that answers what
a grant claims. A check run by the control plane or by the machine would have passed throughout this
outage, since the rule in force admitted the machines' own addresses and it was the containers that were
refused.
The remaining questions below stand, and the record answers two of them: one failure is not news, only
consecutive ones, and a grant that cannot be checked is reported unchecked rather than assumed good.
## Open questions
- Should a grant be checked? The mesh knows the consumer, the provider, the address and the port, so a
reachability check is expressible — but from where: the consumer's machine, as part of a reconcile,
or the provider's?
- What would it cost to be wrong in the other direction? A check that reports a provision broken while
it works is worse than none, because it trains a reader to ignore the report. A provider restarting is
ordinary; a consumer between containers is ordinary.
- What should `status` say about a machine whose modules cannot reach each other? It currently has one
vocabulary for "heard from and current", and that sentence was true the whole time.
- Is there a cheaper signal than a probe? The affected module was logging the failure 6,154 times. The
mesh reads no module's logs and arguably should not — but something a module could *say* about its
own provisions would have surfaced this in minutes.
- Does the same blindness apply to the other direction — a provider that lost a consumer's grant and
is refusing it? Nothing checks that either.