Compare commits

...
Author SHA1 Message Date
jschoubben eba24a72af ADR 0144: anything on a machine may call anything on it, superseding 0143
Everything should be able to call what runs on the same machine, another machine's
service exposed to the private network, and another machine's service exposed
publicly. Three cases; the filter had two.

The first was expressed as the machines' own addresses on the private network. A
caller on the machine carries such an address; a caller in one of its containers
carries a bridge address and matched nothing — measured, same destination and same
machine: src 10.10.0.1 against src 172.17.0.8. The second case worked by accident,
because the tunnel rewrites a caller's address to the sending machine's. Two of
three working is why it read as correct.

0143 answered the wrong question. It proposed verifying each grant from the
consumer's own network position and went to length about which position, because
whether a caller sat in a container changed the answer — and that difference was the
bug. Observing a configuration error is not its remedy. Superseded, and nothing
replaces it; whether the mesh should check a grant is still open in issue 145 and
must stand on its own.

And a module is not a container: 61 of 72 happen to use one, 11 do not, and a rule
reasoning about containers describes most of the mesh rather than the mesh.
2026-09-29 13:32:01 +02:00
mesh-admin 78d4873f4f Merge pull request 'ADR 0143: a consumer verifies the grant it is given' (#180) from decision/0143-a-consumer-verifies-its-grant into main 2026-09-29 11:07:20 +00:00
jschoubben ddfd62edf6 ADR 0143: a consumer verifies the grant it is given
A grant is four facts and a credential — the provision, the machine, the port, and
who the consumer is when it connects — and it is the whole mechanism by which
anything in the mesh reaches anything else. The mesh asserted it and never found
out whether it was true. Issue 145 is what that cost: eleven hours of 'every module
current with its source' while a module could not reach its database.

The consumer verifies it, from its own network position. Not the control plane,
which reaches the address by a path no consumer uses. Not the machine, whose own
packets carried a source address the filter admitted while every container's was
refused — so a host-side check would have passed throughout the outage it exists to
catch. That is inference from the rule that was loaded, stated as such.

A connection and nothing more; speaking each provision's protocol would be a second
implementation of every provision. One failure is not news, only consecutive ones,
and the count is reported rather than the last attempt. A consumer that is not
running reads unchecked, which is a different sentence from broken.

What it costs to be wrong is the constraint on all of it: a check that calls a
working provision broken trains a reader to ignore the report.
2026-09-29 13:07:18 +02:00
mesh-admin f65664640a Merge pull request 'Issue 145: a machine reads healthy while its modules cannot reach each other' (#179) from issue/145-healthy-while-broken into main 2026-09-29 10:37:10 +00:00
jschoubben 492ac7be18 Issue 145: a machine reads healthy while its modules cannot reach each other
Converging the control node closed every path by which a module reached another by
the machine's own name, and it ran for eleven hours while the mesh answered 'all
doing what they were told, all heard from, running what the mesh would send them,
and every module current with its source'.

6,154 database failures in one module's log, beginning at the minute of the flip.
The service accepted TCP and never answered HTTP. Every check the mesh makes
passed, because every check the mesh makes is about the relationship between the
mesh and a machine — applied, current, containers running. None asks whether a
module can reach what it requires, though the mesh composes every grant and so
knows exactly who requires what.

The filter fault is fixed. The eleven hours are the measurement, not the bug.

Also records on 144 that 'more closed than the mesh believes' was harmless only
for what is reached from outside, and an outage for what is reached from within.
2026-09-29 12:37:08 +02:00
mesh-admin 5ac77e3cef Merge pull request 'ADR 0138: reach asks for names on a routed endpoint' (#178) from decision/0138-insight-reach-and-the-proxy into main 2026-09-29 00:50:28 +00:00
jschoubben a619022c35 ADR 0138: a progressive insight — reach asks for names on a routed endpoint
The record says internal and public each mean something to the filter. For an
endpoint the proxy serves, the second half is wrong, and ADR 0045 said so first:
a public service is exposed through the proxy, listening from the mesh, not by
opening its own port.

Found by trying to express one real module, not by review — routed name public
because browsers post to it, machine port private because it serves a dashboard in
cleartext. Under one value for both, saying public would have reopened a port an
operator had just closed. Measured the same evening: the routed name answered from
the internet over TLS while the port was refused from the same place.

Corrects a fact. One statement per endpoint with three things derived from it
stands; the filter column applies to an unrouted endpoint.
2026-09-29 02:50:25 +02:00
mesh-admin 49c065c204 Merge pull request 'Issue 143: correct the diagnosis' (#177) from issue/143-corrected-diagnosis into main 2026-09-29 00:33:13 +00:00
7 changed files with 440 additions and 1 deletions
@@ -99,6 +99,31 @@ composed, so it is not certified.
binding. The per-node source override becomes its reach, widened from the filter alone to the names
and the certificate as well.
## Progressive insight — 2026-09-29, from building it
**Reach does not mean the same thing to the filter for an endpoint the proxy serves.** The decision
above says `internal` means "the filter opens the machine port to the private network" and `public`
means "the filter opens it to anywhere". For a routed endpoint the second half is wrong, and
[ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) already said so before this
record was written: *a public service is exposed through the proxy, not by opening its own port* — it
listens `from: mesh`, only the proxy reaches it, and it is exposed by name.
Found by trying to express one real module, not by review. Its routed name must be public, because
browsers post to it; its machine-side port must not be, because that port serves the dashboard in
cleartext. Under one value driving both, saying "public" would have reopened a port an operator had
just closed. Measured the same evening: that module's routed name answered from the internet over TLS
while its machine-side port was refused from the same place. The port is not the path.
So the reach of a **routed** endpoint asks for names, and its port keeps what the manifest said. The
reach of an **unrouted** endpoint — git over ssh, a mail port, the bus — governs the port, because
there is no name and the port is the only way in. That is the same split this record already draws in
*an endpoint that is not routed is reached but never named*; what it got wrong was carrying the filter
across it.
This corrects a fact, not the decision: one statement per endpoint, three things derived from it and
none of them deciding on its own, all stand. The table in the decision should be read with the filter
column applying to an unrouted endpoint.
## Consequences
- **A manifest gains endpoint names, and a route contribution names an endpoint instead of a port.**
@@ -0,0 +1,135 @@
---
topic: what runs on it
status: superseded
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0010-delivery.md
superseded-by: 02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md
---
# 143. A consumer verifies the grant it is given
## Context
A **grant** is what the mesh writes on a consumer's machine so it can reach a provider. The real one
the forge receives for its database, as it arrives:
```
provision postgres-database
at <the provider's machine, by name>
port the machine port the provider is published on
as the role the provider created for this consumer
```
with the credential sealed in a separate file. Four facts and a password, and they are the whole
mechanism by which anything in the mesh reaches anything else.
**The mesh asserts that claim and never finds out whether it is true.**
[Issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md):
converging a machine dropped the path from a container to a port on its own machine, and for eleven
hours the mesh answered *all doing what they were told, all heard from, every module current with its
source* while a web application logged, six thousand times:
```
connection to server at "<the machine>" (10.10.0.1), port 6852 failed: timeout expired
```
Every check the mesh makes passed, because every check it makes is about the relationship between the
mesh and a machine: the declaration was applied, the digest matched, every container named was running.
None of them asks whether a consumer can reach what it requires — though the mesh composed the grant
and therefore knows the consumer, the machine, the address, the port and the credential.
**And where the check runs decides whether it catches anything.** The rule in force admitted the
machines' own addresses on the private network. A dial from the *machine* to its own address carries
exactly such a source address, so a check run by the host on its own behalf would have matched that rule
and passed — while every container on the machine was refused. This is inference from the rule that was
loaded, not a measurement: the fault was found and fixed before anyone thought to dial from the host.
It is enough to decide the question, because a check whose position differs from the consumer's is
testing something nobody asked about.
## Considered Options
1. **The control plane dials each provision.** Rejected, and it is the tempting one because the control
plane holds every fact. It sits on the provider's machine for most provisions here and reaches the
address by a path no consumer uses; in the measured outage it would have passed throughout.
2. **The host dials on the consumer's behalf, from the machine.** Rejected for the reason above: the
machine's network position is not the consumer's, and the one outage this exists to catch is exactly
a difference between them.
3. **Ask the module.** Rejected: a module is arbitrary software that the mesh does not write. Some could
report on their provisions and most cannot, and a check that covers the modules that opted in tells
nobody anything about the rest.
4. **Read the module's logs.** Rejected: the failure was in a log the whole time, and reading a module's
logs makes the mesh depend on the wording of software it does not control.
5. **The consumer verifies it, from its own network position.** Adopted.
## Decision
**A consumer verifies each grant it is given, from its own network position.** After a reconcile has
applied a grant, the machine opens a connection to the address and port that grant names, from inside
the consumer's own network namespace — the same position the consumer's software dials from, which is
the only position that answers the question the grant asks.
**It is a connection, not a conversation.** Whether the port accepts a connection is what a grant
claims; whether the credential is right, the role exists or the schema is current is the provider's to
answer and the consumer's to discover. A check that spoke each provision's protocol would be a second
implementation of every provision, and would fail for reasons that are not the mesh's.
**One failure is not news.** A provider restarting is ordinary, and so is a consumer between containers.
A grant is reported unreachable only after it has failed on **consecutive** reconciles, and the count is
what the machine reports rather than the last attempt — so a reader can tell "it was briefly away" from
"it has never worked".
**A grant that cannot be checked is said to be unchecked, never assumed good.** A consumer that is not
running has no network position to dial from; that is not a broken grant and must not read as one. It is
also not a verified grant, and the two are different sentences.
**What it costs to be wrong is the constraint on all of it.** A check that reports a working provision
broken trains a reader to ignore the report, which is worse than having none — the fault this
repository keeps finding, one level up. So the threshold is consecutive failures, the check is the
cheapest thing that answers the question, and an unknown is reported as unknown.
**The mesh says it where it says everything else.** A machine's report carries its unreachable grants,
and `status` names them beside what is out of date — so "every module current with its source" stops
being the whole of what the mesh will tell you about a machine whose modules cannot reach each other.
## Consequences
- **The mesh can be wrong out loud.** It has been able to assert a grant and not check it; now a grant
that does not work is a thing the mesh says, and the eleven hours of issue 145 become minutes.
- **The host gains the ability to act from a container's network position**, which it has not needed
before. That is a real capability and the only one this needs.
- **A machine reports something that is not about the declaration.** Everything it reports today is
what it applied and what it holds; this is the first thing it says about whether what it applied
works.
- **A provision with no port is not checked**, because there is nothing to dial. Several are files and
secrets, and saying "checked" about those would be the appearance of verification that this record
exists to remove.
- **What got harder:** a reconcile does more than apply. Every grant adds a connection attempt on a
cadence, which is cheap individually and worth naming: a machine with many consumers dials once per
grant per reconcile.
## How it is checked
- **The outage is caught.** A bed drops the path from a consumer's network position to a provider's
port while leaving the machine's own path to it open — the exact shape of issue 145 — and the grant
reads unreachable. This fails against the previous behaviour, where nothing reported anything, and
against a check run from the machine, which passes while the consumer cannot reach it.
- **A restarting provider is not an outage.** One failed reconcile reports nothing; the count rises and
falls, and the grant reads reachable again without anybody acting.
- **A consumer that is not running reads unchecked, not broken**, asserted separately from the
unreachable case because they are different sentences.
- **A provision with no port is not claimed to be checked.**
- **The report carries the count, not the last attempt**, so "briefly away" and "never worked" are
distinguishable by a reader who sees only the report.
- **`status` names an unreachable grant**, asserted on the output, since a check nothing surfaces is
the same as no check.
## References
- [ADR 0010](0010-delivery.md) — the declaration is owned resources; a grant is one of them
- [ADR 0009](0009-modules-and-the-graph.md) — what a provision and a consumer are
- [issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
— the eleven hours
- [issue 136](../04-ISSUES/136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md) — the
same distance between a declaration and a machine, one level down
@@ -0,0 +1,121 @@
---
topic: what runs on it
status: accepted
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md
supersedes: 02-DECISIONS/0143-a-consumer-verifies-the-grant-it-is-given.md
---
# 144. Anything on a machine may call anything on it, and that is the whole of "local"
## Context
Everything in the mesh should be able to call:
- what runs on the same machine;
- another machine's service over the private network, if that service is exposed there;
- another machine's service over the public network, if it is exposed there.
Three cases. The filter had two of them.
**The first was broken and the break was invisible.** A service exposed to the private network rendered
as the machines' own addresses on it. A caller on the machine carries such an address; a caller inside
one of that machine's containers carries a bridge address and matched nothing. Measured:
```
the machine: local 10.10.0.1 dev lo src 10.10.0.1
a container: 10.10.0.1 via 172.17.0.1 dev eth0 src 172.17.0.8
```
Same destination, same machine, two source addresses. The rule named the first and silently refused the
second, so a module reaching its database on its own machine's name timed out for eleven hours
([issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)).
**The second case works, and by accident.** A caller on another machine reaches the private network over
the tunnel, and arrives carrying that machine's own address — so the rule matches. It would not have
matched the caller's own address either; the tunnel rewrites it. That two of three cases worked is why
this looked correct.
**[ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) answered the wrong question.** Written
hours earlier, it proposed that a consumer verify each grant it is given by opening a connection from
its own network position — and it went to some length about *which* position, because whether a caller
sat in a container changed the answer. That difference was the bug. A verification mechanism would have
reported this outage sooner and would not have prevented it, and the machinery it needed existed only
because the rule was wrong. The remedy for a configuration error is the correct configuration.
**And a module is not a container.** A module is software that delivers one or more services, and it may
do that as a container, an installed package with a unit, a binary, or files something else reads. Of 72
modules in the catalogue, 61 happen to use a container and 11 do not — among them the resolver, the ssh
daemon and the intrusion-prevention module. A rule that reasons about containers describes most of the
mesh and not the mesh.
## Considered Options
1. **A line per service admitting the machine's own callers.** Rejected: it is what was written first,
and it only ever covers the services somebody remembered to think about. It also states, service by
service, a thing that is true of the machine.
2. **Verify each grant from the consumer's position** ([ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md)).
Rejected as a remedy: it observes the fault rather than removing it, and the question it agonised over
— which network position — exists only while the fault does.
3. **Enumerate the addresses a machine's callers may have.** Rejected for the reason no address is named
anywhere in this filter any more ([ADR 0140](0140-the-filter-constrains-what-arrives-from-outside.md)):
a range describes one machine and goes stale in silence.
4. **Local is not filtered, stated once.** Adopted.
## Decision
**Anything on a machine may call anything on that machine, and the filter says so once.** Not per
service, not per port, and not by naming who the callers are: traffic that did not arrive from outside
the machine and did not arrive over the private network is the machine's own, and is admitted. It is
asked by the link the traffic arrived on, because that is a fact about the machine rather than a list
that describes one.
**Local is not a boundary this mesh draws.** Whether a caller is a container, a unit, or the operator's
shell changes nothing, because the thing being decided is "is this the same machine" and the answer does
not depend on the form the caller takes.
**The other two cases are unchanged and are now legible beside it.** A service exposed to the private
network admits the machines on it; a service exposed publicly admits anything. Three cases, three lines,
and a reader can see all three at once.
**[ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) is superseded and nothing replaces it.**
Whether the mesh should check that a grant works is a real question — it reported this machine healthy
for eleven hours — but it is a question about what the mesh can say, not about what it should do, and it
must stand on its own rather than as the remedy for a rule that was wrong. It is not built.
## Consequences
- **The three things everything should be able to call are three lines**, and the first is one line
rather than one per service, so a service added tomorrow is reachable locally without anybody
remembering to say so.
- **A form of module stops mattering to the filter.** The 11 modules that are not containers were never
affected by this bug and were never the reason it was hard to see; they are the reason the rule should
never have mentioned containers.
- **The mesh still cannot say when a grant stops working.** That is the live gap, recorded in issue 145
and no longer pretending to have an answer.
- **What got harder:** nothing. This removes a line per service and replaces it with one.
## How it is checked
- **A caller on the machine reaches a service on it, in the input chain**, asserted on that chain's own
body — because the forward chain carries the same line in the same words, and an assertion on the
whole rendered file passed with the input chain's copy deleted. That is what
[ADR 0137](0137-a-machine-says-which-networks-it-routes.md)'s tests already say to do.
- **It is one rule, not one per service.** Asserted by rendering two services of different reach and
refusing a per-port local line.
- **The three reaches render as three lines**, asserted together, so the whole of what the filter says
about who may call what is one test.
- **The measured case:** from a container on the machine, a service exposed to the private network on
that machine answers. This is the outage, and it fails against the rule this replaces.
## References
- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the filter is the sum
of what its modules listen on
- [ADR 0140](0140-the-filter-constrains-what-arrives-from-outside.md) — why no address is named
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — internal and public,
the other two cases
- [ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) — superseded here
- [issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
+2
View File
@@ -223,6 +223,8 @@ python3 00-META/checks/index.py fail if stale
- **0139** — [A network is forwarded because a module declared it](0139-a-network-is-forwarded-because-a-module-declared-it.md) *(superseded)*
- **0140** — [The filter constrains what arrives from outside, and says nothing about a machine's own guests](0140-the-filter-constrains-what-arrives-from-outside.md)
- **0141** — [The host delivers its own successor, and versions live side by side](0141-the-host-delivers-its-own-successor.md)
- **0143** — [A consumer verifies the grant it is given](0143-a-consumer-verifies-the-grant-it-is-given.md) *(superseded)*
- **0144** — [Anything on a machine may call anything on it, and that is the whole of "local"](0144-anything-on-a-machine-may-call-anything-on-it.md)
### How it is built
+32 -1
View File
@@ -5,7 +5,7 @@ code:
- mesh-controller internal/builder
- mesh-controller cmd/mesh-controller (build, build --behind, push, status)
- mesh-controller internal/inventory/builds.go
updated: 2026-09-21
updated: 2026-09-29
decisions:
- 02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md
- 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md
@@ -226,3 +226,34 @@ when the current failure began and how many reports in a row have said it — th
id, whatever the words; three make the machine stuck, and `status` says so beside the failure. The host keeps trying — stuck is what the mesh
knows, not what the machine is told. *How it is checked:* an inventory test counts three identical
reports, a different one, and a clean apply; the status test asserts the word appears.
## Everything may call what is exposed to it, and local is not a boundary
*2026-09-29, from an outage that ran eleven hours —
[issue 145](../../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
settled by [ADR 0144](../../02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md).*
A grant is four facts and a credential: the provision, the machine, the port, and who the consumer is
when it connects. It is the whole mechanism by which anything in the mesh reaches anything else, and it
rests on three things being callable — what runs on the same machine, another machine's service over the
private network where it is exposed there, and another machine's service over the public network where it
is exposed there.
The filter had two of those. A service exposed to the private network admitted the machines' own addresses
on it; a caller on the machine carries such an address, and a caller inside one of that machine's
containers carries a bridge address and matched nothing. Measured, same destination and same machine:
`src 10.10.0.1` from the machine, `src 172.17.0.8` from a container on it. So a module reaching its
database on its own machine's name timed out for eleven hours while the mesh called the machine healthy.
The second case worked by accident: a caller on another machine arrives over the tunnel carrying that
machine's address, which the rule matched. Two of three working is why this read as correct.
**So local is not a boundary this mesh draws, and the filter says so once.** Traffic that did not arrive
from outside the machine and did not arrive over the private network is the machine's own, and is
admitted — for every service there, not per service. Whether the caller is a container, a unit or a shell
decides nothing, because the question is "is this the same machine".
A verification mechanism was drafted for this and withdrawn. It would have reported the outage sooner and
would not have prevented it, and the part of it that was hard — deciding which network position to check
from — existed only because the rule was wrong. Whether the mesh should check that a grant works is still
open, in issue 145; it is not the remedy for a configuration error.
@@ -57,6 +57,20 @@ enrolment path the design guarantees.
**The safe direction is not the same as the correct one.** Being more closed than intended broke nothing
visible, which is exactly why it went unnoticed for as long as the mesh has been on this machine.
## What it cost, measured later the same day
*2026-09-29.* The predecessor's chain was removed, and something it had been carrying went with it. It
admitted the private ranges wholesale, which is how a container on the machine reached a port declared
for the private network — the mesh's own filter admits the machines' overlay addresses, and a container
comes from a bridge. Every module that reached another by the machine's own name had been relying on the
predecessor's rule without anybody knowing.
That is [issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
and it ran for eleven hours while the mesh reported the machine healthy. The filter is fixed. What this
adds to the account here is that "the machine is more closed than the mesh believes" was not the
harmless direction after all — it was harmless for everything reached from outside, and an outage for
everything reached from within.
## Open questions
- Should the host report every place the machine filters from, rather than one kind — the front-end,
@@ -0,0 +1,111 @@
---
status: located
opened: 2026-09-29
located-in:
- mesh-controller internal/catalogue/filtering.go (fixed for this instance)
- mesh-controller (what status reports, and what it does not ask)
fixed-by:
amended-design: 03-DESIGN/01-to-be/10-delivery.md
---
# 145 — A machine reads healthy while its modules cannot reach each other
## What was observed
Converging the control-node closed every path by which a module on that machine reached another module
by the machine's own name. It ran for **eleven hours**. Throughout, the mesh answered:
```
4 machine(s), all doing what they were told, all heard from,
running what the mesh would send them, and every module current with its source
```
What was actually happening, from one affected module's own log:
```
Doctrine\DBAL\Exception: Failed to connect to the database:
SQLSTATE[08006] connection to server at "novox.internal" (10.10.0.1), port 6852 failed: timeout expired
```
6,154 of them, beginning at the minute of the flip. The web application accepted TCP connections and
never answered an HTTP request; a client waited 35 seconds and gave up. Confirmed from a throwaway
container on the machine: neither the store nor the forge was reachable on the machine's own address.
The cause is [issue 144](../144-the-predecessors-rules-outlive-the-firewall-it-was-found-as/00-report.md)'s
sibling and is fixed: a port declared reachable from the private network admitted the machines' own
overlay addresses, and a container on the machine comes from a bridge address, matching none of them.
What this issue is about is the eleven hours.
**Nothing the mesh reports would have shown it.** Every check the mesh makes passed, because every
check the mesh makes is about the relationship between the mesh and a machine:
- the machine applied what it was sent, and said so;
- its declaration digest matches what the mesh would send;
- every module's source commit matches what the mesh holds;
- every container the declaration names is running.
None of those asks whether a module can reach what it requires. The mesh knows precisely who requires
what — it composes the grants — and never checks that the grant works.
**Nor would an operator's usual look.** The ports were probed from outside and behaved correctly; the
routed services answered; a container's egress to the internet worked. Those are the paths a person
checks after changing a firewall, and all three were fine. The broken path was module-to-module over
the machine's own name, which nothing routine exercises.
## Why it matters beyond this instance
**A mesh that composes a dependency and never tests it can only report on itself.** Every provision the
mesh grants is a claim that a consumer can reach a provider. The mesh asserts that claim, delivers
credentials for it, and has no mechanism that ever finds out. "Every module current with its source"
is a statement about bytes, not about whether anything works.
**The failure was silent in the direction that hides longest.** A service that will not start is
noticed. A service that starts, accepts connections and then cannot reach its database serves errors
under a healthy-looking process, and the machine's own report says the container is running — which it
is.
**It is the same shape as [issue 136](../136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md),
one level up.** There, a module named a program the machine lacked and everything reported success.
Here, the mesh granted a provision the filter refused and everything reported success. Both are the
distance between a declaration and the machine, and in both cases the report was about the declaration.
**And the eleven hours are the measurement, not the bug.** The filter fault was one line and is fixed.
What is not fixed is that nothing in the mesh would have told anybody.
## What was decided
*2026-09-29, the same day, in two steps and the first was wrong.*
The first answer was [ADR 0143](../../02-DECISIONS/0143-a-consumer-verifies-the-grant-it-is-given.md):
the consumer verifies each grant from its own network position, because whether a caller sat in a
container changed whether it could reach the provider. **That difference was the fault**, and the record
is superseded. A verification mechanism would have reported this sooner and would not have prevented it,
and the part of it that was difficult — deciding which network position to check from — existed only
while the rule was wrong.
The remedy is [ADR 0144](../../02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md):
anything on a machine may call anything on it, said once rather than per service, and asked by the link
traffic arrives on rather than the address it carries. Everything should be able to call what runs on the
same machine, another machine's service exposed to the private network, and another machine's service
exposed publicly. The filter had the second and third and expressed the first as a list of addresses that
no container could match.
**The question this issue is actually about is still open.** Nothing in the mesh would have said a grant
had stopped working, and nothing does now. That is not answered by a configuration fix, and it should not
be answered by a mechanism adopted as a remedy for one.
## Open questions
- Should a grant be checked? The mesh knows the consumer, the provider, the address and the port, so a
reachability check is expressible — but from where: the consumer's machine, as part of a reconcile,
or the provider's?
- What would it cost to be wrong in the other direction? A check that reports a provision broken while
it works is worse than none, because it trains a reader to ignore the report. A provider restarting is
ordinary; a consumer between containers is ordinary.
- What should `status` say about a machine whose modules cannot reach each other? It currently has one
vocabulary for "heard from and current", and that sentence was true the whole time.
- Is there a cheaper signal than a probe? The affected module was logging the failure 6,154 times. The
mesh reads no module's logs and arguably should not — but something a module could *say* about its
own provisions would have surfaced this in minutes.
- Does the same blindness apply to the other direction — a provider that lost a consumer's grant and
is refusing it? Nothing checks that either.