Compare commits

..
Author SHA1 Message Date
jschoubben 38482435af ADR 0146: connectivity is checked by name, per hosting form, with a valid certificate
0145's core stands — a module checks what the mesh claims, from where the callers
are — and everything it said about what to dial was wrong.

A raw port is not how anything in this mesh is reached. A real caller resolves a
name, the proxy answers, the proxy reaches the service; dialling a port tests the
last hop of a four-hop path and skips the three where most connectivity lives.

So each hosting form gets its own endpoint, route and name:
connect-docker.<node>.internal for a container, connect-process for the mesh's own
code in a unit it writes, connect-unit for a unit a package ships — and the same
under each public domain. Every machine checks every machine, by name, over TLS,
verifying the certificate against the authority that should have issued it. The
names are the instrument: a failure reads as connect-docker.g14.internal did not
answer.

No name is written anywhere: the machines come from the roster fact, the labels are
the module's, and a machine that joins is checked by the others on the next push.

Five things this needs that the mesh does not have, named rather than assumed —
a module every machine has (ScopeNode is exclusivity, not obligation), a container
running a bundle, a unit a package ships, the public domain in the roster fact, and
issue 129, until which every internal name will correctly fail verification.
2026-09-29 14:01:55 +02:00
mesh-admin cf8a8d78c9 Merge pull request 'ADR 0145: a module checks what the mesh claims is reachable' (#182) from decision/0145-a-module-checks-what-the-mesh-claims into main 2026-09-29 11:44:58 +00:00
jschoubben 9de25994e9 ADR 0145: a module checks what the mesh claims is reachable, and it checks itself
The mesh asserts three things are callable and has never checked any of them. A
module on every machine serves an endpoint of its own and dials every other
machine's, from the position the callers are in — not the host and not the control
plane, both of which reach those addresses by paths no ordinary caller uses and
would have passed throughout the outage.

Its probe is its own endpoint, declared reachable over the private network, so it
is admitted by exactly the rule that governs every internally-exposed service. The
tempting target is a service every machine has, and those are the ones never closed
— ssh above all — which would have passed for all eleven hours.

Resolution and connection reported separately, because they have different owners.
One failure is not a fault, and the count travels with the result. It reports and
repairs nothing.

Built on a scheduled container, a rendered roster fact and a bundle: no new
vocabulary. Adopted on its own merits, which is what 0143 was not.
2026-09-29 13:44:56 +02:00
mesh-admin 64ea47b11d Merge pull request 'ADR 0144: anything on a machine may call anything on it' (#181) from decision/0144-local-is-not-a-boundary into main 2026-09-29 11:32:03 +00:00
jschoubben eba24a72af ADR 0144: anything on a machine may call anything on it, superseding 0143
Everything should be able to call what runs on the same machine, another machine's
service exposed to the private network, and another machine's service exposed
publicly. Three cases; the filter had two.

The first was expressed as the machines' own addresses on the private network. A
caller on the machine carries such an address; a caller in one of its containers
carries a bridge address and matched nothing — measured, same destination and same
machine: src 10.10.0.1 against src 172.17.0.8. The second case worked by accident,
because the tunnel rewrites a caller's address to the sending machine's. Two of
three working is why it read as correct.

0143 answered the wrong question. It proposed verifying each grant from the
consumer's own network position and went to length about which position, because
whether a caller sat in a container changed the answer — and that difference was the
bug. Observing a configuration error is not its remedy. Superseded, and nothing
replaces it; whether the mesh should check a grant is still open in issue 145 and
must stand on its own.

And a module is not a container: 61 of 72 happen to use one, 11 do not, and a rule
reasoning about containers describes most of the mesh rather than the mesh.
2026-09-29 13:32:01 +02:00
mesh-admin 78d4873f4f Merge pull request 'ADR 0143: a consumer verifies the grant it is given' (#180) from decision/0143-a-consumer-verifies-its-grant into main 2026-09-29 11:07:20 +00:00
jschoubben ddfd62edf6 ADR 0143: a consumer verifies the grant it is given
A grant is four facts and a credential — the provision, the machine, the port, and
who the consumer is when it connects — and it is the whole mechanism by which
anything in the mesh reaches anything else. The mesh asserted it and never found
out whether it was true. Issue 145 is what that cost: eleven hours of 'every module
current with its source' while a module could not reach its database.

The consumer verifies it, from its own network position. Not the control plane,
which reaches the address by a path no consumer uses. Not the machine, whose own
packets carried a source address the filter admitted while every container's was
refused — so a host-side check would have passed throughout the outage it exists to
catch. That is inference from the rule that was loaded, stated as such.

A connection and nothing more; speaking each provision's protocol would be a second
implementation of every provision. One failure is not news, only consecutive ones,
and the count is reported rather than the last attempt. A consumer that is not
running reads unchecked, which is a different sentence from broken.

What it costs to be wrong is the constraint on all of it: a check that calls a
working provision broken trains a reader to ignore the report.
2026-09-29 13:07:18 +02:00
mesh-admin f65664640a Merge pull request 'Issue 145: a machine reads healthy while its modules cannot reach each other' (#179) from issue/145-healthy-while-broken into main 2026-09-29 10:37:10 +00:00
7 changed files with 567 additions and 2 deletions
@@ -0,0 +1,135 @@
---
topic: what runs on it
status: superseded
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0010-delivery.md
superseded-by: 02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md
---
# 143. A consumer verifies the grant it is given
## Context
A **grant** is what the mesh writes on a consumer's machine so it can reach a provider. The real one
the forge receives for its database, as it arrives:
```
provision postgres-database
at <the provider's machine, by name>
port the machine port the provider is published on
as the role the provider created for this consumer
```
with the credential sealed in a separate file. Four facts and a password, and they are the whole
mechanism by which anything in the mesh reaches anything else.
**The mesh asserts that claim and never finds out whether it is true.**
[Issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md):
converging a machine dropped the path from a container to a port on its own machine, and for eleven
hours the mesh answered *all doing what they were told, all heard from, every module current with its
source* while a web application logged, six thousand times:
```
connection to server at "<the machine>" (10.10.0.1), port 6852 failed: timeout expired
```
Every check the mesh makes passed, because every check it makes is about the relationship between the
mesh and a machine: the declaration was applied, the digest matched, every container named was running.
None of them asks whether a consumer can reach what it requires — though the mesh composed the grant
and therefore knows the consumer, the machine, the address, the port and the credential.
**And where the check runs decides whether it catches anything.** The rule in force admitted the
machines' own addresses on the private network. A dial from the *machine* to its own address carries
exactly such a source address, so a check run by the host on its own behalf would have matched that rule
and passed — while every container on the machine was refused. This is inference from the rule that was
loaded, not a measurement: the fault was found and fixed before anyone thought to dial from the host.
It is enough to decide the question, because a check whose position differs from the consumer's is
testing something nobody asked about.
## Considered Options
1. **The control plane dials each provision.** Rejected, and it is the tempting one because the control
plane holds every fact. It sits on the provider's machine for most provisions here and reaches the
address by a path no consumer uses; in the measured outage it would have passed throughout.
2. **The host dials on the consumer's behalf, from the machine.** Rejected for the reason above: the
machine's network position is not the consumer's, and the one outage this exists to catch is exactly
a difference between them.
3. **Ask the module.** Rejected: a module is arbitrary software that the mesh does not write. Some could
report on their provisions and most cannot, and a check that covers the modules that opted in tells
nobody anything about the rest.
4. **Read the module's logs.** Rejected: the failure was in a log the whole time, and reading a module's
logs makes the mesh depend on the wording of software it does not control.
5. **The consumer verifies it, from its own network position.** Adopted.
## Decision
**A consumer verifies each grant it is given, from its own network position.** After a reconcile has
applied a grant, the machine opens a connection to the address and port that grant names, from inside
the consumer's own network namespace — the same position the consumer's software dials from, which is
the only position that answers the question the grant asks.
**It is a connection, not a conversation.** Whether the port accepts a connection is what a grant
claims; whether the credential is right, the role exists or the schema is current is the provider's to
answer and the consumer's to discover. A check that spoke each provision's protocol would be a second
implementation of every provision, and would fail for reasons that are not the mesh's.
**One failure is not news.** A provider restarting is ordinary, and so is a consumer between containers.
A grant is reported unreachable only after it has failed on **consecutive** reconciles, and the count is
what the machine reports rather than the last attempt — so a reader can tell "it was briefly away" from
"it has never worked".
**A grant that cannot be checked is said to be unchecked, never assumed good.** A consumer that is not
running has no network position to dial from; that is not a broken grant and must not read as one. It is
also not a verified grant, and the two are different sentences.
**What it costs to be wrong is the constraint on all of it.** A check that reports a working provision
broken trains a reader to ignore the report, which is worse than having none — the fault this
repository keeps finding, one level up. So the threshold is consecutive failures, the check is the
cheapest thing that answers the question, and an unknown is reported as unknown.
**The mesh says it where it says everything else.** A machine's report carries its unreachable grants,
and `status` names them beside what is out of date — so "every module current with its source" stops
being the whole of what the mesh will tell you about a machine whose modules cannot reach each other.
## Consequences
- **The mesh can be wrong out loud.** It has been able to assert a grant and not check it; now a grant
that does not work is a thing the mesh says, and the eleven hours of issue 145 become minutes.
- **The host gains the ability to act from a container's network position**, which it has not needed
before. That is a real capability and the only one this needs.
- **A machine reports something that is not about the declaration.** Everything it reports today is
what it applied and what it holds; this is the first thing it says about whether what it applied
works.
- **A provision with no port is not checked**, because there is nothing to dial. Several are files and
secrets, and saying "checked" about those would be the appearance of verification that this record
exists to remove.
- **What got harder:** a reconcile does more than apply. Every grant adds a connection attempt on a
cadence, which is cheap individually and worth naming: a machine with many consumers dials once per
grant per reconcile.
## How it is checked
- **The outage is caught.** A bed drops the path from a consumer's network position to a provider's
port while leaving the machine's own path to it open — the exact shape of issue 145 — and the grant
reads unreachable. This fails against the previous behaviour, where nothing reported anything, and
against a check run from the machine, which passes while the consumer cannot reach it.
- **A restarting provider is not an outage.** One failed reconcile reports nothing; the count rises and
falls, and the grant reads reachable again without anybody acting.
- **A consumer that is not running reads unchecked, not broken**, asserted separately from the
unreachable case because they are different sentences.
- **A provision with no port is not claimed to be checked.**
- **The report carries the count, not the last attempt**, so "briefly away" and "never worked" are
distinguishable by a reader who sees only the report.
- **`status` names an unreachable grant**, asserted on the output, since a check nothing surfaces is
the same as no check.
## References
- [ADR 0010](0010-delivery.md) — the declaration is owned resources; a grant is one of them
- [ADR 0009](0009-modules-and-the-graph.md) — what a provision and a consumer are
- [issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
— the eleven hours
- [issue 136](../04-ISSUES/136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md) — the
same distance between a declaration and a machine, one level down
@@ -0,0 +1,121 @@
---
topic: what runs on it
status: accepted
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md
supersedes: 02-DECISIONS/0143-a-consumer-verifies-the-grant-it-is-given.md
---
# 144. Anything on a machine may call anything on it, and that is the whole of "local"
## Context
Everything in the mesh should be able to call:
- what runs on the same machine;
- another machine's service over the private network, if that service is exposed there;
- another machine's service over the public network, if it is exposed there.
Three cases. The filter had two of them.
**The first was broken and the break was invisible.** A service exposed to the private network rendered
as the machines' own addresses on it. A caller on the machine carries such an address; a caller inside
one of that machine's containers carries a bridge address and matched nothing. Measured:
```
the machine: local 10.10.0.1 dev lo src 10.10.0.1
a container: 10.10.0.1 via 172.17.0.1 dev eth0 src 172.17.0.8
```
Same destination, same machine, two source addresses. The rule named the first and silently refused the
second, so a module reaching its database on its own machine's name timed out for eleven hours
([issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)).
**The second case works, and by accident.** A caller on another machine reaches the private network over
the tunnel, and arrives carrying that machine's own address — so the rule matches. It would not have
matched the caller's own address either; the tunnel rewrites it. That two of three cases worked is why
this looked correct.
**[ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) answered the wrong question.** Written
hours earlier, it proposed that a consumer verify each grant it is given by opening a connection from
its own network position — and it went to some length about *which* position, because whether a caller
sat in a container changed the answer. That difference was the bug. A verification mechanism would have
reported this outage sooner and would not have prevented it, and the machinery it needed existed only
because the rule was wrong. The remedy for a configuration error is the correct configuration.
**And a module is not a container.** A module is software that delivers one or more services, and it may
do that as a container, an installed package with a unit, a binary, or files something else reads. Of 72
modules in the catalogue, 61 happen to use a container and 11 do not — among them the resolver, the ssh
daemon and the intrusion-prevention module. A rule that reasons about containers describes most of the
mesh and not the mesh.
## Considered Options
1. **A line per service admitting the machine's own callers.** Rejected: it is what was written first,
and it only ever covers the services somebody remembered to think about. It also states, service by
service, a thing that is true of the machine.
2. **Verify each grant from the consumer's position** ([ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md)).
Rejected as a remedy: it observes the fault rather than removing it, and the question it agonised over
— which network position — exists only while the fault does.
3. **Enumerate the addresses a machine's callers may have.** Rejected for the reason no address is named
anywhere in this filter any more ([ADR 0140](0140-the-filter-constrains-what-arrives-from-outside.md)):
a range describes one machine and goes stale in silence.
4. **Local is not filtered, stated once.** Adopted.
## Decision
**Anything on a machine may call anything on that machine, and the filter says so once.** Not per
service, not per port, and not by naming who the callers are: traffic that did not arrive from outside
the machine and did not arrive over the private network is the machine's own, and is admitted. It is
asked by the link the traffic arrived on, because that is a fact about the machine rather than a list
that describes one.
**Local is not a boundary this mesh draws.** Whether a caller is a container, a unit, or the operator's
shell changes nothing, because the thing being decided is "is this the same machine" and the answer does
not depend on the form the caller takes.
**The other two cases are unchanged and are now legible beside it.** A service exposed to the private
network admits the machines on it; a service exposed publicly admits anything. Three cases, three lines,
and a reader can see all three at once.
**[ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) is superseded and nothing replaces it.**
Whether the mesh should check that a grant works is a real question — it reported this machine healthy
for eleven hours — but it is a question about what the mesh can say, not about what it should do, and it
must stand on its own rather than as the remedy for a rule that was wrong. It is not built.
## Consequences
- **The three things everything should be able to call are three lines**, and the first is one line
rather than one per service, so a service added tomorrow is reachable locally without anybody
remembering to say so.
- **A form of module stops mattering to the filter.** The 11 modules that are not containers were never
affected by this bug and were never the reason it was hard to see; they are the reason the rule should
never have mentioned containers.
- **The mesh still cannot say when a grant stops working.** That is the live gap, recorded in issue 145
and no longer pretending to have an answer.
- **What got harder:** nothing. This removes a line per service and replaces it with one.
## How it is checked
- **A caller on the machine reaches a service on it, in the input chain**, asserted on that chain's own
body — because the forward chain carries the same line in the same words, and an assertion on the
whole rendered file passed with the input chain's copy deleted. That is what
[ADR 0137](0137-a-machine-says-which-networks-it-routes.md)'s tests already say to do.
- **It is one rule, not one per service.** Asserted by rendering two services of different reach and
refusing a per-port local line.
- **The three reaches render as three lines**, asserted together, so the whole of what the filter says
about who may call what is one test.
- **The measured case:** from a container on the machine, a service exposed to the private network on
that machine answers. This is the outage, and it fails against the rule this replaces.
## References
- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the filter is the sum
of what its modules listen on
- [ADR 0140](0140-the-filter-constrains-what-arrives-from-outside.md) — why no address is named
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — internal and public,
the other two cases
- [ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) — superseded here
- [issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
@@ -0,0 +1,120 @@
---
topic: what runs on it
status: superseded
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md
superseded-by: 02-DECISIONS/0146-connectivity-is-checked-by-name-per-hosting-form.md
---
# 145. A module checks what the mesh claims is reachable, and it checks itself
## Context
The mesh asserts three things are callable ([ADR 0144](0144-anything-on-a-machine-may-call-anything-on-it.md)):
what runs on the same machine, another machine's service exposed to the private network, and another
machine's service exposed publicly. It has never checked any of them.
[Issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md):
the first of the three was broken for eleven hours and the mesh answered *all heard from, every module
current with its source* throughout. Every check it makes is about the relationship between the mesh and
a machine — applied, current, containers running — and none about whether anything can reach anything.
**A first answer was drafted and withdrawn.** [ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md)
put the check inside the host, verifying each grant from the consumer's network position. It was
superseded because the difference it worked so hard to reproduce — whether a caller sat in a container —
was the bug itself. What survives from it is the part that was right: a check run from the wrong place
proves nothing, and the mesh's own reports are not evidence about the network.
**The mesh already has the shape for this and it is a module.** A module can declare a container that
runs on a cadence ([ADR 0053](0053-a-step-that-runs-on-a-schedule.md), and three modules already use
`*/5 * * * *`), can be given the mesh's roster as a rendered fact — every machine's name, address and
this node's own identity, the same mechanism the resolver and the operator's ssh configuration use — and
can emit what it found on the bus. Nothing new is needed to build this except the module.
**What it must not check is the trap.** The obvious probe target is ssh: present on every machine, never
closed by design. Dialling it would have passed throughout the outage, because ssh is admitted
unconditionally and the thing that broke was a service exposed to the private network. A checker whose
probe is unconditionally open measures the one path that cannot fail, which is the failure this whole
sequence keeps producing — a check that reads as verification and verifies nothing.
## Considered Options
1. **The host verifies each grant** ([ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md)).
Superseded. It needed the host to act from another network position, which is machinery that exists
only while local calls are filtered wrongly.
2. **The control plane dials every node.** Rejected: it sits on one machine and reaches the others by a
path no ordinary caller uses. It would have passed throughout the outage.
3. **Probe an existing service.** Rejected for the target problem above: the services guaranteed on every
machine are the ones that are never closed, so they cannot fail the way the mesh fails.
4. **A module on every machine that serves its own probe and dials the others'.** Adopted.
## Decision
**A module runs on every machine, serves an endpoint of its own, and dials every other machine's.** The
probe is the module's own endpoint, declared reachable over the private network — so the thing being
dialled is admitted by exactly the rule that governs every other internally-exposed service, and fails
when that rule is wrong. A second endpoint, declared public, does the same for the public path where a
machine has one.
**It checks the three cases the mesh claims, by name:**
- its **own machine**, by dialling its own machine's address — the case that broke, and the only one that
distinguishes a caller on the machine from a caller in one of its containers;
- **each other machine over the private network**;
- **each machine's public path**, where one is recorded.
**It resolves before it dials, and says which failed.** A name that does not resolve and a port that does
not answer are different faults with different owners, and a checker that reports one sentence for both
sends a reader to the wrong place.
**It runs where the callers run.** The module's own code in its own container, on the cadence the mesh
already has, from the same position as every other module on that machine. It is not the host and not the
control plane, and that is the whole point.
**It says what it found and nothing else.** It emits results; it repairs nothing, opens nothing and holds
no credential beyond its own. A checker that fixes things is a second control plane.
**One failure is not a fault.** A machine rebooting is ordinary. A path is reported broken after it has
failed on consecutive runs, and the count travels with the result so a reader can tell "briefly away"
from "never worked" — the one thing [ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) got
right and worth keeping.
## Consequences
- **The mesh gains the ability to be wrong out loud about the network.** Eleven hours becomes two runs.
- **It is a module, so it is assigned, built, pushed and reported on like everything else** — no new host
capability, no new vocabulary, nothing in the control plane that has to know about checking.
- **Its own endpoint is the instrument.** That is what makes it able to fail; it also means the checker
must be assigned to a machine before that machine can be checked, and a machine without it is
unchecked rather than healthy.
- **It cannot check what it cannot be told.** The roster gives it machines; it does not give it every
module's endpoints, so this checks the paths the mesh claims and not every grant in the mesh. That is
the honest scope of a first one, and the difference is worth saying rather than growing quietly.
- **What got harder:** one more module on every machine, and a module whose whole purpose is to fail
visibly when something else is wrong. Its own failures will be read as the mesh's, which is the cost of
an instrument.
## How it is checked
- **It catches the measured outage.** A bed closes the path from a container to a service exposed to the
private network on its own machine — issue 145's shape — and the checker reports its own machine
unreachable while every other path still reads reachable. This fails against a probe on a port that is
never closed, which is the wrong target this record exists to name.
- **A machine rebooting is not a fault**: one failed run reports nothing, the count rises and falls.
- **A name that does not resolve is reported as that**, not as a port that did not answer.
- **It reports and does not act**: asserted by giving it a broken path and checking nothing on the machine
changed.
- **A machine without the module reads unchecked**, never healthy — asserted on what the mesh says about
a machine it is not assigned to.
## References
- [ADR 0144](0144-anything-on-a-machine-may-call-anything-on-it.md) — the three things that must be callable
- [ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) — superseded; what survives is that the
position matters
- [ADR 0053](0053-a-step-that-runs-on-a-schedule.md) — the cadence
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — internal and public,
which the probe endpoints declare
- [issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
@@ -0,0 +1,125 @@
---
topic: what runs on it
status: accepted
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0145-a-module-checks-what-the-mesh-claims-is-reachable.md
supersedes: 02-DECISIONS/0145-a-module-checks-what-the-mesh-claims-is-reachable.md
---
# 146. Connectivity is checked by name, per hosting form, with a valid certificate
## Context
[ADR 0145](0145-a-module-checks-what-the-mesh-claims-is-reachable.md) decided that a module checks what
the mesh claims is reachable, from where the callers are, because the mesh reported four machines healthy
for eleven hours while a module could not reach its database
([issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)).
That decision stands. What it got wrong is everything about *what* is dialled.
It dialled a raw port on each machine's address. Three things are wrong with that:
- **A raw port is not how anything in this mesh is reached.** A real caller resolves a name, the proxy
answers it, and the proxy reaches the service. A check that dials a port tests the last hop of a path
with four hops in it, and the three it skips — resolution, the proxy, the certificate — are where most
of the mesh's connectivity actually lives.
- **It tested one hosting form.** A module is software that delivers services, and it may deliver them
from a container, from a unit the mesh writes for its own code, or from a unit a package ships. Those
are three different paths to the same machine, and the outage that produced this was two of them
disagreeing. A probe served one way measures one way.
- **It said nothing about certificates.** An internal name that resolves, routes and answers over TLS
that nothing can verify is not a working path; it is a working path for whoever holds the proxy's
trust and nobody else.
## Decision
**Each hosting form gets its own endpoint, its own route and therefore its own name.** On every machine:
| name | what serves it |
|---|---|
| `connect-docker.<node>.internal` | a container |
| `connect-process.<node>.internal` | the mesh's own code, in a unit the mesh writes |
| `connect-unit.<node>.internal` | a unit a package ships |
and the same set under each machine's public domain where it has one — `connect-docker.<domain>` and its
siblings. The names are the instrument: a failure reads as *`connect-docker.g14.internal` did not answer*,
which says which machine and which hosting form without anybody interpreting anything.
**Every machine checks every machine, by name, over TLS, verifying the certificate.** Not a port, not an
address: resolve the name, connect, complete the handshake, check the certificate against the authority
that should have issued it — the mesh's own for an internal name, a public one for a public name. That is
the whole path a real caller takes, and each step failing is reported as itself.
**No name is written anywhere.** The machines come from the roster the mesh already renders as a fact, and
the labels are the module's. A machine that joins appears in every other machine's roster on the next
push, and they begin checking it without an edit.
**And the module arrives on a machine because the machine exists, not because somebody assigned it.** A
machine that joins and does not have it is worse than unchecked: every other machine is already dialling
its names, so it reads as broken everywhere until someone notices. This is the part the mesh cannot
currently express — see below — and it is the part that makes the rest safe.
**What survives from 0145**, unchanged: it reports and repairs nothing; one failure is not a fault and a
path is broken after consecutive runs with the count travelling with the result; findings are said on the
bus, because a finding in a file on the machine is what this exists to end; and the bus is the one path
that cannot report its own failure, so an emit that does not land is written locally and nowhere else.
## What this needs that the mesh does not have
Named here rather than assumed, because each is a decision of its own and this record is not the place to
make them:
1. **A module that every machine has.** `ScopeNode` means *at most one holder per node* — an exclusivity
rule, not an obligation — and nothing assigns a module at enrolment. Today the resolver, the packet
filter, ssh and intrusion prevention are each assigned per machine by hand, which is the same gap
wearing different clothes.
2. **A container running a module's own bundle.** A `process` runs the mesh's own compiled code with no
image; a `container` needs an image of the module's own, which means a Dockerfile — the thing the
`bundle` artifact exists to abolish. Nothing in the catalogue runs a bundle in a container, so
`connect-docker` has no shape yet.
3. **A unit a package ships, for `connect-unit`.** The `service` resource puts an existing unit into a
state and deliberately installs none, so this form needs a package that serves a port — and naming a
program the machine may not have is
[issue 136](../04-ISSUES/136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md).
4. **A machine's public domain in the roster fact.** The fact carries each machine's name, mesh name,
address and operator account. The public names cannot be composed without the domain.
5. **Something that installs the mesh's own root on a machine.** This is
[issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md), open
since before any of this. Until it is closed, every internal name will fail certificate verification
from every machine — correctly, because nothing can verify it. That is the checker working, and it is
worth saying in advance so the first run is not read as the checker being broken.
## Consequences
- **A failure names the machine and the hosting form.** That is the whole gain over a port: eleven hours
became two runs under 0145, and under this it also becomes one line that says where to look.
- **The checker surfaces issue 129 immediately**, and will report every internal name unverifiable until
it is fixed. A reader must be told that before the first run rather than after.
- **Five things must be built before this is what it says it is**, and until they are, what exists is a
port dial from one position — useful, and not this.
- **What got harder:** a module with three hosting forms of the same trivial service is a strange thing to
read. It is justified only because those three forms are how the mesh actually runs software, and a
checker that tested one of them would keep the class of outage it exists to catch.
## How it is checked
- **A name per hosting form answers from every machine**, asserted by name and not by port.
- **A certificate that does not verify is reported as that**, distinctly from a name that does not resolve
and a port that does not answer — three faults, three owners.
- **A machine that joins is checked by every other machine without an edit**, asserted by adding one to a
bed and looking at what the others dial on their next run.
- **A machine that joins has the module**, which is gap 1 above and is the assertion that cannot be
written yet.
- **The measured outage is still caught**: the path from a container to a service on its own machine is
closed and `connect-docker.<that node>.internal` fails from that machine while the others still pass.
## References
- [ADR 0145](0145-a-module-checks-what-the-mesh-claims-is-reachable.md) — superseded; its core stands
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — the two reaches these
names come from
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — a label plus a domain, which is why no name is written
- [issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md) — what the
internal names will fail on until it is closed
- [issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
+3
View File
@@ -223,6 +223,9 @@ python3 00-META/checks/index.py fail if stale
- **0139** — [A network is forwarded because a module declared it](0139-a-network-is-forwarded-because-a-module-declared-it.md) *(superseded)*
- **0140** — [The filter constrains what arrives from outside, and says nothing about a machine's own guests](0140-the-filter-constrains-what-arrives-from-outside.md)
- **0141** — [The host delivers its own successor, and versions live side by side](0141-the-host-delivers-its-own-successor.md)
- **0143** — [A consumer verifies the grant it is given](0143-a-consumer-verifies-the-grant-it-is-given.md) *(superseded)*
- **0144** — [Anything on a machine may call anything on it, and that is the whole of "local"](0144-anything-on-a-machine-may-call-anything-on-it.md)
- **0145** — [A module checks what the mesh claims is reachable, and it checks itself](0145-a-module-checks-what-the-mesh-claims-is-reachable.md)
### How it is built
+32 -1
View File
@@ -5,7 +5,7 @@ code:
- mesh-controller internal/builder
- mesh-controller cmd/mesh-controller (build, build --behind, push, status)
- mesh-controller internal/inventory/builds.go
updated: 2026-09-21
updated: 2026-09-29
decisions:
- 02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md
- 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md
@@ -226,3 +226,34 @@ when the current failure began and how many reports in a row have said it — th
id, whatever the words; three make the machine stuck, and `status` says so beside the failure. The host keeps trying — stuck is what the mesh
knows, not what the machine is told. *How it is checked:* an inventory test counts three identical
reports, a different one, and a clean apply; the status test asserts the word appears.
## Everything may call what is exposed to it, and local is not a boundary
*2026-09-29, from an outage that ran eleven hours —
[issue 145](../../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
settled by [ADR 0144](../../02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md).*
A grant is four facts and a credential: the provision, the machine, the port, and who the consumer is
when it connects. It is the whole mechanism by which anything in the mesh reaches anything else, and it
rests on three things being callable — what runs on the same machine, another machine's service over the
private network where it is exposed there, and another machine's service over the public network where it
is exposed there.
The filter had two of those. A service exposed to the private network admitted the machines' own addresses
on it; a caller on the machine carries such an address, and a caller inside one of that machine's
containers carries a bridge address and matched nothing. Measured, same destination and same machine:
`src 10.10.0.1` from the machine, `src 172.17.0.8` from a container on it. So a module reaching its
database on its own machine's name timed out for eleven hours while the mesh called the machine healthy.
The second case worked by accident: a caller on another machine arrives over the tunnel carrying that
machine's address, which the rule matched. Two of three working is why this read as correct.
**So local is not a boundary this mesh draws, and the filter says so once.** Traffic that did not arrive
from outside the machine and did not arrive over the private network is the machine's own, and is
admitted — for every service there, not per service. Whether the caller is a container, a unit or a shell
decides nothing, because the question is "is this the same machine".
A verification mechanism was drafted for this and withdrawn. It would have reported the outage sooner and
would not have prevented it, and the part of it that was hard — deciding which network position to check
from — existed only because the rule was wrong. Whether the mesh should check that a grant works is still
open, in issue 145; it is not the remedy for a configuration error.
@@ -5,7 +5,7 @@ located-in:
- mesh-controller internal/catalogue/filtering.go (fixed for this instance)
- mesh-controller (what status reports, and what it does not ask)
fixed-by:
amended-design:
amended-design: 03-DESIGN/01-to-be/10-delivery.md
---
# 145 — A machine reads healthy while its modules cannot reach each other
@@ -72,6 +72,36 @@ distance between a declaration and the machine, and in both cases the report was
**And the eleven hours are the measurement, not the bug.** The filter fault was one line and is fixed.
What is not fixed is that nothing in the mesh would have told anybody.
## What was decided
*2026-09-29, the same day, in two steps and the first was wrong.*
The first answer was [ADR 0143](../../02-DECISIONS/0143-a-consumer-verifies-the-grant-it-is-given.md):
the consumer verifies each grant from its own network position, because whether a caller sat in a
container changed whether it could reach the provider. **That difference was the fault**, and the record
is superseded. A verification mechanism would have reported this sooner and would not have prevented it,
and the part of it that was difficult — deciding which network position to check from — existed only
while the rule was wrong.
The remedy is [ADR 0144](../../02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md):
anything on a machine may call anything on it, said once rather than per service, and asked by the link
traffic arrives on rather than the address it carries. Everything should be able to call what runs on the
same machine, another machine's service exposed to the private network, and another machine's service
exposed publicly. The filter had the second and third and expressed the first as a list of addresses that
no container could match.
**And then the question this issue is actually about was answered on its own terms.**
[ADR 0145](../../02-DECISIONS/0145-a-module-checks-what-the-mesh-claims-is-reachable.md): a module on
every machine serves an endpoint of its own and dials every other machine's, from the position the
callers are in. Its probe is its own endpoint declared reachable over the private network, so it is
admitted by exactly the rule that governs every internally-exposed service and fails when that rule is
wrong — where a probe on a service every machine has would have passed for all eleven hours, because the
services every machine has are the ones never closed.
Adopted on its merits rather than as the remedy for a configuration error, which is what 0143 was and
why it went. The module is written and merged; it is not yet assigned, so every machine currently reads
unchecked.
## Open questions
- Should a grant be checked? The mesh knows the consumer, the provider, the address and the port, so a