Compare commits

..
Author SHA1 Message Date
jschoubben 492ac7be18 Issue 145: a machine reads healthy while its modules cannot reach each other
Converging the control node closed every path by which a module reached another by
the machine's own name, and it ran for eleven hours while the mesh answered 'all
doing what they were told, all heard from, running what the mesh would send them,
and every module current with its source'.

6,154 database failures in one module's log, beginning at the minute of the flip.
The service accepted TCP and never answered HTTP. Every check the mesh makes
passed, because every check the mesh makes is about the relationship between the
mesh and a machine — applied, current, containers running. None asks whether a
module can reach what it requires, though the mesh composes every grant and so
knows exactly who requires what.

The filter fault is fixed. The eleven hours are the measurement, not the bug.

Also records on 144 that 'more closed than the mesh believes' was harmless only
for what is reached from outside, and an outage for what is reached from within.
2026-09-29 12:37:08 +02:00
mesh-admin 5ac77e3cef Merge pull request 'ADR 0138: reach asks for names on a routed endpoint' (#178) from decision/0138-insight-reach-and-the-proxy into main 2026-09-29 00:50:28 +00:00
jschoubben a619022c35 ADR 0138: a progressive insight — reach asks for names on a routed endpoint
The record says internal and public each mean something to the filter. For an
endpoint the proxy serves, the second half is wrong, and ADR 0045 said so first:
a public service is exposed through the proxy, listening from the mesh, not by
opening its own port.

Found by trying to express one real module, not by review — routed name public
because browsers post to it, machine port private because it serves a dashboard in
cleartext. Under one value for both, saying public would have reopened a port an
operator had just closed. Measured the same evening: the routed name answered from
the internet over TLS while the port was refused from the same place.

Corrects a fact. One statement per endpoint with three things derived from it
stands; the filter column applies to an unrouted endpoint.
2026-09-29 02:50:25 +02:00
mesh-admin 49c065c204 Merge pull request 'Issue 143: correct the diagnosis' (#177) from issue/143-corrected-diagnosis into main 2026-09-29 00:33:13 +00:00
jschoubben ddb3980f09 Issue 143: correct the diagnosis — the step exists and did not fire
The first account said the declaration carries no resource that would disable the
found firewall and that nothing implemented the sentence the flip prints. Wrong:
the mechanism is a step in the host's own apply, retireFirewall, and it is careful
— it reads back that the mesh's table is loaded before retiring anything, records
the forward policies first, and verifies ufw reports inactive afterwards.

What is established: ufw was active and enabled two minutes after a flip that
reported the node converged with nothing failed; and the machine's record now
reads disabled_by_mesh: true, written by a reconcile fifty minutes later that
found ufw already inactive because an operator had disabled it by hand.

So the step did not take effect and the record says it did. The candidates are
named rather than chosen — the step is skipped silently when the apply has any
failure, the mesh's table is loaded by a service in the same apply so ordering is
open, and the host's detail lines do not reach the journal, which is why this has
candidates instead of a cause.
2026-09-29 02:33:10 +02:00
mesh-admin 75a8f6abc7 Merge pull request 'Issues 143 and 144: the found firewall is neither retired nor all of it' (#176) from issue/143-and-144-the-found-firewall into main 2026-09-28 23:36:10 +00:00
jschoubben 862253518f Issues 143 and 144: the found firewall is neither retired nor all of it
Both found by converging the control node — the first machine with a firewall to
flip, since the two before it had none.

143: the preview and the flip both say the found firewall is disabled. The node
reported converged, 372 resources applied, nothing failed, and ufw is still
enabled and active. A converged node's declaration carries no resource that would
disable it; the sentence is printed by the command and nothing implements it.

144: ufw was never what filtered the traffic that mattered there. Fifty forwarded
openings converged through it had matched zero packets, while a chain the
predecessor installed in the container runtime's pre-accept hook did the work —
in memory only, recreated by nothing. The mesh's filter now covers that path, so
the machine no longer depends on it, but the chain remains and is the only thing
refusing the bus and the registry, which the design requires reachable from
anywhere so a machine can enrol before it has a private address.
2026-09-29 01:36:08 +02:00
mesh-admin 51ef3eb7e2 Merge pull request 'ADR 0142: the mesh delivers its own components as binaries' (#175) from decision/0142-mesh-delivers-its-own-components into main 2026-09-28 22:51:44 +00:00
4 changed files with 300 additions and 0 deletions
@@ -99,6 +99,31 @@ composed, so it is not certified.
binding. The per-node source override becomes its reach, widened from the filter alone to the names
and the certificate as well.
## Progressive insight — 2026-09-29, from building it
**Reach does not mean the same thing to the filter for an endpoint the proxy serves.** The decision
above says `internal` means "the filter opens the machine port to the private network" and `public`
means "the filter opens it to anywhere". For a routed endpoint the second half is wrong, and
[ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) already said so before this
record was written: *a public service is exposed through the proxy, not by opening its own port* — it
listens `from: mesh`, only the proxy reaches it, and it is exposed by name.
Found by trying to express one real module, not by review. Its routed name must be public, because
browsers post to it; its machine-side port must not be, because that port serves the dashboard in
cleartext. Under one value driving both, saying "public" would have reopened a port an operator had
just closed. Measured the same evening: that module's routed name answered from the internet over TLS
while its machine-side port was refused from the same place. The port is not the path.
So the reach of a **routed** endpoint asks for names, and its port keeps what the manifest said. The
reach of an **unrouted** endpoint — git over ssh, a mail port, the bus — governs the port, because
there is no name and the port is the only way in. That is the same split this record already draws in
*an endpoint that is not routed is reached but never named*; what it got wrong was carrying the filter
across it.
This corrects a fact, not the decision: one statement per endpoint, three things derived from it and
none of them deciding on its own, all stand. The table in the decision should be read with the filter
column applying to an unrouted endpoint.
## Consequences
- **A manifest gains endpoint names, and a route contribution names an endpoint instead of a port.**
@@ -0,0 +1,102 @@
---
status: located
opened: 2026-09-29
located-in:
- mesh-host internal/apply/opening.go (retireFirewall)
- mesh-host internal/apply/apply.go (the condition it is called under)
fixed-by:
amended-design:
---
# 143 — Converging a machine does not retire the firewall it found, and says it does
## What was observed
The control-node was converged on 2026-09-29, the first machine with a found firewall to be flipped —
the two converged before it had none.
The preview said, and the flip repeated:
```
the found firewall (ufw) is disabled, never flushed: its configuration stays on disk
...
sent: the host loads the mesh's filter and disables the firewall it found
```
The mesh then reported the node `converged`, 372 resources applied, nothing failed. Afterwards, on the
machine:
```
systemctl is-enabled ufw -> enabled
systemctl is-active ufw -> active
```
*Corrected 2026-09-29, an hour later, from reading the host rather than the declaration.* **The first
account of this was wrong.** It said the declaration carries no resource that would disable the found
firewall, and that the sentence was printed by the command with nothing implementing it. The
declaration indeed carries no such resource — but the mechanism was never meant to be one. It is a
step in the host's own apply, `retireFirewall`, and it exists, is careful, and is strict: it refuses to
retire anything until it has read back from the machine that the mesh's own table is loaded, it records
the forward policies first so a half-done retirement can be retried, and it verifies ufw reports
inactive afterwards.
What is established is narrower and stranger than "nothing implements it":
- ufw was **active and enabled two minutes after the flip**, and the flip had reported the node
converged with 372 resources applied and nothing failed.
- The machine's own record now reads `disabled_by_mesh: true` — but it was written by a reconcile
*after* an operator disabled ufw by hand, roughly fifty minutes later. A reconcile found ufw already
inactive, asked it to be inactive, read that back, and recorded that the mesh had done it.
- So the step did not take effect at the flip, and the machine's record now says it did.
The candidates are named rather than chosen, because the evidence does not separate them: the step is
called only when the apply had no failures, and a skipped step is silent; the mesh's table is loaded by
a service in the same apply, so whether it was loaded *at the moment the step asked* is an ordering
question; and the host's own detail lines do not reach the journal, so what it decided is not
recoverable after the fact.
## Why it matters beyond this instance
**It is a stated behaviour that does not happen, reported as success** — the fault this repository
exists to catch, and
[ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md) states it as
part of what the flip *is*: "loads the mesh's derived filter in place of its refusal-only table, and
retires the found firewall by disabling it, never by flushing".
**It could only be found on the first machine that had one.** The two machines converged before this
had no firewall to retire, so the step had never run, and nothing reported that it had not. That is
the same shape as [issue 136](../136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md):
a step that is silent when it does nothing.
**The machine is left doubly filtered, which is not what either firewall describes.** Every base chain
at a hook runs and a drop in any is final, so the machine now enforces the *intersection* of the mesh's
derived filter and a rule set left by the system being replaced. Nothing is broken by that today —
measured from outside, mail, the proxy and git-over-ssh answer and the databases and admin interfaces
are refused — but the machine's behaviour is described by neither of the two things claiming to
describe it, and the stale set includes a rule for a broker that no longer exists.
**And returning the node to adopted would be wrong in the other direction.** ADR 0100 says that
restores the found firewall by enabling it again; enabling something that was never disabled is
harmless, but the mesh's belief about which firewall is in force has been wrong in both modes.
## Open questions
- Which side owns retiring it — a resource in the declaration, so it is applied and reported like
everything else, or the flip as an act? A resource seems right: the flip is otherwise entirely
expressed as one, and an act that only the command performs cannot be re-checked on a later
reconcile.
- What should a reconcile do if the found firewall is enabled again by hand, or by a package update?
Convergence is a state, so presumably re-disable it and say so.
- Should the preview say what it *will* do rather than what it does, until a step exists that does it?
The wording was read as evidence twice in one session.
- Is there a check that a sentence the mesh prints corresponds to something that happened? This is the
second time in one session that a printed claim and the machine disagreed.
- **Why did the step not take effect?** It is called only when the apply had no failures, and being
skipped is silent. The mesh's table is loaded by a service in the same apply, so whether it was
loaded when the step asked is an ordering question — and ADR 0100 makes loading it first a
precondition rather than an expectation.
- **A step that records the mesh as having done what an operator did is worse than the omission.** The
record now says the mesh disabled ufw. Nothing distinguishes "we did this" from "we found it already
so". Should it?
- Why do the host's own detail lines not reach the journal? Everything it decided during the flip is
unrecoverable, which is why this account has candidates instead of a cause.
@@ -0,0 +1,84 @@
---
status: located
opened: 2026-09-29
located-in:
- mesh-host internal/apply/opening.go
- mesh-controller cmd/mesh-controller (the converge preview)
fixed-by:
amended-design:
---
# 144 — A predecessor's rules outlive the firewall the mesh found, and the mesh cannot see them
## What was observed
The mesh reports one thing about a machine's existing filtering: `firewall found: ufw`. On the
control-node, ufw was never what filtered the traffic that mattered.
Measured on 2026-09-29, before the machine was converged:
- ufw filters connections *to the machine*. It does not filter connections to a container's published
port, which arrive on the forwarded path where the container runtime accepts them before ufw's
forward chains are reached. Around thirty ports were published that way.
- Every one of the mesh's own forwarded openings, converged through ufw, had matched **zero packets** —
fifty rules in that chain, none ever matched, while the chain itself had passed 1.6 million
established packets. The restrictions read as applied and were inert.
- What actually kept those ports off the internet was a chain the predecessor installed in the
container runtime's own pre-accept hook, allowing the deliberately public ports and the private
ranges and dropping the rest on the outward link. Confirmed from outside: the proxy answered, the
container manager did not.
- That chain exists only in the running kernel. The persisted rule file is the distribution's empty
default, and nothing on disk recreates the chain.
After the flip, the mesh's own filter is loaded and does cover the forwarded path, so the machine no
longer depends on that chain. But **the chain is still there**, and it is now the only thing refusing
two ports the mesh believes are open: the bus and the registry, which
[ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md) requires be
reachable from anywhere so a machine can enrol and pull before it has a private-network address. The
mesh's rendered filter accepts both from anywhere. From outside, both are refused.
## Why it matters beyond this instance
**"The firewall found" is a kind, and filtering is not all in one place.** The host identifies one
front-end and reports it. A machine can carry rules from several sources — the front-end's own, the
container runtime's, an intrusion-prevention chain, and whatever a predecessor installed directly —
and the mesh's account of what filters the machine names exactly one of them.
**So adoption's central promise was half-true in both directions.** What the mesh converged through
the found firewall on the forwarded path did nothing at all, and what did the work was invisible to it.
A machine was reported as filtered by a mechanism that was not filtering.
**And convergence cannot retire what it cannot see.** Even once
[issue 143](../143-converging-does-not-retire-the-firewall-it-found/00-report.md) is fixed and the found
firewall is disabled, this chain remains, silently narrowing the machine below what the mesh's own
filter says. A rule the mesh did not write, cannot list, and will not remove — which today breaks the
enrolment path the design guarantees.
**The safe direction is not the same as the correct one.** Being more closed than intended broke nothing
visible, which is exactly why it went unnoticed for as long as the mesh has been on this machine.
## What it cost, measured later the same day
*2026-09-29.* The predecessor's chain was removed, and something it had been carrying went with it. It
admitted the private ranges wholesale, which is how a container on the machine reached a port declared
for the private network — the mesh's own filter admits the machines' overlay addresses, and a container
comes from a bridge. Every module that reached another by the machine's own name had been relying on the
predecessor's rule without anybody knowing.
That is [issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
and it ran for eleven hours while the mesh reported the machine healthy. The filter is fixed. What this
adds to the account here is that "the machine is more closed than the mesh believes" was not the
harmless direction after all — it was harmless for everything reached from outside, and an outage for
everything reached from within.
## Open questions
- Should the host report every place the machine filters from, rather than one kind — the front-end,
the runtime's hooks, and any chain it does not recognise, named so a person can look?
- What should the mesh do about rules it did not write and does not understand? Reporting them seems
right; removing them cannot be, and leaving them silent is what produced this.
- Does an opening converged through a found firewall need a check that it can actually take effect?
Fifty rules matching nothing would have been visible from the counters at any point.
- Is the bus and the registry being reachable from anywhere still what the mesh wants on a machine that
faces the internet? The design says yes, for enrolment. It deserves asking on its own rather than
being answered by a leftover.
@@ -0,0 +1,89 @@
---
status: located
opened: 2026-09-29
located-in:
- mesh-controller internal/catalogue/filtering.go (fixed for this instance)
- mesh-controller (what status reports, and what it does not ask)
fixed-by:
amended-design:
---
# 145 — A machine reads healthy while its modules cannot reach each other
## What was observed
Converging the control-node closed every path by which a module on that machine reached another module
by the machine's own name. It ran for **eleven hours**. Throughout, the mesh answered:
```
4 machine(s), all doing what they were told, all heard from,
running what the mesh would send them, and every module current with its source
```
What was actually happening, from one affected module's own log:
```
Doctrine\DBAL\Exception: Failed to connect to the database:
SQLSTATE[08006] connection to server at "novox.internal" (10.10.0.1), port 6852 failed: timeout expired
```
6,154 of them, beginning at the minute of the flip. The web application accepted TCP connections and
never answered an HTTP request; a client waited 35 seconds and gave up. Confirmed from a throwaway
container on the machine: neither the store nor the forge was reachable on the machine's own address.
The cause is [issue 144](../144-the-predecessors-rules-outlive-the-firewall-it-was-found-as/00-report.md)'s
sibling and is fixed: a port declared reachable from the private network admitted the machines' own
overlay addresses, and a container on the machine comes from a bridge address, matching none of them.
What this issue is about is the eleven hours.
**Nothing the mesh reports would have shown it.** Every check the mesh makes passed, because every
check the mesh makes is about the relationship between the mesh and a machine:
- the machine applied what it was sent, and said so;
- its declaration digest matches what the mesh would send;
- every module's source commit matches what the mesh holds;
- every container the declaration names is running.
None of those asks whether a module can reach what it requires. The mesh knows precisely who requires
what — it composes the grants — and never checks that the grant works.
**Nor would an operator's usual look.** The ports were probed from outside and behaved correctly; the
routed services answered; a container's egress to the internet worked. Those are the paths a person
checks after changing a firewall, and all three were fine. The broken path was module-to-module over
the machine's own name, which nothing routine exercises.
## Why it matters beyond this instance
**A mesh that composes a dependency and never tests it can only report on itself.** Every provision the
mesh grants is a claim that a consumer can reach a provider. The mesh asserts that claim, delivers
credentials for it, and has no mechanism that ever finds out. "Every module current with its source"
is a statement about bytes, not about whether anything works.
**The failure was silent in the direction that hides longest.** A service that will not start is
noticed. A service that starts, accepts connections and then cannot reach its database serves errors
under a healthy-looking process, and the machine's own report says the container is running — which it
is.
**It is the same shape as [issue 136](../136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md),
one level up.** There, a module named a program the machine lacked and everything reported success.
Here, the mesh granted a provision the filter refused and everything reported success. Both are the
distance between a declaration and the machine, and in both cases the report was about the declaration.
**And the eleven hours are the measurement, not the bug.** The filter fault was one line and is fixed.
What is not fixed is that nothing in the mesh would have told anybody.
## Open questions
- Should a grant be checked? The mesh knows the consumer, the provider, the address and the port, so a
reachability check is expressible — but from where: the consumer's machine, as part of a reconcile,
or the provider's?
- What would it cost to be wrong in the other direction? A check that reports a provision broken while
it works is worse than none, because it trains a reader to ignore the report. A provider restarting is
ordinary; a consumer between containers is ordinary.
- What should `status` say about a machine whose modules cannot reach each other? It currently has one
vocabulary for "heard from and current", and that sentence was true the whole time.
- Is there a cheaper signal than a probe? The affected module was logging the failure 6,154 times. The
mesh reads no module's logs and arguably should not — but something a module could *say* about its
own provisions would have surfaced this in minutes.
- Does the same blindness apply to the other direction — a provider that lost a consumer's grant and
is refusing it? Nothing checks that either.