Issue 063 — the foundation's ports are forwarded, resolved
The broker's port was accepted on input but not forwarded, so a joined node reached a DNAT'd broker only until it restarted. Fixed in mesh-controller (25e42b3), proven by the built-store-cross-node bed adopting the foundation's broker (which restarts it) and the joined node still receiving declarations.
This commit is contained in:
+50
@@ -0,0 +1,50 @@
|
|||||||
|
---
|
||||||
|
status: resolved
|
||||||
|
opened: 2026-09-18
|
||||||
|
located-in:
|
||||||
|
- mesh-controller
|
||||||
|
fixed-by: mesh-controller fix/one-push-is-enough (25e42b3); gated by the built-store-cross-node bed
|
||||||
|
amended-design:
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# 063 — The foundation's ports are accepted on input but not forwarded
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
A joined node reaches the mesh's broker once, then — after the broker restarts — can never reach
|
||||||
|
it again, and so never receives another declaration. Every check between them passes: the port is
|
||||||
|
open, the broker is up, the overlay is established.
|
||||||
|
|
||||||
|
Observed on the built-store-cross-node bed, intermittently. A joined node's host connected to the
|
||||||
|
broker, received its first declaration, and worked. When the foundation's own broker was adopted —
|
||||||
|
which restarts it — the node's link dropped with `CONNECTION_FORCED - Broker shutdown`, then
|
||||||
|
`connection refused`, then `i/o timeout` for ever. The node was left with no declaration on disk
|
||||||
|
at all, and every downstream assertion (here, the registry trust) failed as a consequence of a
|
||||||
|
node that had received nothing.
|
||||||
|
|
||||||
|
The bed passed on the runs where the broker did not happen to restart after the joined node first
|
||||||
|
connected, which is why three green runs preceded the red one on identical code.
|
||||||
|
|
||||||
|
## Why it matters beyond the instance
|
||||||
|
|
||||||
|
The mesh firewall opens the foundation's ports — the broker above all — in the **input** chain,
|
||||||
|
from anywhere, so a node can enrol before it has an overlay address. But the broker is a published
|
||||||
|
container port: a cross-node packet to it is redirected by the runtime and handled in the
|
||||||
|
**forward** chain, which never carried a rule for the foundation. The first connection survived
|
||||||
|
only on its conntrack `established` entry; a broker restart dropped the entry, and the next dial
|
||||||
|
hit the forward chain's default drop.
|
||||||
|
|
||||||
|
This is [issue 047](../047-the-firewall-does-not-cover-published-container-ports/00-report.md)'s
|
||||||
|
lesson — a firewall with no forward coverage says nothing about container ports — reappearing for
|
||||||
|
the one port the whole mesh depends on, because the foundation is not a module and was added to
|
||||||
|
input alone. A rule that works only until the thing it governs restarts is worse than none: it
|
||||||
|
passes every test written before the restart.
|
||||||
|
|
||||||
|
## Open questions
|
||||||
|
|
||||||
|
- Is the broker the only foundation port that resolves to a container, or should every foundation
|
||||||
|
port be assumed to be published and forwarded on principle? (The fix forwards all of them, on
|
||||||
|
that principle.)
|
||||||
|
- Should a bed assert reachability *after* a deliberate broker restart, so "reachable until it
|
||||||
|
bounces" can never again read as "reachable"?
|
||||||
Reference in New Issue
Block a user