A live outage, caused by converging the control node, and undetected for eleven hours while the mesh
reported every machine healthy.
A port declared from: mesh renders as the machines' own addresses on the private network. A container
reaching a port on the machine it runs on comes from a bridge address, so it matches none of them. And
where the container runtime routes directly, that packet is delivered to this machine rather than
forwarded — measured with end-of-chain counters: 368 packets reached the end of the input chain, none
reached the end of the forward chain — so the forward chain's allowance for the machine's own guests
never saw it either.
ADR 0100 requires this to work, in as many words: the store must be reachable "from a container on the
node itself". It was, through the rule the predecessor left behind, which admitted the private ranges
wholesale. Converging replaced that with four overlay addresses and closed it.
What it looked like: nextcloud accepting TCP and never answering HTTP, its own log repeating connection to server at "novox.internal" (10.10.0.1), port 6852 failed: timeout expired — 6,154 times
since the flip. Confirmed from a throwaway container: neither the store nor the forge reachable on the
machine's own address.
So a mesh-scoped port also admits traffic arriving on a link that is neither outward nor the tunnel —
this machine's own guests. Asked for by the link rather than the address, for the reason the forward
chain no longer names an address: a range describes one machine and goes stale in silence. A port open
to everything needs no such line, asserted separately.
Two tests; the first fails with the exact symptom when the line is removed.
How it surfaced: I was verifying that the endpoint-naming change had not altered any rendered route, found
every route identical, and then checked the routed services actually answered. drive.novox.be timed
out. Everything else answered, which is why this had been invisible — the affected modules are the ones
that reach another module by the machine's name, not the ones reached through the proxy.
**A live outage, caused by converging the control node, and undetected for eleven hours while the mesh
reported every machine healthy.**
A port declared `from: mesh` renders as the machines' own addresses on the private network. A container
reaching a port on the machine it runs on comes from a bridge address, so it matches none of them. And
where the container runtime routes directly, that packet is *delivered to this machine* rather than
forwarded — measured with end-of-chain counters: 368 packets reached the end of the input chain, none
reached the end of the forward chain — so the forward chain's allowance for the machine's own guests
never saw it either.
ADR 0100 requires this to work, in as many words: the store must be reachable "from a container on the
node itself". It was, through the rule the predecessor left behind, which admitted the private ranges
wholesale. Converging replaced that with four overlay addresses and closed it.
What it looked like: nextcloud accepting TCP and never answering HTTP, its own log repeating
`connection to server at "novox.internal" (10.10.0.1), port 6852 failed: timeout expired` — 6,154 times
since the flip. Confirmed from a throwaway container: neither the store nor the forge reachable on the
machine's own address.
So a mesh-scoped port also admits traffic arriving on a link that is neither outward nor the tunnel —
this machine's own guests. Asked for by the link rather than the address, for the reason the forward
chain no longer names an address: a range describes one machine and goes stale in silence. A port open
to everything needs no such line, asserted separately.
Two tests; the first fails with the exact symptom when the line is removed.
How it surfaced: I was verifying that the endpoint-naming change had not altered any rendered route, found
every route identical, and then checked the routed services actually answered. `drive.novox.be` timed
out. Everything else answered, which is why this had been invisible — the affected modules are the ones
that reach *another module* by the machine's name, not the ones reached through the proxy.
A port declared from the mesh admitted the machines' own addresses on the private
network. A container reaching a port on the machine it runs on comes from a bridge,
matching none of them — and where the container runtime routes directly, that packet
is delivered to this machine rather than forwarded, so the forward chain's allowance
never saw it either.
ADR 0100 requires this to work: the store is reachable 'from a container on the node
itself'. It was, through a rule the predecessor left, which allowed the private
ranges wholesale. Converging the machine replaced that with the four overlay
addresses and closed it.
Measured, and it was an outage: every module reaching another by its machine's own
name timed out for eleven hours while the mesh reported the machine healthy. A web
application logged 'connection to server at novox.internal (10.10.0.1), port 6852
failed: timeout expired' throughout.
Asked for by the link it arrives on, for the reason the forward chain no longer
names an address: a range describes one machine and goes stale in silence. A port
open to everything needs no such line.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
A live outage, caused by converging the control node, and undetected for eleven hours while the mesh
reported every machine healthy.
A port declared
from: meshrenders as the machines' own addresses on the private network. A containerreaching a port on the machine it runs on comes from a bridge address, so it matches none of them. And
where the container runtime routes directly, that packet is delivered to this machine rather than
forwarded — measured with end-of-chain counters: 368 packets reached the end of the input chain, none
reached the end of the forward chain — so the forward chain's allowance for the machine's own guests
never saw it either.
ADR 0100 requires this to work, in as many words: the store must be reachable "from a container on the
node itself". It was, through the rule the predecessor left behind, which admitted the private ranges
wholesale. Converging replaced that with four overlay addresses and closed it.
What it looked like: nextcloud accepting TCP and never answering HTTP, its own log repeating
connection to server at "novox.internal" (10.10.0.1), port 6852 failed: timeout expired— 6,154 timessince the flip. Confirmed from a throwaway container: neither the store nor the forge reachable on the
machine's own address.
So a mesh-scoped port also admits traffic arriving on a link that is neither outward nor the tunnel —
this machine's own guests. Asked for by the link rather than the address, for the reason the forward
chain no longer names an address: a range describes one machine and goes stale in silence. A port open
to everything needs no such line, asserted separately.
Two tests; the first fails with the exact symptom when the line is removed.
How it surfaced: I was verifying that the endpoint-naming change had not altered any rendered route, found
every route identical, and then checked the routed services actually answered.
drive.novox.betimedout. Everything else answered, which is why this had been invisible — the affected modules are the ones
that reach another module by the machine's name, not the ones reached through the proxy.