Compare commits

..
5 changed files with 218 additions and 0 deletions
@@ -0,0 +1,89 @@
---
topic: the mesh
status: proposed
date: 2026-10-02
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0004-a-node-and-how-it-joins.md
---
# 169. A machine joins through the tunnel, and the bus is never public
## Context
The bus is the one channel every machine depends on: enrolment, every declaration, every tool. The
`nats` module declares it reachable from the mesh only. The controller still opens it to the whole
internet on the machine that runs it, as a *foundation* port that no module declares and nothing may
close ([issue 051](../04-ISSUES/051-the-mesh-cannot-update-what-it-depends-on/00-report.md)).
The reason is joining. [ADR 0004](0004-a-node-and-how-it-joins.md) has a new machine enrol over the bus
**before** it has a tunnel. [ADR 0007](0007-connectivity.md) states it as a requirement: the node
running the broker must be reachable from wherever nodes are, at a stable address.
So the bus listens on the internet permanently, for an event that happens a few times a year. A
sweep of every machine on 2026-10-02 found no client using the public path. Every connection arrives
over the tunnel or from the machine itself. The join token does not use it either: it carries the
controller's configured broker address, a mesh name with the old broker's port.
ADR 0004 already says what a joining machine needs: *an identity, an address, and one peer to reach*.
The tunnel can be that peer, if the hub knows the new machine's key before the machine first knocks.
WireGuard answers nothing to a key it does not know, which is why the tunnel's own port is safe to
leave open where the bus's is not.
## Considered Options
1. **Keep the bus public.** It is authenticated and encrypted, but every exposure of it, and of the
server behind it, is exposure of the one thing everything depends on.
2. **Open the bus publicly only while a join token is live.** Small, and the hub is open only during a
join window. But the window is real, the rule is about time rather than about who may reach the
bus, and the opening and closing are pushes that can fail between them.
3. **The controller makes the new machine's tunnel key and puts it in the token.** One step for the
operator, but the private half leaves a machine it does not belong to. ADR 0004 refuses that for
every key a node holds.
4. **The machine makes its key first, and the token is issued for it.** The machine prints the public
half of its tunnel key. The operator issues the token for that key. The controller gives the
machine its address and adds it as a peer on the hub. The token carries the hub's tunnel endpoint
and key, the machine's address, and the bus's address on the private network. The machine brings
up its tunnel and enrols over it.
## Decision
**Option 4.**
- **A machine makes its own tunnel key before it has a token**, and prints the public half. The private
half never leaves it, as ADR 0004 says of every key a node holds.
- **A token is issued for a tunnel key.** Issuing it assigns the machine's address on the private
network, records the key, and makes the machine a peer of the hub. The hub is sent that before the
token is shown, so the tunnel answers the moment the machine first uses it.
- **The token carries the one peer.** It adds the hub's tunnel endpoint and public key and the
machine's own address. **Where** becomes the bus's address on the private network, which needs no
name resolution.
- **The machine joins through the tunnel.** It brings the tunnel up from the token alone, then enrols
over it exactly as before. The enrolment checks that the key it is offered is the one the token was
issued for.
- **The bus is never public.** It is no longer a foundation port. Its reach is what the `nats` module
declares: the mesh. The tunnel's port stays open, as the one way in.
This changes three things earlier records say. ADR 0004's *where* is the bus's private address, and the
token carries the peer. ADR 0007's requirement that the broker be reachable from wherever nodes are
becomes: **the hub's tunnel is**. Issue 051's broker port stops being a foundation port.
## Consequences
- Joining is two commands on the new machine, with the token issued between them. A token issued for
the wrong key gives a tunnel that never answers, and the machine says so rather than timing out at
the bus.
- An unused token leaves a peer on the hub until it expires. Expiry removes it, the same way it voids
the secret.
- A machine already in the mesh is unaffected: it reaches the bus over its tunnel today.
- The genesis machine, the first one, raises the bus on itself and needs no tunnel to reach it.
## How this is checked
| Rule | Checked by |
|---|---|
| A token is refused without a tunnel key, and carries the hub's peer and the machine's address | a controller test |
| Issuing a token makes the machine a peer of the hub before the token is shown | a controller test over the hub's composed tunnel |
| An expired, unused token's peer is gone from the hub | a controller test |
| Enrolment refuses a tunnel key other than the one the token was issued for | a controller test |
| No machine's filter opens the bus to anywhere | a controller test over the composed filter, and the live sweep from outside the mesh |
| A new machine joins from outside the hub's network with the bus closed to it | the lab, then by hand |
+1
View File
@@ -179,6 +179,7 @@ python3 00-META/checks/index.py fail if stale
- **0163** — [Taking a module over is a comparison: what it compares, what it refuses, and what it carries](0163-taking-a-module-over-is-a-comparison.md)
- **0167** — [A membership carries what its module receives, and who the mesh is](0167-a-membership-carries-what-its-module-receives-and-who-the-mesh-is.md)
- **0168** — [A converged machine is filtered by the mesh alone, and the host says what else refuses](0168-a-converged-machine-is-filtered-by-the-mesh-alone.md)
- **0169** — [A machine joins through the tunnel, and the bus is never public](0169-a-machine-joins-through-the-tunnel-and-the-bus-is-never-public.md) *(proposed)*
### Its tiers, from the bottom up
@@ -0,0 +1,43 @@
---
status: resolved
opened: 2026-10-02
located-in: [mesh-host internal/outward (Links reported only the links carrying a default route)]
fixed-by: mesh-host PR 66 (a link backed by a physical device is named outward, up or down), live 2026-10-02
amended-design: []
---
# 197 — A physical link that is down is not filtered when it comes up
## What was observed
A sweep of every machine's filter on 2026-10-02. A laptop-class machine connected by its radio has a
wired port that was unplugged. Its filter guarded the radio and the tunnel, and accepted everything
arriving on any other link:
```
iifname != { "mesh0", "<radio>" } accept
```
The wired port was not in the list. Plugged in, everything arriving on it would have been accepted,
every port of the machine open to whatever network the cable reached. That would last until the
machine reported again and was pushed a new filter.
## Why it matters
**The filter's one rule about links fails open.** [ADR 0140](../../02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md)
has the filter constrain what arrives from outside, and has the machine say which links face outside.
Everything not named is treated as the machine's own, its containers and bridges. So a link the machine
fails to name is not filtered at all. The host named only the links carrying a default route at the
moment it reported. A cable plugged in later is the ordinary case for a laptop. A second wired network
that never carries the default route, such as a direct link to a storage box, is never named at all.
## Open questions
- A virtual link that faces outside (a VPN client's interface, a USB tether that appears as a virtual
device) has no physical device behind it. It is named only while it carries the default route. Is
that enough?
## Resolved (2026-10-02)
Live on the affected machine after the host was delivered and one more push: its filter now guards the
radio, the tunnel and the unplugged wired port, before anything is plugged into it.
@@ -0,0 +1,14 @@
# Diagnosis
*2026-10-02.*
**Located in `mesh-host` `internal/outward`.** `Links` read the kernel's routing tables and returned the
interfaces carrying a default route. An unplugged port carries none, so it was never reported, and the
controller rendered the filter around the links it was given.
**The fix.** A link faces outside if it carries a default route **or** has a physical device behind it.
The kernel lists every interface under `/sys/class/net`, with a `device` entry for one backed by
hardware. A bridge, a veth, the tunnel and the loopback have none, so they stay the machine's own. The
wired port is now reported up or down, and the filter guards it before anything is plugged in. Tested
with a radio carrying the default route and an unplugged wired port beside a bridge, a veth, the docker
bridge, the tunnel and the loopback: the two physical links are reported, nothing else.
@@ -0,0 +1,71 @@
---
status: resolved
opened: 2026-10-02
located-in: [mesh-catalog modules/dnsmasq (listens on loopback and the machine's mesh address only), the home-server's DNS (a predecessor's dnsmasq configuration the mesh did not own), the home network's DHCP (hands out the home-server as every device's DNS)]
fixed-by: mesh-catalog PR 214 (dnsmasq listens on addresses from a setting; docker's file takes no settings), mesh-controller PR 210 (the settings verb), mesh-catalog PR 215 (unifi network DNS tools), live 2026-10-02
amended-design: []
---
# 198 — The home network's DNS server ran outside the mesh, and the mesh's filter closed it
## What was observed
Every phone on the home Wi-Fi had no internet, while a laptop on the same Wi-Fi did. The router's
DHCP hands every device the home-server's LAN address as its DNS server. The home-server's DNS daemon
was listening on that address, and every query to it timed out. The router itself answered the same
query at once. The laptop worked because it resolves through its own local resolver, not through the
server DHCP names.
## Why it happened
The DNS daemon on the home-server was not the mesh's. It ran under a configuration file a predecessor
generated, listening on loopback, the mesh address and the LAN address. The mesh's `dnsmasq` module was
assigned to the other three machines and not to this one, so no module on the home-server declared
port 53. Its filter opens only what a module declares, so DNS from the LAN was dropped. It started when
the home-server applied the filter this morning, after nine hours of applying nothing
([issue 194](../194-the-hosts-own-former-archive-stops-every-apply/00-report.md)).
Nothing said so. The daemon reported running, the filter applied cleanly, and the mesh had no record
that the home network depended on a service it did not know.
## Why it matters
**A service the mesh does not know is closed by the mesh's filter, by design, and nothing asks whether
something depends on it.** That is the right default for an unknown port. It is the wrong outcome for
the one service a whole network was told to use. The gap is that a machine can run something
important outside the mesh with nothing to show it.
**The mesh's `dnsmasq` could not have served the LAN either.** It listened on loopback and the mesh
address only. The reach of its DNS endpoints opens the filter, but the daemon would not have been
listening on the LAN address anyway.
## Open questions
- Should a machine report the listening services the mesh does not own, the way it reports the links
that face outside? This one would have been visible before the filter closed it.
- The LAN address the home-server answers on is now a setting, beside the reach that opens the filter.
Two statements that must agree. Should reach `public` on a DNS endpoint imply listening beyond the
mesh?
## Resolved (2026-10-02)
The home network was pointed at the gateway for DNS while the fix was built, which got the phones back
within minutes. Then:
- the mesh's `dnsmasq` takes the addresses it listens on beside the machine's from a setting, with
loopback as the mesh-wide default, so no other machine changed;
- the home-server's layer adds its LAN address, and its DNS endpoints' reach is `public`. The router
forwards no DNS, so that means the LAN;
- the module and its sibling `resolv-conf` were assigned to the home-server, replacing the
predecessor's daemon and configuration, which were kept aside;
- the home network was pointed back at the home-server, through a new `unifi` tool.
Checked live: from another machine on the LAN, public names and mesh names both resolve through the
home-server's LAN address, and the mesh and the machine itself resolve as before.
**One fault found on the way, and caught before it reached any machine.** A module's settings are
merged into every mergeable file the module owns. The first attempt therefore put the new setting into
docker's `daemon.json` as well as into dnsmasq's config, and dockerd refuses keys it does not know. The
plan showed it before any push. The change was reverted and redone with docker's file declared to take
no settings. The general fault, a module's settings reaching files they were not meant for, is still
there for any module with more than one file.