Files
hq/01-RESEARCH/004-lab-network/analysis.md
T
jschoubben c0b35652d0 The numbering is the flow: decisions are 02, design is 03
papa-hq reads 01 research -> 03 decision -> 02 design. The order is a
scar, not a choice: 02-DESIGN existed from its initial commit, and when
adr/ was finally promoted on 2026-07-13 it took the next free number
rather than its place in the sequence. By then design was too settled to
renumber.

hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and
02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks
the process in the order it happens: research produces a decision, the
decision authorises a design.

00-GENESIS becomes 00-META, matching papa's rename from the same
restructure.

Every path reference rewritten across documents, frontmatter, playbooks
and skills. All links resolve; all 58 frontmatter blocks parse and their
path fields still point at files that exist.
2026-08-23 18:05:11 +02:00

204 lines
10 KiB
Markdown

---
effort: 004-lab-network
updated: 2026-08-22
---
# Reproducing the mesh network in a lab
Established 2026-08-22 by reading the generating code and the live mesh DB. Every claim is a
file location or a queried row.
---
## 1. The network is data, not configuration
`install.d/mesh-init.sh` and `install.d/adopt.sh` perform **no network configuration at
all** — no WireGuard, no DNS, no firewall, no `/etc/hosts`. Every part of the network layer is
generated by module hooks from mesh-DB rows:
| layer | generated by | from |
|---|---|---|
| WireGuard interface + peers | `modules/wireguard/hooks/index.ts` (`postConfigure`) | `node_wg_keys`, `nodes.site`, `nodes.underlay_addr`, `module_env.WG_ADDRESS` |
| `.internal` name resolution | `modules/dnsmasq-app/hooks/index.ts` | mesh config peers → `internal_domain` + `wg_address` |
| public routing / vhosts | `modules/hal/sdk/src/vhost-gen.ts`, `feature-handlers/vhost.ts` | `vhosts:` manifests + `node_accessors` |
| internal TLS | `modules/mesh-ca/hooks/index.ts` | a singleton CA row in the mesh DB |
**Consequence:** a faithful lab is mostly a matter of writing the right rows. The network that
results is produced by the same code production runs, which is the difference between testing
the network and testing a model of it.
---
## 2. The constraint that decides whether the lab works
`modules/wireguard/hooks/index.ts:225-240` decides, per pair, whether to write an `Endpoint`:
```js
const isPrivate = (a) => /^(10\.|127\.|192\.168\.|172\.(1[6-9]|2\d|3[01])\.)/.test(a);
if (coLocated && underlay) Endpoint = `${underlay}:${port}` // same LAN
else if (underlay && !isPrivate(underlay)) Endpoint = `${underlay}:${port}` // public
// else: no Endpoint — the peer must initiate, and we learn its endpoint from the handshake
```
A simulated public segment addressed out of RFC1918 space — `10.200.0.0/24`, say — makes the
hub's underlay test as **private**. No spoke writes an `Endpoint` for the hub. Nothing can
initiate. **No handshake ever occurs and the mesh silently never forms**, presenting as a
WireGuard fault rather than an addressing choice.
**The simulated public segment must therefore be `203.0.113.0/24`** — TEST-NET-3, reserved by
RFC 5737 for documentation, guaranteed never to route on the real internet, and not matched by
that regex. The code then treats it exactly as it treats a real hosting provider address.
This is the single most important fact in this document.
---
## 3. The topology being reproduced
The shape below is what a mesh of this kind looks like: one node with a routable address, one
publicly named but behind a household NAT, one stationary workstation, one that roams.
Addresses use the documentation ranges of RFC 5737 and RFC 1918 throughout.
| node | profile | site | underlay | WG | accessors |
|---|---|---|---|---|---|
| `anchor` | server | `dc` | `203.0.113.10` (routable) | `10.10.0.1/24` | `anchor.example` (public, primary) + `anchor.internal` (lan) |
| `home-server` | server | `home` | `192.168.1.135` | `10.10.0.2/24` | `home-server.example` (public, primary) + `home-server.internal` (lan) |
| `workstation` | workstation | `home` | `192.168.1.250` | `10.10.0.3/24` | `workstation.internal` (lan, primary) |
| `laptop` | workstation | `NULL` | `NULL` | `10.10.0.4/24` | `laptop.internal` (lan, primary) |
**Hub election is by convention, not by flag.** The hub is the node whose `profile='server'`
*and* whose `WG_ADDRESS` begins `10.10.0.1` (`hooks/index.ts:188`). A lab must assign
`10.10.0.1` to the node it intends as hub or there will be no hub.
**`site` drives direct peering** (`hooks/index.ts:197-203`). Two nodes with the same non-null
`site` peer directly with a `/32`; everything else routes through the hub's `/24`. A `NULL`
site means roaming and hub-only — deliberately, because WireGuard has no failover and a more
specific `/32` route to a dead endpoint blackholes rather than falling back.
The four interesting pairs, all of which the lab must reproduce:
| pair | behaviour | branch taken |
|---|---|---|
| anything → `anchor` | `Endpoint` written | underlay non-private |
| `home-server` ↔ `workstation` | direct peer, LAN endpoints, keepalive | co-located |
| **`anchor` → `home-server`** | **no `Endpoint`; learned from handshake** | not co-located, underlay private |
| `laptop` → anything | hub only, always initiates | `site` is `NULL` |
The third is the one worth building the lab for. The code comments at `hooks/index.ts:206-224`
record what it cost to get right: testing `profile === "server"` was tried and was wrong,
because a home-hosted node **is** a server yet is not publicly reachable — *"role does not
imply reachability; the address does."* An earlier version aimed the hub at that node's public
name, which hairpinned off the household NAT: 1.77 MiB sent, 0 B received, no handshake.
---
## 4. The lab
```
br-wan 203.0.113.0/24 TEST-NET-3 — non-private, so the code treats it as public
│
├── hub 203.0.113.10 profile=server site=dc WG 10.10.0.1
│
└── router VM 203.0.113.1 / 192.168.1.1
│ NAT, plus one forwarded port to reproduce a published-but-NATed node
│
br-lan 192.168.1.0/24 identical to production, same host addresses
├── a 192.168.1.135 profile=server site=home WG 10.10.0.2
└── b 192.168.1.250 profile=workstation site=home WG 10.10.0.3
c — attach to br-lan, or br-wan ("away"), or detach ("asleep")
underlay NULL profile=workstation site=NULL WG 10.10.0.4
```
Kept **byte-identical** to production: the LAN subnet and its host addresses, and the entire
WireGuard plan. Only the public segment is substituted, and only because it must be.
**The router earns its own VM.** It is what makes the published-but-NATed case real: that node
is reachable from outside only through a forwarded port, and the hub must learn its endpoint.
It also gives somewhere to break things — drop the forward and observe whether the mesh
notices or whether the public name simply stops working.
**Names.** A resolver on the wan side is authoritative for the public zone. `.internal` names
need nothing extra: `dnsmasq-app` generates them from mesh config on each node, and writes an
`/etc/hosts` block as a floor underneath, because a node must reach the mesh DB before its own
DNS exists.
---
## 5. What a node needs before any of this works
Rows in the mesh DB — `nodes` (`name`, `user_name`, `profile`, `site`, `underlay_addr`),
`node_accessors`, `module_env.WG_ADDRESS` (the hook hard-fails without it,
`hooks/index.ts:114-116`), and `node_modules` assigning at least `wireguard`, `dnsmasq-app`,
`mesh-ca`, and `traefik` where it serves.
On disk beforehand, because the node must reach the mesh DB before it can read any of the
above: registry database and object-store host and credentials, plus an npm token. This
ordering — contact the mesh before the mesh has configured you — is itself worth reproducing,
and is why the `/etc/hosts` floor exists.
Everything else is generated: the WireGuard keypair locally (the private key never leaves the
node; the public key is published to `node_wg_keys`), the peer list, the DNS records, the TLS
leaf.
---
## 6. Certificates — the lab issues its own
**Settled 2026-08-22: the lab runs its own ACME issuer.**
Public certificates use ACME **HTTP-01** via the reverse proxy, which requires genuine public
reachability, so an isolated lab cannot use the real issuer. Rather than forgo certificate
testing, the lab stands up an ACME server on its wan segment.
**The lab keeps production's two-CA split rather than collapsing it.** Production issues
public names from a public authority and internal names from the mesh CA; a lab with one CA
would hide any bug living in that split. So:
| | production | lab |
|---|---|---|
| public names | a public ACME authority | a test ACME server on the wan segment |
| `.internal` names | `mesh-ca` | `mesh-ca`, unchanged |
**A test issuer is the right shape, not a shortcut.** Purpose-built ACME test servers
deliberately vary their behaviour — validation timing, nonce handling, chain composition — to
expose assumptions a well-behaved authority would let pass. A lab CA that is *too* polite
tests less than the real thing, not more.
### It also exercises the port forward
HTTP-01 means the issuer must reach the node being certified on port 80. In the lab:
- the hub is directly reachable on the wan segment — straightforward
- **the published-but-NATed node is reachable only through the router's forwarded port**
So certificate issuance for that node passes only if the forward is correct. That is exactly
why its certificate works in production, and it makes "the forward is missing" a reproducible
failure rather than a mystery.
### Required change: `caServer` must be configurable
`modules/traefik/docker-compose.yml:17-19` sets the challenge entrypoint, the contact address
and the storage path — but **no `caServer`**, so Traefik defaults to the public authority's
*production* endpoint. Pointing the lab at its own issuer requires adding a `caServer` flag
fed by an environment value, defaulting to production so real nodes are unaffected and the lab
overrides it per node.
Worth noting independently of the lab: aiming at the production endpoint rather than a staging
one means every certificate experiment on a real node consumes production issuance quota, and
a retry loop can exhaust it for a week. The lab issuer removes that exposure.
---
## 7. Incidental finding: `scope:` is read by nothing
Several manifests declare `scope: public` on firewall rules — `wireguard`, `traefik`, `gitea`,
`mailu`, `qbittorrent`. It is **not part of the rule type** (`module-registry.ts:17-29`) and is
**referenced by no code** in the firewall path. Real scoping is done with `from:`, as
`modules/unifi/module.yml:52-93` does deliberately.
So a manifest can appear to restrict a port to the public scope and in fact restrict nothing.
This is the same shape as the rule in `00-META/how-we-build.md` — *an unenforced rule is
indistinguishable from a wrong one, and costs more, because people believe it.*