papa-hq reads 01 research -> 03 decision -> 02 design. The order is a scar, not a choice: 02-DESIGN existed from its initial commit, and when adr/ was finally promoted on 2026-07-13 it took the next free number rather than its place in the sequence. By then design was too settled to renumber. hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and 02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks the process in the order it happens: research produces a decision, the decision authorises a design. 00-GENESIS becomes 00-META, matching papa's rename from the same restructure. Every path reference rewritten across documents, frontmatter, playbooks and skills. All links resolve; all 58 frontmatter blocks parse and their path fields still point at files that exist.
204 lines
10 KiB
Markdown
204 lines
10 KiB
Markdown
---
|
|
effort: 004-lab-network
|
|
updated: 2026-08-22
|
|
---
|
|
|
|
# Reproducing the mesh network in a lab
|
|
|
|
Established 2026-08-22 by reading the generating code and the live mesh DB. Every claim is a
|
|
file location or a queried row.
|
|
|
|
---
|
|
|
|
## 1. The network is data, not configuration
|
|
|
|
`install.d/mesh-init.sh` and `install.d/adopt.sh` perform **no network configuration at
|
|
all** — no WireGuard, no DNS, no firewall, no `/etc/hosts`. Every part of the network layer is
|
|
generated by module hooks from mesh-DB rows:
|
|
|
|
| layer | generated by | from |
|
|
|---|---|---|
|
|
| WireGuard interface + peers | `modules/wireguard/hooks/index.ts` (`postConfigure`) | `node_wg_keys`, `nodes.site`, `nodes.underlay_addr`, `module_env.WG_ADDRESS` |
|
|
| `.internal` name resolution | `modules/dnsmasq-app/hooks/index.ts` | mesh config peers → `internal_domain` + `wg_address` |
|
|
| public routing / vhosts | `modules/hal/sdk/src/vhost-gen.ts`, `feature-handlers/vhost.ts` | `vhosts:` manifests + `node_accessors` |
|
|
| internal TLS | `modules/mesh-ca/hooks/index.ts` | a singleton CA row in the mesh DB |
|
|
|
|
**Consequence:** a faithful lab is mostly a matter of writing the right rows. The network that
|
|
results is produced by the same code production runs, which is the difference between testing
|
|
the network and testing a model of it.
|
|
|
|
---
|
|
|
|
## 2. The constraint that decides whether the lab works
|
|
|
|
`modules/wireguard/hooks/index.ts:225-240` decides, per pair, whether to write an `Endpoint`:
|
|
|
|
```js
|
|
const isPrivate = (a) => /^(10\.|127\.|192\.168\.|172\.(1[6-9]|2\d|3[01])\.)/.test(a);
|
|
|
|
if (coLocated && underlay) Endpoint = `${underlay}:${port}` // same LAN
|
|
else if (underlay && !isPrivate(underlay)) Endpoint = `${underlay}:${port}` // public
|
|
// else: no Endpoint — the peer must initiate, and we learn its endpoint from the handshake
|
|
```
|
|
|
|
A simulated public segment addressed out of RFC1918 space — `10.200.0.0/24`, say — makes the
|
|
hub's underlay test as **private**. No spoke writes an `Endpoint` for the hub. Nothing can
|
|
initiate. **No handshake ever occurs and the mesh silently never forms**, presenting as a
|
|
WireGuard fault rather than an addressing choice.
|
|
|
|
**The simulated public segment must therefore be `203.0.113.0/24`** — TEST-NET-3, reserved by
|
|
RFC 5737 for documentation, guaranteed never to route on the real internet, and not matched by
|
|
that regex. The code then treats it exactly as it treats a real hosting provider address.
|
|
|
|
This is the single most important fact in this document.
|
|
|
|
---
|
|
|
|
## 3. The topology being reproduced
|
|
|
|
The shape below is what a mesh of this kind looks like: one node with a routable address, one
|
|
publicly named but behind a household NAT, one stationary workstation, one that roams.
|
|
Addresses use the documentation ranges of RFC 5737 and RFC 1918 throughout.
|
|
|
|
| node | profile | site | underlay | WG | accessors |
|
|
|---|---|---|---|---|---|
|
|
| `anchor` | server | `dc` | `203.0.113.10` (routable) | `10.10.0.1/24` | `anchor.example` (public, primary) + `anchor.internal` (lan) |
|
|
| `home-server` | server | `home` | `192.168.1.135` | `10.10.0.2/24` | `home-server.example` (public, primary) + `home-server.internal` (lan) |
|
|
| `workstation` | workstation | `home` | `192.168.1.250` | `10.10.0.3/24` | `workstation.internal` (lan, primary) |
|
|
| `laptop` | workstation | `NULL` | `NULL` | `10.10.0.4/24` | `laptop.internal` (lan, primary) |
|
|
|
|
**Hub election is by convention, not by flag.** The hub is the node whose `profile='server'`
|
|
*and* whose `WG_ADDRESS` begins `10.10.0.1` (`hooks/index.ts:188`). A lab must assign
|
|
`10.10.0.1` to the node it intends as hub or there will be no hub.
|
|
|
|
**`site` drives direct peering** (`hooks/index.ts:197-203`). Two nodes with the same non-null
|
|
`site` peer directly with a `/32`; everything else routes through the hub's `/24`. A `NULL`
|
|
site means roaming and hub-only — deliberately, because WireGuard has no failover and a more
|
|
specific `/32` route to a dead endpoint blackholes rather than falling back.
|
|
|
|
The four interesting pairs, all of which the lab must reproduce:
|
|
|
|
| pair | behaviour | branch taken |
|
|
|---|---|---|
|
|
| anything → `anchor` | `Endpoint` written | underlay non-private |
|
|
| `home-server` ↔ `workstation` | direct peer, LAN endpoints, keepalive | co-located |
|
|
| **`anchor` → `home-server`** | **no `Endpoint`; learned from handshake** | not co-located, underlay private |
|
|
| `laptop` → anything | hub only, always initiates | `site` is `NULL` |
|
|
|
|
The third is the one worth building the lab for. The code comments at `hooks/index.ts:206-224`
|
|
record what it cost to get right: testing `profile === "server"` was tried and was wrong,
|
|
because a home-hosted node **is** a server yet is not publicly reachable — *"role does not
|
|
imply reachability; the address does."* An earlier version aimed the hub at that node's public
|
|
name, which hairpinned off the household NAT: 1.77 MiB sent, 0 B received, no handshake.
|
|
|
|
---
|
|
|
|
## 4. The lab
|
|
|
|
```
|
|
br-wan 203.0.113.0/24 TEST-NET-3 — non-private, so the code treats it as public
|
|
│
|
|
├── hub 203.0.113.10 profile=server site=dc WG 10.10.0.1
|
|
│
|
|
└── router VM 203.0.113.1 / 192.168.1.1
|
|
│ NAT, plus one forwarded port to reproduce a published-but-NATed node
|
|
│
|
|
br-lan 192.168.1.0/24 identical to production, same host addresses
|
|
├── a 192.168.1.135 profile=server site=home WG 10.10.0.2
|
|
└── b 192.168.1.250 profile=workstation site=home WG 10.10.0.3
|
|
|
|
c — attach to br-lan, or br-wan ("away"), or detach ("asleep")
|
|
underlay NULL profile=workstation site=NULL WG 10.10.0.4
|
|
```
|
|
|
|
Kept **byte-identical** to production: the LAN subnet and its host addresses, and the entire
|
|
WireGuard plan. Only the public segment is substituted, and only because it must be.
|
|
|
|
**The router earns its own VM.** It is what makes the published-but-NATed case real: that node
|
|
is reachable from outside only through a forwarded port, and the hub must learn its endpoint.
|
|
It also gives somewhere to break things — drop the forward and observe whether the mesh
|
|
notices or whether the public name simply stops working.
|
|
|
|
**Names.** A resolver on the wan side is authoritative for the public zone. `.internal` names
|
|
need nothing extra: `dnsmasq-app` generates them from mesh config on each node, and writes an
|
|
`/etc/hosts` block as a floor underneath, because a node must reach the mesh DB before its own
|
|
DNS exists.
|
|
|
|
---
|
|
|
|
## 5. What a node needs before any of this works
|
|
|
|
Rows in the mesh DB — `nodes` (`name`, `user_name`, `profile`, `site`, `underlay_addr`),
|
|
`node_accessors`, `module_env.WG_ADDRESS` (the hook hard-fails without it,
|
|
`hooks/index.ts:114-116`), and `node_modules` assigning at least `wireguard`, `dnsmasq-app`,
|
|
`mesh-ca`, and `traefik` where it serves.
|
|
|
|
On disk beforehand, because the node must reach the mesh DB before it can read any of the
|
|
above: registry database and object-store host and credentials, plus an npm token. This
|
|
ordering — contact the mesh before the mesh has configured you — is itself worth reproducing,
|
|
and is why the `/etc/hosts` floor exists.
|
|
|
|
Everything else is generated: the WireGuard keypair locally (the private key never leaves the
|
|
node; the public key is published to `node_wg_keys`), the peer list, the DNS records, the TLS
|
|
leaf.
|
|
|
|
---
|
|
|
|
## 6. Certificates — the lab issues its own
|
|
|
|
**Settled 2026-08-22: the lab runs its own ACME issuer.**
|
|
|
|
Public certificates use ACME **HTTP-01** via the reverse proxy, which requires genuine public
|
|
reachability, so an isolated lab cannot use the real issuer. Rather than forgo certificate
|
|
testing, the lab stands up an ACME server on its wan segment.
|
|
|
|
**The lab keeps production's two-CA split rather than collapsing it.** Production issues
|
|
public names from a public authority and internal names from the mesh CA; a lab with one CA
|
|
would hide any bug living in that split. So:
|
|
|
|
| | production | lab |
|
|
|---|---|---|
|
|
| public names | a public ACME authority | a test ACME server on the wan segment |
|
|
| `.internal` names | `mesh-ca` | `mesh-ca`, unchanged |
|
|
|
|
**A test issuer is the right shape, not a shortcut.** Purpose-built ACME test servers
|
|
deliberately vary their behaviour — validation timing, nonce handling, chain composition — to
|
|
expose assumptions a well-behaved authority would let pass. A lab CA that is *too* polite
|
|
tests less than the real thing, not more.
|
|
|
|
### It also exercises the port forward
|
|
|
|
HTTP-01 means the issuer must reach the node being certified on port 80. In the lab:
|
|
|
|
- the hub is directly reachable on the wan segment — straightforward
|
|
- **the published-but-NATed node is reachable only through the router's forwarded port**
|
|
|
|
So certificate issuance for that node passes only if the forward is correct. That is exactly
|
|
why its certificate works in production, and it makes "the forward is missing" a reproducible
|
|
failure rather than a mystery.
|
|
|
|
### Required change: `caServer` must be configurable
|
|
|
|
`modules/traefik/docker-compose.yml:17-19` sets the challenge entrypoint, the contact address
|
|
and the storage path — but **no `caServer`**, so Traefik defaults to the public authority's
|
|
*production* endpoint. Pointing the lab at its own issuer requires adding a `caServer` flag
|
|
fed by an environment value, defaulting to production so real nodes are unaffected and the lab
|
|
overrides it per node.
|
|
|
|
Worth noting independently of the lab: aiming at the production endpoint rather than a staging
|
|
one means every certificate experiment on a real node consumes production issuance quota, and
|
|
a retry loop can exhaust it for a week. The lab issuer removes that exposure.
|
|
|
|
---
|
|
|
|
## 7. Incidental finding: `scope:` is read by nothing
|
|
|
|
Several manifests declare `scope: public` on firewall rules — `wireguard`, `traefik`, `gitea`,
|
|
`mailu`, `qbittorrent`. It is **not part of the rule type** (`module-registry.ts:17-29`) and is
|
|
**referenced by no code** in the firewall path. Real scoping is done with `from:`, as
|
|
`modules/unifi/module.yml:52-93` does deliberately.
|
|
|
|
So a manifest can appear to restrict a port to the public scope and in fact restrict nothing.
|
|
This is the same shape as the rule in `00-META/how-we-build.md` — *an unenforced rule is
|
|
indistinguishable from a wrong one, and costs more, because people believe it.*
|