What the mesh is, what it is becoming, and why. Implementation lives in the code repositories; the reasoning lives here. 00-GENESIS mission, engineering context, effect, and the rules that hold 01-RESEARCH investigations, before they harden into design 02-DESIGN the authoritative specification adr numbered decisions — what was chosen, and what was rejected DECISIONS.md the ledger: every decision, in the order it was taken Written for a reader who is not its author and has no access to the mesh it describes. Addresses use the documentation ranges of RFC 5737 and RFC 1918; nodes are named by role. Single initial commit by intent. The prior history came from a private repository and carried operational detail — a routable address identified as a VPN hub, real domain names, a hosting provider — which sanitising a tip commit would not have removed from the log.
10 KiB
Reproducing the mesh network in a lab
Established 2026-08-22 by reading the generating code and the live mesh DB. Every claim is a file location or a queried row.
1. The network is data, not configuration
install.d/mesh-init.sh and install.d/adopt.sh perform no network configuration at
all — no WireGuard, no DNS, no firewall, no /etc/hosts. Every part of the network layer is
generated by module hooks from mesh-DB rows:
| layer | generated by | from |
|---|---|---|
| WireGuard interface + peers | modules/wireguard/hooks/index.ts (postConfigure) |
node_wg_keys, nodes.site, nodes.underlay_addr, module_env.WG_ADDRESS |
.internal name resolution |
modules/dnsmasq-app/hooks/index.ts |
mesh config peers → internal_domain + wg_address |
| public routing / vhosts | modules/hal/sdk/src/vhost-gen.ts, feature-handlers/vhost.ts |
vhosts: manifests + node_accessors |
| internal TLS | modules/mesh-ca/hooks/index.ts |
a singleton CA row in the mesh DB |
Consequence: a faithful lab is mostly a matter of writing the right rows. The network that results is produced by the same code production runs, which is the difference between testing the network and testing a model of it.
2. The constraint that decides whether the lab works
modules/wireguard/hooks/index.ts:225-240 decides, per pair, whether to write an Endpoint:
const isPrivate = (a) => /^(10\.|127\.|192\.168\.|172\.(1[6-9]|2\d|3[01])\.)/.test(a);
if (coLocated && underlay) Endpoint = `${underlay}:${port}` // same LAN
else if (underlay && !isPrivate(underlay)) Endpoint = `${underlay}:${port}` // public
// else: no Endpoint — the peer must initiate, and we learn its endpoint from the handshake
A simulated public segment addressed out of RFC1918 space — 10.200.0.0/24, say — makes the
hub's underlay test as private. No spoke writes an Endpoint for the hub. Nothing can
initiate. No handshake ever occurs and the mesh silently never forms, presenting as a
WireGuard fault rather than an addressing choice.
The simulated public segment must therefore be 203.0.113.0/24 — TEST-NET-3, reserved by
RFC 5737 for documentation, guaranteed never to route on the real internet, and not matched by
that regex. The code then treats it exactly as it treats a real hosting provider address.
This is the single most important fact in this document.
3. The topology being reproduced
The shape below is what a mesh of this kind looks like: one node with a routable address, one publicly named but behind a household NAT, one stationary workstation, one that roams. Addresses use the documentation ranges of RFC 5737 and RFC 1918 throughout.
| node | profile | site | underlay | WG | accessors |
|---|---|---|---|---|---|
anchor |
server | dc |
203.0.113.10 (routable) |
10.10.0.1/24 |
anchor.example (public, primary) + anchor.internal (lan) |
home-server |
server | home |
192.168.1.135 |
10.10.0.2/24 |
home-server.example (public, primary) + home-server.internal (lan) |
workstation |
workstation | home |
192.168.1.250 |
10.10.0.3/24 |
workstation.internal (lan, primary) |
laptop |
workstation | NULL |
NULL |
10.10.0.4/24 |
laptop.internal (lan, primary) |
Hub election is by convention, not by flag. The hub is the node whose profile='server'
and whose WG_ADDRESS begins 10.10.0.1 (hooks/index.ts:188). A lab must assign
10.10.0.1 to the node it intends as hub or there will be no hub.
site drives direct peering (hooks/index.ts:197-203). Two nodes with the same non-null
site peer directly with a /32; everything else routes through the hub's /24. A NULL
site means roaming and hub-only — deliberately, because WireGuard has no failover and a more
specific /32 route to a dead endpoint blackholes rather than falling back.
The four interesting pairs, all of which the lab must reproduce:
| pair | behaviour | branch taken |
|---|---|---|
anything → anchor |
Endpoint written |
underlay non-private |
home-server ↔ workstation |
direct peer, LAN endpoints, keepalive | co-located |
anchor → home-server |
no Endpoint; learned from handshake |
not co-located, underlay private |
laptop → anything |
hub only, always initiates | site is NULL |
The third is the one worth building the lab for. The code comments at hooks/index.ts:206-224
record what it cost to get right: testing profile === "server" was tried and was wrong,
because a home-hosted node is a server yet is not publicly reachable — "role does not
imply reachability; the address does." An earlier version aimed the hub at that node's public
name, which hairpinned off the household NAT: 1.77 MiB sent, 0 B received, no handshake.
4. The lab
br-wan 203.0.113.0/24 TEST-NET-3 — non-private, so the code treats it as public
│
├── hub 203.0.113.10 profile=server site=dc WG 10.10.0.1
│
└── router VM 203.0.113.1 / 192.168.1.1
│ NAT, plus one forwarded port to reproduce a published-but-NATed node
│
br-lan 192.168.1.0/24 identical to production, same host addresses
├── a 192.168.1.135 profile=server site=home WG 10.10.0.2
└── b 192.168.1.250 profile=workstation site=home WG 10.10.0.3
c — attach to br-lan, or br-wan ("away"), or detach ("asleep")
underlay NULL profile=workstation site=NULL WG 10.10.0.4
Kept byte-identical to production: the LAN subnet and its host addresses, and the entire WireGuard plan. Only the public segment is substituted, and only because it must be.
The router earns its own VM. It is what makes the published-but-NATed case real: that node is reachable from outside only through a forwarded port, and the hub must learn its endpoint. It also gives somewhere to break things — drop the forward and observe whether the mesh notices or whether the public name simply stops working.
Names. A resolver on the wan side is authoritative for the public zone. .internal names
need nothing extra: dnsmasq-app generates them from mesh config on each node, and writes an
/etc/hosts block as a floor underneath, because a node must reach the mesh DB before its own
DNS exists.
5. What a node needs before any of this works
Rows in the mesh DB — nodes (name, user_name, profile, site, underlay_addr),
node_accessors, module_env.WG_ADDRESS (the hook hard-fails without it,
hooks/index.ts:114-116), and node_modules assigning at least wireguard, dnsmasq-app,
mesh-ca, and traefik where it serves.
On disk beforehand, because the node must reach the mesh DB before it can read any of the
above: registry database and object-store host and credentials, plus an npm token. This
ordering — contact the mesh before the mesh has configured you — is itself worth reproducing,
and is why the /etc/hosts floor exists.
Everything else is generated: the WireGuard keypair locally (the private key never leaves the
node; the public key is published to node_wg_keys), the peer list, the DNS records, the TLS
leaf.
6. Certificates — the lab issues its own
Settled 2026-08-22: the lab runs its own ACME issuer.
Public certificates use ACME HTTP-01 via the reverse proxy, which requires genuine public reachability, so an isolated lab cannot use the real issuer. Rather than forgo certificate testing, the lab stands up an ACME server on its wan segment.
The lab keeps production's two-CA split rather than collapsing it. Production issues public names from a public authority and internal names from the mesh CA; a lab with one CA would hide any bug living in that split. So:
| production | lab | |
|---|---|---|
| public names | a public ACME authority | a test ACME server on the wan segment |
.internal names |
mesh-ca |
mesh-ca, unchanged |
A test issuer is the right shape, not a shortcut. Purpose-built ACME test servers deliberately vary their behaviour — validation timing, nonce handling, chain composition — to expose assumptions a well-behaved authority would let pass. A lab CA that is too polite tests less than the real thing, not more.
It also exercises the port forward
HTTP-01 means the issuer must reach the node being certified on port 80. In the lab:
- the hub is directly reachable on the wan segment — straightforward
- the published-but-NATed node is reachable only through the router's forwarded port
So certificate issuance for that node passes only if the forward is correct. That is exactly why its certificate works in production, and it makes "the forward is missing" a reproducible failure rather than a mystery.
Required change: caServer must be configurable
modules/traefik/docker-compose.yml:17-19 sets the challenge entrypoint, the contact address
and the storage path — but no caServer, so Traefik defaults to the public authority's
production endpoint. Pointing the lab at its own issuer requires adding a caServer flag
fed by an environment value, defaulting to production so real nodes are unaffected and the lab
overrides it per node.
Worth noting independently of the lab: aiming at the production endpoint rather than a staging one means every certificate experiment on a real node consumes production issuance quota, and a retry loop can exhaust it for a week. The lab issuer removes that exposure.
7. Incidental finding: scope: is read by nothing
Several manifests declare scope: public on firewall rules — wireguard, traefik, gitea,
mailu, qbittorrent. It is not part of the rule type (module-registry.ts:17-29) and is
referenced by no code in the firewall path. Real scoping is done with from:, as
modules/unifi/module.yml:52-93 does deliberately.
So a manifest can appear to restrict a port to the public scope and in fact restrict nothing.
This is the same shape as the rule in 00-GENESIS/how-we-build.md — an unenforced rule is
indistinguishable from a wrong one, and costs more, because people believe it.