Files
hq/01-RESEARCH/004-lab-network/analysis.md
T
jschoubben c0b35652d0 The numbering is the flow: decisions are 02, design is 03
papa-hq reads 01 research -> 03 decision -> 02 design. The order is a
scar, not a choice: 02-DESIGN existed from its initial commit, and when
adr/ was finally promoted on 2026-07-13 it took the next free number
rather than its place in the sequence. By then design was too settled to
renumber.

hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and
02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks
the process in the order it happens: research produces a decision, the
decision authorises a design.

00-GENESIS becomes 00-META, matching papa's rename from the same
restructure.

Every path reference rewritten across documents, frontmatter, playbooks
and skills. All links resolve; all 58 frontmatter blocks parse and their
path fields still point at files that exist.
2026-08-23 18:05:11 +02:00

10 KiB

effort, updated
effort updated
004-lab-network 2026-08-22

Reproducing the mesh network in a lab

Established 2026-08-22 by reading the generating code and the live mesh DB. Every claim is a file location or a queried row.


1. The network is data, not configuration

install.d/mesh-init.sh and install.d/adopt.sh perform no network configuration at all — no WireGuard, no DNS, no firewall, no /etc/hosts. Every part of the network layer is generated by module hooks from mesh-DB rows:

layer generated by from
WireGuard interface + peers modules/wireguard/hooks/index.ts (postConfigure) node_wg_keys, nodes.site, nodes.underlay_addr, module_env.WG_ADDRESS
.internal name resolution modules/dnsmasq-app/hooks/index.ts mesh config peers → internal_domain + wg_address
public routing / vhosts modules/hal/sdk/src/vhost-gen.ts, feature-handlers/vhost.ts vhosts: manifests + node_accessors
internal TLS modules/mesh-ca/hooks/index.ts a singleton CA row in the mesh DB

Consequence: a faithful lab is mostly a matter of writing the right rows. The network that results is produced by the same code production runs, which is the difference between testing the network and testing a model of it.


2. The constraint that decides whether the lab works

modules/wireguard/hooks/index.ts:225-240 decides, per pair, whether to write an Endpoint:

const isPrivate = (a) => /^(10\.|127\.|192\.168\.|172\.(1[6-9]|2\d|3[01])\.)/.test(a);

if (coLocated && underlay)                    Endpoint = `${underlay}:${port}`   // same LAN
else if (underlay && !isPrivate(underlay))    Endpoint = `${underlay}:${port}`   // public
// else: no Endpoint — the peer must initiate, and we learn its endpoint from the handshake

A simulated public segment addressed out of RFC1918 space — 10.200.0.0/24, say — makes the hub's underlay test as private. No spoke writes an Endpoint for the hub. Nothing can initiate. No handshake ever occurs and the mesh silently never forms, presenting as a WireGuard fault rather than an addressing choice.

The simulated public segment must therefore be 203.0.113.0/24 — TEST-NET-3, reserved by RFC 5737 for documentation, guaranteed never to route on the real internet, and not matched by that regex. The code then treats it exactly as it treats a real hosting provider address.

This is the single most important fact in this document.


3. The topology being reproduced

The shape below is what a mesh of this kind looks like: one node with a routable address, one publicly named but behind a household NAT, one stationary workstation, one that roams. Addresses use the documentation ranges of RFC 5737 and RFC 1918 throughout.

node profile site underlay WG accessors
anchor server dc 203.0.113.10 (routable) 10.10.0.1/24 anchor.example (public, primary) + anchor.internal (lan)
home-server server home 192.168.1.135 10.10.0.2/24 home-server.example (public, primary) + home-server.internal (lan)
workstation workstation home 192.168.1.250 10.10.0.3/24 workstation.internal (lan, primary)
laptop workstation NULL NULL 10.10.0.4/24 laptop.internal (lan, primary)

Hub election is by convention, not by flag. The hub is the node whose profile='server' and whose WG_ADDRESS begins 10.10.0.1 (hooks/index.ts:188). A lab must assign 10.10.0.1 to the node it intends as hub or there will be no hub.

site drives direct peering (hooks/index.ts:197-203). Two nodes with the same non-null site peer directly with a /32; everything else routes through the hub's /24. A NULL site means roaming and hub-only — deliberately, because WireGuard has no failover and a more specific /32 route to a dead endpoint blackholes rather than falling back.

The four interesting pairs, all of which the lab must reproduce:

pair behaviour branch taken
anything → anchor Endpoint written underlay non-private
home-server ↔ workstation direct peer, LAN endpoints, keepalive co-located
anchor → home-server no Endpoint; learned from handshake not co-located, underlay private
laptop → anything hub only, always initiates site is NULL

The third is the one worth building the lab for. The code comments at hooks/index.ts:206-224 record what it cost to get right: testing profile === "server" was tried and was wrong, because a home-hosted node is a server yet is not publicly reachable — "role does not imply reachability; the address does." An earlier version aimed the hub at that node's public name, which hairpinned off the household NAT: 1.77 MiB sent, 0 B received, no handshake.


4. The lab

 br-wan   203.0.113.0/24        TEST-NET-3 — non-private, so the code treats it as public
    │
    ├── hub          203.0.113.10        profile=server  site=dc      WG 10.10.0.1
    │
    └── router VM    203.0.113.1 / 192.168.1.1
           │          NAT, plus one forwarded port to reproduce a published-but-NATed node
           │
      br-lan  192.168.1.0/24     identical to production, same host addresses
           ├── a     192.168.1.135       profile=server  site=home    WG 10.10.0.2
           └── b     192.168.1.250       profile=workstation site=home WG 10.10.0.3

      c — attach to br-lan, or br-wan ("away"), or detach ("asleep")
                     underlay NULL       profile=workstation site=NULL WG 10.10.0.4

Kept byte-identical to production: the LAN subnet and its host addresses, and the entire WireGuard plan. Only the public segment is substituted, and only because it must be.

The router earns its own VM. It is what makes the published-but-NATed case real: that node is reachable from outside only through a forwarded port, and the hub must learn its endpoint. It also gives somewhere to break things — drop the forward and observe whether the mesh notices or whether the public name simply stops working.

Names. A resolver on the wan side is authoritative for the public zone. .internal names need nothing extra: dnsmasq-app generates them from mesh config on each node, and writes an /etc/hosts block as a floor underneath, because a node must reach the mesh DB before its own DNS exists.


5. What a node needs before any of this works

Rows in the mesh DB — nodes (name, user_name, profile, site, underlay_addr), node_accessors, module_env.WG_ADDRESS (the hook hard-fails without it, hooks/index.ts:114-116), and node_modules assigning at least wireguard, dnsmasq-app, mesh-ca, and traefik where it serves.

On disk beforehand, because the node must reach the mesh DB before it can read any of the above: registry database and object-store host and credentials, plus an npm token. This ordering — contact the mesh before the mesh has configured you — is itself worth reproducing, and is why the /etc/hosts floor exists.

Everything else is generated: the WireGuard keypair locally (the private key never leaves the node; the public key is published to node_wg_keys), the peer list, the DNS records, the TLS leaf.


6. Certificates — the lab issues its own

Settled 2026-08-22: the lab runs its own ACME issuer.

Public certificates use ACME HTTP-01 via the reverse proxy, which requires genuine public reachability, so an isolated lab cannot use the real issuer. Rather than forgo certificate testing, the lab stands up an ACME server on its wan segment.

The lab keeps production's two-CA split rather than collapsing it. Production issues public names from a public authority and internal names from the mesh CA; a lab with one CA would hide any bug living in that split. So:

production lab
public names a public ACME authority a test ACME server on the wan segment
.internal names mesh-ca mesh-ca, unchanged

A test issuer is the right shape, not a shortcut. Purpose-built ACME test servers deliberately vary their behaviour — validation timing, nonce handling, chain composition — to expose assumptions a well-behaved authority would let pass. A lab CA that is too polite tests less than the real thing, not more.

It also exercises the port forward

HTTP-01 means the issuer must reach the node being certified on port 80. In the lab:

  • the hub is directly reachable on the wan segment — straightforward
  • the published-but-NATed node is reachable only through the router's forwarded port

So certificate issuance for that node passes only if the forward is correct. That is exactly why its certificate works in production, and it makes "the forward is missing" a reproducible failure rather than a mystery.

Required change: caServer must be configurable

modules/traefik/docker-compose.yml:17-19 sets the challenge entrypoint, the contact address and the storage path — but no caServer, so Traefik defaults to the public authority's production endpoint. Pointing the lab at its own issuer requires adding a caServer flag fed by an environment value, defaulting to production so real nodes are unaffected and the lab overrides it per node.

Worth noting independently of the lab: aiming at the production endpoint rather than a staging one means every certificate experiment on a real node consumes production issuance quota, and a retry loop can exhaust it for a week. The lab issuer removes that exposure.


7. Incidental finding: scope: is read by nothing

Several manifests declare scope: public on firewall rules — wireguard, traefik, gitea, mailu, qbittorrent. It is not part of the rule type (module-registry.ts:17-29) and is referenced by no code in the firewall path. Real scoping is done with from:, as modules/unifi/module.yml:52-93 does deliberately.

So a manifest can appear to restrict a port to the public scope and in fact restrict nothing. This is the same shape as the rule in 00-META/how-we-build.md — an unenforced rule is indistinguishable from a wrong one, and costs more, because people believe it.