Files
hq/01-RESEARCH/004-lab-network/analysis.md
T
jschoubben 702efca6bb Base layer: the mesh as it is, under the mesh as it should be
HQ held only the to-be. Every reader had to already know the system the
decisions were about, and an as-is claim had nowhere to live except inside
an intention.

Adds 02-DESIGN/00-as-is — eleven documents written from the implementation
and the operational record, not from intent, including the parts nobody
would choose again. The two existing designs move under 01-to-be. Layers
are declared in frontmatter and never mix: a design that ships does not
move, its as-is counterpart is written, and both stand.

Back-fills adr/0001-0014 for decisions taken in implementation and never
recorded — the broker, the module abstraction, the mesh database, managed
files, provisioning, migrations, the workspace removal, failing loudly,
the constitution, application placement, linking, the employee model, the
artifact, the three silos. Each marked reconstructed, dated from the
history, and citing the evidence it was recovered from. The two existing
records renumber to 0015 and 0016 so the ledger runs oldest first;
0017 extends 0015 to modules outside the core, principle only — the
domain list is deliberately not invented here.

how-we-build.md becomes the source of the mesh constitution, with a sync
playbook, so the enforced copy stops being the only one that is true.

Process becomes explicit: five playbooks, eight thin skills that defer to
them, a repository map, and AGENTS.md with CLAUDE.md as its include.

The five Observations become 04-ISSUES 001-005 where they can be owned and
closed. 006 is new and uncomfortable: HQ is not indexed into the knowledge
base. That claim is what decision 27 rests on, it was never checked, and
the README now says so instead of repeating it.

Also corrects the ADR index into something generated, the "02-DESIGN is
empty" claim, the VISION.md pointer that did not survive the repo split,
and a note asserting the symlink rule was contradicted — it was a
misreading; the rule forbids hand-made links, the installer links by design.
2026-08-23 03:08:26 +02:00

204 lines
10 KiB
Markdown

---
effort: 004-lab-network
updated: 2026-08-22
---
# Reproducing the mesh network in a lab
Established 2026-08-22 by reading the generating code and the live mesh DB. Every claim is a
file location or a queried row.
---
## 1. The network is data, not configuration
`install.d/mesh-init.sh` and `install.d/adopt.sh` perform **no network configuration at
all** — no WireGuard, no DNS, no firewall, no `/etc/hosts`. Every part of the network layer is
generated by module hooks from mesh-DB rows:
| layer | generated by | from |
|---|---|---|
| WireGuard interface + peers | `modules/wireguard/hooks/index.ts` (`postConfigure`) | `node_wg_keys`, `nodes.site`, `nodes.underlay_addr`, `module_env.WG_ADDRESS` |
| `.internal` name resolution | `modules/dnsmasq-app/hooks/index.ts` | mesh config peers → `internal_domain` + `wg_address` |
| public routing / vhosts | `modules/hal/sdk/src/vhost-gen.ts`, `feature-handlers/vhost.ts` | `vhosts:` manifests + `node_accessors` |
| internal TLS | `modules/mesh-ca/hooks/index.ts` | a singleton CA row in the mesh DB |
**Consequence:** a faithful lab is mostly a matter of writing the right rows. The network that
results is produced by the same code production runs, which is the difference between testing
the network and testing a model of it.
---
## 2. The constraint that decides whether the lab works
`modules/wireguard/hooks/index.ts:225-240` decides, per pair, whether to write an `Endpoint`:
```js
const isPrivate = (a) => /^(10\.|127\.|192\.168\.|172\.(1[6-9]|2\d|3[01])\.)/.test(a);
if (coLocated && underlay) Endpoint = `${underlay}:${port}` // same LAN
else if (underlay && !isPrivate(underlay)) Endpoint = `${underlay}:${port}` // public
// else: no Endpoint — the peer must initiate, and we learn its endpoint from the handshake
```
A simulated public segment addressed out of RFC1918 space — `10.200.0.0/24`, say — makes the
hub's underlay test as **private**. No spoke writes an `Endpoint` for the hub. Nothing can
initiate. **No handshake ever occurs and the mesh silently never forms**, presenting as a
WireGuard fault rather than an addressing choice.
**The simulated public segment must therefore be `203.0.113.0/24`** — TEST-NET-3, reserved by
RFC 5737 for documentation, guaranteed never to route on the real internet, and not matched by
that regex. The code then treats it exactly as it treats a real hosting provider address.
This is the single most important fact in this document.
---
## 3. The topology being reproduced
The shape below is what a mesh of this kind looks like: one node with a routable address, one
publicly named but behind a household NAT, one stationary workstation, one that roams.
Addresses use the documentation ranges of RFC 5737 and RFC 1918 throughout.
| node | profile | site | underlay | WG | accessors |
|---|---|---|---|---|---|
| `anchor` | server | `dc` | `203.0.113.10` (routable) | `10.10.0.1/24` | `anchor.example` (public, primary) + `anchor.internal` (lan) |
| `home-server` | server | `home` | `192.168.1.135` | `10.10.0.2/24` | `home-server.example` (public, primary) + `home-server.internal` (lan) |
| `workstation` | workstation | `home` | `192.168.1.250` | `10.10.0.3/24` | `workstation.internal` (lan, primary) |
| `laptop` | workstation | `NULL` | `NULL` | `10.10.0.4/24` | `laptop.internal` (lan, primary) |
**Hub election is by convention, not by flag.** The hub is the node whose `profile='server'`
*and* whose `WG_ADDRESS` begins `10.10.0.1` (`hooks/index.ts:188`). A lab must assign
`10.10.0.1` to the node it intends as hub or there will be no hub.
**`site` drives direct peering** (`hooks/index.ts:197-203`). Two nodes with the same non-null
`site` peer directly with a `/32`; everything else routes through the hub's `/24`. A `NULL`
site means roaming and hub-only — deliberately, because WireGuard has no failover and a more
specific `/32` route to a dead endpoint blackholes rather than falling back.
The four interesting pairs, all of which the lab must reproduce:
| pair | behaviour | branch taken |
|---|---|---|
| anything → `anchor` | `Endpoint` written | underlay non-private |
| `home-server` ↔ `workstation` | direct peer, LAN endpoints, keepalive | co-located |
| **`anchor` → `home-server`** | **no `Endpoint`; learned from handshake** | not co-located, underlay private |
| `laptop` → anything | hub only, always initiates | `site` is `NULL` |
The third is the one worth building the lab for. The code comments at `hooks/index.ts:206-224`
record what it cost to get right: testing `profile === "server"` was tried and was wrong,
because a home-hosted node **is** a server yet is not publicly reachable — *"role does not
imply reachability; the address does."* An earlier version aimed the hub at that node's public
name, which hairpinned off the household NAT: 1.77 MiB sent, 0 B received, no handshake.
---
## 4. The lab
```
br-wan 203.0.113.0/24 TEST-NET-3 — non-private, so the code treats it as public
│
├── hub 203.0.113.10 profile=server site=dc WG 10.10.0.1
│
└── router VM 203.0.113.1 / 192.168.1.1
│ NAT, plus one forwarded port to reproduce a published-but-NATed node
│
br-lan 192.168.1.0/24 identical to production, same host addresses
├── a 192.168.1.135 profile=server site=home WG 10.10.0.2
└── b 192.168.1.250 profile=workstation site=home WG 10.10.0.3
c — attach to br-lan, or br-wan ("away"), or detach ("asleep")
underlay NULL profile=workstation site=NULL WG 10.10.0.4
```
Kept **byte-identical** to production: the LAN subnet and its host addresses, and the entire
WireGuard plan. Only the public segment is substituted, and only because it must be.
**The router earns its own VM.** It is what makes the published-but-NATed case real: that node
is reachable from outside only through a forwarded port, and the hub must learn its endpoint.
It also gives somewhere to break things — drop the forward and observe whether the mesh
notices or whether the public name simply stops working.
**Names.** A resolver on the wan side is authoritative for the public zone. `.internal` names
need nothing extra: `dnsmasq-app` generates them from mesh config on each node, and writes an
`/etc/hosts` block as a floor underneath, because a node must reach the mesh DB before its own
DNS exists.
---
## 5. What a node needs before any of this works
Rows in the mesh DB — `nodes` (`name`, `user_name`, `profile`, `site`, `underlay_addr`),
`node_accessors`, `module_env.WG_ADDRESS` (the hook hard-fails without it,
`hooks/index.ts:114-116`), and `node_modules` assigning at least `wireguard`, `dnsmasq-app`,
`mesh-ca`, and `traefik` where it serves.
On disk beforehand, because the node must reach the mesh DB before it can read any of the
above: registry database and object-store host and credentials, plus an npm token. This
ordering — contact the mesh before the mesh has configured you — is itself worth reproducing,
and is why the `/etc/hosts` floor exists.
Everything else is generated: the WireGuard keypair locally (the private key never leaves the
node; the public key is published to `node_wg_keys`), the peer list, the DNS records, the TLS
leaf.
---
## 6. Certificates — the lab issues its own
**Settled 2026-08-22: the lab runs its own ACME issuer.**
Public certificates use ACME **HTTP-01** via the reverse proxy, which requires genuine public
reachability, so an isolated lab cannot use the real issuer. Rather than forgo certificate
testing, the lab stands up an ACME server on its wan segment.
**The lab keeps production's two-CA split rather than collapsing it.** Production issues
public names from a public authority and internal names from the mesh CA; a lab with one CA
would hide any bug living in that split. So:
| | production | lab |
|---|---|---|
| public names | a public ACME authority | a test ACME server on the wan segment |
| `.internal` names | `mesh-ca` | `mesh-ca`, unchanged |
**A test issuer is the right shape, not a shortcut.** Purpose-built ACME test servers
deliberately vary their behaviour — validation timing, nonce handling, chain composition — to
expose assumptions a well-behaved authority would let pass. A lab CA that is *too* polite
tests less than the real thing, not more.
### It also exercises the port forward
HTTP-01 means the issuer must reach the node being certified on port 80. In the lab:
- the hub is directly reachable on the wan segment — straightforward
- **the published-but-NATed node is reachable only through the router's forwarded port**
So certificate issuance for that node passes only if the forward is correct. That is exactly
why its certificate works in production, and it makes "the forward is missing" a reproducible
failure rather than a mystery.
### Required change: `caServer` must be configurable
`modules/traefik/docker-compose.yml:17-19` sets the challenge entrypoint, the contact address
and the storage path — but **no `caServer`**, so Traefik defaults to the public authority's
*production* endpoint. Pointing the lab at its own issuer requires adding a `caServer` flag
fed by an environment value, defaulting to production so real nodes are unaffected and the lab
overrides it per node.
Worth noting independently of the lab: aiming at the production endpoint rather than a staging
one means every certificate experiment on a real node consumes production issuance quota, and
a retry loop can exhaust it for a week. The lab issuer removes that exposure.
---
## 7. Incidental finding: `scope:` is read by nothing
Several manifests declare `scope: public` on firewall rules — `wireguard`, `traefik`, `gitea`,
`mailu`, `qbittorrent`. It is **not part of the rule type** (`module-registry.ts:17-29`) and is
**referenced by no code** in the firewall path. Real scoping is done with `from:`, as
`modules/unifi/module.yml:52-93` does deliberately.
So a manifest can appear to restrict a port to the public scope and in fact restrict nothing.
This is the same shape as the rule in `00-GENESIS/how-we-build.md` — *an unenforced rule is
indistinguishable from a wrong one, and costs more, because people believe it.*