HQ held only the to-be. Every reader had to already know the system the decisions were about, and an as-is claim had nowhere to live except inside an intention. Adds 02-DESIGN/00-as-is — eleven documents written from the implementation and the operational record, not from intent, including the parts nobody would choose again. The two existing designs move under 01-to-be. Layers are declared in frontmatter and never mix: a design that ships does not move, its as-is counterpart is written, and both stand. Back-fills adr/0001-0014 for decisions taken in implementation and never recorded — the broker, the module abstraction, the mesh database, managed files, provisioning, migrations, the workspace removal, failing loudly, the constitution, application placement, linking, the employee model, the artifact, the three silos. Each marked reconstructed, dated from the history, and citing the evidence it was recovered from. The two existing records renumber to 0015 and 0016 so the ledger runs oldest first; 0017 extends 0015 to modules outside the core, principle only — the domain list is deliberately not invented here. how-we-build.md becomes the source of the mesh constitution, with a sync playbook, so the enforced copy stops being the only one that is true. Process becomes explicit: five playbooks, eight thin skills that defer to them, a repository map, and AGENTS.md with CLAUDE.md as its include. The five Observations become 04-ISSUES 001-005 where they can be owned and closed. 006 is new and uncomfortable: HQ is not indexed into the knowledge base. That claim is what decision 27 rests on, it was never checked, and the README now says so instead of repeating it. Also corrects the ADR index into something generated, the "02-DESIGN is empty" claim, the VISION.md pointer that did not survive the repo split, and a note asserting the symlink rule was contradicted — it was a misreading; the rule forbids hand-made links, the installer links by design.
204 lines
10 KiB
Markdown
204 lines
10 KiB
Markdown
---
|
|
effort: 004-lab-network
|
|
updated: 2026-08-22
|
|
---
|
|
|
|
# Reproducing the mesh network in a lab
|
|
|
|
Established 2026-08-22 by reading the generating code and the live mesh DB. Every claim is a
|
|
file location or a queried row.
|
|
|
|
---
|
|
|
|
## 1. The network is data, not configuration
|
|
|
|
`install.d/mesh-init.sh` and `install.d/adopt.sh` perform **no network configuration at
|
|
all** — no WireGuard, no DNS, no firewall, no `/etc/hosts`. Every part of the network layer is
|
|
generated by module hooks from mesh-DB rows:
|
|
|
|
| layer | generated by | from |
|
|
|---|---|---|
|
|
| WireGuard interface + peers | `modules/wireguard/hooks/index.ts` (`postConfigure`) | `node_wg_keys`, `nodes.site`, `nodes.underlay_addr`, `module_env.WG_ADDRESS` |
|
|
| `.internal` name resolution | `modules/dnsmasq-app/hooks/index.ts` | mesh config peers → `internal_domain` + `wg_address` |
|
|
| public routing / vhosts | `modules/hal/sdk/src/vhost-gen.ts`, `feature-handlers/vhost.ts` | `vhosts:` manifests + `node_accessors` |
|
|
| internal TLS | `modules/mesh-ca/hooks/index.ts` | a singleton CA row in the mesh DB |
|
|
|
|
**Consequence:** a faithful lab is mostly a matter of writing the right rows. The network that
|
|
results is produced by the same code production runs, which is the difference between testing
|
|
the network and testing a model of it.
|
|
|
|
---
|
|
|
|
## 2. The constraint that decides whether the lab works
|
|
|
|
`modules/wireguard/hooks/index.ts:225-240` decides, per pair, whether to write an `Endpoint`:
|
|
|
|
```js
|
|
const isPrivate = (a) => /^(10\.|127\.|192\.168\.|172\.(1[6-9]|2\d|3[01])\.)/.test(a);
|
|
|
|
if (coLocated && underlay) Endpoint = `${underlay}:${port}` // same LAN
|
|
else if (underlay && !isPrivate(underlay)) Endpoint = `${underlay}:${port}` // public
|
|
// else: no Endpoint — the peer must initiate, and we learn its endpoint from the handshake
|
|
```
|
|
|
|
A simulated public segment addressed out of RFC1918 space — `10.200.0.0/24`, say — makes the
|
|
hub's underlay test as **private**. No spoke writes an `Endpoint` for the hub. Nothing can
|
|
initiate. **No handshake ever occurs and the mesh silently never forms**, presenting as a
|
|
WireGuard fault rather than an addressing choice.
|
|
|
|
**The simulated public segment must therefore be `203.0.113.0/24`** — TEST-NET-3, reserved by
|
|
RFC 5737 for documentation, guaranteed never to route on the real internet, and not matched by
|
|
that regex. The code then treats it exactly as it treats a real hosting provider address.
|
|
|
|
This is the single most important fact in this document.
|
|
|
|
---
|
|
|
|
## 3. The topology being reproduced
|
|
|
|
The shape below is what a mesh of this kind looks like: one node with a routable address, one
|
|
publicly named but behind a household NAT, one stationary workstation, one that roams.
|
|
Addresses use the documentation ranges of RFC 5737 and RFC 1918 throughout.
|
|
|
|
| node | profile | site | underlay | WG | accessors |
|
|
|---|---|---|---|---|---|
|
|
| `anchor` | server | `dc` | `203.0.113.10` (routable) | `10.10.0.1/24` | `anchor.example` (public, primary) + `anchor.internal` (lan) |
|
|
| `home-server` | server | `home` | `192.168.1.135` | `10.10.0.2/24` | `home-server.example` (public, primary) + `home-server.internal` (lan) |
|
|
| `workstation` | workstation | `home` | `192.168.1.250` | `10.10.0.3/24` | `workstation.internal` (lan, primary) |
|
|
| `laptop` | workstation | `NULL` | `NULL` | `10.10.0.4/24` | `laptop.internal` (lan, primary) |
|
|
|
|
**Hub election is by convention, not by flag.** The hub is the node whose `profile='server'`
|
|
*and* whose `WG_ADDRESS` begins `10.10.0.1` (`hooks/index.ts:188`). A lab must assign
|
|
`10.10.0.1` to the node it intends as hub or there will be no hub.
|
|
|
|
**`site` drives direct peering** (`hooks/index.ts:197-203`). Two nodes with the same non-null
|
|
`site` peer directly with a `/32`; everything else routes through the hub's `/24`. A `NULL`
|
|
site means roaming and hub-only — deliberately, because WireGuard has no failover and a more
|
|
specific `/32` route to a dead endpoint blackholes rather than falling back.
|
|
|
|
The four interesting pairs, all of which the lab must reproduce:
|
|
|
|
| pair | behaviour | branch taken |
|
|
|---|---|---|
|
|
| anything → `anchor` | `Endpoint` written | underlay non-private |
|
|
| `home-server` ↔ `workstation` | direct peer, LAN endpoints, keepalive | co-located |
|
|
| **`anchor` → `home-server`** | **no `Endpoint`; learned from handshake** | not co-located, underlay private |
|
|
| `laptop` → anything | hub only, always initiates | `site` is `NULL` |
|
|
|
|
The third is the one worth building the lab for. The code comments at `hooks/index.ts:206-224`
|
|
record what it cost to get right: testing `profile === "server"` was tried and was wrong,
|
|
because a home-hosted node **is** a server yet is not publicly reachable — *"role does not
|
|
imply reachability; the address does."* An earlier version aimed the hub at that node's public
|
|
name, which hairpinned off the household NAT: 1.77 MiB sent, 0 B received, no handshake.
|
|
|
|
---
|
|
|
|
## 4. The lab
|
|
|
|
```
|
|
br-wan 203.0.113.0/24 TEST-NET-3 — non-private, so the code treats it as public
|
|
│
|
|
├── hub 203.0.113.10 profile=server site=dc WG 10.10.0.1
|
|
│
|
|
└── router VM 203.0.113.1 / 192.168.1.1
|
|
│ NAT, plus one forwarded port to reproduce a published-but-NATed node
|
|
│
|
|
br-lan 192.168.1.0/24 identical to production, same host addresses
|
|
├── a 192.168.1.135 profile=server site=home WG 10.10.0.2
|
|
└── b 192.168.1.250 profile=workstation site=home WG 10.10.0.3
|
|
|
|
c — attach to br-lan, or br-wan ("away"), or detach ("asleep")
|
|
underlay NULL profile=workstation site=NULL WG 10.10.0.4
|
|
```
|
|
|
|
Kept **byte-identical** to production: the LAN subnet and its host addresses, and the entire
|
|
WireGuard plan. Only the public segment is substituted, and only because it must be.
|
|
|
|
**The router earns its own VM.** It is what makes the published-but-NATed case real: that node
|
|
is reachable from outside only through a forwarded port, and the hub must learn its endpoint.
|
|
It also gives somewhere to break things — drop the forward and observe whether the mesh
|
|
notices or whether the public name simply stops working.
|
|
|
|
**Names.** A resolver on the wan side is authoritative for the public zone. `.internal` names
|
|
need nothing extra: `dnsmasq-app` generates them from mesh config on each node, and writes an
|
|
`/etc/hosts` block as a floor underneath, because a node must reach the mesh DB before its own
|
|
DNS exists.
|
|
|
|
---
|
|
|
|
## 5. What a node needs before any of this works
|
|
|
|
Rows in the mesh DB — `nodes` (`name`, `user_name`, `profile`, `site`, `underlay_addr`),
|
|
`node_accessors`, `module_env.WG_ADDRESS` (the hook hard-fails without it,
|
|
`hooks/index.ts:114-116`), and `node_modules` assigning at least `wireguard`, `dnsmasq-app`,
|
|
`mesh-ca`, and `traefik` where it serves.
|
|
|
|
On disk beforehand, because the node must reach the mesh DB before it can read any of the
|
|
above: registry database and object-store host and credentials, plus an npm token. This
|
|
ordering — contact the mesh before the mesh has configured you — is itself worth reproducing,
|
|
and is why the `/etc/hosts` floor exists.
|
|
|
|
Everything else is generated: the WireGuard keypair locally (the private key never leaves the
|
|
node; the public key is published to `node_wg_keys`), the peer list, the DNS records, the TLS
|
|
leaf.
|
|
|
|
---
|
|
|
|
## 6. Certificates — the lab issues its own
|
|
|
|
**Settled 2026-08-22: the lab runs its own ACME issuer.**
|
|
|
|
Public certificates use ACME **HTTP-01** via the reverse proxy, which requires genuine public
|
|
reachability, so an isolated lab cannot use the real issuer. Rather than forgo certificate
|
|
testing, the lab stands up an ACME server on its wan segment.
|
|
|
|
**The lab keeps production's two-CA split rather than collapsing it.** Production issues
|
|
public names from a public authority and internal names from the mesh CA; a lab with one CA
|
|
would hide any bug living in that split. So:
|
|
|
|
| | production | lab |
|
|
|---|---|---|
|
|
| public names | a public ACME authority | a test ACME server on the wan segment |
|
|
| `.internal` names | `mesh-ca` | `mesh-ca`, unchanged |
|
|
|
|
**A test issuer is the right shape, not a shortcut.** Purpose-built ACME test servers
|
|
deliberately vary their behaviour — validation timing, nonce handling, chain composition — to
|
|
expose assumptions a well-behaved authority would let pass. A lab CA that is *too* polite
|
|
tests less than the real thing, not more.
|
|
|
|
### It also exercises the port forward
|
|
|
|
HTTP-01 means the issuer must reach the node being certified on port 80. In the lab:
|
|
|
|
- the hub is directly reachable on the wan segment — straightforward
|
|
- **the published-but-NATed node is reachable only through the router's forwarded port**
|
|
|
|
So certificate issuance for that node passes only if the forward is correct. That is exactly
|
|
why its certificate works in production, and it makes "the forward is missing" a reproducible
|
|
failure rather than a mystery.
|
|
|
|
### Required change: `caServer` must be configurable
|
|
|
|
`modules/traefik/docker-compose.yml:17-19` sets the challenge entrypoint, the contact address
|
|
and the storage path — but **no `caServer`**, so Traefik defaults to the public authority's
|
|
*production* endpoint. Pointing the lab at its own issuer requires adding a `caServer` flag
|
|
fed by an environment value, defaulting to production so real nodes are unaffected and the lab
|
|
overrides it per node.
|
|
|
|
Worth noting independently of the lab: aiming at the production endpoint rather than a staging
|
|
one means every certificate experiment on a real node consumes production issuance quota, and
|
|
a retry loop can exhaust it for a week. The lab issuer removes that exposure.
|
|
|
|
---
|
|
|
|
## 7. Incidental finding: `scope:` is read by nothing
|
|
|
|
Several manifests declare `scope: public` on firewall rules — `wireguard`, `traefik`, `gitea`,
|
|
`mailu`, `qbittorrent`. It is **not part of the rule type** (`module-registry.ts:17-29`) and is
|
|
**referenced by no code** in the firewall path. Real scoping is done with `from:`, as
|
|
`modules/unifi/module.yml:52-93` does deliberately.
|
|
|
|
So a manifest can appear to restrict a port to the public scope and in fact restrict nothing.
|
|
This is the same shape as the rule in `00-GENESIS/how-we-build.md` — *an unenforced rule is
|
|
indistinguishable from a wrong one, and costs more, because people believe it.*
|