HQ — the mesh's own documentation
What the mesh is, what it is becoming, and why. Implementation lives in the code repositories; the reasoning lives here. 00-GENESIS mission, engineering context, effect, and the rules that hold 01-RESEARCH investigations, before they harden into design 02-DESIGN the authoritative specification adr numbered decisions — what was chosen, and what was rejected DECISIONS.md the ledger: every decision, in the order it was taken Written for a reader who is not its author and has no access to the mesh it describes. Addresses use the documentation ranges of RFC 5737 and RFC 1918; nodes are named by role. Single initial commit by intent. The prior history came from a private repository and carried operational detail — a routable address identified as a VPN hub, real domain names, a hosting provider — which sanitising a tip commit would not have removed from the log.
This commit is contained in:
@@ -0,0 +1,198 @@
|
||||
# Reproducing the mesh network in a lab
|
||||
|
||||
Established 2026-08-22 by reading the generating code and the live mesh DB. Every claim is a
|
||||
file location or a queried row.
|
||||
|
||||
---
|
||||
|
||||
## 1. The network is data, not configuration
|
||||
|
||||
`install.d/mesh-init.sh` and `install.d/adopt.sh` perform **no network configuration at
|
||||
all** — no WireGuard, no DNS, no firewall, no `/etc/hosts`. Every part of the network layer is
|
||||
generated by module hooks from mesh-DB rows:
|
||||
|
||||
| layer | generated by | from |
|
||||
|---|---|---|
|
||||
| WireGuard interface + peers | `modules/wireguard/hooks/index.ts` (`postConfigure`) | `node_wg_keys`, `nodes.site`, `nodes.underlay_addr`, `module_env.WG_ADDRESS` |
|
||||
| `.internal` name resolution | `modules/dnsmasq-app/hooks/index.ts` | mesh config peers → `internal_domain` + `wg_address` |
|
||||
| public routing / vhosts | `modules/hal/sdk/src/vhost-gen.ts`, `feature-handlers/vhost.ts` | `vhosts:` manifests + `node_accessors` |
|
||||
| internal TLS | `modules/mesh-ca/hooks/index.ts` | a singleton CA row in the mesh DB |
|
||||
|
||||
**Consequence:** a faithful lab is mostly a matter of writing the right rows. The network that
|
||||
results is produced by the same code production runs, which is the difference between testing
|
||||
the network and testing a model of it.
|
||||
|
||||
---
|
||||
|
||||
## 2. The constraint that decides whether the lab works
|
||||
|
||||
`modules/wireguard/hooks/index.ts:225-240` decides, per pair, whether to write an `Endpoint`:
|
||||
|
||||
```js
|
||||
const isPrivate = (a) => /^(10\.|127\.|192\.168\.|172\.(1[6-9]|2\d|3[01])\.)/.test(a);
|
||||
|
||||
if (coLocated && underlay) Endpoint = `${underlay}:${port}` // same LAN
|
||||
else if (underlay && !isPrivate(underlay)) Endpoint = `${underlay}:${port}` // public
|
||||
// else: no Endpoint — the peer must initiate, and we learn its endpoint from the handshake
|
||||
```
|
||||
|
||||
A simulated public segment addressed out of RFC1918 space — `10.200.0.0/24`, say — makes the
|
||||
hub's underlay test as **private**. No spoke writes an `Endpoint` for the hub. Nothing can
|
||||
initiate. **No handshake ever occurs and the mesh silently never forms**, presenting as a
|
||||
WireGuard fault rather than an addressing choice.
|
||||
|
||||
**The simulated public segment must therefore be `203.0.113.0/24`** — TEST-NET-3, reserved by
|
||||
RFC 5737 for documentation, guaranteed never to route on the real internet, and not matched by
|
||||
that regex. The code then treats it exactly as it treats a real hosting provider address.
|
||||
|
||||
This is the single most important fact in this document.
|
||||
|
||||
---
|
||||
|
||||
## 3. The topology being reproduced
|
||||
|
||||
The shape below is what a mesh of this kind looks like: one node with a routable address, one
|
||||
publicly named but behind a household NAT, one stationary workstation, one that roams.
|
||||
Addresses use the documentation ranges of RFC 5737 and RFC 1918 throughout.
|
||||
|
||||
| node | profile | site | underlay | WG | accessors |
|
||||
|---|---|---|---|---|---|
|
||||
| `anchor` | server | `dc` | `203.0.113.10` (routable) | `10.10.0.1/24` | `anchor.example` (public, primary) + `anchor.internal` (lan) |
|
||||
| `home-server` | server | `home` | `192.168.1.135` | `10.10.0.2/24` | `home-server.example` (public, primary) + `home-server.internal` (lan) |
|
||||
| `workstation` | workstation | `home` | `192.168.1.250` | `10.10.0.3/24` | `workstation.internal` (lan, primary) |
|
||||
| `laptop` | workstation | `NULL` | `NULL` | `10.10.0.4/24` | `laptop.internal` (lan, primary) |
|
||||
|
||||
**Hub election is by convention, not by flag.** The hub is the node whose `profile='server'`
|
||||
*and* whose `WG_ADDRESS` begins `10.10.0.1` (`hooks/index.ts:188`). A lab must assign
|
||||
`10.10.0.1` to the node it intends as hub or there will be no hub.
|
||||
|
||||
**`site` drives direct peering** (`hooks/index.ts:197-203`). Two nodes with the same non-null
|
||||
`site` peer directly with a `/32`; everything else routes through the hub's `/24`. A `NULL`
|
||||
site means roaming and hub-only — deliberately, because WireGuard has no failover and a more
|
||||
specific `/32` route to a dead endpoint blackholes rather than falling back.
|
||||
|
||||
The four interesting pairs, all of which the lab must reproduce:
|
||||
|
||||
| pair | behaviour | branch taken |
|
||||
|---|---|---|
|
||||
| anything → `anchor` | `Endpoint` written | underlay non-private |
|
||||
| `home-server` ↔ `workstation` | direct peer, LAN endpoints, keepalive | co-located |
|
||||
| **`anchor` → `home-server`** | **no `Endpoint`; learned from handshake** | not co-located, underlay private |
|
||||
| `laptop` → anything | hub only, always initiates | `site` is `NULL` |
|
||||
|
||||
The third is the one worth building the lab for. The code comments at `hooks/index.ts:206-224`
|
||||
record what it cost to get right: testing `profile === "server"` was tried and was wrong,
|
||||
because a home-hosted node **is** a server yet is not publicly reachable — *"role does not
|
||||
imply reachability; the address does."* An earlier version aimed the hub at that node's public
|
||||
name, which hairpinned off the household NAT: 1.77 MiB sent, 0 B received, no handshake.
|
||||
|
||||
---
|
||||
|
||||
## 4. The lab
|
||||
|
||||
```
|
||||
br-wan 203.0.113.0/24 TEST-NET-3 — non-private, so the code treats it as public
|
||||
│
|
||||
├── hub 203.0.113.10 profile=server site=dc WG 10.10.0.1
|
||||
│
|
||||
└── router VM 203.0.113.1 / 192.168.1.1
|
||||
│ NAT, plus one forwarded port to reproduce a published-but-NATed node
|
||||
│
|
||||
br-lan 192.168.1.0/24 identical to production, same host addresses
|
||||
├── a 192.168.1.135 profile=server site=home WG 10.10.0.2
|
||||
└── b 192.168.1.250 profile=workstation site=home WG 10.10.0.3
|
||||
|
||||
c — attach to br-lan, or br-wan ("away"), or detach ("asleep")
|
||||
underlay NULL profile=workstation site=NULL WG 10.10.0.4
|
||||
```
|
||||
|
||||
Kept **byte-identical** to production: the LAN subnet and its host addresses, and the entire
|
||||
WireGuard plan. Only the public segment is substituted, and only because it must be.
|
||||
|
||||
**The router earns its own VM.** It is what makes the published-but-NATed case real: that node
|
||||
is reachable from outside only through a forwarded port, and the hub must learn its endpoint.
|
||||
It also gives somewhere to break things — drop the forward and observe whether the mesh
|
||||
notices or whether the public name simply stops working.
|
||||
|
||||
**Names.** A resolver on the wan side is authoritative for the public zone. `.internal` names
|
||||
need nothing extra: `dnsmasq-app` generates them from mesh config on each node, and writes an
|
||||
`/etc/hosts` block as a floor underneath, because a node must reach the mesh DB before its own
|
||||
DNS exists.
|
||||
|
||||
---
|
||||
|
||||
## 5. What a node needs before any of this works
|
||||
|
||||
Rows in the mesh DB — `nodes` (`name`, `user_name`, `profile`, `site`, `underlay_addr`),
|
||||
`node_accessors`, `module_env.WG_ADDRESS` (the hook hard-fails without it,
|
||||
`hooks/index.ts:114-116`), and `node_modules` assigning at least `wireguard`, `dnsmasq-app`,
|
||||
`mesh-ca`, and `traefik` where it serves.
|
||||
|
||||
On disk beforehand, because the node must reach the mesh DB before it can read any of the
|
||||
above: registry database and object-store host and credentials, plus an npm token. This
|
||||
ordering — contact the mesh before the mesh has configured you — is itself worth reproducing,
|
||||
and is why the `/etc/hosts` floor exists.
|
||||
|
||||
Everything else is generated: the WireGuard keypair locally (the private key never leaves the
|
||||
node; the public key is published to `node_wg_keys`), the peer list, the DNS records, the TLS
|
||||
leaf.
|
||||
|
||||
---
|
||||
|
||||
## 6. Certificates — the lab issues its own
|
||||
|
||||
**Settled 2026-08-22: the lab runs its own ACME issuer.**
|
||||
|
||||
Public certificates use ACME **HTTP-01** via the reverse proxy, which requires genuine public
|
||||
reachability, so an isolated lab cannot use the real issuer. Rather than forgo certificate
|
||||
testing, the lab stands up an ACME server on its wan segment.
|
||||
|
||||
**The lab keeps production's two-CA split rather than collapsing it.** Production issues
|
||||
public names from a public authority and internal names from the mesh CA; a lab with one CA
|
||||
would hide any bug living in that split. So:
|
||||
|
||||
| | production | lab |
|
||||
|---|---|---|
|
||||
| public names | a public ACME authority | a test ACME server on the wan segment |
|
||||
| `.internal` names | `mesh-ca` | `mesh-ca`, unchanged |
|
||||
|
||||
**A test issuer is the right shape, not a shortcut.** Purpose-built ACME test servers
|
||||
deliberately vary their behaviour — validation timing, nonce handling, chain composition — to
|
||||
expose assumptions a well-behaved authority would let pass. A lab CA that is *too* polite
|
||||
tests less than the real thing, not more.
|
||||
|
||||
### It also exercises the port forward
|
||||
|
||||
HTTP-01 means the issuer must reach the node being certified on port 80. In the lab:
|
||||
|
||||
- the hub is directly reachable on the wan segment — straightforward
|
||||
- **the published-but-NATed node is reachable only through the router's forwarded port**
|
||||
|
||||
So certificate issuance for that node passes only if the forward is correct. That is exactly
|
||||
why its certificate works in production, and it makes "the forward is missing" a reproducible
|
||||
failure rather than a mystery.
|
||||
|
||||
### Required change: `caServer` must be configurable
|
||||
|
||||
`modules/traefik/docker-compose.yml:17-19` sets the challenge entrypoint, the contact address
|
||||
and the storage path — but **no `caServer`**, so Traefik defaults to the public authority's
|
||||
*production* endpoint. Pointing the lab at its own issuer requires adding a `caServer` flag
|
||||
fed by an environment value, defaulting to production so real nodes are unaffected and the lab
|
||||
overrides it per node.
|
||||
|
||||
Worth noting independently of the lab: aiming at the production endpoint rather than a staging
|
||||
one means every certificate experiment on a real node consumes production issuance quota, and
|
||||
a retry loop can exhaust it for a week. The lab issuer removes that exposure.
|
||||
|
||||
---
|
||||
|
||||
## 7. Incidental finding: `scope:` is read by nothing
|
||||
|
||||
Several manifests declare `scope: public` on firewall rules — `wireguard`, `traefik`, `gitea`,
|
||||
`mailu`, `qbittorrent`. It is **not part of the rule type** (`module-registry.ts:17-29`) and is
|
||||
**referenced by no code** in the firewall path. Real scoping is done with `from:`, as
|
||||
`modules/unifi/module.yml:52-93` does deliberately.
|
||||
|
||||
So a manifest can appear to restrict a port to the public scope and in fact restrict nothing.
|
||||
This is the same shape as the rule in `00-GENESIS/how-we-build.md` — *an unenforced rule is
|
||||
indistinguishable from a wrong one, and costs more, because people believe it.*
|
||||
@@ -0,0 +1,44 @@
|
||||
# 004 — Reproducing the mesh network in a lab
|
||||
|
||||
- **Status:** ONGOING — topology established and mapped; not yet stood up
|
||||
- **Initiated by:** jochen, 2026-08-22 — *"the most difficult part of our VM setup will be
|
||||
the networking part"*
|
||||
- **Areas touched:** `modules/wireguard`, `modules/dnsmasq-app`, `modules/traefik`,
|
||||
`modules/mesh-ca`, `node_accessors`, `nodes.site` / `nodes.underlay_addr`.
|
||||
|
||||
## Summary
|
||||
|
||||
The network is **entirely generated from mesh-DB rows by module hooks**. `install.d` performs
|
||||
no network configuration whatsoever — no WireGuard, no DNS, no firewall. That makes a faithful
|
||||
lab primarily a *data* problem rather than a networking problem, and means the lab exercises
|
||||
the real code path instead of a reimplementation of it.
|
||||
|
||||
One constraint decides whether the lab works at all: the WireGuard endpoint rule tests the
|
||||
underlay address against an RFC1918 regex to decide reachability. **A simulated public segment
|
||||
addressed from RFC1918 space silently prevents the mesh from forming** — no endpoint is written
|
||||
for the hub, so nothing can ever initiate. The simulated public segment must therefore use
|
||||
TEST-NET-3 (`203.0.113.0/24`).
|
||||
|
||||
With that one substitution the lab reproduces the production topology exactly, including the
|
||||
case that is hardest to get right: a node that is publicly *named* but sits behind NAT, whose
|
||||
endpoint the hub can only learn from a handshake.
|
||||
|
||||
Detail in [`analysis.md`](analysis.md).
|
||||
|
||||
## Settled
|
||||
|
||||
**The lab issues its own certificates.** Public names are certified by an ACME server on the
|
||||
lab's wan segment; `.internal` names keep the mesh CA. The lab preserves production's two-CA
|
||||
split rather than collapsing it, because a single-CA lab would hide any bug living in that
|
||||
split. It also makes the router's port forward load-bearing — HTTP-01 must reach the
|
||||
published-but-NATed node on port 80, so a broken forward becomes a reproducible certificate
|
||||
failure instead of a mystery.
|
||||
|
||||
Requires one change: `caServer` is not set on the reverse proxy today, so it defaults to the
|
||||
public authority's **production** endpoint. It must become configurable, defaulting to
|
||||
production so real nodes are unaffected.
|
||||
|
||||
## Open
|
||||
|
||||
- Not yet stood up. `incus` is declared in `modules/hal/developer/module.yml` and merged
|
||||
(PR #944); the lab itself is unbuilt.
|
||||
Reference in New Issue
Block a user