A roster fact used to be a name from a closed list, each formatted in Go
here — node-names as a hosts file, node-zones as a resolver's zones. Every
new consumer (ssh's known_hosts, an authorized_keys) meant another formatter
in the control plane, in the consumer's own configuration language.
Now a fact is a path and a Go template over the roster view (this node, the
suffix, and every served name vs the machines). The mesh owns the data; the
module owns the format. /etc/hosts is a template on the network module;
dnsmasq's zones move to dnsmasq. The controller renders and reads neither.
WireGuard stays a computed generator: the overlay is the substrate delivery
rides on, and its config is topology, not a roster projection.
Output is byte-for-byte unchanged, pinned by the hosts golden tests and the
resolver tests that compose the real dnsmasq manifest.
/etc/hosts is the machine's: the distribution's localhost lines, the
operator's own entries, and marked blocks other tools maintain there.
Writing node-names whole replaced all of it the moment the private
network was taken, and every later write by those tools was lost at
the next machine joining. The node-names fact is now emitted with
into: "block", so the host owns only its marked region and keeps the
rest byte for byte. The region holds only the mesh's names: no header
claiming the file, no localhost, no 127.0.1.1 line — the floor was
never the mesh's to write. How a fact is written is a property of the
fact in the closed table; node-zones stays a whole file the mesh owns.
Sequencing: a host older than the block mode refuses the whole
declaration on an unknown into, so every host must be upgraded before
this controller is rolled out.
Carries MTU from the reported tunnel (mesh-host#28) through inventory,
the overlay graph's TakeOver, into the generated config's [Interface].
A tuned path keeps its MTU across the takeover instead of regressing to
1420 and hanging transfers no ping would reveal. Two emit tests; a
tunnel with no MTU writes no line.
A home node behind NAT (no Endpoint → not Reachable) that took over a
tunnel must still listen on that tunnel's port: its LAN peers dial it
there. ListenPort was gated on Reachable, which conflated 'a peer dials
me here' with 'the hub can dial me' — so the takeover guard refused
overlay-up, and the guard's suggested remedy (re-place with an
endpoint) breaks a NAT'd node's path: it stops keepalive and hands the
hub a private LAN address to dial. TakeOver now carries the found
tunnel's port (already known to the controller), and the interface
listens on it when the node is not otherwise reachable. Two tests;
Endpoint-reachable nodes keep the old path unchanged.
Implements novox/hq ADR 0110 and 0111.
The seat set lives in internal/catalogue/seats.go: fourteen seats, each with a scope, what occupying
it delivers, and the record that made it one. A test asserts the count and a decision per entry, so
changing the set means finding the argument, as the host's vocabulary test does. The first set is
every seat already claimed — including the-private-network, which the network module claims from a
manifest composed in this repository's code, not from any module.json — plus npm-package-registry
(ADR 0109) and git (ADR 0111). A test parses every catalogue manifest and this repository's own and
fails on any refused claim, so closing the set refuses nothing in use.
ParseManifest now refuses a claim on a seat the mesh does not define, a seat claimed at another
scope, and a delivering seat claimed by a module that does not provide what it delivers. A
malformed claim is refused once, for being malformed.
Resolution: among several providers of a mesh provision, a pin still wins; then the holder of the
seat that delivers it; then the only provider; otherwise refused as before. ADR 0009's "never
guessed" holds — the seat is the choice made once, mesh-wide, rather than a pin per consumer node.
A provider now carries the module it came from, because a provider is a (node, module) pair and the
pair is what tells a holder from a neighbour on the same machine.
The planner's second pass is now given the first pass's holdings. Without them, a node consuming a
seat-delivered provision was refused there, and a refused node's own claims dropped out of what the
mesh holds — letting a second holder of one of its seats pass unrefused.
`seats [--json]` lists every seat, what it delivers, and each holder, derived from assignments
every time and never stored. Unheld seats are listed. A stored claim outside the set — possible
for a manifest registered before the set closed, since stored manifests are not re-validated — is
shown rather than hidden.
`build --self <owner>/<repo>` builds from a repository on the git seat's holder. The clone URL is
composed at build time from the holder's node and what it serves for git; the recorded source is the
path and the seat (migration 0032), never an address, so a moved forge changes nothing recorded.
Nobody holding the seat refuses self-hosted builds and says so; external URLs are unchanged. An
address passed with --self is refused rather than recorded as a path.
Replaces three foundation tests that defended the builder's carried package binding. The catalogue
removed that binding when the builder began requiring the registry through a real grant, so the
tests were already failing on main; they now assert the builder requires what the npm seat delivers
and carries no copy of its own, and that the forge holds the npm and git seats.
Verified: go vet clean; the whole suite passes against a throwaway Postgres (make postgres), the new
inventory tests included; gofmt clean apart from cmd/mesh-builder/stdout_test.go, which fails on
main too.
hal dnsmasq-app conversion, hq 08-connectivity. Converting the resolver from the module it
replaces made it forward what it cannot answer, which is what the predecessor's does, and
that found two things the controller did not say.
A resolver that forwards must not send a mesh name it does not know upstream: the
`node-zones` fact now carries `local=/<suffix>/` beside the wildcards, written here rather
than in the daemon's configuration because the suffix is the mesh's choice and this file is
the one place the mesh writes what it chose. The default lives in one helper now instead of
being spelled in two functions.
The predecessor points the container runtime's `dns` at the machine's own tunnel address —
a container cannot reach the machine's loopback. A module writing that key needs the
address, and `${machine:at}` is the machine's name; a runtime's resolver list cannot be a
name it would need that resolver to look up. So a module may say `${machine:address}`: what
`at` resolves to, read from the same names the hosts file and the wildcards are written
from, absent — and refused — off the network like `at` is.
The `mesh-resolver` and `resolver-data` constants go: nothing provided or consumed either,
the fact and `mesh-addressing` are the mechanism, and a requirement nothing provides is
refused at resolution.
Tests: the catalogue's dnsmasq, resolv-conf and resolved-split-dns manifests are parsed
and composed as a machine would receive them — fixed upstreams, no-resolv, 127.0.0.1, the
machines file, the runtime's key, the pair that decides what a machine asks refused on one
node; and on a real mesh the resolver's machines file is composed with a wildcard per
machine on the network and composed again without one that left, mirroring the hosts fact.
Review of the ADR 0105 build (hq ADR 0105). Four things it got wrong and one
path it lacked:
- A predecessor spoke's tunnel names one peer, the hub, routed the whole
range; recording refused it and the whole enrolment failed. Range-routed
peers are skipped now — only the hub's peers are ever carried.
- The range and the carried peers were conditions on the node being adopted,
so converging the hub would have renumbered the mesh and dropped the peers
still reaching it. They are facts of the tunnel record now, mode aside; the
takeover alone is declared to an adopted node. Converging the hub is refused
while a carried peer has not enrolled, naming it.
- A push composed a takeover for a hub whose address or endpoint disagreed
with the tunnel, which would have the host stop the found interface and
raise the mesh's where no peer listens. The graph refuses to compose it,
naming both and the placement that fixes it.
- The host's account said taken or not; "found down and the mesh's not up"
read as not taken. Three states now, and an account on every takeover.
- A hub that enrolled before this feature holds a key of its own, and
re-enrolling would rotate every key the mesh sealed credentials to. A node
now rekeys in a report, signed with its identity key over the key it
leaves, the key it takes and the tunnel; the mesh verifies against the live
key, refuses a stale or foreign proof, records key and tunnel, and moves a
hub to the tunnel's address. `overlay show` names the path for a hub that
found no tunnel.
Also: a carried IPv6 peer is routed /128, and identity.ForTest exists so the
link can be tested against a real identity store.
On an adopted hub the private network takes over the tunnel it finds rather
than running beside it (hq ADR 0105): two tunnels leave the mesh's unreachable
through the provider's filter, so no machine can ever join.
The node presents the found tunnel when it enrols, under the key it took as
its own; the inventory records it (node.tunnel, tunnel_peer — migration 0031)
and the mesh composes from it: the overlay's range is the adopted tunnel's,
the hub is placed at the tunnel's address on the tunnel's port, and every
peer the tunnel had is carried in the hub's peer list as a peer of the
tunnel, not a node of the mesh, until a node enrols with that key — which
then keeps the address the tunnel had for it. A fresh node never gets an
address the tunnel holds. The hub's declaration tells the host which unit to
take over; the host's account of carrying it is recorded and shown.
Every reader of the range follows the setting; nothing stores it. A found
tunnel under another key is recorded and not adopted, so ADR 0100's
non-overlap rule keeps applying where a tunnel is left running beside the
mesh's. A lab bed and test skeleton for "How it is checked" are under lab/.
Being on the private network is what grants a machine the right to pull from the mesh's
artifact store, so the module that puts a machine on the network writes the runtime's
trust — a merged /etc/docker/daemon.json naming the store's internal name under
insecure-registries, and a docker.service restart when that fact first lands. The registry
speaks plain HTTP because every path to it is already inside the overlay's encryption; the
provider is found, not configured — whichever module serves artifact-store, on whichever
machine holds it — and with no store on the network nothing is written, which is genesis.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The hosts-file writing moved to the node-names fact; what stays is the suffix
(configurable, MESH_INTERNAL_SUFFIX, defaulting to .internal which IANA reserved
for exactly this), a node's internal name, and the rule for what a node may be
called. Several things compose an internal name, and one of them writing the
suffix differently would be a name nothing answers to.
First attempt at this rewrote the file from memory and silently dropped the
configurable suffix. Restored from the original instead — deleting most of a
file is git surgery, not paraphrase.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
mesh-names, mesh-resolver and the names half of the overlay generators are gone.
They ran no software and could not be swapped for anything, which is the test of
whether something is a module at all — they existed because computed output
needed somewhere to live, and the control plane's only shape for output was a
module.
Now a module says where it wants what the mesh knows:
facts: { node-zones: /etc/mesh-resolver/nodes.conf }
and is given a file, under its own name, applied and removed like anything else
it declares. Two facts exist: node-names (a hosts file — exact names) and
node-zones (every machine as a wildcard, *.homer.internal is homer). Asking for
a fact the mesh does not compute is refused naming what would have worked,
because a daemon that starts and reads a file nobody wrote is a worse way to
find out.
The names ride with the network now: wireguard's manifest asks for node-names
into /etc/hosts, because being on the private network is what gives a machine a
name. networking no longer requires name-resolution — names are not a provision,
and the module that answered it ran nothing.
One behaviour inverted, deliberately: choosing another VPN used to drag
WireGuard in anyway, because only WireGuard provided the addressing the names
module required — the node-scope claim existed to at least make that loud. With
names as a fact there is nothing to drag in: tailscale assigned means tailscale,
alone. The claim still catches two VPNs assigned explicitly.
And a machine the mesh cannot place is left out of both files rather than named
at nothing: a name resolving to nothing hangs a connection, where an unknown
name fails at once and says so. In practice that is only ever a token issued and
not yet used — a machine that has announced itself has an address.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Services are named under the machine they run on — postgres.novox.internal,
plex.ace.internal. The first label is the service and the rest is the node, so
what has to resolve is anything under a node's name. What routes it once it
arrives is a proxy's concern and stays separate.
A hosts file cannot do that. It answers exact names, and a wildcard there would
mean writing down every service in advance — which is the enumeration the
arrangement exists to avoid. novox/hq 08-connectivity named this exact case as
the trigger for needing a resolver rather than a file, and it is the first
thing to meet it.
The mesh writes the data and runs no daemon. A resolver is third-party
software, and third-party software runs on the mesh rather than being of it
(ADR 0001): the mesh has no business shipping one, choosing which one, or
knowing its configuration language. What only the mesh can know is which
machines exist and where they are. A module that runs a resolver requires what
this provides and reads one file, so swapping the daemon changes that module
and nothing here.
Separate from names rather than part of them: a machine with no container
runtime can still have a hosts file, and folding them together would take exact
names away from a machine that cannot run a daemon in order to give it a
wildcard it cannot use either.
A machine with no address is left out. A wildcard pointing at nothing is worse
than no wildcard — every name under it resolves and then hangs, where an
unresolvable name fails at once and says which name it was.
The machine that most needed a firewall was the one that could not have one. A
hub is dialled by every node at other sites and needs its port open; a machine
that is not a hub dials out and needs nothing open. They are the same module,
and `listens` in a manifest is one answer for every machine that runs it — so
the machine a static answer gets wrong is the one facing the public internet.
A generator can now say what it opens, in a second interface rather than a
method on every generator: most have nothing to say here, and requiring an
empty method of each would be a cost paid everywhere for one caller.
The port is the one in the endpoint, which is where the interface takes its
ListenPort from. One source, so a rule set cannot open a port the interface is
not on. Open to everywhere and deliberately: a node at another site is not on
the private network until this port lets it on, so restricting it to the mesh
would be a rule that can never be satisfied by the thing it exists for.
And a generator that cannot say is refused rather than read as silence. Closing
a port on the evidence of a failure to look is how a machine is severed by a
fault somewhere else — and the machine it would sever is the hub, whose only
route to being fixed is the network it just closed.
Connectivity was code beside the module system doing the module system's
job: every machine with an address was on the private network and there
was no way to keep one off.
A manifest can now say its resources are computed by the control plane,
which is what a peer list needs — it is derived from every machine at
once, so nothing could be written in advance. The network is a module
from there on: assigned, resolved, settled, and absent from a machine
nobody gave it to.
Three modules rather than one, because WireGuard is one VPN of several:
mesh-wireguard provides private-network, mesh-addressing
claims the-private-network, one per node
mesh-names provides name-resolution, requires mesh-addressing
networking requires both, and ships no files of its own
The last is the point. Most people want the network up and do not want
to choose a VPN, so `assign networking` takes the only answer to each
requirement silently. The day the catalogue holds a second one there are
two answers, the resolver refuses and names them, and choosing is
assigning the one you want. No flavor field, nothing to configure.
Names left the WireGuard declaration for their own module. They would be
identical over a different private network, and bundling them made one
module out of two things.
Three faults the walk found:
- choosing tailscale still installed WireGuard, dragged back in by the
names needing the mesh's own addresses. Caught now by a claim: running
two VPNs is fine, being *the* mesh network is singular.
- a requirement wanted by two modules was reported twice, identically.
- "this mesh has no hub" was reported when the real cause was that a
node could not be resolved at all. It now names the node and the why.
And a test that asserts the manifests actually shipped, after the claim
went missing from the real one while every test stayed green.
Step 5 of the connectivity order. Every node's internal name resolves to its
overlay address, on every node, computed centrally because it needs every node
at once.
Under `.internal`, which IANA reserved for exactly this in 2024 -- a name there
can never collide with a public one, so an internal name that leaks into a
public resolver fails rather than reaching a stranger's machine. The suffix is
settable for a mesh that wants its own.
Delivered in the same declaration as the peer list rather than a second one. A
node holding the peers and not the names, or the reverse, is half on the
network for as long as that lasts.
This is not the /etc/hosts floor the design removes. That floor existed because
a node had to reach the mesh's database before its own DNS worked -- a fallback
for a circularity that is now gone. This is the mechanism: the complete set of
names, generated whole and owned by the mesh, rather than a patch written
underneath something else. A resolver daemon becomes necessary when names are
wanted that are not one-per-node, and that is not yet true.
A node resolves its own name to its overlay address rather than a loopback,
because a service binding to the name it was given would otherwise listen
somewhere nothing else can reach -- and the failure would appear on every other
machine rather than that one.
A node with no address gets no name. A name resolving to nothing is worse than
no name: connecting to an address that does not answer hangs, where a name that
does not resolve fails at once and says which name it was.
Found while writing it: a test asserting every file in the declaration is mode
0600 would have forced /etc/hosts to 0600 and broken every lookup on the
machine, to protect a file that is not secret.
Verified in the lab: three machines, nine name lookups, each resolving to the
right overlay address and reaching it.
Three machines across two sites, two of them behind no reachable address, all
nine paths open. The mesh computes the graph, delivers it as a declaration, and
the nodes bring it up.
Every fault below looked like success from inside the mesh: the graph was
right, the files were right, the services were up, every node reported it had
applied. None was reachable by reasoning.
A running interface does not re-read its configuration. A node joins, every
existing node's peer list changes, the file is replaced -- and the service is
already running, so nothing reloads it. Fixed as declared state rather than a
command: the service must reflect the file. A command to restart would be an
action, and the link may not carry one. The host refused exactly that, which is
how this shape was arrived at.
A hub sharing a site with a spoke appeared twice in that spoke's peer list --
once as a direct peer, once as the route of last resort. WireGuard takes one
entry per key and refuses the file. The ordinary shape of a small mesh, and in
none of the tests written before it ran.
Two nodes at one site that neither can be dialled were peered directly. Nobody
opens the path, and the direct route is more specific than the hub's, so it
wins and blackholes -- this design's own warning arriving in its
implementation. They now route through the hub unless one end can be dialled.
And Docker sets the FORWARD policy to DROP, so a hub with ip_forward enabled
carried nothing between its spokes. The substrate at tier 1 silently breaks the
network at tier 2, and nothing in either tier's state says so. The hub inserts
its own rule above those chains and removes it on the way down.
Two weak tests found by injection along the way: one asserted the keepalive
rule only against the hub, whose peer entries happen not to set that field at
all, so it tested an absence; the other checked the firewall rules by looking
for FORWARD anywhere, which the PostDown line satisfies on its own.
The first thing the control plane decides rather than relays. Every node's peer
list is derived from every node at once, which is what makes this control-plane
work by definition: no node has that view.
A hub, with direct peering between nodes at the same site. Not a full mesh, and
the reason is a property of WireGuard rather than a preference -- there is no
failover, so a more specific route to a dead endpoint blackholes instead of
falling back. A node gets exactly one path to any peer, because two would mean
one of them silently swallowing traffic. A roaming node is hub-only for the
same reason.
Reachability and the hub are declared, never inferred from an address. The
address is evidence and is not the fact: carrier-grade NAT looks public and is
not, a routable address behind a closed firewall looks public and is not, and
the regular expression that used to decide it got the lab wrong too. Hub
election by address prefix failed silently when nobody knew the convention.
No private key travels, and that is the whole design. The node generated its
own keypair and kept the private half; the configuration points at a file the
node wrote, using WireGuard's own PostUp. So the control plane composes a
complete configuration for a node it cannot pretend to be -- it knows every
public key and holds none of the private ones.
Delivered as an ordinary declaration: a package, a file and a service. The host
does not know what a private network is and does not learn one. There is a test
holding that line, because the moment connectivity needs a new shape in tier 0
is the moment the host stops being small enough to trust.
The generated file is written to be read: each peer says why it is there, a
peer with no endpoint says why it has none, and the header says not to edit it
-- an edit survives until the graph next changes and then vanishes, which is
worse than never being applied, because the machine works and then stops and
nothing changed that anybody remembers.
Fault injection found one weak test. The keepalive rule was asserted only
against the hub, whose peer entries happen not to set the field at all, so it
was testing an absence rather than the rule. It now checks two direct peers
where one is reachable and one is not.