fb87f9f7d6dcaa3dd8e708573c039d5f38ea3005
15
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
97448194ac |
Seats are a closed set, a seat's holder answers for what it delivers, and a build source may live on the git seat
Implements novox/hq ADR 0110 and 0111. The seat set lives in internal/catalogue/seats.go: fourteen seats, each with a scope, what occupying it delivers, and the record that made it one. A test asserts the count and a decision per entry, so changing the set means finding the argument, as the host's vocabulary test does. The first set is every seat already claimed — including the-private-network, which the network module claims from a manifest composed in this repository's code, not from any module.json — plus npm-package-registry (ADR 0109) and git (ADR 0111). A test parses every catalogue manifest and this repository's own and fails on any refused claim, so closing the set refuses nothing in use. ParseManifest now refuses a claim on a seat the mesh does not define, a seat claimed at another scope, and a delivering seat claimed by a module that does not provide what it delivers. A malformed claim is refused once, for being malformed. Resolution: among several providers of a mesh provision, a pin still wins; then the holder of the seat that delivers it; then the only provider; otherwise refused as before. ADR 0009's "never guessed" holds — the seat is the choice made once, mesh-wide, rather than a pin per consumer node. A provider now carries the module it came from, because a provider is a (node, module) pair and the pair is what tells a holder from a neighbour on the same machine. The planner's second pass is now given the first pass's holdings. Without them, a node consuming a seat-delivered provision was refused there, and a refused node's own claims dropped out of what the mesh holds — letting a second holder of one of its seats pass unrefused. `seats [--json]` lists every seat, what it delivers, and each holder, derived from assignments every time and never stored. Unheld seats are listed. A stored claim outside the set — possible for a manifest registered before the set closed, since stored manifests are not re-validated — is shown rather than hidden. `build --self <owner>/<repo>` builds from a repository on the git seat's holder. The clone URL is composed at build time from the holder's node and what it serves for git; the recorded source is the path and the seat (migration 0032), never an address, so a moved forge changes nothing recorded. Nobody holding the seat refuses self-hosted builds and says so; external URLs are unchanged. An address passed with --self is refused rather than recorded as a path. Replaces three foundation tests that defended the builder's carried package binding. The catalogue removed that binding when the builder began requiring the registry through a real grant, so the tests were already failing on main; they now assert the builder requires what the npm seat delivers and carries no copy of its own, and that the forge holds the npm and git seats. Verified: go vet clean; the whole suite passes against a throwaway Postgres (make postgres), the new inventory tests included; gofmt clean apart from cmd/mesh-builder/stdout_test.go, which fails on main too. |
||
|
|
0d8264ff55 |
Give the resolver the mesh's suffix as a local domain and a module its machine's address
hal dnsmasq-app conversion, hq 08-connectivity. Converting the resolver from the module it
replaces made it forward what it cannot answer, which is what the predecessor's does, and
that found two things the controller did not say.
A resolver that forwards must not send a mesh name it does not know upstream: the
`node-zones` fact now carries `local=/<suffix>/` beside the wildcards, written here rather
than in the daemon's configuration because the suffix is the mesh's choice and this file is
the one place the mesh writes what it chose. The default lives in one helper now instead of
being spelled in two functions.
The predecessor points the container runtime's `dns` at the machine's own tunnel address —
a container cannot reach the machine's loopback. A module writing that key needs the
address, and `${machine:at}` is the machine's name; a runtime's resolver list cannot be a
name it would need that resolver to look up. So a module may say `${machine:address}`: what
`at` resolves to, read from the same names the hosts file and the wildcards are written
from, absent — and refused — off the network like `at` is.
The `mesh-resolver` and `resolver-data` constants go: nothing provided or consumed either,
the fact and `mesh-addressing` are the mechanism, and a requirement nothing provides is
refused at resolution.
Tests: the catalogue's dnsmasq, resolv-conf and resolved-split-dns manifests are parsed
and composed as a machine would receive them — fixed upstreams, no-resolv, 127.0.0.1, the
machines file, the runtime's key, the pair that decides what a machine asks refused on one
node; and on a real mesh the resolver's machines file is composed with a wildcard per
machine on the network and composed again without one that left, mirroring the hosts fact.
|
||
|
|
4566c5c9aa |
Adopt the tunnel as a mesh fact, refuse a mismatched takeover, and rekey after enrolment
Review of the ADR 0105 build (hq ADR 0105). Four things it got wrong and one path it lacked: - A predecessor spoke's tunnel names one peer, the hub, routed the whole range; recording refused it and the whole enrolment failed. Range-routed peers are skipped now — only the hub's peers are ever carried. - The range and the carried peers were conditions on the node being adopted, so converging the hub would have renumbered the mesh and dropped the peers still reaching it. They are facts of the tunnel record now, mode aside; the takeover alone is declared to an adopted node. Converging the hub is refused while a carried peer has not enrolled, naming it. - A push composed a takeover for a hub whose address or endpoint disagreed with the tunnel, which would have the host stop the found interface and raise the mesh's where no peer listens. The graph refuses to compose it, naming both and the placement that fixes it. - The host's account said taken or not; "found down and the mesh's not up" read as not taken. Three states now, and an account on every takeover. - A hub that enrolled before this feature holds a key of its own, and re-enrolling would rotate every key the mesh sealed credentials to. A node now rekeys in a report, signed with its identity key over the key it leaves, the key it takes and the tunnel; the mesh verifies against the live key, refuses a stale or foreign proof, records key and tunnel, and moves a hub to the tunnel's address. `overlay show` names the path for a hub that found no tunnel. Also: a carried IPv6 peer is routed /128, and identity.ForTest exists so the link can be tested against a real identity store. |
||
|
|
3c836f0abb |
Adopt the predecessor's tunnel in place: its range, its address, its peers
On an adopted hub the private network takes over the tunnel it finds rather than running beside it (hq ADR 0105): two tunnels leave the mesh's unreachable through the provider's filter, so no machine can ever join. The node presents the found tunnel when it enrols, under the key it took as its own; the inventory records it (node.tunnel, tunnel_peer — migration 0031) and the mesh composes from it: the overlay's range is the adopted tunnel's, the hub is placed at the tunnel's address on the tunnel's port, and every peer the tunnel had is carried in the hub's peer list as a peer of the tunnel, not a node of the mesh, until a node enrols with that key — which then keeps the address the tunnel had for it. A fresh node never gets an address the tunnel holds. The hub's declaration tells the host which unit to take over; the host's account of carrying it is recorded and shown. Every reader of the range follows the setting; nothing stores it. A found tunnel under another key is recorded and not adopted, so ADR 0100's non-overlap rule keeps applying where a tunnel is left running beside the mesh's. A lab bed and test skeleton for "How it is checked" are under lab/. |
||
|
|
28b7fb81ba | Write the registry's trust into the runtime's file and reload the runtime instead of restarting it; prefix reload-on like restart-on (hq ADR 0102) | ||
|
|
0e9035d479 |
The network carries the registry trust (ADR 0082, issues 042/048)
Being on the private network is what grants a machine the right to pull from the mesh's artifact store, so the module that puts a machine on the network writes the runtime's trust — a merged /etc/docker/daemon.json naming the store's internal name under insecure-registries, and a docker.service restart when that fact first lands. The registry speaks plain HTTP because every path to it is already inside the overlay's encryption; the provider is found, not configured — whichever module serves artifact-store, on whichever machine holds it — and with no store on the network nothing is written, which is genesis. https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx |
||
|
|
c3b88b9148 |
Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo becomes mesh-controller, the seat the-controller, and the store+broker pair the foundation (embedded base bundles, default template and example lock renamed with their go:embed directives). No behaviour change — a pure vocabulary rename. Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx |
||
|
|
b8cacbaf4d |
What remains of names.go is the naming
The hosts-file writing moved to the node-names fact; what stays is the suffix (configurable, MESH_INTERNAL_SUFFIX, defaulting to .internal which IANA reserved for exactly this), a node's internal name, and the rule for what a node may be called. Several things compose an internal name, and one of them writing the suffix differently would be a name nothing answers to. First attempt at this rewrote the file from memory and silently dropped the configurable suffix. Restored from the original instead — deleting most of a file is git surgery, not paraphrase. Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx |
||
|
|
fcdb065660 |
The mesh's knowledge is a fact a module asks for, not three modules
mesh-names, mesh-resolver and the names half of the overlay generators are gone.
They ran no software and could not be swapped for anything, which is the test of
whether something is a module at all — they existed because computed output
needed somewhere to live, and the control plane's only shape for output was a
module.
Now a module says where it wants what the mesh knows:
facts: { node-zones: /etc/mesh-resolver/nodes.conf }
and is given a file, under its own name, applied and removed like anything else
it declares. Two facts exist: node-names (a hosts file — exact names) and
node-zones (every machine as a wildcard, *.homer.internal is homer). Asking for
a fact the mesh does not compute is refused naming what would have worked,
because a daemon that starts and reads a file nobody wrote is a worse way to
find out.
The names ride with the network now: wireguard's manifest asks for node-names
into /etc/hosts, because being on the private network is what gives a machine a
name. networking no longer requires name-resolution — names are not a provision,
and the module that answered it ran nothing.
One behaviour inverted, deliberately: choosing another VPN used to drag
WireGuard in anyway, because only WireGuard provided the addressing the names
module required — the node-scope claim existed to at least make that loud. With
names as a fact there is nothing to drag in: tailscale assigned means tailscale,
alone. The claim still catches two VPNs assigned explicitly.
And a machine the mesh cannot place is left out of both files rather than named
at nothing: a name resolving to nothing hangs a connection, where an unknown
name fails at once and says so. In practice that is only ever a token issued and
not yet used — a machine that has announced itself has an address.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
|
||
|
|
3954157555 |
The mesh computes every name under a machine, for a resolver to answer
Services are named under the machine they run on — postgres.novox.internal, plex.ace.internal. The first label is the service and the rest is the node, so what has to resolve is anything under a node's name. What routes it once it arrives is a proxy's concern and stays separate. A hosts file cannot do that. It answers exact names, and a wildcard there would mean writing down every service in advance — which is the enumeration the arrangement exists to avoid. novox/hq 08-connectivity named this exact case as the trigger for needing a resolver rather than a file, and it is the first thing to meet it. The mesh writes the data and runs no daemon. A resolver is third-party software, and third-party software runs on the mesh rather than being of it (ADR 0001): the mesh has no business shipping one, choosing which one, or knowing its configuration language. What only the mesh can know is which machines exist and where they are. A module that runs a resolver requires what this provides and reads one file, so swapping the daemon changes that module and nothing here. Separate from names rather than part of them: a machine with no container runtime can still have a hosts file, and folding them together would take exact names away from a machine that cannot run a daemon in order to give it a wildcard it cannot use either. A machine with no address is left out. A wildcard pointing at nothing is worse than no wildcard — every name under it resolves and then hangs, where an unresolvable name fails at once and says which name it was. |
||
|
|
092109debc |
A computed module says what its machine opens, so a hub can be filtered
The machine that most needed a firewall was the one that could not have one. A hub is dialled by every node at other sites and needs its port open; a machine that is not a hub dials out and needs nothing open. They are the same module, and `listens` in a manifest is one answer for every machine that runs it — so the machine a static answer gets wrong is the one facing the public internet. A generator can now say what it opens, in a second interface rather than a method on every generator: most have nothing to say here, and requiring an empty method of each would be a cost paid everywhere for one caller. The port is the one in the endpoint, which is where the interface takes its ListenPort from. One source, so a rule set cannot open a port the interface is not on. Open to everywhere and deliberately: a node at another site is not on the private network until this port lets it on, so restricting it to the mesh would be a rule that can never be satisfied by the thing it exists for. And a generator that cannot say is refused rather than read as silence. Closing a port on the evidence of a failure to look is how a machine is severed by a fault somewhere else — and the machine it would sever is the hub, whose only route to being fixed is the network it just closed. |
||
|
|
44d134ba25 |
Networking is a module, and a domain module is how you avoid choosing
Connectivity was code beside the module system doing the module system's
job: every machine with an address was on the private network and there
was no way to keep one off.
A manifest can now say its resources are computed by the control plane,
which is what a peer list needs — it is derived from every machine at
once, so nothing could be written in advance. The network is a module
from there on: assigned, resolved, settled, and absent from a machine
nobody gave it to.
Three modules rather than one, because WireGuard is one VPN of several:
mesh-wireguard provides private-network, mesh-addressing
claims the-private-network, one per node
mesh-names provides name-resolution, requires mesh-addressing
networking requires both, and ships no files of its own
The last is the point. Most people want the network up and do not want
to choose a VPN, so `assign networking` takes the only answer to each
requirement silently. The day the catalogue holds a second one there are
two answers, the resolver refuses and names them, and choosing is
assigning the one you want. No flavor field, nothing to configure.
Names left the WireGuard declaration for their own module. They would be
identical over a different private network, and bundling them made one
module out of two things.
Three faults the walk found:
- choosing tailscale still installed WireGuard, dragged back in by the
names needing the mesh's own addresses. Caught now by a claim: running
two VPNs is fine, being *the* mesh network is singular.
- a requirement wanted by two modules was reported twice, identically.
- "this mesh has no hub" was reported when the real cause was that a
node could not be resolved at all. It now names the node and the why.
And a test that asserts the manifests actually shipped, after the claim
went missing from the real one while every test stayed green.
|
||
|
|
fc1417be72 |
Names, from the same graph as the network
Step 5 of the connectivity order. Every node's internal name resolves to its overlay address, on every node, computed centrally because it needs every node at once. Under `.internal`, which IANA reserved for exactly this in 2024 -- a name there can never collide with a public one, so an internal name that leaks into a public resolver fails rather than reaching a stranger's machine. The suffix is settable for a mesh that wants its own. Delivered in the same declaration as the peer list rather than a second one. A node holding the peers and not the names, or the reverse, is half on the network for as long as that lasts. This is not the /etc/hosts floor the design removes. That floor existed because a node had to reach the mesh's database before its own DNS worked -- a fallback for a circularity that is now gone. This is the mechanism: the complete set of names, generated whole and owned by the mesh, rather than a patch written underneath something else. A resolver daemon becomes necessary when names are wanted that are not one-per-node, and that is not yet true. A node resolves its own name to its overlay address rather than a loopback, because a service binding to the name it was given would otherwise listen somewhere nothing else can reach -- and the failure would appear on every other machine rather than that one. A node with no address gets no name. A name resolving to nothing is worse than no name: connecting to an address that does not answer hangs, where a name that does not resolve fails at once and says which name it was. Found while writing it: a test asserting every file in the declaration is mode 0600 would have forced /etc/hosts to 0600 and broken every lookup on the machine, to protect a file that is not secret. Verified in the lab: three machines, nine name lookups, each resolving to the right overlay address and reaching it. |
||
|
|
8b974deb42 |
A working private network, and four reasons it did not work
Three machines across two sites, two of them behind no reachable address, all nine paths open. The mesh computes the graph, delivers it as a declaration, and the nodes bring it up. Every fault below looked like success from inside the mesh: the graph was right, the files were right, the services were up, every node reported it had applied. None was reachable by reasoning. A running interface does not re-read its configuration. A node joins, every existing node's peer list changes, the file is replaced -- and the service is already running, so nothing reloads it. Fixed as declared state rather than a command: the service must reflect the file. A command to restart would be an action, and the link may not carry one. The host refused exactly that, which is how this shape was arrived at. A hub sharing a site with a spoke appeared twice in that spoke's peer list -- once as a direct peer, once as the route of last resort. WireGuard takes one entry per key and refuses the file. The ordinary shape of a small mesh, and in none of the tests written before it ran. Two nodes at one site that neither can be dialled were peered directly. Nobody opens the path, and the direct route is more specific than the hub's, so it wins and blackholes -- this design's own warning arriving in its implementation. They now route through the hub unless one end can be dialled. And Docker sets the FORWARD policy to DROP, so a hub with ip_forward enabled carried nothing between its spokes. The substrate at tier 1 silently breaks the network at tier 2, and nothing in either tier's state says so. The hub inserts its own rule above those chains and removes it on the way down. Two weak tests found by injection along the way: one asserted the keepalive rule only against the hub, whose peer entries happen not to set that field at all, so it tested an absence; the other checked the firewall rules by looking for FORWARD anywhere, which the PostDown line satisfies on its own. |
||
|
|
f44e73d286 |
The mesh computes a private network it cannot impersonate
The first thing the control plane decides rather than relays. Every node's peer list is derived from every node at once, which is what makes this control-plane work by definition: no node has that view. A hub, with direct peering between nodes at the same site. Not a full mesh, and the reason is a property of WireGuard rather than a preference -- there is no failover, so a more specific route to a dead endpoint blackholes instead of falling back. A node gets exactly one path to any peer, because two would mean one of them silently swallowing traffic. A roaming node is hub-only for the same reason. Reachability and the hub are declared, never inferred from an address. The address is evidence and is not the fact: carrier-grade NAT looks public and is not, a routable address behind a closed firewall looks public and is not, and the regular expression that used to decide it got the lab wrong too. Hub election by address prefix failed silently when nobody knew the convention. No private key travels, and that is the whole design. The node generated its own keypair and kept the private half; the configuration points at a file the node wrote, using WireGuard's own PostUp. So the control plane composes a complete configuration for a node it cannot pretend to be -- it knows every public key and holds none of the private ones. Delivered as an ordinary declaration: a package, a file and a service. The host does not know what a private network is and does not learn one. There is a test holding that line, because the moment connectivity needs a new shape in tier 0 is the moment the host stops being small enough to trust. The generated file is written to be read: each peer says why it is there, a peer with no endpoint says why it has none, and the header says not to edit it -- an edit survives until the graph next changes and then vanishes, which is worse than never being applied, because the machine works and then stops and nothing changed that anybody remembers. Fault injection found one weak test. The keepalive rule was asserted only against the hub, whose peer entries happen not to set the field at all, so it was testing an absence rather than the rule. It now checks two direct peers where one is reachable and one is not. |