a82bfb41f2fa8122da0aeea203f441318f61de3a
124
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0bbb5c6838 |
Builds have a history, and failures are rows like any other
A build result was answered to whoever asked and kept nowhere. So "when did this last build", "why did it fail" and "which machine built what is running" had no answer, and a build nobody was waiting for was reported into the void — which is the same as not reporting it. Failures are recorded too, and that is the point rather than a detail: a failed build that leaves no trace is indistinguishable from one nobody asked for, and the difference is the whole of whether somebody should be looking at something. A build that never learned what it was building keeps the repository, because that is what a person goes and looks at. Recording is idempotent on the correlation id, because a result can arrive twice — as the answer to whoever asked, and on the exchange when nobody was. Two rows would show one build as two, and which is real is not answerable afterwards. The serving control plane now binds `built` as well, so results from builds it did not ask for are kept. It refuses them loudly when it has nowhere to put them rather than dropping them, so the broker's own counters show something arriving that nothing handles. `builds [<module>]` reads it: what happened lately across the mesh, or what has happened to one module — the first asked after something goes wrong, the second when deciding whether to trust something. What was published is kept with the build, so a digest traces back to what made it without holding the manifest twice in a place that can disagree with the first. |
||
|
|
421fe73dce |
The mesh builds: a machine takes the work, and the catalogue shows it
A build is work, not state. Everything else the control plane sends a node is a declaration — this is what you should be — reconciled forever. A build happens once and is finished. Putting it in a declaration would mean rebuilding on every reconcile, or a declaration carrying "and I already did this", which is state about an event rather than about a machine. So it travels on its own queue and the answer comes back correlated. One queue, so several build machines share the work and each request is done exactly once — which a per-machine routing key would not give. mesh-builder is the program a build machine runs. Not the control plane, which must not run commands on a machine; not the host, which would then need a container runtime and git everywhere to do something almost no machine will ever do. It holds its own broker credential and nothing else. Three properties that are decisions: - a request is acknowledged only once the answer is away, so a builder that dies mid-build leaves the work for another machine rather than losing it with nobody ever hearing why - one build at a time. Five at once against one runtime finishes all five slower than it would have finished the first, and the queue is what shares work between machines - a failure is a RESULT. A build that fails silently is indistinguishable from a builder that is not running, and those want different responses And `module list` is a catalogue: what exists, at which version, built from which commit or handed over by hand or shipped with the control plane, whether it is behind its source, and which machines run it. All of that was recorded from the first build and none of it was shown, so "is this current?" could only be answered by reading the database. Proven against a real broker, registry and store: the mesh asked, a builder consumed, built, published, answered; the manifest was recorded with its commit; the source moved and the catalogue said "behind"; rebuilding caught it up with a new digest because the content changed. |
||
|
|
7d033ad9f6 |
Publish to the registry, and a command that builds a repository
One store, and it is the registry the bootstrap already pulls from. An OCI registry is a content-addressed blob store that also understands images: PUT a blob and it is retrievable at /v2/<name>/blobs/sha256:… for ever, by digest. An archive is a content-addressed blob. A second store beside it was considered and is the right answer for objects that are mutable, need per-reader access, or are not build output — somebody's uploads, a backup, a thing with a lifecycle. None of that describes a digest-pinned archive, and running a second service to hold one kind of immutable blob is two things to run, two to back up, and two ways for an artifact to be missing. Overturnable by reading: the manifest carries a URL and a digest, and neither says what served it. `build <repository>` clones, reads module.json, builds what it declares, publishes, and records the manifest with the commit it came from. It is a command rather than something the control plane does on its own, because building runs things on a machine and what the control plane may send a machine is bounded by the declaration language. This is the shape the builder module takes when it is given work over the broker. Proven end to end on a real repository and a real registry: a shell module with a package, a user and a dotfile archive built, published, fetched back at the digest it declared, rebuilt to the same digest, and its manifest accepted by the host's own parser — including `user` and `archive`, which did not exist this morning. A tag is never accepted as a pin, and a blob already stored is not sent again — it is named by its content, so re-uploading asks the registry to store what it already has under the name it already has. |
||
|
|
02d1020bce |
The token carries the node's name
Found by raising a mesh end to end. The broker account a joining node authenticates as is named after the node, and exists before that machine has been told anything — so the node has to know its name before the mesh can tell it. Without it, enrolment fails at the broker with an empty username, which says nothing about why. Not a secret, and the issuer already knows it. The wire-format test now covers it, so a rename on either side fails in both repositories rather than at enrolment on a real machine. |
||
|
|
c3046dcf56 |
A provider is told who its consumers are, and a reference provisioner
Contributions were node-local, so a mesh-scoped provider — the one case that most needs them — never heard from its consumers. A database was given a password and no idea what to create it for. Cross-node consumers now reach the provider's `receives` file, merged in with the ones on its own machine: from the provider's side they are the same thing, and a provider that had to read two lists would read one of them. Each names the file its credential is in rather than carrying it, because the mesh discarded the value and could not put it there. The readable half therefore stays readable. And examples/postgres-provisioner, which is the last step: it reads what the host wrote and makes PostgreSQL accept it. Explicitly not part of the control plane — the control plane decides and never touches a machine. This runs on the machine and touches it, and a real one ships with the module that ships PostgreSQL. It lives here because this is where the contract is defined, written as something that runs so it can be read. It reconciles rather than applying a change, because it is never told what changed. Three things that follow, and each is a fault somebody has shipped: - the password is set every time, not only on creation, or a rotation reports success and changes nothing - what it made and nobody asks for any more is revoked, or a departed consumer keeps a working login for ever - what it did not make is left alone, or it cannot be run on a database that predates it Proven in the lab against a real PostgreSQL, each assertion confirmed to fail with the behaviour removed. The suite is in mesh-lab, which also records the two ways the test itself was wrong first. |
||
|
|
20f78cd5f1 |
Credentials the mesh delivers and cannot read
HAL keeps env vars in the registry, encrypted at rest. Its own tooling records what that bought and what it did not. `secret_locate` matches by value rather than by name — because the same password sits in mesh_provisions, in module_env, in each node's .env in plain text, and inside every connection string composed from it, and its documentation says those URL copies "are often the only copies actually in use". And a query against the encrypted column returns zero rows and proves nothing, so auditing moved to the decrypted copies on the nodes. Two faults there, and encryption at rest addresses neither: the control plane can read what it stores, so a copy of the database is a copy of every credential; and one secret has many homes with nothing tracking them. So here the mesh generates a password, seals it to each end with keys those nodes generated, stores both blobs, and discards the plaintext. It cannot read what it holds. Neither can the broker relaying it. And nothing is composed centrally — a connection string is assembled on the machine that needs one — so no copy is ever minted in a shape nothing tracks. `Compromise of a node is compromise of that node` (ADR 0004) is now true of secrets, not only of identity. Two files rather than one, because the mesh cannot compose a document containing a value it discarded: `binds` carries the readable facts, `secrets` carries the credential alone. The readable half stays readable in the declaration; the secret half changes only when the secret does, which makes restart-on precise. The provider gets a directory, one file per consumer, for the same reason. It is made once and kept — regenerating per declaration would restart both ends on every push, and the password a provider was told to create would never be the one its consumer was given. It is remade when either end's sealing key changes, and both ends learn the new one in the same push, so there is no window where half the mesh holds a dead credential. Two tests found passing for the wrong reason, both caught because their injection came back clean: - the provider's copy was asserted non-empty, which reads the same whichever column is selected. It now opens the blob with the provider's own key. - RotateSecret deleted and re-created; the re-create was dead, because the next read makes one anyway. Removed, and a second path to the same act is how two ends come to disagree. And one real fault: three places built a declaration, and the one behind `--json` predated credentials, so it silently produced a declaration missing them — a difference between what `plan` showed and what anything reading `--json` got. There is one path now. |
||
|
|
c4782ae2fd |
An app is told where its database is
Knowing that a machine needs the anchor's database is useless to the
program that needs it unless the program is told. It knew; nothing was
written anywhere it could read.
Two fields, mirroring contributes/receives in the other direction:
serves: {database: {port: 5432, driver: postgres}} on the provider
binds: {database: /etc/app/database.json} on the consumer
The provider says what a consumer needs to know; the mesh adds the half
only it has — which machine, and what that machine is called on the
private network. The file says, in itself, that it carries no credential
and why. A missing field looks like a bug; a stated absence looks like a
boundary.
Binding something answered on this machine writes nothing. A file saying
"it is on this node" is a fact nobody needs and one more thing to keep
true.
And two machines that share no private network are refused rather than
wired together. An app here and a database there with no path between
them is a mesh that reports itself configured and does not work — the
failure surfaces as a connection timing out, which is the slowest place
to find it. This is checkable now only because the network became
something a machine is given rather than something it has by having an
address.
One fault, found by running it: working out who is on the private network
resolved the mesh, and resolving the mesh asks who is on the private
network. It hung for two minutes. The comment above the function said not
to do that and the function did it anyway; it now resolves each node
locally, which is the right answer to the question regardless — whether a
machine is on the network depends on what it was assigned, not on what it
takes from others.
|
||
|
|
d4064122d6 |
Where the answer to a requirement is allowed to live
Two different things were both written `requires`. A shell, a display
server and a private network have to be on the machine that needs them.
A database does not — it runs somewhere and is reached over the network.
Both were answered the same way, so requiring a database installed
PostgreSQL on every machine that ran a web application.
What a module provides now carries a scope, the same idea claims already
use, written short in the ordinary case:
"provides": ["shell"]
"provides": [{"name": "database", "scope": "mesh"}]
A mesh-scoped requirement is answered by finding the node already running
it — never by installing it here. Choosing a machine to put a database on
is a decision with consequences, and nothing resolving a web application
should make it silently. With nothing anywhere it refuses and says which
module to assign; with two it refuses and says how to choose.
Choosing is `pin <node> <provision> <from>`, kept per node because that
is the granularity the choice has. A pin at a machine that does not
provide it refuses rather than falling back — a fallback would quietly
move somebody's data. One provider does not overrule a pin either.
Resolving a node now needs to know what the others offer, and working
that out needs them resolved, so it is two passes: the first answers only
what each node offers, the second answers everything. Nothing is ever
declared from the first.
A node's plan says what it takes from elsewhere. It is the only part of a
set that stops working when a different machine goes away, and nothing
else in that output would have said so. It is also where a credential
will hang once there is a mechanism for handing one back.
One test found passing for the wrong reason: it read pins through a join
on the provider, which hides a dangling row whether or not it was cleaned
up. It counts rows now, and bites when the cascade is removed.
|
||
|
|
5a3a87e8c3 |
A module can tell its provider what it needs
`requires` said a thing must be there. It never said what to do with it,
so a web application requiring a reverse proxy had nowhere to put "this
name, this port". The two modules that needed it most went round the
outside and opened a connection to the control plane's database, which is
why every node holds a credential to it permanently.
Two fields close it:
contributes: {reverse-proxy: {host: board, port: 8080}}
receives: {reverse-proxy: /etc/traefik/dynamic/mesh.json}
The control plane collects every contribution on a node and writes them
to the path the provider named, ordered by module so the file does not
churn. Contributing to something is requiring it — asking to be published
means a publisher must exist, and a module that had to say both would
eventually say one.
The control plane does not know what a reverse proxy is and does not
write one's configuration. It delivers facts; the module turns them into
whatever it runs. That is why swapping the proxy touches nothing that
publishes through it, and why the host needs no new vocabulary — a
received file is a file.
Settings reach a contribution the same way they reach a file, because a
hostname is exactly what differs between one mesh and the next.
Two things found by running it:
- the file had a `//` header, so it said "do not edit" to a person and
failed to parse for the program meant to read it. The note is inside
the document now.
- a provider with no consumers gets an empty file rather than none. It
cannot otherwise tell "nothing asked for me" from "the mesh never
wrote it", and those want different responses.
Also `plan <node> --json`, which is how the declaration gets handed to
the host's own parser.
|
||
|
|
44d134ba25 |
Networking is a module, and a domain module is how you avoid choosing
Connectivity was code beside the module system doing the module system's
job: every machine with an address was on the private network and there
was no way to keep one off.
A manifest can now say its resources are computed by the control plane,
which is what a peer list needs — it is derived from every machine at
once, so nothing could be written in advance. The network is a module
from there on: assigned, resolved, settled, and absent from a machine
nobody gave it to.
Three modules rather than one, because WireGuard is one VPN of several:
mesh-wireguard provides private-network, mesh-addressing
claims the-private-network, one per node
mesh-names provides name-resolution, requires mesh-addressing
networking requires both, and ships no files of its own
The last is the point. Most people want the network up and do not want
to choose a VPN, so `assign networking` takes the only answer to each
requirement silently. The day the catalogue holds a second one there are
two answers, the resolver refuses and names them, and choosing is
assigning the one you want. No flavor field, nothing to configure.
Names left the WireGuard declaration for their own module. They would be
identical over a different private network, and bundling them made one
module out of two things.
Three faults the walk found:
- choosing tailscale still installed WireGuard, dragged back in by the
names needing the mesh's own addresses. Caught now by a claim: running
two VPNs is fine, being *the* mesh network is singular.
- a requirement wanted by two modules was reported twice, identically.
- "this mesh has no hub" was reported when the real cause was that a
node could not be resolved at all. It now names the node and the why.
And a test that asserts the manifests actually shipped, after the claim
went missing from the real one while every test stayed green.
|
||
|
|
65ade756f2 |
Settings: changing a module's config without editing its file
Managed files are generated and never edited, so somebody's intention about one has to live where the generator can see it. It does now: the module ships defaults, settings go over the top by key, and the file is produced from both. Upstream can rewrite its half freely and the keys somebody chose survive. Two layers, both from the start. The mesh's settings for a module, then one machine's over those. A node that differs is expressed by differing, rather than by restating everything the rest already say -- which would pin all of it against future changes for no reason. An override beats a default and there is nothing to resolve. A setting is a statement about that key made deliberately; the default was only ever what to do in the absence of one. So when upstream changes a key somebody has set, there is no conflict, no merge markers, and nothing to ask. Nested blocks merge and lists are replaced whole. Setting one field of a block must not delete its siblings, or every setting would restate the whole block and pin all of it. A list that merged element-wise could neither be shortened nor reordered, and there is no correct guess about which element is "the same one". A module can keep specific keys for itself -- a socket path its own code depends on -- and setting one is REFUSED rather than ignored. A setting quietly dropped is somebody believing they changed something. Settings that reach nothing are named at the moment they would be used, not discovered later by the machine not behaving differently. `plan --files` prints what a machine would be given before it is sent, because "1 resource" does not tell you whether the merge landed. One test kept with a note that it does not defend this code: output stability comes from Go's encoder sorting map keys, so it passes with the merging removed. Worth having as the thing that would catch a change of encoder, but it is not evidence about anything written here, and it was checked. |
||
|
|
653e232f1c |
The mesh knows where a module came from, and whether it is behind
Delivery is a comparison, not a pipeline: the control plane holds what source
exists and what has been built from it, and the difference is the work. Both
halves are written down now, so "is this current" is a question about two
columns rather than something you find out by building.
`status` answers "did my change go out?", which ADR 0010 names as the real risk
of replacing a pipeline with a comparison -- it is answerable today by opening
a pipeline, and something had to replace that.
zsh holds 4f2a9c1e, source has 9e3b7d2a
running on laptop
The machines are the point. A module being out of date is a fact about the
catalogue; which machines are running last week's version is the thing with
consequences.
Three things this had to get right.
A module with no source is never behind -- it was handed over directly, which
is how a one-off arrives, and saying "out of date" about it would be inventing
a comparison against nothing.
A source nobody has checked is not behind either. Reporting it as behind would
put every module on the list the moment provenance was recorded, which makes
the list say nothing. Fault injection found this: my first test passed with the
guard removed, because both halves were empty strings and compared equal. The
case that actually needed it -- a known commit and an unknown head -- was
untested.
And handing over a manifest by hand does not erase where the module normally
comes from. Fixing something in a hurry is legitimate; silently forgetting its
origin is not, because that record is the only thing that would say afterwards
that a machine is running something nobody can rebuild.
Also fixed the flag parsing, which stopped at the first positional argument and
silently ignored every flag after it -- so `module add thing.json --source x`
recorded no source at all and said it had succeeded. The host's own parser
documents this exact footgun and I wrote it again anyway.
|
||
|
|
931a3a19a5 |
Taking a module off a node takes it off the machine
The half of the module system that was built and never proved. Unassigning i3 removed i3's file AND xorg's, because xorg was only there to satisfy i3 -- the node's own record agrees, and the resolution the mesh sends no longer mentions either. That works because a declaration removes what the mesh previously declared and nothing else, which is 04-ISSUES/010's fix carrying its weight here: the substrate the machine raised for itself is untouched by any of it. Tests for the storage layer, which had none. The ones worth naming: A module a machine is running cannot be forgotten -- not a fault, it means the mesh would lose the ability to describe what is on that machine. Removing a node DOES take its assignments, and the asymmetry is deliberate: a node that is gone cannot be running anything. A node that has never reported has NO capabilities rather than all of them. That refuses anything needing one, which is wrong but visible -- where assuming it can do everything would assign work it cannot do and find out on the machine. And a capability the node reported as ABSENT is not counted: reading the list without the verdict would let a module onto a machine that said no. `overlay push` is gone, replaced by `push`, which sends a node its network and its modules as one declaration. Two commands that overlap is how a mesh ends up half-configured by whichever was run. The old name answers with where to go, and answers before opening a database -- needing one would turn a redirect into a connection error. |
||
|
|
409cd16a09 |
The mesh decides what a node runs
The gap that has been named at the end of every report for a week. Until now a
declaration came from a person handing over a file; now it comes from what was
assigned, resolved against the catalogue, and the control plane is deciding
rather than relaying.
Everything from the module conversation, built and run on real machines:
assign laptop i3 -> accepted, brings xorg, because nothing else provides
it and there was no choice to make
assign laptop sway -> refused: xorg and wayland both claim the-seat
assign laptop editor -> refused: three modules provide a shell -- bash,
fish, zsh -- choose one
assign laptop zsh -> accepted, and the editor's requirement is answered
bash, fish beside it -> fine, nothing is claimed
Claims rather than pairwise exclusion, so a third display server would say what
it claims and need no edit to xorg or wayland. Scoped to node, site or mesh:
two DHCP servers at one site collide and at two sites do not, and the mesh-wide
one is the hub said as a claim instead of hard-coded.
Some conflicts cost no manifest field at all. The refusal above names the seat
AND the two files, because the mesh already holds every resource of every
module -- neither i3 nor sway knows the other exists.
Resource identities carry their module, so two modules may both call something
"config" without the second silently replacing the first. What a service
reflects is qualified the same way, or it would name a resource that no longer
exists and stop being restarted when its own configuration changes.
Nothing is sent until every node resolves. A push that configured three and
refused on the fourth would leave the mesh in a state nobody asked for, and the
fourth is exactly where a claim collision appears.
One real flaw found by using it rather than by testing it: assigning zsh did
not satisfy a requirement for a shell. Requirements were counted against the
catalogue without first asking what the set already offers, so "choose one and
assign it" named three modules and then ignored the one you chose. The remedy
was useless and every test passed.
|
||
|
|
f0cff88172 |
The mesh knows who is out of touch
09-the-node-lifecycle asks for this in as many words -- *how long it has been disconnected is a fact the mesh must hold, and nothing holds it today. Without it, a node running last month's assignments looks exactly like one that is current.* Now it holds it. `node list` says "here", "out of touch 4m", or "never spoken", and the third is kept distinct from the second on purpose: a node that has never spoken did not finish joining, and a node last heard from a month ago is running a month-old picture of the mesh. Those need different responses from a person. A bare word that a node is there moves last_seen and touches nothing else. It is not an account of what the machine holds, and recording it as one would replace the recovery copy with an empty list every minute -- so a rebuilding node would then be told it owns nothing and remove whatever it found. There is a test for exactly that. Heard is silent in the log. A node saying it is there every minute would fill the log with the ordinary case, and a log where the ordinary case is loud is a log nobody reads. Verified in the lab across the threshold, both directions. |
||
|
|
fc1417be72 |
Names, from the same graph as the network
Step 5 of the connectivity order. Every node's internal name resolves to its overlay address, on every node, computed centrally because it needs every node at once. Under `.internal`, which IANA reserved for exactly this in 2024 -- a name there can never collide with a public one, so an internal name that leaks into a public resolver fails rather than reaching a stranger's machine. The suffix is settable for a mesh that wants its own. Delivered in the same declaration as the peer list rather than a second one. A node holding the peers and not the names, or the reverse, is half on the network for as long as that lasts. This is not the /etc/hosts floor the design removes. That floor existed because a node had to reach the mesh's database before its own DNS worked -- a fallback for a circularity that is now gone. This is the mechanism: the complete set of names, generated whole and owned by the mesh, rather than a patch written underneath something else. A resolver daemon becomes necessary when names are wanted that are not one-per-node, and that is not yet true. A node resolves its own name to its overlay address rather than a loopback, because a service binding to the name it was given would otherwise listen somewhere nothing else can reach -- and the failure would appear on every other machine rather than that one. A node with no address gets no name. A name resolving to nothing is worse than no name: connecting to an address that does not answer hangs, where a name that does not resolve fails at once and says which name it was. Found while writing it: a test asserting every file in the declaration is mode 0600 would have forced /etc/hosts to 0600 and broken every lookup on the machine, to protect a file that is not secret. Verified in the lab: three machines, nine name lookups, each resolving to the right overlay address and reaching it. |
||
|
|
8b974deb42 |
A working private network, and four reasons it did not work
Three machines across two sites, two of them behind no reachable address, all nine paths open. The mesh computes the graph, delivers it as a declaration, and the nodes bring it up. Every fault below looked like success from inside the mesh: the graph was right, the files were right, the services were up, every node reported it had applied. None was reachable by reasoning. A running interface does not re-read its configuration. A node joins, every existing node's peer list changes, the file is replaced -- and the service is already running, so nothing reloads it. Fixed as declared state rather than a command: the service must reflect the file. A command to restart would be an action, and the link may not carry one. The host refused exactly that, which is how this shape was arrived at. A hub sharing a site with a spoke appeared twice in that spoke's peer list -- once as a direct peer, once as the route of last resort. WireGuard takes one entry per key and refuses the file. The ordinary shape of a small mesh, and in none of the tests written before it ran. Two nodes at one site that neither can be dialled were peered directly. Nobody opens the path, and the direct route is more specific than the hub's, so it wins and blackholes -- this design's own warning arriving in its implementation. They now route through the hub unless one end can be dialled. And Docker sets the FORWARD policy to DROP, so a hub with ip_forward enabled carried nothing between its spokes. The substrate at tier 1 silently breaks the network at tier 2, and nothing in either tier's state says so. The hub inserts its own rule above those chains and removes it on the way down. Two weak tests found by injection along the way: one asserted the keepalive rule only against the hub, whose peer entries happen not to set that field at all, so it tested an absence; the other checked the firewall rules by looking for FORWARD anywhere, which the PostDown line satisfies on its own. |
||
|
|
f563ababa1 |
The mesh keeps a copy of what each node owns
novox/hq 09-the-node-lifecycle asks for this and it was missing: the host reports what it owns and the mesh keeps the last report. A backup, never a source -- nothing decides anything from it, and a node that disagrees with it wins, because the node is the one that can see the machine. Its point is the orphans. A node that loses its state file currently strands whatever it applied: nothing on the machine knows those resources were the mesh's doing, so nothing removes them. With this, a rebuilt node receives both the declaration and the record of what it previously owned. Never reported and reported nothing are kept apart, and that is the whole care in it. A node that applied nothing holds nothing; a node that has never spoken is unknown -- and handing back an empty list for the second would tell a rebuilding node it owns nothing and have it remove whatever it found. The age comes back with the answer rather than being left for the caller to go and find. An answer about a machine is worth much less without one, and this repository has already been bitten by a cache with no age on it. A refusal or a partial failure moves last_seen and nothing else: neither is an account of what the machine holds, and recording one as though it were would tell a rebuilding node to remove what it still has. |
||
|
|
bbbc860188 |
The control plane declares, and hears back
`declare` sends a node a signed declaration; `serve` now also consumes reports. Signed over the exact bytes published, which is what the node verifies. Anything re-encoding in between would sign one thing and check another, and a difference in key order alone would have a node refuse a declaration that was genuinely the mesh's. Sent to the node's queue directly rather than through the exchange: a declaration is for one node, and routing by name through a shared exchange means a binding per node that nothing removes when a node is retired. Enrolment now issues the node its own broker password, replacing the token's secret, and tells it the broker address, the fingerprint and the signing key -- so a node can reconnect after a restart without a person and a new token, which is what makes disconnection ordinary rather than a crisis. A report is a statement, not a write. What a node says it applied is its own account of its own machine, kept as a copy for recovery rather than as a source. |
||
|
|
46e760fc94 |
The control plane serves, and a node can join
There was no chicken-and-egg to solve. The mesh runs the broker, so it creates the node's account when it issues the token, and the one-time secret is that account's password. A joining node's first connection is already authenticated; enrolment is what it says once it is in. I had been treating this as a decision that needed taking, and it did not. The account is per node and scoped: it may read its own queue, write to the one exchange, and configure nothing else. The patterns are anchored and the node name is constrained to characters that cannot widen them, because a name carrying a dot or a star would silently let that node read everybody's queues. `serve` is the control plane running: one connection, one queue, one consumer. One deliberately -- two consumers on a queue get round-robined and each receives half of what it expects, which has happened on this project before, between a module's daemon and its capability server. Enrolment spends the token first, in the single statement that both finds and marks it, and only then records the key. That order is the order things become irreversible: recording a key for a node whose token turned out to be spent would leave the mesh believing a machine that never had the right to join. Refusals are one message for every reason. The log says which, where an operator can see it; the node is told only that the token cannot be used. Verified in the lab, on a sealed machine, through the whole first-node path. |
||
|
|
ea6569277d |
A token with all four parts
Given the broker's address and its certificate, mesh-control now issues a token carrying everything ADR 0004 asks for: where to connect, what to expect there, whose signature to believe afterwards, and a one-time right to join. Verified by decoding one and checking the fingerprint against `openssl x509 | sha256sum` -- they match. The fingerprint is derived from the certificate on disk and never configured. A configured pin can drift from the certificate it describes, and a drifted pin is worse than none: every node issued a token during the drift refuses to connect, and the failure looks like an attack rather than a mistake. Computed over DER, which is what a client sees on the wire. Hashing the PEM text instead would mean the same certificate, re-wrapped with different line endings, produced a different pin -- there is a test for exactly that, and one for pointing this at tls.key by mistake, which would otherwise produce a confident pin over the wrong file. Having no broker stays a state rather than a failure: a control plane holds records and a signing key without one. Having half a broker is refused, because a token with an address and nothing to check it against invites a node to trust whatever answers. Fault injection caught the same weak test I wrote earlier in the day -- asking whether something failed rather than why, so deleting the guard changed nothing because it failed one line later anyway. Both are now asserted on the reason. |
||
|
|
7553af6c5a |
The control plane's signing key, and a second context to hold it
Everything is blocked on what a node presents to prove which node it is. This builds the other direction, which is not blocked: what a node believes. identity is the second of the seven contexts. It holds an Ed25519 signing key the control plane generates once, whose public half now travels in every enrolment token. A node believes a declaration because it carries a signature that key made -- pinning only the broker would make the control plane's authority transitive, and since the host applies whatever the link delivers, a compromised broker forging declarations is the whole machine. Establishing the key is idempotent, and it has to be: a second key generated by a restart is a mesh where every node holds the wrong public half, so every declaration is refused by every node with nothing visibly wrong. The guarantee is a partial unique index plus a read-back, not the check before the insert -- six processes racing to establish all agree on one key, and there is a test that runs them. Tokens are now one line of base64 carrying three of their four parts. The missing two are the broker's address and its certificate fingerprint, both step 5 of the bootstrap. The command prints the token and names what is missing rather than emitting something that looks usable. The second context also tests a claim this repository had made and never checked: that a context reaches only its own store. Two databases, two credentials, no setting that reaches both. Running migrate with one stops and names the grant it lacks -- verified, not asserted. Assembling a token needs a node record from one and a key from the other, and neither reads the other's store; the process holding both grants asks each for its part. 45 tests, none skipped. Fault injection found one test whose property is enforced somewhere other than where I injected -- idempotency comes from the database constraint, not from the early return, which is what the code comment already said. |
||
|
|
66768208d2 |
Node records, and the right to join once
The next step after the schema: inventory now holds node records and enrolment tokens, and mesh-control has the commands to work with them. A token is issued for a node record, which is where re-enrolment gets decided -- what an identity binds to is settled when the token is made, not when it is presented, so the machine presenting one does not need to know whether it is joining or returning. What the token guarantees, each with a test confirmed to fail when the behaviour is removed: the secret is 256 random bits, shown once and stored only as a hash; it works exactly once; it stops working when it expires; issuing again for a node invalidates the outstanding one, because two live tokens are two machines able to join as the same node. Redemption is a single statement that finds and spends together, so eight concurrent attempts on one secret produce exactly one winner rather than a race between a check and a write. Refusals are deliberately identical for unknown, spent and expired. Somebody guessing must not learn which guess was a real token that had merely aged out. SHA-256 rather than a password hash, and that is a choice not a shortcut: the secret is high-entropy random, so there is nothing to guess and a slow hash would buy nothing while making every redemption expensive. It stops before what a node receives in exchange. What a machine presents afterwards to prove it is that node is not decided anywhere, and a migration is the most expensive place here to guess. So a token carries one of the four things ADR 0004 requires. The command prints the secret and then says exactly that -- the broker's address, its certificate fingerprint and the control plane's signing identity do not exist yet. Better than emitting something that looks complete and silently cannot be used. |
||
|
|
306c4ca13b |
The control plane, as far as identity
Tier 2 exists now. It holds one context of seven, inventory, and does one thing with it: brings its schema up to date. That is step 3 of the substrate bootstrap -- the step the first node cannot get past. Verified against a real PostgreSQL, with the built binary: applied 0001-nodes, reported 'already up to date' on the second run, and the node table is there with the index and the unique constraint the migration asks for. Written in Go, and the image is FROM scratch holding one file. Confirmed by unpacking it. That is the whole argument of ADR 0024: the bundle pins this image by digest and runs it where nothing can check it, so everything in it is something a person has to audit before trusting a first node. Exclusive store ownership is built as a rule about credentials rather than about intentions. There is no mesh-wide connection setting and no way to ask for one -- a context reads MESH_STORE_<ITS OWN NAME> and holds nothing else, so reaching another context's store needs a new variable, which is visible in the declaration that runs it. The migration runner is mostly refusals: an edited migration that already ran, a migration numbered below one that has run, duplicate numbers, misnamed files, empty files. All stop rather than warn, because at the moment any of them is true nobody knows what the database holds. It stops before identity, deliberately. What a node presents to prove who it is has not been decided anywhere, and a migration is the most expensive place in this system to guess. Two tests did not defend what they claimed, and both are fixed rather than removed. One asked only whether Open returned an error, which it did either way -- a bad context name and a missing credential both fail, so deleting the name check changed nothing. The other claimed to prove the migration runs in a transaction, but PostgreSQL already wraps a multi-statement query in one of its own, so it passed with the transaction taken out. What the transaction actually buys is that the schema change and the row recording it commit together, and there is now a test for that which fails when they are split. |