fcdb065660851b17f2bcf1f558a782146f45f539
23
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2b82872ac3 |
A build keeps what it was told
The builder announces a build with the resolved manifest, the path inside the repository, and every artifact it stood on. The control plane received all of it and kept none of it. That was survivable while the catalogue heard the same announcement directly. It stops being survivable the moment the catalogue was not there to hear it — which on a fresh mesh is always, and always for the same modules: the shared base, the store the catalogue runs on, and the catalogue itself are each necessarily built BEFORE the catalogue exists to hear about them. The graph's foundation is the part the graph never sees. Replaying those builds needs what they said, not a summary. Without the manifest there are no requires/provides edges; without `against` there are no build edges, which are the ones that answer "a base moved, what must be rebuilt". A replay carrying neither would restore the module list and leave the question the catalogue exists for still wrong, while looking fixed. Kept null rather than empty where a build predates this, so a replay can say it is holding nothing instead of inventing an empty declaration for a module that certainly had one. And `built_against`, not `built_on`: that column exists and means the machine, which is a different fact about a different subject. Toward novox/hq 04-ISSUES/050. Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx |
||
|
|
588aa424e2 |
The mesh acts on what the catalogue decided a build meant
The builder says what it built and the catalogue decides whether that was an upgrade. Only the control plane knows which machines run the thing, so it is the one that acts — and what it does is a choice somebody recorded, not a behaviour compiled in: record that they are behind, or send it, one machine at a time or together. Recording is the absence of an action rather than a second path: a machine not running what the mesh would send it is already something the mesh reports. Defaulted to recording. A mesh that rolls out everything it builds the moment it builds it is reasonable to want and a bad thing to arrive by default — the first module to inherit it would be the control plane, upgrading itself out from under the push applying it. |
||
|
|
f151de103f |
Build a module from a repository and a path within it
The builder cloned a repository and read the manifest at its root, which means one repository per module. Nothing we have is shaped that way, so the builder could be asked to build nothing that exists (novox/hq ADR 0069). The path travels the whole way — named when asking, carried in the request, used to read the manifest and as the context everything is produced from, echoed back in the result, and recorded as part of where a module came from. Without that last part the mesh could notice a module was behind its source and then be unable to rebuild it, which is the worst of both. A path climbing out of the clone is refused: a machine whose job is building other people's repositories must not read whatever else is on its disk. Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx |
||
|
|
c4030947b0 |
The routing record is 0066, not 0056
0056 is 'the authority is the control plane, not a database'. A citation pointing at the wrong decision is worse than none: it reads as corroboration. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF |
||
|
|
232862315c |
catalogue: compose a route's name from a label and its node's domain, and resolve it in-mesh
A public route used to carry its whole hostname as a literal in the module manifest, so running the same catalogue against a different domain meant overriding that literal on every routed module, per node. The mesh was, in effect, holding a map of names to services: the one thing it should never hold, because the subdomain is the operator's choice and the domain is the node's. Compose instead. A route contribution carries a `label` (the subdomain); a node carries its `public_domain` as node-level configuration; the mesh joins `<label>.<public-domain>` and grants exactly that, interpreting neither half. Held as a node property beside the node's other node-level facts (endpoint, site, overlay address), not in a module's settings — the ADR calls it node-level, and the settings table is keyed per module. Additive, so an unmigrated catalogue keeps working: a contribution that still carries a full `name` and no `label` passes through unchanged, and the catalogue can migrate module by module. A labelled contribution on a node with no public domain composes nothing, reading downstream as a route that named no host. And propagate: each granted route name is published into internal resolution mesh-wide, mapped to the node that serves it, alongside the `<node>.internal` names every container already gets. So a container — and an internal ACME validator, which cannot complete a challenge for a name it cannot reach — resolves a routed name to the proxy that serves it. Name-agnostic throughout: the mesh propagates whatever names it was told to serve and knows nothing about what they mean. novox/hq 02-DECISIONS/0056 Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF |
||
|
|
1b63e21c0f |
Caught up is an equality, not an ordering
The report carries the digest of the declaration it applied (mesh-host 8211d8b), and the mesh stores it beside the outcome. `reported` rows in the status JSON now say `current`: whether the machine's last word names the declaration last sent. Not derivable from the timestamps beside it, which is why they were not enough: an apply begun under the previous declaration reports after the next send — newer, and still about the old words. The lab lost exactly that race between one test's closing push and the next test's opening one. Empty digests — every host from before reports carried one — read as not current, which errs toward waiting rather than toward asserting on files that are not there yet. |
||
|
|
41f7c51032 |
Assign around what a machine already holds
The other half of ADR 0038, and what 04-ISSUES/028 was actually about. A module can now avoid colliding with another module; until this it could not avoid colliding with the mesh itself. The substrate is not a module. A node raises it from the bundle it carries before any mesh exists, so the control plane had never heard of the store, the broker, or its own container — and handed a database module 5432, which the store already had. So the machine says. The host records what each resource binds, distinguishing what it carried from what the mesh sent — a distinction that already existed so the two never remove each other — and reports the carried ones. The node states and this context writes, which is the shape of every message between them. What the declaration binds, not what is open. A machine's open ports are a moving target, and assigning around them would mean a port that was free when it was asked for and taken when it was used. Replaced whole each time rather than merged: a machine that gave a port back must be believed about that too, and a set that only grows keeps a port reserved for something no longer there. Tested against a real database, and the tests bite — removing the check hands the module 20000, which the machine had said it holds. |
||
|
|
1f5b70a995 |
The mesh assigns the port, and a module says it once
novox/hq ADR 0038. A module cannot choose a port: it is written once and assigned anywhere, so any number it picks is a guess about a machine it has never seen. A database module met the mesh's own store on 5432 and was told, by a container runtime three layers down, that the port was already allocated. The number used to appear three times in every module — the rule set, what a consumer is told, and what the runtime publishes — agreeing only because one person wrote all three. Now it appears once, in `listens`, and the other two are derived: the container publishes `20000:5432`, the consumer is told 20000, and the rule set opens 20000. An assignment is made once and kept, as a credential is. A port that moved on every declaration would restart both ends each time and hand a consumer a number that was true when it was read. Ports the protocol fixes — mail on 25, submission on 587, DNS on 53 — say so, and are then claims: one holder per machine, and the second is refused by name at assignment. That is the mechanism the mesh already has for what is singular on a machine, pointed at ports. A mapping written the long way is left exactly as it is. Some things must be pinned by hand, and quietly overruling somebody who wrote both halves would be worse than not offering the short form. Still open, and known: the substrate is not a module, so the mesh has never heard of its own store and cannot yet assign around it. That is what 028 will still be about after this. |
||
|
|
0af3ea1acf |
A consumer is a module on a machine, not a machine
novox/hq 04-ISSUES/022. A credential was keyed by provision, consumer node and provider node, so "who is asking" was answered by naming a host. The node this mesh exists to take over runs eight modules against one database server. The symptom had two halves and only one was loud. The provider refused, naming the modules and explaining they would share one credential, which reads as a decision rather than a limit. The consumer did not refuse: it resolved cleanly, wrote one module's credential file and left the others absent — a service that starts and cannot authenticate, with nothing saying why. That is 021 again on a different axis. Three modules wanting one database produced one need, carrying whichever module mentioned it first, because the resolution walk is a work-list over names. The fan-out now happens in one place, after the walk. The record path already did this correctly and said why: a consumer here is a module on a machine. It is the same rule. Downstream: the secret's key gains the consuming module, the grant file is named after both halves, needs are matched by provision and module rather than provision alone, and the provisioners name the role and the access key after the module. The refusal in ContributionsTo is gone because there is nothing left to refuse. Worth stating plainly: without that refusal, gitea's login would have opened keycloak's database. From the provisioner's side it created exactly what it was asked to create. Existing secrets are discarded rather than backfilled. They cannot say which module they were for, and a secret is remade and delivered to both ends on the next push — so this costs one rotation and invents nothing. Also guards the role name against PostgreSQL's 63-byte truncation, which is a notice rather than an error and would reintroduce exactly this collision at a length nobody tests. Three faults injected — the fan-out removed, needs matched by name alone, the grant file named after the machine — each caught. |
||
|
|
5a28434ba8 |
"Behind" means not running what the mesh would send
It meant "failed or refused". So a machine that applied cleanly and whose declaration has since changed was not behind — and novox/hq ADR 0010's question, did my change go out?, was answerable exactly for the machines that broke. For every machine that worked, the answer was silence whether the change had gone out or not, which is the thing replacing a pipeline was supposed not to cost. The mesh now records a digest of what it last sent each machine. A digest rather than the declaration: it can compute what a machine should be at any moment, and keeping a copy would be a second account of it able to disagree with the first. What cannot be recomputed is what was actually sent. Recorded after the send, not before — a digest kept for something that failed to send would make the machine look current for a declaration it never received. Never told stays separate from out of date. The remedy is the same push and the situations are not alike: nobody has ever asked that machine to be anything. And a machine the mesh could not work out is not reported as waiting, because saying so would invent a comparison — that is `plan`'s answer to give. `status` says it and `push --behind` sends it, or the flag would know something the person reading the status does not. |
||
|
|
58c8ab7747 |
A secret the mesh was given is not one the mesh can reinvent
Two kinds live in module_secret and they behaved identically, which is right for one of them. A made secret is the mesh's: when a node regenerates its sealing key the mesh makes another and nothing is lost, because nothing else ever knew the old one. An accepted secret is not. A broker account's password exists because the broker was told about it. Regenerating one puts 32 random bytes where a working credential was — and the machine applies it, reports success, and the program reading it fails to authenticate somewhere else entirely, with the mesh insisting the secret was delivered, which it was. The row now records where the value came from, and a rejoined machine asking for an accepted one is refused with the remedy named: issue it again. No amount of pushing produces a password the broker has never heard of. Found while making the builder a module, which is the first thing to hold one. |
||
|
|
c37d368f65 |
A module may need a secret of its own, and the provisioner watches
Two things, both found by trying to write a real postgres module and discovering it could not be said. A database has a superuser password, a broker an administrator, a registry an account. None of them is *for* anybody — they are not the credential a consumer is given, and the mechanism that hands those out has a consumer in the middle of it. So a module may declare what it needs and where to put it, and the mesh generates one per node, seals it, and reads it no more than it reads any other. Per node, deliberately: a module running on three machines has three passwords. One in the manifest instead would put the same secret on every machine that ever runs it, in a file anybody can read, for ever. Made once and kept, or a running database would be handed a password it was not started with; remade when the machine's sealing key changes, like everything else sealed here. A need declared and not made is refused rather than skipped, because a module whose own credential is silently absent starts, fails to authenticate, and the reason is three layers from the machine reporting it. And the provisioner can watch. That is what lets it be a module rather than a binary somebody places: run once, it needs invoking after every declaration by a timer or a unit wired to a file; watching, it is an ordinary long-running service the host already supervises. It polls rather than watching the filesystem, because the host writes atomically — the file is replaced, so a watch on the path stops seeing anything after the first replacement, and a watcher that silently stops working is worse than a poll. Credentials are compared by digest and never held: this runs for as long as the machine is up. |
||
|
|
9681b288aa |
Keep what each machine did, so status can say what is wrong
A node reports back after applying a declaration: it worked, some of it failed, or the whole thing was refused. A refusal or a failure moved last_seen and the reason went to a log line — so "which machine is not doing what it was told" had no answer the next morning, which is the question a mesh exists to answer. Refused and failed are kept as different things, because they are different situations with different remedies: refused means the machine is exactly as it was and what is wrong is in what was sent; failed means it is in a state nobody declared and what is wrong is on the machine. One word for both would make the record say less than the node did. One row per node, replaced. The question is the machine's current state — "this failed an hour ago and then succeeded" is not a machine anybody needs to look at, and a table of every report would bury the ones that matter under the ones that do not. `status` now answers three questions in the order somebody asks them: is anything broken, is anything not answering, is anything out of date. The first has consequences now, the third is a plan for later, and a status leading with the third would bury the first. A machine that has never spoken is reported as quiet rather than as broken — new, switched off and unreachable are not the same as tried and could not. The mapping from a report to an outcome had no test at all, which the injection caught: it is the code deciding which of those situations a machine is in. It has four now, including that a partial report never becomes the account of what the machine holds — the fault that destroyed a substrate once. |
||
|
|
0bbb5c6838 |
Builds have a history, and failures are rows like any other
A build result was answered to whoever asked and kept nowhere. So "when did this last build", "why did it fail" and "which machine built what is running" had no answer, and a build nobody was waiting for was reported into the void — which is the same as not reporting it. Failures are recorded too, and that is the point rather than a detail: a failed build that leaves no trace is indistinguishable from one nobody asked for, and the difference is the whole of whether somebody should be looking at something. A build that never learned what it was building keeps the repository, because that is what a person goes and looks at. Recording is idempotent on the correlation id, because a result can arrive twice — as the answer to whoever asked, and on the exchange when nobody was. Two rows would show one build as two, and which is real is not answerable afterwards. The serving control plane now binds `built` as well, so results from builds it did not ask for are kept. It refuses them loudly when it has nowhere to put them rather than dropping them, so the broker's own counters show something arriving that nothing handles. `builds [<module>]` reads it: what happened lately across the mesh, or what has happened to one module — the first asked after something goes wrong, the second when deciding whether to trust something. What was published is kept with the build, so a digest traces back to what made it without holding the manifest twice in a place that can disagree with the first. |
||
|
|
20f78cd5f1 |
Credentials the mesh delivers and cannot read
HAL keeps env vars in the registry, encrypted at rest. Its own tooling records what that bought and what it did not. `secret_locate` matches by value rather than by name — because the same password sits in mesh_provisions, in module_env, in each node's .env in plain text, and inside every connection string composed from it, and its documentation says those URL copies "are often the only copies actually in use". And a query against the encrypted column returns zero rows and proves nothing, so auditing moved to the decrypted copies on the nodes. Two faults there, and encryption at rest addresses neither: the control plane can read what it stores, so a copy of the database is a copy of every credential; and one secret has many homes with nothing tracking them. So here the mesh generates a password, seals it to each end with keys those nodes generated, stores both blobs, and discards the plaintext. It cannot read what it holds. Neither can the broker relaying it. And nothing is composed centrally — a connection string is assembled on the machine that needs one — so no copy is ever minted in a shape nothing tracks. `Compromise of a node is compromise of that node` (ADR 0004) is now true of secrets, not only of identity. Two files rather than one, because the mesh cannot compose a document containing a value it discarded: `binds` carries the readable facts, `secrets` carries the credential alone. The readable half stays readable in the declaration; the secret half changes only when the secret does, which makes restart-on precise. The provider gets a directory, one file per consumer, for the same reason. It is made once and kept — regenerating per declaration would restart both ends on every push, and the password a provider was told to create would never be the one its consumer was given. It is remade when either end's sealing key changes, and both ends learn the new one in the same push, so there is no window where half the mesh holds a dead credential. Two tests found passing for the wrong reason, both caught because their injection came back clean: - the provider's copy was asserted non-empty, which reads the same whichever column is selected. It now opens the blob with the provider's own key. - RotateSecret deleted and re-created; the re-create was dead, because the next read makes one anyway. Removed, and a second path to the same act is how two ends come to disagree. And one real fault: three places built a declaration, and the one behind `--json` predated credentials, so it silently produced a declaration missing them — a difference between what `plan` showed and what anything reading `--json` got. There is one path now. |
||
|
|
d4064122d6 |
Where the answer to a requirement is allowed to live
Two different things were both written `requires`. A shell, a display
server and a private network have to be on the machine that needs them.
A database does not — it runs somewhere and is reached over the network.
Both were answered the same way, so requiring a database installed
PostgreSQL on every machine that ran a web application.
What a module provides now carries a scope, the same idea claims already
use, written short in the ordinary case:
"provides": ["shell"]
"provides": [{"name": "database", "scope": "mesh"}]
A mesh-scoped requirement is answered by finding the node already running
it — never by installing it here. Choosing a machine to put a database on
is a decision with consequences, and nothing resolving a web application
should make it silently. With nothing anywhere it refuses and says which
module to assign; with two it refuses and says how to choose.
Choosing is `pin <node> <provision> <from>`, kept per node because that
is the granularity the choice has. A pin at a machine that does not
provide it refuses rather than falling back — a fallback would quietly
move somebody's data. One provider does not overrule a pin either.
Resolving a node now needs to know what the others offer, and working
that out needs them resolved, so it is two passes: the first answers only
what each node offers, the second answers everything. Nothing is ever
declared from the first.
A node's plan says what it takes from elsewhere. It is the only part of a
set that stops working when a different machine goes away, and nothing
else in that output would have said so. It is also where a credential
will hang once there is a mechanism for handing one back.
One test found passing for the wrong reason: it read pins through a join
on the provider, which hides a dangling row whether or not it was cleaned
up. It counts rows now, and bites when the cascade is removed.
|
||
|
|
65ade756f2 |
Settings: changing a module's config without editing its file
Managed files are generated and never edited, so somebody's intention about one has to live where the generator can see it. It does now: the module ships defaults, settings go over the top by key, and the file is produced from both. Upstream can rewrite its half freely and the keys somebody chose survive. Two layers, both from the start. The mesh's settings for a module, then one machine's over those. A node that differs is expressed by differing, rather than by restating everything the rest already say -- which would pin all of it against future changes for no reason. An override beats a default and there is nothing to resolve. A setting is a statement about that key made deliberately; the default was only ever what to do in the absence of one. So when upstream changes a key somebody has set, there is no conflict, no merge markers, and nothing to ask. Nested blocks merge and lists are replaced whole. Setting one field of a block must not delete its siblings, or every setting would restate the whole block and pin all of it. A list that merged element-wise could neither be shortened nor reordered, and there is no correct guess about which element is "the same one". A module can keep specific keys for itself -- a socket path its own code depends on -- and setting one is REFUSED rather than ignored. A setting quietly dropped is somebody believing they changed something. Settings that reach nothing are named at the moment they would be used, not discovered later by the machine not behaving differently. `plan --files` prints what a machine would be given before it is sent, because "1 resource" does not tell you whether the merge landed. One test kept with a note that it does not defend this code: output stability comes from Go's encoder sorting map keys, so it passes with the merging removed. Worth having as the thing that would catch a change of encoder, but it is not evidence about anything written here, and it was checked. |
||
|
|
653e232f1c |
The mesh knows where a module came from, and whether it is behind
Delivery is a comparison, not a pipeline: the control plane holds what source
exists and what has been built from it, and the difference is the work. Both
halves are written down now, so "is this current" is a question about two
columns rather than something you find out by building.
`status` answers "did my change go out?", which ADR 0010 names as the real risk
of replacing a pipeline with a comparison -- it is answerable today by opening
a pipeline, and something had to replace that.
zsh holds 4f2a9c1e, source has 9e3b7d2a
running on laptop
The machines are the point. A module being out of date is a fact about the
catalogue; which machines are running last week's version is the thing with
consequences.
Three things this had to get right.
A module with no source is never behind -- it was handed over directly, which
is how a one-off arrives, and saying "out of date" about it would be inventing
a comparison against nothing.
A source nobody has checked is not behind either. Reporting it as behind would
put every module on the list the moment provenance was recorded, which makes
the list say nothing. Fault injection found this: my first test passed with the
guard removed, because both halves were empty strings and compared equal. The
case that actually needed it -- a known commit and an unknown head -- was
untested.
And handing over a manifest by hand does not erase where the module normally
comes from. Fixing something in a hurry is legitimate; silently forgetting its
origin is not, because that record is the only thing that would say afterwards
that a machine is running something nobody can rebuild.
Also fixed the flag parsing, which stopped at the first positional argument and
silently ignored every flag after it -- so `module add thing.json --source x`
recorded no source at all and said it had succeeded. The host's own parser
documents this exact footgun and I wrote it again anyway.
|
||
|
|
409cd16a09 |
The mesh decides what a node runs
The gap that has been named at the end of every report for a week. Until now a
declaration came from a person handing over a file; now it comes from what was
assigned, resolved against the catalogue, and the control plane is deciding
rather than relaying.
Everything from the module conversation, built and run on real machines:
assign laptop i3 -> accepted, brings xorg, because nothing else provides
it and there was no choice to make
assign laptop sway -> refused: xorg and wayland both claim the-seat
assign laptop editor -> refused: three modules provide a shell -- bash,
fish, zsh -- choose one
assign laptop zsh -> accepted, and the editor's requirement is answered
bash, fish beside it -> fine, nothing is claimed
Claims rather than pairwise exclusion, so a third display server would say what
it claims and need no edit to xorg or wayland. Scoped to node, site or mesh:
two DHCP servers at one site collide and at two sites do not, and the mesh-wide
one is the hub said as a claim instead of hard-coded.
Some conflicts cost no manifest field at all. The refusal above names the seat
AND the two files, because the mesh already holds every resource of every
module -- neither i3 nor sway knows the other exists.
Resource identities carry their module, so two modules may both call something
"config" without the second silently replacing the first. What a service
reflects is qualified the same way, or it would name a resource that no longer
exists and stop being restarted when its own configuration changes.
Nothing is sent until every node resolves. A push that configured three and
refused on the fourth would leave the mesh in a state nobody asked for, and the
fourth is exactly where a claim collision appears.
One real flaw found by using it rather than by testing it: assigning zsh did
not satisfy a requirement for a shell. Requirements were counted against the
catalogue without first asking what the set already offers, so "choose one and
assign it" named three modules and then ignored the one you chose. The remedy
was useless and every test passed.
|
||
|
|
f44e73d286 |
The mesh computes a private network it cannot impersonate
The first thing the control plane decides rather than relays. Every node's peer list is derived from every node at once, which is what makes this control-plane work by definition: no node has that view. A hub, with direct peering between nodes at the same site. Not a full mesh, and the reason is a property of WireGuard rather than a preference -- there is no failover, so a more specific route to a dead endpoint blackholes instead of falling back. A node gets exactly one path to any peer, because two would mean one of them silently swallowing traffic. A roaming node is hub-only for the same reason. Reachability and the hub are declared, never inferred from an address. The address is evidence and is not the fact: carrier-grade NAT looks public and is not, a routable address behind a closed firewall looks public and is not, and the regular expression that used to decide it got the lab wrong too. Hub election by address prefix failed silently when nobody knew the convention. No private key travels, and that is the whole design. The node generated its own keypair and kept the private half; the configuration points at a file the node wrote, using WireGuard's own PostUp. So the control plane composes a complete configuration for a node it cannot pretend to be -- it knows every public key and holds none of the private ones. Delivered as an ordinary declaration: a package, a file and a service. The host does not know what a private network is and does not learn one. There is a test holding that line, because the moment connectivity needs a new shape in tier 0 is the moment the host stops being small enough to trust. The generated file is written to be read: each peer says why it is there, a peer with no endpoint says why it has none, and the header says not to edit it -- an edit survives until the graph next changes and then vanishes, which is worse than never being applied, because the machine works and then stops and nothing changed that anybody remembers. Fault injection found one weak test. The keepalive rule was asserted only against the hub, whose peer entries happen not to set the field at all, so it was testing an absence rather than the rule. It now checks two direct peers where one is reachable and one is not. |
||
|
|
f563ababa1 |
The mesh keeps a copy of what each node owns
novox/hq 09-the-node-lifecycle asks for this and it was missing: the host reports what it owns and the mesh keeps the last report. A backup, never a source -- nothing decides anything from it, and a node that disagrees with it wins, because the node is the one that can see the machine. Its point is the orphans. A node that loses its state file currently strands whatever it applied: nothing on the machine knows those resources were the mesh's doing, so nothing removes them. With this, a rebuilt node receives both the declaration and the record of what it previously owned. Never reported and reported nothing are kept apart, and that is the whole care in it. A node that applied nothing holds nothing; a node that has never spoken is unknown -- and handing back an empty list for the second would tell a rebuilding node it owns nothing and have it remove whatever it found. The age comes back with the answer rather than being left for the caller to go and find. An answer about a machine is worth much less without one, and this repository has already been bitten by a cache with no age on it. A refusal or a partial failure moves last_seen and nothing else: neither is an account of what the machine holds, and recording one as though it were would tell a rebuilding node to remove what it still has. |
||
|
|
66768208d2 |
Node records, and the right to join once
The next step after the schema: inventory now holds node records and enrolment tokens, and mesh-control has the commands to work with them. A token is issued for a node record, which is where re-enrolment gets decided -- what an identity binds to is settled when the token is made, not when it is presented, so the machine presenting one does not need to know whether it is joining or returning. What the token guarantees, each with a test confirmed to fail when the behaviour is removed: the secret is 256 random bits, shown once and stored only as a hash; it works exactly once; it stops working when it expires; issuing again for a node invalidates the outstanding one, because two live tokens are two machines able to join as the same node. Redemption is a single statement that finds and spends together, so eight concurrent attempts on one secret produce exactly one winner rather than a race between a check and a write. Refusals are deliberately identical for unknown, spent and expired. Somebody guessing must not learn which guess was a real token that had merely aged out. SHA-256 rather than a password hash, and that is a choice not a shortcut: the secret is high-entropy random, so there is nothing to guess and a slow hash would buy nothing while making every redemption expensive. It stops before what a node receives in exchange. What a machine presents afterwards to prove it is that node is not decided anywhere, and a migration is the most expensive place here to guess. So a token carries one of the four things ADR 0004 requires. The command prints the secret and then says exactly that -- the broker's address, its certificate fingerprint and the control plane's signing identity do not exist yet. Better than emitting something that looks complete and silently cannot be used. |
||
|
|
306c4ca13b |
The control plane, as far as identity
Tier 2 exists now. It holds one context of seven, inventory, and does one thing with it: brings its schema up to date. That is step 3 of the substrate bootstrap -- the step the first node cannot get past. Verified against a real PostgreSQL, with the built binary: applied 0001-nodes, reported 'already up to date' on the second run, and the node table is there with the index and the unique constraint the migration asks for. Written in Go, and the image is FROM scratch holding one file. Confirmed by unpacking it. That is the whole argument of ADR 0024: the bundle pins this image by digest and runs it where nothing can check it, so everything in it is something a person has to audit before trusting a first node. Exclusive store ownership is built as a rule about credentials rather than about intentions. There is no mesh-wide connection setting and no way to ask for one -- a context reads MESH_STORE_<ITS OWN NAME> and holds nothing else, so reaching another context's store needs a new variable, which is visible in the declaration that runs it. The migration runner is mostly refusals: an edited migration that already ran, a migration numbered below one that has run, duplicate numbers, misnamed files, empty files. All stop rather than warn, because at the moment any of them is true nobody knows what the database holds. It stops before identity, deliberately. What a node presents to prove who it is has not been decided anywhere, and a migration is the most expensive place in this system to guess. Two tests did not defend what they claimed, and both are fixed rather than removed. One asked only whether Open returned an error, which it did either way -- a bad context name and a missing credential both fail, so deleting the name check changed nothing. The other claimed to prove the migration runs in a transaction, but PostgreSQL already wraps a multi-statement query in one of its own, so it passed with the transaction taken out. What the transaction actually buys is that the schema change and the row recording it commit together, and there is now a test for that which fails when they are split. |