da31bcb11eeba6a84f154ae4dc7e86c8bb969ffb
506
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6d3bb18546 |
A node is known by a key it generated
The thing I had been calling blocked for weeks, built in an afternoon once it was pointed out that it was already decided. 08-connectivity says of the overlay keys: each node generates its own keypair, the private half never leaves the machine, the public half is published -- and says outright this IS ADR 0004's "a node holds its own identity". Nobody had applied it to node identity itself. identity now holds the public half of each node's key. Only the public half, which is the property worth having: a copy of this database is a list of who to believe, not a set of credentials, so compromise of a node really is compromise of only that node. Exactly one key is live per node, and re-enrolment revokes the one it replaced in the same transaction -- two live identities is the stolen-laptop case with the replaced machine still believed. Fault injection was worth the time here. Three findings. The unique index was defended by no test at all: sequential enrolment is already safe because the code revokes before inserting, so removing the constraint changed nothing. The constraint only matters when two enrolments race, and there is now a test that runs six at once and fails without it. My injection harness also lied to me. One injection matched nothing, changed no file, and reported NO BITE identically to a real one -- so a test that defends nothing and an injection that does nothing look the same. The harness now checksums the files and says NO-OP when they did not change. And one honest NO BITE left standing: making the key lookup return a zero key for an unknown node does not fail the test, because the signature check refuses it a line later. Two independent mechanisms, not a placebo. 61 tests, none skipped. |
||
|
|
ea6569277d |
A token with all four parts
Given the broker's address and its certificate, mesh-control now issues a token carrying everything ADR 0004 asks for: where to connect, what to expect there, whose signature to believe afterwards, and a one-time right to join. Verified by decoding one and checking the fingerprint against `openssl x509 | sha256sum` -- they match. The fingerprint is derived from the certificate on disk and never configured. A configured pin can drift from the certificate it describes, and a drifted pin is worse than none: every node issued a token during the drift refuses to connect, and the failure looks like an attack rather than a mistake. Computed over DER, which is what a client sees on the wire. Hashing the PEM text instead would mean the same certificate, re-wrapped with different line endings, produced a different pin -- there is a test for exactly that, and one for pointing this at tls.key by mistake, which would otherwise produce a confident pin over the wrong file. Having no broker stays a state rather than a failure: a control plane holds records and a signing key without one. Having half a broker is refused, because a token with an address and nothing to check it against invites a node to trust whatever answers. Fault injection caught the same weak test I wrote earlier in the day -- asking whether something failed rather than why, so deleting the guard changed nothing because it failed one line later anyway. Both are now asserted on the reason. |
||
|
|
7553af6c5a |
The control plane's signing key, and a second context to hold it
Everything is blocked on what a node presents to prove which node it is. This builds the other direction, which is not blocked: what a node believes. identity is the second of the seven contexts. It holds an Ed25519 signing key the control plane generates once, whose public half now travels in every enrolment token. A node believes a declaration because it carries a signature that key made -- pinning only the broker would make the control plane's authority transitive, and since the host applies whatever the link delivers, a compromised broker forging declarations is the whole machine. Establishing the key is idempotent, and it has to be: a second key generated by a restart is a mesh where every node holds the wrong public half, so every declaration is refused by every node with nothing visibly wrong. The guarantee is a partial unique index plus a read-back, not the check before the insert -- six processes racing to establish all agree on one key, and there is a test that runs them. Tokens are now one line of base64 carrying three of their four parts. The missing two are the broker's address and its certificate fingerprint, both step 5 of the bootstrap. The command prints the token and names what is missing rather than emitting something that looks usable. The second context also tests a claim this repository had made and never checked: that a context reaches only its own store. Two databases, two credentials, no setting that reaches both. Running migrate with one stops and names the grant it lacks -- verified, not asserted. Assembling a token needs a node record from one and a key from the other, and neither reads the other's store; the process holding both grants asks each for its part. 45 tests, none skipped. Fault injection found one test whose property is enforced somewhere other than where I injected -- idempotency comes from the database constraint, not from the early return, which is what the code comment already said. |
||
|
|
66768208d2 |
Node records, and the right to join once
The next step after the schema: inventory now holds node records and enrolment tokens, and mesh-control has the commands to work with them. A token is issued for a node record, which is where re-enrolment gets decided -- what an identity binds to is settled when the token is made, not when it is presented, so the machine presenting one does not need to know whether it is joining or returning. What the token guarantees, each with a test confirmed to fail when the behaviour is removed: the secret is 256 random bits, shown once and stored only as a hash; it works exactly once; it stops working when it expires; issuing again for a node invalidates the outstanding one, because two live tokens are two machines able to join as the same node. Redemption is a single statement that finds and spends together, so eight concurrent attempts on one secret produce exactly one winner rather than a race between a check and a write. Refusals are deliberately identical for unknown, spent and expired. Somebody guessing must not learn which guess was a real token that had merely aged out. SHA-256 rather than a password hash, and that is a choice not a shortcut: the secret is high-entropy random, so there is nothing to guess and a slow hash would buy nothing while making every redemption expensive. It stops before what a node receives in exchange. What a machine presents afterwards to prove it is that node is not decided anywhere, and a migration is the most expensive place here to guess. So a token carries one of the four things ADR 0004 requires. The command prints the secret and then says exactly that -- the broker's address, its certificate fingerprint and the control plane's signing identity do not exist yet. Better than emitting something that looks complete and silently cannot be used. |
||
|
|
3b5861282c |
Point at 0006 and 0008 rather than a record of their own
The language and the bundle ordering went into ADR 0006, where the substrate and the control plane already live, and the store-per-context mechanics into 0008, which already decided the rule. Nothing changed but where the reasoning is kept. |
||
|
|
306c4ca13b |
The control plane, as far as identity
Tier 2 exists now. It holds one context of seven, inventory, and does one thing with it: brings its schema up to date. That is step 3 of the substrate bootstrap -- the step the first node cannot get past. Verified against a real PostgreSQL, with the built binary: applied 0001-nodes, reported 'already up to date' on the second run, and the node table is there with the index and the unique constraint the migration asks for. Written in Go, and the image is FROM scratch holding one file. Confirmed by unpacking it. That is the whole argument of ADR 0024: the bundle pins this image by digest and runs it where nothing can check it, so everything in it is something a person has to audit before trusting a first node. Exclusive store ownership is built as a rule about credentials rather than about intentions. There is no mesh-wide connection setting and no way to ask for one -- a context reads MESH_STORE_<ITS OWN NAME> and holds nothing else, so reaching another context's store needs a new variable, which is visible in the declaration that runs it. The migration runner is mostly refusals: an edited migration that already ran, a migration numbered below one that has run, duplicate numbers, misnamed files, empty files. All stop rather than warn, because at the moment any of them is true nobody knows what the database holds. It stops before identity, deliberately. What a node presents to prove who it is has not been decided anywhere, and a migration is the most expensive place in this system to guess. Two tests did not defend what they claimed, and both are fixed rather than removed. One asked only whether Open returned an error, which it did either way -- a bad context name and a missing credential both fail, so deleting the name check changed nothing. The other claimed to prove the migration runs in a transaction, but PostgreSQL already wraps a multi-statement query in one of its own, so it passed with the transaction taken out. What the transaction actually buys is that the schema change and the row recording it commit together, and there is now a test for that which fails when they are split. |