e81f35297938e9a886a204dc9c2be29d1d9145fc
5
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f04d00b411 |
The proxy can obtain a public certificate, and asks staging by default
Work breakdown 1.4. The mesh's own authority certifies internal names and always did; a name reachable from outside needs one the world already trusts, and there was no ACME anywhere in this repository. Uses acme/autocert from x/crypto, which was already a dependency — one indirect addition (x/net, for idna) and no new direct one. Three things worth more than the feature: **Staging is the default** (novox/hq 04-ISSUES/004). Production issuance is rate-limited per domain and per account and does not replenish quickly. Defaulting to production would leave the safe path depending on remembering to opt out, on exactly the work most likely to iterate. A staging certificate is trusted by no browser, so the mistake announces itself on the first request rather than a fortnight later. **A certificate is only asked for on a name the mesh routes here.** Without that policy, anything that can reach the port and send a name triggers an order for it — a scan becomes a stream of failed orders against the account's rate limit, and the proxy looks healthy throughout. What it may certify is what it was told to serve. **A private issuer is trusted by naming a file, never by skipping verification.** Skip would still apply on the day this points at a public issuer, and nothing would say so. TLS is opt-in: without TLS_LISTEN the proxy serves plain HTTP exactly as before, which is what an internal-only mesh wants. With it and no cache, it refuses rather than defaulting — every restart would otherwise order new certificates, silently, until the rate limit says it does not. |
||
|
|
c3046dcf56 |
A provider is told who its consumers are, and a reference provisioner
Contributions were node-local, so a mesh-scoped provider — the one case that most needs them — never heard from its consumers. A database was given a password and no idea what to create it for. Cross-node consumers now reach the provider's `receives` file, merged in with the ones on its own machine: from the provider's side they are the same thing, and a provider that had to read two lists would read one of them. Each names the file its credential is in rather than carrying it, because the mesh discarded the value and could not put it there. The readable half therefore stays readable. And examples/postgres-provisioner, which is the last step: it reads what the host wrote and makes PostgreSQL accept it. Explicitly not part of the control plane — the control plane decides and never touches a machine. This runs on the machine and touches it, and a real one ships with the module that ships PostgreSQL. It lives here because this is where the contract is defined, written as something that runs so it can be read. It reconciles rather than applying a change, because it is never told what changed. Three things that follow, and each is a fault somebody has shipped: - the password is set every time, not only on creation, or a rotation reports success and changes nothing - what it made and nobody asks for any more is revoked, or a departed consumer keeps a working login for ever - what it did not make is left alone, or it cannot be run on a database that predates it Proven in the lab against a real PostgreSQL, each assertion confirmed to fail with the behaviour removed. The suite is in mesh-lab, which also records the two ways the test itself was wrong first. |
||
|
|
20f78cd5f1 |
Credentials the mesh delivers and cannot read
HAL keeps env vars in the registry, encrypted at rest. Its own tooling records what that bought and what it did not. `secret_locate` matches by value rather than by name — because the same password sits in mesh_provisions, in module_env, in each node's .env in plain text, and inside every connection string composed from it, and its documentation says those URL copies "are often the only copies actually in use". And a query against the encrypted column returns zero rows and proves nothing, so auditing moved to the decrypted copies on the nodes. Two faults there, and encryption at rest addresses neither: the control plane can read what it stores, so a copy of the database is a copy of every credential; and one secret has many homes with nothing tracking them. So here the mesh generates a password, seals it to each end with keys those nodes generated, stores both blobs, and discards the plaintext. It cannot read what it holds. Neither can the broker relaying it. And nothing is composed centrally — a connection string is assembled on the machine that needs one — so no copy is ever minted in a shape nothing tracks. `Compromise of a node is compromise of that node` (ADR 0004) is now true of secrets, not only of identity. Two files rather than one, because the mesh cannot compose a document containing a value it discarded: `binds` carries the readable facts, `secrets` carries the credential alone. The readable half stays readable in the declaration; the secret half changes only when the secret does, which makes restart-on precise. The provider gets a directory, one file per consumer, for the same reason. It is made once and kept — regenerating per declaration would restart both ends on every push, and the password a provider was told to create would never be the one its consumer was given. It is remade when either end's sealing key changes, and both ends learn the new one in the same push, so there is no window where half the mesh holds a dead credential. Two tests found passing for the wrong reason, both caught because their injection came back clean: - the provider's copy was asserted non-empty, which reads the same whichever column is selected. It now opens the blob with the provider's own key. - RotateSecret deleted and re-created; the re-create was dead, because the next read makes one anyway. Removed, and a second path to the same act is how two ends come to disagree. And one real fault: three places built a declaration, and the one behind `--json` predated credentials, so it silently produced a declaration missing them — a difference between what `plan` showed and what anything reading `--json` got. There is one path now. |
||
|
|
46e760fc94 |
The control plane serves, and a node can join
There was no chicken-and-egg to solve. The mesh runs the broker, so it creates the node's account when it issues the token, and the one-time secret is that account's password. A joining node's first connection is already authenticated; enrolment is what it says once it is in. I had been treating this as a decision that needed taking, and it did not. The account is per node and scoped: it may read its own queue, write to the one exchange, and configure nothing else. The patterns are anchored and the node name is constrained to characters that cannot widen them, because a name carrying a dot or a star would silently let that node read everybody's queues. `serve` is the control plane running: one connection, one queue, one consumer. One deliberately -- two consumers on a queue get round-robined and each receives half of what it expects, which has happened on this project before, between a module's daemon and its capability server. Enrolment spends the token first, in the single statement that both finds and marks it, and only then records the key. That order is the order things become irreversible: recording a key for a node whose token turned out to be spent would leave the mesh believing a machine that never had the right to join. Refusals are one message for every reason. The log says which, where an operator can see it; the node is told only that the token cannot be used. Verified in the lab, on a sealed machine, through the whole first-node path. |
||
|
|
306c4ca13b |
The control plane, as far as identity
Tier 2 exists now. It holds one context of seven, inventory, and does one thing with it: brings its schema up to date. That is step 3 of the substrate bootstrap -- the step the first node cannot get past. Verified against a real PostgreSQL, with the built binary: applied 0001-nodes, reported 'already up to date' on the second run, and the node table is there with the index and the unique constraint the migration asks for. Written in Go, and the image is FROM scratch holding one file. Confirmed by unpacking it. That is the whole argument of ADR 0024: the bundle pins this image by digest and runs it where nothing can check it, so everything in it is something a person has to audit before trusting a first node. Exclusive store ownership is built as a rule about credentials rather than about intentions. There is no mesh-wide connection setting and no way to ask for one -- a context reads MESH_STORE_<ITS OWN NAME> and holds nothing else, so reaching another context's store needs a new variable, which is visible in the declaration that runs it. The migration runner is mostly refusals: an edited migration that already ran, a migration numbered below one that has run, duplicate numbers, misnamed files, empty files. All stop rather than warn, because at the moment any of them is true nobody knows what the database holds. It stops before identity, deliberately. What a node presents to prove who it is has not been decided anywhere, and a migration is the most expensive place in this system to guess. Two tests did not defend what they claimed, and both are fixed rather than removed. One asked only whether Open returned an error, which it did either way -- a bad context name and a missing credential both fail, so deleting the name check changed nothing. The other claimed to prove the migration runs in a transaction, but PostgreSQL already wraps a multi-statement query in one of its own, so it passed with the transaction taken out. What the transaction actually buys is that the schema change and the row recording it commit together, and there is now a test for that which fails when they are split. |