53eb000a84fb8886402f5fc51cdcc85e69c8fed5
6
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0af3ea1acf |
A consumer is a module on a machine, not a machine
novox/hq 04-ISSUES/022. A credential was keyed by provision, consumer node and provider node, so "who is asking" was answered by naming a host. The node this mesh exists to take over runs eight modules against one database server. The symptom had two halves and only one was loud. The provider refused, naming the modules and explaining they would share one credential, which reads as a decision rather than a limit. The consumer did not refuse: it resolved cleanly, wrote one module's credential file and left the others absent — a service that starts and cannot authenticate, with nothing saying why. That is 021 again on a different axis. Three modules wanting one database produced one need, carrying whichever module mentioned it first, because the resolution walk is a work-list over names. The fan-out now happens in one place, after the walk. The record path already did this correctly and said why: a consumer here is a module on a machine. It is the same rule. Downstream: the secret's key gains the consuming module, the grant file is named after both halves, needs are matched by provision and module rather than provision alone, and the provisioners name the role and the access key after the module. The refusal in ContributionsTo is gone because there is nothing left to refuse. Worth stating plainly: without that refusal, gitea's login would have opened keycloak's database. From the provisioner's side it created exactly what it was asked to create. Existing secrets are discarded rather than backfilled. They cannot say which module they were for, and a secret is remade and delivered to both ends on the next push — so this costs one rotation and invents nothing. Also guards the role name against PostgreSQL's 63-byte truncation, which is a notice rather than an error and would reintroduce exactly this collision at a length nobody tests. Three faults injected — the fan-out removed, needs matched by name alone, the grant file named after the machine — each caught. |
||
|
|
ebcfd37b92 |
Rotate a credential and move both ends together
The invariant novox/hq ADR 0001 records as unowned, and it was measurably false in HAL: a provision documented as never rotating minted a new password on every adoption and updated only the provider's row. Consumers on three nodes held dead credentials for two days while the mesh reported success. Nothing enumerated who held the old one. Three things make that impossible here. The holders are a set the mesh can name — each pair has its own credential, so rotating one consumer touches one role and the affected list is a query rather than an assumption. Both ends are pushed by this command rather than a later one, because leaving the sending to whoever remembered is the fault exactly. And it is all-or-nothing: if any affected machine cannot be resolved, nothing is sent and the old credential keeps working, which is a mesh that has not rotated rather than one that has half-rotated. The window is stated rather than hidden: a role's password changes on the provider and the file changes on the consumer, and they cannot be simultaneous. The provisioner now takes its superuser password from the file the mesh wrote, which is how the mesh delivers one. Passing it through the environment needed a person in the middle of the one path that exists so there is not one — and put a superuser password where `docker inspect` prints it. |
||
|
|
58c8ab7747 |
A secret the mesh was given is not one the mesh can reinvent
Two kinds live in module_secret and they behaved identically, which is right for one of them. A made secret is the mesh's: when a node regenerates its sealing key the mesh makes another and nothing is lost, because nothing else ever knew the old one. An accepted secret is not. A broker account's password exists because the broker was told about it. Regenerating one puts 32 random bytes where a working credential was — and the machine applies it, reports success, and the program reading it fails to authenticate somewhere else entirely, with the mesh insisting the secret was delivered, which it was. The row now records where the value came from, and a rejoined machine asking for an accepted one is refused with the remedy named: issue it again. No amount of pushing produces a password the broker has never heard of. Found while making the builder a module, which is the first thing to hold one. |
||
|
|
0262873254 |
status --json, so a board has something to read
A board reads through interfaces and holds nothing. Everything it needs is already answered — as text, for people, which is not something a page can read. `--json` rather than a serving API, because nothing needs one yet: whatever serves a board runs the command, and the constraint holds either way — the board never touches a context's store. An API is the larger thing and should wait until something asks for it. Both forms are gathered from the same reads before either says anything, so they answer the same questions rather than being two implementations that can drift. That was not true of the first version: the JSON printed after the text, because the branch was too late. Four properties, each asserted and each confirmed to fail when removed: - refused and failed stay distinct all the way out. They are fixed in different places, so one word for both sends half a page's readers to the wrong one — and how much DID apply is carried, since "three of eight" and "none of eight" are different machines - a machine that never spoke carries no time at all, rather than a zero one that any page would format as a date in 1970 - nothing is null. A page distinguishing "no machines are wrong" from "this field is missing" has to handle both, and null is the one that gets forgotten - no field is named like a secret. Everything here comes from records that hold no readable one, but a shape a page is built against is exactly where one would eventually be added for convenience |
||
|
|
c37d368f65 |
A module may need a secret of its own, and the provisioner watches
Two things, both found by trying to write a real postgres module and discovering it could not be said. A database has a superuser password, a broker an administrator, a registry an account. None of them is *for* anybody — they are not the credential a consumer is given, and the mechanism that hands those out has a consumer in the middle of it. So a module may declare what it needs and where to put it, and the mesh generates one per node, seals it, and reads it no more than it reads any other. Per node, deliberately: a module running on three machines has three passwords. One in the manifest instead would put the same secret on every machine that ever runs it, in a file anybody can read, for ever. Made once and kept, or a running database would be handed a password it was not started with; remade when the machine's sealing key changes, like everything else sealed here. A need declared and not made is refused rather than skipped, because a module whose own credential is silently absent starts, fails to authenticate, and the reason is three layers from the machine reporting it. And the provisioner can watch. That is what lets it be a module rather than a binary somebody places: run once, it needs invoking after every declaration by a timer or a unit wired to a file; watching, it is an ordinary long-running service the host already supervises. It polls rather than watching the filesystem, because the host writes atomically — the file is replaced, so a watch on the path stops seeing anything after the first replacement, and a watcher that silently stops working is worse than a poll. Credentials are compared by digest and never held: this runs for as long as the machine is up. |
||
|
|
20f78cd5f1 |
Credentials the mesh delivers and cannot read
HAL keeps env vars in the registry, encrypted at rest. Its own tooling records what that bought and what it did not. `secret_locate` matches by value rather than by name — because the same password sits in mesh_provisions, in module_env, in each node's .env in plain text, and inside every connection string composed from it, and its documentation says those URL copies "are often the only copies actually in use". And a query against the encrypted column returns zero rows and proves nothing, so auditing moved to the decrypted copies on the nodes. Two faults there, and encryption at rest addresses neither: the control plane can read what it stores, so a copy of the database is a copy of every credential; and one secret has many homes with nothing tracking them. So here the mesh generates a password, seals it to each end with keys those nodes generated, stores both blobs, and discards the plaintext. It cannot read what it holds. Neither can the broker relaying it. And nothing is composed centrally — a connection string is assembled on the machine that needs one — so no copy is ever minted in a shape nothing tracks. `Compromise of a node is compromise of that node` (ADR 0004) is now true of secrets, not only of identity. Two files rather than one, because the mesh cannot compose a document containing a value it discarded: `binds` carries the readable facts, `secrets` carries the credential alone. The readable half stays readable in the declaration; the secret half changes only when the secret does, which makes restart-on precise. The provider gets a directory, one file per consumer, for the same reason. It is made once and kept — regenerating per declaration would restart both ends on every push, and the password a provider was told to create would never be the one its consumer was given. It is remade when either end's sealing key changes, and both ends learn the new one in the same push, so there is no window where half the mesh holds a dead credential. Two tests found passing for the wrong reason, both caught because their injection came back clean: - the provider's copy was asserted non-empty, which reads the same whichever column is selected. It now opens the blob with the provider's own key. - RotateSecret deleted and re-created; the re-create was dead, because the next read makes one anyway. Removed, and a second path to the same act is how two ends come to disagree. And one real fault: three places built a declaration, and the one behind `--json` predated credentials, so it silently produced a declaration missing them — a difference between what `plan` showed and what anything reading `--json` got. There is one path now. |