--- layer: to-be status: in-progress code: - mesh-controller internal/inventory/secrets.go - mesh-controller cmd/mesh-controller/rotate.go - mesh-controller examples/postgres-provisioner updated: 2026-09-21 decisions: - 02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md - 02-DECISIONS/0009-modules-and-the-graph.md - 02-DECISIONS/0085-a-secret-is-a-provision.md - 02-DECISIONS/0086-a-secret-reaches-a-process-as-a-file.md --- # 13 — Credentials, and moving them *Written 2026-08-31, when rotation was built. The delivery half was already proven; this is the half [ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) records as unowned, and it was **measurably false** in the system being replaced.* ## What went wrong before, precisely `provision_ensure`, documented as *NEVER rotates an existing secret*, minted a new password on every adoption and updated **only the provider's row**. Consumers on three nodes held dead credentials for two days. Two rows for one provision were written 216 ms apart, so at most one could have matched the live role. The mesh reported success throughout. Three separate faults, and it is worth naming them apart because they have different fixes: | | | |---|---| | **one credential, many holders** | rotating it was necessarily a fan-out, and nothing enumerated who held it | | **the record moved and the consumers did not** | the change and the delivery were different acts, and only the first happened | | **nothing said so** | the mesh could not tell a rotated credential from a working one, so nobody looked | ## What replaces it **Every pair has its own credential.** A provision between one consumer and one provider is one password, made once and kept. So rotating a credential touches one role and leaves every other consumer alone — and *who holds this* is a query rather than an assumption. That alone removes the first fault: there is no shared secret to fan out. **A consumer is a module on a machine, not a machine.** This was written as though a pair were two machines, and built that way, and it was wrong in a way that only shows on a real node ([`022`](../../04-ISSUES/022-one-credential-per-node-per-provision-not-per-module/00-report.md)): a machine running several services against one database server had one credential between them. The provider refused to plan at all, and the consuming node did not refuse — it gave the first module a credential and the rest nothing. Two modules on one node are as separate as two on different nodes. They are different containers, with different data, and one login opening both is the thing this page exists to prevent. It is also what makes withdrawal possible: one role per machine cannot express *this module no longer has a login and the others still do*. **The change and the delivery are one command.** `rotate` discards the credential and sends both ends, and it does the sending itself. Leaving that to whoever remembers is the second fault exactly, and the interval in which it goes wrong is unbounded — two days, in the recorded case. **It is all-or-nothing.** If any affected machine cannot be resolved, nothing is sent and the old credential keeps working. A mesh that has not rotated is far better than one that has half-rotated, and the difference is that the first is obvious. **The window is stated rather than hidden.** A role's password changes on the provider and the file changes on the consumer, and those cannot be simultaneous. So there is an interval in which a consumer cannot authenticate, and the honest thing is to make it as short as the broker allows and to say it exists. `status` names who is still behind. ## The provider makes it true, and the mesh cannot The mesh generated the password, sealed it to the machine that must accept it, and **discarded the plaintext** — so it cannot tell a database to start accepting it. Something on that machine reads what the host wrote and makes it true. That something is part of the module, not part of the controller. **The controller decides and never touches a machine; a provisioner runs on the machine and touches it.** Two files, because the mesh could not compose a document containing a value it does not have: | | | |---|---| | a manifest | every consumer, what it asked for, and where its credential is | | one file per consumer | that consumer's password, alone | **Both, or neither works.** A provisioner given the passwords and not the manifest finds a directory of unexplained secrets and reports that nothing has been granted — which is true, and reads exactly like a credential that was never delivered. That has now happened once, here. **The password a provisioner uses is itself a file the mesh wrote.** Passing it through the environment needs a person in the middle of the one path that exists so there is not one, and puts a superuser password where `docker inspect` prints it. ## A secret reaches a process as a file The care above ends at the container's door if the module then hands the value to the runtime as an environment variable: `docker inspect` prints it, and the process's `/proc` entry holds it for anything on the machine that can talk to the runtime. So a secret reaches a process **as a file** ([ADR 0086](../../02-DECISIONS/0086-a-secret-reaches-a-process-as-a-file.md)): the module mounts the file the host wrote and points the program at it, and the mesh's own programs take a `_FILE` twin for every variable that carries a credential. Software that reads only its environment is not forbidden; it is **declared**, on the container, with a reason, so the manifest says which secrets are exposed that way. **Checked** at composition: the catalogue engine refuses a `${secret:…}` in a container's `env` outright, and a secret-carrying file named in `env-file` unless the container carries the declaration. ## How it is checked Not by comparing two files. **Two ends holding a matching string proves they agree, not that either is right** — the recorded fault produced two ends that agreed with each other and not with the database. So the check is three logins against a real PostgreSQL, from the consumer's own machine, over the private network: 1. the delivered credential authenticates 2. after rotation, the new one authenticates 3. **the one that was rotated away does not** The third is what makes it a rotation rather than an addition. Without it the check passes against a provider that added a password and removed nothing. **Not over loopback.** `pg_hba` trusts anything there, so every password looks correct — a deliberately wrong one returned a row for an afternoon before that was noticed. ## The secret that is not a pair Everything on this page is about the credential *between a consumer and a provider* — the login one module uses against another. A mesh also holds secrets that are not that: a value a single module needs for **its own** use, and a value only an operator can supply. Those had no owner, and so no rotation — the gap this page's machinery could not reach because there was no pair to rotate. [ADR 0085](../../02-DECISIONS/0085-a-secret-is-a-provision.md) gives them one. Such a secret is provided by a **vault** module ([24](24-the-secrets-vault.md)): a module requires a `secret` provision, and the credential of that consumer↔vault pair is an ordinary pair credential — so it rotates, and is queried for who holds it, through exactly the machinery described above, unchanged. The rotation a module's own secret lacks is not a second mechanism; it is this one, pointed at a secret the vault provides. What this page proves for a database password holds, by construction, for a secret from the vault.