Files
hq/03-DESIGN/01-to-be/13-credentials-and-their-rotation.md
T
jschoubben 2baf22ac43 A pair is a module and a provider, not two machines
Amends the credentials page, which said "every pair has its own
credential" and meant two machines. Built that way, it was wrong in a
way that only shows on a real node: a machine running several services
against one database server had one credential between them, so the
provider refused to plan at all and the consuming node quietly gave the
first module a credential and the rest nothing.

The page already argues the case against itself — one credential with
many holders is the first of the three faults it was written to remove.
It just drew the boundary at the machine.

Two modules on one node are as separate as two on different nodes, and
one login opening both is what this page exists to prevent. It is also
what makes withdrawal possible: one role per machine cannot say that
this module has lost its login and the others still have theirs.
2026-09-01 02:46:37 +02:00

108 lines
5.5 KiB
Markdown

---
layer: to-be
status: in-progress
code:
- mesh-control internal/inventory/secrets.go
- mesh-control cmd/mesh-control/rotate.go
- mesh-control examples/postgres-provisioner
updated: 2026-09-01
decisions:
- 02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md
- 02-DECISIONS/0009-modules-and-the-graph.md
---
# 13 — Credentials, and moving them
*Written 2026-08-31, when rotation was built. The delivery half was already proven; this is the
half [ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) records as
unowned, and it was **measurably false** in the system being replaced.*
## What went wrong before, precisely
`provision_ensure`, documented as *NEVER rotates an existing secret*, minted a new password on
every adoption and updated **only the provider's row**. Consumers on three nodes held dead
credentials for two days. Two rows for one provision were written 216 ms apart, so at most one
could have matched the live role. The mesh reported success throughout.
Three separate faults, and it is worth naming them apart because they have different fixes:
| | |
|---|---|
| **one credential, many holders** | rotating it was necessarily a fan-out, and nothing enumerated who held it |
| **the record moved and the consumers did not** | the change and the delivery were different acts, and only the first happened |
| **nothing said so** | the mesh could not tell a rotated credential from a working one, so nobody looked |
## What replaces it
**Every pair has its own credential.** A provision between one consumer and one provider is one
password, made once and kept. So rotating a credential touches one role and leaves every other
consumer alone — and *who holds this* is a query rather than an assumption. That alone removes the
first fault: there is no shared secret to fan out.
**A consumer is a module on a machine, not a machine.** This was written as though a pair were two
machines, and built that way, and it was wrong in a way that only shows on a real node
([`022`](../../04-ISSUES/022-one-credential-per-node-per-provision-not-per-module/00-report.md)):
a machine running several services against one database server had one credential between them.
The provider refused to plan at all, and the consuming node did not refuse — it gave the first
module a credential and the rest nothing.
Two modules on one node are as separate as two on different nodes. They are different containers,
with different data, and one login opening both is the thing this page exists to prevent. It is
also what makes withdrawal possible: one role per machine cannot express *this module no longer
has a login and the others still do*.
**The change and the delivery are one command.** `rotate` discards the credential and sends both
ends, and it does the sending itself. Leaving that to whoever remembers is the second fault
exactly, and the interval in which it goes wrong is unbounded — two days, in the recorded case.
**It is all-or-nothing.** If any affected machine cannot be resolved, nothing is sent and the old
credential keeps working. A mesh that has not rotated is far better than one that has
half-rotated, and the difference is that the first is obvious.
**The window is stated rather than hidden.** A role's password changes on the provider and the file
changes on the consumer, and those cannot be simultaneous. So there is an interval in which a
consumer cannot authenticate, and the honest thing is to make it as short as the broker allows and
to say it exists. `status` names who is still behind.
## The provider makes it true, and the mesh cannot
The mesh generated the password, sealed it to the machine that must accept it, and **discarded the
plaintext** — so it cannot tell a database to start accepting it. Something on that machine reads
what the host wrote and makes it true.
That something is part of the module, not part of the control plane. **The control plane decides
and never touches a machine; a provisioner runs on the machine and touches it.** Two files, because
the mesh could not compose a document containing a value it does not have:
| | |
|---|---|
| a manifest | every consumer, what it asked for, and where its credential is |
| one file per consumer | that consumer's password, alone |
**Both, or neither works.** A provisioner given the passwords and not the manifest finds a
directory of unexplained secrets and reports that nothing has been granted — which is true, and
reads exactly like a credential that was never delivered. That has now happened once, here.
**The password a provisioner uses is itself a file the mesh wrote.** Passing it through the
environment needs a person in the middle of the one path that exists so there is not one, and puts
a superuser password where `docker inspect` prints it.
## How it is checked
Not by comparing two files. **Two ends holding a matching string proves they agree, not that either
is right** — the recorded fault produced two ends that agreed with each other and not with the
database.
So the check is three logins against a real PostgreSQL, from the consumer's own machine, over the
private network:
1. the delivered credential authenticates
2. after rotation, the new one authenticates
3. **the one that was rotated away does not**
The third is what makes it a rotation rather than an addition. Without it the check passes against
a provider that added a password and removed nothing.
**Not over loopback.** `pg_hba` trusts anything there, so every password looks correct — a
deliberately wrong one returned a row for an afternoon before that was noticed.