Commit Graph
3 Commits
Author SHA1 Message Date
jschoubben 0af3ea1acf A consumer is a module on a machine, not a machine
novox/hq 04-ISSUES/022. A credential was keyed by provision, consumer
node and provider node, so "who is asking" was answered by naming a
host. The node this mesh exists to take over runs eight modules against
one database server.

The symptom had two halves and only one was loud. The provider refused,
naming the modules and explaining they would share one credential, which
reads as a decision rather than a limit. The consumer did not refuse: it
resolved cleanly, wrote one module's credential file and left the others
absent — a service that starts and cannot authenticate, with nothing
saying why. That is 021 again on a different axis.

Three modules wanting one database produced one need, carrying whichever
module mentioned it first, because the resolution walk is a work-list
over names. The fan-out now happens in one place, after the walk. The
record path already did this correctly and said why: a consumer here is
a module on a machine. It is the same rule.

Downstream: the secret's key gains the consuming module, the grant file
is named after both halves, needs are matched by provision and module
rather than provision alone, and the provisioners name the role and the
access key after the module. The refusal in ContributionsTo is gone
because there is nothing left to refuse.

Worth stating plainly: without that refusal, gitea's login would have
opened keycloak's database. From the provisioner's side it created
exactly what it was asked to create.

Existing secrets are discarded rather than backfilled. They cannot say
which module they were for, and a secret is remade and delivered to both
ends on the next push — so this costs one rotation and invents nothing.

Also guards the role name against PostgreSQL's 63-byte truncation, which
is a notice rather than an error and would reintroduce exactly this
collision at a length nobody tests.

Three faults injected — the fan-out removed, needs matched by name
alone, the grant file named after the machine — each caught.
2026-09-01 02:40:09 +02:00
jschoubben 6eeb10f066 A command opens each store once, not once per machine
Working out what a machine should be reaches the identity context for its
certificate and the licence context for its model access. Both were opened —
and waited on — inside functions called for every node in a push. Two machines
hid it. Fifty would be fifty connect-and-wait cycles for data that does not
change while the push runs.

So a command holds what it has open, and passes it. Each context is opened on
first use rather than up front, because most commands need one and paying to
reach three would be the same waste from the other side.

The contexts stay separate, which is the point: this is one struct holding
three connections to three databases, not one connection to a shared one. No
context reaches another's store, and each still holds only its own credential
(novox/hq ADR 0008).

A pure move again — the gate is green before and after, and no test changed.
2026-08-31 13:35:39 +02:00
jschoubben ebcfd37b92 Rotate a credential and move both ends together
The invariant novox/hq ADR 0001 records as unowned, and it was measurably
false in HAL: a provision documented as never rotating minted a new password on
every adoption and updated only the provider's row. Consumers on three nodes
held dead credentials for two days while the mesh reported success. Nothing
enumerated who held the old one.

Three things make that impossible here. The holders are a set the mesh can name
— each pair has its own credential, so rotating one consumer touches one role
and the affected list is a query rather than an assumption. Both ends are
pushed by this command rather than a later one, because leaving the sending to
whoever remembered is the fault exactly. And it is all-or-nothing: if any
affected machine cannot be resolved, nothing is sent and the old credential
keeps working, which is a mesh that has not rotated rather than one that has
half-rotated.

The window is stated rather than hidden: a role's password changes on the
provider and the file changes on the consumer, and they cannot be simultaneous.

The provisioner now takes its superuser password from the file the mesh wrote,
which is how the mesh delivers one. Passing it through the environment needed a
person in the middle of the one path that exists so there is not one — and put
a superuser password where `docker inspect` prints it.
2026-08-31 02:38:52 +02:00