novox/hq 04-ISSUES/022. A credential was keyed by provision, consumer
node and provider node, so "who is asking" was answered by naming a
host. The node this mesh exists to take over runs eight modules against
one database server.
The symptom had two halves and only one was loud. The provider refused,
naming the modules and explaining they would share one credential, which
reads as a decision rather than a limit. The consumer did not refuse: it
resolved cleanly, wrote one module's credential file and left the others
absent — a service that starts and cannot authenticate, with nothing
saying why. That is 021 again on a different axis.
Three modules wanting one database produced one need, carrying whichever
module mentioned it first, because the resolution walk is a work-list
over names. The fan-out now happens in one place, after the walk. The
record path already did this correctly and said why: a consumer here is
a module on a machine. It is the same rule.
Downstream: the secret's key gains the consuming module, the grant file
is named after both halves, needs are matched by provision and module
rather than provision alone, and the provisioners name the role and the
access key after the module. The refusal in ContributionsTo is gone
because there is nothing left to refuse.
Worth stating plainly: without that refusal, gitea's login would have
opened keycloak's database. From the provisioner's side it created
exactly what it was asked to create.
Existing secrets are discarded rather than backfilled. They cannot say
which module they were for, and a secret is remade and delivered to both
ends on the next push — so this costs one rotation and invents nothing.
Also guards the role name against PostgreSQL's 63-byte truncation, which
is a notice rather than an error and would reintroduce exactly this
collision at a length nobody tests.
Three faults injected — the fan-out removed, needs matched by name
alone, the grant file named after the machine — each caught.
Phase 1.1 of the work breakdown. The finding that shaped it came before
any code: **the control plane special-cases nothing.** provides,
requires, contributes and grants are entirely name-agnostic, so asking
for a bucket needed no change to the mesh at all — only a provider that
answers. What was missing was the last step, where something on the
machine turns a delivered secret into a key that works.
Named `s3-bucket` by ADR 0027's test: a consumer's code is written
against the S3 API, and swapping one store for another does not break
it, so the coupling is to the protocol rather than the product — which
is what the substrate design already said about AMQP, S3 and OCI.
Proven on a real store, 7 assertions: a generated secret becomes a
working key; rotation makes the new one work and the old one stop; a
consumer that goes away loses its key; a key nobody here made is left
alone; a manifest naming a credential that was never written is refused;
an unusable bucket name is refused naming the consumer that asked.
**And the one a database does not need.** One PostgreSQL server holds
separate databases and the product enforces the boundary; one object
store holds every bucket behind one endpoint, so a consumer being unable
to reach another's is a policy somebody wrote. A policy granting
arn:aws:s3:::* would pass every other test in the file, so the unit
tests assert what the policy does NOT say.
It drives the vendor's command line rather than an SDK: the admin API
encrypts its request bodies, which is why a separate admin library
exists, and pulling that in would add a system-metrics dependency tree
to a repository with none in order to create a user.