Files
mesh-controller/internal/inventory/migrations/0015-a-consumer-is-a-module-on-a-machine.sql
jschoubben 0af3ea1acf A consumer is a module on a machine, not a machine
novox/hq 04-ISSUES/022. A credential was keyed by provision, consumer
node and provider node, so "who is asking" was answered by naming a
host. The node this mesh exists to take over runs eight modules against
one database server.

The symptom had two halves and only one was loud. The provider refused,
naming the modules and explaining they would share one credential, which
reads as a decision rather than a limit. The consumer did not refuse: it
resolved cleanly, wrote one module's credential file and left the others
absent — a service that starts and cannot authenticate, with nothing
saying why. That is 021 again on a different axis.

Three modules wanting one database produced one need, carrying whichever
module mentioned it first, because the resolution walk is a work-list
over names. The fan-out now happens in one place, after the walk. The
record path already did this correctly and said why: a consumer here is
a module on a machine. It is the same rule.

Downstream: the secret's key gains the consuming module, the grant file
is named after both halves, needs are matched by provision and module
rather than provision alone, and the provisioners name the role and the
access key after the module. The refusal in ContributionsTo is gone
because there is nothing left to refuse.

Worth stating plainly: without that refusal, gitea's login would have
opened keycloak's database. From the provisioner's side it created
exactly what it was asked to create.

Existing secrets are discarded rather than backfilled. They cannot say
which module they were for, and a secret is remade and delivered to both
ends on the next push — so this costs one rotation and invents nothing.

Also guards the role name against PostgreSQL's 63-byte truncation, which
is a notice rather than an error and would reintroduce exactly this
collision at a length nobody tests.

Three faults injected — the fan-out removed, needs matched by name
alone, the grant file named after the machine — each caught.
2026-09-01 02:40:09 +02:00

32 lines
1.8 KiB
SQL

-- A credential belongs to a consumer, and a consumer is a module on a machine.
--
-- novox/hq 04-ISSUES/022. The key was (provision, consumer node, provider node), so "who is
-- asking" was answered by naming a host. A node running three modules against one database server
-- had one credential between them: the provisioner created one role, `mesh_<node>`, owning every
-- database it was asked for, and gitea's login opened keycloak's data. Nothing anywhere would
-- have said so -- from the provisioner's side it created exactly what it was asked to create.
--
-- **Two modules on one node are as separate as two on different nodes.** They are different
-- containers, on different networks, with different data. This is the same correction as 021,
-- which found the machine wrongly treated as a trust boundary; here it was wrongly treated as an
-- identity.
--
-- It also restores withdrawal. One role per node cannot express "this module no longer has a
-- login and the others still do", so a consumer that went away kept a working credential for as
-- long as any other consumer on that machine remained.
alter table secret add column consumer_module text references module(name) on delete cascade;
-- Existing rows cannot say which module they were for, because at the time nothing recorded it.
--
-- **Discarded rather than guessed.** A secret is remade on the next declaration and reaches both
-- ends in the same push, which is exactly what rotation does -- so this costs one rotation and
-- nothing else. Backfilling with "whichever module resolves first" would be inventing an answer
-- to the question this migration exists because nobody could answer.
delete from secret;
alter table secret alter column consumer_module set not null;
alter table secret drop constraint secret_pkey;
alter table secret add primary key (name, consumer, consumer_module, provider);