Commit Graph
4 Commits
Author SHA1 Message Date
jschoubben c3046dcf56 A provider is told who its consumers are, and a reference provisioner
Contributions were node-local, so a mesh-scoped provider — the one case
that most needs them — never heard from its consumers. A database was
given a password and no idea what to create it for.

Cross-node consumers now reach the provider's `receives` file, merged in
with the ones on its own machine: from the provider's side they are the
same thing, and a provider that had to read two lists would read one of
them. Each names the file its credential is in rather than carrying it,
because the mesh discarded the value and could not put it there. The
readable half therefore stays readable.

And examples/postgres-provisioner, which is the last step: it reads what
the host wrote and makes PostgreSQL accept it. Explicitly not part of the
control plane — the control plane decides and never touches a machine.
This runs on the machine and touches it, and a real one ships with the
module that ships PostgreSQL. It lives here because this is where the
contract is defined, written as something that runs so it can be read.

It reconciles rather than applying a change, because it is never told
what changed. Three things that follow, and each is a fault somebody has
shipped:

- the password is set every time, not only on creation, or a rotation
  reports success and changes nothing
- what it made and nobody asks for any more is revoked, or a departed
  consumer keeps a working login for ever
- what it did not make is left alone, or it cannot be run on a database
  that predates it

Proven in the lab against a real PostgreSQL, each assertion confirmed to
fail with the behaviour removed. The suite is in mesh-lab, which also
records the two ways the test itself was wrong first.
2026-08-30 01:31:25 +02:00
jschoubben 20f78cd5f1 Credentials the mesh delivers and cannot read
HAL keeps env vars in the registry, encrypted at rest. Its own tooling
records what that bought and what it did not. `secret_locate` matches by
value rather than by name — because the same password sits in
mesh_provisions, in module_env, in each node's .env in plain text, and
inside every connection string composed from it, and its documentation
says those URL copies "are often the only copies actually in use". And a
query against the encrypted column returns zero rows and proves nothing,
so auditing moved to the decrypted copies on the nodes.

Two faults there, and encryption at rest addresses neither: the control
plane can read what it stores, so a copy of the database is a copy of
every credential; and one secret has many homes with nothing tracking
them.

So here the mesh generates a password, seals it to each end with keys
those nodes generated, stores both blobs, and discards the plaintext. It
cannot read what it holds. Neither can the broker relaying it. And
nothing is composed centrally — a connection string is assembled on the
machine that needs one — so no copy is ever minted in a shape nothing
tracks. `Compromise of a node is compromise of that node` (ADR 0004) is
now true of secrets, not only of identity.

Two files rather than one, because the mesh cannot compose a document
containing a value it discarded: `binds` carries the readable facts,
`secrets` carries the credential alone. The readable half stays readable
in the declaration; the secret half changes only when the secret does,
which makes restart-on precise. The provider gets a directory, one file
per consumer, for the same reason.

It is made once and kept — regenerating per declaration would restart
both ends on every push, and the password a provider was told to create
would never be the one its consumer was given. It is remade when either
end's sealing key changes, and both ends learn the new one in the same
push, so there is no window where half the mesh holds a dead credential.

Two tests found passing for the wrong reason, both caught because their
injection came back clean:

- the provider's copy was asserted non-empty, which reads the same
  whichever column is selected. It now opens the blob with the
  provider's own key.
- RotateSecret deleted and re-created; the re-create was dead, because
  the next read makes one anyway. Removed, and a second path to the same
  act is how two ends come to disagree.

And one real fault: three places built a declaration, and the one behind
`--json` predated credentials, so it silently produced a declaration
missing them — a difference between what `plan` showed and what anything
reading `--json` got. There is one path now.
2026-08-30 00:21:18 +02:00
jschoubben 46e760fc94 The control plane serves, and a node can join
There was no chicken-and-egg to solve. The mesh runs the broker, so it creates
the node's account when it issues the token, and the one-time secret is that
account's password. A joining node's first connection is already authenticated;
enrolment is what it says once it is in. I had been treating this as a decision
that needed taking, and it did not.

The account is per node and scoped: it may read its own queue, write to the one
exchange, and configure nothing else. The patterns are anchored and the node
name is constrained to characters that cannot widen them, because a name
carrying a dot or a star would silently let that node read everybody's queues.

`serve` is the control plane running: one connection, one queue, one consumer.
One deliberately -- two consumers on a queue get round-robined and each receives
half of what it expects, which has happened on this project before, between a
module's daemon and its capability server.

Enrolment spends the token first, in the single statement that both finds and
marks it, and only then records the key. That order is the order things become
irreversible: recording a key for a node whose token turned out to be spent
would leave the mesh believing a machine that never had the right to join.

Refusals are one message for every reason. The log says which, where an
operator can see it; the node is told only that the token cannot be used.

Verified in the lab, on a sealed machine, through the whole first-node path.
2026-08-29 16:03:14 +02:00
jschoubben 306c4ca13b The control plane, as far as identity
Tier 2 exists now. It holds one context of seven, inventory, and does one
thing with it: brings its schema up to date. That is step 3 of the substrate
bootstrap -- the step the first node cannot get past.

Verified against a real PostgreSQL, with the built binary: applied 0001-nodes,
reported 'already up to date' on the second run, and the node table is there
with the index and the unique constraint the migration asks for.

Written in Go, and the image is FROM scratch holding one file. Confirmed by
unpacking it. That is the whole argument of ADR 0024: the bundle pins this
image by digest and runs it where nothing can check it, so everything in it is
something a person has to audit before trusting a first node.

Exclusive store ownership is built as a rule about credentials rather than
about intentions. There is no mesh-wide connection setting and no way to ask
for one -- a context reads MESH_STORE_<ITS OWN NAME> and holds nothing else, so
reaching another context's store needs a new variable, which is visible in the
declaration that runs it.

The migration runner is mostly refusals: an edited migration that already ran,
a migration numbered below one that has run, duplicate numbers, misnamed files,
empty files. All stop rather than warn, because at the moment any of them is
true nobody knows what the database holds.

It stops before identity, deliberately. What a node presents to prove who it is
has not been decided anywhere, and a migration is the most expensive place in
this system to guess.

Two tests did not defend what they claimed, and both are fixed rather than
removed. One asked only whether Open returned an error, which it did either way
-- a bad context name and a missing credential both fail, so deleting the name
check changed nothing. The other claimed to prove the migration runs in a
transaction, but PostgreSQL already wraps a multi-statement query in one of its
own, so it passed with the transaction taken out. What the transaction actually
buys is that the schema change and the row recording it commit together, and
there is now a test for that which fails when they are split.
2026-08-29 02:44:09 +02:00