Commit Graph
8 Commits
Author SHA1 Message Date
jschoubben aecac5bda2 One bus: the AMQP transport is gone from the controller
The mesh runs on the seat's bus alone (novox/hq ADR 0131, design 28 task 5.5). The old
transport's consume loop, build request, tool ask, management API and account scoping are
deleted, and the bus switch with them; the controller connects to the broker seat and to
nothing else. The store-window tests keep their assertions on a bus-less fake, and the tests
that only made sense for the old transport's in-memory holding go with it.
2026-09-28 03:36:16 +02:00
jschoubben 1f36787d75 The mesh's four streams, asserted on every start
Task 1.4. The foundation set only — a seat's streams come at registration
and a module's consumers at assignment, neither of which has happened at
genesis (ADR 0118).

Asserted rather than created: a stream that was deleted, or a mesh raised
from a backup, must converge rather than run without the guarantee its
messages assume.

Two things the definitions have to get right, both tested:
- CONTROL names its subjects instead of taking mesh.control.>, because
  heartbeats live under that prefix and a stream of them competes for
  retention with the messages that matter
- EVENTS filters on the event token, which is why that token exists; a
  filter over a module's whole namespace would persist every tool call

Overlapping filters are refused where the set is written: NATS accepts two
streams matching one subject and stores the message twice under two
retentions, which nothing reports.

Adds nats.go as a dependency; it pulled golang.org/x/* forward. Full suite
green.
2026-09-26 21:02:18 +02:00
jschoubben c3b88b9148 Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:40:40 +02:00
jschoubben f04d00b411 The proxy can obtain a public certificate, and asks staging by default
Work breakdown 1.4. The mesh's own authority certifies internal names
and always did; a name reachable from outside needs one the world
already trusts, and there was no ACME anywhere in this repository.

Uses acme/autocert from x/crypto, which was already a dependency — one
indirect addition (x/net, for idna) and no new direct one.

Three things worth more than the feature:

**Staging is the default** (novox/hq 04-ISSUES/004). Production issuance
is rate-limited per domain and per account and does not replenish
quickly. Defaulting to production would leave the safe path depending on
remembering to opt out, on exactly the work most likely to iterate. A
staging certificate is trusted by no browser, so the mistake announces
itself on the first request rather than a fortnight later.

**A certificate is only asked for on a name the mesh routes here.**
Without that policy, anything that can reach the port and send a name
triggers an order for it — a scan becomes a stream of failed orders
against the account's rate limit, and the proxy looks healthy
throughout. What it may certify is what it was told to serve.

**A private issuer is trusted by naming a file, never by skipping
verification.** Skip would still apply on the day this points at a
public issuer, and nothing would say so.

TLS is opt-in: without TLS_LISTEN the proxy serves plain HTTP exactly as
before, which is what an internal-only mesh wants. With it and no cache,
it refuses rather than defaulting — every restart would otherwise order
new certificates, silently, until the rate limit says it does not.
2026-08-31 19:35:34 +02:00
jschoubben c3046dcf56 A provider is told who its consumers are, and a reference provisioner
Contributions were node-local, so a mesh-scoped provider — the one case
that most needs them — never heard from its consumers. A database was
given a password and no idea what to create it for.

Cross-node consumers now reach the provider's `receives` file, merged in
with the ones on its own machine: from the provider's side they are the
same thing, and a provider that had to read two lists would read one of
them. Each names the file its credential is in rather than carrying it,
because the mesh discarded the value and could not put it there. The
readable half therefore stays readable.

And examples/postgres-provisioner, which is the last step: it reads what
the host wrote and makes PostgreSQL accept it. Explicitly not part of the
control plane — the control plane decides and never touches a machine.
This runs on the machine and touches it, and a real one ships with the
module that ships PostgreSQL. It lives here because this is where the
contract is defined, written as something that runs so it can be read.

It reconciles rather than applying a change, because it is never told
what changed. Three things that follow, and each is a fault somebody has
shipped:

- the password is set every time, not only on creation, or a rotation
  reports success and changes nothing
- what it made and nobody asks for any more is revoked, or a departed
  consumer keeps a working login for ever
- what it did not make is left alone, or it cannot be run on a database
  that predates it

Proven in the lab against a real PostgreSQL, each assertion confirmed to
fail with the behaviour removed. The suite is in mesh-lab, which also
records the two ways the test itself was wrong first.
2026-08-30 01:31:25 +02:00
jschoubben 20f78cd5f1 Credentials the mesh delivers and cannot read
HAL keeps env vars in the registry, encrypted at rest. Its own tooling
records what that bought and what it did not. `secret_locate` matches by
value rather than by name — because the same password sits in
mesh_provisions, in module_env, in each node's .env in plain text, and
inside every connection string composed from it, and its documentation
says those URL copies "are often the only copies actually in use". And a
query against the encrypted column returns zero rows and proves nothing,
so auditing moved to the decrypted copies on the nodes.

Two faults there, and encryption at rest addresses neither: the control
plane can read what it stores, so a copy of the database is a copy of
every credential; and one secret has many homes with nothing tracking
them.

So here the mesh generates a password, seals it to each end with keys
those nodes generated, stores both blobs, and discards the plaintext. It
cannot read what it holds. Neither can the broker relaying it. And
nothing is composed centrally — a connection string is assembled on the
machine that needs one — so no copy is ever minted in a shape nothing
tracks. `Compromise of a node is compromise of that node` (ADR 0004) is
now true of secrets, not only of identity.

Two files rather than one, because the mesh cannot compose a document
containing a value it discarded: `binds` carries the readable facts,
`secrets` carries the credential alone. The readable half stays readable
in the declaration; the secret half changes only when the secret does,
which makes restart-on precise. The provider gets a directory, one file
per consumer, for the same reason.

It is made once and kept — regenerating per declaration would restart
both ends on every push, and the password a provider was told to create
would never be the one its consumer was given. It is remade when either
end's sealing key changes, and both ends learn the new one in the same
push, so there is no window where half the mesh holds a dead credential.

Two tests found passing for the wrong reason, both caught because their
injection came back clean:

- the provider's copy was asserted non-empty, which reads the same
  whichever column is selected. It now opens the blob with the
  provider's own key.
- RotateSecret deleted and re-created; the re-create was dead, because
  the next read makes one anyway. Removed, and a second path to the same
  act is how two ends come to disagree.

And one real fault: three places built a declaration, and the one behind
`--json` predated credentials, so it silently produced a declaration
missing them — a difference between what `plan` showed and what anything
reading `--json` got. There is one path now.
2026-08-30 00:21:18 +02:00
jschoubben 46e760fc94 The control plane serves, and a node can join
There was no chicken-and-egg to solve. The mesh runs the broker, so it creates
the node's account when it issues the token, and the one-time secret is that
account's password. A joining node's first connection is already authenticated;
enrolment is what it says once it is in. I had been treating this as a decision
that needed taking, and it did not.

The account is per node and scoped: it may read its own queue, write to the one
exchange, and configure nothing else. The patterns are anchored and the node
name is constrained to characters that cannot widen them, because a name
carrying a dot or a star would silently let that node read everybody's queues.

`serve` is the control plane running: one connection, one queue, one consumer.
One deliberately -- two consumers on a queue get round-robined and each receives
half of what it expects, which has happened on this project before, between a
module's daemon and its capability server.

Enrolment spends the token first, in the single statement that both finds and
marks it, and only then records the key. That order is the order things become
irreversible: recording a key for a node whose token turned out to be spent
would leave the mesh believing a machine that never had the right to join.

Refusals are one message for every reason. The log says which, where an
operator can see it; the node is told only that the token cannot be used.

Verified in the lab, on a sealed machine, through the whole first-node path.
2026-08-29 16:03:14 +02:00
jschoubben 306c4ca13b The control plane, as far as identity
Tier 2 exists now. It holds one context of seven, inventory, and does one
thing with it: brings its schema up to date. That is step 3 of the substrate
bootstrap -- the step the first node cannot get past.

Verified against a real PostgreSQL, with the built binary: applied 0001-nodes,
reported 'already up to date' on the second run, and the node table is there
with the index and the unique constraint the migration asks for.

Written in Go, and the image is FROM scratch holding one file. Confirmed by
unpacking it. That is the whole argument of ADR 0024: the bundle pins this
image by digest and runs it where nothing can check it, so everything in it is
something a person has to audit before trusting a first node.

Exclusive store ownership is built as a rule about credentials rather than
about intentions. There is no mesh-wide connection setting and no way to ask
for one -- a context reads MESH_STORE_<ITS OWN NAME> and holds nothing else, so
reaching another context's store needs a new variable, which is visible in the
declaration that runs it.

The migration runner is mostly refusals: an edited migration that already ran,
a migration numbered below one that has run, duplicate numbers, misnamed files,
empty files. All stop rather than warn, because at the moment any of them is
true nobody knows what the database holds.

It stops before identity, deliberately. What a node presents to prove who it is
has not been decided anywhere, and a migration is the most expensive place in
this system to guess.

Two tests did not defend what they claimed, and both are fixed rather than
removed. One asked only whether Open returned an error, which it did either way
-- a bad context name and a missing credential both fail, so deleting the name
check changed nothing. The other claimed to prove the migration runs in a
transaction, but PostgreSQL already wraps a multi-statement query in one of its
own, so it passed with the transaction taken out. What the transaction actually
buys is that the schema change and the row recording it commit together, and
there is now a test for that which fails when they are split.
2026-08-29 02:44:09 +02:00