Commit Graph
100 Commits
Author SHA1 Message Date
jschoubben 3ffddff8ed The builder announces what it built, and what it was built on top of
Answering and announcing are different acts. The reply goes to whoever asked and
is correlated to their request; the announcement says to the whole mesh that a
module now exists at a commit, which is what the catalogue places in the module
graph (novox/hq ADR 0072). A build nobody asked for still has to be announced, or
the graph knows less than the registry does.

What it was built on top of is read out of the build's own inputs rather than
declared, because a declared list drifts from what the code actually uses
(ADR 0009). These are artifact references, which is what a build input names;
resolving them to module-versions is the catalogue's work, since it is what knows
which module-version published which artifact.

Events ride the topic exchange, not the direct one nodes speak over, so the
builder's account is granted both: it must be able to answer and to announce.
The envelope is the sdk's, reproduced exactly — a second shape would be a second
thing for consumers to handle, and they are written against the first.

Announcing is not allowed to fail a build. The work was done and was answered; a
build reported as failed because saying so failed is a lie about it.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-13 00:57:41 +02:00
jschoubben f151de103f Build a module from a repository and a path within it
The builder cloned a repository and read the manifest at its root, which means one
repository per module. Nothing we have is shaped that way, so the builder could be
asked to build nothing that exists (novox/hq ADR 0069).

The path travels the whole way — named when asking, carried in the request, used
to read the manifest and as the context everything is produced from, echoed back
in the result, and recorded as part of where a module came from. Without that last
part the mesh could notice a module was behind its source and then be unable to
rebuild it, which is the worst of both.

A path climbing out of the clone is refused: a machine whose job is building other
people's repositories must not read whatever else is on its disk.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 16:45:50 +02:00
jschoubben c4030947b0 The routing record is 0066, not 0056
0056 is 'the authority is the control plane, not a database'. A citation
pointing at the wrong decision is worse than none: it reads as corroboration.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:09:56 +02:00
jschoubben 522d8be925 Read a context's store connection from a file, not only from the environment
A store connection string carries a password, and the control plane took it
from MESH_STORE_<CONTEXT> — an environment variable, which is readable in
`docker inspect`, in the process's own /proc entry, and in whatever composed
it. Every other module in the catalogue is given secret material as a file the
mesh sealed to the machine and the host wrote.

That difference is what stopped the control plane from being an ordinary module
(novox/hq ADR 0067). A manifest can put a sealed value into a file's `content`
with ${secret:…}; it has no substitution into a container's `env` at all. So a
control-plane manifest could be written with the password in it, or without the
setting — neither honest. The fix is not to change the manifest format but to
let the control plane read what everything else reads: a file.

MESH_STORE_<CONTEXT>_FILE names one. Exactly one of the two may be set; both is
refused rather than settled by precedence, because whichever won, the other
would still read as the setting in force and the process would be writing to a
store nobody expects. Trailing whitespace is trimmed — a file written by a
person or by a filled-in placeholder ends in a newline, and a newline inside a
URL is rejected several layers from anything that could explain it. Leading
whitespace is left, being a mangled value rather than a habit.

The no-leak property is kept and extended: a file that cannot be read names its
path, never its contents, and the parse failure now names whichever source was
used because a variable name and a path are not the secret.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:40:37 +02:00
jschoubben c6029504e4 Merge branch 'fix/settings-assign-and-cli' into feat/adr-0056-compose-and-propagate 2026-09-10 21:17:46 +02:00
jschoubben 72f30e9c29 overlay: placing a node with nothing said no longer unplaces it
The sibling of node public-domain, and the worse one: a placement is three facts
declared together, so an invocation that said none of them took all three away —
the endpoint every other machine dials, the site, and the hub. A mesh whose hub
was placed that way has no paths left, at the moment somebody was trying to look
at it.

--nothing keeps the real case (a machine that roams and opens every path itself)
sayable, by name.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:17:09 +02:00
jschoubben 0be227bd7e status: one machine that cannot be worked out no longer takes the answer from the rest
`status --json` emitted no JSON at all when a single node was unresolvable. A blocked
node is not on the private network, and a mesh whose hub is that node has no hub — which
came back through the reading as a refusal, so `status` printed nothing and `status
--json` put multi-line prose on stderr and not one byte on stdout. A machine-readable
interface that stops being machine-readable exactly when something is wrong is one nobody
can build an alarm on.

Why a machine cannot be worked out is read as data now, per machine, through the same
whoResolves the private network is built from — so this and the network agree about who
could not be resolved rather than deciding it twice. The private network failing to
compute is kept as a note beside it instead of ending the read: it is almost always a
consequence of those same refusals, and every question that does not depend on it is
still answered.

It reaches all three ways of saying it, from the one reading: the text form leads with it
because a machine here is in none of the answers below, the JSON carries `unresolved`
(always a list, never null) and `network`, and the page has a section of its own.

That also closes a silent success. A machine that resolves to nothing has nothing
computed for it, so there is nothing to compare it against and nothing it can be behind —
it appeared in no answer at all, and `status` reported a mesh where nothing could be sent
anywhere as "all doing what they were told".

statusAsJSON takes the whole reading now rather than a growing argument list, which is
what let an answer be added to the text form and forgotten here. The two are one
function's output in two shapes and must not be able to differ about what was asked.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:12:06 +02:00
jschoubben 946fddd622 catalogue: a module may name the machine it was assigned to
An authority inside the mesh is reached at <machine>.internal, so its own
certificate must be issued for that name — and it is the one module that cannot
be told its name by a binding, because it provides rather than requires. Written
as a literal it would be one deployment's machine name in a manifest, which is
what ADR 0056 exists to remove.

${machine:name} and ${machine:at}, beside the bound values and refused the same
way. An address the machine does not have is named here rather than discovered
later as a certificate nobody can verify.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:12:04 +02:00
jschoubben acbb8cf6da node public-domain: asking what it is no longer takes it away
`node public-domain <name>` cleared the domain. It reads like a question — it is exactly
what anybody types to find out what the answer is — and it silently took every routed
name the node had. There is no output that makes up for that: by the time it prints, the
fact is gone, and the mesh cannot tell a person what a domain used to be.

The bare form reports now. Clearing is still a real thing to want — a machine that stops
facing the outside composes no names, and lab-versus-production is this one setting
(novox/hq ADR 0056) — so it keeps a way to be said, by name: `--clear`. A domain and
`--clear` together are refused rather than one of them silently winning.

The other `node` subcommands were checked. `add`, `list` and `show` write nothing they
were not asked to, so there is nothing to make consistent with.

`overlay place <node>` with no flags has the same shape — it clears the endpoint, the
site and the hub flag — and is deliberately left alone here. It is a verb rather than a
question and every caller passes flags, so the fix is a different judgement and belongs
in its own change.

The usage text gains the three forms, and `module forget`'s new flag, neither of which
it named before.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:11:48 +02:00
jschoubben d032fe6e8d assign: what it checks is the mesh, not the one machine
An assignment was verified by resolving the node it was made on. The verify that matters
is resolution over all of them: a module offering a mesh-scoped provision stops offering
it the moment its own node stops resolving, so an assignment could be reported as fine
while it took that provision away from every consumer elsewhere. Those consumers were
then told "nothing in this mesh provides it", naming as the remedy a module that was
already assigned — a wrong answer about a machine nobody had touched.

That is novox/hq 04-ISSUES/017's shape exactly: an action succeeds into a state its own
verify rejects, and it does so because the action's own test is not the test the verify
uses. 017's remedy was to make them the same test, and this makes them the same test.

The assignment is still kept, and that is the other half of the decision. Assignment is
not an ordering: a consumer assigned before its provider does not resolve for as long as
it takes to assign the provider, and refusing the first half of a pair would make the
order somebody types two commands in part of the mesh's rules. So `assign` and `unassign`
now name every OTHER machine that cannot be worked out as things stand, in the mesh's own
words, beside whatever they already said about this one. It reports the state and never
claims causation — saying "this assignment broke laptop" would mean resolving the whole
mesh twice and would still be a guess about which of several changes did it.

It costs a resolution per machine. Assignment is a person typing a command, and being
told which machines this just blocked is worth more than the milliseconds.

This is what a four-node raise read as "bumping a module's version broke provider
recognition". It was neither the version nor the provider: nothing in this codebase reads
the version column, every lookup is keyed on the module name alone, and a test in the
previous commit now says so. It was one machine's set of assignments, and nothing said so.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:11:33 +02:00
jschoubben f5f860fd2b inventory: forgetting a module says what goes with it, and refuses until told
`module forget` cascaded. The settings, the module's own secrets and the ports the mesh
chose all name the module by a foreign key that cascades, so removing the row took all
three and reported "forgotten" — an action succeeding into a state its own verify would
reject (novox/hq 04-ISSUES/017). A sealed secret is not recoverable afterwards, because
the mesh discarded the plaintext when it made it.

It now reads what it would destroy, names each thing one at a time, and refuses.
`--and-what-it-holds` is how somebody says they mean it, and the removal then reports
what went — this being the only record that any of it ever existed.

Reported as "operator settings do not persist, because re-registering a module
cascade-deletes them". Half of that is wrong, and the test now says so out loud: the
upsert is on the name, so `module add` at a new version leaves the settings, the secrets
and the ports exactly where they were. The command that destroyed them was `forget`, and
a wrong belief about which command destroys data is expensive in both directions — it
sends people looking for a fault that is not there, and leaves the real one unexamined.

Checked by internal/inventory/forget_test.go, which writes all three, re-registers the
module at a new version, reads them back, and only then tries to forget it.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:11:15 +02:00
jschoubben 9fab0b731a catalogue: a contribution reaches the port its own machine published, from any machine
The co-located fix could not reach a grant assembled for a consumer on another
machine: ContributionsFrom never sees a port map, so the proxy was told the
workload's software port and dialled a number that machine never published. The
consumer's own assignments are fetched where the grant is built and applied
there. The same fault as 038, one node over.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:08:09 +02:00
jschoubben e78c849002 catalogue: a provider that names a provision and says nothing is not the answer
servedOnThisMachine stopped at the first module whose `serves` mentioned the
provision, even when that entry was empty and there was therefore no fact to give
a consumer. here() had always kept looking in that case, and a set where one
module names a provision without describing it and another describes it is exactly
where the difference shows. Restore the search.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:04:16 +02:00
jschoubben 76fcbda9a7 route-proxy: an ACME account belongs to the authority that issued it
autocert keeps its account key at one fixed name, `acme_account+key`, in whatever
directory it is given, and reuses it for ever. That is right while the authority
stays the same and silently wrong the moment it does not. Re-initialising the
internal CA makes a new authority with a new root: it has never heard of the
account in the cache, rejects every use of it, and autocert has no path back from
that. Nothing is re-registered, no order ever reaches the CA, and issuance stops
with nothing saying why — until somebody guesses that deleting the cache directory
by hand is the answer.

Name the directory after the authority instead of sharing one between all of them:
a digest of the ACME directory URL and the root this proxy was told to verify it
with. A re-initialised CA has a new root, the mesh delivers it as a changed bundle,
the proxy restarts on that file and lands in a directory with no account in it, so
autocert registers afresh and orders again. The healing is that "is this account
still valid" never has to be asked — an account is only ever found where it is
still valid, which needs no error codes, no probe at startup and no network call
that can itself fail.

It closes a latent one of the same shape: pointing ACME_DIRECTORY at production
after testing against staging reused the staging account, because the cache had no
idea the two were different.

Trailing whitespace around the delivered root is not a new authority — the mesh
writes that file, and a newline coming or going must not throw away an account.
Old directories are left on disk, unused: they hold the only copy of certificates
that may still be valid, and this program is not the thing that should decide a
certificate is finished with.

A proxy already holding certificates orders them once more on the first start after
this, because its account moves. Free against the lab's own CA and against staging;
one issuance per name against a public authority.

novox/hq ADR 0056

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:02:00 +02:00
jschoubben 54cb915c4e push: one machine that cannot be worked out no longer holds back the mesh
A whole-mesh push refused outright the moment any single node failed to resolve:
every node was composed first, and one entry in `refusals` returned "nothing was
sent" before the send loop ran. So on the ADR 0056 lab, one module on the anchor
requiring a provision nobody had assigned a provider for — `nothing provides
"acme-ca", wanted by route-proxy` — stopped every OTHER machine from being sent
anything. Nothing converged anywhere, and the machines that went unconverged were
the ones with nothing wrong with them. The failure and the punishment were on
different machines.

This is the rule a92c11b established one level down, where an un-hostable module
stopped taking down the healthy modules beside it, applied one level up: the blast
radius of a fault is the thing that has it. A node whose declaration composes is
sent; a node whose does not is named, with its reason, and the push still ends
non-zero — skipping is not succeeding, and a command that exits cleanly having
missed a machine is a command that lies. The message now says how many machines
WERE sent, because "nothing was sent" was the claim that had become untrue.

sendTo keeps its all-or-nothing rule, and its comment now says that push
deliberately does not share it: a rotation reaching the consumer and refusing on
the provider leaves one end holding a credential the other has never heard of,
which is a real coupling between two named machines. A whole-mesh push has no such
coupling and never did.

Composing is split out of pushCommand so the rule can be tested without a broker
and a database.

novox/hq ADR 0056

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 20:59:21 +02:00
jschoubben b824c65ab0 catalogue: a co-located provider's served values, and a co-located contribution's port
Two more of one fault, and the fault is the same as 04-ISSUES/038: the same-node
path diverging from the cross-node one.

The mesh works out what a provider on ANOTHER machine serves by walking that node
— reading its manifest with that machine's port assignments, then settling the
result with that node's settings layers — before offering it to a consumer. A
provider on the consumer's OWN machine never passes through that walk, so every
step of it had to be repeated in resolve.go's servedHere and declaration.go's
here(). 038 repeated the port. Nothing repeated the settling.

So a served value the operator supplied reached a co-located consumer as the
manifest's empty default. On the ADR 0056 anchor that value is an internal CA's
root: step-ca and route-proxy on one node, route-proxy's binding carrying
root: "", an empty CA bundle written, a silent fall back to the system trust
store, and issuance stopping with nothing saying why. The same step-ca on another
node would have worked.

The second is the mirror direction. gitea declares a bare container port 3000 and
the machine publishes it as 20000:3000, but gitea's route CONTRIBUTION still said
3000 — so the proxy beside it dialled a port nothing listens on and answered 502.
038 fixed what a consumer is TOLD about a provider; this is what a workload TELLS
a provider about itself. The redirect uses the CONTRIBUTING module's assignment,
because the port is the workload's, not the proxy's; a contribution carried here
from another machine is left exactly as it is, its port being that machine's to
assign.

Both are settled in Declaration, which is the first moment the machine's ports and
the provider's settings both exist. That also removes an order dependence: the
resolver built its same-node needs mid-walk, from whichever modules had been
chosen by the time the requirement came up and in whatever order a map iterated,
so what a co-located binding carried depended on the order somebody happened to
assign things in. Re-deriving from the finished closure does not.

servedHere keeps its job — deciding whether a same-node provider serves anything
at all, which is what makes the need exist — and now says that its values are
provisional.

novox/hq ADR 0056

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 20:57:23 +02:00
jschoubben 0fb2ab7716 catalogue: composeName handles the apex label '@' (bare public domain)
An empty label composed nothing, so a module served at the bare domain (a node's
own site) had to keep a full name — the one route the label model could not
express. The zone-file convention '@' now composes to the public domain itself,
no leading dot, so the apex is a label like any other. Test added.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 00:27:11 +02:00
jschoubben 98f5d34610 route-proxy: an empty ACME_CA_BUNDLE means the system trust store
Piece A of ADR 0056 (selectable issuer). A provider that serves an empty root —
public-acme, whose root already ships in the OS trust store — leaves route-proxy's
CA bundle file existing but empty, because the mesh writes it unconditionally from
${bound:acme-ca:root}. Read that as "trust the system roots", the same as an unset
bundle, instead of failing with "holds no certificate this can trust". A bundle
that holds bytes but no parseable certificate is still refused.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 00:22:06 +02:00
jschoubben 232862315c catalogue: compose a route's name from a label and its node's domain, and resolve it in-mesh
A public route used to carry its whole hostname as a literal in the module
manifest, so running the same catalogue against a different domain meant
overriding that literal on every routed module, per node. The mesh was, in
effect, holding a map of names to services: the one thing it should never hold,
because the subdomain is the operator's choice and the domain is the node's.

Compose instead. A route contribution carries a `label` (the subdomain); a node
carries its `public_domain` as node-level configuration; the mesh joins
`<label>.<public-domain>` and grants exactly that, interpreting neither half.
Held as a node property beside the node's other node-level facts (endpoint,
site, overlay address), not in a module's settings — the ADR calls it
node-level, and the settings table is keyed per module.

Additive, so an unmigrated catalogue keeps working: a contribution that still
carries a full `name` and no `label` passes through unchanged, and the catalogue
can migrate module by module. A labelled contribution on a node with no public
domain composes nothing, reading downstream as a route that named no host.

And propagate: each granted route name is published into internal resolution
mesh-wide, mapped to the node that serves it, alongside the `<node>.internal`
names every container already gets. So a container — and an internal ACME
validator, which cannot complete a challenge for a name it cannot reach —
resolves a routed name to the proxy that serves it. Name-agnostic throughout:
the mesh propagates whatever names it was told to serve and knows nothing about
what they mean.

novox/hq 02-DECISIONS/0056

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 23:27:35 +02:00
jschoubben c147a26138 catalogue: announce a same-node provider at the port it is published on
A co-located consumer of a `from: mesh` provision was told the port the
provider module DECLARED, not the host port the mesh assigned and published
it on. The same-node served facts are settled while resolving (servedHere,
here()), before a bare `ports` mapping is assigned its host port, so they
carried the declared number; only the cross-node path re-derived them after
assignment. So the provider was published on <node>.internal:<assigned> while
its own-machine consumer dialled <node>.internal:<declared>, where nothing
listens — the ordinary small-mesh case, and the one the fix for issue 018
(announce the same-node provider at all) left one promise short of kept.

Redirect same-node needs to the machine's assignment in Declaration, where the
port map is known, exactly as plan.go already does cross-node. The publish bind
is unchanged (all interfaces, scoped to the mesh by the listen's firewall rule);
only the announced port is corrected.

novox/hq 04-ISSUES/038

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 23:02:50 +02:00
jschoubben 9e29772ed4 Merge pull request 'resolve: one un-hostable assignment no longer takes down a whole node's push' (#19) from feat/assign-hostability into main 2026-09-08 18:45:21 +02:00
jschoubben a92c11be12 Resolve: one un-hostable assignment no longer refuses the whole node
A module a person assigns to a machine that cannot host it — its declared
capability has no detector there, as fail2ban does on a host with no firewall —
made Resolve refuse the entire node, so a whole-node push refused to send the
healthy modules beside it too. One module on the wrong machine took down every
other module on that node.

Assign already keeps such an assignment on purpose (it is what a person meant,
and acts.go says so), so the fix is on the resolve/push side: a directly-assigned
module the machine cannot host is left out of the closure and reported as
un-applied on the Resolution, rather than refusing the set. The healthy modules
still resolve, declare, and converge. A module that is *required* by something
running here and cannot be hosted still refuses — that set is genuinely
incoherent — so the distinction is who wanted it.

assign, plan and push now name the un-applied module and the missing capability,
via a shared WrongMachine message, so it is neither silently dropped nor fatal.

Reconciled two tests that encoded the old whole-node refusal for directly-assigned
un-hostable modules; added coverage for the healthy-modules-still-converge case
and the required-un-hostable-still-refuses distinction.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-08 18:37:41 +02:00
jschoubben 474382c141 Merge pull request 'catalogue: deliver a keyless same-node provider's served facts' (#18) from feat/serve-here-keyless into main 2026-09-07 05:25:57 +02:00
jschoubben b78a911e34 catalogue: deliver a keyless same-node provider's served facts
A node-scope provider that answers a requirement on the same machine and
serves connection facts (a port) but mints no credential delivered
nothing to a co-located consumer. resolve.go only built the delivering
Needed when brokered[want] was set — true only for mesh-scope providers;
a node-scope keyless provider set local[want] instead and fell through,
so knownFor saw no binding and boundInto refused the consumer's
${bound:model-access:port} file.

Deliver the served facts as a need whenever the same-node answer serves a
non-empty set, with a loopback fallback for the address when the node is
off the private network — the reachability rule does not apply to two ends
on one machine. The brokered (credentialed, mesh-scope) path is untouched.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 05:08:04 +02:00
jschoubben e5eaa0478a Merge pull request 'adapters: openai is a static-key vendor' (#17) from feat/openai-access into main 2026-09-07 04:25:53 +02:00
jschoubben 55753d5cfd adapters: openai is a static-key vendor
One registry line — `"openai": StaticKey`. The generic static-key adapter
already serves any vendor (accept is the vendor-independent seal, deliver is the
value unchanged, and refresh/identity/usage are not implemented), so a second
static-key vendor is data, not code. The shape-selection test now asserts
For("openai") reports static-key.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 04:25:26 +02:00
jschoubben e03833ec40 Merge pull request 'catalogue: give host-network containers the mesh's names too' (#16) from feat/usage-store into main 2026-09-07 04:08:19 +02:00
jschoubben 958bef56c7 catalogue: give host-network containers the mesh's names too
A container with `network: host` was skipped when the mesh injects its
`<node>.internal` names, on the belief it "shares the machine's hosts file
already". It does not: `docker run --network host` still gives the container
its own /etc/hosts (localhost and its own id only), so every internal name the
mesh wrote is invisible inside it, and a client that dials one gets EAI_AGAIN.

This surfaced with the first host-network consumer to dial a provider by the
`.internal` address the mesh hands it as `${bound:...:at}` (the model-usage
store reaching its postgres). The remedy is the same `--add-host` every other
container already gets — the runtime accepts it with `--network host`
(verified against Docker) and mesh-host emits it for any network mode.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 03:44:28 +02:00
jschoubben d9c4818e6d Merge pull request 'model-access: refreshable-grant — manager holds the refresh token, control plane never reads it' (#15) from feat/model-access-submit into main 2026-09-07 02:48:49 +02:00
jschoubben f00077f376 secrets: regenerate the module sealed-box cross-check fixture
The manager module's seal now runs over the audited tweetnacl-sealedbox-js (mesh-catalog
anthropic-manager) rather than a hand-transcribed NaCl. Regenerate the fixture's sealed value
from that new seal() over the same node key pair and plaintext, so the fixture is the new
library's output. crypto_box_seal is randomised, so the blob differs; the wire format does
not. TestModuleSealedBoxOpensInGo still opens it under box.OpenAnonymous and recovers the
plaintext, proving TS(tweetnacl-sealedbox-js) to Go interop.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 02:17:54 +02:00
jschoubben 33fd28ffa6 licences: deliver the refresh token by the ordinary sealed path, not a bespoke envelope
The refreshable-grant refresh token no longer rides a custom at-rest envelope that a
module opens with a node private key. A module is never given a node's private sealing
key, so that path could not exist -- the gap Phase C hit.

Instead the refresh token is a credential sealed to the MANAGER holder with the same
anonymous box (secrets.Seal / crypto_box_seal) every credential uses, stored as one
sealed blob, and delivered by the existing host-unseal-and-mount: the host opens it with
the node's real key and mounts the cleartext at the manager module's bound path, exactly
as a consumer's db password is delivered.

  - refresh_grant now stores { sealed, manager_key }, dropping the AtRest token/wrapped_key
    columns; internal/secrets/atrest.go is retired (nothing else used it).
  - the licence records its manager as (node, module); KeyFor delivers the refresh token to
    the manager holder and the access token to consumers, disambiguated by module so the two
    can co-locate. Accept and the reseal skip the manager holder.
  - the manager holder is delivered the node's PUBLIC sealing key in its bound facts, so the
    module can re-seal a rotated refresh token with no private key of its own; the
    declaration tolerates its empty pre-adoption secret rather than refusing.
  - SubmitRefresh / set-grant take a sealed blob, never a refresh token in the clear.

The invariant holds unchanged: the control plane never reads the refresh token, and no node
but the manager holds it. A committed cross-language test proves the TypeScript module seal
opens under Go box.OpenAnonymous (the host's Unseal) -- both are NaCl crypto_box_seal.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 01:55:08 +02:00
jschoubben 8e0c22fc2e licences: submit-refresh, the module-produced refresh entry point
Phase C of model-access (ADR 0050). Refresh above calls an in-process
VendorRefresher, which would open the at-rest envelope inside the control
plane's own process. Anthropic must not: its refresh runs on the manager
node. So add SubmitRefresh, the companion that publishes a refresh a
manager node already performed -- it is given only the new access token in
the clear (sealed per holder, as any accepted key) and an opaque re-sealed
refresh envelope (stored unopened). The refresh token in the clear never
crosses this boundary. The reseal-and-publish half is extracted and shared
with Refresh, so the sealing logic is one implementation.

CLI: licence grant (print the opaque envelope), set-grant (store a
module-produced envelope -- adoption), submit-refresh (access token +
optional rotated envelope). Tests defend that the manager alone opens the
refresh token and the control plane never holds it in the clear.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 01:00:27 +02:00
jschoubben 59faa150ac Merge pull request 'model-access B: the refreshable-grant machinery (ADR 0050 carve-out)' (#14) from feat/model-access-refreshable into main 2026-09-07 00:25:15 +02:00
jschoubben 2e33c5e80e model access B: refreshable-grant machinery — manager, at-rest refresh token, refresh flow
The ADR 0050 carve-out, built generic and vendor-neutral. A refreshable-grant
licence records one manager node; that node holds the refresh token encrypted at
rest, access tokens are still sealed per holder, and the refresh token is never in
a holder's delivery. Bounded on the three stated axes: refreshable-grant vendors
only, the refresh token only, the manager node only. Anthropic's actual OAuth
refresh stays a Phase-C plug-in behind a clean seam.

- New at-rest crypto (secrets.SealAtRest/OpenAtRest): envelope encryption distinct
  from the per-holder anonymous-box seal. The refresh token is under a symmetric
  data key (secretbox); the data key is wrapped to the manager node's public
  sealing key. The database alone holds ciphertext and a wrapped key with no
  private half to open either — only the manager node reads it back.

- Refreshable-grant adapter dispatch: anthropic is now refreshable-grant,
  anthropic-api-key the static-key second case. The adapter implements the
  Refresher seam by delegating to an injected VendorRefresher (the Phase-C plug,
  none shipped). static-key is untouched. The type assertion to Refresher is what
  gates the carve-out to refreshable-grant vendors.

- Refresh lease/rotate/publish flow (Licences.Refresh): a transaction-scoped
  advisory lock is the single-refresher lease; the new access token comes from the
  vendor refresh, is sealed per holder (secrets.Seal, as Accept does) and delivered
  on the next push — doc 13's reseal-and-publish half, all-or-nothing. The refresh
  token stays put, re-encrypted at rest only if the vendor rotated it.

- Manager and refresh_grant schema: consolidated into migrations/0001 and carried
  by a new incremental 0003 (the dual-write rule).

- 17 new tests, including the four security checks: KeyFor never carries the
  refresh token, a static key has no manager and cannot be refreshed, the at-rest
  token needs the manager's key, and a refresh delivers a new sealed access token.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 00:23:35 +02:00
jschoubben 163200c4dd Merge pull request 'model-access A: provider→vendor + adapter dispatch + static-key (ADR 0050 foundation)' (#13) from feat/model-access-vendor into main 2026-09-07 00:02:27 +02:00
jschoubben ddb41baaf4 Model access is vendor-agnostic: rename provider→vendor, add adapter seam (Phase A)
ADR 0050 Phase A. Rename the licence's `provider` field to `vendor` — the
inventory already uses "provider" for which node answers a brokered provision,
and one word must not carry two facts — and route the licence layer's sealing
and delivery through a per-vendor adapter selected by that field.

The rename touches the Go struct/params/SQL in internal/licences, the operator
CLI, and the schema: 0001 (the consolidated schema) now creates the column as
`vendor`; a new guarded 0002 renames it on a database that predates the change,
and is a no-op on a fresh one.

The adapter (internal/licences/adapters) has a `shape` and the two verbs a
static-key vendor needs — accept (the generic anonymous-box seal) and deliver
(the sealed blob unchanged). refresh/identity/usage are named as optional
capability interfaces so the refreshable-grant seam exists before its code.
A registry maps vendor→shape (anthropic→static-key for now, with a Phase-B
TODO to swap it to refreshable-grant); an unknown vendor is refused clearly.

Behaviour is unchanged from the operator's view except the field name.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 23:49:12 +02:00
jschoubben 671fb4f8f3 Merge pull request 'fix: a require-only consumer of a parameterless provision still asks (mint gap)' (#12) from fix/require-only-mint into main 2026-09-06 23:21:25 +02:00
jschoubben d0ef659824 fix: a require-only consumer of a parameterless provision still asks (and is minted a credential)
A consumer that requires a provision whose serves names no consumer key (redis-cache, amqp)
contributes no payload, but it still ASKS for it. ContributionsFrom keyed 'asks' on
contributions alone, so such a consumer's grant got From='' — read as withdrawn — and the
provider never created its account. redis-cache consumers (e.g. baserow) were silently
unprovisioned, tolerated only by their embedded fallback. A module asks iff it still requires
the provision, whether or not it hands anything up. Regression test added.

Found by the lavinmq AMQP provider bed (given:[] for a require-only amqp consumer); fix
lab-proven green there.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 23:20:58 +02:00
jschoubben b69fbc0e86 Merge pull request 'schedule: parse + validate a cron on a container, refuse run-once+schedule (ADR 0053)' (#11) from feat/schedule-container into main 2026-09-06 14:27:41 +02:00
jschoubben 78b8b6e256 catalogue: carry and validate a container schedule (ADR 0053)
A container may declare schedule: "<cron>", the recurring twin of
run-once. The resolver already carries a resource's keys through
untouched, so schedule reaches the rendered host declaration on its own;
what belongs here is refusing, near its author, what the host would
otherwise refuse far away.

The manifest parser refuses a schedule that is not a string, one that is
not a well-formed five-field cron (cron.go: fields, ranges, *, comma,
dash, slash), and the contradictory pair run-once + schedule -- a
container runs once and gates, or on a cadence, or stays up, never two.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 14:08:51 +02:00
jschoubben dc65440e31 Merge pull request 'fix: a ${secret:} placeholder must fill from the file-owner's credential (provider-seal-key gate)' (#10) from fix/secret-per-consumer into main 2026-09-06 13:47:22 +02:00
jschoubben 99a753994e fix: fill a ptr-secret placeholder from the file-owner's credential, not the last consumer's
The provider-seal-key gate: on a node with two modules requiring the same provision (baserow
and letta both consuming postgres), sealedFor matched a need by provision NAME alone, so a
file's ${secret:X} placeholder took whichever consumer's sealed credential came last in
r.Needs -- the OTHER module's password. baserow was handed letta's password and could not
authenticate. The secrets:-map delivery path already guards this (For == m.Module, novox/hq
04-ISSUES/022); the ${secret:...} placeholder path did not. Added the same guard.

Also dedups the contributions file: when provider and consumer are co-located, grantsFor
enumerates the same-node consumer, so a consumer was emitted twice into the provider's
receives file (once full with its grant, once partial). The m.Contributes loop now skips a
(provision, module) the grants loop already carried; non-grant contributions (routes) still emit.

Regression test added: two consumers of one provision each get their own credential. Proven
end-to-end on a two-node lab install (mesh-lab assigned-two-node-db): baserow and letta on one
node, substrate on another, each authenticates with its own minted password.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 13:46:55 +02:00
jschoubben 3b0f17b8fd Merge pull request 'run-once: a container the host runs to completion (ADR 0052)' (#9) from feat/lifecycle-run-once into main 2026-09-06 00:05:06 +02:00
jschoubben f96c247c5f catalogue: run-once is a step the host runs to completion (ADR 0052)
A container may be marked `run-once: true` — a step the host runs to completion,
gating whatever the declaration places after it. The control plane's part is
small: the field is carried to the host unchanged (containers pass through as
maps), and the step keeps its author-order position ahead of the container it
gates, because the gate is declaration order, not a resolved dependency
(ADR 0005).

The manifest parser refuses a run-once that is not a boolean and the pair
run-once + restart-on (contradictory lifecycles) — near the manifest rather than
far away on the machine, the same lesson the action ban records. Three unit
tests; go build ./... and go test ./... green.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 23:57:33 +02:00
jschoubben 6bb9434298 Merge pull request 'A module accesses operator-owned data, it does not own it (ADR 0051)' (#8) from feat/shared-data-access into main 2026-09-05 22:48:11 +02:00
jschoubben bcdde475cd Merge pull request 'A bare alive moves last_seen and nothing else (convergence race fix)' (#7) from fix/service-only-converge into main 2026-09-05 22:47:23 +02:00
jschoubben cae9a3e54f A bare alive moves last_seen and nothing else
A node says it is there every minute and describes what it applied rarely,
and both went through Heard, which wrote every one down as a report. So a
bare alive replaced the node's last real apply with an empty one -- clearing
the declaration digest `current` is measured against, the carried ports a
push assigns around, and the clean-or-failed outcome. A node that had just
caught up read as behind within the minute, and never converged.

Whether it converged in time was a race the node's own apply set: the link's
one loop applies a declaration to completion before it can send the pending
heartbeat, so a fast apply (catalogue-small) leaves the digest standing the
~60s until the next beat -- long enough for the lab to see `current` -- while
a heavy wave whose apply outran the first beat (mongodb + unifi + marrytts)
had the alive fire milliseconds after the report and never showed `current`
at all, timing out settle even at 1200s.

Heard now returns after moving last_seen for a report that carries no account
of what the machine did -- nothing applied, nothing refused, nothing failed,
which is exactly a bare alive. A real report always carries one. This is what
the commit that began hearing alives said it did and did not: "a bare word
that a node is there moves last_seen and touches nothing else."

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 22:37:53 +02:00
jschoubben aeb65a3e1d catalogue: a module accesses operator-owned data, and does not own it
04-ISSUES/036: the media stack is several modules that must share the
library and download directories on one machine, but the manifest could
only say "a directory I own". Six modules each declared the same paths as
their own resources, and the resolver's duplicate-owner refusal — right
in general — would refuse the stack's only sensible assignment the first
time two of them landed on one node.

Add an `accesses` field: a pre-existing, operator-owned path a module is
granted use of but does not own (novox/hq ADR 0051). Distinct from a
`directory` resource on every axis the host acts on — the mesh creates,
chowns and reconciles a directory; it mounts an access and owns nothing.
An access is not a resource, so it never enters the duplicate-owner map
and several modules may name one path with no conflict. What is refused
is the contradiction: a path one module owns and another accesses.

Rendered into the declaration as an `access` resource, before the
container that mounts it, so the host can find it present or refuse
clearly. Unit tests cover co-resolution (the exact 036 case), the
unchanged owner-vs-owner refusal, the owner-vs-accessor refusal, and
access validation.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 22:10:58 +02:00
jschoubben 59fcb41855 Merge pull request 'Resync hq ADR references (0044-0054 -> 0039-0049)' (#6) from feat/adr-ref-resync into main 2026-09-05 12:47:52 +02:00
jschoubben 87c193b202 Resync hq ADR references 0044-0054 -> 0039-0049 after the hq record reconciliation 2026-09-05 12:47:13 +02:00
jschoubben 5541949b59 Merge pull request 'broker: a module account scopes its tool serve queues + mesh.rpc (ADR 0052)' (#5) from events/tool-account-scope into initialization 2026-09-05 03:06:51 +02:00
jschoubben 9b7ba2e20c identity: a consumer's identity fits the tightest backend, via a slug (ADR 0054)
A module may declare a short `slug`; the mesh derives mesh_<node>_<slug|name> and
refuses at assignment (naming the slug as the remedy) when it would still overflow —
identityLimit is now 20, an S3 access key's, the tightest of the backends a login
reaches (04-ISSUES/010). The slug rides the grant so the provider derives the same
login the consumer does, even across nodes. CheckIdentity is now wired, in grantsFor.

Also, the minted secret shrinks to 40 chars (30 bytes) from 43: an S3 secret key is
8-40, the same fit-the-tightest-backend rule on the credential's other half.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 02:51:39 +02:00
jschoubben b1bf1659d9 broker: a module account scopes its tool serve queues and mesh.rpc (ADR 0052)
CreateModuleAccount now also grants serve.<module>.* (declare, bind, consume its
own tool queues) and mesh.rpc (bind them on, publish replies) — so a module can
serve its tools and reply, scoped to exactly its own, and no other module's. The
broker tests still hold a module out of another's queue.
2026-09-04 21:56:05 +02:00
jschoubben b306c74467 filtering: a per-node 'expose' setting overrides a listen's source (ADR 0051)
listens.from was a manifest constant — one value for every node a module runs
on. Now a per-node setting overrides it: {"expose": {"5432": "anywhere"}} makes
postgres public on the machine it is set for while it stays from:mesh elsewhere,
and the firewall (ADR 0050) is computed from the effective source. Exposure()
validates it — a port the module does not listen on, or a source that is not
mesh/anywhere/machine, is refused rather than reaching nothing; UnusedSettings
knows 'expose' is a real destination. Tested: default mesh, setting opens it to
anywhere, bad settings refused.
2026-09-04 21:12:33 +02:00
jschoubben 931ca6f01e cli: module issue seals the node and module into the credential
So the runtime knows the identity the account was scoped to, without a
manifest naming the node.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 01:53:23 +02:00
jschoubben 64496f0335 broker: the substrate pre-declares a consumer's dead-lettered queue (ADR 0048)
LavinMQ refuses a non-administrator declaring a queue with a dead-letter
exchange, so a scoped module cannot make its own. EnsureModuleQueue declares
<node>.<module>.events with its DLX as the mesh, and 'module issue' does so
for a consuming module — the runtime then passively checks it rather than
declaring. Verified against a real broker: the scoped account binds and
consumes the pre-declared queue.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 01:51:17 +02:00
jschoubben f41d280e66 cli: module issue — deliver a module its scoped broker account (ADR 0048)
'module issue <module> --node <m>' looks up the module's emits/consumes from
the catalogue, ensures the bus exchanges exist, creates its scoped account
(CreateModuleAccount), and seals an amqps {url,fingerprint} to the node as the
module's broker own-secret — the same delivery as 'builder issue', now generic.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 01:33:03 +02:00
jschoubben 47bbb0cca6 broker: a generic module account, scoped by emits and consumes (ADR 0048)
CreateModuleAccount gives an assigned module its own broker account whose
permissions ARE its manifest: declare and read its own <node>.<module>.events
queue, read the events exchange to bind onto if it consumes, write the events
exchange only if it emits. The account name carries the node (sealed per
machine), the permissions carry the module (one cannot read another's queue).
The builder becomes one instance of this rule rather than a separate kind.

EnsureEventExchanges declares the bus the substrate owns — mesh.events,
mesh.rpc, mesh.events.dead + a retention queue — idempotently, since a module
account may not declare an exchange.

Scope tested as patterns (no broker needed), and every management call verified
against a real LavinMQ. Honest limit recorded in the code: LavinMQ has no topic
permissions, so ADR 0047's emit-origin reservation (module.<self>.*) is stamped
by the sdk, not enforced by the broker; a pure consumer like the audit logger
is unaffected.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 01:31:19 +02:00
jschoubben 2c06c0d663 manifest: emits and consumes — the event relationship
A module declares the event types it emits and the patterns it consumes,
parallel to provides/requires (novox/hq ADR 0046). Events span module.*,
mesh.* and node.* sources; a consumes for an event nothing emits is a
dangling edge. Fields only here; the dangling-edge check and the runtime
wiring follow.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-03 23:44:30 +02:00
jschoubben 68a9235792 The mount gate knows a facility from a directory
Portainer mounts the container runtime's socket, and the catalogue's
mount gate refused it — rightly by its own lights, since nothing in the
manifest distinguishes a machine facility from the module's data. That
distinction is 04-ISSUES/026's open question, so the gate now carries
the one facility the catalogue mounts as a named exception beside the
citation, one line per facility, never a pattern.

Also the confession: the previous commit landed with this gate red,
because a pipeline's tail swallowed go test's exit code. The gate was
right and the process around it briefly was not.
2026-09-02 02:09:10 +02:00
jschoubben d278edabe0 Three more: nodered, icecast, portainer
Node-RED and Icecast are the plain shapes — a data directory owned by
the number inside, generated passwords in a host-written env file, one
declared port each.

Portainer mounts the container runtime's socket, which is the mount
04-ISSUES/026 is reopened about: a machine facility, not the module's
data, spelled today exactly like a data directory. It is converted as
it runs now rather than held hostage to that vocabulary — the same
mount the builder already carries — and it will be the second citation
when 026 gets its answer.

Letta stays unconverted for now: the arrangement being replaced pins a
year-old image of a fast-moving project, and converting a pin nobody
would keep is not fidelity. It wants a fresh look at what version to
run, which is a decision and not a translation.

All pinned by real digests, resolved on this workstation today.
2026-09-02 02:08:33 +02:00
jschoubben 1c4e10e0d8 The cache keeps no ACL file, and the watch keeps the cache
The lab's diagnostics said it in one line: AUTH called without any
password configured for the default user. With an aclfile configured,
redis takes the default user from the file and quietly ignores
requirepass — so the empty seed this module shipped left the store
without any password at all, politely refusing the credential the mesh
had sealed for it.

And the file could never have worked here anyway: it was host-declared
content, which the host reconciles, so every re-apply would have wiped
what ACL SAVE wrote — a fight between two reconcilers with the tenants
as the ball.

So no file. requirepass alone does what it says, and durability moves
to the watch, which now checks the store and not only its inputs: a
restarted store comes back empty and is re-granted within a tick,
because reconciling is against reality, not against a diff of
instructions. A run that failed leaves last empty, so the next tick
retries instead of believing the inputs were handled.
2026-09-02 01:23:38 +02:00
jschoubben 1b63e21c0f Caught up is an equality, not an ordering
The report carries the digest of the declaration it applied (mesh-host
8211d8b), and the mesh stores it beside the outcome. `reported` rows in
the status JSON now say `current`: whether the machine's last word
names the declaration last sent.

Not derivable from the timestamps beside it, which is why they were
not enough: an apply begun under the previous declaration reports
after the next send — newer, and still about the old words. The lab
lost exactly that race between one test's closing push and the next
test's opening one.

Empty digests — every host from before reports carried one — read as
not current, which errs toward waiting rather than toward asserting on
files that are not there yet.
2026-09-02 00:02:44 +02:00
jschoubben 649ce9bc3b Directories belong to the number that runs inside
The lab named both failures in one run: the store restarted forever on
a conf file it could not read, and the forge could not traverse into
the directory that held its files. Both are the same fault — a file the
mesh declares root-owned, consumed by a container process that dropped
to a uid the machine has never heard of.

The forge's data now belongs to 1000, the user its container runs as.
The store's conf, ACL file and data belong to 999, which is what redis
becomes after its entrypoint drops privileges. The package registry's
conf directory belongs to 10001, which writes htpasswd into it.

Made expressible by the host in the commit beside this one: an owner
may be numeric, because a container's user has no name on the machine.
2026-09-01 23:45:17 +02:00
jschoubben d85be55e3c The media tail: bazarr, jackett, nzbget, tautulli, ombi
The same shape as the four before them — a config directory of their
own, the shared library paths they actually touch, one owner and mode
across every sharer — which is the point of doing them together: five
manifests that differ only in name, port and which shelves they read
prove the shape is a shape and not a coincidence.

All pinned by real digests, resolved on this workstation today. The
catalogue stands at 27.
2026-09-01 23:26:29 +02:00
jschoubben 1dce1fc5e7 The media four: plex, sonarr, radarr, qbittorrent
The shape they share is the interesting part: a media library is a fact
about the machine that several modules mount, so each declares exactly
the paths it touches, with one owner and mode across all of them. The
host reconciles a directory rather than creating it, so the second
module ensuring a path the first already ensured is maintenance, not a
fight — and ADR 0030 keeps a shared path alive when one of its sharers
is unassigned, because it holds what the mesh did not put there.

Plex runs on the machine's own network like home-assistant, and for the
same reason: discovery is the point. The other three publish, so the
mesh chooses their machine ports.

All pinned by real digests, resolved on this workstation today. The
catalogue stands at 22, of which the arrangement being replaced ran 95
— but its number counts workstation ricing and one-machine tooling
beside services, and only the services convert to manifests like these.
2026-09-01 23:23:43 +02:00
jschoubben 0ba52d03a3 Two more: nextcloud and home-assistant
Nextcloud is the first consumer of two provisions at once: a database
from the mesh's postgres and primary storage in a bucket from the
mesh's own object store — which is how the arrangement being replaced
ran it, minus the bundled MariaDB it no longer needs. Every credential
in one env file the host writes; the manifest holds placeholders and
the mesh holds nothing readable.

Home automation runs on the machine's own network, because discovering
devices is the point and a bridge would hide them — so nothing is
published, the declared port is the bound port, and the rule set opens
exactly it. The case MachineSide was corrected for, in the catalogue.

Both pinned by real digests, resolved on this workstation today.
2026-09-01 23:09:40 +02:00
jschoubben c0107b8572 Status says when each machine last reported, beside when it was sent
"Not waiting" says the declaration is current, not that the machine
finished applying it: the sent digest is recorded at send. So a test
that pushed, saw waiting clear, and asked the machine what it was
running found containers that did not exist yet — the certificate fix
made compositions stable, and the settling that used to fail first had
been hiding the gap behind it.

The mesh already held the missing half: every machine's last report,
with its time. It just was not in the JSON. `reported` now sets each
machine's last word beside when the current declaration went to it, and
"has it caught up" becomes a comparison of two timestamps the mesh
recorded itself — a report newer than the send means the machine acted
on what was sent; older means it is still working, which waiting alone
cannot distinguish.
2026-09-01 23:02:26 +02:00
jschoubben ef7750f816 Three more: searxng, influxdb, verdaccio
Converted from the arrangement being replaced. The search engine brings
its own valkey on its own network — a sidecar is just a second
container resource. The time-series database initialises itself from
two generated secrets, and its data and config directories are declared
with the owner the image runs as. The package registry's configuration
is a declared file rather than a merged one, which is the position
16-module-coverage takes on config merging: the module knows its own
format because it wrote the rest of the file.

Two were read and deliberately not converted, which is worth recording
where the next person will look:

n8n builds a custom image, so it is a module with a repository rather
than a manifest in this catalogue — where a module's own code lives is
ADR 0037's question, and pretending otherwise here would prejudge it.

mosquitto authenticates from a hashed password file that only
mosquitto_passwd can write, and the mesh delivers plaintext sealed
files — so an honest conversion needs a small provisioner, the same
shape as the cache's. Without one, the manifest would compose a broker
nobody can log in to, which is exactly the kind of module that parses,
resolves, and stops on the machine.

All images pinned by real digests, resolved on this workstation today.
2026-09-01 22:49:18 +02:00
jschoubben 8c4a478b09 Four more of the catalogue: registry, redis, umami, grafana
Converted from the arrangement being replaced, in its shapes rather
than theirs.

The registry is the manifest the lab already proved, promoted: names
its image by digest and is never built (04-ISSUES/029), provides the
artifact store, claims it once per machine.

Redis is the third provision after a database and a bucket, and the
first whose tenancy is a pattern in a shared keyspace rather than a
namespace something else enforces. Its provisioner mirrors the postgres
one's contract line for line — the manifest, the sealed per-consumer
files, the mark, the withdrawal of orphans — and speaks RESP directly:
five commands are needed, and a client library large enough to hide
them would be most of the program's size. A prefix Redis would read as
a pattern is refused, because the grant must mean what the manifest
said; grants are persisted with ACL SAVE, or said loudly, because a
cache that forgets its tenants on restart reports success until then.

Umami asks the mesh for its database and a generated app secret, and
carries no state of its own — the arrangement being replaced ran a
bundled second postgres beside it. Grafana keeps its dashboards in a
declared directory with the image's own owner. Both listen on 3000, as
does the forge — which is the mesh's port assignment earning its keep.

Traefik is deliberately not converted: the mesh's route provider is
mesh-route-proxy, which speaks route grants natively, and a traefik
that consumed them would be an adapter nobody has written pretending
to be a conversion.

All images pinned by real digests, resolved on this workstation today.
2026-09-01 22:44:53 +02:00
jschoubben 38d4e77cec A certificate is issued once and kept
Every signing carries a fresh random serial, so a mesh that signed per
composition composed a different declaration every time it was asked
what a machine should be. Every machine carrying a certificate then
stood eternally "waiting" — pushed seconds ago and already behind — and
the forge test, the first to wait for settledness on such a machine,
failed four runs in a row wearing three other faults' clothes.

Found live on a kept mesh, which is what settled it: two plans seconds
apart, identical to the byte but for one serial, in the certificate
file. Deduction had four theories; the diff had one line.

The keeping columns had existed since the serving key's migration —
"and what was issued for it" — and were written by nothing, the same
shape ReleasePorts was found in this morning.

Kept beside the serving key it certifies, and it stands while the name,
the key and the clock agree: a node rejoining with a new key or renamed
gets a fresh signing, exactly as if nothing were kept, and so does one
whose certificate is into its last stretch of life. The port's rule and
the secret's, applied to the third thing composed fresh each time.
2026-09-01 22:26:50 +02:00
jschoubben b70f0d626a What review found in the port machinery, fixed
Three faults, one file split. All from reading, all verified to bite.

Unassign now releases the module's ports. ReleasePorts existed, said
"for when it is unassigned" in its own comment, and was called by
nothing — so a fixed port stayed claimed in the name of a module that
was gone, and the next module needing it was refused by a ghost.
Kept-once-chosen is a promise about a module that is still here.

MachineSide reads addressed mappings. "127.0.0.1:8080:80" was split at
the first colon, "127.0.0.1" failed to parse as a port, and the mapping
was silently skipped — putting the filter back on the declared port,
the exact fault the function was written to end. The machine side is
the second-from-last part, which is the reading the host already
applies, and the substrate bundle writes that shape today.

An allocation race answers in the mesh's words. Two concurrent picks of
the same port used to surface as a Postgres constraint violation,
verbatim. The table has two keys, so the collision is one of two facts:
the racer was this same assignment — then its answer is the answer,
kept-once-chosen does not care who chose — or another module took the
machine port, and an unfixed pick is simply made again against the
moved free list. A fixed port that lost the race is refused by name.
Told apart by re-reading the row, not by the constraint's name, so this
does not couple to the migration's spelling.

And the artifact-store cycle tests moved to bootstrap_cycle_test.go;
machineside_test.go had quietly become three subjects.
2026-09-01 21:54:05 +02:00
jschoubben 8174f5c41e plan is the send without the sending, so it allocates
Confining allocation to the push path took `plan` with it, and `plan`
belongs on the other side: it is a person asking what a push would do to
one named machine, so the port it shows and the secret it seals must be
the ones a push would use. Both are kept once chosen, so showing
numbers a later push would replace answers a question nobody asked.

Caught by the lab: composing the real modules stopped producing
postgres's sealed superuser, because nothing had minted it and the
read-only path correctly declined to.

The line is not question versus command. It is a person asking once
about one machine, against the mesh asking continuously about all of
them — the second is what hung, and the second is what reads.
2026-09-01 21:32:57 +02:00
jschoubben be62f49eab What provides the artifact store cannot be delivered through it
Refused where it is written: a module that provides `artifact-store` and
also builds artifacts asks the mesh to put an artifact into the thing
that artifact is needed to create. Building publishes to the store, and
the builder will not start without one — "a built artifact nobody can
fetch is not built".

This is the question the substrate record asks of every candidate: can
it grant itself the thing it provides? The store cannot create its own
database, the broker cannot create its own virtual host, and a registry
cannot grant itself a repository. The first two are why they are in the
bundle. This is the same sentence, unenforced.

So such a module names its image, exactly as the bundle names the three
a first node starts from. One that wants an interface or a tool server
beside it is a second module, mirrored the ordinary way once the first
is running — a real limit, and better said here than discovered on a
mesh new enough that nobody is watching it.

Refused at the manifest because the alternative is a build that never
returns.

The provision name is now a constant. A string compared in one place is
a convention; a string a rule turns on is a fact.
2026-09-01 21:18:14 +02:00
jschoubben 4d6ec5b10c One object store, not two
object-store.json and minio.json described the same thing: same image,
same provision at the same scope, same provisioner. Not two
implementations a person could choose between — one module written
twice. Assigning both to a node would have collided on `s3-bucket`.

It exists because it was written first, to pair with photos.json for the
README's worked edge, and minio.json was the fuller version of the same
module written later. Nobody removed the first.

The pair test keeps its point and now reads the surviving one. Checked
across the rest: this was the only duplicate.
2026-09-01 21:06:17 +02:00
jschoubben 0d975a051d Asking what the mesh would send must not change it
`status` hung. It composes a declaration for every node to answer *is
this machine running what I would send it*, and composing one assigns
each module a machine port — so the question wrote to the database, and
wrote to the same rows as the machine it was asking about.

`port_assignment` is unique on (node, machine). Two transactions
inserting the same port do not race, they queue: the second waits on the
index until the first commits. A status polled every two seconds while a
node applies is two writers on those rows, and the poll stopped
returning rather than returning something wrong — which is the better
failure of the two, and still a failure.

The latent version of this was there before anything polled: two
compositions running at once could both allocate.

So allocation belongs to the send path alone. The mesh chooses a port
when it commits to sending one; every other caller reads what was
chosen. A module with nothing assigned has never been sent, which is
precisely what "waiting" means — the read needs no number to be right
about that, and inventing one would make the answer worse.

Named rather than passed as a bare bool: at three call sites, `true` and
`false` say nothing about which of these two things is meant.

Checked by the lab, which now polls status throughout an apply.
2026-09-01 20:05:08 +02:00
jschoubben 83c6a2f244 Withdraw the mount check: it refuses the builder
The rule was right about data and wrong about everything else. The
builder mounts the container runtime's socket, which is not its data,
does not belong to it, and must not be declared as one of its
directories — and the check refused the builder's own manifest.

Caught by the lab, though not honestly: the run was already going when
this went in, so the builder binary was rebuilt mid-run with the check
compiled into it and the failure was mine, not the mesh's. Confirmed
against the manifest directly rather than inferred from the log.

What it was protecting is real and stands — the fourteen mounts are all
declared. But enforcing it needs a way to tell "the directory my data
lives in" from "a machine facility I was granted", and the mesh has no
vocabulary for the second. `capabilities` is the closest thing and does
not name paths. That is a design decision, so it goes back to
04-ISSUES/026 rather than being invented here to make a check pass.

The check that every real manifest still parses is kept. It costs
nothing and it is how the next attempt at this finds out sooner.
2026-09-01 19:45:36 +02:00
jschoubben 53eb000a84 A container may not mount a path the module never declared
Closes the half of 04-ISSUES/026 that would otherwise come back. The
fourteen mounts across the forge, the mail system, the store and the
object store are all declared now — but nothing said they had to be, so
they were right by coincidence and the next volume added would not be.

A bind mount whose source does not exist is created by the container
runtime, as root, with a mode it picks. So `owner` and `mode` — which
exist precisely so a module can say who its data belongs to — were
silently not applied to the only directories holding data.

And the rule written for exactly this case did not reach them. A
directory the mesh declared and no longer wants is kept, not removed,
when it holds anything the mesh did not put there (ADR 0030). That is
the answer to *what happens to my data when a module goes away*, and it
is written in terms of declared directories: an undeclared one sits
outside it, because the mesh does not know it is there.

Refused where it is written rather than on the machine, which cannot
tell the difference — by the time the host sees the mount it is being
asked to make a directory, which it is perfectly able to do. The fault
is in the manifest, so it is named at the manifest. Same argument as the
action refusal directly above it.

A path under a declared directory counts as declared, as do the files a
module already names: its own secrets, its grants, what it receives.

Every real manifest is checked to still parse, and the refusal bites.
2026-09-01 19:36:01 +02:00
jschoubben c67f836185 The mesh may only move a port it actually publishes
The lab caught this: a module declaring a port and running no container
had its rule set opened on 20000 while its service sat on 9101. The
firewall reported success and blocked the thing it was told to admit,
which is the precise failure the filtering comment warns about, arrived
at from the other side.

Assignment was applied to every declared port. But a container's mapping
is the thing that translates, and where there is none the software binds
what it binds — the mesh choosing a number does not move the service, it
only makes the mesh wrong about where it is.

The declaration side already knew this: publishedOn rewrites container
ports and nothing else. Filtering did not, so the two disagreed about
the same fact. MachineSide is now the one derivation both follow.

It also fixes a second case nobody had hit yet: a mapping the manifest
wrote itself, like the mail system's 7080:80. That is passed through
untouched when composing, so assigning it a machine port would have
opened a rule on a port the container does not publish. Either side of
such a mapping now names it, and the host side is the answer — a module
may read `listens` as what its software binds or as what the machine
exposes, and both readings want the same number.

Recorded either way, assigned or not: the map means where this module's
port is on this machine, and every reader needs that answer regardless
of who chose it.

Tests bite — making it always assignable reproduces the lab failure.
2026-09-01 19:31:05 +02:00
jschoubben 41f7c51032 Assign around what a machine already holds
The other half of ADR 0038, and what 04-ISSUES/028 was actually about.
A module can now avoid colliding with another module; until this it
could not avoid colliding with the mesh itself.

The substrate is not a module. A node raises it from the bundle it
carries before any mesh exists, so the control plane had never heard of
the store, the broker, or its own container — and handed a database
module 5432, which the store already had.

So the machine says. The host records what each resource binds,
distinguishing what it carried from what the mesh sent — a distinction
that already existed so the two never remove each other — and reports
the carried ones. The node states and this context writes, which is the
shape of every message between them.

What the declaration binds, not what is open. A machine's open ports are
a moving target, and assigning around them would mean a port that was
free when it was asked for and taken when it was used.

Replaced whole each time rather than merged: a machine that gave a port
back must be believed about that too, and a set that only grows keeps a
port reserved for something no longer there.

Tested against a real database, and the tests bite — removing the check
hands the module 20000, which the machine had said it holds.
2026-09-01 18:32:38 +02:00
jschoubben 1f5b70a995 The mesh assigns the port, and a module says it once
novox/hq ADR 0038. A module cannot choose a port: it is written once and
assigned anywhere, so any number it picks is a guess about a machine it
has never seen. A database module met the mesh's own store on 5432 and
was told, by a container runtime three layers down, that the port was
already allocated.

The number used to appear three times in every module — the rule set,
what a consumer is told, and what the runtime publishes — agreeing only
because one person wrote all three. Now it appears once, in `listens`,
and the other two are derived: the container publishes `20000:5432`, the
consumer is told 20000, and the rule set opens 20000.

An assignment is made once and kept, as a credential is. A port that
moved on every declaration would restart both ends each time and hand a
consumer a number that was true when it was read.

Ports the protocol fixes — mail on 25, submission on 587, DNS on 53 —
say so, and are then claims: one holder per machine, and the second is
refused by name at assignment. That is the mechanism the mesh already
has for what is singular on a machine, pointed at ports.

A mapping written the long way is left exactly as it is. Some things
must be pinned by hand, and quietly overruling somebody who wrote both
halves would be worse than not offering the short form.

Still open, and known: the substrate is not a module, so the mesh has
never heard of its own store and cannot yet assign around it. That is
what 028 will still be about after this.
2026-09-01 17:52:53 +02:00
jschoubben a5d85266d0 A container does not take restart-on, and nine of them did
This is what stopped the forge. The host refused the whole declaration:

  resource "postgres.server": a container does not use "restart-on",
  and it is set. Refused rather than ignored

`restart-on` belongs to a service. I put it on containers this morning
so one would pick up a rotated credential — nine times across seven
modules — and nothing between the manifest and the machine said a word.
The control plane composed it happily; the parser accepted it; the
manifest tests passed. The only thing that knew was the host, five steps
downstream, and hearing from it cost a seventeen-minute run.

The host was right twice over. It refused, and it refused *everything*,
because applying the parts it understood would leave a machine that
looks configured and is not. One misplaced key therefore stops a module
dead, which is the correct severity and an argument for catching it
where it is written.

So the shapes and their keys are now written down here and checked. They
are duplicated from another repository deliberately — this is its wire
format, like the shape of a grant file — and a contract with two copies
and no check is a contract until somebody edits one.

What this does not fix is why I reached for it: a container cannot
follow a file. Filed separately.
2026-09-01 17:19:04 +02:00
jschoubben f5b03e1474 Declare the directories that hold the data
novox/hq 04-ISSUES/026. Four modules mounted fourteen host paths that no
resource declared — the mail spool, the databases, the object store's
data. Each would be created by the container runtime as root, with a
mode nobody chose, so `owner` and `mode` went unapplied on exactly the
directories that matter.

The worse half: a directory the mesh declared and no longer wants is
kept rather than removed when it holds anything the mesh did not put
there. That rule is the answer to what happens to data when a module
goes away, and it is written in terms of declared directories. An
undeclared one is not covered. So the one rule guarding against data
loss reached the configuration directories, which are cheap to lose, and
missed the data directories, which are why the rule exists.

The cause is worth naming. These manifests were written by reading the
arrangement being replaced and carrying its compose files across —
service, image, ports, volumes, environment. The container shape can
express all of that, which is what made the transliteration feel like
progress. A shape that can express a compose file gets filled in like
one, and a volume line borrowed from compose declares no owner, no mode
and no intent.

Declared parent-first, because the host applies in the order written and
does not sort. The check is mechanical now, because a person comparing
volumes against directories by hand is the process that produced this.

Still open, and bigger: whether these paths are where a module's data
should live at all. They were inherited whole, and they decide what a
person backs up.
2026-09-01 16:09:23 +02:00
jschoubben ee3cc1b6f4 Pin the example modules to images that exist
novox/hq 04-ISSUES/025. Every image reference in every example module
was sixty-four zeros — eighteen of them across five modules. Each
parsed, resolved, and composed into a declaration a host accepts, and
none could ever have started: the machine reaches `docker pull` and
stops. That is why those modules were written and not running, and no
check saw it because every check passed.

The host validates the shape of a reference and nothing more, which is
correct: verifying a digest exists means reaching a registry, and that
is the one thing a host must never have to do. So the last place that
could catch this is the wrong place to try.

The guard therefore sits where a declaration is composed, not where a
manifest is parsed. A file in a repository is allowed to await a pin —
the design already says the manifest in a repository names artifacts
while the manifest the mesh holds names digests, and the bundle works
exactly that way. What must never happen is a placeholder reaching a
machine, and composing is the last moment before one does.

Twelve third-party images resolved to real digests without pulling
anything, which is also the mechanism the open issue needs. Two
discoveries came free: mailu publishes to ghcr rather than Docker Hub,
so seven references named repositories that do not exist at all; and it
renamed roundcube to webmail, so that one would have failed even with
the right registry.

What stays a placeholder is the mesh's own provisioner images, which
genuinely have no digest until built and pushed — the bundle's problem,
legitimately unresolved here. The stand-in consumer now stands in with
a real image rather than an invented one.
2026-09-01 15:13:33 +02:00
jschoubben 2835f41a64 Stop committing a 12 MB binary I added by accident today
The lab is pointed at a path for the builder it should write, and I
pointed it at the repository root instead of build/, which .gitignore
already covers. Two of today's commits carry the compiled binary as a
result.

Untracked and ignored by name, so the same slip does not land it again.
2026-09-01 03:21:45 +02:00
jschoubben e49586646b A comment claimed a test that does not exist
I wrote that a test asserts the control plane's placeholder expression
and the host's still agree. None does, and none in this repository could
— a unit test here can only assert what this repository already
believes.

That is precisely the thing this project refuses to tolerate: a stated
rule with no way to check it, which costs more than no rule because
people believe it. Written by me, today, in the same file that closes a
gap of the same kind.

What actually proves it is the lab, and the comment now says so.
2026-09-01 03:15:22 +02:00
jschoubben a4090014f3 An example may not name an image nothing builds
Found by reading the manifests rather than by running them. Two of the
provisioner images the examples name had no way to be produced: the
object store's had a Dockerfile and no target, and Keycloak's did not
exist at all — no image, no Dockerfile, no program.

A module naming an image nothing produces resolves, plans, pushes and
stops on the machine at `docker pull`, which is the fault arriving as
far from its cause as it can get.

The object store's target is added. Keycloak's provisioner is removed
from its manifest, because writing a manifest for a program that does
not exist is the same mistake as the .env files: it parses, it resolves,
and it could never work.

That makes keycloak's manifest true about today — a server the mesh
runs, with its database and its admin credential — and it makes the gap
loud. Keycloak no longer claims to provide oidc-client, so a consumer
asking for one is refused at plan time by name, rather than resolving
cleanly and never having a client created.

The check covers only images beginning `mesh-`. Postgres and the rest
come from a registry and are somebody else's to build; what this bounds
is the set this repository is responsible for and might forget.
2026-09-01 03:12:49 +02:00
jschoubben e5243cd753 Ask every consumer for usable configuration, not just the one in hand
The keycloak check was written while keycloak was the module being
worked on, which is how a check ends up proving one thing about one
file. It now runs over every example that requires something, and asks
the two questions that matter for all of them: that no ${bound:...}
reached the machine as a value, and that anything named PASSWORD is
still a hole only the host can fill.

The first is the one worth having. A placeholder written through is read
as a value by whatever parses the file — a connection to a host called
"${bound:postgres-database:at}" — and the failure names neither the
module nor the mesh.

Modules whose requirements nothing in the examples answers are logged
and passed over, because that is a fact about the example set rather
than about them.
2026-09-01 03:10:43 +02:00
jschoubben 71f77617e3 Omit a consumer's identity where there is none
A contribution that is not a credential grant — a module offering
something to another on its own machine — has nobody to be identified
to, and was carrying an empty `as`. A field that is always present and
usually empty teaches a reader to ignore it, including when it is not.
2026-09-01 03:09:18 +02:00
jschoubben 122680b554 A consumer can write its own connection string
novox/hq 04-ISSUES/023. A consumer was given its password, the address,
the port and where its credential lives, and still could not connect —
the user name was invented by the provisioner and recorded nowhere, and
the rest sat in a JSON binding that a program reading KEY=value cannot
use.

Both halves have the same cause: the mesh knew something and did not say
it.

**Who a consumer is, said once.** The provisioner used to derive
mesh_<node>_<module> and that string existed nowhere else — not in the
control plane, not in the binding, and above all not at the consumer,
which has to present it. Now the mesh derives it once and sends it to
both ends, so they agree by construction rather than by two conventions
that were the same on the day they were written. The provisioners refuse
to invent one if the mesh says nothing, because falling back to a name
of their own would create a role the consumer would never guess and
everything would report success.

**Bound values reach the file that needs them.** ${bound:provision:key}
is the symmetric twin of the sealed placeholder, and simpler: these
values are not secret, so the control plane fills them in before sending
and the host gains no field and learns no format. It stays
name-agnostic — at, as and from are true of any provision, and every
other key comes from what the provider said it serves.

The asymmetry it removes was backwards. The secret is the hard case,
because the mesh must not be able to read it, and the secret was the
part that already arrived.

Keycloak and Gitea now produce complete connections, asserted from the
manifests on disk rather than from fixtures: every part filled, no
placeholder surviving as a value, and the password still a hole only the
host can close. Three faults injected, each caught.
2026-09-01 03:03:07 +02:00
jschoubben 96f90ab986 Refuse an action where it was written, not on the machine
A module may not declare an action: the link may not carry a command to
run, and that bound is what limits a compromised control plane to shapes
it cannot turn into arbitrary code (novox/hq ADR 0005). The host
enforces it, correctly and in the right place.

But a module's resources reach a machine over the link, so a manifest
carrying an action was accepted here, stored, resolved, planned and
pushed — and refused on the machine, in the host's log, with nothing
connecting it back to the manifest that caused it.

The rule held. It was just unusable, which is the same shape as the
network shape earlier today: the refusal was right, arrived far from its
cause, and nobody was reading the log.

The refusal names the rule and what to do instead, because "you may not"
with no alternative is where a module author stops.

Found while checking a claim I had written in the coverage document —
that a module cannot declare one. It could; it just could not deliver
it. The document is corrected.
2026-09-01 02:55:56 +02:00
jschoubben be2dca27ab The example modules put their credentials where the programs read them
Every one of these declared `own-secrets` pointing at a path called
`.env` and then mounted it as `env-file`. The file's whole content is
the password. Docker reads that as a malformed line and the container
starts with no password set — which is not a failure to start, it is a
service running with the wrong credential.

They parsed, they resolved, and none of them could ever have worked.
That is what a manifest checked only by the parser buys.

Each now keeps the sealed file as what it is — a password, alone — and
declares a file beside it whose content says ${secret:name}. The host
fills the hole on the machine, which is the only place both halves
exist. The provisioners mount the bare file, because they read a
password file and always did.

Two tests, both driven from the manifests on disk rather than from
fixtures: every ${secret:x} must name something the module declared, and
nothing may read a bare password file as an env file. Injecting the
shipped bug reproduces it word for word.

Keycloak, Gitea and Mailu still cannot connect to their databases, for
the reason in 04-ISSUES/023 — the user name is the provisioner's
invention and the bound values cannot reach a config file. Their own
credentials are right now; that half was independent and is done.
2026-09-01 02:52:00 +02:00
jschoubben 0af3ea1acf A consumer is a module on a machine, not a machine
novox/hq 04-ISSUES/022. A credential was keyed by provision, consumer
node and provider node, so "who is asking" was answered by naming a
host. The node this mesh exists to take over runs eight modules against
one database server.

The symptom had two halves and only one was loud. The provider refused,
naming the modules and explaining they would share one credential, which
reads as a decision rather than a limit. The consumer did not refuse: it
resolved cleanly, wrote one module's credential file and left the others
absent — a service that starts and cannot authenticate, with nothing
saying why. That is 021 again on a different axis.

Three modules wanting one database produced one need, carrying whichever
module mentioned it first, because the resolution walk is a work-list
over names. The fan-out now happens in one place, after the walk. The
record path already did this correctly and said why: a consumer here is
a module on a machine. It is the same rule.

Downstream: the secret's key gains the consuming module, the grant file
is named after both halves, needs are matched by provision and module
rather than provision alone, and the provisioners name the role and the
access key after the module. The refusal in ContributionsTo is gone
because there is nothing left to refuse.

Worth stating plainly: without that refusal, gitea's login would have
opened keycloak's database. From the provisioner's side it created
exactly what it was asked to create.

Existing secrets are discarded rather than backfilled. They cannot say
which module they were for, and a secret is remade and delivered to both
ends on the next push — so this costs one rotation and invents nothing.

Also guards the role name against PostgreSQL's 63-byte truncation, which
is a notice rather than an error and would reintroduce exactly this
collision at a length nobody tests.

Three faults injected — the fan-out removed, needs matched by name
alone, the grant file named after the machine — each caught.
2026-09-01 02:40:09 +02:00
jschoubben 1314be5282 A file may hold a credential where its content says one belongs
The gap that stopped keycloak and gitea from starting. A granted
credential arrives as a file whose entire content is the password, which
is what a program reading a password file wants — and most programs do
not read one. They read KEY=value, or a JSON document with the token at
an attribute inside it. A module in that position could be handed the
bare value or nothing, and both are useless.

The host has been able to do this all along: content with ${secret:name}
in it, sealed values beside it, substitution on the machine, which is
the only place both halves exist. Nothing filled the values in, so the
hole could be written and never closed and the host refused the file.
That refusal was correct and the feature was unreachable.

A module reaches its own secrets and the credentials it was granted —
both things it wrote in its own manifest — and nothing else. Naming
another module's is refused: two modules on one machine are as separate
as two on different machines, and letting one read the other's
credential by guessing a name would end that to save writing a file.

Filling runs after settings, which is the whole reason it sits where it
does. A setting is how a placeholder gets into a JSON document in the
first place — the desktop client that reads its token from an attribute,
not an environment variable. Before the merge that file's content is
"{}" and asks for nothing.

Tested through Declaration rather than through the helper. Three times
in this repository a test asserted on a helper while the code calling it
was wrong, and each time the injected fault stayed silent. Three faults
injected here — the call removed, the call moved before settings, and
the module boundary widened — each caught by the test meant for it.
2026-09-01 02:28:50 +02:00
jschoubben df62bb57e5 A requirement answered on this machine is still a requirement
novox/hq 04-ISSUES/021. Two modules where one provided what the other
required, on one node, resolved cleanly with zero needs: no credential
was made, the consumer's secret file was never written, and whatever
read it would fail somewhere else entirely. Nothing was refused and
nothing was reported.

The world a node resolves against is every OTHER node, so a provider on
the same machine never became a Needed, and the credential loop walks
Needs. Every step reasonable, the sum a silent gap.

It survived because everything proven until now was cross-machine —
the interesting case for a mesh and the rare one in practice. The first
module to want a database on its own machine was the first real one.

The assumption underneath was that a local consumer needs no credential,
which holds for a process reaching a unix socket where the system can
vouch for the caller. It does not hold for containers, which is how
nearly everything here runs: the consumer reaches the provider over TCP
from its own container and the database asks for a password exactly as
it would from another machine. **The machine stops being a trust
boundary once both ends are containers.**

A brokered provision answered here is now a need naming this node, and
carries what the provider serves — which a local provider never
contributes through the world. A name nothing grants is unchanged: a
shell answered here is answered, and nothing more is owed. Both
directions tested, both injections bite.
2026-09-01 02:15:26 +02:00
jschoubben d0511ee3fe The command API, which refuses everything until it knows who is asking
novox/hq ADR 0035: one implementation, several surfaces, and a surface
holds no decisions. The act of assigning — including that an assignment
which does not resolve is kept and still refused — moved into acts.go,
and the command line now calls it too. Two surfaces, one refusal, in the
same words.

It will not run without --issuer, and refuses at start rather than per
request so it is found by whoever ran it rather than by whoever finds
it. There is no flag that removes the check.

The authenticator is honest about what it is: no token can be verified
until an identity provider exists, because that is a module and none is
running, so every request is refused and told that the command line
still works. A surface that functioned without authentication would be
one somebody left running — and the board this stands behind is
published on a public name.

Four refusals, four tests. The last one first asserted "not 200", which
passed because a request with no database fails at the store anyway — it
proved nothing about whether the input was checked. It now asserts the
specific refusal, and bites when the check is removed.
2026-09-01 01:55:56 +02:00
jschoubben 8cf1ecf6a5 secret accept: see the flag that comes after the arguments
The command refused every real invocation. Go's flag package stops
parsing at the first non-flag argument, so with the positionals first —
the order that reads correctly — `--from -` stayed among them and the
count check rejected it.

The host's own parser carries a note about this exact fault, and the
version it describes is worse: there a flag somebody passed was silently
ignored and the command succeeded anyway. This one at least refused.

The tests did not catch it because every case in them was a rejection.
The command was broken in the only way that matters — it refused what it
is for — and the suite was green. The lab found it at the first call.

Two tests now: the helper, and the command itself with a --from naming a
file that is not there, so the complaint must be about the file rather
than about usage. The second exists because injecting against the first
stayed silent: testing the helper alone left the command free to ignore
it entirely.
2026-08-31 23:05:10 +02:00
jschoubben 7e9c28fd9e secret accept — carry a value the mesh did not make
The entry point for adopting something already running, and the half
that was missing. The store has carried the distinction since the
beginning — a module secret records whether it was `made` or `accepted`,
and refuses to invent a replacement for the second — and
AcceptSecretForModule existed, with exactly one caller: the broker
account issued to a build machine. Nothing else could write one.

Without it every module secret is generated, which against a database
that already exists puts 32 random bytes where a working credential was.
The machine applies it, reports success, and whatever reads it fails to
authenticate somewhere else entirely, with the mesh insisting the secret
was delivered — which it was.

The value is read from a file or from standard input, never from an
argument: a value on the command line is in the shell's history and in
the process list. Same path a model-access key already takes, and no new
dependency — the first version reached for x/term and the existing one
needed nothing.

Sealed on the way in, plaintext discarded, and not printed back. The
only difference from a generated secret is where the value came from.

Two rules with a test each, and the second is the one that would have
been got wrong: only the line ending is removed, never surrounding
space. Trimming both ends is the obvious thing and would deliver a
password chosen with a leading space as a different password, silently.

Both were briefly untested for different reasons — the trimming lived
where no test could reach it, and then a -run filter matched neither
test. Extracted, and injected against the whole suite.
2026-08-31 21:57:57 +02:00
jschoubben f04d00b411 The proxy can obtain a public certificate, and asks staging by default
Work breakdown 1.4. The mesh's own authority certifies internal names
and always did; a name reachable from outside needs one the world
already trusts, and there was no ACME anywhere in this repository.

Uses acme/autocert from x/crypto, which was already a dependency — one
indirect addition (x/net, for idna) and no new direct one.

Three things worth more than the feature:

**Staging is the default** (novox/hq 04-ISSUES/004). Production issuance
is rate-limited per domain and per account and does not replenish
quickly. Defaulting to production would leave the safe path depending on
remembering to opt out, on exactly the work most likely to iterate. A
staging certificate is trusted by no browser, so the mistake announces
itself on the first request rather than a fortnight later.

**A certificate is only asked for on a name the mesh routes here.**
Without that policy, anything that can reach the port and send a name
triggers an order for it — a scan becomes a stream of failed orders
against the account's rate limit, and the proxy looks healthy
throughout. What it may certify is what it was told to serve.

**A private issuer is trusted by naming a file, never by skipping
verification.** Skip would still apply on the day this points at a
public issuer, and nothing would say so.

TLS is opt-in: without TLS_LISTEN the proxy serves plain HTTP exactly as
before, which is what an internal-only mesh wants. With it and no cache,
it refuses rather than defaulting — every restart would otherwise order
new certificates, silently, until the rate limit says it does not.
2026-08-31 19:35:34 +02:00
jschoubben 16faadfe52 A session is a consumer of a licence, and (node, module) already names one
Work breakdown 1.2. Two sessions run on the control-plane node — the
node's own and the mesh's (novox/hq ADR 0026) — so a machine stopped
being a usable answer to "whose licence is this".

14-model-access.md called per-module-per-machine "a step toward it and
not it", and that is true of a worker: many run on one machine from one
module, so the pair cannot name them apart. It is not true of a session.
The two sessions are two modules — the same mechanism started in
different context roots, and a context root is what a module delivers —
so (node, module) tells them apart and nothing needed adding.

Checked rather than argued: different licences on one machine, each with
its own key, and a session on no licence is not handed the other's.

The third test exists because a fault injection stayed silent. The first
two put the sessions on different licences, so the licence alone
disambiguates and the module argument is never load-bearing — removing
it from the query changed nothing and everything still passed. Two
sessions on the SAME licence is the case that needs the pair to be the
identity: releasing one must leave the other, and a machine-shaped
answer takes both.
2026-08-31 18:41:19 +02:00