Commit Graph
366 Commits
Author SHA1 Message Date
jschoubben d032fe6e8d assign: what it checks is the mesh, not the one machine
An assignment was verified by resolving the node it was made on. The verify that matters
is resolution over all of them: a module offering a mesh-scoped provision stops offering
it the moment its own node stops resolving, so an assignment could be reported as fine
while it took that provision away from every consumer elsewhere. Those consumers were
then told "nothing in this mesh provides it", naming as the remedy a module that was
already assigned — a wrong answer about a machine nobody had touched.

That is novox/hq 04-ISSUES/017's shape exactly: an action succeeds into a state its own
verify rejects, and it does so because the action's own test is not the test the verify
uses. 017's remedy was to make them the same test, and this makes them the same test.

The assignment is still kept, and that is the other half of the decision. Assignment is
not an ordering: a consumer assigned before its provider does not resolve for as long as
it takes to assign the provider, and refusing the first half of a pair would make the
order somebody types two commands in part of the mesh's rules. So `assign` and `unassign`
now name every OTHER machine that cannot be worked out as things stand, in the mesh's own
words, beside whatever they already said about this one. It reports the state and never
claims causation — saying "this assignment broke laptop" would mean resolving the whole
mesh twice and would still be a guess about which of several changes did it.

It costs a resolution per machine. Assignment is a person typing a command, and being
told which machines this just blocked is worth more than the milliseconds.

This is what a four-node raise read as "bumping a module's version broke provider
recognition". It was neither the version nor the provider: nothing in this codebase reads
the version column, every lookup is keyed on the module name alone, and a test in the
previous commit now says so. It was one machine's set of assignments, and nothing said so.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:11:33 +02:00
jschoubben f5f860fd2b inventory: forgetting a module says what goes with it, and refuses until told
`module forget` cascaded. The settings, the module's own secrets and the ports the mesh
chose all name the module by a foreign key that cascades, so removing the row took all
three and reported "forgotten" — an action succeeding into a state its own verify would
reject (novox/hq 04-ISSUES/017). A sealed secret is not recoverable afterwards, because
the mesh discarded the plaintext when it made it.

It now reads what it would destroy, names each thing one at a time, and refuses.
`--and-what-it-holds` is how somebody says they mean it, and the removal then reports
what went — this being the only record that any of it ever existed.

Reported as "operator settings do not persist, because re-registering a module
cascade-deletes them". Half of that is wrong, and the test now says so out loud: the
upsert is on the name, so `module add` at a new version leaves the settings, the secrets
and the ports exactly where they were. The command that destroyed them was `forget`, and
a wrong belief about which command destroys data is expensive in both directions — it
sends people looking for a fault that is not there, and leaves the real one unexamined.

Checked by internal/inventory/forget_test.go, which writes all three, re-registers the
module at a new version, reads them back, and only then tries to forget it.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:11:15 +02:00
jschoubben 9fab0b731a catalogue: a contribution reaches the port its own machine published, from any machine
The co-located fix could not reach a grant assembled for a consumer on another
machine: ContributionsFrom never sees a port map, so the proxy was told the
workload's software port and dialled a number that machine never published. The
consumer's own assignments are fetched where the grant is built and applied
there. The same fault as 038, one node over.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:08:09 +02:00
jschoubben e78c849002 catalogue: a provider that names a provision and says nothing is not the answer
servedOnThisMachine stopped at the first module whose `serves` mentioned the
provision, even when that entry was empty and there was therefore no fact to give
a consumer. here() had always kept looking in that case, and a set where one
module names a provision without describing it and another describes it is exactly
where the difference shows. Restore the search.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:04:16 +02:00
jschoubben 76fcbda9a7 route-proxy: an ACME account belongs to the authority that issued it
autocert keeps its account key at one fixed name, `acme_account+key`, in whatever
directory it is given, and reuses it for ever. That is right while the authority
stays the same and silently wrong the moment it does not. Re-initialising the
internal CA makes a new authority with a new root: it has never heard of the
account in the cache, rejects every use of it, and autocert has no path back from
that. Nothing is re-registered, no order ever reaches the CA, and issuance stops
with nothing saying why — until somebody guesses that deleting the cache directory
by hand is the answer.

Name the directory after the authority instead of sharing one between all of them:
a digest of the ACME directory URL and the root this proxy was told to verify it
with. A re-initialised CA has a new root, the mesh delivers it as a changed bundle,
the proxy restarts on that file and lands in a directory with no account in it, so
autocert registers afresh and orders again. The healing is that "is this account
still valid" never has to be asked — an account is only ever found where it is
still valid, which needs no error codes, no probe at startup and no network call
that can itself fail.

It closes a latent one of the same shape: pointing ACME_DIRECTORY at production
after testing against staging reused the staging account, because the cache had no
idea the two were different.

Trailing whitespace around the delivered root is not a new authority — the mesh
writes that file, and a newline coming or going must not throw away an account.
Old directories are left on disk, unused: they hold the only copy of certificates
that may still be valid, and this program is not the thing that should decide a
certificate is finished with.

A proxy already holding certificates orders them once more on the first start after
this, because its account moves. Free against the lab's own CA and against staging;
one issuance per name against a public authority.

novox/hq ADR 0056

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:02:00 +02:00
jschoubben 54cb915c4e push: one machine that cannot be worked out no longer holds back the mesh
A whole-mesh push refused outright the moment any single node failed to resolve:
every node was composed first, and one entry in `refusals` returned "nothing was
sent" before the send loop ran. So on the ADR 0056 lab, one module on the anchor
requiring a provision nobody had assigned a provider for — `nothing provides
"acme-ca", wanted by route-proxy` — stopped every OTHER machine from being sent
anything. Nothing converged anywhere, and the machines that went unconverged were
the ones with nothing wrong with them. The failure and the punishment were on
different machines.

This is the rule a92c11b established one level down, where an un-hostable module
stopped taking down the healthy modules beside it, applied one level up: the blast
radius of a fault is the thing that has it. A node whose declaration composes is
sent; a node whose does not is named, with its reason, and the push still ends
non-zero — skipping is not succeeding, and a command that exits cleanly having
missed a machine is a command that lies. The message now says how many machines
WERE sent, because "nothing was sent" was the claim that had become untrue.

sendTo keeps its all-or-nothing rule, and its comment now says that push
deliberately does not share it: a rotation reaching the consumer and refusing on
the provider leaves one end holding a credential the other has never heard of,
which is a real coupling between two named machines. A whole-mesh push has no such
coupling and never did.

Composing is split out of pushCommand so the rule can be tested without a broker
and a database.

novox/hq ADR 0056

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 20:59:21 +02:00
jschoubben b824c65ab0 catalogue: a co-located provider's served values, and a co-located contribution's port
Two more of one fault, and the fault is the same as 04-ISSUES/038: the same-node
path diverging from the cross-node one.

The mesh works out what a provider on ANOTHER machine serves by walking that node
— reading its manifest with that machine's port assignments, then settling the
result with that node's settings layers — before offering it to a consumer. A
provider on the consumer's OWN machine never passes through that walk, so every
step of it had to be repeated in resolve.go's servedHere and declaration.go's
here(). 038 repeated the port. Nothing repeated the settling.

So a served value the operator supplied reached a co-located consumer as the
manifest's empty default. On the ADR 0056 anchor that value is an internal CA's
root: step-ca and route-proxy on one node, route-proxy's binding carrying
root: "", an empty CA bundle written, a silent fall back to the system trust
store, and issuance stopping with nothing saying why. The same step-ca on another
node would have worked.

The second is the mirror direction. gitea declares a bare container port 3000 and
the machine publishes it as 20000:3000, but gitea's route CONTRIBUTION still said
3000 — so the proxy beside it dialled a port nothing listens on and answered 502.
038 fixed what a consumer is TOLD about a provider; this is what a workload TELLS
a provider about itself. The redirect uses the CONTRIBUTING module's assignment,
because the port is the workload's, not the proxy's; a contribution carried here
from another machine is left exactly as it is, its port being that machine's to
assign.

Both are settled in Declaration, which is the first moment the machine's ports and
the provider's settings both exist. That also removes an order dependence: the
resolver built its same-node needs mid-walk, from whichever modules had been
chosen by the time the requirement came up and in whatever order a map iterated,
so what a co-located binding carried depended on the order somebody happened to
assign things in. Re-deriving from the finished closure does not.

servedHere keeps its job — deciding whether a same-node provider serves anything
at all, which is what makes the need exist — and now says that its values are
provisional.

novox/hq ADR 0056

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 20:57:23 +02:00
jschoubben 0fb2ab7716 catalogue: composeName handles the apex label '@' (bare public domain)
An empty label composed nothing, so a module served at the bare domain (a node's
own site) had to keep a full name — the one route the label model could not
express. The zone-file convention '@' now composes to the public domain itself,
no leading dot, so the apex is a label like any other. Test added.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 00:27:11 +02:00
jschoubben 98f5d34610 route-proxy: an empty ACME_CA_BUNDLE means the system trust store
Piece A of ADR 0056 (selectable issuer). A provider that serves an empty root —
public-acme, whose root already ships in the OS trust store — leaves route-proxy's
CA bundle file existing but empty, because the mesh writes it unconditionally from
${bound:acme-ca:root}. Read that as "trust the system roots", the same as an unset
bundle, instead of failing with "holds no certificate this can trust". A bundle
that holds bytes but no parseable certificate is still refused.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 00:22:06 +02:00
jschoubben 232862315c catalogue: compose a route's name from a label and its node's domain, and resolve it in-mesh
A public route used to carry its whole hostname as a literal in the module
manifest, so running the same catalogue against a different domain meant
overriding that literal on every routed module, per node. The mesh was, in
effect, holding a map of names to services: the one thing it should never hold,
because the subdomain is the operator's choice and the domain is the node's.

Compose instead. A route contribution carries a `label` (the subdomain); a node
carries its `public_domain` as node-level configuration; the mesh joins
`<label>.<public-domain>` and grants exactly that, interpreting neither half.
Held as a node property beside the node's other node-level facts (endpoint,
site, overlay address), not in a module's settings — the ADR calls it
node-level, and the settings table is keyed per module.

Additive, so an unmigrated catalogue keeps working: a contribution that still
carries a full `name` and no `label` passes through unchanged, and the catalogue
can migrate module by module. A labelled contribution on a node with no public
domain composes nothing, reading downstream as a route that named no host.

And propagate: each granted route name is published into internal resolution
mesh-wide, mapped to the node that serves it, alongside the `<node>.internal`
names every container already gets. So a container — and an internal ACME
validator, which cannot complete a challenge for a name it cannot reach —
resolves a routed name to the proxy that serves it. Name-agnostic throughout:
the mesh propagates whatever names it was told to serve and knows nothing about
what they mean.

novox/hq 02-DECISIONS/0056

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 23:27:35 +02:00
jschoubben c147a26138 catalogue: announce a same-node provider at the port it is published on
A co-located consumer of a `from: mesh` provision was told the port the
provider module DECLARED, not the host port the mesh assigned and published
it on. The same-node served facts are settled while resolving (servedHere,
here()), before a bare `ports` mapping is assigned its host port, so they
carried the declared number; only the cross-node path re-derived them after
assignment. So the provider was published on <node>.internal:<assigned> while
its own-machine consumer dialled <node>.internal:<declared>, where nothing
listens — the ordinary small-mesh case, and the one the fix for issue 018
(announce the same-node provider at all) left one promise short of kept.

Redirect same-node needs to the machine's assignment in Declaration, where the
port map is known, exactly as plan.go already does cross-node. The publish bind
is unchanged (all interfaces, scoped to the mesh by the listen's firewall rule);
only the announced port is corrected.

novox/hq 04-ISSUES/038

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 23:02:50 +02:00
jschoubben 9e29772ed4 Merge pull request 'resolve: one un-hostable assignment no longer takes down a whole node's push' (#19) from feat/assign-hostability into main 2026-09-08 18:45:21 +02:00
jschoubben a92c11be12 Resolve: one un-hostable assignment no longer refuses the whole node
A module a person assigns to a machine that cannot host it — its declared
capability has no detector there, as fail2ban does on a host with no firewall —
made Resolve refuse the entire node, so a whole-node push refused to send the
healthy modules beside it too. One module on the wrong machine took down every
other module on that node.

Assign already keeps such an assignment on purpose (it is what a person meant,
and acts.go says so), so the fix is on the resolve/push side: a directly-assigned
module the machine cannot host is left out of the closure and reported as
un-applied on the Resolution, rather than refusing the set. The healthy modules
still resolve, declare, and converge. A module that is *required* by something
running here and cannot be hosted still refuses — that set is genuinely
incoherent — so the distinction is who wanted it.

assign, plan and push now name the un-applied module and the missing capability,
via a shared WrongMachine message, so it is neither silently dropped nor fatal.

Reconciled two tests that encoded the old whole-node refusal for directly-assigned
un-hostable modules; added coverage for the healthy-modules-still-converge case
and the required-un-hostable-still-refuses distinction.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-08 18:37:41 +02:00
jschoubben 474382c141 Merge pull request 'catalogue: deliver a keyless same-node provider's served facts' (#18) from feat/serve-here-keyless into main 2026-09-07 05:25:57 +02:00
jschoubben b78a911e34 catalogue: deliver a keyless same-node provider's served facts
A node-scope provider that answers a requirement on the same machine and
serves connection facts (a port) but mints no credential delivered
nothing to a co-located consumer. resolve.go only built the delivering
Needed when brokered[want] was set — true only for mesh-scope providers;
a node-scope keyless provider set local[want] instead and fell through,
so knownFor saw no binding and boundInto refused the consumer's
${bound:model-access:port} file.

Deliver the served facts as a need whenever the same-node answer serves a
non-empty set, with a loopback fallback for the address when the node is
off the private network — the reachability rule does not apply to two ends
on one machine. The brokered (credentialed, mesh-scope) path is untouched.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 05:08:04 +02:00
jschoubben e5eaa0478a Merge pull request 'adapters: openai is a static-key vendor' (#17) from feat/openai-access into main 2026-09-07 04:25:53 +02:00
jschoubben 55753d5cfd adapters: openai is a static-key vendor
One registry line — `"openai": StaticKey`. The generic static-key adapter
already serves any vendor (accept is the vendor-independent seal, deliver is the
value unchanged, and refresh/identity/usage are not implemented), so a second
static-key vendor is data, not code. The shape-selection test now asserts
For("openai") reports static-key.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 04:25:26 +02:00
jschoubben e03833ec40 Merge pull request 'catalogue: give host-network containers the mesh's names too' (#16) from feat/usage-store into main 2026-09-07 04:08:19 +02:00
jschoubben 958bef56c7 catalogue: give host-network containers the mesh's names too
A container with `network: host` was skipped when the mesh injects its
`<node>.internal` names, on the belief it "shares the machine's hosts file
already". It does not: `docker run --network host` still gives the container
its own /etc/hosts (localhost and its own id only), so every internal name the
mesh wrote is invisible inside it, and a client that dials one gets EAI_AGAIN.

This surfaced with the first host-network consumer to dial a provider by the
`.internal` address the mesh hands it as `${bound:...:at}` (the model-usage
store reaching its postgres). The remedy is the same `--add-host` every other
container already gets — the runtime accepts it with `--network host`
(verified against Docker) and mesh-host emits it for any network mode.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 03:44:28 +02:00
jschoubben d9c4818e6d Merge pull request 'model-access: refreshable-grant — manager holds the refresh token, control plane never reads it' (#15) from feat/model-access-submit into main 2026-09-07 02:48:49 +02:00
jschoubben f00077f376 secrets: regenerate the module sealed-box cross-check fixture
The manager module's seal now runs over the audited tweetnacl-sealedbox-js (mesh-catalog
anthropic-manager) rather than a hand-transcribed NaCl. Regenerate the fixture's sealed value
from that new seal() over the same node key pair and plaintext, so the fixture is the new
library's output. crypto_box_seal is randomised, so the blob differs; the wire format does
not. TestModuleSealedBoxOpensInGo still opens it under box.OpenAnonymous and recovers the
plaintext, proving TS(tweetnacl-sealedbox-js) to Go interop.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 02:17:54 +02:00
jschoubben 33fd28ffa6 licences: deliver the refresh token by the ordinary sealed path, not a bespoke envelope
The refreshable-grant refresh token no longer rides a custom at-rest envelope that a
module opens with a node private key. A module is never given a node's private sealing
key, so that path could not exist -- the gap Phase C hit.

Instead the refresh token is a credential sealed to the MANAGER holder with the same
anonymous box (secrets.Seal / crypto_box_seal) every credential uses, stored as one
sealed blob, and delivered by the existing host-unseal-and-mount: the host opens it with
the node's real key and mounts the cleartext at the manager module's bound path, exactly
as a consumer's db password is delivered.

  - refresh_grant now stores { sealed, manager_key }, dropping the AtRest token/wrapped_key
    columns; internal/secrets/atrest.go is retired (nothing else used it).
  - the licence records its manager as (node, module); KeyFor delivers the refresh token to
    the manager holder and the access token to consumers, disambiguated by module so the two
    can co-locate. Accept and the reseal skip the manager holder.
  - the manager holder is delivered the node's PUBLIC sealing key in its bound facts, so the
    module can re-seal a rotated refresh token with no private key of its own; the
    declaration tolerates its empty pre-adoption secret rather than refusing.
  - SubmitRefresh / set-grant take a sealed blob, never a refresh token in the clear.

The invariant holds unchanged: the control plane never reads the refresh token, and no node
but the manager holds it. A committed cross-language test proves the TypeScript module seal
opens under Go box.OpenAnonymous (the host's Unseal) -- both are NaCl crypto_box_seal.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 01:55:08 +02:00
jschoubben 8e0c22fc2e licences: submit-refresh, the module-produced refresh entry point
Phase C of model-access (ADR 0050). Refresh above calls an in-process
VendorRefresher, which would open the at-rest envelope inside the control
plane's own process. Anthropic must not: its refresh runs on the manager
node. So add SubmitRefresh, the companion that publishes a refresh a
manager node already performed -- it is given only the new access token in
the clear (sealed per holder, as any accepted key) and an opaque re-sealed
refresh envelope (stored unopened). The refresh token in the clear never
crosses this boundary. The reseal-and-publish half is extracted and shared
with Refresh, so the sealing logic is one implementation.

CLI: licence grant (print the opaque envelope), set-grant (store a
module-produced envelope -- adoption), submit-refresh (access token +
optional rotated envelope). Tests defend that the manager alone opens the
refresh token and the control plane never holds it in the clear.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 01:00:27 +02:00
jschoubben 59faa150ac Merge pull request 'model-access B: the refreshable-grant machinery (ADR 0050 carve-out)' (#14) from feat/model-access-refreshable into main 2026-09-07 00:25:15 +02:00
jschoubben 2e33c5e80e model access B: refreshable-grant machinery — manager, at-rest refresh token, refresh flow
The ADR 0050 carve-out, built generic and vendor-neutral. A refreshable-grant
licence records one manager node; that node holds the refresh token encrypted at
rest, access tokens are still sealed per holder, and the refresh token is never in
a holder's delivery. Bounded on the three stated axes: refreshable-grant vendors
only, the refresh token only, the manager node only. Anthropic's actual OAuth
refresh stays a Phase-C plug-in behind a clean seam.

- New at-rest crypto (secrets.SealAtRest/OpenAtRest): envelope encryption distinct
  from the per-holder anonymous-box seal. The refresh token is under a symmetric
  data key (secretbox); the data key is wrapped to the manager node's public
  sealing key. The database alone holds ciphertext and a wrapped key with no
  private half to open either — only the manager node reads it back.

- Refreshable-grant adapter dispatch: anthropic is now refreshable-grant,
  anthropic-api-key the static-key second case. The adapter implements the
  Refresher seam by delegating to an injected VendorRefresher (the Phase-C plug,
  none shipped). static-key is untouched. The type assertion to Refresher is what
  gates the carve-out to refreshable-grant vendors.

- Refresh lease/rotate/publish flow (Licences.Refresh): a transaction-scoped
  advisory lock is the single-refresher lease; the new access token comes from the
  vendor refresh, is sealed per holder (secrets.Seal, as Accept does) and delivered
  on the next push — doc 13's reseal-and-publish half, all-or-nothing. The refresh
  token stays put, re-encrypted at rest only if the vendor rotated it.

- Manager and refresh_grant schema: consolidated into migrations/0001 and carried
  by a new incremental 0003 (the dual-write rule).

- 17 new tests, including the four security checks: KeyFor never carries the
  refresh token, a static key has no manager and cannot be refreshed, the at-rest
  token needs the manager's key, and a refresh delivers a new sealed access token.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 00:23:35 +02:00
jschoubben 163200c4dd Merge pull request 'model-access A: provider→vendor + adapter dispatch + static-key (ADR 0050 foundation)' (#13) from feat/model-access-vendor into main 2026-09-07 00:02:27 +02:00
jschoubben ddb41baaf4 Model access is vendor-agnostic: rename provider→vendor, add adapter seam (Phase A)
ADR 0050 Phase A. Rename the licence's `provider` field to `vendor` — the
inventory already uses "provider" for which node answers a brokered provision,
and one word must not carry two facts — and route the licence layer's sealing
and delivery through a per-vendor adapter selected by that field.

The rename touches the Go struct/params/SQL in internal/licences, the operator
CLI, and the schema: 0001 (the consolidated schema) now creates the column as
`vendor`; a new guarded 0002 renames it on a database that predates the change,
and is a no-op on a fresh one.

The adapter (internal/licences/adapters) has a `shape` and the two verbs a
static-key vendor needs — accept (the generic anonymous-box seal) and deliver
(the sealed blob unchanged). refresh/identity/usage are named as optional
capability interfaces so the refreshable-grant seam exists before its code.
A registry maps vendor→shape (anthropic→static-key for now, with a Phase-B
TODO to swap it to refreshable-grant); an unknown vendor is refused clearly.

Behaviour is unchanged from the operator's view except the field name.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 23:49:12 +02:00
jschoubben 671fb4f8f3 Merge pull request 'fix: a require-only consumer of a parameterless provision still asks (mint gap)' (#12) from fix/require-only-mint into main 2026-09-06 23:21:25 +02:00
jschoubben d0ef659824 fix: a require-only consumer of a parameterless provision still asks (and is minted a credential)
A consumer that requires a provision whose serves names no consumer key (redis-cache, amqp)
contributes no payload, but it still ASKS for it. ContributionsFrom keyed 'asks' on
contributions alone, so such a consumer's grant got From='' — read as withdrawn — and the
provider never created its account. redis-cache consumers (e.g. baserow) were silently
unprovisioned, tolerated only by their embedded fallback. A module asks iff it still requires
the provision, whether or not it hands anything up. Regression test added.

Found by the lavinmq AMQP provider bed (given:[] for a require-only amqp consumer); fix
lab-proven green there.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 23:20:58 +02:00
jschoubben b69fbc0e86 Merge pull request 'schedule: parse + validate a cron on a container, refuse run-once+schedule (ADR 0053)' (#11) from feat/schedule-container into main 2026-09-06 14:27:41 +02:00
jschoubben 78b8b6e256 catalogue: carry and validate a container schedule (ADR 0053)
A container may declare schedule: "<cron>", the recurring twin of
run-once. The resolver already carries a resource's keys through
untouched, so schedule reaches the rendered host declaration on its own;
what belongs here is refusing, near its author, what the host would
otherwise refuse far away.

The manifest parser refuses a schedule that is not a string, one that is
not a well-formed five-field cron (cron.go: fields, ranges, *, comma,
dash, slash), and the contradictory pair run-once + schedule -- a
container runs once and gates, or on a cadence, or stays up, never two.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 14:08:51 +02:00
jschoubben dc65440e31 Merge pull request 'fix: a ${secret:} placeholder must fill from the file-owner's credential (provider-seal-key gate)' (#10) from fix/secret-per-consumer into main 2026-09-06 13:47:22 +02:00
jschoubben 99a753994e fix: fill a ptr-secret placeholder from the file-owner's credential, not the last consumer's
The provider-seal-key gate: on a node with two modules requiring the same provision (baserow
and letta both consuming postgres), sealedFor matched a need by provision NAME alone, so a
file's ${secret:X} placeholder took whichever consumer's sealed credential came last in
r.Needs -- the OTHER module's password. baserow was handed letta's password and could not
authenticate. The secrets:-map delivery path already guards this (For == m.Module, novox/hq
04-ISSUES/022); the ${secret:...} placeholder path did not. Added the same guard.

Also dedups the contributions file: when provider and consumer are co-located, grantsFor
enumerates the same-node consumer, so a consumer was emitted twice into the provider's
receives file (once full with its grant, once partial). The m.Contributes loop now skips a
(provision, module) the grants loop already carried; non-grant contributions (routes) still emit.

Regression test added: two consumers of one provision each get their own credential. Proven
end-to-end on a two-node lab install (mesh-lab assigned-two-node-db): baserow and letta on one
node, substrate on another, each authenticates with its own minted password.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 13:46:55 +02:00
jschoubben 3b0f17b8fd Merge pull request 'run-once: a container the host runs to completion (ADR 0052)' (#9) from feat/lifecycle-run-once into main 2026-09-06 00:05:06 +02:00
jschoubben f96c247c5f catalogue: run-once is a step the host runs to completion (ADR 0052)
A container may be marked `run-once: true` — a step the host runs to completion,
gating whatever the declaration places after it. The control plane's part is
small: the field is carried to the host unchanged (containers pass through as
maps), and the step keeps its author-order position ahead of the container it
gates, because the gate is declaration order, not a resolved dependency
(ADR 0005).

The manifest parser refuses a run-once that is not a boolean and the pair
run-once + restart-on (contradictory lifecycles) — near the manifest rather than
far away on the machine, the same lesson the action ban records. Three unit
tests; go build ./... and go test ./... green.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 23:57:33 +02:00
jschoubben 6bb9434298 Merge pull request 'A module accesses operator-owned data, it does not own it (ADR 0051)' (#8) from feat/shared-data-access into main 2026-09-05 22:48:11 +02:00
jschoubben bcdde475cd Merge pull request 'A bare alive moves last_seen and nothing else (convergence race fix)' (#7) from fix/service-only-converge into main 2026-09-05 22:47:23 +02:00
jschoubben cae9a3e54f A bare alive moves last_seen and nothing else
A node says it is there every minute and describes what it applied rarely,
and both went through Heard, which wrote every one down as a report. So a
bare alive replaced the node's last real apply with an empty one -- clearing
the declaration digest `current` is measured against, the carried ports a
push assigns around, and the clean-or-failed outcome. A node that had just
caught up read as behind within the minute, and never converged.

Whether it converged in time was a race the node's own apply set: the link's
one loop applies a declaration to completion before it can send the pending
heartbeat, so a fast apply (catalogue-small) leaves the digest standing the
~60s until the next beat -- long enough for the lab to see `current` -- while
a heavy wave whose apply outran the first beat (mongodb + unifi + marrytts)
had the alive fire milliseconds after the report and never showed `current`
at all, timing out settle even at 1200s.

Heard now returns after moving last_seen for a report that carries no account
of what the machine did -- nothing applied, nothing refused, nothing failed,
which is exactly a bare alive. A real report always carries one. This is what
the commit that began hearing alives said it did and did not: "a bare word
that a node is there moves last_seen and touches nothing else."

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 22:37:53 +02:00
jschoubben aeb65a3e1d catalogue: a module accesses operator-owned data, and does not own it
04-ISSUES/036: the media stack is several modules that must share the
library and download directories on one machine, but the manifest could
only say "a directory I own". Six modules each declared the same paths as
their own resources, and the resolver's duplicate-owner refusal — right
in general — would refuse the stack's only sensible assignment the first
time two of them landed on one node.

Add an `accesses` field: a pre-existing, operator-owned path a module is
granted use of but does not own (novox/hq ADR 0051). Distinct from a
`directory` resource on every axis the host acts on — the mesh creates,
chowns and reconciles a directory; it mounts an access and owns nothing.
An access is not a resource, so it never enters the duplicate-owner map
and several modules may name one path with no conflict. What is refused
is the contradiction: a path one module owns and another accesses.

Rendered into the declaration as an `access` resource, before the
container that mounts it, so the host can find it present or refuse
clearly. Unit tests cover co-resolution (the exact 036 case), the
unchanged owner-vs-owner refusal, the owner-vs-accessor refusal, and
access validation.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 22:10:58 +02:00
jschoubben 59fcb41855 Merge pull request 'Resync hq ADR references (0044-0054 -> 0039-0049)' (#6) from feat/adr-ref-resync into main 2026-09-05 12:47:52 +02:00
jschoubben 87c193b202 Resync hq ADR references 0044-0054 -> 0039-0049 after the hq record reconciliation 2026-09-05 12:47:13 +02:00
jschoubben 5541949b59 Merge pull request 'broker: a module account scopes its tool serve queues + mesh.rpc (ADR 0052)' (#5) from events/tool-account-scope into initialization 2026-09-05 03:06:51 +02:00
jschoubben 9b7ba2e20c identity: a consumer's identity fits the tightest backend, via a slug (ADR 0054)
A module may declare a short `slug`; the mesh derives mesh_<node>_<slug|name> and
refuses at assignment (naming the slug as the remedy) when it would still overflow —
identityLimit is now 20, an S3 access key's, the tightest of the backends a login
reaches (04-ISSUES/010). The slug rides the grant so the provider derives the same
login the consumer does, even across nodes. CheckIdentity is now wired, in grantsFor.

Also, the minted secret shrinks to 40 chars (30 bytes) from 43: an S3 secret key is
8-40, the same fit-the-tightest-backend rule on the credential's other half.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 02:51:39 +02:00
jschoubben b1bf1659d9 broker: a module account scopes its tool serve queues and mesh.rpc (ADR 0052)
CreateModuleAccount now also grants serve.<module>.* (declare, bind, consume its
own tool queues) and mesh.rpc (bind them on, publish replies) — so a module can
serve its tools and reply, scoped to exactly its own, and no other module's. The
broker tests still hold a module out of another's queue.
2026-09-04 21:56:05 +02:00
jschoubben b306c74467 filtering: a per-node 'expose' setting overrides a listen's source (ADR 0051)
listens.from was a manifest constant — one value for every node a module runs
on. Now a per-node setting overrides it: {"expose": {"5432": "anywhere"}} makes
postgres public on the machine it is set for while it stays from:mesh elsewhere,
and the firewall (ADR 0050) is computed from the effective source. Exposure()
validates it — a port the module does not listen on, or a source that is not
mesh/anywhere/machine, is refused rather than reaching nothing; UnusedSettings
knows 'expose' is a real destination. Tested: default mesh, setting opens it to
anywhere, bad settings refused.
2026-09-04 21:12:33 +02:00
jschoubben 931ca6f01e cli: module issue seals the node and module into the credential
So the runtime knows the identity the account was scoped to, without a
manifest naming the node.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 01:53:23 +02:00
jschoubben 64496f0335 broker: the substrate pre-declares a consumer's dead-lettered queue (ADR 0048)
LavinMQ refuses a non-administrator declaring a queue with a dead-letter
exchange, so a scoped module cannot make its own. EnsureModuleQueue declares
<node>.<module>.events with its DLX as the mesh, and 'module issue' does so
for a consuming module — the runtime then passively checks it rather than
declaring. Verified against a real broker: the scoped account binds and
consumes the pre-declared queue.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 01:51:17 +02:00
jschoubben f41d280e66 cli: module issue — deliver a module its scoped broker account (ADR 0048)
'module issue <module> --node <m>' looks up the module's emits/consumes from
the catalogue, ensures the bus exchanges exist, creates its scoped account
(CreateModuleAccount), and seals an amqps {url,fingerprint} to the node as the
module's broker own-secret — the same delivery as 'builder issue', now generic.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 01:33:03 +02:00
jschoubben 47bbb0cca6 broker: a generic module account, scoped by emits and consumes (ADR 0048)
CreateModuleAccount gives an assigned module its own broker account whose
permissions ARE its manifest: declare and read its own <node>.<module>.events
queue, read the events exchange to bind onto if it consumes, write the events
exchange only if it emits. The account name carries the node (sealed per
machine), the permissions carry the module (one cannot read another's queue).
The builder becomes one instance of this rule rather than a separate kind.

EnsureEventExchanges declares the bus the substrate owns — mesh.events,
mesh.rpc, mesh.events.dead + a retention queue — idempotently, since a module
account may not declare an exchange.

Scope tested as patterns (no broker needed), and every management call verified
against a real LavinMQ. Honest limit recorded in the code: LavinMQ has no topic
permissions, so ADR 0047's emit-origin reservation (module.<self>.*) is stamped
by the sdk, not enforced by the broker; a pure consumer like the audit logger
is unaffected.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 01:31:19 +02:00
jschoubben 2c06c0d663 manifest: emits and consumes — the event relationship
A module declares the event types it emits and the patterns it consumes,
parallel to provides/requires (novox/hq ADR 0046). Events span module.*,
mesh.* and node.* sources; a consumes for an event nothing emits is a
dangling edge. Fields only here; the dangling-edge check and the runtime
wiring follow.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-03 23:44:30 +02:00