Commit Graph
73 Commits
Author SHA1 Message Date
jschoubben ee1b8ffe24 1.7, second half: the user list read out of the mesh's records
The derivation had nothing feeding it. `BusRecords` reads what it needs — the
machines, what each runs, every manifest, and which machines hold a live token —
and turns it into the records the composer derives from.

**A module's authority comes from its manifest, not from its assignment.** The
assignment says where it runs; what it may say is what it declared. So the two are
read together and the manifest decides, which is also why a seat's protocol is
gathered across the whole catalogue rather than from one manifest: a seat is
declared by one module and held by another, and that is the whole reason a seat
exists.

Three things checked against a real store, each a user that would be wrong in a
way nothing reports:

- A module assigned to a machine becomes a user with exactly the authority it
  declared, including the protocol of a seat some *other* module declared — a
  module granted nothing on a seat it was assigned to send to would fail on its
  first publish with an authorisation error that says nothing about a seat.
- Only a machine holding a live token gets an enrolment user. One outliving its
  token is a right to join that nobody issued.
- A module assigned and absent from the catalogue is refused rather than composed
  with an empty permission list. The catalogue already refuses to forget an
  assigned module, so this is the second line — and it earns its place there,
  because relying on another package's invariant is how a rule ends up enforced by
  nothing.

People are left empty rather than guessed at: the account model is built and
`operator issue` is not, so there is nobody to derive yet.
2026-09-27 01:49:09 +02:00
jschoubben 0560c792d8 1.7, first half: the mesh can say who its bus users are, and hold their keys
Two pieces the composer has been waiting for since it was written.

**The credential has to outlive its own minting.** On the bus the mesh runs on
today an account is a management call: mint a password, hand it over, seal the
plaintext to whoever will use it, keep nothing — which works because the broker
remembers. Here the users are one file, rewritten whenever any of it changes, so
keeping nothing would mean the first person's access change silently blanking
every module's password. So a bus user's bcrypt hash is now recorded, keyed by the
username the file needs, and the plaintext comes back exactly once. Verified
against a real store that the hash verifies the password it was made from, that
the password itself is not in there, that minting again rotates rather than adds,
and that forgetting a node takes its host's and its modules' credentials with it.

**Permissions are not stored, and that is the point.** Only the credential is
kept. Authority is derived from what each module declares, every time the file is
written (ADR 0043) — a stored permission list would be a second account of a
user's authority, able to disagree with the records it came from, and both would
look internally consistent while they did.

`Users` derives the list: the controller always first and always present, one
user per node, one per module per node, one per live token, one per person. Two
users with one name is refused where both can be named, rather than left to be
whichever one the server happened to read. A user the mesh has never minted a
password for is *named* rather than dropped or written as a user anybody is:
that is an ordinary situation with an obvious remedy, and the caller decides
whether a partial file is worth writing.

What remains of 1.7: delivering the file to the node that runs the server, and
minting at enrolment and assignment — which is transport-coupled, because a node
on the old bus must not be handed a credential for the new one.
2026-09-27 01:44:53 +02:00
jschoubben 06cf3c04e5 The consume side behind a seam, and the window wiring into the loop
The outbound half went behind `Bus` and the transport stopped reaching its
callers; this is the other half, and the larger one. Every handler took
`amqp.Delivery`, so the serving loop could not move to another bus without
moving enrolment, reports, builds, upgrades and catch-up with it in one breath.

`Control` states one message in the mesh's words — took it, dropped it, or held
it for the store — and `Inbound` is where messages come from. The AMQP
implementation is today's loop moved rather than changed: same queues, same
prefetch, same holding, because the mesh is running on it and a bus nothing
speaks yet is no reason to alter the one every node is on.

The window (window.go) is now what decides, instead of the conditions that were
inlined in the loop. Two things that surfaced in the wiring:

**Supersession is asked before the store, not after.** A report about a
declaration the mesh has moved past would otherwise wait out a restarting store
to be written and then overwrite what the node is doing now.

**Half of a report is not about a declaration, and that half is never stale.**
What the machine *is* — the tunnel it took over, the ports its own bundle
holds, what an adopted node found, a node moving its overlay key — reaches the
mesh on a report and nowhere else. A rekey set aside as stale is a node whose
overlay key never moves, and no retry is coming, because the node said it once.
So staleness is asked only of a report that is purely an apply's account.

The one thing holding-in-memory can do that holding-in-the-server cannot is
named rather than hidden: `About` sets aside a held message when a newer one
about the same thing arrives, and the bus being built ignores it because the
digest answers the same question.
2026-09-27 00:44:16 +02:00
jochen 97448194ac Seats are a closed set, a seat's holder answers for what it delivers, and a build source may live on the git seat
Implements novox/hq ADR 0110 and 0111.

The seat set lives in internal/catalogue/seats.go: fourteen seats, each with a scope, what occupying
it delivers, and the record that made it one. A test asserts the count and a decision per entry, so
changing the set means finding the argument, as the host's vocabulary test does. The first set is
every seat already claimed — including the-private-network, which the network module claims from a
manifest composed in this repository's code, not from any module.json — plus npm-package-registry
(ADR 0109) and git (ADR 0111). A test parses every catalogue manifest and this repository's own and
fails on any refused claim, so closing the set refuses nothing in use.

ParseManifest now refuses a claim on a seat the mesh does not define, a seat claimed at another
scope, and a delivering seat claimed by a module that does not provide what it delivers. A
malformed claim is refused once, for being malformed.

Resolution: among several providers of a mesh provision, a pin still wins; then the holder of the
seat that delivers it; then the only provider; otherwise refused as before. ADR 0009's "never
guessed" holds — the seat is the choice made once, mesh-wide, rather than a pin per consumer node.
A provider now carries the module it came from, because a provider is a (node, module) pair and the
pair is what tells a holder from a neighbour on the same machine.

The planner's second pass is now given the first pass's holdings. Without them, a node consuming a
seat-delivered provision was refused there, and a refused node's own claims dropped out of what the
mesh holds — letting a second holder of one of its seats pass unrefused.

`seats [--json]` lists every seat, what it delivers, and each holder, derived from assignments
every time and never stored. Unheld seats are listed. A stored claim outside the set — possible
for a manifest registered before the set closed, since stored manifests are not re-validated — is
shown rather than hidden.

`build --self <owner>/<repo>` builds from a repository on the git seat's holder. The clone URL is
composed at build time from the holder's node and what it serves for git; the recorded source is the
path and the seat (migration 0032), never an address, so a moved forge changes nothing recorded.
Nobody holding the seat refuses self-hosted builds and says so; external URLs are unchanged. An
address passed with --self is refused rather than recorded as a path.

Replaces three foundation tests that defended the builder's carried package binding. The catalogue
removed that binding when the builder began requiring the registry through a real grant, so the
tests were already failing on main; they now assert the builder requires what the npm seat delivers
and carries no copy of its own, and that the forge holds the npm and git seats.

Verified: go vet clean; the whole suite passes against a throwaway Postgres (make postgres), the new
inventory tests included; gofmt clean apart from cmd/mesh-builder/stdout_test.go, which fails on
main too.
2026-09-25 20:48:10 +02:00
jschoubben 4566c5c9aa Adopt the tunnel as a mesh fact, refuse a mismatched takeover, and rekey after enrolment
Review of the ADR 0105 build (hq ADR 0105). Four things it got wrong and one
path it lacked:

- A predecessor spoke's tunnel names one peer, the hub, routed the whole
  range; recording refused it and the whole enrolment failed. Range-routed
  peers are skipped now — only the hub's peers are ever carried.
- The range and the carried peers were conditions on the node being adopted,
  so converging the hub would have renumbered the mesh and dropped the peers
  still reaching it. They are facts of the tunnel record now, mode aside; the
  takeover alone is declared to an adopted node. Converging the hub is refused
  while a carried peer has not enrolled, naming it.
- A push composed a takeover for a hub whose address or endpoint disagreed
  with the tunnel, which would have the host stop the found interface and
  raise the mesh's where no peer listens. The graph refuses to compose it,
  naming both and the placement that fixes it.
- The host's account said taken or not; "found down and the mesh's not up"
  read as not taken. Three states now, and an account on every takeover.
- A hub that enrolled before this feature holds a key of its own, and
  re-enrolling would rotate every key the mesh sealed credentials to. A node
  now rekeys in a report, signed with its identity key over the key it
  leaves, the key it takes and the tunnel; the mesh verifies against the live
  key, refuses a stale or foreign proof, records key and tunnel, and moves a
  hub to the tunnel's address. `overlay show` names the path for a hub that
  found no tunnel.

Also: a carried IPv6 peer is routed /128, and identity.ForTest exists so the
link can be tested against a real identity store.
2026-09-24 00:02:07 +02:00
jschoubben 3c836f0abb Adopt the predecessor's tunnel in place: its range, its address, its peers
On an adopted hub the private network takes over the tunnel it finds rather
than running beside it (hq ADR 0105): two tunnels leave the mesh's unreachable
through the provider's filter, so no machine can ever join.

The node presents the found tunnel when it enrols, under the key it took as
its own; the inventory records it (node.tunnel, tunnel_peer — migration 0031)
and the mesh composes from it: the overlay's range is the adopted tunnel's,
the hub is placed at the tunnel's address on the tunnel's port, and every
peer the tunnel had is carried in the hub's peer list as a peer of the
tunnel, not a node of the mesh, until a node enrols with that key — which
then keeps the address the tunnel had for it. A fresh node never gets an
address the tunnel holds. The hub's declaration tells the host which unit to
take over; the host's account of carrying it is recorded and shown.

Every reader of the range follows the setting; nothing stores it. A found
tunnel under another key is recorded and not adopted, so ADR 0100's
non-overlap rule keeps applying where a tunnel is left running beside the
mesh's. A lab bed and test skeleton for "How it is checked" are under lab/.
2026-09-23 23:26:34 +02:00
jschoubben 2cf8739a84 Clear the whole account an adopted node gave when it converges (hq ADR 0100) 2026-09-22 19:45:35 +02:00
jschoubben 87ecc9326e Refuse a port given for the whole mesh where it is set, not at every node's composition (hq ADR 0100) 2026-09-22 19:45:35 +02:00
jschoubben 65187d4de0 Wait for a held node without pinning a pool connection, and give up after a bounded wait naming it (hq ADR 0100) 2026-09-22 18:31:10 +02:00
jschoubben 689c2c6d33 Hold a node from composing to sending, so a push composed before converge or adopt is never sent after it (hq ADR 0100) 2026-09-22 18:10:04 +02:00
jschoubben 4bb19c9e40 Give a machine port one holder: refuse ssh's, another module's and a doubled one, and release the assignment a given port replaces (hq ADR 0100) 2026-09-22 18:05:31 +02:00
jschoubben dbf62b5212 Read the foundation's ports from each node's settings, wherever a port is used (hq ADR 0100) 2026-09-22 17:31:07 +02:00
jschoubben 28894fa5bd Keep what an adopted node reports holding, the firewall it found and what is reachable on it (hq ADR 0100) 2026-09-22 17:23:09 +02:00
jschoubben 1c32af6a22 Record whether a node is adopted and which modules were taken on it (hq ADR 0100) 2026-09-22 17:14:05 +02:00
jschoubben 4567fa666c Review of 083: finishing an enrolment whose token was spent takes proof of the key's private half, a live lease and a first delivery — a public key alone cannot replay a spent token; shutdown leaves held messages for the broker; identical builds supersede; what is held leaves room in the prefetch 2026-09-22 14:33:13 +02:00
jschoubben a3b7e830c8 Review of 083: a message the store cannot take is held and retried on a ticker, not slept on, so enrolments are answered meanwhile; a newer one per subject supersedes; the password is replaced after the spend; the same presenter may finish after a lost answer; upgrades retry only on the store 2026-09-22 14:23:03 +02:00
jschoubben 1a41b88ed3 Nothing the control queue carries is lost while the store restarts: an enrolment claims its token and spends it last, and is asked to try again; build results, upgrades and catch-ups are handed back, bounded (novox/hq issue 083) 2026-09-22 14:06:28 +02:00
jschoubben 3fd0c37d22 Review of 082: a store killed mid-conversation is an outage and a wrong password is not; one report holds the queue at most two minutes, then is let go loudly; shutdown does not wait out the pause 2026-09-22 13:34:04 +02:00
jschoubben 047830a6b5 A report the store cannot take yet is handed back to the broker, not acknowledged and lost (novox/hq issue 082) 2026-09-22 13:25:13 +02:00
jschoubben 0d1f2a4152 Review: module issue's pre-check factored and tested; a pair delivery for a requirement the module keeps no secret for is refused; builder issue's usage says the module comes first 2026-09-21 23:58:36 +02:00
jschoubben 35c5c2bb9b secret accept is refused for a name the module does not declare, and a pair delivery for a requirement or local it has not got (novox/hq issue 078) 2026-09-21 23:31:06 +02:00
jschoubben 9f3790dcda Review: one local name is still a local name; a local name is unique; recovery knows it; recipes read as instructions; a tag before a digest; ask fails at once when nothing serves
A secrets object with one local name delivered no file. Two requirements could share
a local name. secret recover and the export could not tell two locals apart. The
recipe check missed continued lines and read heredoc bodies as bases. repo:tag@digest
kept the tag in the repository. ask now publishes mandatory, so a tool nothing serves
is said at once rather than after the wait.
2026-09-21 21:03:22 +02:00
jschoubben 6ae4ae1dba A module may hold several secrets from one provider, each a pair of its own
secrets: maps a requirement to several files under local names. Each local name is
its own need, its own pair credential (the pair is keyed on it: migration 0027),
its own file on the consumer, its own holder at the provider (the identity with the
local name after it) and rotates apart from the others. The plain shape is
unchanged and every existing row is the credential it was (novox/hq 04-ISSUES/069,
ADR 0094).
2026-09-21 20:28:16 +02:00
jschoubben 53a79a17a1 Review: a failure is the same by resource id, not by the host's words; bound files and /run/docker.sock are declared
The host's error text may carry a duration or a counter, and a resource looping on
it would never have read as stuck. The previous row is read and compared here.
Stuck needs a start to say. A container may mount the file a binding lands in; the
runtime socket is declared under both of its spellings; the catalogue-wide test
takes MESH_CATALOG.
2026-09-21 19:23:02 +02:00
jschoubben 396e05bb65 An operator delivers a pair credential, and the mesh never replaces it
secret accept grows --provider: the value is sealed to the consumer's node, the
provider's node and the operator's key, and the pair records origin 'accepted'.
An accepted pair is not remade when a key changes (the mesh does not hold the
value; the read is refused naming the remedy) and rotate refuses it (accepting a
new value is the rotation). The vault's third species has its entry
(novox/hq 04-ISSUES/070, ADR 0092).
2026-09-21 17:50:41 +02:00
jschoubben fcaad7271a A failure that repeats is said to be stuck
The mesh kept one report per machine, replaced, so a resource nothing can ever apply
looked like a failure that had just happened, every reconcile interval, for ever.
The row now keeps when the current failure began and how many reports in a row have
said it — the same outcome, refusal and failed resources; anything different starts
again and a clean apply clears it. Three make the machine stuck, and status says so
beside the failure, in words and in JSON (novox/hq 04-ISSUES/065, ADR 0090).
2026-09-21 17:43:24 +02:00
jschoubben 77e6c1a684 Recoverable means sealed to the current operator key; recovery names the provider
From review: the export counted any operator-sealed row as recoverable, so a
secret sealed to a replaced key was reported as openable with the current one;
replacing the key counted orphans in one table of two; and a pair credential
held from two providers was recovered as whichever row came first. The export
now lists what the current key opens, what an earlier key opens, and what has
no copy; `secret recover` takes --provider and refuses ambiguity; files that
must not exist are created exclusively; one constructor builds the export for
the operator's file and the vault's disk alike.
2026-09-21 01:16:32 +02:00
jschoubben 565f144a20 A pair credential is sealed to the operator key too
The secret the vault provides a module is the credential of the consumer↔vault
pair, and so is every credential a provider grants; sealing only own secrets
to the operator left exactly those unrecoverable. Same column, same call; the
export and `secret recover` address a pair by consumer node, module and the
provision's name, and say which kind each entry is.
2026-09-21 00:36:16 +02:00
jschoubben e140ed5d0b An operator key, a second seal on every own secret, and the vault keeps the export
novox/hq ADR 0085, amended: the mesh's root secrets — the store's superuser,
the broker's administrator, every secret a module holds for itself — were
sealed to a node key and nothing else, so a lost node took them with it.
Now the mesh records an operator's public sealing key and seals every own
secret to it as well, minted or accepted. The private half is written once
by `operator key new` to a file the operator keeps off the mesh; the mesh
holds one more blob per secret that it cannot open.

`secret recover` opens a secret with that key, to a 0600 file, from the
store or from an export; `secret export` writes every operator-sealed copy
as ciphertext. A module that `keeps` (the vault) is handed that export as a
declared file on its own disk, so recovery survives the store.

Secrets made before the key exists have no operator copy and are said so —
the plaintext was discarded — until each is issued again.
2026-09-20 23:56:49 +02:00
jschoubben c3b88b9148 Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:40:40 +02:00
jschoubben 3ae7b88c6d A catalogue asks for what it was not there to hear
Its event queue is durable, so a running catalogue misses nothing. What it cannot
have is what was announced before it first ran — and on a fresh mesh that is never
arbitrary: the shared base, the store the catalogue runs on, and the catalogue
itself are each necessarily built BEFORE a catalogue exists to hear about them.
The graph's foundation is the part it never sees.

So it says it is catching up, and the control plane re-announces what it
recorded, oldest first, marked as a replay. Oldest first because a graph is built
in the order things happened: registering a module that stands on a base before
the base would point an edge at a version nothing has seen, and the shape of a
fresh mesh guarantees the base is both first and the one that was missed.

The replayer hands announcements back rather than publishing them, because the
wire belongs to the link package and a replay building its own events could drift
from what the builder emits — the one thing it must match exactly, since the
catalogue has a single handler for both.

Its own queue and its own consumer: two consumers on one queue split its
messages, and a catch-up request going to whichever half was not listening is a
gap that looks like a working mesh.

Toward novox/hq 04-ISSUES/050.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 01:26:58 +02:00
jschoubben 2b82872ac3 A build keeps what it was told
The builder announces a build with the resolved manifest, the path inside the
repository, and every artifact it stood on. The control plane received all of it
and kept none of it.

That was survivable while the catalogue heard the same announcement directly. It
stops being survivable the moment the catalogue was not there to hear it — which
on a fresh mesh is always, and always for the same modules: the shared base, the
store the catalogue runs on, and the catalogue itself are each necessarily built
BEFORE the catalogue exists to hear about them. The graph's foundation is the
part the graph never sees.

Replaying those builds needs what they said, not a summary. Without the manifest
there are no requires/provides edges; without `against` there are no build edges,
which are the ones that answer "a base moved, what must be rebuilt". A replay
carrying neither would restore the module list and leave the question the
catalogue exists for still wrong, while looking fixed.

Kept null rather than empty where a build predates this, so a replay can say it
is holding nothing instead of inventing an empty declaration for a module that
certainly had one. And `built_against`, not `built_on`: that column exists and
means the machine, which is a different fact about a different subject.

Toward novox/hq 04-ISSUES/050.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 01:22:34 +02:00
jschoubben 25d2fe1308 Only a thing built and never run is unassignable
The check read "declares no resources" as "runs nowhere", and those are not
the same. The private network declares no resources either — the control plane
computes them when it composes a machine's declaration — and it is assigned to
every machine that has to reach another one. Refusing it stopped a four-machine
bed at its first assignment.

The signal is narrower: it builds an artifact and places nothing. Made a
function of its own, because a judgement with a wrong answer this expensive
should be testable without a database — nothing guarded it, which is how it
shipped.
2026-09-14 17:32:43 +02:00
jschoubben 95a9da8bdb A module that puts nothing on a machine cannot be assigned to one
The image every module in the scripted toolchain is compiled on top of is
registered as a module so the mesh can build, version and depend on it. It is
not one: nothing about it belongs on a machine. Assigning it succeeded, the
machine was sent a declaration containing nothing of it, and everything
reported success — the operator had said run this here and the mesh had agreed
to something it cannot do.

A module whose resources are worked out per node is asked about separately, so
it stays assignable, which is the point of it.
2026-09-14 11:08:32 +02:00
jschoubben cfe2816495 A module names the module its build stands on, not a copy of it
A fingerprint written into a recipe names one particular copy of the base — the
copy on whichever machine the person typing it was using. On any other mesh that
copy has never existed, so the build stops on its first line with a message
about an image nobody can look up. Three modules in the catalogue were in
exactly that state, and the line each of them replaced was equally dead.

A module now names the module and artifact instead, and the mesh answers with
what it holds. The builder is still a thing that clones, builds and answers: the
answer travels with the question, because only the mesh knows what it has.

A base the mesh has not built is refused before anything is built, naming which
module has to exist first.
2026-09-13 23:53:22 +02:00
jschoubben 588aa424e2 The mesh acts on what the catalogue decided a build meant
The builder says what it built and the catalogue decides whether that was an
upgrade. Only the control plane knows which machines run the thing, so it is
the one that acts — and what it does is a choice somebody recorded, not a
behaviour compiled in: record that they are behind, or send it, one machine at
a time or together.

Recording is the absence of an action rather than a second path: a machine not
running what the mesh would send it is already something the mesh reports.

Defaulted to recording. A mesh that rolls out everything it builds the moment
it builds it is reasonable to want and a bad thing to arrive by default — the
first module to inherit it would be the control plane, upgrading itself out
from under the push applying it.
2026-09-13 01:55:07 +02:00
jschoubben f151de103f Build a module from a repository and a path within it
The builder cloned a repository and read the manifest at its root, which means one
repository per module. Nothing we have is shaped that way, so the builder could be
asked to build nothing that exists (novox/hq ADR 0069).

The path travels the whole way — named when asking, carried in the request, used
to read the manifest and as the context everything is produced from, echoed back
in the result, and recorded as part of where a module came from. Without that last
part the mesh could notice a module was behind its source and then be unable to
rebuild it, which is the worst of both.

A path climbing out of the clone is refused: a machine whose job is building other
people's repositories must not read whatever else is on its disk.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 16:45:50 +02:00
jschoubben c4030947b0 The routing record is 0066, not 0056
0056 is 'the authority is the control plane, not a database'. A citation
pointing at the wrong decision is worse than none: it reads as corroboration.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:09:56 +02:00
jschoubben f5f860fd2b inventory: forgetting a module says what goes with it, and refuses until told
`module forget` cascaded. The settings, the module's own secrets and the ports the mesh
chose all name the module by a foreign key that cascades, so removing the row took all
three and reported "forgotten" — an action succeeding into a state its own verify would
reject (novox/hq 04-ISSUES/017). A sealed secret is not recoverable afterwards, because
the mesh discarded the plaintext when it made it.

It now reads what it would destroy, names each thing one at a time, and refuses.
`--and-what-it-holds` is how somebody says they mean it, and the removal then reports
what went — this being the only record that any of it ever existed.

Reported as "operator settings do not persist, because re-registering a module
cascade-deletes them". Half of that is wrong, and the test now says so out loud: the
upsert is on the name, so `module add` at a new version leaves the settings, the secrets
and the ports exactly where they were. The command that destroyed them was `forget`, and
a wrong belief about which command destroys data is expensive in both directions — it
sends people looking for a fault that is not there, and leaves the real one unexamined.

Checked by internal/inventory/forget_test.go, which writes all three, re-registers the
module at a new version, reads them back, and only then tries to forget it.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:11:15 +02:00
jschoubben 232862315c catalogue: compose a route's name from a label and its node's domain, and resolve it in-mesh
A public route used to carry its whole hostname as a literal in the module
manifest, so running the same catalogue against a different domain meant
overriding that literal on every routed module, per node. The mesh was, in
effect, holding a map of names to services: the one thing it should never hold,
because the subdomain is the operator's choice and the domain is the node's.

Compose instead. A route contribution carries a `label` (the subdomain); a node
carries its `public_domain` as node-level configuration; the mesh joins
`<label>.<public-domain>` and grants exactly that, interpreting neither half.
Held as a node property beside the node's other node-level facts (endpoint,
site, overlay address), not in a module's settings — the ADR calls it
node-level, and the settings table is keyed per module.

Additive, so an unmigrated catalogue keeps working: a contribution that still
carries a full `name` and no `label` passes through unchanged, and the catalogue
can migrate module by module. A labelled contribution on a node with no public
domain composes nothing, reading downstream as a route that named no host.

And propagate: each granted route name is published into internal resolution
mesh-wide, mapped to the node that serves it, alongside the `<node>.internal`
names every container already gets. So a container — and an internal ACME
validator, which cannot complete a challenge for a name it cannot reach —
resolves a routed name to the proxy that serves it. Name-agnostic throughout:
the mesh propagates whatever names it was told to serve and knows nothing about
what they mean.

novox/hq 02-DECISIONS/0056

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 23:27:35 +02:00
jschoubben 1b63e21c0f Caught up is an equality, not an ordering
The report carries the digest of the declaration it applied (mesh-host
8211d8b), and the mesh stores it beside the outcome. `reported` rows in
the status JSON now say `current`: whether the machine's last word
names the declaration last sent.

Not derivable from the timestamps beside it, which is why they were
not enough: an apply begun under the previous declaration reports
after the next send — newer, and still about the old words. The lab
lost exactly that race between one test's closing push and the next
test's opening one.

Empty digests — every host from before reports carried one — read as
not current, which errs toward waiting rather than toward asserting on
files that are not there yet.
2026-09-02 00:02:44 +02:00
jschoubben c0107b8572 Status says when each machine last reported, beside when it was sent
"Not waiting" says the declaration is current, not that the machine
finished applying it: the sent digest is recorded at send. So a test
that pushed, saw waiting clear, and asked the machine what it was
running found containers that did not exist yet — the certificate fix
made compositions stable, and the settling that used to fail first had
been hiding the gap behind it.

The mesh already held the missing half: every machine's last report,
with its time. It just was not in the JSON. `reported` now sets each
machine's last word beside when the current declaration went to it, and
"has it caught up" becomes a comparison of two timestamps the mesh
recorded itself — a report newer than the send means the machine acted
on what was sent; older means it is still working, which waiting alone
cannot distinguish.
2026-09-01 23:02:26 +02:00
jschoubben b70f0d626a What review found in the port machinery, fixed
Three faults, one file split. All from reading, all verified to bite.

Unassign now releases the module's ports. ReleasePorts existed, said
"for when it is unassigned" in its own comment, and was called by
nothing — so a fixed port stayed claimed in the name of a module that
was gone, and the next module needing it was refused by a ghost.
Kept-once-chosen is a promise about a module that is still here.

MachineSide reads addressed mappings. "127.0.0.1:8080:80" was split at
the first colon, "127.0.0.1" failed to parse as a port, and the mapping
was silently skipped — putting the filter back on the declared port,
the exact fault the function was written to end. The machine side is
the second-from-last part, which is the reading the host already
applies, and the substrate bundle writes that shape today.

An allocation race answers in the mesh's words. Two concurrent picks of
the same port used to surface as a Postgres constraint violation,
verbatim. The table has two keys, so the collision is one of two facts:
the racer was this same assignment — then its answer is the answer,
kept-once-chosen does not care who chose — or another module took the
machine port, and an unfixed pick is simply made again against the
moved free list. A fixed port that lost the race is refused by name.
Told apart by re-reading the row, not by the constraint's name, so this
does not couple to the migration's spelling.

And the artifact-store cycle tests moved to bootstrap_cycle_test.go;
machineside_test.go had quietly become three subjects.
2026-09-01 21:54:05 +02:00
jschoubben 4d6ec5b10c One object store, not two
object-store.json and minio.json described the same thing: same image,
same provision at the same scope, same provisioner. Not two
implementations a person could choose between — one module written
twice. Assigning both to a node would have collided on `s3-bucket`.

It exists because it was written first, to pair with photos.json for the
README's worked edge, and minio.json was the fuller version of the same
module written later. Nobody removed the first.

The pair test keeps its point and now reads the surviving one. Checked
across the rest: this was the only duplicate.
2026-09-01 21:06:17 +02:00
jschoubben 41f7c51032 Assign around what a machine already holds
The other half of ADR 0038, and what 04-ISSUES/028 was actually about.
A module can now avoid colliding with another module; until this it
could not avoid colliding with the mesh itself.

The substrate is not a module. A node raises it from the bundle it
carries before any mesh exists, so the control plane had never heard of
the store, the broker, or its own container — and handed a database
module 5432, which the store already had.

So the machine says. The host records what each resource binds,
distinguishing what it carried from what the mesh sent — a distinction
that already existed so the two never remove each other — and reports
the carried ones. The node states and this context writes, which is the
shape of every message between them.

What the declaration binds, not what is open. A machine's open ports are
a moving target, and assigning around them would mean a port that was
free when it was asked for and taken when it was used.

Replaced whole each time rather than merged: a machine that gave a port
back must be believed about that too, and a set that only grows keeps a
port reserved for something no longer there.

Tested against a real database, and the tests bite — removing the check
hands the module 20000, which the machine had said it holds.
2026-09-01 18:32:38 +02:00
jschoubben 1f5b70a995 The mesh assigns the port, and a module says it once
novox/hq ADR 0038. A module cannot choose a port: it is written once and
assigned anywhere, so any number it picks is a guess about a machine it
has never seen. A database module met the mesh's own store on 5432 and
was told, by a container runtime three layers down, that the port was
already allocated.

The number used to appear three times in every module — the rule set,
what a consumer is told, and what the runtime publishes — agreeing only
because one person wrote all three. Now it appears once, in `listens`,
and the other two are derived: the container publishes `20000:5432`, the
consumer is told 20000, and the rule set opens 20000.

An assignment is made once and kept, as a credential is. A port that
moved on every declaration would restart both ends each time and hand a
consumer a number that was true when it was read.

Ports the protocol fixes — mail on 25, submission on 587, DNS on 53 —
say so, and are then claims: one holder per machine, and the second is
refused by name at assignment. That is the mechanism the mesh already
has for what is singular on a machine, pointed at ports.

A mapping written the long way is left exactly as it is. Some things
must be pinned by hand, and quietly overruling somebody who wrote both
halves would be worse than not offering the short form.

Still open, and known: the substrate is not a module, so the mesh has
never heard of its own store and cannot yet assign around it. That is
what 028 will still be about after this.
2026-09-01 17:52:53 +02:00
jschoubben 0af3ea1acf A consumer is a module on a machine, not a machine
novox/hq 04-ISSUES/022. A credential was keyed by provision, consumer
node and provider node, so "who is asking" was answered by naming a
host. The node this mesh exists to take over runs eight modules against
one database server.

The symptom had two halves and only one was loud. The provider refused,
naming the modules and explaining they would share one credential, which
reads as a decision rather than a limit. The consumer did not refuse: it
resolved cleanly, wrote one module's credential file and left the others
absent — a service that starts and cannot authenticate, with nothing
saying why. That is 021 again on a different axis.

Three modules wanting one database produced one need, carrying whichever
module mentioned it first, because the resolution walk is a work-list
over names. The fan-out now happens in one place, after the walk. The
record path already did this correctly and said why: a consumer here is
a module on a machine. It is the same rule.

Downstream: the secret's key gains the consuming module, the grant file
is named after both halves, needs are matched by provision and module
rather than provision alone, and the provisioners name the role and the
access key after the module. The refusal in ContributionsTo is gone
because there is nothing left to refuse.

Worth stating plainly: without that refusal, gitea's login would have
opened keycloak's database. From the provisioner's side it created
exactly what it was asked to create.

Existing secrets are discarded rather than backfilled. They cannot say
which module they were for, and a secret is remade and delivered to both
ends on the next push — so this costs one rotation and invents nothing.

Also guards the role name against PostgreSQL's 63-byte truncation, which
is a notice rather than an error and would reintroduce exactly this
collision at a length nobody tests.

Three faults injected — the fan-out removed, needs matched by name
alone, the grant file named after the machine — each caught.
2026-09-01 02:40:09 +02:00
jschoubben ee84b624b1 A provision names the engine, because a consumer is coupled to one
Provisions were named after roles: provides "database", requires
"database". Nothing distinguished engines, so a module written against
PostgreSQL could be matched to a provider of SQL Server, resolve as
satisfied, deploy, and fail on its first query — with nothing
connecting that error back to a match made elsewhere by something that
believed it had done its job.

The failure is in the direction that hides. Refusing on ambiguity
exists precisely so this does not happen, and the generic name walked
around it: with one provider of each name nothing is ambiguous, so
nothing is asked.

How it got in: every resolver test had exactly one provider per name,
so no mismatch was expressible and none was caught. The fixtures agreed
with the design — the same fault as the imagined test output in
04-ISSUES/005, at the level of a name.

Refused rather than documented, because the old naming *was* the
documented convention. Providing database/db/sql/sql-database is now a
parse error naming what to write instead.

The rule is about coupling, not specificity everywhere: route and
resolver stay role-named, because a consumer genuinely cannot tell
which proxy answered. novox/hq ADR 0027.
2026-08-31 17:12:46 +02:00
jschoubben 5a28434ba8 "Behind" means not running what the mesh would send
It meant "failed or refused". So a machine that applied cleanly and whose
declaration has since changed was not behind — and novox/hq ADR 0010's
question, did my change go out?, was answerable exactly for the machines that
broke. For every machine that worked, the answer was silence whether the change
had gone out or not, which is the thing replacing a pipeline was supposed not
to cost.

The mesh now records a digest of what it last sent each machine. A digest
rather than the declaration: it can compute what a machine should be at any
moment, and keeping a copy would be a second account of it able to disagree
with the first. What cannot be recomputed is what was actually sent.

Recorded after the send, not before — a digest kept for something that failed
to send would make the machine look current for a declaration it never
received.

Never told stays separate from out of date. The remedy is the same push and the
situations are not alike: nobody has ever asked that machine to be anything.
And a machine the mesh could not work out is not reported as waiting, because
saying so would invent a comparison — that is `plan`'s answer to give.

`status` says it and `push --behind` sends it, or the flag would know something
the person reading the status does not.
2026-08-31 05:17:21 +02:00
jschoubben dfca21fa55 Keep what a machine said about itself, not just the yes
novox/hq ADR 0009: a capability's presence gates an assignment and its detail
carries a value — seat: card1-DP-1, an architecture, an amount of memory. So
'can this run here' and 'what should it be configured as' are one fact read two
ways, and the mesh was keeping the first read and discarding the second.

The reason an absent capability is absent went the same way, which is the case
a person most needs: 'this machine has no container runtime' is the answer and
'docker is not installed' is why, and only the machine knows why.

`node show` says it back. Never reported and reported nothing stay different
things there — one machine has not run the host, the other ran it and can do
nothing, and those send a person to different places.
2026-08-31 04:53:45 +02:00