Commit Graph
251 Commits
Author SHA1 Message Date
jschoubben 7180a273a2 The enrolment user is per token, and it has an inbox
Design 25 §6 says an enrolling node subscribes the inbox its own token derives.
It had none: `sub` was empty, so a node would publish its request and wait out
its timeout against a mesh that had answered — the handshake could not have
completed.

And there was one shared `enrolment` user, which cannot carry that inbox at all:
a permission belongs to a user, so an inbox per token means a user per token.
Named after the node, which **is** the token's id — a token is issued for a node
record, the mesh holds one live claim per record, and the node's name is the one
identifier both sides have before anything else is agreed. It is also exactly
what the other transport does, where the account is named after the node and the
secret is its password.

A nameless enrolment user is now refused rather than composed into
`_INBOX.enrol..>`: an empty subject token, and worse, one every nameless
enrolment user would share — which is one machine able to read the credentials
sealed to another.

Still to wire: something that composes one of these per live token. Nothing
composes enrolment users yet, on either bus — on the old one the account is made
imperatively through the broker's management API when a token is issued, and
here there is no management API, so issuing a token has to recompose the server's
configuration. That is the remaining half of enrolment on the new bus.
2026-09-27 01:31:34 +02:00
jschoubben e65b3950cc A node's declaration consumer, which only the mesh can make
A host's account reaches no part of the JetStream API — correctly, because the
controller is the only writer of consumer definitions — so the object a node
reads its declarations through has to be waiting before the host binds to it,
and nothing created one. Named after the node, because the node's own ack grant
is `$JS.ACK.NODES.<node>.>` and a consumer named anything else is one the host
cannot acknowledge a delivery from.

No max-deliver, and a five-minute ack wait: a declaration is settled only after
the node has applied it and reported, which is minutes on a machine pulling
images, and the stream holds exactly one message per node — so there is nothing
to dead-letter, only one message to redeliver for as long as that node is away.

Asserted on start as well as created at enrolment, for the reason the streams
are: a mesh raised from a restored backup has node records and no consumers, and
a node whose consumer is missing hears nothing while everything else about it
looks correct.

The test sets the consumer against the grant the node actually gets, because
each of the three ways of getting it wrong is silent: a wrong name cannot ack, a
wrong filter reads another node's declarations, and a pull consumer is one a host
has no authority to bind.
2026-09-27 01:24:46 +02:00
jschoubben 88bef39952 The consume side on NATS, and the window held by the server
The other implementation behind the seam, so the store-window guarantee now has
both: one loop, one message at a time, the same window deciding. What differs is
where a held message lives, and that is the whole point of the move — the AMQP
side keeps an unacknowledged delivery in this process, bounded by the prefetch
and lost if the controller stops; this keeps eight bytes saying when the window
opened, and the message stays the server's.

Checked against a running server, seven claims that reasoning cannot answer: a
report is heard and leaves the work queue; one the store cannot take is naked
with a delay, stays in the stream, and is recorded when the store returns; one
about a superseded declaration is settled without being acted on; one the store
never takes is let go once the bound passes; a heartbeat is heard and nothing is
persisted; and the enrolment answer reaches the address the request carried in
its payload — the test design 25 §2 asks for, so the reason for that field
cannot quietly become folklore.

Three things the wiring forced into the open:

**The controller could not have consumed a module event.** Its permissions
granted no event subject to subscribe and no ack subject on the events stream,
so every announcement would have been redelivered for ever, refused by the list
it already had. Both narrow: each followed subject named, not `mesh.mod.*.>`.

**The controller's consumers are not derived.** It files no manifest, so its
authority cannot come from a declaration that does not exist; they sit beside the
mesh's own streams and are asserted the same way. No max-deliver on CONTROL —
the window's bound is the controller's, and a server that dead-lettered first
would discard the push the stream exists to protect.

**Channels, not callbacks.** The library would run a handler on its own
goroutine, and the window's bookkeeping is unlocked because the AMQP loop never
had two.
2026-09-27 00:53:33 +02:00
jschoubben 06cf3c04e5 The consume side behind a seam, and the window wiring into the loop
The outbound half went behind `Bus` and the transport stopped reaching its
callers; this is the other half, and the larger one. Every handler took
`amqp.Delivery`, so the serving loop could not move to another bus without
moving enrolment, reports, builds, upgrades and catch-up with it in one breath.

`Control` states one message in the mesh's words — took it, dropped it, or held
it for the store — and `Inbound` is where messages come from. The AMQP
implementation is today's loop moved rather than changed: same queues, same
prefetch, same holding, because the mesh is running on it and a bus nothing
speaks yet is no reason to alter the one every node is on.

The window (window.go) is now what decides, instead of the conditions that were
inlined in the loop. Two things that surfaced in the wiring:

**Supersession is asked before the store, not after.** A report about a
declaration the mesh has moved past would otherwise wait out a restarting store
to be written and then overwrite what the node is doing now.

**Half of a report is not about a declaration, and that half is never stale.**
What the machine *is* — the tunnel it took over, the ports its own bundle
holds, what an adopted node found, a node moving its overlay key — reaches the
mesh on a report and nowhere else. A rekey set aside as stale is a node whose
overlay key never moves, and no retry is coming, because the node said it once.
So staleness is asked only of a report that is purely an apply's account.

The one thing holding-in-memory can do that holding-in-the-server cannot is
named rather than hidden: `About` sets aside a held message when a newer one
about the same thing arrives, and the bus being built ignores it because the
digest answers the same question.
2026-09-27 00:44:16 +02:00
jschoubben 7a8a19b11b A person's account (step 4.4, the account half)
Design 25 §7. A person is not a module and holds no seat: nothing is
addressed to them, nothing is delivered to them, and they have no durable
consumer. What they have is permission to ask, as a list of tools or `*`
for an administrator.

Four properties the tests hold it to, each of which is a way of being
wrong that would not announce itself: a person reaches nothing but tools,
so one cannot claim a module said something; no ack subject, because
authority over a consumer that does not exist is authority nobody would
audit; no allow_responses, because a person who can answer a request is
impersonating a module on a bus where anyone may serve a tool; and two
people do not share an inbox.
2026-09-27 00:17:52 +02:00
jschoubben ce6ac057f6 gofmt the probe 2026-09-27 00:14:47 +02:00
jschoubben abce68fdf0 Verify that a reply address does not survive a stream
Design 25 §2 builds enrolment around carrying the reply subject in the
payload, because a JetStream consumer claims the transport Reply field for
its own ack address. The whole handshake rests on it, so it is checked:
the caller asked for _INBOX.LCr3M83q... and the consumer saw
$JS.ACK.PROBE.probe_consumer... The design was right, and the workaround
is necessary rather than defensive.

Worth having as a test rather than a note: if a future server version
stopped doing this, enrolment would keep working and the reason for the
payload field would quietly become folklore.
2026-09-27 00:14:32 +02:00
jschoubben cc019908c2 Asking a tool goes through the seam, and loses two problems
Step 3.4. On the new bus there is no reply queue to declare and no
correlation to check: each account is granted one inbox prefix and no
other, so an answer cannot reach the wrong asker. That settles a cost
build.go records having paid — on a shared reply exchange every asker saw
every result, which is why the correlation was checked rather than assumed.

And a tool nobody serves says so at once rather than after the whole wait.
The difference between "that module is down" and "that tool is slow" is the
first thing a person asking wants, and both tests are against a real server
because both are claims about what the server does, not about this code.

RequestBuild stays as it is, and is a different shape on the new bus rather
than the same one: a build takes minutes, so it is work submitted to a
queue with the outcome returning to a reply subject the request carries —
the pattern design 25 §2 already sets for anything crossing a stream. It
touches the builder too, so it goes with that conversion.
2026-09-27 00:11:30 +02:00
jschoubben d0a9abcb1c The store window as a decision, and the problem moving it to the server
introduces

Step 3.4, the consume side's hard part. The guarantee (ADR 0083) is that a
push the controller cannot record because its store is restarting is held
and retried — never dropped, never falsely acknowledged. Keeping the
delivery unacknowledged in memory becomes a nak with a delay: the server
holds it, the controller keeps no list of parked messages, and a controller
that restarts mid-window loses nothing it was holding.

That is a plain win and it introduces one problem. Holding in memory let
the controller drop an older report when a newer one for the same node
arrived, "because acting on it after the newer would undo the newer". A
naked message is the server's and comes back whatever happened meanwhile,
so the older report is redelivered after the newer was applied.

The answer was already in the message. A report carries `Declared`, the
digest of the declaration it is about, which exists because an earlier
attempt to order reports by time lost the race it invited. So supersession
stops being something the controller remembers and becomes something it
checks — the same shape as a node refusing a superseded declaration by
sequence (issue 107): ordering settled by what a message says, not by when
it arrived.

Pure, so the guarantee is testable without a bus, a store or a clock. Nine
tests, including that staleness is decided before the store is waited on —
a redelivery that lost its race must not hold a slot a current message
needs.
2026-09-27 00:05:52 +02:00
jschoubben 6c12780abe Describe the bus on its own terms
Comments framed the new bus by what it replaces — a comparison in almost
every explanation, which reads as though NATS were a variant of the old
thing rather than the mesh's nervous system. Removed throughout, and
OverAMQP becomes OverCurrent: the seam's two sides are the bus the mesh
runs on today and the one being built, not two protocols.

What remains is the client library's own package name, which is its name.
2026-09-26 23:51:00 +02:00
jschoubben 92d87b0082 gofmt bus.go — import grouping 2026-09-26 23:47:26 +02:00
jschoubben 2fad32767e The controller's outbound link behind a seam, with both transports
Step 3.4, first half. Every one of these took an *amqp.Channel, so the
transport reached every caller and swapping it meant touching all of them.
The seam turned out to be small — the controller sends exactly two kinds of
message that expect no answer — which is the same measurement that said
this bus could be replaced at all.

Bus is stated in the mesh's words, not a transport's: PublishEvent and
PublishDeclaration. Two implementations, both shipping, because steps 1 to
4 leave every node on AMQP and the NATS one is selected at the rollout.
Both ship is also what makes them comparable: one conformance fixture holds
both to the same envelope, and the NATS one is checked against a real
server reading back from the stream rather than from the code that wrote it.

Still on *amqp.Channel: RequestBuild and Ask, which carry reply-queue
machinery, and the whole consume side — the control loop, enrolment, serve.
2026-09-26 23:47:15 +02:00
jschoubben 1801f1178e Hold the Go emitter to the shared fixtures
Every required header set, each value in the pinned shape, and the subject
derived the same way. Read from the sdk's conformance directory by sibling
path, never copied.
2026-09-26 23:40:59 +02:00
jschoubben bf82805cfd A manifest holds no subject, checked across all 72
Design 29 §1's load-bearing rule: a module names its events, tools and
seats locally and the mesh derives the subject, so reorganising the subject
space leaves every manifest correct. It held by construction, and a rule
held by construction is one a later field breaks quietly.
2026-09-26 23:34:42 +02:00
jschoubben 173c8c7c21 Rename the mesh's seats to mesh-*, keeping their interfaces
novox/hq ADR 0118: the prefix is the reservation rule, so a module
declaring any mesh-* name is refused and there is no reserved-names list to
drift. Ten seats renamed in the table, the manifests that claim them, the
controller's own shipped manifests, and the tests.

Not the migration 0118 expected: a holding is derived at resolution from
manifests and never stored, so nothing recorded points at an old name. A
kept rename table tells a manifest written against one what it became —
kept rather than retired, because a module lives in its own repository and
may be registered long after the catalogue stopped using it.

**A seat is not the interface it delivers.** The git seat became mesh-git
and the git provision did not; likewise the package registry. A blanket
replace renamed both, and the failure read "the package registry is served
on <nil>", which does not say "you renamed an interface". A test now pins
every seat against what it delivers, and that neither name is also the
other.
2026-09-26 23:07:09 +02:00
jschoubben aa74bd86ca Derive a seat's stream and a module's consumer, and wire JetStream
Task 3.9's other half and 1.4's missing client. The derivation is pure and
unit-tested; only "does the server accept this" needs one running, behind
MESH_TEST_NATS so the ordinary suite stays offline.

A seat's work queue is created at registration, not assignment, so work
queues until a holder appears — a stream created at assignment would make
"the holder is not here yet" mean "your messages are gone". Named after the
seat, because the holder can change and the queued work must not care.

A holder's worker uses a queue group even though the seat guarantees one
holder: the seat is authority, the queue group is delivery, and tying them
together means the day somebody allows two holders every message is
processed twice with nothing reporting it.

One consumer per module carrying every filter, because its ack permission is
derived from its name.

And a real bug the live server caught: a durable name may not contain a dot,
but an ack subject is $JS.ACK.<stream>.<consumer>, so the single string that
read correctly inside the permission was rejected as a consumer name. Split
in two, beside the permission that has to match. Unfixed, the symptom would
have been every message redelivered forever with a permission list that
looks right — which is the failure design 25 §4 warns about.
2026-09-26 22:28:42 +02:00
jschoubben 7232d6df4b A module may declare its own seats (novox/hq ADR 0118)
The manifest carries seats and uses; registration refuses a mesh-* name, a
duplicate declarer, an undeclared uses or claim, a seat with no protocol, a
scope mismatch, and a holder that does not answer what its seat promises.

The parser stops judging unknown claim names, because it cannot: another
module may declare that seat, and one manifest cannot tell. The test that
encoded the old rule is rewritten to assert the refusal at registration, and
a new one pins the case the parser could not have distinguished.

Tools are declared for the first time, under their own key — serves already
means a provision's facts.
2026-09-26 22:17:17 +02:00
jschoubben 112cbd294d Per-subject caps on EVENTS, and a comment corrected against the server
NATS refuses overlapping streams rather than double-storing, which is the
opposite of what the Overlaps comment claimed. The check still earns its
place — it names both streams at composition rather than one at apply — and
the refusal is what rules out a shared stream beside per-module ones.
2026-09-26 21:49:55 +02:00
jschoubben 553814b6eb The mesh bus seat is one per mesh, and the amqp broker does not contend
Step 2.3 of novox/hq ADR 0116. The refusal is the resolver's existing one;
these pin it for this seat, including that a different bus implementation is
refused for the same reason — the property that lets the bus be replaced.
2026-09-26 21:40:03 +02:00
jschoubben eeb0560fc2 The mesh-broker seat delivers mesh-bus (novox/hq ADR 0120) 2026-09-26 21:17:21 +02:00
jschoubben 1f36787d75 The mesh's four streams, asserted on every start
Task 1.4. The foundation set only — a seat's streams come at registration
and a module's consumers at assignment, neither of which has happened at
genesis (ADR 0118).

Asserted rather than created: a stream that was deleted, or a mesh raised
from a backup, must converge rather than run without the guarantee its
messages assume.

Two things the definitions have to get right, both tested:
- CONTROL names its subjects instead of taking mesh.control.>, because
  heartbeats live under that prefix and a stream of them competes for
  retention with the messages that matter
- EVENTS filters on the event token, which is why that token exists; a
  filter over a module's whole namespace would persist every tool call

Overlapping filters are refused where the set is written: NATS accepts two
streams matching one subject and stores the message twice under two
retentions, which nothing reports.

Adds nats.go as a dependency; it pulled golang.org/x/* forward. Full suite
green.
2026-09-26 21:02:18 +02:00
jschoubben c753f9d5c0 Compose the bus's accounts instead of calling a management API
Task 1.3 of novox/hq ADR 0116. On AMQP an account was an HTTP call; on NATS
it is text the controller composes and the server reloads (ADR 0106). Pure,
so the mesh's whole authority model is testable as strings.

NATS closes a gap management.go recorded rather than hid: LavinMQ has no
topic permissions, so an emitter was granted the events exchange whole and
ADR 0042's origin reservation was "stamped by the sdk, not enforced here".
Per-subject permissions make it the server's refusal.

Two things found by composing a real file rather than reading the design:

- a scoped inbox leaves a responder unable to reply, because the answer goes
  to the caller's inbox. allow_responses is the answer — one reply to the
  subject of a message actually received — and only principals that serve
  are granted it. Recorded in design 25 §4.
- composition must be deterministic: the module's entrypoint reloads on the
  file's digest, so an order-dependent composer would reload the whole bus
  on every controller restart. Covered by a test.

The golden fixture is the exact text `nats-server -t` accepts, so the syntax
is the server's rather than one we invented.
2026-09-26 20:58:49 +02:00
jschoubben fb87f9f7d6 The mesh-broker seat delivers nothing
The bus is the only broker (novox/hq ADR 0117): messaging is subjects on it,
scoped by a module's own emits/consumes, not a server handed out as a
provision. So the seat joins mesh-controller and the-catalogue in delivering
no interface. Full suite green.
2026-09-26 19:34:14 +02:00
jschoubben fda47558d5 A collision is two places, not two spellings
checkResources compared paths as written, so ${dir:state}/server.env —
the same characters in every module, a different directory in each —
refused the first two placed modules that met. Paths are placed before
they are compared, under the default root, which keeps every real
collision: distinct modules' places are distinct under any one root,
and a module stating another's placed root is caught because a pathless
directory now owns its placed path in the comparison too.
2026-09-26 18:25:28 +02:00
jschoubben 2e3b13c0f8 The assignment's own root is a place, and the manifest's maps are placed
Slice two of ADR 0112. A pathless directory saying place "." is the
assignment's one directory, <root>/<module> — to-be 27's shape — and
place never reaches the host, which parses strictly. The maps naming
where bindings, credentials and contributions land (binds, secrets,
own-secrets, receives, grants) fill against the placed directories at
composition, into fresh maps and a fresh module slice, because one
resolution composes for many nodes. The five absolute-path checks on
those maps accept a placed reference — resolution makes it absolute
before anything reads it — while certificate, operator-keeps and
accesses paths stay absolute-only: those are the operator's or another
vocabulary's. unknownDirRefs scans the maps too, and validates place
itself: only on a directory, only ".", never beside a stated path.

Found by the foundation tests validating the sibling catalogue: the
first conversion's blanket replace turned /var/lib/gitea/database.json
into ${dir:data}base.json — which resolves to the right path by pure
string concatenation. Production was saved by a coincidence; the
catalogue cleanup that follows spells it ${dir:state}/database.json.
2026-09-26 18:18:13 +02:00
jschoubben d2cbc9dbdc A directory the mesh places: ${dir:<id>} and the pathless directory resource
The first executable slice of ADR 0112 / to-be 27, sized to what the
operator settled tonight: a module definition names no host path for
its own data. A directory resource may omit path; composition resolves
it to <root>/<module>/<id>, the root a node's setting on Rendering with
/var/lib as the default — which reproduces exactly the layout novox
converged to by hand. ${dir:<id>} names the place from a resource's
path, content, mounts, environment and env-files, the same shape as
${bound:…}. A directory that states a path keeps it and still answers
by name — that is the adopted-data placement, mssql its live case.

Resolved in the controller at composition, so the wire format and the
host change not at all; a reference naming no directory refuses at the
manifest and again at composition; nested fills are rebuilt, never
written into the manifest's own maps, because one manifest composes
for many nodes.
2026-09-26 17:51:54 +02:00
jschoubben 7e42380dcd Merge main 2026-09-26 14:29:00 +02:00
jschoubben 6ac9013d6e builder: a clone may offer the forge's credential, through git's own store
A private repository could not be built: the builder clones anonymously,
and had no way to say who it is. It already holds exactly one credential
to exactly the right place — the package-registry binding and its sealed
secret, one gitea user whose password answers npm and git alike — so a
clone now offers that, and nothing new is minted or carried.

Offered, never pushed: the credential is written as a git
credential-store file (0600, in the workspace, never argv) and named
with -c credential.helper, so git itself decides when it applies — only
on an authentication challenge, and only for the URL it was written
for, scheme, host and port included. A public repository clones exactly
as before; a repository on any other host is never shown it. The same
store rides along on an artifact's own context clone, so a private
module with a private context builds too.
2026-09-25 21:47:32 +02:00
jochen 97448194ac Seats are a closed set, a seat's holder answers for what it delivers, and a build source may live on the git seat
Implements novox/hq ADR 0110 and 0111.

The seat set lives in internal/catalogue/seats.go: fourteen seats, each with a scope, what occupying
it delivers, and the record that made it one. A test asserts the count and a decision per entry, so
changing the set means finding the argument, as the host's vocabulary test does. The first set is
every seat already claimed — including the-private-network, which the network module claims from a
manifest composed in this repository's code, not from any module.json — plus npm-package-registry
(ADR 0109) and git (ADR 0111). A test parses every catalogue manifest and this repository's own and
fails on any refused claim, so closing the set refuses nothing in use.

ParseManifest now refuses a claim on a seat the mesh does not define, a seat claimed at another
scope, and a delivering seat claimed by a module that does not provide what it delivers. A
malformed claim is refused once, for being malformed.

Resolution: among several providers of a mesh provision, a pin still wins; then the holder of the
seat that delivers it; then the only provider; otherwise refused as before. ADR 0009's "never
guessed" holds — the seat is the choice made once, mesh-wide, rather than a pin per consumer node.
A provider now carries the module it came from, because a provider is a (node, module) pair and the
pair is what tells a holder from a neighbour on the same machine.

The planner's second pass is now given the first pass's holdings. Without them, a node consuming a
seat-delivered provision was refused there, and a refused node's own claims dropped out of what the
mesh holds — letting a second holder of one of its seats pass unrefused.

`seats [--json]` lists every seat, what it delivers, and each holder, derived from assignments
every time and never stored. Unheld seats are listed. A stored claim outside the set — possible
for a manifest registered before the set closed, since stored manifests are not re-validated — is
shown rather than hidden.

`build --self <owner>/<repo>` builds from a repository on the git seat's holder. The clone URL is
composed at build time from the holder's node and what it serves for git; the recorded source is the
path and the seat (migration 0032), never an address, so a moved forge changes nothing recorded.
Nobody holding the seat refuses self-hosted builds and says so; external URLs are unchanged. An
address passed with --self is refused rather than recorded as a path.

Replaces three foundation tests that defended the builder's carried package binding. The catalogue
removed that binding when the builder began requiring the registry through a real grant, so the
tests were already failing on main; they now assert the builder requires what the npm seat delivers
and carries no copy of its own, and that the forge holds the npm and git seats.

Verified: go vet clean; the whole suite passes against a throwaway Postgres (make postgres), the new
inventory tests included; gofmt clean apart from cmd/mesh-builder/stdout_test.go, which fails on
main too.
2026-09-25 20:48:10 +02:00
jschoubben 20e57c3f51 an image artifact may name its own build context, apart from the module's repository
route-proxy's own Dockerfile documents the shape it has always needed and
never had: 'the proxy source is not vendored here... the build context is
the mesh-controller repository root, and this Dockerfile compiles
./examples/route-proxy from it.' Nothing in the mesh could do that — the
build command clones one repository and builds every artifact from
within it, so route-proxy has never once been built through the pipeline,
consistent with it never having been assigned anywhere. Found attempting
exactly that build tonight: 'stat go.mod: file does not exist', because
the context was mesh-catalog, which does not have one.

An image artifact may now carry a context: {repository, ref}, cloned
fresh alongside the module's own tree. The recipe (Dockerfile) is still
read from the module's own directory, at the module's own commit — only
docker build's own context argument moves. Packaging and source stay
exactly as separate as route-proxy's own comment already said they were,
now for real.
2026-09-25 17:39:27 +02:00
jschoubben 752abaa81d Merge pull request 'A route composes its internal-network alias too, not only its public name' (#60) from feat/route-carries-internal-alias into main 2026-09-25 15:24:01 +00:00
jschoubben af26ed2e07 gitea's ssh port test matched a manifest mistake; resolver test used Names, not Machines
TestTheForgesSshPortIsGivenByTheNumberTheForgeCallsIt exercised a settings
override from '2222' to 222 — but 2222 was never a real port anywhere,
just a mistake in gitea's own manifest (fixed alongside this: listens.port
is now 22, the container's real internal sshd port, matching every other
module's convention, and ports declares 222:22 directly — 222 has always
been the real, fixed public git-ssh port, needing no per-node override).
Split into two tests: the fixed default with no override, and a genuine
override case for a hypothetical node whose predecessor used a different
number, keyed correctly by 22.

TestTheResolverAndWhatAsksItComposeOnOneMachine set Rendering.Names but
FactNodeZones reads Rendering.Machines (novox/hq issue 111 split the two
apart: every name the mesh serves vs. the machines subset) — a loose end
from that merge, not exercised until now. Both are the same map in this
test's scenario, so both fields are set.
2026-09-25 17:07:24 +02:00
jschoubben f996a6707e a route composes its internal-network alias too, not only its public name
Every cutover done on novox tonight (drive, files, files-api, git,
keycloak, umami) dropped the <label>.<node>.internal alias HAL always
paired with the public hostname — found only when the operator tested it
by hand. Not a security boundary (a predecessor proxy served both as a
convenience, reaching a service over the VPN without a public TLS round
trip, not as access control), so restoring it is composing the same
convenience the same way the public name already is: <label> joined to
the node's own private address (r.At), independently of whether a public
domain exists to join the other half to.

composeName's signature changes (publicDomain, internalDomain) but its
shape does not — additive, label-gated, apex-aware, exactly mirroring the
public half it already did. A contribution the mesh writes both names
into is the entire fix; route-adapter and route-proxy pick up internal-
name whenever they're updated to serve it, not before, so this alone
changes nothing about what is live on any node yet.
2026-09-25 16:51:38 +02:00
jschoubben 8fa5443862 contributes: a module's grant carries no value where it contributed several times
ContributionsFrom settled to whichever of a module's several contributions to
one requirement sorted first, arbitrarily — the grant minted for it then
carried that contribution's label and port under a credential the OTHER
contribution's consumer never sees, and collided with that same
contribution's own entry from contributions() besides.

Confirmed live: minio's two route contributions (files-api, files) produced
three entries in route-adapter's received file — files-api twice, once
credentialed and once not, files not credentialed at all. Every
single-contribution module (gitea, keycloak, umami) already mints an unused
credential for `route` too — route never needs one, by its own
documentation — but with exactly one contribution to match there was nothing
to collide with, so it never surfaced.

Where a module contributes more than once, there is no single value to
settle on. The module still asks, still gets its one credential — a pair
credential is not a place for a label or a port anyway — and each named
contribution reaches the provider on its own, unchanged.

No cleanup needed for the secret already minted live for minio+route: the
sealed blob is a random pair credential unrelated to Values, which is
recomputed fresh on every plan/push regardless.
2026-09-24 18:54:35 +02:00
jschoubben f4bcb320fe contributes: a module may answer one requirement several times
A module's contributes was map[string]map[string]any — one JSON object key
per requirement, structurally exactly one contribution to "route" ever.
minio needs two public hostnames (the S3 API and the console), which is
two different contributions to route from one module, and nothing let it
say so.

This is the same shape of problem ADR 0094 solved for secrets (a module
needing several values from one provider that gives one per pair):
contributes now accepts either the ordinary {label, port} object, or an
object of local names to several such objects. Detected per requirement
key by what's inside, since (unlike secrets' string-vs-object split) both
shapes are JSON objects: an ordinary contribution's fields are scalars, the
several-instance shape is local-name -> object. Confirmed against every
module.json in mesh-catalog before relying on that split.

Both route-proxy and the migration-era route-adapter already key generated
routers off the composed hostname (Values["name"]), not the module name,
so two contributions with the same From reach them as two independent
routes with no changes needed on the receiving side.
2026-09-24 18:36:08 +02:00
jschoubben 2277583e99 Tell the resolver the machines, not the names the mesh merely serves
The map the control plane hands a resolution holds both: the machines, and every name
the mesh was told to route to whichever machine serves it. A container's hosts wants all
of it, so a routed name resolves to the proxy. A resolver's zones want only the machines:
told the mesh's suffix is its own it answers authoritatively for everything under it and
forwards none of it, so a routed name with the suffix appended — drive.example.test.internal
— is a name nobody will ever ask for, standing beside the machines and looking as real.

Found composing the resolver's first assignment on a live machine, before pushing it.
hq issue 111.
2026-09-24 01:31:11 +02:00
jschoubben 0d8264ff55 Give the resolver the mesh's suffix as a local domain and a module its machine's address
hal dnsmasq-app conversion, hq 08-connectivity. Converting the resolver from the module it
replaces made it forward what it cannot answer, which is what the predecessor's does, and
that found two things the controller did not say.

A resolver that forwards must not send a mesh name it does not know upstream: the
`node-zones` fact now carries `local=/<suffix>/` beside the wildcards, written here rather
than in the daemon's configuration because the suffix is the mesh's choice and this file is
the one place the mesh writes what it chose. The default lives in one helper now instead of
being spelled in two functions.

The predecessor points the container runtime's `dns` at the machine's own tunnel address —
a container cannot reach the machine's loopback. A module writing that key needs the
address, and `${machine:at}` is the machine's name; a runtime's resolver list cannot be a
name it would need that resolver to look up. So a module may say `${machine:address}`: what
`at` resolves to, read from the same names the hosts file and the wildcards are written
from, absent — and refused — off the network like `at` is.

The `mesh-resolver` and `resolver-data` constants go: nothing provided or consumed either,
the fact and `mesh-addressing` are the mechanism, and a requirement nothing provides is
refused at resolution.

Tests: the catalogue's dnsmasq, resolv-conf and resolved-split-dns manifests are parsed
and composed as a machine would receive them — fixed upstreams, no-resolv, 127.0.0.1, the
machines file, the runtime's key, the pair that decides what a machine asks refused on one
node; and on a real mesh the resolver's machines file is composed with a wildcard per
machine on the network and composed again without one that left, mirroring the hosts fact.
2026-09-24 01:10:15 +02:00
jschoubben 6073e94a4f Merge pull request 'Adopt the predecessor's tunnel in place: its range, its address, its peers (hq ADR 0105)' (#49) from feat/adopt-the-tunnel into main 2026-09-23 22:38:31 +00:00
jschoubben 4566c5c9aa Adopt the tunnel as a mesh fact, refuse a mismatched takeover, and rekey after enrolment
Review of the ADR 0105 build (hq ADR 0105). Four things it got wrong and one
path it lacked:

- A predecessor spoke's tunnel names one peer, the hub, routed the whole
  range; recording refused it and the whole enrolment failed. Range-routed
  peers are skipped now — only the hub's peers are ever carried.
- The range and the carried peers were conditions on the node being adopted,
  so converging the hub would have renumbered the mesh and dropped the peers
  still reaching it. They are facts of the tunnel record now, mode aside; the
  takeover alone is declared to an adopted node. Converging the hub is refused
  while a carried peer has not enrolled, naming it.
- A push composed a takeover for a hub whose address or endpoint disagreed
  with the tunnel, which would have the host stop the found interface and
  raise the mesh's where no peer listens. The graph refuses to compose it,
  naming both and the placement that fixes it.
- The host's account said taken or not; "found down and the mesh's not up"
  read as not taken. Three states now, and an account on every takeover.
- A hub that enrolled before this feature holds a key of its own, and
  re-enrolling would rotate every key the mesh sealed credentials to. A node
  now rekeys in a report, signed with its identity key over the key it
  leaves, the key it takes and the tunnel; the mesh verifies against the live
  key, refuses a stale or foreign proof, records key and tunnel, and moves a
  hub to the tunnel's address. `overlay show` names the path for a hub that
  found no tunnel.

Also: a carried IPv6 peer is routed /128, and identity.ForTest exists so the
link can be tested against a real identity store.
2026-09-24 00:02:07 +02:00
jschoubben 7ef7669c0c Merge pull request 'An address is read from the node's settings where it is used, never recorded with a port (hq issue 102)' (#50) from fix/addresses-follow-the-node into main 2026-09-23 21:55:03 +00:00
jschoubben cdd3638312 The control plane's own manifest says where the node put the store and the broker
Beside each sealed connection genesis wrote, the port this machine put the
seat's holder at: `${seat:mesh-store:5432}` for the three stores,
`${seat:mesh-broker:…}` for the bus, the plain AMQP port and the management API.
Filled from the node's settings when the control plane composes its own
declaration; empty — the sealed value stands — when the mesh has nothing to add.

On its own, after the commit before it is built and running: a control plane
that does not know the placeholder passes it through as the value, and this
manifest is composed by whatever control plane is running when it is pushed.
The reader ignores an unfilled placeholder either way, and a test holds it to
ignoring exactly what this manifest says.

novox/hq 04-ISSUES/102
2026-09-23 23:49:55 +02:00
jschoubben e07b56ce43 An address is read from the node's settings where it is used, never recorded with a port
Three readers did not follow a moved foundation port (novox/hq 04-ISSUES/102),
and each took the control-node down in its own way: the control plane's own
store and broker connections, sealed at genesis with the port inside; and every
build the mesh ever recorded, kept as `<registry>:<port>/<module>/<artifact>@…`.

The control plane cannot open its own sealed connections to move a port, and it
cannot bind the store as a consumer would — a binding mints a credential. So its
settings get a third twin, `NAME_PORT`, read on top of the sealed value by the
store, the broker, the management API and the bus connection, and filled into
its container by a placeholder that names a seat, `${seat:mesh-store:5432}`,
from the node's given or mesh-assigned ports — never the manifest's number, and
empty when the mesh has nothing to add, so what genesis wrote stands. A value
that is still a placeholder is nothing said, aloud: the manifest naming it lands
in the next commit, once every control plane that composes it knows it.

A build is now recorded by digest and path — `artifact-store://<module>/<artifact>@…`
— and the store's address is composed in where a reference is used: the
declaration, the trust file, the bases a build is handed, a replay to the
catalogue. Over the network as `<node>.internal:<port>`; on the store's own node
before any network exists — every genesis push before its "network" step — by
loopback. A reference recorded before this, with an address, is re-routed the
same way when the mesh built it. The trust file and every provider's address
come from one derivation: the node's given port, over the mesh's assignment,
over the manifest's number.

novox/hq 04-ISSUES/102
2026-09-23 23:49:31 +02:00
jschoubben 26690d89f1 The artifact store's seat is one per mesh, and the test says so from the catalogue
Review of the registry work found the seat node-scoped: a second `distribution` on another
machine resolved cleanly there, and only afterwards did the mesh notice `artifact-store`
offered by two nodes, with every consumer elsewhere refusing to choose. A node-scoped
requirement with one candidate installs that candidate, so anything that wanted the store
beside it would have raised a fresh, empty store on the wrong machine first.

The claim is mesh-scoped in mesh-catalog now; this holds the catalogue's manifest to it —
a second store anywhere is refused by name, where it is assigned.
2026-09-23 23:40:53 +02:00
jschoubben 3c836f0abb Adopt the predecessor's tunnel in place: its range, its address, its peers
On an adopted hub the private network takes over the tunnel it finds rather
than running beside it (hq ADR 0105): two tunnels leave the mesh's unreachable
through the provider's filter, so no machine can ever join.

The node presents the found tunnel when it enrols, under the key it took as
its own; the inventory records it (node.tunnel, tunnel_peer — migration 0031)
and the mesh composes from it: the overlay's range is the adopted tunnel's,
the hub is placed at the tunnel's address on the tunnel's port, and every
peer the tunnel had is carried in the hub's peer list as a peer of the
tunnel, not a node of the mesh, until a node enrols with that key — which
then keeps the address the tunnel had for it. A fresh node never gets an
address the tunnel holds. The hub's declaration tells the host which unit to
take over; the host's account of carrying it is recorded and shown.

Every reader of the range follows the setting; nothing stores it. A found
tunnel under another key is recorded and not adopted, so ADR 0100's
non-overlap rule keeps applying where a tunnel is left running beside the
mesh's. A lab bed and test skeleton for "How it is checked" are under lab/.
2026-09-23 23:26:34 +02:00
jschoubben 7d6f37af54 One entry per mapping, under the name the module itself uses
Review found the first pass aliased its answer under both ends of a mapping, which is
wrong wherever two mappings share a number: the alias lands on a key belonging to another
mapping, the later write wins, and the filter and the container then disagree — the very
fault this change exists to close. Two reproduced cases: a module publishing 8080:80 beside
9090:8080 had an explicit setting silently overwritten; a module publishing 4001:80 beside
4002:80 composed both containers onto one machine port, where before it was safely refused.

Now a mapping's answer is filed once, under the end the module names in its listens — the
number the plan, the filter, the openings, the guard and the consumer all ask for — and a
key that names two mappings is refused in the same words as a setting that does.

Also: the guard assertion in the end-to-end test failed open when the resource was absent;
the plan-mirroring helper now says it stands in only where the plan does not allocate, and
the assertions it feeds are narrowed to the port under test.
2026-09-23 02:32:28 +02:00
jschoubben 58644fd282 A node may move a port a module publishes as a mapping's machine side
A module publishing `2222:22` — the machine's own ssh daemon holds 22, so
the module takes 2222 and says so in `listens` — could not be moved. The
setting was read against the last segment of each mapping alone, so the
number the module uses everywhere else was refused as a port it does not
publish, and the node's every push failed for as long as the setting was
stored. The one key that was accepted, the container's own port, was then
read only when the container's mapping was rewritten: the mapping moved
and the ports map, the filter, the adopted node's openings, its guard and
what a consumer is told all stayed on the number the software had left.

Either end of a mapping now names it, and a given port comes back under
both, so every reader finds the same number under the key it holds.
Ambiguity is refused where it is real — one number naming two different
mappings, or the two ends of one mapping given two different numbers.
2026-09-23 02:08:35 +02:00
jschoubben f0049190d7 Prove the untouched environment is the same map, not an equal one 2026-09-22 23:29:28 +02:00
jschoubben 7352c846dd A module is told its port in a container's environment too (hq issue 088)
${port:…} answered only inside a file's content, and the one place a module
routinely writes its own address is a container's `env` — where the literal is
wrong on every node whose assignment differs from the manifest's number, and
wrong again on a node given that port as a setting (ADR 0100). Nothing checked
it: the value is a string like any other, and it fails at runtime, on one node.

Filled by the control plane, like a bound value: a port is not secret, so there
is nothing for the host to be the only witness of and it learns no new field.
That is the line ADR 0086 draws — its objection is to a secret being in an
environment at all, not to who fills one in — so a port crosses it and a
credential still does not. Same guard as before: a port the module never said it
listens on is refused, now naming the container and the variable.

The env map is the catalogue's, shared by every node running the module, and the
resource around it is a shallow copy, so a filled value goes into a fresh map —
otherwise the first node composed writes its own port into the manifest and
every node after it is told that one.

Inert on the catalogue as it stands: ${port:…} is written in one other place in
it, a file. Renamed off _files, which this no longer is.
2026-09-22 22:36:28 +02:00
jschoubben 729537e745 Tie the builder's carried binding to what the forge serves, and hold at shut
The builder carries a binding because at genesis nothing provides
`package-registry` to resolve one from; once the forge is a module the same
consumer is told what the forge serves. Nothing held the two to the same number,
so the catalogue could drift into dialling one port before the forge is assigned
and another after.

And `at` is now protected, for the reason it had to be: a setting that moves it
points the builder, and the registry password it sends, at a host somebody else
chose.

novox/hq 04-ISSUES/085
2026-09-22 21:56:51 +02:00
jschoubben 0e413e3e7a Hold the forge's port to the same rule as every other provider's
The package registry was the one foundation port not resolved from what its
module serves. Nothing in the controller had to change for it — `ports` on the
forge moves its container, what it serves and what consumers are told, and the
builder's carried binding is settable like any other mergeable file — but
nothing said so, which is how it came to be special in the first place.

Two tests over the catalogue's own manifests: the forge's port is given on a
node and reaches what it serves, and the builder's carried binding takes the
port from the node while keeping who the binding is with.

novox/hq 04-ISSUES/085
2026-09-22 21:40:08 +02:00