Commit Graph
68 Commits
Author SHA1 Message Date
jschoubben 19d2725c13 A module can name the mesh's range: ${machine:mesh-range} (novox/hq ADR 0112)
A module cannot know the private network's CIDR — it is a per-mesh value chosen
at genesis — but sometimes must name it: an intrusion filter that must never
ban a tunnel peer. Carry the overlay range on the Rendering and offer it as the
machine fact mesh-range, the same way a machine's own address is offered, so the
module names it rather than hardcoding a value (data is the mesh's). Absent when
the mesh has no range. Enables the fail2ban ignoreip fix.
2026-09-27 16:52:00 +02:00
jschoubben 2e3b13c0f8 The assignment's own root is a place, and the manifest's maps are placed
Slice two of ADR 0112. A pathless directory saying place "." is the
assignment's one directory, <root>/<module> — to-be 27's shape — and
place never reaches the host, which parses strictly. The maps naming
where bindings, credentials and contributions land (binds, secrets,
own-secrets, receives, grants) fill against the placed directories at
composition, into fresh maps and a fresh module slice, because one
resolution composes for many nodes. The five absolute-path checks on
those maps accept a placed reference — resolution makes it absolute
before anything reads it — while certificate, operator-keeps and
accesses paths stay absolute-only: those are the operator's or another
vocabulary's. unknownDirRefs scans the maps too, and validates place
itself: only on a directory, only ".", never beside a stated path.

Found by the foundation tests validating the sibling catalogue: the
first conversion's blanket replace turned /var/lib/gitea/database.json
into ${dir:data}base.json — which resolves to the right path by pure
string concatenation. Production was saved by a coincidence; the
catalogue cleanup that follows spells it ${dir:state}/database.json.
2026-09-26 18:18:13 +02:00
jschoubben d2cbc9dbdc A directory the mesh places: ${dir:<id>} and the pathless directory resource
The first executable slice of ADR 0112 / to-be 27, sized to what the
operator settled tonight: a module definition names no host path for
its own data. A directory resource may omit path; composition resolves
it to <root>/<module>/<id>, the root a node's setting on Rendering with
/var/lib as the default — which reproduces exactly the layout novox
converged to by hand. ${dir:<id>} names the place from a resource's
path, content, mounts, environment and env-files, the same shape as
${bound:…}. A directory that states a path keeps it and still answers
by name — that is the adopted-data placement, mssql its live case.

Resolved in the controller at composition, so the wire format and the
host change not at all; a reference naming no directory refuses at the
manifest and again at composition; nested fills are rebuilt, never
written into the manifest's own maps, because one manifest composes
for many nodes.
2026-09-26 17:51:54 +02:00
jschoubben f996a6707e a route composes its internal-network alias too, not only its public name
Every cutover done on novox tonight (drive, files, files-api, git,
keycloak, umami) dropped the <label>.<node>.internal alias HAL always
paired with the public hostname — found only when the operator tested it
by hand. Not a security boundary (a predecessor proxy served both as a
convenience, reaching a service over the VPN without a public TLS round
trip, not as access control), so restoring it is composing the same
convenience the same way the public name already is: <label> joined to
the node's own private address (r.At), independently of whether a public
domain exists to join the other half to.

composeName's signature changes (publicDomain, internalDomain) but its
shape does not — additive, label-gated, apex-aware, exactly mirroring the
public half it already did. A contribution the mesh writes both names
into is the entire fix; route-adapter and route-proxy pick up internal-
name whenever they're updated to serve it, not before, so this alone
changes nothing about what is live on any node yet.
2026-09-25 16:51:38 +02:00
jschoubben 8fa5443862 contributes: a module's grant carries no value where it contributed several times
ContributionsFrom settled to whichever of a module's several contributions to
one requirement sorted first, arbitrarily — the grant minted for it then
carried that contribution's label and port under a credential the OTHER
contribution's consumer never sees, and collided with that same
contribution's own entry from contributions() besides.

Confirmed live: minio's two route contributions (files-api, files) produced
three entries in route-adapter's received file — files-api twice, once
credentialed and once not, files not credentialed at all. Every
single-contribution module (gitea, keycloak, umami) already mints an unused
credential for `route` too — route never needs one, by its own
documentation — but with exactly one contribution to match there was nothing
to collide with, so it never surfaced.

Where a module contributes more than once, there is no single value to
settle on. The module still asks, still gets its one credential — a pair
credential is not a place for a label or a port anyway — and each named
contribution reaches the provider on its own, unchanged.

No cleanup needed for the secret already minted live for minio+route: the
sealed blob is a random pair credential unrelated to Values, which is
recomputed fresh on every plan/push regardless.
2026-09-24 18:54:35 +02:00
jschoubben f4bcb320fe contributes: a module may answer one requirement several times
A module's contributes was map[string]map[string]any — one JSON object key
per requirement, structurally exactly one contribution to "route" ever.
minio needs two public hostnames (the S3 API and the console), which is
two different contributions to route from one module, and nothing let it
say so.

This is the same shape of problem ADR 0094 solved for secrets (a module
needing several values from one provider that gives one per pair):
contributes now accepts either the ordinary {label, port} object, or an
object of local names to several such objects. Detected per requirement
key by what's inside, since (unlike secrets' string-vs-object split) both
shapes are JSON objects: an ordinary contribution's fields are scalars, the
several-instance shape is local-name -> object. Confirmed against every
module.json in mesh-catalog before relying on that split.

Both route-proxy and the migration-era route-adapter already key generated
routers off the composed hostname (Values["name"]), not the module name,
so two contributions with the same From reach them as two independent
routes with no changes needed on the receiving side.
2026-09-24 18:36:08 +02:00
jschoubben 2277583e99 Tell the resolver the machines, not the names the mesh merely serves
The map the control plane hands a resolution holds both: the machines, and every name
the mesh was told to route to whichever machine serves it. A container's hosts wants all
of it, so a routed name resolves to the proxy. A resolver's zones want only the machines:
told the mesh's suffix is its own it answers authoritatively for everything under it and
forwards none of it, so a routed name with the suffix appended — drive.example.test.internal
— is a name nobody will ever ask for, standing beside the machines and looking as real.

Found composing the resolver's first assignment on a live machine, before pushing it.
hq issue 111.
2026-09-24 01:31:11 +02:00
jschoubben 0d8264ff55 Give the resolver the mesh's suffix as a local domain and a module its machine's address
hal dnsmasq-app conversion, hq 08-connectivity. Converting the resolver from the module it
replaces made it forward what it cannot answer, which is what the predecessor's does, and
that found two things the controller did not say.

A resolver that forwards must not send a mesh name it does not know upstream: the
`node-zones` fact now carries `local=/<suffix>/` beside the wildcards, written here rather
than in the daemon's configuration because the suffix is the mesh's choice and this file is
the one place the mesh writes what it chose. The default lives in one helper now instead of
being spelled in two functions.

The predecessor points the container runtime's `dns` at the machine's own tunnel address —
a container cannot reach the machine's loopback. A module writing that key needs the
address, and `${machine:at}` is the machine's name; a runtime's resolver list cannot be a
name it would need that resolver to look up. So a module may say `${machine:address}`: what
`at` resolves to, read from the same names the hosts file and the wildcards are written
from, absent — and refused — off the network like `at` is.

The `mesh-resolver` and `resolver-data` constants go: nothing provided or consumed either,
the fact and `mesh-addressing` are the mechanism, and a requirement nothing provides is
refused at resolution.

Tests: the catalogue's dnsmasq, resolv-conf and resolved-split-dns manifests are parsed
and composed as a machine would receive them — fixed upstreams, no-resolv, 127.0.0.1, the
machines file, the runtime's key, the pair that decides what a machine asks refused on one
node; and on a real mesh the resolver's machines file is composed with a wildcard per
machine on the network and composed again without one that left, mirroring the hosts fact.
2026-09-24 01:10:15 +02:00
jschoubben e07b56ce43 An address is read from the node's settings where it is used, never recorded with a port
Three readers did not follow a moved foundation port (novox/hq 04-ISSUES/102),
and each took the control-node down in its own way: the control plane's own
store and broker connections, sealed at genesis with the port inside; and every
build the mesh ever recorded, kept as `<registry>:<port>/<module>/<artifact>@…`.

The control plane cannot open its own sealed connections to move a port, and it
cannot bind the store as a consumer would — a binding mints a credential. So its
settings get a third twin, `NAME_PORT`, read on top of the sealed value by the
store, the broker, the management API and the bus connection, and filled into
its container by a placeholder that names a seat, `${seat:mesh-store:5432}`,
from the node's given or mesh-assigned ports — never the manifest's number, and
empty when the mesh has nothing to add, so what genesis wrote stands. A value
that is still a placeholder is nothing said, aloud: the manifest naming it lands
in the next commit, once every control plane that composes it knows it.

A build is now recorded by digest and path — `artifact-store://<module>/<artifact>@…`
— and the store's address is composed in where a reference is used: the
declaration, the trust file, the bases a build is handed, a replay to the
catalogue. Over the network as `<node>.internal:<port>`; on the store's own node
before any network exists — every genesis push before its "network" step — by
loopback. A reference recorded before this, with an address, is re-routed the
same way when the mesh built it. The trust file and every provider's address
come from one derivation: the node's given port, over the mesh's assignment,
over the manifest's number.

novox/hq 04-ISSUES/102
2026-09-23 23:49:31 +02:00
jschoubben 58644fd282 A node may move a port a module publishes as a mapping's machine side
A module publishing `2222:22` — the machine's own ssh daemon holds 22, so
the module takes 2222 and says so in `listens` — could not be moved. The
setting was read against the last segment of each mapping alone, so the
number the module uses everywhere else was refused as a port it does not
publish, and the node's every push failed for as long as the setting was
stored. The one key that was accepted, the container's own port, was then
read only when the container's mapping was rewritten: the mapping moved
and the ports map, the filter, the adopted node's openings, its guard and
what a consumer is told all stayed on the number the software had left.

Either end of a mapping now names it, and a given port comes back under
both, so every reader finds the same number under the key it holds.
Ambiguity is refused where it is real — one number naming two different
mappings, or the two ends of one mapping given two different numbers.
2026-09-23 02:08:35 +02:00
jschoubben 0f3eedd163 Do not guard a port this node is told to open to everyone (hq ADR 0103) 2026-09-22 19:44:35 +02:00
jschoubben 8db66e9532 Derive the guard from taken modules only: their published private-network ports and their manifests' guards (hq ADR 0103) 2026-09-22 17:58:37 +02:00
jschoubben 28b7fb81ba Write the registry's trust into the runtime's file and reload the runtime instead of restarting it; prefix reload-on like restart-on (hq ADR 0102) 2026-09-22 17:51:38 +02:00
jschoubben dbf62b5212 Read the foundation's ports from each node's settings, wherever a port is used (hq ADR 0100) 2026-09-22 17:31:07 +02:00
jschoubben 1ea84f8b8f Take a module, converge a node after a preview, and return it to adopted, from the command line and the API (hq ADR 0100) 2026-09-22 17:27:35 +02:00
jschoubben c3b1617693 Declare openings and a refusal-only guard on an adopted node in place of the filter (hq ADR 0100) 2026-09-22 17:21:44 +02:00
jschoubben a86a6c2974 Carry an adopted node's mode and taken modules in every declaration, from one marshaller (hq ADR 0100) 2026-09-22 17:17:54 +02:00
jschoubben 8152665298 The facts write the suffix the control plane composed the names with, handed down rather than written twice (review of issue 079); the fixture is keyed as production keys it 2026-09-22 01:27:50 +02:00
jschoubben 9f3790dcda Review: one local name is still a local name; a local name is unique; recovery knows it; recipes read as instructions; a tag before a digest; ask fails at once when nothing serves
A secrets object with one local name delivered no file. Two requirements could share
a local name. secret recover and the export could not tell two locals apart. The
recipe check missed continued lines and read heredoc bodies as bases. repo:tag@digest
kept the tag in the repository. ask now publishes mandatory, so a tool nothing serves
is said at once rather than after the wait.
2026-09-21 21:03:22 +02:00
jschoubben 5049d201c6 A provider keeps one grant file per holder, with the local name in its id and path
The lab's vault refused a declaration naming two files with one identity: the
two secrets of one consumer. The holder's suffix is in the resource id and the path now.
2026-09-21 20:41:19 +02:00
jschoubben 6ae4ae1dba A module may hold several secrets from one provider, each a pair of its own
secrets: maps a requirement to several files under local names. Each local name is
its own need, its own pair credential (the pair is keyed on it: migration 0027),
its own file on the consumer, its own holder at the provider (the identity with the
local name after it) and rotates apart from the others. The plain shape is
unchanged and every existing row is the credential it was (novox/hq 04-ISSUES/069,
ADR 0094).
2026-09-21 20:28:16 +02:00
jschoubben 69bb0fcb67 A module names who its secret files belong to (secrets-owner)
The control plane runs as 65534 and crash-looped on permission denied the
first time its credentials were mounted as files the host wrote as root at
0600 — the env-file shape hid this because the daemon reads an env-file on
the host side. The composer now gives a module's secret files the owner the
manifest names.
2026-09-21 10:19:00 +02:00
jschoubben 4531f2244f A secret reaches a process as a file (ADR 0086)
The broker settings take a _FILE twin like the store connections; the
catalogue engine refuses a secret placeholder in a container's env and a
secret-carrying env-file unless the container says why with
secrets-in-environment, which stays in the catalogue and never reaches the
machine.
2026-09-21 10:10:33 +02:00
jschoubben 77e6c1a684 Recoverable means sealed to the current operator key; recovery names the provider
From review: the export counted any operator-sealed row as recoverable, so a
secret sealed to a replaced key was reported as openable with the current one;
replacing the key counted orphans in one table of two; and a pair credential
held from two providers was recovered as whichever row came first. The export
now lists what the current key opens, what an earlier key opens, and what has
no copy; `secret recover` takes --provider and refuses ambiguity; files that
must not exist are created exclusively; one constructor builds the export for
the operator's file and the vault's disk alike.
2026-09-21 01:16:32 +02:00
jschoubben 565f144a20 A pair credential is sealed to the operator key too
The secret the vault provides a module is the credential of the consumer↔vault
pair, and so is every credential a provider grants; sealing only own secrets
to the operator left exactly those unrecoverable. Same column, same call; the
export and `secret recover` address a pair by consumer node, module and the
provision's name, and say which kind each entry is.
2026-09-21 00:36:16 +02:00
jschoubben e140ed5d0b An operator key, a second seal on every own secret, and the vault keeps the export
novox/hq ADR 0085, amended: the mesh's root secrets — the store's superuser,
the broker's administrator, every secret a module holds for itself — were
sealed to a node key and nothing else, so a lost node took them with it.
Now the mesh records an operator's public sealing key and seals every own
secret to it as well, minted or accepted. The private half is written once
by `operator key new` to a file the operator keeps off the mesh; the mesh
holds one more blob per secret that it cannot open.

`secret recover` opens a secret with that key, to a 0600 file, from the
store or from an export; `secret export` writes every operator-sealed copy
as ciphertext. A module that `keeps` (the vault) is handed that export as a
declared file on its own disk, so recovery survives the store.

Secrets made before the key exists have no operator copy and are said so —
the plaintext was discarded — until each is issued again.
2026-09-20 23:56:49 +02:00
jschoubben bc289cbe04 The restart-on rename reads both list shapes
A resource composed in code carries restart-on as []string; the rename
only read []any, so the overlay's registry-trust reload kept its bare
reference, pointed at nothing, and the runtime was never restarted —
the trust was on disk and not in the daemon, with every check passing.
Diagnosed on the built-store-cross-node bed, run 8 (issues 042/048).
2026-09-17 23:35:12 +02:00
jschoubben c3b88b9148 Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:40:40 +02:00
jschoubben fcdb065660 The mesh's knowledge is a fact a module asks for, not three modules
mesh-names, mesh-resolver and the names half of the overlay generators are gone.
They ran no software and could not be swapped for anything, which is the test of
whether something is a module at all — they existed because computed output
needed somewhere to live, and the control plane's only shape for output was a
module.

Now a module says where it wants what the mesh knows:

  facts: { node-zones: /etc/mesh-resolver/nodes.conf }

and is given a file, under its own name, applied and removed like anything else
it declares. Two facts exist: node-names (a hosts file — exact names) and
node-zones (every machine as a wildcard, *.homer.internal is homer). Asking for
a fact the mesh does not compute is refused naming what would have worked,
because a daemon that starts and reads a file nobody wrote is a worse way to
find out.

The names ride with the network now: wireguard's manifest asks for node-names
into /etc/hosts, because being on the private network is what gives a machine a
name. networking no longer requires name-resolution — names are not a provision,
and the module that answered it ran nothing.

One behaviour inverted, deliberately: choosing another VPN used to drag
WireGuard in anyway, because only WireGuard provided the addressing the names
module required — the node-scope claim existed to at least make that loud. With
names as a fact there is nothing to drag in: tailscale assigned means tailscale,
alone. The claim still catches two VPNs assigned explicitly.

And a machine the mesh cannot place is left out of both files rather than named
at nothing: a name resolving to nothing hangs a connection, where an unknown
name fails at once and says so. In practice that is only ever a token issued and
not yet used — a machine that has announced itself has an address.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 21:31:18 +02:00
jschoubben 385bc5b9bf A module can be told which port it was given
The mesh assigns the machine-side port and a module does not choose one (ADR
0038). For a container that is invisible: the mesh rewrites ports into
assigned:wanted, the software binds the number it always bound, and the machine
publishes another.

A process has no such layer. It runs on the machine, there is nothing to rewrite,
and it binds whatever its configuration says. So every process bound the number
written in its own config, two modules declaring the same one would collide, and
the mesh's whole reason for assigning ports was defeated by the resource kind
that most needs it — introduced, by me, three commits ago.

So a module asks. ${port:8080} is "the machine-side port you gave me for the 8080
I said I listen on", written into its own configuration exactly as an address it
was bound to is.

Asking about a port it never declared is refused, and the refusal says what it
did declare: the module is asking about something the mesh has no opinion on, and
answering would put a guess into a configuration file as a port number. With
nothing assigned yet it is told what it asked for, so a mesh that has made no
assignment still composes something coherent rather than writing a zero.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 13:20:01 +02:00
jschoubben dda001d64b The firewall opens the port the mesh itself runs on
Rules are derived from what modules declare they listen on, and the substrate is
not a module. So the broker's port — the one every machine dials to enrol and to
receive every declaration it is ever sent — appeared in no ruleset the mesh has
ever generated.

Nothing caught it because a mesh of one never dials its own broker across the
network: the ruleset looks complete right up until a second machine tries to
join a firewalled anchor and is refused by the packet filter, during enrolment,
before the mesh can report anything about it. Assigning the firewall before
joining machines is both the natural order and the one that breaks.

It is a floor for the same reason ssh is. A machine nobody can reach cannot be
repaired; a machine the mesh cannot reach cannot be managed. Neither is a thing
any module asks for and neither may be derived away.

From anywhere rather than from the private network, deliberately: a node enrols
BEFORE it has an address on that network, so narrowing the rule to it would close
the door being knocked on.

The port is read from the broker this control plane was told about, so the
address handed out in a token and the port a machine must accept on stay one
fact. A mesh never told about a broker gets no such rule, rather than a broken
one — and cannot issue tokens either, which is where that surfaces.

Closes novox/hq 04-ISSUES/052.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 00:54:31 +02:00
jschoubben c8d8211385 The firewall governs what is forwarded, and never closes ssh
Two faults, opposite directions, both in issue 047.

There was no forward chain, on the reasoning that dropping there stops every
container the runtime allowed. The first half is true; the conclusion was not. A
published port is redirected and then forwarded, so it never reaches the input
chain — the firewall was silent about the ports most worth protecting. The way
through is the one the system being replaced already used: deny by default, then
allow the runtime's own networks explicitly. A forwarded rule matches what the
client originally asked for, because the destination has been rewritten by the
time the chain sees it.

And ssh is now a floor nothing derives. Every other line comes from what is
assigned, which is the point — but a mesh part-way through adopting a machine
has been assigned almost nothing, so what it computed was a chain that shut the
port used to fix it. From the mesh always; from outside on a machine that faces
outward, because that is the way in when the private network is what broke.

Rehearsed on three machines: a docker-published port declared mesh-only is now
reachable from inside the mesh and refused from outside. Before, it was
reachable from both.
2026-09-14 15:27:10 +02:00
jschoubben 21abd948aa Say that a module is unbuilt, rather than letting a machine call it malformed
A container naming an artifact is a module saying the mesh builds this. Until a
build publishes one there is nothing to run — and what reached the machine was
an unresolved field, which its language has no room for, so it refused the whole
declaration and reported that a container does not use "artifact". That reads
as a broken manifest. It is not broken, it is unbuilt, and only the mesh can
tell those apart.

Found by the four-machine bed, which assigns modules the mesh has not built.
2026-09-13 04:56:00 +02:00
jschoubben c4030947b0 The routing record is 0066, not 0056
0056 is 'the authority is the control plane, not a database'. A citation
pointing at the wrong decision is worse than none: it reads as corroboration.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:09:56 +02:00
jschoubben 946fddd622 catalogue: a module may name the machine it was assigned to
An authority inside the mesh is reached at <machine>.internal, so its own
certificate must be issued for that name — and it is the one module that cannot
be told its name by a binding, because it provides rather than requires. Written
as a literal it would be one deployment's machine name in a manifest, which is
what ADR 0056 exists to remove.

${machine:name} and ${machine:at}, beside the bound values and refused the same
way. An address the machine does not have is named here rather than discovered
later as a certificate nobody can verify.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:12:04 +02:00
jschoubben 9fab0b731a catalogue: a contribution reaches the port its own machine published, from any machine
The co-located fix could not reach a grant assembled for a consumer on another
machine: ContributionsFrom never sees a port map, so the proxy was told the
workload's software port and dialled a number that machine never published. The
consumer's own assignments are fetched where the grant is built and applied
there. The same fault as 038, one node over.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:08:09 +02:00
jschoubben e78c849002 catalogue: a provider that names a provision and says nothing is not the answer
servedOnThisMachine stopped at the first module whose `serves` mentioned the
provision, even when that entry was empty and there was therefore no fact to give
a consumer. here() had always kept looking in that case, and a set where one
module names a provision without describing it and another describes it is exactly
where the difference shows. Restore the search.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:04:16 +02:00
jschoubben b824c65ab0 catalogue: a co-located provider's served values, and a co-located contribution's port
Two more of one fault, and the fault is the same as 04-ISSUES/038: the same-node
path diverging from the cross-node one.

The mesh works out what a provider on ANOTHER machine serves by walking that node
— reading its manifest with that machine's port assignments, then settling the
result with that node's settings layers — before offering it to a consumer. A
provider on the consumer's OWN machine never passes through that walk, so every
step of it had to be repeated in resolve.go's servedHere and declaration.go's
here(). 038 repeated the port. Nothing repeated the settling.

So a served value the operator supplied reached a co-located consumer as the
manifest's empty default. On the ADR 0056 anchor that value is an internal CA's
root: step-ca and route-proxy on one node, route-proxy's binding carrying
root: "", an empty CA bundle written, a silent fall back to the system trust
store, and issuance stopping with nothing saying why. The same step-ca on another
node would have worked.

The second is the mirror direction. gitea declares a bare container port 3000 and
the machine publishes it as 20000:3000, but gitea's route CONTRIBUTION still said
3000 — so the proxy beside it dialled a port nothing listens on and answered 502.
038 fixed what a consumer is TOLD about a provider; this is what a workload TELLS
a provider about itself. The redirect uses the CONTRIBUTING module's assignment,
because the port is the workload's, not the proxy's; a contribution carried here
from another machine is left exactly as it is, its port being that machine's to
assign.

Both are settled in Declaration, which is the first moment the machine's ports and
the provider's settings both exist. That also removes an order dependence: the
resolver built its same-node needs mid-walk, from whichever modules had been
chosen by the time the requirement came up and in whatever order a map iterated,
so what a co-located binding carried depended on the order somebody happened to
assign things in. Re-deriving from the finished closure does not.

servedHere keeps its job — deciding whether a same-node provider serves anything
at all, which is what makes the need exist — and now says that its values are
provisional.

novox/hq ADR 0056

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 20:57:23 +02:00
jschoubben 0fb2ab7716 catalogue: composeName handles the apex label '@' (bare public domain)
An empty label composed nothing, so a module served at the bare domain (a node's
own site) had to keep a full name — the one route the label model could not
express. The zone-file convention '@' now composes to the public domain itself,
no leading dot, so the apex is a label like any other. Test added.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 00:27:11 +02:00
jschoubben 232862315c catalogue: compose a route's name from a label and its node's domain, and resolve it in-mesh
A public route used to carry its whole hostname as a literal in the module
manifest, so running the same catalogue against a different domain meant
overriding that literal on every routed module, per node. The mesh was, in
effect, holding a map of names to services: the one thing it should never hold,
because the subdomain is the operator's choice and the domain is the node's.

Compose instead. A route contribution carries a `label` (the subdomain); a node
carries its `public_domain` as node-level configuration; the mesh joins
`<label>.<public-domain>` and grants exactly that, interpreting neither half.
Held as a node property beside the node's other node-level facts (endpoint,
site, overlay address), not in a module's settings — the ADR calls it
node-level, and the settings table is keyed per module.

Additive, so an unmigrated catalogue keeps working: a contribution that still
carries a full `name` and no `label` passes through unchanged, and the catalogue
can migrate module by module. A labelled contribution on a node with no public
domain composes nothing, reading downstream as a route that named no host.

And propagate: each granted route name is published into internal resolution
mesh-wide, mapped to the node that serves it, alongside the `<node>.internal`
names every container already gets. So a container — and an internal ACME
validator, which cannot complete a challenge for a name it cannot reach —
resolves a routed name to the proxy that serves it. Name-agnostic throughout:
the mesh propagates whatever names it was told to serve and knows nothing about
what they mean.

novox/hq 02-DECISIONS/0056

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 23:27:35 +02:00
jschoubben c147a26138 catalogue: announce a same-node provider at the port it is published on
A co-located consumer of a `from: mesh` provision was told the port the
provider module DECLARED, not the host port the mesh assigned and published
it on. The same-node served facts are settled while resolving (servedHere,
here()), before a bare `ports` mapping is assigned its host port, so they
carried the declared number; only the cross-node path re-derived them after
assignment. So the provider was published on <node>.internal:<assigned> while
its own-machine consumer dialled <node>.internal:<declared>, where nothing
listens — the ordinary small-mesh case, and the one the fix for issue 018
(announce the same-node provider at all) left one promise short of kept.

Redirect same-node needs to the machine's assignment in Declaration, where the
port map is known, exactly as plan.go already does cross-node. The publish bind
is unchanged (all interfaces, scoped to the mesh by the listen's firewall rule);
only the announced port is corrected.

novox/hq 04-ISSUES/038

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 23:02:50 +02:00
jschoubben 958bef56c7 catalogue: give host-network containers the mesh's names too
A container with `network: host` was skipped when the mesh injects its
`<node>.internal` names, on the belief it "shares the machine's hosts file
already". It does not: `docker run --network host` still gives the container
its own /etc/hosts (localhost and its own id only), so every internal name the
mesh wrote is invisible inside it, and a client that dials one gets EAI_AGAIN.

This surfaced with the first host-network consumer to dial a provider by the
`.internal` address the mesh hands it as `${bound:...:at}` (the model-usage
store reaching its postgres). The remedy is the same `--add-host` every other
container already gets — the runtime accepts it with `--network host`
(verified against Docker) and mesh-host emits it for any network mode.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 03:44:28 +02:00
jschoubben 33fd28ffa6 licences: deliver the refresh token by the ordinary sealed path, not a bespoke envelope
The refreshable-grant refresh token no longer rides a custom at-rest envelope that a
module opens with a node private key. A module is never given a node's private sealing
key, so that path could not exist -- the gap Phase C hit.

Instead the refresh token is a credential sealed to the MANAGER holder with the same
anonymous box (secrets.Seal / crypto_box_seal) every credential uses, stored as one
sealed blob, and delivered by the existing host-unseal-and-mount: the host opens it with
the node's real key and mounts the cleartext at the manager module's bound path, exactly
as a consumer's db password is delivered.

  - refresh_grant now stores { sealed, manager_key }, dropping the AtRest token/wrapped_key
    columns; internal/secrets/atrest.go is retired (nothing else used it).
  - the licence records its manager as (node, module); KeyFor delivers the refresh token to
    the manager holder and the access token to consumers, disambiguated by module so the two
    can co-locate. Accept and the reseal skip the manager holder.
  - the manager holder is delivered the node's PUBLIC sealing key in its bound facts, so the
    module can re-seal a rotated refresh token with no private key of its own; the
    declaration tolerates its empty pre-adoption secret rather than refusing.
  - SubmitRefresh / set-grant take a sealed blob, never a refresh token in the clear.

The invariant holds unchanged: the control plane never reads the refresh token, and no node
but the manager holds it. A committed cross-language test proves the TypeScript module seal
opens under Go box.OpenAnonymous (the host's Unseal) -- both are NaCl crypto_box_seal.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 01:55:08 +02:00
jschoubben d0ef659824 fix: a require-only consumer of a parameterless provision still asks (and is minted a credential)
A consumer that requires a provision whose serves names no consumer key (redis-cache, amqp)
contributes no payload, but it still ASKS for it. ContributionsFrom keyed 'asks' on
contributions alone, so such a consumer's grant got From='' — read as withdrawn — and the
provider never created its account. redis-cache consumers (e.g. baserow) were silently
unprovisioned, tolerated only by their embedded fallback. A module asks iff it still requires
the provision, whether or not it hands anything up. Regression test added.

Found by the lavinmq AMQP provider bed (given:[] for a require-only amqp consumer); fix
lab-proven green there.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 23:20:58 +02:00
jschoubben 99a753994e fix: fill a ptr-secret placeholder from the file-owner's credential, not the last consumer's
The provider-seal-key gate: on a node with two modules requiring the same provision (baserow
and letta both consuming postgres), sealedFor matched a need by provision NAME alone, so a
file's ${secret:X} placeholder took whichever consumer's sealed credential came last in
r.Needs -- the OTHER module's password. baserow was handed letta's password and could not
authenticate. The secrets:-map delivery path already guards this (For == m.Module, novox/hq
04-ISSUES/022); the ${secret:...} placeholder path did not. Added the same guard.

Also dedups the contributions file: when provider and consumer are co-located, grantsFor
enumerates the same-node consumer, so a consumer was emitted twice into the provider's
receives file (once full with its grant, once partial). The m.Contributes loop now skips a
(provision, module) the grants loop already carried; non-grant contributions (routes) still emit.

Regression test added: two consumers of one provision each get their own credential. Proven
end-to-end on a two-node lab install (mesh-lab assigned-two-node-db): baserow and letta on one
node, substrate on another, each authenticates with its own minted password.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 13:46:55 +02:00
jschoubben aeb65a3e1d catalogue: a module accesses operator-owned data, and does not own it
04-ISSUES/036: the media stack is several modules that must share the
library and download directories on one machine, but the manifest could
only say "a directory I own". Six modules each declared the same paths as
their own resources, and the resolver's duplicate-owner refusal — right
in general — would refuse the stack's only sensible assignment the first
time two of them landed on one node.

Add an `accesses` field: a pre-existing, operator-owned path a module is
granted use of but does not own (novox/hq ADR 0051). Distinct from a
`directory` resource on every axis the host acts on — the mesh creates,
chowns and reconciles a directory; it mounts an access and owns nothing.
An access is not a resource, so it never enters the duplicate-owner map
and several modules may name one path with no conflict. What is refused
is the contradiction: a path one module owns and another accesses.

Rendered into the declaration as an `access` resource, before the
container that mounts it, so the host can find it present or refuse
clearly. Unit tests cover co-resolution (the exact 036 case), the
unchanged owner-vs-owner refusal, the owner-vs-accessor refusal, and
access validation.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 22:10:58 +02:00
jschoubben 87c193b202 Resync hq ADR references 0044-0054 -> 0039-0049 after the hq record reconciliation 2026-09-05 12:47:13 +02:00
jschoubben 9b7ba2e20c identity: a consumer's identity fits the tightest backend, via a slug (ADR 0054)
A module may declare a short `slug`; the mesh derives mesh_<node>_<slug|name> and
refuses at assignment (naming the slug as the remedy) when it would still overflow —
identityLimit is now 20, an S3 access key's, the tightest of the backends a login
reaches (04-ISSUES/010). The slug rides the grant so the provider derives the same
login the consumer does, even across nodes. CheckIdentity is now wired, in grantsFor.

Also, the minted secret shrinks to 40 chars (30 bytes) from 43: an S3 secret key is
8-40, the same fit-the-tightest-backend rule on the credential's other half.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 02:51:39 +02:00
jschoubben b306c74467 filtering: a per-node 'expose' setting overrides a listen's source (ADR 0051)
listens.from was a manifest constant — one value for every node a module runs
on. Now a per-node setting overrides it: {"expose": {"5432": "anywhere"}} makes
postgres public on the machine it is set for while it stays from:mesh elsewhere,
and the firewall (ADR 0050) is computed from the effective source. Exposure()
validates it — a port the module does not listen on, or a source that is not
mesh/anywhere/machine, is refused rather than reaching nothing; UnusedSettings
knows 'expose' is a real destination. Tested: default mesh, setting opens it to
anywhere, bad settings refused.
2026-09-04 21:12:33 +02:00
jschoubben 1f5b70a995 The mesh assigns the port, and a module says it once
novox/hq ADR 0038. A module cannot choose a port: it is written once and
assigned anywhere, so any number it picks is a guess about a machine it
has never seen. A database module met the mesh's own store on 5432 and
was told, by a container runtime three layers down, that the port was
already allocated.

The number used to appear three times in every module — the rule set,
what a consumer is told, and what the runtime publishes — agreeing only
because one person wrote all three. Now it appears once, in `listens`,
and the other two are derived: the container publishes `20000:5432`, the
consumer is told 20000, and the rule set opens 20000.

An assignment is made once and kept, as a credential is. A port that
moved on every declaration would restart both ends each time and hand a
consumer a number that was true when it was read.

Ports the protocol fixes — mail on 25, submission on 587, DNS on 53 —
say so, and are then claims: one holder per machine, and the second is
refused by name at assignment. That is the mechanism the mesh already
has for what is singular on a machine, pointed at ports.

A mapping written the long way is left exactly as it is. Some things
must be pinned by hand, and quietly overruling somebody who wrote both
halves would be worse than not offering the short form.

Still open, and known: the substrate is not a module, so the mesh has
never heard of its own store and cannot yet assign around it. That is
what 028 will still be about after this.
2026-09-01 17:52:53 +02:00