novox/hq ADR 0085, amended: the mesh's root secrets — the store's superuser,
the broker's administrator, every secret a module holds for itself — were
sealed to a node key and nothing else, so a lost node took them with it.
Now the mesh records an operator's public sealing key and seals every own
secret to it as well, minted or accepted. The private half is written once
by `operator key new` to a file the operator keeps off the mesh; the mesh
holds one more blob per secret that it cannot open.
`secret recover` opens a secret with that key, to a 0600 file, from the
store or from an export; `secret export` writes every operator-sealed copy
as ciphertext. A module that `keeps` (the vault) is handed that export as a
declared file on its own disk, so recovery survives the store.
Secrets made before the key exists have no operator copy and are said so —
the plaintext was discarded — until each is issued again.
The modules the mesh runs were under examples/modules/, which framed
the real catalogue as illustrations of a control-plane package. They
are neither examples nor the control plane's — they are the mesh's own
catalogue, and they now live in their own repository (novox/mesh-catalog),
consumed as a build source like any other.
The engine that reads them stays here (internal/catalogue): the control
plane owns the manifest contract; the data does not belong beside it.
Removed with them: modules_test.go and parseall_test.go, which validated
the example manifests against the parser. That validation logically
follows the catalogue to mesh-catalog, but it imports internal/catalogue,
so re-homing it needs the parser exported from internal/ first — a
deliberate follow-up, not done here. Until then the pipeline is the gate,
and internal/catalogue's own inline tests still cover the parser.
Answers novox/hq ADR 0030's open tier-4 question — where the catalogue
lives — in favour of one flat mesh-catalog repository.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The broker's amqps port is opened from anywhere so a node can enrol
before it has an overlay address — but only in the input chain. The
broker is a published container port, so a cross-node dial is DNAT'd
and forwarded, never reaching input; it survived on the first
connection's conntrack entry and no more. Adopting the foundation's own
broker restarts it, dropping that entry, after which a joined node
could never receive another declaration. The forward chain now carries
the foundation ports too, from anywhere, matching their input rule.
Intermittent in the built-store-cross-node bed: it passed whenever the
broker did not happen to restart after the joined node first connected.
An adversarial review of the 055 fix found it encoded the wrong invariants, latent while
every mesh keeps its broker on the hub. Now: the address is the overlay name of the node
ASSIGNED a module claiming the mesh-broker seat (the hub stands in only while nothing holds
the seat — genesis); "on the overlay" is what whereEveryoneIs answers (resolved the
networking module), not "has an address"; a portless genesis address defaults to 5671
instead of silently disabling the path; a second `overlay place --hub` is refused rather
than last-write-wins; and `overlay place` says that earlier credentials keep their old
address. A test now binds the controller's own module.json to its seat, so deleting the
claim fails the suite.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A resource composed in code carries restart-on as []string; the rename
only read []any, so the overlay's registry-trust reload kept its bare
reference, pointed at nothing, and the runtime was never restarted —
the trust was on disk and not in the daemon, with every check passing.
Diagnosed on the built-store-cross-node bed, run 8 (issues 042/048).
Renames the module's own claim the-controller -> mesh-controller (the seat is the server,
ADR 0079), and adds TestAFoundationModuleCannotBeRaisedOnASecondNode asserting each
foundation module's second assignment is refused with 'one per mesh'. Closes hq issue 056.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A new 'package' artifact kind builds a module's own code on a public base image
and publishes it to the mesh's package registry by version (hq ADR 0076) — the
SDK above all, which the toolchain is built from and so cannot be built in the
toolchain. The credential a build needs to resolve or publish packages is
rendered as an .npmrc (basic auth, hq ADR 0048) and given to an image build as a
buildkit secret, never a layer, so a token is not baked into the toolchain image.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
mesh-names, mesh-resolver and the names half of the overlay generators are gone.
They ran no software and could not be swapped for anything, which is the test of
whether something is a module at all — they existed because computed output
needed somewhere to live, and the control plane's only shape for output was a
module.
Now a module says where it wants what the mesh knows:
facts: { node-zones: /etc/mesh-resolver/nodes.conf }
and is given a file, under its own name, applied and removed like anything else
it declares. Two facts exist: node-names (a hosts file — exact names) and
node-zones (every machine as a wildcard, *.homer.internal is homer). Asking for
a fact the mesh does not compute is refused naming what would have worked,
because a daemon that starts and reads a file nobody wrote is a worse way to
find out.
The names ride with the network now: wireguard's manifest asks for node-names
into /etc/hosts, because being on the private network is what gives a machine a
name. networking no longer requires name-resolution — names are not a provision,
and the module that answered it ran nothing.
One behaviour inverted, deliberately: choosing another VPN used to drag
WireGuard in anyway, because only WireGuard provided the addressing the names
module required — the node-scope claim existed to at least make that loud. With
names as a fact there is nothing to drag in: tailscale assigned means tailscale,
alone. The claim still catches two VPNs assigned explicitly.
And a machine the mesh cannot place is left out of both files rather than named
at nothing: a name resolving to nothing hangs a connection, where an unknown
name fails at once and says so. In practice that is only ever a token issued and
not yet used — a machine that has announced itself has an address.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The mesh assigns the machine-side port and a module does not choose one (ADR
0038). For a container that is invisible: the mesh rewrites ports into
assigned:wanted, the software binds the number it always bound, and the machine
publishes another.
A process has no such layer. It runs on the machine, there is nothing to rewrite,
and it binds whatever its configuration says. So every process bound the number
written in its own config, two modules declaring the same one would collide, and
the mesh's whole reason for assigning ports was defeated by the resource kind
that most needs it — introduced, by me, three commits ago.
So a module asks. ${port:8080} is "the machine-side port you gave me for the 8080
I said I listen on", written into its own configuration exactly as an address it
was bound to is.
Asking about a port it never declared is refused, and the refusal says what it
did declare: the module is asking about something the mesh has no opinion on, and
answering would put a guess into a configuration file as a port number. With
nothing assigned yet it is told what it asked for, so a mesh that has made no
assignment still composes something coherent rather than writing a zero.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The bundle recipe: the one that both builds and packs. An archive packs a
directory as it stands, so shipping compiled output meant compiling somewhere
first — which meant a Dockerfile repeating the same incantation in every module.
Two base arguments with no defaults, a working directory chosen so the SDK
resolves upward, the compiler invoked by absolute path because the usual symlink
is resolved away when the base image is assembled, a second stage, an environment
variable naming the entrypoints. Most of the catalogue is unconverted and that is
why; two conversions done in one session were each wrong twice with a working
example open in the next window.
A bundle says a language and a list of entrypoints. The mesh knows what the
language implies. Anything a module could override there it would be writing a
Dockerfile to override, so a toolchain is deliberately not configurable.
Declared rather than inferred, both of them: guessing the language from which
files are present makes a build depend on a directory listing, and guessing the
entrypoints makes it change meaning when somebody adds a helper.
A toolchain the mesh does not hold is refused before anything is compiled, naming
what to build first — the same treatment a missing base already gets, because it
is the same question and somebody can answer it. A language the mesh does not
build is refused saying what would have worked, since the author is usually one
word away.
The list of languages is closed and adding to it is a decision. Every language is
another implementation of the contracts every module shares, and those change
rarely and cascade when they do (ADR 0039) — a mesh whose SDKs disagree about the
envelope fails by ignoring messages rather than by failing to compile.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Rules are derived from what modules declare they listen on, and the substrate is
not a module. So the broker's port — the one every machine dials to enrol and to
receive every declaration it is ever sent — appeared in no ruleset the mesh has
ever generated.
Nothing caught it because a mesh of one never dials its own broker across the
network: the ruleset looks complete right up until a second machine tries to
join a firewalled anchor and is refused by the packet filter, during enrolment,
before the mesh can report anything about it. Assigning the firewall before
joining machines is both the natural order and the one that breaks.
It is a floor for the same reason ssh is. A machine nobody can reach cannot be
repaired; a machine the mesh cannot reach cannot be managed. Neither is a thing
any module asks for and neither may be derived away.
From anywhere rather than from the private network, deliberately: a node enrols
BEFORE it has an address on that network, so narrowing the rule to it would close
the door being knocked on.
The port is read from the broker this control plane was told about, so the
address handed out in a token and the port a machine must accept on stay one
fact. A mesh never told about a broker gets no such rule, rather than a broken
one — and cannot issue tokens either, which is where that surfaces.
Closes novox/hq 04-ISSUES/052.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Two sibling branches resolve a provision answered by the consumer's own machine.
The one for a node-scoped provider falls back to loopback when the machine is on
no private network, with a comment saying why and a test holding it. The one for
a mesh-scoped provider passed node.At straight through, and nothing noticed
because nothing had yet composed a host out of it.
The mesh's own artifact store is mesh-scoped and sits on the same machine as the
builder that pushes to it. Give the builder the address from its binding and it
gets MESH_REGISTRY=:5000 — a name with no host, written into its environment
without complaint. It surfaces much later as
":5000/mesh-tools/build" is not a valid repository/tag
which is a message about a tag for a fault in how a binding was resolved, on a
machine several steps from the decision.
A machine off the private network still reaches itself, which is what the
neighbouring branch already said. The test fails without the fix, showing the
empty address rather than only the symptom.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The floor allowed ssh from the mesh's addresses, and from everywhere on a
machine that faces outward. On a machine the mesh knows no addresses for it
emitted neither — so the chain dropped by default and ssh was simply shut.
That is the first machine anybody adopts: reached over the network, with the
port needed to fix it closed by the act of adopting it. Found by reading the
rules off a live machine rather than trusting the generator.
Two faults, opposite directions, both in issue 047.
There was no forward chain, on the reasoning that dropping there stops every
container the runtime allowed. The first half is true; the conclusion was not. A
published port is redirected and then forwarded, so it never reaches the input
chain — the firewall was silent about the ports most worth protecting. The way
through is the one the system being replaced already used: deny by default, then
allow the runtime's own networks explicitly. A forwarded rule matches what the
client originally asked for, because the destination has been rewritten by the
time the chain sees it.
And ssh is now a floor nothing derives. Every other line comes from what is
assigned, which is the point — but a mesh part-way through adopting a machine
has been assigned almost nothing, so what it computed was a chain that shut the
port used to fix it. From the mesh always; from outside on a machine that faces
outward, because that is the way in when the private network is what broke.
Rehearsed on three machines: a docker-published port declared mesh-only is now
reachable from inside the mesh and refused from outside. Before, it was
reachable from both.
A fingerprint written into a recipe names one particular copy of the base — the
copy on whichever machine the person typing it was using. On any other mesh that
copy has never existed, so the build stops on its first line with a message
about an image nobody can look up. Three modules in the catalogue were in
exactly that state, and the line each of them replaced was equally dead.
A module now names the module and artifact instead, and the mesh answers with
what it holds. The builder is still a thing that clones, builds and answers: the
answer travels with the question, because only the mesh knows what it has.
A base the mesh has not built is refused before anything is built, naming which
module has to exist first.
A container naming an artifact is a module saying the mesh builds this. Until a
build publishes one there is nothing to run — and what reached the machine was
an unresolved field, which its language has no room for, so it refused the whole
declaration and reported that a container does not use "artifact". That reads
as a broken manifest. It is not broken, it is unbuilt, and only the mesh can
tell those apart.
Found by the four-machine bed, which assigns modules the mesh has not built.
0056 is 'the authority is the control plane, not a database'. A citation
pointing at the wrong decision is worse than none: it reads as corroboration.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
An authority inside the mesh is reached at <machine>.internal, so its own
certificate must be issued for that name — and it is the one module that cannot
be told its name by a binding, because it provides rather than requires. Written
as a literal it would be one deployment's machine name in a manifest, which is
what ADR 0056 exists to remove.
${machine:name} and ${machine:at}, beside the bound values and refused the same
way. An address the machine does not have is named here rather than discovered
later as a certificate nobody can verify.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The co-located fix could not reach a grant assembled for a consumer on another
machine: ContributionsFrom never sees a port map, so the proxy was told the
workload's software port and dialled a number that machine never published. The
consumer's own assignments are fetched where the grant is built and applied
there. The same fault as 038, one node over.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
servedOnThisMachine stopped at the first module whose `serves` mentioned the
provision, even when that entry was empty and there was therefore no fact to give
a consumer. here() had always kept looking in that case, and a set where one
module names a provision without describing it and another describes it is exactly
where the difference shows. Restore the search.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Two more of one fault, and the fault is the same as 04-ISSUES/038: the same-node
path diverging from the cross-node one.
The mesh works out what a provider on ANOTHER machine serves by walking that node
— reading its manifest with that machine's port assignments, then settling the
result with that node's settings layers — before offering it to a consumer. A
provider on the consumer's OWN machine never passes through that walk, so every
step of it had to be repeated in resolve.go's servedHere and declaration.go's
here(). 038 repeated the port. Nothing repeated the settling.
So a served value the operator supplied reached a co-located consumer as the
manifest's empty default. On the ADR 0056 anchor that value is an internal CA's
root: step-ca and route-proxy on one node, route-proxy's binding carrying
root: "", an empty CA bundle written, a silent fall back to the system trust
store, and issuance stopping with nothing saying why. The same step-ca on another
node would have worked.
The second is the mirror direction. gitea declares a bare container port 3000 and
the machine publishes it as 20000:3000, but gitea's route CONTRIBUTION still said
3000 — so the proxy beside it dialled a port nothing listens on and answered 502.
038 fixed what a consumer is TOLD about a provider; this is what a workload TELLS
a provider about itself. The redirect uses the CONTRIBUTING module's assignment,
because the port is the workload's, not the proxy's; a contribution carried here
from another machine is left exactly as it is, its port being that machine's to
assign.
Both are settled in Declaration, which is the first moment the machine's ports and
the provider's settings both exist. That also removes an order dependence: the
resolver built its same-node needs mid-walk, from whichever modules had been
chosen by the time the requirement came up and in whatever order a map iterated,
so what a co-located binding carried depended on the order somebody happened to
assign things in. Re-deriving from the finished closure does not.
servedHere keeps its job — deciding whether a same-node provider serves anything
at all, which is what makes the need exist — and now says that its values are
provisional.
novox/hq ADR 0056
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
An empty label composed nothing, so a module served at the bare domain (a node's
own site) had to keep a full name — the one route the label model could not
express. The zone-file convention '@' now composes to the public domain itself,
no leading dot, so the apex is a label like any other. Test added.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A public route used to carry its whole hostname as a literal in the module
manifest, so running the same catalogue against a different domain meant
overriding that literal on every routed module, per node. The mesh was, in
effect, holding a map of names to services: the one thing it should never hold,
because the subdomain is the operator's choice and the domain is the node's.
Compose instead. A route contribution carries a `label` (the subdomain); a node
carries its `public_domain` as node-level configuration; the mesh joins
`<label>.<public-domain>` and grants exactly that, interpreting neither half.
Held as a node property beside the node's other node-level facts (endpoint,
site, overlay address), not in a module's settings — the ADR calls it
node-level, and the settings table is keyed per module.
Additive, so an unmigrated catalogue keeps working: a contribution that still
carries a full `name` and no `label` passes through unchanged, and the catalogue
can migrate module by module. A labelled contribution on a node with no public
domain composes nothing, reading downstream as a route that named no host.
And propagate: each granted route name is published into internal resolution
mesh-wide, mapped to the node that serves it, alongside the `<node>.internal`
names every container already gets. So a container — and an internal ACME
validator, which cannot complete a challenge for a name it cannot reach —
resolves a routed name to the proxy that serves it. Name-agnostic throughout:
the mesh propagates whatever names it was told to serve and knows nothing about
what they mean.
novox/hq 02-DECISIONS/0056
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A co-located consumer of a `from: mesh` provision was told the port the
provider module DECLARED, not the host port the mesh assigned and published
it on. The same-node served facts are settled while resolving (servedHere,
here()), before a bare `ports` mapping is assigned its host port, so they
carried the declared number; only the cross-node path re-derived them after
assignment. So the provider was published on <node>.internal:<assigned> while
its own-machine consumer dialled <node>.internal:<declared>, where nothing
listens — the ordinary small-mesh case, and the one the fix for issue 018
(announce the same-node provider at all) left one promise short of kept.
Redirect same-node needs to the machine's assignment in Declaration, where the
port map is known, exactly as plan.go already does cross-node. The publish bind
is unchanged (all interfaces, scoped to the mesh by the listen's firewall rule);
only the announced port is corrected.
novox/hq 04-ISSUES/038
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A module a person assigns to a machine that cannot host it — its declared
capability has no detector there, as fail2ban does on a host with no firewall —
made Resolve refuse the entire node, so a whole-node push refused to send the
healthy modules beside it too. One module on the wrong machine took down every
other module on that node.
Assign already keeps such an assignment on purpose (it is what a person meant,
and acts.go says so), so the fix is on the resolve/push side: a directly-assigned
module the machine cannot host is left out of the closure and reported as
un-applied on the Resolution, rather than refusing the set. The healthy modules
still resolve, declare, and converge. A module that is *required* by something
running here and cannot be hosted still refuses — that set is genuinely
incoherent — so the distinction is who wanted it.
assign, plan and push now name the un-applied module and the missing capability,
via a shared WrongMachine message, so it is neither silently dropped nor fatal.
Reconciled two tests that encoded the old whole-node refusal for directly-assigned
un-hostable modules; added coverage for the healthy-modules-still-converge case
and the required-un-hostable-still-refuses distinction.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A node-scope provider that answers a requirement on the same machine and
serves connection facts (a port) but mints no credential delivered
nothing to a co-located consumer. resolve.go only built the delivering
Needed when brokered[want] was set — true only for mesh-scope providers;
a node-scope keyless provider set local[want] instead and fell through,
so knownFor saw no binding and boundInto refused the consumer's
${bound:model-access:port} file.
Deliver the served facts as a need whenever the same-node answer serves a
non-empty set, with a loopback fallback for the address when the node is
off the private network — the reachability rule does not apply to two ends
on one machine. The brokered (credentialed, mesh-scope) path is untouched.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A container with `network: host` was skipped when the mesh injects its
`<node>.internal` names, on the belief it "shares the machine's hosts file
already". It does not: `docker run --network host` still gives the container
its own /etc/hosts (localhost and its own id only), so every internal name the
mesh wrote is invisible inside it, and a client that dials one gets EAI_AGAIN.
This surfaced with the first host-network consumer to dial a provider by the
`.internal` address the mesh hands it as `${bound:...:at}` (the model-usage
store reaching its postgres). The remedy is the same `--add-host` every other
container already gets — the runtime accepts it with `--network host`
(verified against Docker) and mesh-host emits it for any network mode.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The refreshable-grant refresh token no longer rides a custom at-rest envelope that a
module opens with a node private key. A module is never given a node's private sealing
key, so that path could not exist -- the gap Phase C hit.
Instead the refresh token is a credential sealed to the MANAGER holder with the same
anonymous box (secrets.Seal / crypto_box_seal) every credential uses, stored as one
sealed blob, and delivered by the existing host-unseal-and-mount: the host opens it with
the node's real key and mounts the cleartext at the manager module's bound path, exactly
as a consumer's db password is delivered.
- refresh_grant now stores { sealed, manager_key }, dropping the AtRest token/wrapped_key
columns; internal/secrets/atrest.go is retired (nothing else used it).
- the licence records its manager as (node, module); KeyFor delivers the refresh token to
the manager holder and the access token to consumers, disambiguated by module so the two
can co-locate. Accept and the reseal skip the manager holder.
- the manager holder is delivered the node's PUBLIC sealing key in its bound facts, so the
module can re-seal a rotated refresh token with no private key of its own; the
declaration tolerates its empty pre-adoption secret rather than refusing.
- SubmitRefresh / set-grant take a sealed blob, never a refresh token in the clear.
The invariant holds unchanged: the control plane never reads the refresh token, and no node
but the manager holds it. A committed cross-language test proves the TypeScript module seal
opens under Go box.OpenAnonymous (the host's Unseal) -- both are NaCl crypto_box_seal.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A consumer that requires a provision whose serves names no consumer key (redis-cache, amqp)
contributes no payload, but it still ASKS for it. ContributionsFrom keyed 'asks' on
contributions alone, so such a consumer's grant got From='' — read as withdrawn — and the
provider never created its account. redis-cache consumers (e.g. baserow) were silently
unprovisioned, tolerated only by their embedded fallback. A module asks iff it still requires
the provision, whether or not it hands anything up. Regression test added.
Found by the lavinmq AMQP provider bed (given:[] for a require-only amqp consumer); fix
lab-proven green there.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A container may declare schedule: "<cron>", the recurring twin of
run-once. The resolver already carries a resource's keys through
untouched, so schedule reaches the rendered host declaration on its own;
what belongs here is refusing, near its author, what the host would
otherwise refuse far away.
The manifest parser refuses a schedule that is not a string, one that is
not a well-formed five-field cron (cron.go: fields, ranges, *, comma,
dash, slash), and the contradictory pair run-once + schedule -- a
container runs once and gates, or on a cadence, or stays up, never two.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The provider-seal-key gate: on a node with two modules requiring the same provision (baserow
and letta both consuming postgres), sealedFor matched a need by provision NAME alone, so a
file's ${secret:X} placeholder took whichever consumer's sealed credential came last in
r.Needs -- the OTHER module's password. baserow was handed letta's password and could not
authenticate. The secrets:-map delivery path already guards this (For == m.Module, novox/hq
04-ISSUES/022); the ${secret:...} placeholder path did not. Added the same guard.
Also dedups the contributions file: when provider and consumer are co-located, grantsFor
enumerates the same-node consumer, so a consumer was emitted twice into the provider's
receives file (once full with its grant, once partial). The m.Contributes loop now skips a
(provision, module) the grants loop already carried; non-grant contributions (routes) still emit.
Regression test added: two consumers of one provision each get their own credential. Proven
end-to-end on a two-node lab install (mesh-lab assigned-two-node-db): baserow and letta on one
node, substrate on another, each authenticates with its own minted password.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A container may be marked `run-once: true` — a step the host runs to completion,
gating whatever the declaration places after it. The control plane's part is
small: the field is carried to the host unchanged (containers pass through as
maps), and the step keeps its author-order position ahead of the container it
gates, because the gate is declaration order, not a resolved dependency
(ADR 0005).
The manifest parser refuses a run-once that is not a boolean and the pair
run-once + restart-on (contradictory lifecycles) — near the manifest rather than
far away on the machine, the same lesson the action ban records. Three unit
tests; go build ./... and go test ./... green.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
04-ISSUES/036: the media stack is several modules that must share the
library and download directories on one machine, but the manifest could
only say "a directory I own". Six modules each declared the same paths as
their own resources, and the resolver's duplicate-owner refusal — right
in general — would refuse the stack's only sensible assignment the first
time two of them landed on one node.
Add an `accesses` field: a pre-existing, operator-owned path a module is
granted use of but does not own (novox/hq ADR 0051). Distinct from a
`directory` resource on every axis the host acts on — the mesh creates,
chowns and reconciles a directory; it mounts an access and owns nothing.
An access is not a resource, so it never enters the duplicate-owner map
and several modules may name one path with no conflict. What is refused
is the contradiction: a path one module owns and another accesses.
Rendered into the declaration as an `access` resource, before the
container that mounts it, so the host can find it present or refuse
clearly. Unit tests cover co-resolution (the exact 036 case), the
unchanged owner-vs-owner refusal, the owner-vs-accessor refusal, and
access validation.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A module may declare a short `slug`; the mesh derives mesh_<node>_<slug|name> and
refuses at assignment (naming the slug as the remedy) when it would still overflow —
identityLimit is now 20, an S3 access key's, the tightest of the backends a login
reaches (04-ISSUES/010). The slug rides the grant so the provider derives the same
login the consumer does, even across nodes. CheckIdentity is now wired, in grantsFor.
Also, the minted secret shrinks to 40 chars (30 bytes) from 43: an S3 secret key is
8-40, the same fit-the-tightest-backend rule on the credential's other half.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
listens.from was a manifest constant — one value for every node a module runs
on. Now a per-node setting overrides it: {"expose": {"5432": "anywhere"}} makes
postgres public on the machine it is set for while it stays from:mesh elsewhere,
and the firewall (ADR 0050) is computed from the effective source. Exposure()
validates it — a port the module does not listen on, or a source that is not
mesh/anywhere/machine, is refused rather than reaching nothing; UnusedSettings
knows 'expose' is a real destination. Tested: default mesh, setting opens it to
anywhere, bad settings refused.
A module declares the event types it emits and the patterns it consumes,
parallel to provides/requires (novox/hq ADR 0046). Events span module.*,
mesh.* and node.* sources; a consumes for an event nothing emits is a
dangling edge. Fields only here; the dangling-edge check and the runtime
wiring follow.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Three faults, one file split. All from reading, all verified to bite.
Unassign now releases the module's ports. ReleasePorts existed, said
"for when it is unassigned" in its own comment, and was called by
nothing — so a fixed port stayed claimed in the name of a module that
was gone, and the next module needing it was refused by a ghost.
Kept-once-chosen is a promise about a module that is still here.
MachineSide reads addressed mappings. "127.0.0.1:8080:80" was split at
the first colon, "127.0.0.1" failed to parse as a port, and the mapping
was silently skipped — putting the filter back on the declared port,
the exact fault the function was written to end. The machine side is
the second-from-last part, which is the reading the host already
applies, and the substrate bundle writes that shape today.
An allocation race answers in the mesh's words. Two concurrent picks of
the same port used to surface as a Postgres constraint violation,
verbatim. The table has two keys, so the collision is one of two facts:
the racer was this same assignment — then its answer is the answer,
kept-once-chosen does not care who chose — or another module took the
machine port, and an unfixed pick is simply made again against the
moved free list. A fixed port that lost the race is refused by name.
Told apart by re-reading the row, not by the constraint's name, so this
does not couple to the migration's spelling.
And the artifact-store cycle tests moved to bootstrap_cycle_test.go;
machineside_test.go had quietly become three subjects.
Refused where it is written: a module that provides `artifact-store` and
also builds artifacts asks the mesh to put an artifact into the thing
that artifact is needed to create. Building publishes to the store, and
the builder will not start without one — "a built artifact nobody can
fetch is not built".
This is the question the substrate record asks of every candidate: can
it grant itself the thing it provides? The store cannot create its own
database, the broker cannot create its own virtual host, and a registry
cannot grant itself a repository. The first two are why they are in the
bundle. This is the same sentence, unenforced.
So such a module names its image, exactly as the bundle names the three
a first node starts from. One that wants an interface or a tool server
beside it is a second module, mirrored the ordinary way once the first
is running — a real limit, and better said here than discovered on a
mesh new enough that nobody is watching it.
Refused at the manifest because the alternative is a build that never
returns.
The provision name is now a constant. A string compared in one place is
a convention; a string a rule turns on is a fact.
The rule was right about data and wrong about everything else. The
builder mounts the container runtime's socket, which is not its data,
does not belong to it, and must not be declared as one of its
directories — and the check refused the builder's own manifest.
Caught by the lab, though not honestly: the run was already going when
this went in, so the builder binary was rebuilt mid-run with the check
compiled into it and the failure was mine, not the mesh's. Confirmed
against the manifest directly rather than inferred from the log.
What it was protecting is real and stands — the fourteen mounts are all
declared. But enforcing it needs a way to tell "the directory my data
lives in" from "a machine facility I was granted", and the mesh has no
vocabulary for the second. `capabilities` is the closest thing and does
not name paths. That is a design decision, so it goes back to
04-ISSUES/026 rather than being invented here to make a check pass.
The check that every real manifest still parses is kept. It costs
nothing and it is how the next attempt at this finds out sooner.
Closes the half of 04-ISSUES/026 that would otherwise come back. The
fourteen mounts across the forge, the mail system, the store and the
object store are all declared now — but nothing said they had to be, so
they were right by coincidence and the next volume added would not be.
A bind mount whose source does not exist is created by the container
runtime, as root, with a mode it picks. So `owner` and `mode` — which
exist precisely so a module can say who its data belongs to — were
silently not applied to the only directories holding data.
And the rule written for exactly this case did not reach them. A
directory the mesh declared and no longer wants is kept, not removed,
when it holds anything the mesh did not put there (ADR 0030). That is
the answer to *what happens to my data when a module goes away*, and it
is written in terms of declared directories: an undeclared one sits
outside it, because the mesh does not know it is there.
Refused where it is written rather than on the machine, which cannot
tell the difference — by the time the host sees the mount it is being
asked to make a directory, which it is perfectly able to do. The fault
is in the manifest, so it is named at the manifest. Same argument as the
action refusal directly above it.
A path under a declared directory counts as declared, as do the files a
module already names: its own secrets, its grants, what it receives.
Every real manifest is checked to still parse, and the refusal bites.
The lab caught this: a module declaring a port and running no container
had its rule set opened on 20000 while its service sat on 9101. The
firewall reported success and blocked the thing it was told to admit,
which is the precise failure the filtering comment warns about, arrived
at from the other side.
Assignment was applied to every declared port. But a container's mapping
is the thing that translates, and where there is none the software binds
what it binds — the mesh choosing a number does not move the service, it
only makes the mesh wrong about where it is.
The declaration side already knew this: publishedOn rewrites container
ports and nothing else. Filtering did not, so the two disagreed about
the same fact. MachineSide is now the one derivation both follow.
It also fixes a second case nobody had hit yet: a mapping the manifest
wrote itself, like the mail system's 7080:80. That is passed through
untouched when composing, so assigning it a machine port would have
opened a rule on a port the container does not publish. Either side of
such a mapping now names it, and the host side is the answer — a module
may read `listens` as what its software binds or as what the machine
exposes, and both readings want the same number.
Recorded either way, assigned or not: the map means where this module's
port is on this machine, and every reader needs that answer regardless
of who chose it.
Tests bite — making it always assignable reproduces the lab failure.
novox/hq ADR 0038. A module cannot choose a port: it is written once and
assigned anywhere, so any number it picks is a guess about a machine it
has never seen. A database module met the mesh's own store on 5432 and
was told, by a container runtime three layers down, that the port was
already allocated.
The number used to appear three times in every module — the rule set,
what a consumer is told, and what the runtime publishes — agreeing only
because one person wrote all three. Now it appears once, in `listens`,
and the other two are derived: the container publishes `20000:5432`, the
consumer is told 20000, and the rule set opens 20000.
An assignment is made once and kept, as a credential is. A port that
moved on every declaration would restart both ends each time and hand a
consumer a number that was true when it was read.
Ports the protocol fixes — mail on 25, submission on 587, DNS on 53 —
say so, and are then claims: one holder per machine, and the second is
refused by name at assignment. That is the mechanism the mesh already
has for what is singular on a machine, pointed at ports.
A mapping written the long way is left exactly as it is. Some things
must be pinned by hand, and quietly overruling somebody who wrote both
halves would be worse than not offering the short form.
Still open, and known: the substrate is not a module, so the mesh has
never heard of its own store and cannot yet assign around it. That is
what 028 will still be about after this.
novox/hq 04-ISSUES/025. Every image reference in every example module
was sixty-four zeros — eighteen of them across five modules. Each
parsed, resolved, and composed into a declaration a host accepts, and
none could ever have started: the machine reaches `docker pull` and
stops. That is why those modules were written and not running, and no
check saw it because every check passed.
The host validates the shape of a reference and nothing more, which is
correct: verifying a digest exists means reaching a registry, and that
is the one thing a host must never have to do. So the last place that
could catch this is the wrong place to try.
The guard therefore sits where a declaration is composed, not where a
manifest is parsed. A file in a repository is allowed to await a pin —
the design already says the manifest in a repository names artifacts
while the manifest the mesh holds names digests, and the bundle works
exactly that way. What must never happen is a placeholder reaching a
machine, and composing is the last moment before one does.
Twelve third-party images resolved to real digests without pulling
anything, which is also the mechanism the open issue needs. Two
discoveries came free: mailu publishes to ghcr rather than Docker Hub,
so seven references named repositories that do not exist at all; and it
renamed roundcube to webmail, so that one would have failed even with
the right registry.
What stays a placeholder is the mesh's own provisioner images, which
genuinely have no digest until built and pushed — the bundle's problem,
legitimately unresolved here. The stand-in consumer now stands in with
a real image rather than an invented one.
I wrote that a test asserts the control plane's placeholder expression
and the host's still agree. None does, and none in this repository could
— a unit test here can only assert what this repository already
believes.
That is precisely the thing this project refuses to tolerate: a stated
rule with no way to check it, which costs more than no rule because
people believe it. Written by me, today, in the same file that closes a
gap of the same kind.
What actually proves it is the lab, and the comment now says so.
A contribution that is not a credential grant — a module offering
something to another on its own machine — has nobody to be identified
to, and was carrying an empty `as`. A field that is always present and
usually empty teaches a reader to ignore it, including when it is not.
novox/hq 04-ISSUES/023. A consumer was given its password, the address,
the port and where its credential lives, and still could not connect —
the user name was invented by the provisioner and recorded nowhere, and
the rest sat in a JSON binding that a program reading KEY=value cannot
use.
Both halves have the same cause: the mesh knew something and did not say
it.
**Who a consumer is, said once.** The provisioner used to derive
mesh_<node>_<module> and that string existed nowhere else — not in the
control plane, not in the binding, and above all not at the consumer,
which has to present it. Now the mesh derives it once and sends it to
both ends, so they agree by construction rather than by two conventions
that were the same on the day they were written. The provisioners refuse
to invent one if the mesh says nothing, because falling back to a name
of their own would create a role the consumer would never guess and
everything would report success.
**Bound values reach the file that needs them.** ${bound:provision:key}
is the symmetric twin of the sealed placeholder, and simpler: these
values are not secret, so the control plane fills them in before sending
and the host gains no field and learns no format. It stays
name-agnostic — at, as and from are true of any provision, and every
other key comes from what the provider said it serves.
The asymmetry it removes was backwards. The secret is the hard case,
because the mesh must not be able to read it, and the secret was the
part that already arrived.
Keycloak and Gitea now produce complete connections, asserted from the
manifests on disk rather than from fixtures: every part filled, no
placeholder surviving as a value, and the password still a hole only the
host can close. Three faults injected, each caught.
A module may not declare an action: the link may not carry a command to
run, and that bound is what limits a compromised control plane to shapes
it cannot turn into arbitrary code (novox/hq ADR 0005). The host
enforces it, correctly and in the right place.
But a module's resources reach a machine over the link, so a manifest
carrying an action was accepted here, stored, resolved, planned and
pushed — and refused on the machine, in the host's log, with nothing
connecting it back to the manifest that caused it.
The rule held. It was just unusable, which is the same shape as the
network shape earlier today: the refusal was right, arrived far from its
cause, and nobody was reading the log.
The refusal names the rule and what to do instead, because "you may not"
with no alternative is where a module author stops.
Found while checking a claim I had written in the coverage document —
that a module cannot declare one. It could; it just could not deliver
it. The document is corrected.