A secrets object with one local name delivered no file. Two requirements could share
a local name. secret recover and the export could not tell two locals apart. The
recipe check missed continued lines and read heredoc bodies as bases. repo:tag@digest
kept the tag in the repository. ask now publishes mandatory, so a tool nothing serves
is said at once rather than after the wait.
build.on takes {arg, image@sha256:…} beside {arg, module, artifact}: the image is
copied into the mesh's registry before the build (ADR 0096) and the recipe reads the
copy from the argument. A FROM or COPY --from naming a registry image the manifest
did not declare is refused before the build, naming it and the remedy; stages,
declared arguments and scratch are not fetches (novox/hq 04-ISSUES/064, ADR 0097).
The lab's vault refused a declaration naming two files with one identity: the
two secrets of one consumer. The holder's suffix is in the resource id and the path now.
Two consumers of one same-node provision produced two raw needs and, fanned out per
consumer, four — the same credential twice for each. Harmless, since a pair is one
row however often it is asked for, and wrong all the same.
The lab's two-secrets consumer was given one credential and no file: the expansion
ran on the resolver's walk over names, on whichever module mentioned the provision
first, and the per-consumer pass copied that. It expands in that pass now, and a
test has two consumers of one provision, one keeping one file and one keeping two.
secrets: maps a requirement to several files under local names. Each local name is
its own need, its own pair credential (the pair is keyed on it: migration 0027),
its own file on the consumer, its own holder at the provider (the identity with the
local name after it) and rotates apart from the others. The plain shape is
unchanged and every existing row is the credential it was (novox/hq 04-ISSUES/069,
ADR 0094).
The host's error text may carry a duration or a counter, and a resource looping on
it would never have read as stuck. The previous row is read and compared here.
Stuck needs a start to say. A container may mount the file a binding lands in; the
runtime socket is declared under both of its spellings; the catalogue-wide test
takes MESH_CATALOG.
Brought back from 53eb000, withdrawn because it refused the builder's mount of the
container runtime's socket. A path is declared as the module's own (a directory or
file resource, or where a secret, grant or contribution lands), as the operator's
(an accesses entry, ADR 0051), or as the machine's (a facility a declared capability
grants: container-runtime grants its socket). Every catalogue manifest passes, and
a test says so (novox/hq 04-ISSUES/026, ADR 0091).
The control plane runs as 65534 and crash-looped on permission denied the
first time its credentials were mounted as files the host wrote as root at
0600 — the env-file shape hid this because the daemon reads an env-file on
the host side. The composer now gives a module's secret files the owner the
manifest names.
The broker settings take a _FILE twin like the store connections; the
catalogue engine refuses a secret placeholder in a container's env and a
secret-carrying env-file unless the container says why with
secrets-in-environment, which stays in the catalogue and never reaches the
machine.
From review: the export counted any operator-sealed row as recoverable, so a
secret sealed to a replaced key was reported as openable with the current one;
replacing the key counted orphans in one table of two; and a pair credential
held from two providers was recovered as whichever row came first. The export
now lists what the current key opens, what an earlier key opens, and what has
no copy; `secret recover` takes --provider and refuses ambiguity; files that
must not exist are created exclusively; one constructor builds the export for
the operator's file and the vault's disk alike.
The secret the vault provides a module is the credential of the consumer↔vault
pair, and so is every credential a provider grants; sealing only own secrets
to the operator left exactly those unrecoverable. Same column, same call; the
export and `secret recover` address a pair by consumer node, module and the
provision's name, and say which kind each entry is.
novox/hq ADR 0085, amended: the mesh's root secrets — the store's superuser,
the broker's administrator, every secret a module holds for itself — were
sealed to a node key and nothing else, so a lost node took them with it.
Now the mesh records an operator's public sealing key and seals every own
secret to it as well, minted or accepted. The private half is written once
by `operator key new` to a file the operator keeps off the mesh; the mesh
holds one more blob per secret that it cannot open.
`secret recover` opens a secret with that key, to a 0600 file, from the
store or from an export; `secret export` writes every operator-sealed copy
as ciphertext. A module that `keeps` (the vault) is handed that export as a
declared file on its own disk, so recovery survives the store.
Secrets made before the key exists have no operator copy and are said so —
the plaintext was discarded — until each is issued again.
The modules the mesh runs were under examples/modules/, which framed
the real catalogue as illustrations of a control-plane package. They
are neither examples nor the control plane's — they are the mesh's own
catalogue, and they now live in their own repository (novox/mesh-catalog),
consumed as a build source like any other.
The engine that reads them stays here (internal/catalogue): the control
plane owns the manifest contract; the data does not belong beside it.
Removed with them: modules_test.go and parseall_test.go, which validated
the example manifests against the parser. That validation logically
follows the catalogue to mesh-catalog, but it imports internal/catalogue,
so re-homing it needs the parser exported from internal/ first — a
deliberate follow-up, not done here. Until then the pipeline is the gate,
and internal/catalogue's own inline tests still cover the parser.
Answers novox/hq ADR 0030's open tier-4 question — where the catalogue
lives — in favour of one flat mesh-catalog repository.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The broker's amqps port is opened from anywhere so a node can enrol
before it has an overlay address — but only in the input chain. The
broker is a published container port, so a cross-node dial is DNAT'd
and forwarded, never reaching input; it survived on the first
connection's conntrack entry and no more. Adopting the foundation's own
broker restarts it, dropping that entry, after which a joined node
could never receive another declaration. The forward chain now carries
the foundation ports too, from anywhere, matching their input rule.
Intermittent in the built-store-cross-node bed: it passed whenever the
broker did not happen to restart after the joined node first connected.
An adversarial review of the 055 fix found it encoded the wrong invariants, latent while
every mesh keeps its broker on the hub. Now: the address is the overlay name of the node
ASSIGNED a module claiming the mesh-broker seat (the hub stands in only while nothing holds
the seat — genesis); "on the overlay" is what whereEveryoneIs answers (resolved the
networking module), not "has an address"; a portless genesis address defaults to 5671
instead of silently disabling the path; a second `overlay place --hub` is refused rather
than last-write-wins; and `overlay place` says that earlier credentials keep their old
address. A test now binds the controller's own module.json to its seat, so deleting the
claim fails the suite.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A resource composed in code carries restart-on as []string; the rename
only read []any, so the overlay's registry-trust reload kept its bare
reference, pointed at nothing, and the runtime was never restarted —
the trust was on disk and not in the daemon, with every check passing.
Diagnosed on the built-store-cross-node bed, run 8 (issues 042/048).
Renames the module's own claim the-controller -> mesh-controller (the seat is the server,
ADR 0079), and adds TestAFoundationModuleCannotBeRaisedOnASecondNode asserting each
foundation module's second assignment is refused with 'one per mesh'. Closes hq issue 056.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A new 'package' artifact kind builds a module's own code on a public base image
and publishes it to the mesh's package registry by version (hq ADR 0076) — the
SDK above all, which the toolchain is built from and so cannot be built in the
toolchain. The credential a build needs to resolve or publish packages is
rendered as an .npmrc (basic auth, hq ADR 0048) and given to an image build as a
buildkit secret, never a layer, so a token is not baked into the toolchain image.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
mesh-names, mesh-resolver and the names half of the overlay generators are gone.
They ran no software and could not be swapped for anything, which is the test of
whether something is a module at all — they existed because computed output
needed somewhere to live, and the control plane's only shape for output was a
module.
Now a module says where it wants what the mesh knows:
facts: { node-zones: /etc/mesh-resolver/nodes.conf }
and is given a file, under its own name, applied and removed like anything else
it declares. Two facts exist: node-names (a hosts file — exact names) and
node-zones (every machine as a wildcard, *.homer.internal is homer). Asking for
a fact the mesh does not compute is refused naming what would have worked,
because a daemon that starts and reads a file nobody wrote is a worse way to
find out.
The names ride with the network now: wireguard's manifest asks for node-names
into /etc/hosts, because being on the private network is what gives a machine a
name. networking no longer requires name-resolution — names are not a provision,
and the module that answered it ran nothing.
One behaviour inverted, deliberately: choosing another VPN used to drag
WireGuard in anyway, because only WireGuard provided the addressing the names
module required — the node-scope claim existed to at least make that loud. With
names as a fact there is nothing to drag in: tailscale assigned means tailscale,
alone. The claim still catches two VPNs assigned explicitly.
And a machine the mesh cannot place is left out of both files rather than named
at nothing: a name resolving to nothing hangs a connection, where an unknown
name fails at once and says so. In practice that is only ever a token issued and
not yet used — a machine that has announced itself has an address.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The mesh assigns the machine-side port and a module does not choose one (ADR
0038). For a container that is invisible: the mesh rewrites ports into
assigned:wanted, the software binds the number it always bound, and the machine
publishes another.
A process has no such layer. It runs on the machine, there is nothing to rewrite,
and it binds whatever its configuration says. So every process bound the number
written in its own config, two modules declaring the same one would collide, and
the mesh's whole reason for assigning ports was defeated by the resource kind
that most needs it — introduced, by me, three commits ago.
So a module asks. ${port:8080} is "the machine-side port you gave me for the 8080
I said I listen on", written into its own configuration exactly as an address it
was bound to is.
Asking about a port it never declared is refused, and the refusal says what it
did declare: the module is asking about something the mesh has no opinion on, and
answering would put a guess into a configuration file as a port number. With
nothing assigned yet it is told what it asked for, so a mesh that has made no
assignment still composes something coherent rather than writing a zero.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The bundle recipe: the one that both builds and packs. An archive packs a
directory as it stands, so shipping compiled output meant compiling somewhere
first — which meant a Dockerfile repeating the same incantation in every module.
Two base arguments with no defaults, a working directory chosen so the SDK
resolves upward, the compiler invoked by absolute path because the usual symlink
is resolved away when the base image is assembled, a second stage, an environment
variable naming the entrypoints. Most of the catalogue is unconverted and that is
why; two conversions done in one session were each wrong twice with a working
example open in the next window.
A bundle says a language and a list of entrypoints. The mesh knows what the
language implies. Anything a module could override there it would be writing a
Dockerfile to override, so a toolchain is deliberately not configurable.
Declared rather than inferred, both of them: guessing the language from which
files are present makes a build depend on a directory listing, and guessing the
entrypoints makes it change meaning when somebody adds a helper.
A toolchain the mesh does not hold is refused before anything is compiled, naming
what to build first — the same treatment a missing base already gets, because it
is the same question and somebody can answer it. A language the mesh does not
build is refused saying what would have worked, since the author is usually one
word away.
The list of languages is closed and adding to it is a decision. Every language is
another implementation of the contracts every module shares, and those change
rarely and cascade when they do (ADR 0039) — a mesh whose SDKs disagree about the
envelope fails by ignoring messages rather than by failing to compile.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Rules are derived from what modules declare they listen on, and the substrate is
not a module. So the broker's port — the one every machine dials to enrol and to
receive every declaration it is ever sent — appeared in no ruleset the mesh has
ever generated.
Nothing caught it because a mesh of one never dials its own broker across the
network: the ruleset looks complete right up until a second machine tries to
join a firewalled anchor and is refused by the packet filter, during enrolment,
before the mesh can report anything about it. Assigning the firewall before
joining machines is both the natural order and the one that breaks.
It is a floor for the same reason ssh is. A machine nobody can reach cannot be
repaired; a machine the mesh cannot reach cannot be managed. Neither is a thing
any module asks for and neither may be derived away.
From anywhere rather than from the private network, deliberately: a node enrols
BEFORE it has an address on that network, so narrowing the rule to it would close
the door being knocked on.
The port is read from the broker this control plane was told about, so the
address handed out in a token and the port a machine must accept on stay one
fact. A mesh never told about a broker gets no such rule, rather than a broken
one — and cannot issue tokens either, which is where that surfaces.
Closes novox/hq 04-ISSUES/052.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Two sibling branches resolve a provision answered by the consumer's own machine.
The one for a node-scoped provider falls back to loopback when the machine is on
no private network, with a comment saying why and a test holding it. The one for
a mesh-scoped provider passed node.At straight through, and nothing noticed
because nothing had yet composed a host out of it.
The mesh's own artifact store is mesh-scoped and sits on the same machine as the
builder that pushes to it. Give the builder the address from its binding and it
gets MESH_REGISTRY=:5000 — a name with no host, written into its environment
without complaint. It surfaces much later as
":5000/mesh-tools/build" is not a valid repository/tag
which is a message about a tag for a fault in how a binding was resolved, on a
machine several steps from the decision.
A machine off the private network still reaches itself, which is what the
neighbouring branch already said. The test fails without the fix, showing the
empty address rather than only the symptom.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The floor allowed ssh from the mesh's addresses, and from everywhere on a
machine that faces outward. On a machine the mesh knows no addresses for it
emitted neither — so the chain dropped by default and ssh was simply shut.
That is the first machine anybody adopts: reached over the network, with the
port needed to fix it closed by the act of adopting it. Found by reading the
rules off a live machine rather than trusting the generator.
Two faults, opposite directions, both in issue 047.
There was no forward chain, on the reasoning that dropping there stops every
container the runtime allowed. The first half is true; the conclusion was not. A
published port is redirected and then forwarded, so it never reaches the input
chain — the firewall was silent about the ports most worth protecting. The way
through is the one the system being replaced already used: deny by default, then
allow the runtime's own networks explicitly. A forwarded rule matches what the
client originally asked for, because the destination has been rewritten by the
time the chain sees it.
And ssh is now a floor nothing derives. Every other line comes from what is
assigned, which is the point — but a mesh part-way through adopting a machine
has been assigned almost nothing, so what it computed was a chain that shut the
port used to fix it. From the mesh always; from outside on a machine that faces
outward, because that is the way in when the private network is what broke.
Rehearsed on three machines: a docker-published port declared mesh-only is now
reachable from inside the mesh and refused from outside. Before, it was
reachable from both.
A fingerprint written into a recipe names one particular copy of the base — the
copy on whichever machine the person typing it was using. On any other mesh that
copy has never existed, so the build stops on its first line with a message
about an image nobody can look up. Three modules in the catalogue were in
exactly that state, and the line each of them replaced was equally dead.
A module now names the module and artifact instead, and the mesh answers with
what it holds. The builder is still a thing that clones, builds and answers: the
answer travels with the question, because only the mesh knows what it has.
A base the mesh has not built is refused before anything is built, naming which
module has to exist first.
A container naming an artifact is a module saying the mesh builds this. Until a
build publishes one there is nothing to run — and what reached the machine was
an unresolved field, which its language has no room for, so it refused the whole
declaration and reported that a container does not use "artifact". That reads
as a broken manifest. It is not broken, it is unbuilt, and only the mesh can
tell those apart.
Found by the four-machine bed, which assigns modules the mesh has not built.
0056 is 'the authority is the control plane, not a database'. A citation
pointing at the wrong decision is worse than none: it reads as corroboration.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
An authority inside the mesh is reached at <machine>.internal, so its own
certificate must be issued for that name — and it is the one module that cannot
be told its name by a binding, because it provides rather than requires. Written
as a literal it would be one deployment's machine name in a manifest, which is
what ADR 0056 exists to remove.
${machine:name} and ${machine:at}, beside the bound values and refused the same
way. An address the machine does not have is named here rather than discovered
later as a certificate nobody can verify.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The co-located fix could not reach a grant assembled for a consumer on another
machine: ContributionsFrom never sees a port map, so the proxy was told the
workload's software port and dialled a number that machine never published. The
consumer's own assignments are fetched where the grant is built and applied
there. The same fault as 038, one node over.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
servedOnThisMachine stopped at the first module whose `serves` mentioned the
provision, even when that entry was empty and there was therefore no fact to give
a consumer. here() had always kept looking in that case, and a set where one
module names a provision without describing it and another describes it is exactly
where the difference shows. Restore the search.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF