The mesh's own images declare theirs now, so the refusal ADR 0097 deferred is live.
The builder's own image and the examples take arguments with defaults; make builds
them, not the mesh.
A secrets object with one local name delivered no file. Two requirements could share
a local name. secret recover and the export could not tell two locals apart. The
recipe check missed continued lines and read heredoc bodies as bases. repo:tag@digest
kept the tag in the repository. ask now publishes mandatory, so a tool nothing serves
is said at once rather than after the wait.
A module serves tools under an account scoped to exactly that, and nothing else in
the mesh held an account that could ask one. The control plane does: ask publishes
on the RPC exchange with a private reply queue bound under its own name, checks the
correlation, prints the answer, and exits non-zero for a tool that answered with an
error or a module that never answered (novox/hq 04-ISSUES/049, ADR 0095).
secrets: maps a requirement to several files under local names. Each local name is
its own need, its own pair credential (the pair is keyed on it: migration 0027),
its own file on the consumer, its own holder at the provider (the identity with the
local name after it) and rotates apart from the others. The plain shape is
unchanged and every existing row is the credential it was (novox/hq 04-ISSUES/069,
ADR 0094).
secret accept grows --provider: the value is sealed to the consumer's node, the
provider's node and the operator's key, and the pair records origin 'accepted'.
An accepted pair is not remade when a key changes (the mesh does not hold the
value; the read is refused naming the remedy) and rotate refuses it (accepting a
new value is the rotation). The vault's third species has its entry
(novox/hq 04-ISSUES/070, ADR 0092).
The mesh kept one report per machine, replaced, so a resource nothing can ever apply
looked like a failure that had just happened, every reconcile interval, for ever.
The row now keeps when the current failure began and how many reports in a row have
said it — the same outcome, refusal and failed resources; anything different starts
again and a clean apply clears it. Three make the machine stuck, and status says so
beside the failure, in words and in JSON (novox/hq 04-ISSUES/065, ADR 0090).
From review: the export counted any operator-sealed row as recoverable, so a
secret sealed to a replaced key was reported as openable with the current one;
replacing the key counted orphans in one table of two; and a pair credential
held from two providers was recovered as whichever row came first. The export
now lists what the current key opens, what an earlier key opens, and what has
no copy; `secret recover` takes --provider and refuses ambiguity; files that
must not exist are created exclusively; one constructor builds the export for
the operator's file and the vault's disk alike.
novox/hq ADR 0085, amended: the mesh's root secrets — the store's superuser,
the broker's administrator, every secret a module holds for itself — were
sealed to a node key and nothing else, so a lost node took them with it.
Now the mesh records an operator's public sealing key and seals every own
secret to it as well, minted or accepted. The private half is written once
by `operator key new` to a file the operator keeps off the mesh; the mesh
holds one more blob per secret that it cannot open.
`secret recover` opens a secret with that key, to a 0600 file, from the
store or from an export; `secret export` writes every operator-sealed copy
as ciphertext. A module that `keeps` (the vault) is handed that export as a
declared file on its own disk, so recovery survives the store.
Secrets made before the key exists have no operator copy and are said so —
the plaintext was discarded — until each is issued again.
Two robustness fixes to the ADR 0083 cascade, from an adversarial review:
- It routed swept machines through sendTo, which is all-or-nothing — so
one swept machine's compose error failed the operator's named push and
skipped its --wait, the intolerance the main path exists to avoid
(ADR 0066). It now composes them through composeEach, exactly as the
named send does: a machine that cannot be worked out is a refusal in
the final report, and the rest are still sent. composeEach's tolerance
is already covered by TestOneUnresolvableNodeStillLetsTheRestBeSent.
- The fixed 4-round cap could stop a real cascade short in silence. The
loop is now bounded by the node count (a node is flushed once and never
revisited, so it cannot run longer) and says so if the guard is ever
hit, rather than passing over an unfinished cascade quietly.
Scope is unchanged: a named push still flushes every machine left behind,
per ADR 0083 as accepted.
The first cut compared a before/after snapshot of the named push — but
the provision is minted at assign or module-issue, before push runs, so
by push time the provider is already behind with no delta to detect.
Fixed to flush machines whose declaration differs from what they were
last SENT (the same Waiting path --behind uses), which is the honest
meaning of 'one push leaves the mesh consistent' (ADR 0083). Verified
live on a kept two-node mesh: pushing the consumer populates the
provider's grant and the vhost is minted.
A provision is minted while composing the consumer's node, and the
provider's grant list is a pure read of secrets already issued — so
pushing the consumer left the provider blind until somebody pushed it
again, with no signal to. A named push now captures what every machine
should be before composing, recomputes after, and sends the machines
whose declaration changed because of this push — by name, never
silently, converging over bounded rounds.
An adversarial review of the 055 fix found it encoded the wrong invariants, latent while
every mesh keeps its broker on the hub. Now: the address is the overlay name of the node
ASSIGNED a module claiming the mesh-broker seat (the hub stands in only while nothing holds
the seat — genesis); "on the overlay" is what whereEveryoneIs answers (resolved the
networking module), not "has an address"; a portless genesis address defaults to 5671
instead of silently disabling the path; a second `overlay place --hub` is refused rather
than last-write-wins; and `overlay place` says that earlier credentials keep their old
address. A test now binds the controller's own module.json to its seat, so deleting the
claim fails the suite.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
artifactStoreOnNetwork collapsed a failed inventory read into 'no
store', so a hiccup composed a declaration without the registry trust,
delivered by a push that reported success — and nothing recomposed the
machine until the next push. Seen once in three fresh runs of the
built-store-cross-node bed (run 11). Refused loudly instead: 'no store'
now only ever means the mesh has none.
Being on the private network is what grants a machine the right to pull from the mesh's
artifact store, so the module that puts a machine on the network writes the runtime's
trust — a merged /etc/docker/daemon.json naming the store's internal name under
insecure-registries, and a docker.service restart when that fact first lands. The registry
speaks plain HTTP because every path to it is already inside the overlay's encryption; the
provider is found, not configured — whichever module serves artifact-store, on whichever
machine holds it — and with no store on the network nothing is written, which is genesis.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The broker credential handed to a module named the genesis MESH_BROKER_ADDRESS — the
broker's public endpoint. That is reachable from the control-node itself but not routed
to another node, whose firewall admits only the overlay (from:mesh); a consumer on a
joined node timed out fetching the broker's certificate and never connected.
brokerReachableAt returns the broker's address as the given node can reach it: a node on
the overlay gets the hub's `.internal` name (which every node resolves and the firewall
admits, the fingerprint pin making the host swap safe for TLS); a node not yet on the
overlay — at genesis, before any `overlay place`, when the builder's account is issued —
keeps the genesis address, so bring-up is unchanged. Both credential paths (module issue
and builder issue) use it. This is the reachability half of novox/hq issue 055.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Opt-in (default 0, unchanged behaviour). A named push can wait until the node
reports it applied exactly this declaration, so an interactive push means "the
node is now what it was told". Not made the default: a push that recreates the
control plane would kill the waiting command running inside it.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A new 'package' artifact kind builds a module's own code on a public base image
and publishes it to the mesh's package registry by version (hq ADR 0076) — the
SDK above all, which the toolchain is built from and so cannot be built in the
toolchain. The credential a build needs to resolve or publish packages is
rendered as an .npmrc (basic auth, hq ADR 0048) and given to an image build as a
buildkit secret, never a layer, so a token is not baked into the toolchain image.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The logging added a line to stdout, and the genesis path parses the builder's
stdout as JSON — so the first log line broke the parse with "invalid character
'c'", the c from "[clone]". A build that had worked stopped working because of a
print statement.
The installer's runner captures stdout alone (cmd.Output), and the contract was
already stdout=result, stderr=everything else. The fix is to honour it: every
builder diagnostic — the step log, the per-command echo, the module path's own
lines — goes to stderr. Stdout carries only once.go's result JSON.
And a unit test now fails if any fmt.Print to stdout appears in the two builder
command files, except the three that belong there: the result, --version, and
--help. A guard, because this was invisible until a 20-minute run hit it, and the
same class of mistake should fail in milliseconds next time.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A build was silent from clone to publish, so a build in progress, one that failed
quietly, and a request that never arrived all looked identical — which cost a long
diagnosis against a running mesh chasing "the handler never fired".
Now: the handler announces a request the instant it lands. Build logs each phase
— clone, commit, manifest, bases, each artifact starting and finishing with what
it produced, resolve, done — through a Log callback that is nil-safe, so the tests
that pass none still build. And the Command runner echoes every command before it
runs, with where and how long it took, because on a hang the last line is exactly
the command it is stuck inside: "git clone waiting on a network that will not
answer" rather than "the builder did nothing".
The unreadable-request path prints to stdout now too, not stderr, so it shows in
docker logs without splitting streams — the split is what hid it.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
mesh-names, mesh-resolver and the names half of the overlay generators are gone.
They ran no software and could not be swapped for anything, which is the test of
whether something is a module at all — they existed because computed output
needed somewhere to live, and the control plane's only shape for output was a
module.
Now a module says where it wants what the mesh knows:
facts: { node-zones: /etc/mesh-resolver/nodes.conf }
and is given a file, under its own name, applied and removed like anything else
it declares. Two facts exist: node-names (a hosts file — exact names) and
node-zones (every machine as a wildcard, *.homer.internal is homer). Asking for
a fact the mesh does not compute is refused naming what would have worked,
because a daemon that starts and reads a file nobody wrote is a worse way to
find out.
The names ride with the network now: wireguard's manifest asks for node-names
into /etc/hosts, because being on the private network is what gives a machine a
name. networking no longer requires name-resolution — names are not a provision,
and the module that answered it ran nothing.
One behaviour inverted, deliberately: choosing another VPN used to drag
WireGuard in anyway, because only WireGuard provided the addressing the names
module required — the node-scope claim existed to at least make that loud. With
names as a fact there is nothing to drag in: tailscale assigned means tailscale,
alone. The claim still catches two VPNs assigned explicitly.
And a machine the mesh cannot place is left out of both files rather than named
at nothing: a name resolving to nothing hangs a connection, where an unknown
name fails at once and says so. In practice that is only ever a token issued and
not yet used — a machine that has announced itself has an address.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Its event queue is durable, so a running catalogue misses nothing. What it cannot
have is what was announced before it first ran — and on a fresh mesh that is never
arbitrary: the shared base, the store the catalogue runs on, and the catalogue
itself are each necessarily built BEFORE a catalogue exists to hear about them.
The graph's foundation is the part it never sees.
So it says it is catching up, and the control plane re-announces what it
recorded, oldest first, marked as a replay. Oldest first because a graph is built
in the order things happened: registering a module that stands on a base before
the base would point an edge at a version nothing has seen, and the shape of a
fresh mesh guarantees the base is both first and the one that was missed.
The replayer hands announcements back rather than publishing them, because the
wire belongs to the link package and a replay building its own events could drift
from what the builder emits — the one thing it must match exactly, since the
catalogue has a single handler for both.
Its own queue and its own consumer: two consumers on one queue split its
messages, and a catch-up request going to whichever half was not listening is a
gap that looks like a working mesh.
Toward novox/hq 04-ISSUES/050.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The builder announces a build with the resolved manifest, the path inside the
repository, and every artifact it stood on. The control plane received all of it
and kept none of it.
That was survivable while the catalogue heard the same announcement directly. It
stops being survivable the moment the catalogue was not there to hear it — which
on a fresh mesh is always, and always for the same modules: the shared base, the
store the catalogue runs on, and the catalogue itself are each necessarily built
BEFORE the catalogue exists to hear about them. The graph's foundation is the
part the graph never sees.
Replaying those builds needs what they said, not a summary. Without the manifest
there are no requires/provides edges; without `against` there are no build edges,
which are the ones that answer "a base moved, what must be rebuilt". A replay
carrying neither would restore the module list and leave the question the
catalogue exists for still wrong, while looking fixed.
Kept null rather than empty where a build predates this, so a replay can say it
is holding nothing instead of inventing an empty declaration for a module that
certainly had one. And `built_against`, not `built_on`: that column exists and
means the machine, which is a different fact about a different subject.
Toward novox/hq 04-ISSUES/050.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Rules are derived from what modules declare they listen on, and the substrate is
not a module. So the broker's port — the one every machine dials to enrol and to
receive every declaration it is ever sent — appeared in no ruleset the mesh has
ever generated.
Nothing caught it because a mesh of one never dials its own broker across the
network: the ruleset looks complete right up until a second machine tries to
join a firewalled anchor and is refused by the packet filter, during enrolment,
before the mesh can report anything about it. Assigning the firewall before
joining machines is both the natural order and the one that breaks.
It is a floor for the same reason ssh is. A machine nobody can reach cannot be
repaired; a machine the mesh cannot reach cannot be managed. Neither is a thing
any module asks for and neither may be derived away.
From anywhere rather than from the private network, deliberately: a node enrols
BEFORE it has an address on that network, so narrowing the rule to it would close
the door being knocked on.
The port is read from the broker this control plane was told about, so the
address handed out in a token and the port a machine must accept on stay one
fact. A mesh never told about a broker gets no such rule, rather than a broken
one — and cannot issue tokens either, which is where that surfaces.
Closes novox/hq 04-ISSUES/052.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A fingerprint written into a recipe names one particular copy of the base — the
copy on whichever machine the person typing it was using. On any other mesh that
copy has never existed, so the build stops on its first line with a message
about an image nobody can look up. Three modules in the catalogue were in
exactly that state, and the line each of them replaced was equally dead.
A module now names the module and artifact instead, and the mesh answers with
what it holds. The builder is still a thing that clones, builds and answers: the
answer travels with the question, because only the mesh knows what it has.
A base the mesh has not built is refused before anything is built, naming which
module has to exist first.
This is how a mesh is raised: the installer carries this program and runs it
once, before anything exists, to produce the control plane from the same
repository and path every later rebuild will use. What raises the mesh is then
the same thing that maintains it, rather than a second mechanism exercised once
per new mesh — which is how often enough to rot.
With nowhere to publish, an image stays in the machine's own runtime and is
named by the digest of its own configuration: the same identity the installer
has always used for the image it carried.
The builder says what it built and the catalogue decides whether that was an
upgrade. Only the control plane knows which machines run the thing, so it is
the one that acts — and what it does is a choice somebody recorded, not a
behaviour compiled in: record that they are behind, or send it, one machine at
a time or together.
Recording is the absence of an action rather than a second path: a machine not
running what the mesh would send it is already something the mesh reports.
Defaulted to recording. A mesh that rolls out everything it builds the moment
it builds it is reasonable to want and a bad thing to arrive by default — the
first module to inherit it would be the control plane, upgrading itself out
from under the push applying it.
An event whose origin reads a container id names something no other module
can look up. The mesh already knows the answer, and a module's environment
file is a file resource, so ${machine:name} reaches it with no composer change.
Answering and announcing are different acts. The reply goes to whoever asked and
is correlated to their request; the announcement says to the whole mesh that a
module now exists at a commit, which is what the catalogue places in the module
graph (novox/hq ADR 0072). A build nobody asked for still has to be announced, or
the graph knows less than the registry does.
What it was built on top of is read out of the build's own inputs rather than
declared, because a declared list drifts from what the code actually uses
(ADR 0009). These are artifact references, which is what a build input names;
resolving them to module-versions is the catalogue's work, since it is what knows
which module-version published which artifact.
Events ride the topic exchange, not the direct one nodes speak over, so the
builder's account is granted both: it must be able to answer and to announce.
The envelope is the sdk's, reproduced exactly — a second shape would be a second
thing for consumers to handle, and they are written against the first.
Announcing is not allowed to fail a build. The work was done and was answered; a
build reported as failed because saying so failed is a lie about it.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The builder cloned a repository and read the manifest at its root, which means one
repository per module. Nothing we have is shaped that way, so the builder could be
asked to build nothing that exists (novox/hq ADR 0069).
The path travels the whole way — named when asking, carried in the request, used
to read the manifest and as the context everything is produced from, echoed back
in the result, and recorded as part of where a module came from. Without that last
part the mesh could notice a module was behind its source and then be unable to
rebuild it, which is the worst of both.
A path climbing out of the clone is refused: a machine whose job is building other
people's repositories must not read whatever else is on its disk.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
0056 is 'the authority is the control plane, not a database'. A citation
pointing at the wrong decision is worse than none: it reads as corroboration.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A store connection string carries a password, and the control plane took it
from MESH_STORE_<CONTEXT> — an environment variable, which is readable in
`docker inspect`, in the process's own /proc entry, and in whatever composed
it. Every other module in the catalogue is given secret material as a file the
mesh sealed to the machine and the host wrote.
That difference is what stopped the control plane from being an ordinary module
(novox/hq ADR 0067). A manifest can put a sealed value into a file's `content`
with ${secret:…}; it has no substitution into a container's `env` at all. So a
control-plane manifest could be written with the password in it, or without the
setting — neither honest. The fix is not to change the manifest format but to
let the control plane read what everything else reads: a file.
MESH_STORE_<CONTEXT>_FILE names one. Exactly one of the two may be set; both is
refused rather than settled by precedence, because whichever won, the other
would still read as the setting in force and the process would be writing to a
store nobody expects. Trailing whitespace is trimmed — a file written by a
person or by a filled-in placeholder ends in a newline, and a newline inside a
URL is rejected several layers from anything that could explain it. Leading
whitespace is left, being a mangled value rather than a habit.
The no-leak property is kept and extended: a file that cannot be read names its
path, never its contents, and the parse failure now names whichever source was
used because a variable name and a path are not the secret.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The sibling of node public-domain, and the worse one: a placement is three facts
declared together, so an invocation that said none of them took all three away —
the endpoint every other machine dials, the site, and the hub. A mesh whose hub
was placed that way has no paths left, at the moment somebody was trying to look
at it.
--nothing keeps the real case (a machine that roams and opens every path itself)
sayable, by name.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
`status --json` emitted no JSON at all when a single node was unresolvable. A blocked
node is not on the private network, and a mesh whose hub is that node has no hub — which
came back through the reading as a refusal, so `status` printed nothing and `status
--json` put multi-line prose on stderr and not one byte on stdout. A machine-readable
interface that stops being machine-readable exactly when something is wrong is one nobody
can build an alarm on.
Why a machine cannot be worked out is read as data now, per machine, through the same
whoResolves the private network is built from — so this and the network agree about who
could not be resolved rather than deciding it twice. The private network failing to
compute is kept as a note beside it instead of ending the read: it is almost always a
consequence of those same refusals, and every question that does not depend on it is
still answered.
It reaches all three ways of saying it, from the one reading: the text form leads with it
because a machine here is in none of the answers below, the JSON carries `unresolved`
(always a list, never null) and `network`, and the page has a section of its own.
That also closes a silent success. A machine that resolves to nothing has nothing
computed for it, so there is nothing to compare it against and nothing it can be behind —
it appeared in no answer at all, and `status` reported a mesh where nothing could be sent
anywhere as "all doing what they were told".
statusAsJSON takes the whole reading now rather than a growing argument list, which is
what let an answer be added to the text form and forgotten here. The two are one
function's output in two shapes and must not be able to differ about what was asked.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
`node public-domain <name>` cleared the domain. It reads like a question — it is exactly
what anybody types to find out what the answer is — and it silently took every routed
name the node had. There is no output that makes up for that: by the time it prints, the
fact is gone, and the mesh cannot tell a person what a domain used to be.
The bare form reports now. Clearing is still a real thing to want — a machine that stops
facing the outside composes no names, and lab-versus-production is this one setting
(novox/hq ADR 0056) — so it keeps a way to be said, by name: `--clear`. A domain and
`--clear` together are refused rather than one of them silently winning.
The other `node` subcommands were checked. `add`, `list` and `show` write nothing they
were not asked to, so there is nothing to make consistent with.
`overlay place <node>` with no flags has the same shape — it clears the endpoint, the
site and the hub flag — and is deliberately left alone here. It is a verb rather than a
question and every caller passes flags, so the fix is a different judgement and belongs
in its own change.
The usage text gains the three forms, and `module forget`'s new flag, neither of which
it named before.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
An assignment was verified by resolving the node it was made on. The verify that matters
is resolution over all of them: a module offering a mesh-scoped provision stops offering
it the moment its own node stops resolving, so an assignment could be reported as fine
while it took that provision away from every consumer elsewhere. Those consumers were
then told "nothing in this mesh provides it", naming as the remedy a module that was
already assigned — a wrong answer about a machine nobody had touched.
That is novox/hq 04-ISSUES/017's shape exactly: an action succeeds into a state its own
verify rejects, and it does so because the action's own test is not the test the verify
uses. 017's remedy was to make them the same test, and this makes them the same test.
The assignment is still kept, and that is the other half of the decision. Assignment is
not an ordering: a consumer assigned before its provider does not resolve for as long as
it takes to assign the provider, and refusing the first half of a pair would make the
order somebody types two commands in part of the mesh's rules. So `assign` and `unassign`
now name every OTHER machine that cannot be worked out as things stand, in the mesh's own
words, beside whatever they already said about this one. It reports the state and never
claims causation — saying "this assignment broke laptop" would mean resolving the whole
mesh twice and would still be a guess about which of several changes did it.
It costs a resolution per machine. Assignment is a person typing a command, and being
told which machines this just blocked is worth more than the milliseconds.
This is what a four-node raise read as "bumping a module's version broke provider
recognition". It was neither the version nor the provider: nothing in this codebase reads
the version column, every lookup is keyed on the module name alone, and a test in the
previous commit now says so. It was one machine's set of assignments, and nothing said so.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF