On an adopted hub the private network takes over the tunnel it finds rather
than running beside it (hq ADR 0105): two tunnels leave the mesh's unreachable
through the provider's filter, so no machine can ever join.
The node presents the found tunnel when it enrols, under the key it took as
its own; the inventory records it (node.tunnel, tunnel_peer — migration 0031)
and the mesh composes from it: the overlay's range is the adopted tunnel's,
the hub is placed at the tunnel's address on the tunnel's port, and every
peer the tunnel had is carried in the hub's peer list as a peer of the
tunnel, not a node of the mesh, until a node enrols with that key — which
then keeps the address the tunnel had for it. A fresh node never gets an
address the tunnel holds. The hub's declaration tells the host which unit to
take over; the host's account of carrying it is recorded and shown.
Every reader of the range follows the setting; nothing stores it. A found
tunnel under another key is recorded and not adopted, so ADR 0100's
non-overlap rule keeps applying where a tunnel is left running beside the
mesh's. A lab bed and test skeleton for "How it is checked" are under lab/.
Review found the first pass aliased its answer under both ends of a mapping, which is
wrong wherever two mappings share a number: the alias lands on a key belonging to another
mapping, the later write wins, and the filter and the container then disagree — the very
fault this change exists to close. Two reproduced cases: a module publishing 8080:80 beside
9090:8080 had an explicit setting silently overwritten; a module publishing 4001:80 beside
4002:80 composed both containers onto one machine port, where before it was safely refused.
Now a mapping's answer is filed once, under the end the module names in its listens — the
number the plan, the filter, the openings, the guard and the consumer all ask for — and a
key that names two mappings is refused in the same words as a setting that does.
Also: the guard assertion in the end-to-end test failed open when the resource was absent;
the plan-mirroring helper now says it stands in only where the plan does not allocate, and
the assertions it feeds are narrowed to the port under test.
A module publishing `2222:22` — the machine's own ssh daemon holds 22, so
the module takes 2222 and says so in `listens` — could not be moved. The
setting was read against the last segment of each mapping alone, so the
number the module uses everywhere else was refused as a port it does not
publish, and the node's every push failed for as long as the setting was
stored. The one key that was accepted, the container's own port, was then
read only when the container's mapping was rewritten: the mapping moved
and the ports map, the filter, the adopted node's openings, its guard and
what a consumer is told all stayed on the number the software had left.
Either end of a mapping now names it, and a given port comes back under
both, so every reader finds the same number under the key it holds.
Ambiguity is refused where it is real — one number naming two different
mappings, or the two ends of one mapping given two different numbers.
The mesh's own images declare theirs now, so the refusal ADR 0097 deferred is live.
The builder's own image and the examples take arguments with defaults; make builds
them, not the mesh.
A secrets object with one local name delivered no file. Two requirements could share
a local name. secret recover and the export could not tell two locals apart. The
recipe check missed continued lines and read heredoc bodies as bases. repo:tag@digest
kept the tag in the repository. ask now publishes mandatory, so a tool nothing serves
is said at once rather than after the wait.
A module serves tools under an account scoped to exactly that, and nothing else in
the mesh held an account that could ask one. The control plane does: ask publishes
on the RPC exchange with a private reply queue bound under its own name, checks the
correlation, prints the answer, and exits non-zero for a tool that answered with an
error or a module that never answered (novox/hq 04-ISSUES/049, ADR 0095).
secrets: maps a requirement to several files under local names. Each local name is
its own need, its own pair credential (the pair is keyed on it: migration 0027),
its own file on the consumer, its own holder at the provider (the identity with the
local name after it) and rotates apart from the others. The plain shape is
unchanged and every existing row is the credential it was (novox/hq 04-ISSUES/069,
ADR 0094).
secret accept grows --provider: the value is sealed to the consumer's node, the
provider's node and the operator's key, and the pair records origin 'accepted'.
An accepted pair is not remade when a key changes (the mesh does not hold the
value; the read is refused naming the remedy) and rotate refuses it (accepting a
new value is the rotation). The vault's third species has its entry
(novox/hq 04-ISSUES/070, ADR 0092).
The mesh kept one report per machine, replaced, so a resource nothing can ever apply
looked like a failure that had just happened, every reconcile interval, for ever.
The row now keeps when the current failure began and how many reports in a row have
said it — the same outcome, refusal and failed resources; anything different starts
again and a clean apply clears it. Three make the machine stuck, and status says so
beside the failure, in words and in JSON (novox/hq 04-ISSUES/065, ADR 0090).
From review: the export counted any operator-sealed row as recoverable, so a
secret sealed to a replaced key was reported as openable with the current one;
replacing the key counted orphans in one table of two; and a pair credential
held from two providers was recovered as whichever row came first. The export
now lists what the current key opens, what an earlier key opens, and what has
no copy; `secret recover` takes --provider and refuses ambiguity; files that
must not exist are created exclusively; one constructor builds the export for
the operator's file and the vault's disk alike.
novox/hq ADR 0085, amended: the mesh's root secrets — the store's superuser,
the broker's administrator, every secret a module holds for itself — were
sealed to a node key and nothing else, so a lost node took them with it.
Now the mesh records an operator's public sealing key and seals every own
secret to it as well, minted or accepted. The private half is written once
by `operator key new` to a file the operator keeps off the mesh; the mesh
holds one more blob per secret that it cannot open.
`secret recover` opens a secret with that key, to a 0600 file, from the
store or from an export; `secret export` writes every operator-sealed copy
as ciphertext. A module that `keeps` (the vault) is handed that export as a
declared file on its own disk, so recovery survives the store.
Secrets made before the key exists have no operator copy and are said so —
the plaintext was discarded — until each is issued again.
Two robustness fixes to the ADR 0083 cascade, from an adversarial review:
- It routed swept machines through sendTo, which is all-or-nothing — so
one swept machine's compose error failed the operator's named push and
skipped its --wait, the intolerance the main path exists to avoid
(ADR 0066). It now composes them through composeEach, exactly as the
named send does: a machine that cannot be worked out is a refusal in
the final report, and the rest are still sent. composeEach's tolerance
is already covered by TestOneUnresolvableNodeStillLetsTheRestBeSent.
- The fixed 4-round cap could stop a real cascade short in silence. The
loop is now bounded by the node count (a node is flushed once and never
revisited, so it cannot run longer) and says so if the guard is ever
hit, rather than passing over an unfinished cascade quietly.
Scope is unchanged: a named push still flushes every machine left behind,
per ADR 0083 as accepted.
The first cut compared a before/after snapshot of the named push — but
the provision is minted at assign or module-issue, before push runs, so
by push time the provider is already behind with no delta to detect.
Fixed to flush machines whose declaration differs from what they were
last SENT (the same Waiting path --behind uses), which is the honest
meaning of 'one push leaves the mesh consistent' (ADR 0083). Verified
live on a kept two-node mesh: pushing the consumer populates the
provider's grant and the vhost is minted.
A provision is minted while composing the consumer's node, and the
provider's grant list is a pure read of secrets already issued — so
pushing the consumer left the provider blind until somebody pushed it
again, with no signal to. A named push now captures what every machine
should be before composing, recomputes after, and sends the machines
whose declaration changed because of this push — by name, never
silently, converging over bounded rounds.
An adversarial review of the 055 fix found it encoded the wrong invariants, latent while
every mesh keeps its broker on the hub. Now: the address is the overlay name of the node
ASSIGNED a module claiming the mesh-broker seat (the hub stands in only while nothing holds
the seat — genesis); "on the overlay" is what whereEveryoneIs answers (resolved the
networking module), not "has an address"; a portless genesis address defaults to 5671
instead of silently disabling the path; a second `overlay place --hub` is refused rather
than last-write-wins; and `overlay place` says that earlier credentials keep their old
address. A test now binds the controller's own module.json to its seat, so deleting the
claim fails the suite.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
artifactStoreOnNetwork collapsed a failed inventory read into 'no
store', so a hiccup composed a declaration without the registry trust,
delivered by a push that reported success — and nothing recomposed the
machine until the next push. Seen once in three fresh runs of the
built-store-cross-node bed (run 11). Refused loudly instead: 'no store'
now only ever means the mesh has none.
Being on the private network is what grants a machine the right to pull from the mesh's
artifact store, so the module that puts a machine on the network writes the runtime's
trust — a merged /etc/docker/daemon.json naming the store's internal name under
insecure-registries, and a docker.service restart when that fact first lands. The registry
speaks plain HTTP because every path to it is already inside the overlay's encryption; the
provider is found, not configured — whichever module serves artifact-store, on whichever
machine holds it — and with no store on the network nothing is written, which is genesis.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The broker credential handed to a module named the genesis MESH_BROKER_ADDRESS — the
broker's public endpoint. That is reachable from the control-node itself but not routed
to another node, whose firewall admits only the overlay (from:mesh); a consumer on a
joined node timed out fetching the broker's certificate and never connected.
brokerReachableAt returns the broker's address as the given node can reach it: a node on
the overlay gets the hub's `.internal` name (which every node resolves and the firewall
admits, the fingerprint pin making the host swap safe for TLS); a node not yet on the
overlay — at genesis, before any `overlay place`, when the builder's account is issued —
keeps the genesis address, so bring-up is unchanged. Both credential paths (module issue
and builder issue) use it. This is the reachability half of novox/hq issue 055.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Opt-in (default 0, unchanged behaviour). A named push can wait until the node
reports it applied exactly this declaration, so an interactive push means "the
node is now what it was told". Not made the default: a push that recreates the
control plane would kill the waiting command running inside it.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A new 'package' artifact kind builds a module's own code on a public base image
and publishes it to the mesh's package registry by version (hq ADR 0076) — the
SDK above all, which the toolchain is built from and so cannot be built in the
toolchain. The credential a build needs to resolve or publish packages is
rendered as an .npmrc (basic auth, hq ADR 0048) and given to an image build as a
buildkit secret, never a layer, so a token is not baked into the toolchain image.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The logging added a line to stdout, and the genesis path parses the builder's
stdout as JSON — so the first log line broke the parse with "invalid character
'c'", the c from "[clone]". A build that had worked stopped working because of a
print statement.
The installer's runner captures stdout alone (cmd.Output), and the contract was
already stdout=result, stderr=everything else. The fix is to honour it: every
builder diagnostic — the step log, the per-command echo, the module path's own
lines — goes to stderr. Stdout carries only once.go's result JSON.
And a unit test now fails if any fmt.Print to stdout appears in the two builder
command files, except the three that belong there: the result, --version, and
--help. A guard, because this was invisible until a 20-minute run hit it, and the
same class of mistake should fail in milliseconds next time.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A build was silent from clone to publish, so a build in progress, one that failed
quietly, and a request that never arrived all looked identical — which cost a long
diagnosis against a running mesh chasing "the handler never fired".
Now: the handler announces a request the instant it lands. Build logs each phase
— clone, commit, manifest, bases, each artifact starting and finishing with what
it produced, resolve, done — through a Log callback that is nil-safe, so the tests
that pass none still build. And the Command runner echoes every command before it
runs, with where and how long it took, because on a hang the last line is exactly
the command it is stuck inside: "git clone waiting on a network that will not
answer" rather than "the builder did nothing".
The unreadable-request path prints to stdout now too, not stderr, so it shows in
docker logs without splitting streams — the split is what hid it.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
mesh-names, mesh-resolver and the names half of the overlay generators are gone.
They ran no software and could not be swapped for anything, which is the test of
whether something is a module at all — they existed because computed output
needed somewhere to live, and the control plane's only shape for output was a
module.
Now a module says where it wants what the mesh knows:
facts: { node-zones: /etc/mesh-resolver/nodes.conf }
and is given a file, under its own name, applied and removed like anything else
it declares. Two facts exist: node-names (a hosts file — exact names) and
node-zones (every machine as a wildcard, *.homer.internal is homer). Asking for
a fact the mesh does not compute is refused naming what would have worked,
because a daemon that starts and reads a file nobody wrote is a worse way to
find out.
The names ride with the network now: wireguard's manifest asks for node-names
into /etc/hosts, because being on the private network is what gives a machine a
name. networking no longer requires name-resolution — names are not a provision,
and the module that answered it ran nothing.
One behaviour inverted, deliberately: choosing another VPN used to drag
WireGuard in anyway, because only WireGuard provided the addressing the names
module required — the node-scope claim existed to at least make that loud. With
names as a fact there is nothing to drag in: tailscale assigned means tailscale,
alone. The claim still catches two VPNs assigned explicitly.
And a machine the mesh cannot place is left out of both files rather than named
at nothing: a name resolving to nothing hangs a connection, where an unknown
name fails at once and says so. In practice that is only ever a token issued and
not yet used — a machine that has announced itself has an address.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx