Every module.json already declares a why for each port under listens,
but plan only ever used it to build the firewall's rule set — nothing
printed it. An operator deciding whether to assign a module had no way
to see what it would open without reading the manifest by hand.
plan <node> now prints each assigned module's listens entries — port,
protocol, source, and its why — right under the module line, so the
same text that feeds the firewall is visible at the point someone is
actually deciding whether to open it.
The map the control plane hands a resolution holds both: the machines, and every name
the mesh was told to route to whichever machine serves it. A container's hosts wants all
of it, so a routed name resolves to the proxy. A resolver's zones want only the machines:
told the mesh's suffix is its own it answers authoritatively for everything under it and
forwards none of it, so a routed name with the suffix appended — drive.example.test.internal
— is a name nobody will ever ask for, standing beside the machines and looking as real.
Found composing the resolver's first assignment on a live machine, before pushing it.
hq issue 111.
hal dnsmasq-app conversion, hq 08-connectivity. Converting the resolver from the module it
replaces made it forward what it cannot answer, which is what the predecessor's does, and
that found two things the controller did not say.
A resolver that forwards must not send a mesh name it does not know upstream: the
`node-zones` fact now carries `local=/<suffix>/` beside the wildcards, written here rather
than in the daemon's configuration because the suffix is the mesh's choice and this file is
the one place the mesh writes what it chose. The default lives in one helper now instead of
being spelled in two functions.
The predecessor points the container runtime's `dns` at the machine's own tunnel address —
a container cannot reach the machine's loopback. A module writing that key needs the
address, and `${machine:at}` is the machine's name; a runtime's resolver list cannot be a
name it would need that resolver to look up. So a module may say `${machine:address}`: what
`at` resolves to, read from the same names the hosts file and the wildcards are written
from, absent — and refused — off the network like `at` is.
The `mesh-resolver` and `resolver-data` constants go: nothing provided or consumed either,
the fact and `mesh-addressing` are the mechanism, and a requirement nothing provides is
refused at resolution.
Tests: the catalogue's dnsmasq, resolv-conf and resolved-split-dns manifests are parsed
and composed as a machine would receive them — fixed upstreams, no-resolv, 127.0.0.1, the
machines file, the runtime's key, the pair that decides what a machine asks refused on one
node; and on a real mesh the resolver's machines file is composed with a wildcard per
machine on the network and composed again without one that left, mirroring the hosts fact.
Review of the ADR 0105 build (hq ADR 0105). Four things it got wrong and one
path it lacked:
- A predecessor spoke's tunnel names one peer, the hub, routed the whole
range; recording refused it and the whole enrolment failed. Range-routed
peers are skipped now — only the hub's peers are ever carried.
- The range and the carried peers were conditions on the node being adopted,
so converging the hub would have renumbered the mesh and dropped the peers
still reaching it. They are facts of the tunnel record now, mode aside; the
takeover alone is declared to an adopted node. Converging the hub is refused
while a carried peer has not enrolled, naming it.
- A push composed a takeover for a hub whose address or endpoint disagreed
with the tunnel, which would have the host stop the found interface and
raise the mesh's where no peer listens. The graph refuses to compose it,
naming both and the placement that fixes it.
- The host's account said taken or not; "found down and the mesh's not up"
read as not taken. Three states now, and an account on every takeover.
- A hub that enrolled before this feature holds a key of its own, and
re-enrolling would rotate every key the mesh sealed credentials to. A node
now rekeys in a report, signed with its identity key over the key it
leaves, the key it takes and the tunnel; the mesh verifies against the live
key, refuses a stale or foreign proof, records key and tunnel, and moves a
hub to the tunnel's address. `overlay show` names the path for a hub that
found no tunnel.
Also: a carried IPv6 peer is routed /128, and identity.ForTest exists so the
link can be tested against a real identity store.
Three readers did not follow a moved foundation port (novox/hq 04-ISSUES/102),
and each took the control-node down in its own way: the control plane's own
store and broker connections, sealed at genesis with the port inside; and every
build the mesh ever recorded, kept as `<registry>:<port>/<module>/<artifact>@…`.
The control plane cannot open its own sealed connections to move a port, and it
cannot bind the store as a consumer would — a binding mints a credential. So its
settings get a third twin, `NAME_PORT`, read on top of the sealed value by the
store, the broker, the management API and the bus connection, and filled into
its container by a placeholder that names a seat, `${seat:mesh-store:5432}`,
from the node's given or mesh-assigned ports — never the manifest's number, and
empty when the mesh has nothing to add, so what genesis wrote stands. A value
that is still a placeholder is nothing said, aloud: the manifest naming it lands
in the next commit, once every control plane that composes it knows it.
A build is now recorded by digest and path — `artifact-store://<module>/<artifact>@…`
— and the store's address is composed in where a reference is used: the
declaration, the trust file, the bases a build is handed, a replay to the
catalogue. Over the network as `<node>.internal:<port>`; on the store's own node
before any network exists — every genesis push before its "network" step — by
loopback. A reference recorded before this, with an address, is re-routed the
same way when the mesh built it. The trust file and every provider's address
come from one derivation: the node's given port, over the mesh's assignment,
over the manifest's number.
novox/hq 04-ISSUES/102
On an adopted hub the private network takes over the tunnel it finds rather
than running beside it (hq ADR 0105): two tunnels leave the mesh's unreachable
through the provider's filter, so no machine can ever join.
The node presents the found tunnel when it enrols, under the key it took as
its own; the inventory records it (node.tunnel, tunnel_peer — migration 0031)
and the mesh composes from it: the overlay's range is the adopted tunnel's,
the hub is placed at the tunnel's address on the tunnel's port, and every
peer the tunnel had is carried in the hub's peer list as a peer of the
tunnel, not a node of the mesh, until a node enrols with that key — which
then keeps the address the tunnel had for it. A fresh node never gets an
address the tunnel holds. The hub's declaration tells the host which unit to
take over; the host's account of carrying it is recorded and shown.
Every reader of the range follows the setting; nothing stores it. A found
tunnel under another key is recorded and not adopted, so ADR 0100's
non-overlap rule keeps applying where a tunnel is left running beside the
mesh's. A lab bed and test skeleton for "How it is checked" are under lab/.
Review found the first pass aliased its answer under both ends of a mapping, which is
wrong wherever two mappings share a number: the alias lands on a key belonging to another
mapping, the later write wins, and the filter and the container then disagree — the very
fault this change exists to close. Two reproduced cases: a module publishing 8080:80 beside
9090:8080 had an explicit setting silently overwritten; a module publishing 4001:80 beside
4002:80 composed both containers onto one machine port, where before it was safely refused.
Now a mapping's answer is filed once, under the end the module names in its listens — the
number the plan, the filter, the openings, the guard and the consumer all ask for — and a
key that names two mappings is refused in the same words as a setting that does.
Also: the guard assertion in the end-to-end test failed open when the resource was absent;
the plan-mirroring helper now says it stands in only where the plan does not allocate, and
the assertions it feeds are narrowed to the port under test.
A module publishing `2222:22` — the machine's own ssh daemon holds 22, so
the module takes 2222 and says so in `listens` — could not be moved. The
setting was read against the last segment of each mapping alone, so the
number the module uses everywhere else was refused as a port it does not
publish, and the node's every push failed for as long as the setting was
stored. The one key that was accepted, the container's own port, was then
read only when the container's mapping was rewritten: the mapping moved
and the ports map, the filter, the adopted node's openings, its guard and
what a consumer is told all stayed on the number the software had left.
Either end of a mapping now names it, and a given port comes back under
both, so every reader finds the same number under the key it holds.
Ambiguity is refused where it is real — one number naming two different
mappings, or the two ends of one mapping given two different numbers.
The mesh's own images declare theirs now, so the refusal ADR 0097 deferred is live.
The builder's own image and the examples take arguments with defaults; make builds
them, not the mesh.
A secrets object with one local name delivered no file. Two requirements could share
a local name. secret recover and the export could not tell two locals apart. The
recipe check missed continued lines and read heredoc bodies as bases. repo:tag@digest
kept the tag in the repository. ask now publishes mandatory, so a tool nothing serves
is said at once rather than after the wait.
A module serves tools under an account scoped to exactly that, and nothing else in
the mesh held an account that could ask one. The control plane does: ask publishes
on the RPC exchange with a private reply queue bound under its own name, checks the
correlation, prints the answer, and exits non-zero for a tool that answered with an
error or a module that never answered (novox/hq 04-ISSUES/049, ADR 0095).
secrets: maps a requirement to several files under local names. Each local name is
its own need, its own pair credential (the pair is keyed on it: migration 0027),
its own file on the consumer, its own holder at the provider (the identity with the
local name after it) and rotates apart from the others. The plain shape is
unchanged and every existing row is the credential it was (novox/hq 04-ISSUES/069,
ADR 0094).
secret accept grows --provider: the value is sealed to the consumer's node, the
provider's node and the operator's key, and the pair records origin 'accepted'.
An accepted pair is not remade when a key changes (the mesh does not hold the
value; the read is refused naming the remedy) and rotate refuses it (accepting a
new value is the rotation). The vault's third species has its entry
(novox/hq 04-ISSUES/070, ADR 0092).
The mesh kept one report per machine, replaced, so a resource nothing can ever apply
looked like a failure that had just happened, every reconcile interval, for ever.
The row now keeps when the current failure began and how many reports in a row have
said it — the same outcome, refusal and failed resources; anything different starts
again and a clean apply clears it. Three make the machine stuck, and status says so
beside the failure, in words and in JSON (novox/hq 04-ISSUES/065, ADR 0090).
From review: the export counted any operator-sealed row as recoverable, so a
secret sealed to a replaced key was reported as openable with the current one;
replacing the key counted orphans in one table of two; and a pair credential
held from two providers was recovered as whichever row came first. The export
now lists what the current key opens, what an earlier key opens, and what has
no copy; `secret recover` takes --provider and refuses ambiguity; files that
must not exist are created exclusively; one constructor builds the export for
the operator's file and the vault's disk alike.
novox/hq ADR 0085, amended: the mesh's root secrets — the store's superuser,
the broker's administrator, every secret a module holds for itself — were
sealed to a node key and nothing else, so a lost node took them with it.
Now the mesh records an operator's public sealing key and seals every own
secret to it as well, minted or accepted. The private half is written once
by `operator key new` to a file the operator keeps off the mesh; the mesh
holds one more blob per secret that it cannot open.
`secret recover` opens a secret with that key, to a 0600 file, from the
store or from an export; `secret export` writes every operator-sealed copy
as ciphertext. A module that `keeps` (the vault) is handed that export as a
declared file on its own disk, so recovery survives the store.
Secrets made before the key exists have no operator copy and are said so —
the plaintext was discarded — until each is issued again.
Two robustness fixes to the ADR 0083 cascade, from an adversarial review:
- It routed swept machines through sendTo, which is all-or-nothing — so
one swept machine's compose error failed the operator's named push and
skipped its --wait, the intolerance the main path exists to avoid
(ADR 0066). It now composes them through composeEach, exactly as the
named send does: a machine that cannot be worked out is a refusal in
the final report, and the rest are still sent. composeEach's tolerance
is already covered by TestOneUnresolvableNodeStillLetsTheRestBeSent.
- The fixed 4-round cap could stop a real cascade short in silence. The
loop is now bounded by the node count (a node is flushed once and never
revisited, so it cannot run longer) and says so if the guard is ever
hit, rather than passing over an unfinished cascade quietly.
Scope is unchanged: a named push still flushes every machine left behind,
per ADR 0083 as accepted.
The first cut compared a before/after snapshot of the named push — but
the provision is minted at assign or module-issue, before push runs, so
by push time the provider is already behind with no delta to detect.
Fixed to flush machines whose declaration differs from what they were
last SENT (the same Waiting path --behind uses), which is the honest
meaning of 'one push leaves the mesh consistent' (ADR 0083). Verified
live on a kept two-node mesh: pushing the consumer populates the
provider's grant and the vhost is minted.
A provision is minted while composing the consumer's node, and the
provider's grant list is a pure read of secrets already issued — so
pushing the consumer left the provider blind until somebody pushed it
again, with no signal to. A named push now captures what every machine
should be before composing, recomputes after, and sends the machines
whose declaration changed because of this push — by name, never
silently, converging over bounded rounds.
An adversarial review of the 055 fix found it encoded the wrong invariants, latent while
every mesh keeps its broker on the hub. Now: the address is the overlay name of the node
ASSIGNED a module claiming the mesh-broker seat (the hub stands in only while nothing holds
the seat — genesis); "on the overlay" is what whereEveryoneIs answers (resolved the
networking module), not "has an address"; a portless genesis address defaults to 5671
instead of silently disabling the path; a second `overlay place --hub` is refused rather
than last-write-wins; and `overlay place` says that earlier credentials keep their old
address. A test now binds the controller's own module.json to its seat, so deleting the
claim fails the suite.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
artifactStoreOnNetwork collapsed a failed inventory read into 'no
store', so a hiccup composed a declaration without the registry trust,
delivered by a push that reported success — and nothing recomposed the
machine until the next push. Seen once in three fresh runs of the
built-store-cross-node bed (run 11). Refused loudly instead: 'no store'
now only ever means the mesh has none.
Being on the private network is what grants a machine the right to pull from the mesh's
artifact store, so the module that puts a machine on the network writes the runtime's
trust — a merged /etc/docker/daemon.json naming the store's internal name under
insecure-registries, and a docker.service restart when that fact first lands. The registry
speaks plain HTTP because every path to it is already inside the overlay's encryption; the
provider is found, not configured — whichever module serves artifact-store, on whichever
machine holds it — and with no store on the network nothing is written, which is genesis.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The broker credential handed to a module named the genesis MESH_BROKER_ADDRESS — the
broker's public endpoint. That is reachable from the control-node itself but not routed
to another node, whose firewall admits only the overlay (from:mesh); a consumer on a
joined node timed out fetching the broker's certificate and never connected.
brokerReachableAt returns the broker's address as the given node can reach it: a node on
the overlay gets the hub's `.internal` name (which every node resolves and the firewall
admits, the fingerprint pin making the host swap safe for TLS); a node not yet on the
overlay — at genesis, before any `overlay place`, when the builder's account is issued —
keeps the genesis address, so bring-up is unchanged. Both credential paths (module issue
and builder issue) use it. This is the reachability half of novox/hq issue 055.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx