The first cut compared a before/after snapshot of the named push — but
the provision is minted at assign or module-issue, before push runs, so
by push time the provider is already behind with no delta to detect.
Fixed to flush machines whose declaration differs from what they were
last SENT (the same Waiting path --behind uses), which is the honest
meaning of 'one push leaves the mesh consistent' (ADR 0083). Verified
live on a kept two-node mesh: pushing the consumer populates the
provider's grant and the vhost is minted.
The broker's amqps port is opened from anywhere so a node can enrol
before it has an overlay address — but only in the input chain. The
broker is a published container port, so a cross-node dial is DNAT'd
and forwarded, never reaching input; it survived on the first
connection's conntrack entry and no more. Adopting the foundation's own
broker restarts it, dropping that entry, after which a joined node
could never receive another declaration. The forward chain now carries
the foundation ports too, from anywhere, matching their input rule.
Intermittent in the built-store-cross-node bed: it passed whenever the
broker did not happen to restart after the joined node first connected.
A provision is minted while composing the consumer's node, and the
provider's grant list is a pure read of secrets already issued — so
pushing the consumer left the provider blind until somebody pushed it
again, with no signal to. A named push now captures what every machine
should be before composing, recomputes after, and sends the machines
whose declaration changed because of this push — by name, never
silently, converging over bounded rounds.
An adversarial review of the 055 fix found it encoded the wrong invariants, latent while
every mesh keeps its broker on the hub. Now: the address is the overlay name of the node
ASSIGNED a module claiming the mesh-broker seat (the hub stands in only while nothing holds
the seat — genesis); "on the overlay" is what whereEveryoneIs answers (resolved the
networking module), not "has an address"; a portless genesis address defaults to 5671
instead of silently disabling the path; a second `overlay place --hub` is refused rather
than last-write-wins; and `overlay place` says that earlier credentials keep their old
address. A test now binds the controller's own module.json to its seat, so deleting the
claim fails the suite.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
artifactStoreOnNetwork collapsed a failed inventory read into 'no
store', so a hiccup composed a declaration without the registry trust,
delivered by a push that reported success — and nothing recomposed the
machine until the next push. Seen once in three fresh runs of the
built-store-cross-node bed (run 11). Refused loudly instead: 'no store'
now only ever means the mesh has none.
A resource composed in code carries restart-on as []string; the rename
only read []any, so the overlay's registry-trust reload kept its bare
reference, pointed at nothing, and the runtime was never restarted —
the trust was on disk and not in the daemon, with every check passing.
Diagnosed on the built-store-cross-node bed, run 8 (issues 042/048).
Being on the private network is what grants a machine the right to pull from the mesh's
artifact store, so the module that puts a machine on the network writes the runtime's
trust — a merged /etc/docker/daemon.json naming the store's internal name under
insecure-registries, and a docker.service restart when that fact first lands. The registry
speaks plain HTTP because every path to it is already inside the overlay's encryption; the
provider is found, not configured — whichever module serves artifact-store, on whichever
machine holds it — and with no store on the network nothing is written, which is genesis.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Renames the module's own claim the-controller -> mesh-controller (the seat is the server,
ADR 0079), and adds TestAFoundationModuleCannotBeRaisedOnASecondNode asserting each
foundation module's second assignment is refused with 'one per mesh'. Closes hq issue 056.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The broker credential handed to a module named the genesis MESH_BROKER_ADDRESS — the
broker's public endpoint. That is reachable from the control-node itself but not routed
to another node, whose firewall admits only the overlay (from:mesh); a consumer on a
joined node timed out fetching the broker's certificate and never connected.
brokerReachableAt returns the broker's address as the given node can reach it: a node on
the overlay gets the hub's `.internal` name (which every node resolves and the firewall
admits, the fingerprint pin making the host swap safe for TLS); a node not yet on the
overlay — at genesis, before any `overlay place`, when the builder's account is issued —
keeps the genesis address, so bring-up is unchanged. Both credential paths (module issue
and builder issue) use it. This is the reachability half of novox/hq issue 055.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Opt-in (default 0, unchanged behaviour). A named push can wait until the node
reports it applied exactly this declaration, so an interactive push means "the
node is now what it was told". Not made the default: a push that recreates the
control plane would kill the waiting command running inside it.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A per-run .npmrc in every build context put a changing credential in COPY . . of
modules that resolve no mesh package — a non-deterministic image (a needless
rollout every build, which recreated the control plane) and a credential in a
build stage. Now it is written only for a package artifact or an image whose
Dockerfile names .npmrc.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A new 'package' artifact kind builds a module's own code on a public base image
and publishes it to the mesh's package registry by version (hq ADR 0076) — the
SDK above all, which the toolchain is built from and so cannot be built in the
toolchain. The credential a build needs to resolve or publish packages is
rendered as an .npmrc (basic auth, hq ADR 0048) and given to an image build as a
buildkit secret, never a layer, so a token is not baked into the toolchain image.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The logging added a line to stdout, and the genesis path parses the builder's
stdout as JSON — so the first log line broke the parse with "invalid character
'c'", the c from "[clone]". A build that had worked stopped working because of a
print statement.
The installer's runner captures stdout alone (cmd.Output), and the contract was
already stdout=result, stderr=everything else. The fix is to honour it: every
builder diagnostic — the step log, the per-command echo, the module path's own
lines — goes to stderr. Stdout carries only once.go's result JSON.
And a unit test now fails if any fmt.Print to stdout appears in the two builder
command files, except the three that belong there: the result, --version, and
--help. A guard, because this was invisible until a 20-minute run hit it, and the
same class of mistake should fail in milliseconds next time.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A build was silent from clone to publish, so a build in progress, one that failed
quietly, and a request that never arrived all looked identical — which cost a long
diagnosis against a running mesh chasing "the handler never fired".
Now: the handler announces a request the instant it lands. Build logs each phase
— clone, commit, manifest, bases, each artifact starting and finishing with what
it produced, resolve, done — through a Log callback that is nil-safe, so the tests
that pass none still build. And the Command runner echoes every command before it
runs, with where and how long it took, because on a hang the last line is exactly
the command it is stuck inside: "git clone waiting on a network that will not
answer" rather than "the builder did nothing".
The unreadable-request path prints to stdout now too, not stderr, so it shows in
docker logs without splitting streams — the split is what hid it.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The hosts-file writing moved to the node-names fact; what stays is the suffix
(configurable, MESH_INTERNAL_SUFFIX, defaulting to .internal which IANA reserved
for exactly this), a node's internal name, and the rule for what a node may be
called. Several things compose an internal name, and one of them writing the
suffix differently would be a name nothing answers to.
First attempt at this rewrote the file from memory and silently dropped the
configurable suffix. Restored from the original instead — deleting most of a
file is git surgery, not paraphrase.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
mesh-names, mesh-resolver and the names half of the overlay generators are gone.
They ran no software and could not be swapped for anything, which is the test of
whether something is a module at all — they existed because computed output
needed somewhere to live, and the control plane's only shape for output was a
module.
Now a module says where it wants what the mesh knows:
facts: { node-zones: /etc/mesh-resolver/nodes.conf }
and is given a file, under its own name, applied and removed like anything else
it declares. Two facts exist: node-names (a hosts file — exact names) and
node-zones (every machine as a wildcard, *.homer.internal is homer). Asking for
a fact the mesh does not compute is refused naming what would have worked,
because a daemon that starts and reads a file nobody wrote is a worse way to
find out.
The names ride with the network now: wireguard's manifest asks for node-names
into /etc/hosts, because being on the private network is what gives a machine a
name. networking no longer requires name-resolution — names are not a provision,
and the module that answered it ran nothing.
One behaviour inverted, deliberately: choosing another VPN used to drag
WireGuard in anyway, because only WireGuard provided the addressing the names
module required — the node-scope claim existed to at least make that loud. With
names as a fact there is nothing to drag in: tailscale assigned means tailscale,
alone. The claim still catches two VPNs assigned explicitly.
And a machine the mesh cannot place is left out of both files rather than named
at nothing: a name resolving to nothing hangs a connection, where an unknown
name fails at once and says so. In practice that is only ever a token issued and
not yet used — a machine that has announced itself has an address.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The mesh assigns the machine-side port and a module does not choose one (ADR
0038). For a container that is invisible: the mesh rewrites ports into
assigned:wanted, the software binds the number it always bound, and the machine
publishes another.
A process has no such layer. It runs on the machine, there is nothing to rewrite,
and it binds whatever its configuration says. So every process bound the number
written in its own config, two modules declaring the same one would collide, and
the mesh's whole reason for assigning ports was defeated by the resource kind
that most needs it — introduced, by me, three commits ago.
So a module asks. ${port:8080} is "the machine-side port you gave me for the 8080
I said I listen on", written into its own configuration exactly as an address it
was bound to is.
Asking about a port it never declared is refused, and the refusal says what it
did declare: the module is asking about something the mesh has no opinion on, and
answering would put a guess into a configuration file as a port number. With
nothing assigned yet it is told what it asked for, so a mesh that has made no
assignment still composes something coherent rather than writing a zero.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A module is one piece of software and may still carry a daemon in one language,
tools in another and a package in a third. The first cut compiled every bundle
into the toolchain's single output directory, so two of them would have
overwritten each other and then been packed together — one artifact containing
both, published twice.
So output is a property of the artifact, not of the toolchain, and the toolchain
says how it is told where to write rather than where it writes. Under a directory
named for the build rather than beside the source, so a pack never sweeps up the
module's own working files.
A second toolchain is declared so the multi-language path is exercised rather
than asserted — a list with one entry cannot fail the way a list with four will.
And the fake compiler in the tests now writes where it was TOLD to. One that
always wrote to a fixed place would have passed whether or not each artifact got
its own directory, which is the whole of what these tests are for.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The bundle recipe: the one that both builds and packs. An archive packs a
directory as it stands, so shipping compiled output meant compiling somewhere
first — which meant a Dockerfile repeating the same incantation in every module.
Two base arguments with no defaults, a working directory chosen so the SDK
resolves upward, the compiler invoked by absolute path because the usual symlink
is resolved away when the base image is assembled, a second stage, an environment
variable naming the entrypoints. Most of the catalogue is unconverted and that is
why; two conversions done in one session were each wrong twice with a working
example open in the next window.
A bundle says a language and a list of entrypoints. The mesh knows what the
language implies. Anything a module could override there it would be writing a
Dockerfile to override, so a toolchain is deliberately not configurable.
Declared rather than inferred, both of them: guessing the language from which
files are present makes a build depend on a directory listing, and guessing the
entrypoints makes it change meaning when somebody adds a helper.
A toolchain the mesh does not hold is refused before anything is compiled, naming
what to build first — the same treatment a missing base already gets, because it
is the same question and somebody can answer it. A language the mesh does not
build is refused saying what would have worked, since the author is usually one
word away.
The list of languages is closed and adding to it is a decision. Every language is
another implementation of the contracts every module shares, and those change
rarely and cascade when they do (ADR 0039) — a mesh whose SDKs disagree about the
envelope fails by ignoring messages rather than by failing to compile.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Its event queue is durable, so a running catalogue misses nothing. What it cannot
have is what was announced before it first ran — and on a fresh mesh that is never
arbitrary: the shared base, the store the catalogue runs on, and the catalogue
itself are each necessarily built BEFORE a catalogue exists to hear about them.
The graph's foundation is the part it never sees.
So it says it is catching up, and the control plane re-announces what it
recorded, oldest first, marked as a replay. Oldest first because a graph is built
in the order things happened: registering a module that stands on a base before
the base would point an edge at a version nothing has seen, and the shape of a
fresh mesh guarantees the base is both first and the one that was missed.
The replayer hands announcements back rather than publishing them, because the
wire belongs to the link package and a replay building its own events could drift
from what the builder emits — the one thing it must match exactly, since the
catalogue has a single handler for both.
Its own queue and its own consumer: two consumers on one queue split its
messages, and a catch-up request going to whichever half was not listening is a
gap that looks like a working mesh.
Toward novox/hq 04-ISSUES/050.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The builder announces a build with the resolved manifest, the path inside the
repository, and every artifact it stood on. The control plane received all of it
and kept none of it.
That was survivable while the catalogue heard the same announcement directly. It
stops being survivable the moment the catalogue was not there to hear it — which
on a fresh mesh is always, and always for the same modules: the shared base, the
store the catalogue runs on, and the catalogue itself are each necessarily built
BEFORE the catalogue exists to hear about them. The graph's foundation is the
part the graph never sees.
Replaying those builds needs what they said, not a summary. Without the manifest
there are no requires/provides edges; without `against` there are no build edges,
which are the ones that answer "a base moved, what must be rebuilt". A replay
carrying neither would restore the module list and leave the question the
catalogue exists for still wrong, while looking fixed.
Kept null rather than empty where a build predates this, so a replay can say it
is holding nothing instead of inventing an empty declaration for a module that
certainly had one. And `built_against`, not `built_on`: that column exists and
means the machine, which is a different fact about a different subject.
Toward novox/hq 04-ISSUES/050.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Rules are derived from what modules declare they listen on, and the substrate is
not a module. So the broker's port — the one every machine dials to enrol and to
receive every declaration it is ever sent — appeared in no ruleset the mesh has
ever generated.
Nothing caught it because a mesh of one never dials its own broker across the
network: the ruleset looks complete right up until a second machine tries to
join a firewalled anchor and is refused by the packet filter, during enrolment,
before the mesh can report anything about it. Assigning the firewall before
joining machines is both the natural order and the one that breaks.
It is a floor for the same reason ssh is. A machine nobody can reach cannot be
repaired; a machine the mesh cannot reach cannot be managed. Neither is a thing
any module asks for and neither may be derived away.
From anywhere rather than from the private network, deliberately: a node enrols
BEFORE it has an address on that network, so narrowing the rule to it would close
the door being knocked on.
The port is read from the broker this control plane was told about, so the
address handed out in a token and the port a machine must accept on stay one
fact. A mesh never told about a broker gets no such rule, rather than a broken
one — and cannot issue tokens either, which is where that surfaces.
Closes novox/hq 04-ISSUES/052.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Two sibling branches resolve a provision answered by the consumer's own machine.
The one for a node-scoped provider falls back to loopback when the machine is on
no private network, with a comment saying why and a test holding it. The one for
a mesh-scoped provider passed node.At straight through, and nothing noticed
because nothing had yet composed a host out of it.
The mesh's own artifact store is mesh-scoped and sits on the same machine as the
builder that pushes to it. Give the builder the address from its binding and it
gets MESH_REGISTRY=:5000 — a name with no host, written into its environment
without complaint. It surfaces much later as
":5000/mesh-tools/build" is not a valid repository/tag
which is a message about a tag for a fault in how a binding was resolved, on a
machine several steps from the decision.
A machine off the private network still reaches itself, which is what the
neighbouring branch already said. The test fails without the fix, showing the
empty address rather than only the symptom.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The check read "declares no resources" as "runs nowhere", and those are not
the same. The private network declares no resources either — the control plane
computes them when it composes a machine's declaration — and it is assigned to
every machine that has to reach another one. Refusing it stopped a four-machine
bed at its first assignment.
The signal is narrower: it builds an artifact and places nothing. Made a
function of its own, because a judgement with a wrong answer this expensive
should be testable without a database — nothing guarded it, which is how it
shipped.
The floor allowed ssh from the mesh's addresses, and from everywhere on a
machine that faces outward. On a machine the mesh knows no addresses for it
emitted neither — so the chain dropped by default and ssh was simply shut.
That is the first machine anybody adopts: reached over the network, with the
port needed to fix it closed by the act of adopting it. Found by reading the
rules off a live machine rather than trusting the generator.
Two faults, opposite directions, both in issue 047.
There was no forward chain, on the reasoning that dropping there stops every
container the runtime allowed. The first half is true; the conclusion was not. A
published port is redirected and then forwarded, so it never reaches the input
chain — the firewall was silent about the ports most worth protecting. The way
through is the one the system being replaced already used: deny by default, then
allow the runtime's own networks explicitly. A forwarded rule matches what the
client originally asked for, because the destination has been rewritten by the
time the chain sees it.
And ssh is now a floor nothing derives. Every other line comes from what is
assigned, which is the point — but a mesh part-way through adopting a machine
has been assigned almost nothing, so what it computed was a chain that shut the
port used to fix it. From the mesh always; from outside on a machine that faces
outward, because that is the way in when the private network is what broke.
Rehearsed on three machines: a docker-published port declared mesh-only is now
reachable from inside the mesh and refused from outside. Before, it was
reachable from both.
The image every module in the scripted toolchain is compiled on top of is
registered as a module so the mesh can build, version and depend on it. It is
not one: nothing about it belongs on a machine. Assigning it succeeded, the
machine was sent a declaration containing nothing of it, and everything
reported success — the operator had said run this here and the mesh had agreed
to something it cannot do.
A module whose resources are worked out per node is asked about separately, so
it stays assignable, which is the point of it.
A fingerprint written into a recipe names one particular copy of the base — the
copy on whichever machine the person typing it was using. On any other mesh that
copy has never existed, so the build stops on its first line with a message
about an image nobody can look up. Three modules in the catalogue were in
exactly that state, and the line each of them replaced was equally dead.
A module now names the module and artifact instead, and the mesh answers with
what it holds. The builder is still a thing that clones, builds and answers: the
answer travels with the question, because only the mesh knows what it has.
A base the mesh has not built is refused before anything is built, naming which
module has to exist first.
A container naming an artifact is a module saying the mesh builds this. Until a
build publishes one there is nothing to run — and what reached the machine was
an unresolved field, which its language has no room for, so it refused the whole
declaration and reported that a container does not use "artifact". That reads
as a broken manifest. It is not broken, it is unbuilt, and only the mesh can
tell those apart.
Found by the four-machine bed, which assigns modules the mesh has not built.
A module is a repository and a path within it, and the manifest sits at that
path. This one sat in the catalogue instead, so the thing that says what the
control plane is and the thing it is made of lived in different repositories
and could drift apart with nothing to notice.
It also names an artifact it builds rather than a placeholder digest somebody
fills in, which is what lets the mesh build its own control plane.
This is how a mesh is raised: the installer carries this program and runs it
once, before anything exists, to produce the control plane from the same
repository and path every later rebuild will use. What raises the mesh is then
the same thing that maintains it, rather than a second mechanism exercised once
per new mesh — which is how often enough to rot.
With nowhere to publish, an image stays in the machine's own runtime and is
named by the digest of its own configuration: the same identity the installer
has always used for the image it carried.
The builder says what it built and the catalogue decides whether that was an
upgrade. Only the control plane knows which machines run the thing, so it is
the one that acts — and what it does is a choice somebody recorded, not a
behaviour compiled in: record that they are behind, or send it, one machine at
a time or together.
Recording is the absence of an action rather than a second path: a machine not
running what the mesh would send it is already something the mesh reports.
Defaulted to recording. A mesh that rolls out everything it builds the moment
it builds it is reasonable to want and a bad thing to arrive by default — the
first module to inherit it would be the control plane, upgrading itself out
from under the push applying it.
An event whose origin reads a container id names something no other module
can look up. The mesh already knows the answer, and a module's environment
file is a file resource, so ${machine:name} reaches it with no composer change.
Answering and announcing are different acts. The reply goes to whoever asked and
is correlated to their request; the announcement says to the whole mesh that a
module now exists at a commit, which is what the catalogue places in the module
graph (novox/hq ADR 0072). A build nobody asked for still has to be announced, or
the graph knows less than the registry does.
What it was built on top of is read out of the build's own inputs rather than
declared, because a declared list drifts from what the code actually uses
(ADR 0009). These are artifact references, which is what a build input names;
resolving them to module-versions is the catalogue's work, since it is what knows
which module-version published which artifact.
Events ride the topic exchange, not the direct one nodes speak over, so the
builder's account is granted both: it must be able to answer and to announce.
The envelope is the sdk's, reproduced exactly — a second shape would be a second
thing for consumers to handle, and they are written against the first.
Announcing is not allowed to fail a build. The work was done and was answered; a
build reported as failed because saying so failed is a lie about it.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx