novox/hq 04-ISSUES/142, and the second of the two things ADR 0141's own
insight named: "a version cannot reach the path". A component is unpacked
into a directory named for its version so it can read its own version from
its path — and an archive named a fixed path in the manifest with nothing
interpolating the build into it, so nothing could ask for
.../versions/<version>/ and every machine took a hand-placed fallback.
A resource using an archive or a bundle may now say ${version} in any of
its values. No artifact name in the reference: the resource already says
which artifact it is for, and a second name is a second thing to keep in
step.
The version is the artifact's digest, short, and not the commit. Two
builds of one commit are meant to be the same bytes — every toolchain here
is -trimpath for that reason — so a content-addressed version means an
unchanged build resolves to the path it already had. A commit-named path
would move for an identical binary and recreate everything that reads it.
An image is refused one, with a reason: an image is not unpacked, so it
has no versioned place. Left alone it would reach a machine as literal
text and be created as a directory called ${version}.
novox/hq 04-ISSUES/142. A bundle is refused if it names what it is built
from, because a bundle is the module's own directory compiled whole and
naming a source would be describing its own build. That reason holds for
an interpreted language and cannot hold for a compiled one.
A repository written in Go carries several commands — the host and its
bootstrap live in one — and "the module's own directory" is then not a
package at all. So a compiled bundle may say which package, and says the
module root by saying nothing. The refusal stands for every interpreted
bundle, which is what it was written for.
Found by writing the host's manifest, which is the first bundle in a
compiled language this mesh has had.
novox/hq 04-ISSUES/142, and ADR 0141's own progressive insight naming
this as the first of two things missing: "nothing can compile it". The
toolchain list was a closed set of typescript and python, whose warning —
every language is another implementation of the contracts modules share —
does not attach to Go. Go is how the host, the control plane and the
builder are written, and none of them is a module in that sense: the host
is what APPLIES modules.
The toolchain names mesh-tools-go as its base rather than pinning an
upstream release here (ADR 0142, 0044): named and not pinned means the
mesh answers with the copy it holds, and moving compiler is a build
instead of an edit to this file.
Two things beyond the list also assumed one language, and both would have
failed after the entry was added:
sourcesFor turned every entrypoint into a `.ts` file. The extension is
the toolchain's now — one language's file extension written into the
code that serves every language is a wall the next one hits.
The output directory was the compiler's to create. tsc --outDir makes
one; go build -o writes into a directory and does not make it, failing
with a message about a path rather than about a build. Made here for
every toolchain, because which compilers are forgiving is not something
a reader should have to know.
And a toolchain now says what it is pointed at: a file list from the
module's entrypoints, or the one package the artifact is built `from`.
Pointing `go build` at a file list builds a program out of exactly those
files and ignores the rest of the package — a missing symbol rather than a
legible refusal.
Static and -trimpath: what a machine holds is a file, not a container, so
a binary needing a libc it did not bring is a delivery that works until a
machine differs; and a version comes from where a component sits rather
than from its linker, so two builds of one commit are the same bytes.
novox/hq 04-ISSUES/087. The version I shipped this morning said "N
machine(s) run an older host than another machine does" and worked it out
by comparing versions as strings. A host reports its version as a commit.
Commits have no order.
On the live mesh it named the three machines running the NEWER host as the
ones behind: `ced54d4` sorts above `04a27ca` and means nothing. An
arbitrary lexicographic result, presented as a fact, about the one thing
this was built to make trustworthy.
It now reports the split — which machines run which version — and claims
no ordering:
4 machine(s) do not all run the same host:
04a27ca g14, novox, shanks
ced54d4 ace
a host refuses a declaration carrying a field it does not know, whole
— so the mesh may send only what every one of these understands. Which
of them is newer is not readable from a commit; that needs a version
the host reports as ordered
More useful as well as more honest: the reader sees who is on which side
of the split, which is what decides whether a field can be sent.
A report that confidently says the opposite of the truth is worse than one
that says less — which is the subject of 04-ISSUES/145, arriving by my own
door within an hour of my closing it.
novox/hq 04-ISSUES/145. "N machine(s), all doing what they were told, all
heard from, running what the mesh would send them, and every module
current with its source" was true for eleven hours of a mesh in which no
module could reach another. An operator read it, and every routine check
they made afterwards — ports from outside, routed services, egress —
passed, because the broken path was module-to-module over the machine's
own name and nothing exercises that.
Every question the sentence answers is about the mesh and a machine
agreeing: applied what it was sent, matches what would be sent, built from
what the source has. None dials a provision, and the mesh composes every
one of those grants itself. So the sentence now says so, in the reader's
way, immediately below it.
This is not the check ADR 0146 describes and does not pretend to be. It
closes the distance between "the machines are as the mesh described them"
and "it works" by naming it, which is where the eleven hours went.
Also: printStatus is separated from the asking, so its exact words can be
read by a test with no store, bus or machine. Those words have been acted
on and been misleading twice — here, and a held module reading as a
machine doing what it was told (04-ISSUES/125) — which makes them the
thing worth holding still.
novox/hq 04-ISSUES/087. A host refuses a declaration carrying a field it
does not know, and refuses it WHOLE — deliberately, because that keeps a
half-understood declaration off a machine. It makes every new declaration
field a flag day: hosts first, then the controller. The mesh had no record
of which host any machine ran, so that order was kept by somebody
remembering it, and a machine that refused for this reason reported a
failure with nothing saying why.
The machine has reported its host version since ADR 0141. The
controller's own copy of the report did not have the field, so it was
unmarshalled into nothing and thrown away on arrival. It has it now,
records it, and shows it in `node show` — "not reported" rather than
blank, because a machine that has not said is not a machine running
nothing.
Status says which machines run an older host than another machine does,
and which is newest. Deliberately disagreement rather than staleness:
nothing delivers a host version yet (ADR 0141, accepted and not built), so
the mesh holds no canonical current version and cannot honestly say a
machine is behind THE host. What it can say is that the oldest host in the
mesh is what the mesh may send.
A machine that has reported nothing is left out rather than called
behind. Versions compare as strings, which suits the timestamps and
commits this mesh uses and is wrong for a scheme where "10" sorts before
"9" — said in the code, at the place that would have to learn.
novox/hq 04-ISSUES/125. A module assigned to a machine and never taken
runs none of what it declares. Status had no vocabulary for it: the
machine was heard from, current, and doing what it was told, so the mesh
printed "all doing what they were told" — which was true, and was acted
on, and every public name on the machine went dark.
Status now names each module a machine is holding rather than running,
per machine and with a count, read from what the MACHINE reported rather
than from the mesh's take-time listing — the machine is the only thing
that knows what it found. The JSON form carries the same rows, absent
rather than empty when nothing is held.
And a hold suppresses the all-well sentence, where being adopted does
not: adopted is a mode somebody chose, a module assigned and never taken
is a half-finished action with nothing left to finish it. The condition
is now a named function so the rule lives in one place and a test binds
to the real thing rather than a copy of it.
untakenModules raises a read it cannot make rather than answering "holding
nothing" from a failed query, which is the shape this whole issue is.
novox/hq 04-ISSUES/156. Issue 146 put the stream into a push consumer's
delivery subject. The server will not move that subject while a
subscriber is bound, and answers `consumer name already in use` — a
message about the name, for a conflict about the subject. A node is bound
to its declaration consumer the whole time it is up: that IS a node
listening. So every node consumer in a running mesh became one the
assertion could not bring to match, and the control plane crash-looped on
the assertion it makes before it serves. A fresh mesh showed nothing,
because nothing was bound.
Kept rather than deleted and re-made. Re-making moves the subject, and a
holder may not be allowed to subscribe to the new one yet: the wider
grant travels in the bus's user list, which this same control plane
composes and a machine applies minutes later. On the live mesh the nodes
are granted `_DELIVER.<node>` and not `_DELIVER.<node>.>`, so re-making
would have silenced every machine — worse than the collision it fixes,
and harder to undo.
Kept rather than fatal, which is what 146's change intended and did not
do. The bare subject still delivers, and collides only where one holder
has two consumers of one name. That is the controller's own pair, and the
controller is not bound to them while it asserts, so those do move.
Also: an existing consumer's deliver policy is carried across rather than
reasserted, because the server refuses to change it and where a consumer
starts is its history.
Two tests against a real server: a consumer with a subscriber bound keeps
its subject, is reported, and still delivers; one with nothing bound
moves, so 146's fix still applies where it matters.
Three gatherers walk every node and pass over one whose plan will not
compose, so that one broken set does not cost the rest. They read a
plain error to mean that, and so read a store that was briefly
unreachable as a machine running nothing.
On the roster of routed names that is not a degraded answer but a false
one: it states to every machine at once that another machine's names do
not exist. Because the roster is part of every container's identity, a
control node replaced every container it ran — its own store, the
registry, the edge, mail, the bus — on a six-minute cycle for hours. The
loop closed through the store this is read from: each pass restarted it,
the read failed, one name left the roster, and the roster changing is
every container changing.
planFor now marks the two failures that really are the node's own — its
set not composing, and a setting that reaches nothing — and the three
gatherers pass over those and only those. Every other failure is raised,
naming the machine and the read, because a mesh-wide refusal with
nothing named in it is the other way to lose an evening.
novox/hq 04-ISSUES/152, and 151 for why a changed roster is a changed
container.
novox/hq 04-ISSUES/146. A push consumer delivers onto an ordinary subject and
everything subscribed to it gets a copy. The controller holds a consumer called
'controller' on CONTROL and another called 'controller' on EVENTS, and both were
given _DELIVER.controller — so the one process, holding both subscriptions,
acted on every message twice.
Measured: one enrolment published, one message in the stream, one delivery, no
redelivery, and the controller enrolled the machine twice — the second minting a
credential that replaced the one the machine had just been handed, which is why
it then reconnected for ever as a user whose password the mesh had rotated. Every
report and every followed event doubled the same way, silently.
The stream goes in the subject because the pair is what identifies a consumer.
A subscriber's permission gains the same shape, keeping the bare name so an
existing consumer keeps working until the next assertion moves it.
It was broken and stayed broken: the Dockerfile's fallback base is a Go older
than go.mod asks for, so every hand build died at 'go mod download' with
'go.mod requires go >= 1.26.0'. The pipeline never saw it because the pipeline
passes the declared base in, so the cost fell entirely on whoever built the
image themselves and had to find the digest by hand (novox/hq 04-ISSUES/146).
Read from module.json rather than written here as well, so the two cannot
disagree, and refused outright if the manifest declares none.
novox/hq 04-ISSUES/146. The composed user list names an enrolment user for
every machine with a live token and nothing minted a credential for it, so the
composer left it out as a user with no password — and every enrolment since the
mesh moved to this bus was refused before the mesh heard of it. The comment
above the issuing code already said the account is created before the token is
handed over; now it is. Recorded rather than minted, because the token's secret
is the password.
And 'broker accounts', which composes the same list the declaration carries and
writes it to standard output. For genesis, where no declaration can reach the
machine running the bus because that machine is not yet a node. It says what it
composed; whoever is raising the machine places it. A control plane that wrote
the file itself would have to learn where the bus keeps its configuration and
how to make it reload, which is the module's knowledge.
novox/hq 04-ISSUES/146. The foundation made it by running openssl inside the
broker's image, which worked while the broker was one that carried it and
stopped the day the bus changed: the new one has a shell and no openssl, so
the step exited 127 and no mesh could be raised. No other image the bundle
names has it either, so there was nothing to substitute.
broker certificate --into <dir> writes the pair, --check is the step's verify.
Self-signed on purpose — a host pins this server's exact certificate (ADR
0004) and at genesis there is no authority to ask — and made once, because a
second certificate is one every host that pinned the first no longer believes.
The key is written before the certificate, so an interruption never leaves
something that looks finished.
novox/hq ADR 0147. ca-trust carries a script and a unit; the one thing
neither can state is where the authority is, because that is a fact about
the mesh. This checks the rendering — the script fetches from the bound
address and is executable, and the unit runs it both ways, install and
remove. The verification itself is the lab's.
Local is not a boundary this mesh draws. A service here is callable by everything
else here, whatever form either takes — a package with a unit, a binary, a
container. Whether a caller sits in a container was never meant to change the
answer, and the only reason it did was that this chain asked about addresses: a
caller on the machine carries the machine's address, a caller in one of its
containers carries a bridge address, and a rule naming the former silently refused
the latter.
One rule for every service here, replacing the line-per-port added an hour ago,
which only ever covered the ports somebody remembered to think about. The three
reaches are now three lines: on this machine, over the private network, from
anywhere.
The tests assert per chain body, because the forward chain carries the same line in
the same words and an assertion on the whole file passed with the input chain's copy
deleted — which is what ADR 0137's own tests say to do and this file was not doing.
A port declared from the mesh admitted the machines' own addresses on the private
network. A container reaching a port on the machine it runs on comes from a bridge,
matching none of them — and where the container runtime routes directly, that packet
is delivered to this machine rather than forwarded, so the forward chain's allowance
never saw it either.
ADR 0100 requires this to work: the store is reachable 'from a container on the node
itself'. It was, through a rule the predecessor left, which allowed the private
ranges wholesale. Converging the machine replaced that with the four overlay
addresses and closed it.
Measured, and it was an outage: every module reaching another by its machine's own
name timed out for eleven hours while the mesh reported the machine healthy. A web
application logged 'connection to server at novox.internal (10.10.0.1), port 6852
failed: timeout expired' throughout.
Asked for by the link it arrives on, for the reason the forward chain no longer
names an address: a range describes one machine and goes stale in silence. A port
open to everything needs no such line.
Everything downstream reads the port: the provider is told where to reach the
consumer, and the redirection that turns a declared port into the number the
machine published is keyed on it. A route naming only its endpoint left the proxy
with no port at all, and a proxy with no port has nothing to dial.
Caught after the catalogue had already been changed to name endpoints and before
the mesh picked those manifests up, which is the only reason nothing broke: every
module's manifest is behind its source right now, so the plan still renders from
the old shape.
The declared port, not the machine one — the redirection happens later and is keyed
on the declared number, so filling in the machine port here would be redirected
twice or not at all. A route that repeats a port keeps it.
A test that reads every manifest in a catalogue checkout and parses it with the
control plane's own parser, rather than asserting against a fixture: whether the
manifests as written are accepted is the question, and a copy of one proves nothing
about the other seventy-one.
Skipped unless MESH_CATALOGUE names a checkout, so it costs nothing in ordinary
runs and is there when the catalogue changes shape. It also refuses to pass if no
endpoint is named, because a run that validated nothing would otherwise read as
success.
novox/hq ADR 0138, completing it. One block per endpoint instead of three keys
joined by a port number:
{"endpoints": {"web": {"port": 20009, "label": "cinema", "reach": "both"},
"stream": {"reach": "internal"}}}
Which machine port it lands on, the subdomain a proxy serves it under, and how far
it reaches are the three things an operator says when a module is assigned, and they
were said in ports, in the route's label and in reach — each keyed by the port. A
module with two endpoints of different shapes could only be configured by a reader
who knew which number was which.
Every field is optional; a block that says only a reach leaves the port to the mesh
and the label to the module, which is the ordinary case. A name the module does not
declare is refused, and the refusal lists what it does declare. A port or a reach
said both here and through the older key is refused rather than merged — two places
saying one thing is what this key exists to end, and merging would follow whichever
was read last.
Eight tests. The reach assertion deliberately narrows what the manifest says, because
a reach that agrees with the manifest proves nothing about whether the block was read
— which I found by writing the weaker version first and watching a revert not fail.
novox/hq ADR 0138's remaining half, and the words ship one release before any
manifest uses them.
A port number is not a name. Three facts have to be said about an endpoint when a
module is assigned — which machine port it lands on, the subdomain a proxy serves
it under, and how far it reaches — and they were said in three places keyed by the
port. A module with two endpoints of different shapes cannot be configured that way
without a reader joining numbers by hand: a web surface behind the proxy, whose
port only the proxy need reach, and a protocol port clients dial directly because
the client expects that number.
So a listen carries a name, lowercase and unique within the module, and a route
names the endpoint it serves instead of repeating its port. Two endpoints with one
name are refused, because an assignment configuring one would silently configure
whichever the mesh read last. A route naming an endpoint the module does not declare
is refused where it is written rather than resolving to no port and serving nothing.
An unnamed endpoint stays valid and a route repeating a port still resolves, which
is every module in the catalogue today.
A routed endpoint's port is how the proxy reaches it and nothing else (ADR 0045):
a public service listens from the mesh, only the proxy reaches it, and it is
exposed by name. So reach on a routed endpoint asks for names, and the port keeps
what the manifest said; on an unrouted one — git over ssh, a mail port, the bus —
it governs the port, because there is no name and the port is the only way in.
Found by trying to express a real module rather than by review: routed name public
because browsers post to it, machine-side port private because it serves a
dashboard in cleartext. Under one value for both there was no way to say it, and
'public' would have reopened a port narrowed an hour earlier.
novox/hq ADR 0138, corrected in place the same day.
novox/hq ADR 0138. Reachability was settled three times over: the filter read a
listen's source with expose able to override it; the proxy composed a public name
and an internal name for every route it was given, because it could; and the
certificate authority followed from which names existed. Each was defensible and
the combination was unstated, so "this endpoint must not be public" could not be
written and was enforced by nothing — while a public certificate for that name was
obtained anyway. Measured on the control node: an identity provider holding a
90-day public certificate and a 24-hour internal one, neither asked for.
`reach` is one value per endpoint, per node — machine, internal, public or both —
and the filter's source and the composed names both follow it. The authority needs
no work: the proxy already asks the public authority for a route's own name and its
internal authority for the internal one, so controlling the names controls the
authority.
Joined by the port, which a route already names: 35 of the catalogue's 36 route
entries name a port the same module declares a listen on, and the one that does not
is a path-level refusal — a rule about a name rather than an endpoint, left alone.
Nothing said composes both names and follows the manifest's `from`, so every mesh
already running is unchanged until an assignment speaks. A port that says both
reach and expose is refused: they say the same thing in different words, and the
filter would follow one while the names followed the other.
The forward chain blocked everything passing through the machine and then allowed
the machine's own containers back by naming their address ranges: 172.16.0.0/12 and
192.168.128.0/17 fixed here, the rest recorded per machine by 0043. Every way of
keeping that list correct fails — a constant describes one machine, a recorded range
goes stale in silence and cannot tell a network the mesh made from one a predecessor
left behind, and generating it from the modules would put half the rule set on the
machine.
The mesh has no position on a container reaching outward: that is not a port opened
to anybody. So both chains are written around the links traffic arrives on. What did
not arrive from outside is accepted in one line; what did meets the declared rules.
The tunnel is named beside the outward links rather than treated as inside, or a port
nothing declares would be reachable from every machine in the mesh.
A machine that has not reported an outward link is sent no filter and keeps the one
it has, refused where a person reads it rather than as a rule set that will not load.
Removes the two constants, `node networks`, and the column behind it. novox/hq ADR
0140, superseding 0137 and 0139.
First of the steps in novox/hq ADR 0142, and it ships alone: a new manifest word
reaches the builder and the controller one release before any manifest uses it.
A toolchain deliberately accepts nothing from the module — anything a module
could override there it would be writing a Dockerfile to override — and yet a
compiled binary is per operating system, pinned at link time so a host refuses to
touch a machine it was not built for (ADR 0005). The way out is that the target
belongs to the artifact: one artifact per system, one build each, recipe still the
mesh's.
A bundle in a language that compiles to a binary must name a system, or it would
be built for whatever the build machine happened to be — which reads as portable
and is not. A bundle in a language that runs anywhere may not name one, because a
system that decides nothing reads as though it did. The list is the host's own
names, not a compiler's: the difference between two of them is a C library rather
than a kernel.
Nothing declares a system yet, and no toolchain compiles to a binary yet, so this
changes no build.
The derived filter denies forwarding by default and then allows the container
runtime's two default pools, named in this code with a comment saying a machine
configured otherwise needs to say so -- and no way to say it. So the filter was
right on a machine using the defaults and silently wrong on any other.
Measured today: flipping a workstation to the derived filter cut egress for five
of its container networks and for every network its test beds create, because
those come from ranges the defaults do not cover. Nothing reported a fault; the
guests just could not reach anything, while the machine reported it had applied
what it was told.
A node-level fact beside the public domain, because the machine routes them and
the module that loads the filter may be replaced. Added to the defaults, never
replacing them. Their guests also keep address and name service, without which a
network does not work at all, and the converge preview now says what a machine
routes instead of leaving it to a sentence about what it cannot preview.
A module's declaration and its consumer are derived from the same records, and only one of them
followed a push: the consumers were raised when the control plane started serving, so a module that
gained a `consumes` was sent a declaration it could act on and a consumer that never delivered the
event — with nothing anywhere saying the two disagreed. Found on review: the catalogue's own replay
subscription was recorded, granted and never delivered.
Everything the raise does is idempotent, so a push may do it.
The grant permitted the control plane to state what it applied and the seat said nothing about it, so
the check that every derived subscription has an owner found the catalogue subscribing to a subject
nothing publishes — which is exactly the fault that check exists for, pointed at me.
A seat carries the protocol of its role (novox/hq ADR 0129), so the facts are the mesh-controller
seat's `emits`. That is also what lets another module declare it consumes them. No accepts, so no work
queue is raised for the seat — only what its holder may say. The agreement test now holds all three
places to one another: the seat, the grant, and the words the mesh states them with.
A second Accept value would have overwritten the first, because the header was Set per value rather
than Added. One value is all any caller passes today, so nothing was wrong — but a helper that
quietly keeps only the last of what it was given is a trap for whoever passes two.
And a replayed announcement that could not be written was published as an empty body: a fact on the
mesh that says nothing, which the reader can only log and drop. It is now said and skipped, because a
body that cannot be marshalled is this program's fault rather than the bus's.
The pipeline was observable from a merge to an artifact and went dark where it touched a machine: a
node's report is control traffic only the control plane reads, so nothing said which version a
machine runs, or that it refused to (novox/hq ADR 0134). The control plane now states both under the
seat it holds — a role's events belong to the role and keep their address when the holder is
replaced — and only when the report is news, because a machine reconciles every minute and a fact per
report would be a fact per minute per machine.
Whether a report is news is the store's answer: it holds the previous one, so the listener returns it
and the server states the fact. That also gives the catch-up replay a subject the controller may
publish: it was published as a module's event from a module called "control-plane", which does not
exist, so the controller's own account refused it and every catalogue that asked what it missed was
answered with nothing.
A resource's id is `<module>.<its own id>` and a module's name may itself contain a dot — novox.be is
one — so the owner of a resource is everything before the *last* dot. The preparation step's id used a
dot, which made its owner unreadable by that rule; it uses a hyphen, and the id says what it belongs
to whichever way a reader splits it.
The check that skips copying a base the mesh already holds asked with no Accept header, and a
registry answers a manifest only in a media type the caller named: the same digest answered 200 with
the manifest types and 404 without them. So the builder concluded it held nothing, copied every
vendor base again, and exhausted the public hub's pull limit a second time today.
The test could not have caught it, because the fake registry answered a manifest HEAD regardless of
Accept — more permissive than the thing it stands in for. It is now as strict as a real registry, and
fails without the fix.
Now that every parser on the mesh knows the word, the control plane's manifest says it. Its schema
stops being a special case: the mesh derives the step from its own resource and gates its server on
it, which is the failure of novox/hq 04-ISSUES/133 closed by the mechanism rather than by a
hand-written step in one manifest.
A manifest word has to reach every parser before a manifest carries it. The builder refused
mesh-controller's manifest with `unknown field "prepares"` until it was rebuilt; then the running
control plane could not read the manifest in the build result either, so the build was recorded with
no module and the version never moved. Strict parsing is deliberate (novox/hq 04-ISSUES/003), so the
word ships first and a manifest uses it next: this takes `prepares` back out of the control plane's
own manifest, leaving the code that understands it, and the manifest says it again once this is
running everywhere.
The mesh derives the preparation from the module's own resource instead of each module hand-writing
a step beside it (novox/hq ADR 0135). A manifest says one word — `prepares` — and the mesh runs that
module's own program in its preparation mode, in the module's own context: the same image, the same
environment, the same mounts, because it is the same code. A published port and a fixed address are
taken away rather than copied, since the version being replaced still holds them.
One word for every kind of module: a Go binary receives `prepare` as its argument, a bundle receives
it through the runtime whose entry takes the same word. The control plane answers it like anything
else — its own schema stops being a special case, and its hand-written step is gone.
The mesh replaced its own control plane with a build carrying a migration, applied none of it, and
then refused every build it recorded for three quarters of an hour while reporting itself healthy
(novox/hq 04-ISSUES/133). The module now declares the step ADR 0052 prescribes: a run-once
`migrate` before the server, re-run whenever the image moves because the image is part of a step's
digest, and gating — a migration that fails stops the new server from starting rather than letting
it serve against a schema it does not have.
Three faults the mesh's own logs showed this morning. A handler that outlives the acknowledgement
window was handed its message again while it was still working: acting on a merge builds modules,
minutes against a thirty-second window, so one merge ran the whole catalogue five times over. The
transport now says the work is in progress while it runs, which is where the window belongs.
Everything the mesh hands out — a token, a membership, a person's credential — took its address
from the enrolment setting, which on a mesh that has moved still names the broker it moved from:
the first person issued after the move was handed the retired broker's port. There is one bus, and
its address is the one the control plane is connected to.
And `operator issue` documented an argument order its parser refused.
Three faults in one path. A merge rebuilt every module built from the repository, so one change in
a repository holding twenty-six of them meant twenty-six builds. A merge into a repository a module
only *packages* source from rebuilt nothing — two modules are built from the control plane's own
repository and neither had ever been rebuilt when it moved — because the manifest the mesh keeps
carries no build section, so a build now says which repositories it read and the mesh keeps that
beside what it stood on. And a module handed over by hand could record a repository with no
directory inside it, which is a module nothing can ever rebuild (novox/hq 04-ISSUES/131, /132).
A change inside no module's own directory is a change to what they share, and everything built from
that repository is rebuilt: rebuilding too much is the safe direction, because the fault this whole
path exists for is a mesh that believes it is current and is not.
The forge announces what it finds merged, and an old merge surfacing late moved the recorded head
backwards and rebuilt everything built from that repository, once per old merge. A merge made
before the source was last seen is history; one that says nothing about when is taken as news.
The catalogue now carries when each source was last seen. The bus-records test follows #116:
a module's tools are every one under its own name.
A base is named by digest, and a digest the mesh's registry holds under the module's repository
is the same bytes whatever upstream would say. Asked on every build, the public hub's anonymous
pull limit was reached on the first merge that rebuilt a whole catalogue, and every module whose
base lives there failed on a copy it did not need.
Every module moved onto the bus by the rollout was issued on the old one, so none had a consumer
waiting; and the grant named a push delivery a runtime's client never binds, while the pull it
does make — asking about its consumer, asking it for messages — was refused. The consumers a
module's declarations imply are now raised whenever the bus is, and the grant is the pull.
The control plane is the way in for tool calls (novox/hq ADR 0095): a person or an agent asks
through it, so it alone may publish to every module's tool subject. The first ask on the new bus
was refused the publish.
Every module that served a tool was refused the subscription on the new bus: the grant listed
tools from a manifest field no module fills, because the tools a module serves are what its code
answers and a second copy of that list would be a second source of truth. The grant is now the
module's own tool namespace; nothing else may subscribe it, a caller is still granted per tool by
name, and a module may answer what it was asked.
The probe connected bare, and a bus that requires TLS and a user refused it at the handshake —
so the check reported the standing server as absent. It now dials with the controller's own
credential and pin, which is the one fact the check is there to report.
The test pinned what a build stood on to the address the builder pulled from; the edge names
another module's artifact and is kept the way that artifact is (novox/hq 04-ISSUES/102).
Bases reach a recipe as build arguments, so the digest was never in the file the builder read
edges from: no build on the mesh recorded what it stood on, and 'build --on', the bases-first
order and the merge follow-up all walked a graph with no edges (novox/hq 04-ISSUES/131). The
builder now reports every base it resolved; the controller records them by artifact path and
reads the newest build's edges from the store, since a recorded manifest carries no build.on.
The mesh runs on the seat's bus alone (novox/hq ADR 0131, design 28 task 5.5). The old
transport's consume loop, build request, tool ask, management API and account scoping are
deleted, and the bus switch with them; the controller connects to the broker seat and to
nothing else. The store-window tests keep their assertions on a bus-less fake, and the tests
that only made sense for the old transport's in-memory holding go with it.
The decoder named the forge's merge subject as the fourth thing followed and the
list was three long: every message that fell through to that switch panicked the
control plane (2026-09-28). The entry was written and lost between two attempts
at the same edit. A test now walks the list; the composed grants and the genesis
template carry the subject.
The controller follows the forge's merges (novox/hq 04-ISSUES/131). For each
module recorded as built from that repository and branch it records the move to
the merge commit and builds it — bases first, because a module built before the
module it stands on is built against the old one and reports success, and a base
that fails stops what stands on it. Nothing is pushed here: what a finished build
does to the machines running the module stays the upgrade's decision.
Two more things the same ordering gives: `build --behind` builds bases first, and
`build --on <module>` rebuilds everything that stands on a module — the rebuild a
changed base needs, which "behind" does not see because their sources did not
move.
The seat table has name, scope, delivers and decision, and the protocol ADR 0129
gave a seat lives only in the compiled defaults; loading the rows dropped it, so
no role's work queue was ever raised and the first build submitted over the new
bus met "no response from stream". Until the table gains the columns, a row with
no protocol keeps the compiled one of its name. And the holder of a seat is
granted what taking work from its queue needs — asking about the worker consumer
it binds, and acknowledging on it — which the first machine to try was refused.
The control plane's own seat placeholders no longer include the old bus's port,
which the switch removed with the variable.
Two halves of one gap the first build over the new bus met. The machine decided
its bus from a variable its container never received, so the credential the mesh
sealed to it went unread; a credential for the new bus names the bus by scheme and
carries user, password and fingerprint beside the address, and that is enough to
dial it, pinned. And the roles' work queues were raised with no holders, so the
consumer a machine binds to take work was never created: the holders are read
from the catalogue and the handover record, as the resolver reads them.
A push consumer delivers on _DELIVER.<its name>, and a client bound to it
subscribes exactly that. No principal was granted it, and the server refused
every one the first time it bound a consumer: the control plane, each machine,
and a module would have been next. Each kind is granted its own consumers'
delivery subjects and no other's. The line announcing the raised bus printed the
URL with the credential in it; the address alone now.
Two refusals the first live connections met. A user in the MESH account was told
"JetStream not enabled for account" the first time it bound a consumer: with
accounts defined, JetStream is enabled per account, not only globally — the
account's setting, which the mesh owns, not the server's block, which it does not.
And the control plane's client used a random inbox prefix where it is granted
exactly _INBOX.<its user>.>, so the server's first answer could not reach it. The
prefix now follows from the user in the URL, for every principal that dials so.
The bus presents the mesh's own certificate, which names nothing a public verifier
accepts; the client verified by name and failed against a bus that was answering
("certificate is not valid for any names", 2026-09-28). It now pins the leaf's
fingerprint from MESH_BROKER_CERTIFICATE, as every host does. And a connection
error named the whole URL, password included — the address alone now.
The seams were there and nothing chose a side: serve, push, ask and build all
opened the old bus's connection and declared over its channel, whatever
MESH_BUS_NATS said. So the switch moved every host and left the control plane
unable to follow — "this control plane has no MESH_BROKER_AMQP" with the new bus
named and standing (2026-09-28). That was task 4.3 of design 28, still open.
One place now decides: connectLink reads the switch, refuses both buses named at
once, raises the new bus's streams and this controller's consumers when it is
handed the inventory, and opens the link over whichever bus it is on. Every
caller that sent a declaration or asked a tool through the old channel goes
through the server's bus instead, which the new transport has and the channel is
not. OverNats is that outbound: a declaration is a JetStream publish into the
node's own subject, an event is announced on the subject its name derives to, a
tool is request and reply on the module's tool subject.
Binding to a consumer asks the server about it and hears the answer on the
client's inbox; hearing a declaration acknowledges it. A machine's user was granted
none of that and was refused the first time one dialled a permissioned server:
"this node cannot read its declarations". Its inbox is its own prefix, which the
host now sets.
The control plane is a module too, and its broker secret is the old bus's
credential it is still using while the mint runs. Writing the new bus's blob there
cut the mesh off from its own old bus mid-move. Its new-bus credential is the
controller principal's bus secret; the module principal is skipped.
MESH_BUS_NATS_FILE named /run/secrets/bus and nothing put a file there: the
manifest binds each secret explicitly, and the switch added the secret and the
variable but not the bind. Found live — the control plane came up on the new bus
and could not read its own credential.
Without them, a machine running the next holder of a seat beside the current one
resolves as two holders, is refused, and drops out of the map — and with it the
address every other machine composes for what it offers. Found live: the control
node vanished from the private network the moment the new bus was assigned beside
the old one, and nothing on any machine could be composed.
BareAddress adds a scheme where none was, so the host it yielded carried one and
every URL built from it carried two. Caught before a push: the host is now taken
with no scheme and no port, and every URL adds its own. `rollout mint --again`
mints every credential afresh for exactly this case — a mint that was wrong before
anything received it.
The map of who is where on the private network lists the machines placed around
the hub, not the hub — and the control node is the hub, and runs the new bus.
Found on the first live run: refused for having no address. The address every
machine already dials the current bus at is the same machine, so that host with
the new port is what they are told.
One environment variable moves it (design 25): MESH_BUS_NATS, read from the `bus`
secret `rollout mint` sealed to its machine, and the two that named the old bus go,
because being told about both is refused at start. Registered and built ahead of
the push that flips it, so the flip and the bus it flips to arrive in one
declaration — the same push that starts the new bus and hands this machine its
membership. Deliberately not merged until every credential is minted.
`rollout mint` gives every principal the new bus will have a credential it does
not yet have and puts each where its owner reads it: a machine's as a membership
— bus address, fingerprint, password, transport — sealed into its declaration
(migration 0041, the `bus-membership` resource the host reads after applying); a
module's as its broker secret, through the same delivery `module issue` uses; the
control plane's own as its `bus` secret. Idempotent, and worked out from where the
bus's module is assigned rather than from this process's environment, because this
process is still on the old bus when it runs and must be.
This is the half of design 28 task 5.2 the first live attempt found missing: a
credential was minted only at enrolment, at `module issue` and for a person, so no
machine already enrolled could ever be moved. `rollout check` was right to refuse;
now there is something to run first.
On a mesh that predates the record, every handover has to begin by writing down who
already holds the seat — otherwise the next holder cannot be assigned beside it,
because two eligible claimants with nothing on record are refused. Found on the
live mesh minutes after 0040 moved the bus seat's row: the standing holder no
longer satisfies what the seat delivers, on purpose, and so could not be recorded,
and so nothing could stand beside it.
Recording who already holds is not making a new holder. Derivation never read
what the seat delivers, so the standing holder holds regardless; when nothing is
on record and the named assignment claims the seat at its scope, only that claim
is checked. Every change of holder is still judged in full.
Two halves of novox/hq ADR 0131. Migration 0040 moves the mesh-broker row from
`amqp` to `mesh-bus`, so the seat's holder answers for the mesh's bus and not for
the wire protocol the old broker spoke — which is what let only the retiring
broker hold the seat that names the bus. Safe under the current holder: the
control plane composes its own address through the seat by name and the overview
derives holders by name; only registration and provision resolution read the
column. What must not happen in between is re-registering the current holder.
And the parser refuses a manifest that provides or requires `amqp`, each refusal
saying what to do instead: a module reaches the mesh's bus through the sdk and
depends on the seat, not on a protocol. A whole-catalogue test asserts nothing
beside this checkout names it; the three modules that did are removed there.
Two tests that used the old broker as a fixture now use the module that replaces
it or a manifest this package owns.
`seat <name> --to <node>/<module>` makes one assignment the holder of a seat in
the same write that removes the previous one. The row is new (migration 0039);
without one, the resolver derives the holder as it always did — the sole eligible
assignment, two refused — so nothing changes for a mesh that never hands a seat
over. With one, the recorded assignment holds and any other whose module could
hold the seat is eligible and silent: not refused, not holding. That is what lets
the next holder run beside the current one until the switch (hq design 26, design
28 task 5.3, ADR 0131).
Why: the controller finds its own bus through a seat, and the day that seat was
left with nobody in it — because two eligible holders could not coexist and the
old one's claim was taken away — the control plane looped for two hours while
every service stayed up. A handover that is never empty in between is the fix,
not a workaround for it.
`CanHold` is the one judgement of whether a module may hold a seat — claims it at
its scope, provides what it delivers, against the store's row — shared by
registration and the handover so they cannot drift apart. The holding belongs to
the assignment and goes when it does, so a seat never points at nothing running.
Tests: the resolver with and without a record, on the same and another machine,
under a former name; the store's row replaced not added, refused for an
unassigned target, removed with its assignment; CanHold's four answers and that
they follow the store. Full suite green against a real NATS and store.
`rollout check` said the old broker stays running as an ordinary provider of
amqp, and this was not its retirement. That was ADR 0127, which ADR 0131 has
superseded: AMQP is not a provision, so once every machine reports on the new bus
nothing of the mesh speaks to the old broker and its module is unassigned. The
plan says so, as its last step.
The flag that made the "it stays" line conditional is gone with the line — there
is no case in which the broker is kept. The test that pinned the opposite now
pins this, and says which record changed under it. The stale citation of a
record numbered 0119 is corrected while here.
A claim on a seat was checked against `SeatNamed` inside `ParseManifest`, and the
build machine parses manifests too. It has no store, so there it answered from the
set compiled into the binary — a copy of data the control plane owns (ADR 0122).
When the two disagreed, that copy won where it mattered. The store's row said the
bus seat answers for `amqp`; the binary's said `mesh-bus`; and a holder that
provides `amqp` was refused at build time for not providing `mesh-bus`. The seat
went unheld, the controller lost the address it composes through that seat, and the
control plane crash-looped on a bus that was healthy the whole time.
So the two checks that read the set — a seat's scope, and what its holder must
provide — move to CatalogueProblems, which runs only in the control plane and only
after UseSeats has replaced the set with the store's. The parser keeps what it can
judge from the manifest alone, the reserved-namespace rule included.
A test pins it: the same manifest, two different values in the store, and the answer
follows the store both times. It fails if the check moves back.
`nats.go v1.54.0` declares `go >= 1.26`, so `go mod tidy` puts the directive at
1.26.0 and will not leave it at 1.25. The base image is pinned by digest at
1.25.14-alpine, so nothing in this repository compiles against it — caught when the
build machine's own build failed with "go.mod requires go >= 1.26.0 (running go
1.25.14)".
Moved to 1.26.8-alpine by digest, same flavour as before. The pin stays a digest:
the point of pinning is that the compiler does not change underneath a build, and
that is still true one version along.
Two other manifests carry the same pin for the same reason — they compile this
repository's code — and move with it in the catalogue.
Both branches changed the seat set from the same starting point, so every number
collided and every `mesh-*` name existed twice. The trunk's numbers and names win:
this branch's records became 0129/0130 and its migrations 0037/0038, and the
hardcoded rename map gave way to the trunk's `seat_alias` table — a rename is a
row now (ADR 0122), not a recompile.
Three of my checks were wrong and the merge is what showed it:
A seat with an empty protocol is a marker, not an incomplete declaration. Most
node-scoped seats are markers — which module is this machine's packet filter —
and refusing one refused most of the set, the showcase module included. A
mistyped field name is already refused by the parser, so an empty protocol was
written as one deliberately.
A claim on a seat this manifest does not declare is not the parser's to judge. A
module may hold a seat another module declared; that is the whole reason ADR 0126
has callers name the seat and not its provider. Whether the seat exists is a fact
about the catalogue, so the refusal is at registration, where every declaration
is in view.
And a seat may share a name with the provision it delivers. `git`, the npm
registry and the artifact store still do, because renaming a delivering seat
cascades to every consumer requiring it, with a window where a holder stops
resolving mid-flight. The trunk deferred exactly those three on purpose.
Full suite green against a real NATS and store.