Step 3.4. On the new bus there is no reply queue to declare and no
correlation to check: each account is granted one inbox prefix and no
other, so an answer cannot reach the wrong asker. That settles a cost
build.go records having paid — on a shared reply exchange every asker saw
every result, which is why the correlation was checked rather than assumed.
And a tool nobody serves says so at once rather than after the whole wait.
The difference between "that module is down" and "that tool is slow" is the
first thing a person asking wants, and both tests are against a real server
because both are claims about what the server does, not about this code.
RequestBuild stays as it is, and is a different shape on the new bus rather
than the same one: a build takes minutes, so it is work submitted to a
queue with the outcome returning to a reply subject the request carries —
the pattern design 25 §2 already sets for anything crossing a stream. It
touches the builder too, so it goes with that conversion.
introduces
Step 3.4, the consume side's hard part. The guarantee (ADR 0083) is that a
push the controller cannot record because its store is restarting is held
and retried — never dropped, never falsely acknowledged. Keeping the
delivery unacknowledged in memory becomes a nak with a delay: the server
holds it, the controller keeps no list of parked messages, and a controller
that restarts mid-window loses nothing it was holding.
That is a plain win and it introduces one problem. Holding in memory let
the controller drop an older report when a newer one for the same node
arrived, "because acting on it after the newer would undo the newer". A
naked message is the server's and comes back whatever happened meanwhile,
so the older report is redelivered after the newer was applied.
The answer was already in the message. A report carries `Declared`, the
digest of the declaration it is about, which exists because an earlier
attempt to order reports by time lost the race it invited. So supersession
stops being something the controller remembers and becomes something it
checks — the same shape as a node refusing a superseded declaration by
sequence (issue 107): ordering settled by what a message says, not by when
it arrived.
Pure, so the guarantee is testable without a bus, a store or a clock. Nine
tests, including that staleness is decided before the store is waited on —
a redelivery that lost its race must not hold a slot a current message
needs.
Comments framed the new bus by what it replaces — a comparison in almost
every explanation, which reads as though NATS were a variant of the old
thing rather than the mesh's nervous system. Removed throughout, and
OverAMQP becomes OverCurrent: the seam's two sides are the bus the mesh
runs on today and the one being built, not two protocols.
What remains is the client library's own package name, which is its name.
Step 3.4, first half. Every one of these took an *amqp.Channel, so the
transport reached every caller and swapping it meant touching all of them.
The seam turned out to be small — the controller sends exactly two kinds of
message that expect no answer — which is the same measurement that said
this bus could be replaced at all.
Bus is stated in the mesh's words, not a transport's: PublishEvent and
PublishDeclaration. Two implementations, both shipping, because steps 1 to
4 leave every node on AMQP and the NATS one is selected at the rollout.
Both ship is also what makes them comparable: one conformance fixture holds
both to the same envelope, and the NATS one is checked against a real
server reading back from the stream rather than from the code that wrote it.
Still on *amqp.Channel: RequestBuild and Ask, which carry reply-queue
machinery, and the whole consume side — the control loop, enrolment, serve.
Every required header set, each value in the pinned shape, and the subject
derived the same way. Read from the sdk's conformance directory by sibling
path, never copied.
Review of the ADR 0105 build (hq ADR 0105). Four things it got wrong and one
path it lacked:
- A predecessor spoke's tunnel names one peer, the hub, routed the whole
range; recording refused it and the whole enrolment failed. Range-routed
peers are skipped now — only the hub's peers are ever carried.
- The range and the carried peers were conditions on the node being adopted,
so converging the hub would have renumbered the mesh and dropped the peers
still reaching it. They are facts of the tunnel record now, mode aside; the
takeover alone is declared to an adopted node. Converging the hub is refused
while a carried peer has not enrolled, naming it.
- A push composed a takeover for a hub whose address or endpoint disagreed
with the tunnel, which would have the host stop the found interface and
raise the mesh's where no peer listens. The graph refuses to compose it,
naming both and the placement that fixes it.
- The host's account said taken or not; "found down and the mesh's not up"
read as not taken. Three states now, and an account on every takeover.
- A hub that enrolled before this feature holds a key of its own, and
re-enrolling would rotate every key the mesh sealed credentials to. A node
now rekeys in a report, signed with its identity key over the key it
leaves, the key it takes and the tunnel; the mesh verifies against the live
key, refuses a stale or foreign proof, records key and tunnel, and moves a
hub to the tunnel's address. `overlay show` names the path for a hub that
found no tunnel.
Also: a carried IPv6 peer is routed /128, and identity.ForTest exists so the
link can be tested against a real identity store.
Three readers did not follow a moved foundation port (novox/hq 04-ISSUES/102),
and each took the control-node down in its own way: the control plane's own
store and broker connections, sealed at genesis with the port inside; and every
build the mesh ever recorded, kept as `<registry>:<port>/<module>/<artifact>@…`.
The control plane cannot open its own sealed connections to move a port, and it
cannot bind the store as a consumer would — a binding mints a credential. So its
settings get a third twin, `NAME_PORT`, read on top of the sealed value by the
store, the broker, the management API and the bus connection, and filled into
its container by a placeholder that names a seat, `${seat:mesh-store:5432}`,
from the node's given or mesh-assigned ports — never the manifest's number, and
empty when the mesh has nothing to add, so what genesis wrote stands. A value
that is still a placeholder is nothing said, aloud: the manifest naming it lands
in the next commit, once every control plane that composes it knows it.
A build is now recorded by digest and path — `artifact-store://<module>/<artifact>@…`
— and the store's address is composed in where a reference is used: the
declaration, the trust file, the bases a build is handed, a replay to the
catalogue. Over the network as `<node>.internal:<port>`; on the store's own node
before any network exists — every genesis push before its "network" step — by
loopback. A reference recorded before this, with an address, is re-routed the
same way when the mesh built it. The trust file and every provider's address
come from one derivation: the node's given port, over the mesh's assignment,
over the manifest's number.
novox/hq 04-ISSUES/102
On an adopted hub the private network takes over the tunnel it finds rather
than running beside it (hq ADR 0105): two tunnels leave the mesh's unreachable
through the provider's filter, so no machine can ever join.
The node presents the found tunnel when it enrols, under the key it took as
its own; the inventory records it (node.tunnel, tunnel_peer — migration 0031)
and the mesh composes from it: the overlay's range is the adopted tunnel's,
the hub is placed at the tunnel's address on the tunnel's port, and every
peer the tunnel had is carried in the hub's peer list as a peer of the
tunnel, not a node of the mesh, until a node enrols with that key — which
then keeps the address the tunnel had for it. A fresh node never gets an
address the tunnel holds. The hub's declaration tells the host which unit to
take over; the host's account of carrying it is recorded and shown.
Every reader of the range follows the setting; nothing stores it. A found
tunnel under another key is recorded and not adopted, so ADR 0100's
non-overlap rule keeps applying where a tunnel is left running beside the
mesh's. A lab bed and test skeleton for "How it is checked" are under lab/.
A secrets object with one local name delivered no file. Two requirements could share
a local name. secret recover and the export could not tell two locals apart. The
recipe check missed continued lines and read heredoc bodies as bases. repo:tag@digest
kept the tag in the repository. ask now publishes mandatory, so a tool nothing serves
is said at once rather than after the wait.
A module serves tools under an account scoped to exactly that, and nothing else in
the mesh held an account that could ask one. The control plane does: ask publishes
on the RPC exchange with a private reply queue bound under its own name, checks the
correlation, prints the answer, and exits non-zero for a tool that answered with an
error or a module that never answered (novox/hq 04-ISSUES/049, ADR 0095).
secrets: maps a requirement to several files under local names. Each local name is
its own need, its own pair credential (the pair is keyed on it: migration 0027),
its own file on the consumer, its own holder at the provider (the identity with the
local name after it) and rotates apart from the others. The plain shape is
unchanged and every existing row is the credential it was (novox/hq 04-ISSUES/069,
ADR 0094).
The broker settings take a _FILE twin like the store connections; the
catalogue engine refuses a secret placeholder in a container's env and a
secret-carrying env-file unless the container says why with
secrets-in-environment, which stays in the catalogue and never reaches the
machine.
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Its event queue is durable, so a running catalogue misses nothing. What it cannot
have is what was announced before it first ran — and on a fresh mesh that is never
arbitrary: the shared base, the store the catalogue runs on, and the catalogue
itself are each necessarily built BEFORE a catalogue exists to hear about them.
The graph's foundation is the part it never sees.
So it says it is catching up, and the control plane re-announces what it
recorded, oldest first, marked as a replay. Oldest first because a graph is built
in the order things happened: registering a module that stands on a base before
the base would point an edge at a version nothing has seen, and the shape of a
fresh mesh guarantees the base is both first and the one that was missed.
The replayer hands announcements back rather than publishing them, because the
wire belongs to the link package and a replay building its own events could drift
from what the builder emits — the one thing it must match exactly, since the
catalogue has a single handler for both.
Its own queue and its own consumer: two consumers on one queue split its
messages, and a catch-up request going to whichever half was not listening is a
gap that looks like a working mesh.
Toward novox/hq 04-ISSUES/050.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A fingerprint written into a recipe names one particular copy of the base — the
copy on whichever machine the person typing it was using. On any other mesh that
copy has never existed, so the build stops on its first line with a message
about an image nobody can look up. Three modules in the catalogue were in
exactly that state, and the line each of them replaced was equally dead.
A module now names the module and artifact instead, and the mesh answers with
what it holds. The builder is still a thing that clones, builds and answers: the
answer travels with the question, because only the mesh knows what it has.
A base the mesh has not built is refused before anything is built, naming which
module has to exist first.
The builder says what it built and the catalogue decides whether that was an
upgrade. Only the control plane knows which machines run the thing, so it is
the one that acts — and what it does is a choice somebody recorded, not a
behaviour compiled in: record that they are behind, or send it, one machine at
a time or together.
Recording is the absence of an action rather than a second path: a machine not
running what the mesh would send it is already something the mesh reports.
Defaulted to recording. A mesh that rolls out everything it builds the moment
it builds it is reasonable to want and a bad thing to arrive by default — the
first module to inherit it would be the control plane, upgrading itself out
from under the push applying it.
Answering and announcing are different acts. The reply goes to whoever asked and
is correlated to their request; the announcement says to the whole mesh that a
module now exists at a commit, which is what the catalogue places in the module
graph (novox/hq ADR 0072). A build nobody asked for still has to be announced, or
the graph knows less than the registry does.
What it was built on top of is read out of the build's own inputs rather than
declared, because a declared list drifts from what the code actually uses
(ADR 0009). These are artifact references, which is what a build input names;
resolving them to module-versions is the catalogue's work, since it is what knows
which module-version published which artifact.
Events ride the topic exchange, not the direct one nodes speak over, so the
builder's account is granted both: it must be able to answer and to announce.
The envelope is the sdk's, reproduced exactly — a second shape would be a second
thing for consumers to handle, and they are written against the first.
Announcing is not allowed to fail a build. The work was done and was answered; a
build reported as failed because saying so failed is a lie about it.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The builder cloned a repository and read the manifest at its root, which means one
repository per module. Nothing we have is shaped that way, so the builder could be
asked to build nothing that exists (novox/hq ADR 0069).
The path travels the whole way — named when asking, carried in the request, used
to read the manifest and as the context everything is produced from, echoed back
in the result, and recorded as part of where a module came from. Without that last
part the mesh could notice a module was behind its source and then be unable to
rebuild it, which is the worst of both.
A path climbing out of the clone is refused: a machine whose job is building other
people's repositories must not read whatever else is on its disk.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A node says it is there every minute and describes what it applied rarely,
and both went through Heard, which wrote every one down as a report. So a
bare alive replaced the node's last real apply with an empty one -- clearing
the declaration digest `current` is measured against, the carried ports a
push assigns around, and the clean-or-failed outcome. A node that had just
caught up read as behind within the minute, and never converged.
Whether it converged in time was a race the node's own apply set: the link's
one loop applies a declaration to completion before it can send the pending
heartbeat, so a fast apply (catalogue-small) leaves the digest standing the
~60s until the next beat -- long enough for the lab to see `current` -- while
a heavy wave whose apply outran the first beat (mongodb + unifi + marrytts)
had the alive fire milliseconds after the report and never showed `current`
at all, timing out settle even at 1200s.
Heard now returns after moving last_seen for a report that carries no account
of what the machine did -- nothing applied, nothing refused, nothing failed,
which is exactly a bare alive. A real report always carries one. This is what
the commit that began hearing alives said it did and did not: "a bare word
that a node is there moves last_seen and touches nothing else."
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The report carries the digest of the declaration it applied (mesh-host
8211d8b), and the mesh stores it beside the outcome. `reported` rows in
the status JSON now say `current`: whether the machine's last word
names the declaration last sent.
Not derivable from the timestamps beside it, which is why they were
not enough: an apply begun under the previous declaration reports
after the next send — newer, and still about the old words. The lab
lost exactly that race between one test's closing push and the next
test's opening one.
Empty digests — every host from before reports carried one — read as
not current, which errs toward waiting rather than toward asserting on
files that are not there yet.
The other half of ADR 0038, and what 04-ISSUES/028 was actually about.
A module can now avoid colliding with another module; until this it
could not avoid colliding with the mesh itself.
The substrate is not a module. A node raises it from the bundle it
carries before any mesh exists, so the control plane had never heard of
the store, the broker, or its own container — and handed a database
module 5432, which the store already had.
So the machine says. The host records what each resource binds,
distinguishing what it carried from what the mesh sent — a distinction
that already existed so the two never remove each other — and reports
the carried ones. The node states and this context writes, which is the
shape of every message between them.
What the declaration binds, not what is open. A machine's open ports are
a moving target, and assigning around them would mean a port that was
free when it was asked for and taken when it was used.
Replaced whole each time rather than merged: a machine that gave a port
back must be believed about that too, and a set that only grows keeps a
port reserved for something no longer there.
Tested against a real database, and the tests bite — removing the check
hands the module 20000, which the machine had said it holds.
novox/hq 04-ISSUES/022. A credential was keyed by provision, consumer
node and provider node, so "who is asking" was answered by naming a
host. The node this mesh exists to take over runs eight modules against
one database server.
The symptom had two halves and only one was loud. The provider refused,
naming the modules and explaining they would share one credential, which
reads as a decision rather than a limit. The consumer did not refuse: it
resolved cleanly, wrote one module's credential file and left the others
absent — a service that starts and cannot authenticate, with nothing
saying why. That is 021 again on a different axis.
Three modules wanting one database produced one need, carrying whichever
module mentioned it first, because the resolution walk is a work-list
over names. The fan-out now happens in one place, after the walk. The
record path already did this correctly and said why: a consumer here is
a module on a machine. It is the same rule.
Downstream: the secret's key gains the consuming module, the grant file
is named after both halves, needs are matched by provision and module
rather than provision alone, and the provisioners name the role and the
access key after the module. The refusal in ContributionsTo is gone
because there is nothing left to refuse.
Worth stating plainly: without that refusal, gitea's login would have
opened keycloak's database. From the provisioner's side it created
exactly what it was asked to create.
Existing secrets are discarded rather than backfilled. They cannot say
which module they were for, and a secret is remade and delivered to both
ends on the next push — so this costs one rotation and invents nothing.
Also guards the role name against PostgreSQL's 63-byte truncation, which
is a notice rather than an error and would reintroduce exactly this
collision at a length nobody tests.
Three faults injected — the fan-out removed, needs matched by name
alone, the grant file named after the machine — each caught.
Provisions were named after roles: provides "database", requires
"database". Nothing distinguished engines, so a module written against
PostgreSQL could be matched to a provider of SQL Server, resolve as
satisfied, deploy, and fail on its first query — with nothing
connecting that error back to a match made elsewhere by something that
believed it had done its job.
The failure is in the direction that hides. Refusing on ambiguity
exists precisely so this does not happen, and the generic name walked
around it: with one provider of each name nothing is ambiguous, so
nothing is asked.
How it got in: every resolver test had exactly one provider per name,
so no mismatch was expressible and none was caught. The fixtures agreed
with the design — the same fault as the imagined test output in
04-ISSUES/005, at the level of a name.
Refused rather than documented, because the old naming *was* the
documented convention. Providing database/db/sql/sql-database is now a
parse error naming what to write instead.
The rule is about coupling, not specificity everywhere: route and
resolver stay role-named, because a consumer genuinely cannot tell
which proxy answered. novox/hq ADR 0027.
08-connectivity keeps two authorities apart on purpose: a public one for
names the outside world reaches, and the mesh's own for names only the
mesh knows. Nothing implemented the second, so anything between machines
was plaintext or trust-on-first-use — which the design refuses everywhere
else.
A node now generates a fourth key at enrolment and reports the public
half. A fourth, because a key used for two purposes is one rotation away
from breaking the other: the identity key signs messages to the mesh and
would do for TLS, and reusing it would mean rotating a node's identity
every time its certificate is replaced.
**Nothing secret travels and nothing is sealed.** A certificate authority
says "this name belongs to the holder of this key", so the mesh signs a
public half it cannot use, and the certificate it issues is public. A
module asks for one and is given the certificate and, if it wants,
the mesh's own — the private key is a path to a file the machine already
has, the same arrangement the private network's key uses.
Asserted by verifying rather than inspecting, because a certificate that
parses and does not chain fails at the moment something connects:
- what the mesh issues verifies against the mesh, for the name asked for
- the name is in the subject alternative names, since a certificate
carrying it only in the common name is refused by every modern client
- it certifies the key the node generated and no other
- another mesh's certificate does not verify, which is the whole point of
two authorities being separate
- the authority cannot sign another authority — one that could is one
that can be delegated without anybody deciding to
- two control planes starting together agree on one authority, or a mesh
has certificates half its machines refuse
Certificates last ten years, which is a choice: a short life needs
something to renew it, and a renewal that fails silently is a mesh that
stops trusting itself on a date nobody wrote down. What makes one
replaceable is that the mesh reissues on demand, not that it expires.
A node reports back after applying a declaration: it worked, some of it
failed, or the whole thing was refused. A refusal or a failure moved
last_seen and the reason went to a log line — so "which machine is not
doing what it was told" had no answer the next morning, which is the
question a mesh exists to answer.
Refused and failed are kept as different things, because they are
different situations with different remedies: refused means the machine
is exactly as it was and what is wrong is in what was sent; failed means
it is in a state nobody declared and what is wrong is on the machine. One
word for both would make the record say less than the node did.
One row per node, replaced. The question is the machine's current state —
"this failed an hour ago and then succeeded" is not a machine anybody
needs to look at, and a table of every report would bury the ones that
matter under the ones that do not.
`status` now answers three questions in the order somebody asks them: is
anything broken, is anything not answering, is anything out of date. The
first has consequences now, the third is a plan for later, and a status
leading with the third would bury the first. A machine that has never
spoken is reported as quiet rather than as broken — new, switched off and
unreachable are not the same as tried and could not.
The mapping from a report to an outcome had no test at all, which the
injection caught: it is the code deciding which of those situations a
machine is in. It has four now, including that a partial report never
becomes the account of what the machine holds — the fault that destroyed
a substrate once.
The builder was documented as holding its own broker credential and
nothing else, and nothing issued one — so in practice it used whatever it
was handed, which was the broker's administrative account. A program
documented as holding its own credential and given somebody else's is
worse than one with no story at all.
`builder issue <name>` creates an account that may read the build queue
and write to the mesh exchange. Not a node account: a build machine is
not a node, and a node's queue carries its declarations.
Two faults found by running it, both about the answer path:
- the reply queue was left for the broker to name, and the account was
scoped to `amq.gen-*` — one broker's convention. The builder built,
could not answer, and the connection closed. Reply queues are named
here now, deterministically.
- the answer then went via the DEFAULT exchange, where permission is
granted per exchange rather than per queue. A builder allowed to use it
could publish into any node's queue, which is the privilege a build
machine most obviously should not have. Answers go through the mesh
exchange, which it already may use, and an asker binds its reply queue
to the same key and filters by correlation.
Verified against a real broker: a builder cannot consume a node's queue
and cannot publish to the default exchange. That check nearly reported
the opposite — an unconfirmed publish is asynchronous, so the refusal
arrives as a channel close afterwards and a naive test sees success. With
publisher confirms it is immediate. A negative security assertion made
against an asynchronous call is not an assertion.
Redelivery was observed working while fixing this: builders that died
before answering left their work on the queue, and the next builder did
all of it.
Also: the queue and exchange names exist in both `broker` and `link`,
because `link` imports `broker`. A test in an external package keeps them
agreeing — a builder scoped to a queue nothing publishes to takes no work
and says nothing about why.
A build result was answered to whoever asked and kept nowhere. So "when
did this last build", "why did it fail" and "which machine built what is
running" had no answer, and a build nobody was waiting for was reported
into the void — which is the same as not reporting it.
Failures are recorded too, and that is the point rather than a detail: a
failed build that leaves no trace is indistinguishable from one nobody
asked for, and the difference is the whole of whether somebody should be
looking at something. A build that never learned what it was building
keeps the repository, because that is what a person goes and looks at.
Recording is idempotent on the correlation id, because a result can
arrive twice — as the answer to whoever asked, and on the exchange when
nobody was. Two rows would show one build as two, and which is real is
not answerable afterwards.
The serving control plane now binds `built` as well, so results from
builds it did not ask for are kept. It refuses them loudly when it has
nowhere to put them rather than dropping them, so the broker's own
counters show something arriving that nothing handles.
`builds [<module>]` reads it: what happened lately across the mesh, or
what has happened to one module — the first asked after something goes
wrong, the second when deciding whether to trust something.
What was published is kept with the build, so a digest traces back to
what made it without holding the manifest twice in a place that can
disagree with the first.
A build is work, not state. Everything else the control plane sends a
node is a declaration — this is what you should be — reconciled forever.
A build happens once and is finished. Putting it in a declaration would
mean rebuilding on every reconcile, or a declaration carrying "and I
already did this", which is state about an event rather than about a
machine.
So it travels on its own queue and the answer comes back correlated. One
queue, so several build machines share the work and each request is done
exactly once — which a per-machine routing key would not give.
mesh-builder is the program a build machine runs. Not the control plane,
which must not run commands on a machine; not the host, which would then
need a container runtime and git everywhere to do something almost no
machine will ever do. It holds its own broker credential and nothing
else.
Three properties that are decisions:
- a request is acknowledged only once the answer is away, so a builder
that dies mid-build leaves the work for another machine rather than
losing it with nobody ever hearing why
- one build at a time. Five at once against one runtime finishes all five
slower than it would have finished the first, and the queue is what
shares work between machines
- a failure is a RESULT. A build that fails silently is
indistinguishable from a builder that is not running, and those want
different responses
And `module list` is a catalogue: what exists, at which version, built
from which commit or handed over by hand or shipped with the control
plane, whether it is behind its source, and which machines run it. All of
that was recorded from the first build and none of it was shown, so "is
this current?" could only be answered by reading the database.
Proven against a real broker, registry and store: the mesh asked, a
builder consumed, built, published, answered; the manifest was recorded
with its commit; the source moved and the catalogue said "behind";
rebuilding caught it up with a new digest because the content changed.
The enrolment request is a struct in each repository. A node now reports
a third key — the one its secrets are sealed to — and that wiring had
unit tests on each side and had never been run across the join. A field
renamed on one side fails silently: enrolment succeeds, the key is
absent, and the node looks joined until the first thing sealed to it
cannot be opened, by which point nobody is looking at enrolment.
So the host's suite writes a real request and this one reads it, the same
way the declaration check already runs in the other direction. Both skip
with a reason when the neighbour is not checked out.
It does more than compare shapes: it seals something to the key that
arrived and opens it with the private half the host kept. Confirmed to
fail three ways — a renamed field, a value that is not a key, and a key
that is present, correctly named and simply somebody else's. Only the
last needs the sealing step, and it is the one a shape check would pass.
Also `inventory.ForTest`, because the check lives beside the link and a
second copy of the throwaway-database helper would be a second thing to
keep true.
HAL keeps env vars in the registry, encrypted at rest. Its own tooling
records what that bought and what it did not. `secret_locate` matches by
value rather than by name — because the same password sits in
mesh_provisions, in module_env, in each node's .env in plain text, and
inside every connection string composed from it, and its documentation
says those URL copies "are often the only copies actually in use". And a
query against the encrypted column returns zero rows and proves nothing,
so auditing moved to the decrypted copies on the nodes.
Two faults there, and encryption at rest addresses neither: the control
plane can read what it stores, so a copy of the database is a copy of
every credential; and one secret has many homes with nothing tracking
them.
So here the mesh generates a password, seals it to each end with keys
those nodes generated, stores both blobs, and discards the plaintext. It
cannot read what it holds. Neither can the broker relaying it. And
nothing is composed centrally — a connection string is assembled on the
machine that needs one — so no copy is ever minted in a shape nothing
tracks. `Compromise of a node is compromise of that node` (ADR 0004) is
now true of secrets, not only of identity.
Two files rather than one, because the mesh cannot compose a document
containing a value it discarded: `binds` carries the readable facts,
`secrets` carries the credential alone. The readable half stays readable
in the declaration; the secret half changes only when the secret does,
which makes restart-on precise. The provider gets a directory, one file
per consumer, for the same reason.
It is made once and kept — regenerating per declaration would restart
both ends on every push, and the password a provider was told to create
would never be the one its consumer was given. It is remade when either
end's sealing key changes, and both ends learn the new one in the same
push, so there is no window where half the mesh holds a dead credential.
Two tests found passing for the wrong reason, both caught because their
injection came back clean:
- the provider's copy was asserted non-empty, which reads the same
whichever column is selected. It now opens the blob with the
provider's own key.
- RotateSecret deleted and re-created; the re-create was dead, because
the next read makes one anyway. Removed, and a second path to the same
act is how two ends come to disagree.
And one real fault: three places built a declaration, and the one behind
`--json` predated credentials, so it silently produced a declaration
missing them — a difference between what `plan` showed and what anything
reading `--json` got. There is one path now.
09-the-node-lifecycle asks for this in as many words -- *how long it has been
disconnected is a fact the mesh must hold, and nothing holds it today. Without
it, a node running last month's assignments looks exactly like one that is
current.* Now it holds it.
`node list` says "here", "out of touch 4m", or "never spoken", and the third is
kept distinct from the second on purpose: a node that has never spoken did not
finish joining, and a node last heard from a month ago is running a month-old
picture of the mesh. Those need different responses from a person.
A bare word that a node is there moves last_seen and touches nothing else. It
is not an account of what the machine holds, and recording it as one would
replace the recovery copy with an empty list every minute -- so a rebuilding
node would then be told it owns nothing and remove whatever it found. There is
a test for exactly that.
Heard is silent in the log. A node saying it is there every minute would fill
the log with the ordinary case, and a log where the ordinary case is loud is a
log nobody reads.
Verified in the lab across the threshold, both directions.
Three machines across two sites, two of them behind no reachable address, all
nine paths open. The mesh computes the graph, delivers it as a declaration, and
the nodes bring it up.
Every fault below looked like success from inside the mesh: the graph was
right, the files were right, the services were up, every node reported it had
applied. None was reachable by reasoning.
A running interface does not re-read its configuration. A node joins, every
existing node's peer list changes, the file is replaced -- and the service is
already running, so nothing reloads it. Fixed as declared state rather than a
command: the service must reflect the file. A command to restart would be an
action, and the link may not carry one. The host refused exactly that, which is
how this shape was arrived at.
A hub sharing a site with a spoke appeared twice in that spoke's peer list --
once as a direct peer, once as the route of last resort. WireGuard takes one
entry per key and refuses the file. The ordinary shape of a small mesh, and in
none of the tests written before it ran.
Two nodes at one site that neither can be dialled were peered directly. Nobody
opens the path, and the direct route is more specific than the hub's, so it
wins and blackholes -- this design's own warning arriving in its
implementation. They now route through the hub unless one end can be dialled.
And Docker sets the FORWARD policy to DROP, so a hub with ip_forward enabled
carried nothing between its spokes. The substrate at tier 1 silently breaks the
network at tier 2, and nothing in either tier's state says so. The hub inserts
its own rule above those chains and removes it on the way down.
Two weak tests found by injection along the way: one asserted the keepalive
rule only against the hub, whose peer entries happen not to set that field at
all, so it tested an absence; the other checked the firewall rules by looking
for FORWARD anywhere, which the PostDown line satisfies on its own.
novox/hq 09-the-node-lifecycle asks for this and it was missing: the host
reports what it owns and the mesh keeps the last report. A backup, never a
source -- nothing decides anything from it, and a node that disagrees with it
wins, because the node is the one that can see the machine.
Its point is the orphans. A node that loses its state file currently strands
whatever it applied: nothing on the machine knows those resources were the
mesh's doing, so nothing removes them. With this, a rebuilt node receives both
the declaration and the record of what it previously owned.
Never reported and reported nothing are kept apart, and that is the whole care
in it. A node that applied nothing holds nothing; a node that has never spoken
is unknown -- and handing back an empty list for the second would tell a
rebuilding node it owns nothing and have it remove whatever it found.
The age comes back with the answer rather than being left for the caller to go
and find. An answer about a machine is worth much less without one, and this
repository has already been bitten by a cache with no age on it.
A refusal or a partial failure moves last_seen and nothing else: neither is an
account of what the machine holds, and recording one as though it were would
tell a rebuilding node to remove what it still has.
The control queue was bound to enrol and not to report, so every report a node
sent was accepted by the broker, matched no binding, and dropped. The publisher
saw success and the consumer saw nothing, for an afternoon.
The refactor that was meant to bind both never applied -- it left behind a
helper nothing called, which compiled and passed vet. The loop is now where the
bind is, so there is one place to forget rather than two.
`declare` sends a node a signed declaration; `serve` now also consumes reports.
Signed over the exact bytes published, which is what the node verifies. Anything
re-encoding in between would sign one thing and check another, and a difference
in key order alone would have a node refuse a declaration that was genuinely
the mesh's.
Sent to the node's queue directly rather than through the exchange: a
declaration is for one node, and routing by name through a shared exchange
means a binding per node that nothing removes when a node is retired.
Enrolment now issues the node its own broker password, replacing the token's
secret, and tells it the broker address, the fingerprint and the signing key --
so a node can reconnect after a restart without a person and a new token, which
is what makes disconnection ordinary rather than a crisis.
A report is a statement, not a write. What a node says it applied is its own
account of its own machine, kept as a copy for recovery rather than as a source.
There was no chicken-and-egg to solve. The mesh runs the broker, so it creates
the node's account when it issues the token, and the one-time secret is that
account's password. A joining node's first connection is already authenticated;
enrolment is what it says once it is in. I had been treating this as a decision
that needed taking, and it did not.
The account is per node and scoped: it may read its own queue, write to the one
exchange, and configure nothing else. The patterns are anchored and the node
name is constrained to characters that cannot widen them, because a name
carrying a dot or a star would silently let that node read everybody's queues.
`serve` is the control plane running: one connection, one queue, one consumer.
One deliberately -- two consumers on a queue get round-robined and each receives
half of what it expects, which has happened on this project before, between a
module's daemon and its capability server.
Enrolment spends the token first, in the single statement that both finds and
marks it, and only then records the key. That order is the order things become
irreversible: recording a key for a node whose token turned out to be spent
would leave the mesh believing a machine that never had the right to join.
Refusals are one message for every reason. The log says which, where an
operator can see it; the node is told only that the token cannot be used.
Verified in the lab, on a sealed machine, through the whole first-node path.