A secrets object with one local name delivered no file. Two requirements could share
a local name. secret recover and the export could not tell two locals apart. The
recipe check missed continued lines and read heredoc bodies as bases. repo:tag@digest
kept the tag in the repository. ask now publishes mandatory, so a tool nothing serves
is said at once rather than after the wait.
secrets: maps a requirement to several files under local names. Each local name is
its own need, its own pair credential (the pair is keyed on it: migration 0027),
its own file on the consumer, its own holder at the provider (the identity with the
local name after it) and rotates apart from the others. The plain shape is
unchanged and every existing row is the credential it was (novox/hq 04-ISSUES/069,
ADR 0094).
The host's error text may carry a duration or a counter, and a resource looping on
it would never have read as stuck. The previous row is read and compared here.
Stuck needs a start to say. A container may mount the file a binding lands in; the
runtime socket is declared under both of its spellings; the catalogue-wide test
takes MESH_CATALOG.
secret accept grows --provider: the value is sealed to the consumer's node, the
provider's node and the operator's key, and the pair records origin 'accepted'.
An accepted pair is not remade when a key changes (the mesh does not hold the
value; the read is refused naming the remedy) and rotate refuses it (accepting a
new value is the rotation). The vault's third species has its entry
(novox/hq 04-ISSUES/070, ADR 0092).
The mesh kept one report per machine, replaced, so a resource nothing can ever apply
looked like a failure that had just happened, every reconcile interval, for ever.
The row now keeps when the current failure began and how many reports in a row have
said it — the same outcome, refusal and failed resources; anything different starts
again and a clean apply clears it. Three make the machine stuck, and status says so
beside the failure, in words and in JSON (novox/hq 04-ISSUES/065, ADR 0090).
From review: the export counted any operator-sealed row as recoverable, so a
secret sealed to a replaced key was reported as openable with the current one;
replacing the key counted orphans in one table of two; and a pair credential
held from two providers was recovered as whichever row came first. The export
now lists what the current key opens, what an earlier key opens, and what has
no copy; `secret recover` takes --provider and refuses ambiguity; files that
must not exist are created exclusively; one constructor builds the export for
the operator's file and the vault's disk alike.
The secret the vault provides a module is the credential of the consumer↔vault
pair, and so is every credential a provider grants; sealing only own secrets
to the operator left exactly those unrecoverable. Same column, same call; the
export and `secret recover` address a pair by consumer node, module and the
provision's name, and say which kind each entry is.
novox/hq ADR 0085, amended: the mesh's root secrets — the store's superuser,
the broker's administrator, every secret a module holds for itself — were
sealed to a node key and nothing else, so a lost node took them with it.
Now the mesh records an operator's public sealing key and seals every own
secret to it as well, minted or accepted. The private half is written once
by `operator key new` to a file the operator keeps off the mesh; the mesh
holds one more blob per secret that it cannot open.
`secret recover` opens a secret with that key, to a 0600 file, from the
store or from an export; `secret export` writes every operator-sealed copy
as ciphertext. A module that `keeps` (the vault) is handed that export as a
declared file on its own disk, so recovery survives the store.
Secrets made before the key exists have no operator copy and are said so —
the plaintext was discarded — until each is issued again.
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Its event queue is durable, so a running catalogue misses nothing. What it cannot
have is what was announced before it first ran — and on a fresh mesh that is never
arbitrary: the shared base, the store the catalogue runs on, and the catalogue
itself are each necessarily built BEFORE a catalogue exists to hear about them.
The graph's foundation is the part it never sees.
So it says it is catching up, and the control plane re-announces what it
recorded, oldest first, marked as a replay. Oldest first because a graph is built
in the order things happened: registering a module that stands on a base before
the base would point an edge at a version nothing has seen, and the shape of a
fresh mesh guarantees the base is both first and the one that was missed.
The replayer hands announcements back rather than publishing them, because the
wire belongs to the link package and a replay building its own events could drift
from what the builder emits — the one thing it must match exactly, since the
catalogue has a single handler for both.
Its own queue and its own consumer: two consumers on one queue split its
messages, and a catch-up request going to whichever half was not listening is a
gap that looks like a working mesh.
Toward novox/hq 04-ISSUES/050.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The builder announces a build with the resolved manifest, the path inside the
repository, and every artifact it stood on. The control plane received all of it
and kept none of it.
That was survivable while the catalogue heard the same announcement directly. It
stops being survivable the moment the catalogue was not there to hear it — which
on a fresh mesh is always, and always for the same modules: the shared base, the
store the catalogue runs on, and the catalogue itself are each necessarily built
BEFORE the catalogue exists to hear about them. The graph's foundation is the
part the graph never sees.
Replaying those builds needs what they said, not a summary. Without the manifest
there are no requires/provides edges; without `against` there are no build edges,
which are the ones that answer "a base moved, what must be rebuilt". A replay
carrying neither would restore the module list and leave the question the
catalogue exists for still wrong, while looking fixed.
Kept null rather than empty where a build predates this, so a replay can say it
is holding nothing instead of inventing an empty declaration for a module that
certainly had one. And `built_against`, not `built_on`: that column exists and
means the machine, which is a different fact about a different subject.
Toward novox/hq 04-ISSUES/050.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The check read "declares no resources" as "runs nowhere", and those are not
the same. The private network declares no resources either — the control plane
computes them when it composes a machine's declaration — and it is assigned to
every machine that has to reach another one. Refusing it stopped a four-machine
bed at its first assignment.
The signal is narrower: it builds an artifact and places nothing. Made a
function of its own, because a judgement with a wrong answer this expensive
should be testable without a database — nothing guarded it, which is how it
shipped.
The image every module in the scripted toolchain is compiled on top of is
registered as a module so the mesh can build, version and depend on it. It is
not one: nothing about it belongs on a machine. Assigning it succeeded, the
machine was sent a declaration containing nothing of it, and everything
reported success — the operator had said run this here and the mesh had agreed
to something it cannot do.
A module whose resources are worked out per node is asked about separately, so
it stays assignable, which is the point of it.
A fingerprint written into a recipe names one particular copy of the base — the
copy on whichever machine the person typing it was using. On any other mesh that
copy has never existed, so the build stops on its first line with a message
about an image nobody can look up. Three modules in the catalogue were in
exactly that state, and the line each of them replaced was equally dead.
A module now names the module and artifact instead, and the mesh answers with
what it holds. The builder is still a thing that clones, builds and answers: the
answer travels with the question, because only the mesh knows what it has.
A base the mesh has not built is refused before anything is built, naming which
module has to exist first.
The builder says what it built and the catalogue decides whether that was an
upgrade. Only the control plane knows which machines run the thing, so it is
the one that acts — and what it does is a choice somebody recorded, not a
behaviour compiled in: record that they are behind, or send it, one machine at
a time or together.
Recording is the absence of an action rather than a second path: a machine not
running what the mesh would send it is already something the mesh reports.
Defaulted to recording. A mesh that rolls out everything it builds the moment
it builds it is reasonable to want and a bad thing to arrive by default — the
first module to inherit it would be the control plane, upgrading itself out
from under the push applying it.
The builder cloned a repository and read the manifest at its root, which means one
repository per module. Nothing we have is shaped that way, so the builder could be
asked to build nothing that exists (novox/hq ADR 0069).
The path travels the whole way — named when asking, carried in the request, used
to read the manifest and as the context everything is produced from, echoed back
in the result, and recorded as part of where a module came from. Without that last
part the mesh could notice a module was behind its source and then be unable to
rebuild it, which is the worst of both.
A path climbing out of the clone is refused: a machine whose job is building other
people's repositories must not read whatever else is on its disk.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
0056 is 'the authority is the control plane, not a database'. A citation
pointing at the wrong decision is worse than none: it reads as corroboration.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
`module forget` cascaded. The settings, the module's own secrets and the ports the mesh
chose all name the module by a foreign key that cascades, so removing the row took all
three and reported "forgotten" — an action succeeding into a state its own verify would
reject (novox/hq 04-ISSUES/017). A sealed secret is not recoverable afterwards, because
the mesh discarded the plaintext when it made it.
It now reads what it would destroy, names each thing one at a time, and refuses.
`--and-what-it-holds` is how somebody says they mean it, and the removal then reports
what went — this being the only record that any of it ever existed.
Reported as "operator settings do not persist, because re-registering a module
cascade-deletes them". Half of that is wrong, and the test now says so out loud: the
upsert is on the name, so `module add` at a new version leaves the settings, the secrets
and the ports exactly where they were. The command that destroyed them was `forget`, and
a wrong belief about which command destroys data is expensive in both directions — it
sends people looking for a fault that is not there, and leaves the real one unexamined.
Checked by internal/inventory/forget_test.go, which writes all three, re-registers the
module at a new version, reads them back, and only then tries to forget it.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A public route used to carry its whole hostname as a literal in the module
manifest, so running the same catalogue against a different domain meant
overriding that literal on every routed module, per node. The mesh was, in
effect, holding a map of names to services: the one thing it should never hold,
because the subdomain is the operator's choice and the domain is the node's.
Compose instead. A route contribution carries a `label` (the subdomain); a node
carries its `public_domain` as node-level configuration; the mesh joins
`<label>.<public-domain>` and grants exactly that, interpreting neither half.
Held as a node property beside the node's other node-level facts (endpoint,
site, overlay address), not in a module's settings — the ADR calls it
node-level, and the settings table is keyed per module.
Additive, so an unmigrated catalogue keeps working: a contribution that still
carries a full `name` and no `label` passes through unchanged, and the catalogue
can migrate module by module. A labelled contribution on a node with no public
domain composes nothing, reading downstream as a route that named no host.
And propagate: each granted route name is published into internal resolution
mesh-wide, mapped to the node that serves it, alongside the `<node>.internal`
names every container already gets. So a container — and an internal ACME
validator, which cannot complete a challenge for a name it cannot reach —
resolves a routed name to the proxy that serves it. Name-agnostic throughout:
the mesh propagates whatever names it was told to serve and knows nothing about
what they mean.
novox/hq 02-DECISIONS/0056
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The report carries the digest of the declaration it applied (mesh-host
8211d8b), and the mesh stores it beside the outcome. `reported` rows in
the status JSON now say `current`: whether the machine's last word
names the declaration last sent.
Not derivable from the timestamps beside it, which is why they were
not enough: an apply begun under the previous declaration reports
after the next send — newer, and still about the old words. The lab
lost exactly that race between one test's closing push and the next
test's opening one.
Empty digests — every host from before reports carried one — read as
not current, which errs toward waiting rather than toward asserting on
files that are not there yet.
"Not waiting" says the declaration is current, not that the machine
finished applying it: the sent digest is recorded at send. So a test
that pushed, saw waiting clear, and asked the machine what it was
running found containers that did not exist yet — the certificate fix
made compositions stable, and the settling that used to fail first had
been hiding the gap behind it.
The mesh already held the missing half: every machine's last report,
with its time. It just was not in the JSON. `reported` now sets each
machine's last word beside when the current declaration went to it, and
"has it caught up" becomes a comparison of two timestamps the mesh
recorded itself — a report newer than the send means the machine acted
on what was sent; older means it is still working, which waiting alone
cannot distinguish.
Three faults, one file split. All from reading, all verified to bite.
Unassign now releases the module's ports. ReleasePorts existed, said
"for when it is unassigned" in its own comment, and was called by
nothing — so a fixed port stayed claimed in the name of a module that
was gone, and the next module needing it was refused by a ghost.
Kept-once-chosen is a promise about a module that is still here.
MachineSide reads addressed mappings. "127.0.0.1:8080:80" was split at
the first colon, "127.0.0.1" failed to parse as a port, and the mapping
was silently skipped — putting the filter back on the declared port,
the exact fault the function was written to end. The machine side is
the second-from-last part, which is the reading the host already
applies, and the substrate bundle writes that shape today.
An allocation race answers in the mesh's words. Two concurrent picks of
the same port used to surface as a Postgres constraint violation,
verbatim. The table has two keys, so the collision is one of two facts:
the racer was this same assignment — then its answer is the answer,
kept-once-chosen does not care who chose — or another module took the
machine port, and an unfixed pick is simply made again against the
moved free list. A fixed port that lost the race is refused by name.
Told apart by re-reading the row, not by the constraint's name, so this
does not couple to the migration's spelling.
And the artifact-store cycle tests moved to bootstrap_cycle_test.go;
machineside_test.go had quietly become three subjects.
object-store.json and minio.json described the same thing: same image,
same provision at the same scope, same provisioner. Not two
implementations a person could choose between — one module written
twice. Assigning both to a node would have collided on `s3-bucket`.
It exists because it was written first, to pair with photos.json for the
README's worked edge, and minio.json was the fuller version of the same
module written later. Nobody removed the first.
The pair test keeps its point and now reads the surviving one. Checked
across the rest: this was the only duplicate.
The other half of ADR 0038, and what 04-ISSUES/028 was actually about.
A module can now avoid colliding with another module; until this it
could not avoid colliding with the mesh itself.
The substrate is not a module. A node raises it from the bundle it
carries before any mesh exists, so the control plane had never heard of
the store, the broker, or its own container — and handed a database
module 5432, which the store already had.
So the machine says. The host records what each resource binds,
distinguishing what it carried from what the mesh sent — a distinction
that already existed so the two never remove each other — and reports
the carried ones. The node states and this context writes, which is the
shape of every message between them.
What the declaration binds, not what is open. A machine's open ports are
a moving target, and assigning around them would mean a port that was
free when it was asked for and taken when it was used.
Replaced whole each time rather than merged: a machine that gave a port
back must be believed about that too, and a set that only grows keeps a
port reserved for something no longer there.
Tested against a real database, and the tests bite — removing the check
hands the module 20000, which the machine had said it holds.
novox/hq ADR 0038. A module cannot choose a port: it is written once and
assigned anywhere, so any number it picks is a guess about a machine it
has never seen. A database module met the mesh's own store on 5432 and
was told, by a container runtime three layers down, that the port was
already allocated.
The number used to appear three times in every module — the rule set,
what a consumer is told, and what the runtime publishes — agreeing only
because one person wrote all three. Now it appears once, in `listens`,
and the other two are derived: the container publishes `20000:5432`, the
consumer is told 20000, and the rule set opens 20000.
An assignment is made once and kept, as a credential is. A port that
moved on every declaration would restart both ends each time and hand a
consumer a number that was true when it was read.
Ports the protocol fixes — mail on 25, submission on 587, DNS on 53 —
say so, and are then claims: one holder per machine, and the second is
refused by name at assignment. That is the mechanism the mesh already
has for what is singular on a machine, pointed at ports.
A mapping written the long way is left exactly as it is. Some things
must be pinned by hand, and quietly overruling somebody who wrote both
halves would be worse than not offering the short form.
Still open, and known: the substrate is not a module, so the mesh has
never heard of its own store and cannot yet assign around it. That is
what 028 will still be about after this.
novox/hq 04-ISSUES/022. A credential was keyed by provision, consumer
node and provider node, so "who is asking" was answered by naming a
host. The node this mesh exists to take over runs eight modules against
one database server.
The symptom had two halves and only one was loud. The provider refused,
naming the modules and explaining they would share one credential, which
reads as a decision rather than a limit. The consumer did not refuse: it
resolved cleanly, wrote one module's credential file and left the others
absent — a service that starts and cannot authenticate, with nothing
saying why. That is 021 again on a different axis.
Three modules wanting one database produced one need, carrying whichever
module mentioned it first, because the resolution walk is a work-list
over names. The fan-out now happens in one place, after the walk. The
record path already did this correctly and said why: a consumer here is
a module on a machine. It is the same rule.
Downstream: the secret's key gains the consuming module, the grant file
is named after both halves, needs are matched by provision and module
rather than provision alone, and the provisioners name the role and the
access key after the module. The refusal in ContributionsTo is gone
because there is nothing left to refuse.
Worth stating plainly: without that refusal, gitea's login would have
opened keycloak's database. From the provisioner's side it created
exactly what it was asked to create.
Existing secrets are discarded rather than backfilled. They cannot say
which module they were for, and a secret is remade and delivered to both
ends on the next push — so this costs one rotation and invents nothing.
Also guards the role name against PostgreSQL's 63-byte truncation, which
is a notice rather than an error and would reintroduce exactly this
collision at a length nobody tests.
Three faults injected — the fan-out removed, needs matched by name
alone, the grant file named after the machine — each caught.
Provisions were named after roles: provides "database", requires
"database". Nothing distinguished engines, so a module written against
PostgreSQL could be matched to a provider of SQL Server, resolve as
satisfied, deploy, and fail on its first query — with nothing
connecting that error back to a match made elsewhere by something that
believed it had done its job.
The failure is in the direction that hides. Refusing on ambiguity
exists precisely so this does not happen, and the generic name walked
around it: with one provider of each name nothing is ambiguous, so
nothing is asked.
How it got in: every resolver test had exactly one provider per name,
so no mismatch was expressible and none was caught. The fixtures agreed
with the design — the same fault as the imagined test output in
04-ISSUES/005, at the level of a name.
Refused rather than documented, because the old naming *was* the
documented convention. Providing database/db/sql/sql-database is now a
parse error naming what to write instead.
The rule is about coupling, not specificity everywhere: route and
resolver stay role-named, because a consumer genuinely cannot tell
which proxy answered. novox/hq ADR 0027.
It meant "failed or refused". So a machine that applied cleanly and whose
declaration has since changed was not behind — and novox/hq ADR 0010's
question, did my change go out?, was answerable exactly for the machines that
broke. For every machine that worked, the answer was silence whether the change
had gone out or not, which is the thing replacing a pipeline was supposed not
to cost.
The mesh now records a digest of what it last sent each machine. A digest
rather than the declaration: it can compute what a machine should be at any
moment, and keeping a copy would be a second account of it able to disagree
with the first. What cannot be recomputed is what was actually sent.
Recorded after the send, not before — a digest kept for something that failed
to send would make the machine look current for a declaration it never
received.
Never told stays separate from out of date. The remedy is the same push and the
situations are not alike: nobody has ever asked that machine to be anything.
And a machine the mesh could not work out is not reported as waiting, because
saying so would invent a comparison — that is `plan`'s answer to give.
`status` says it and `push --behind` sends it, or the flag would know something
the person reading the status does not.
novox/hq ADR 0009: a capability's presence gates an assignment and its detail
carries a value — seat: card1-DP-1, an architecture, an amount of memory. So
'can this run here' and 'what should it be configured as' are one fact read two
ways, and the mesh was keeping the first read and discarding the second.
The reason an absent capability is absent went the same way, which is the case
a person most needs: 'this machine has no container runtime' is the answer and
'docker is not installed' is why, and only the machine knows why.
`node show` says it back. Never reported and reported nothing stay different
things there — one machine has not run the host, the other ran it and can do
nothing, and those send a person to different places.
The invariant novox/hq ADR 0001 records as unowned, and it was measurably
false in HAL: a provision documented as never rotating minted a new password on
every adoption and updated only the provider's row. Consumers on three nodes
held dead credentials for two days while the mesh reported success. Nothing
enumerated who held the old one.
Three things make that impossible here. The holders are a set the mesh can name
— each pair has its own credential, so rotating one consumer touches one role
and the affected list is a query rather than an assumption. Both ends are
pushed by this command rather than a later one, because leaving the sending to
whoever remembered is the fault exactly. And it is all-or-nothing: if any
affected machine cannot be resolved, nothing is sent and the old credential
keeps working, which is a mesh that has not rotated rather than one that has
half-rotated.
The window is stated rather than hidden: a role's password changes on the
provider and the file changes on the consumer, and they cannot be simultaneous.
The provisioner now takes its superuser password from the file the mesh wrote,
which is how the mesh delivers one. Passing it through the environment needed a
person in the middle of the one path that exists so there is not one — and put
a superuser password where `docker inspect` prints it.
Two kinds live in module_secret and they behaved identically, which is right
for one of them. A made secret is the mesh's: when a node regenerates its
sealing key the mesh makes another and nothing is lost, because nothing else
ever knew the old one.
An accepted secret is not. A broker account's password exists because the
broker was told about it. Regenerating one puts 32 random bytes where a working
credential was — and the machine applies it, reports success, and the program
reading it fails to authenticate somewhere else entirely, with the mesh
insisting the secret was delivered, which it was.
The row now records where the value came from, and a rejoined machine asking
for an accepted one is refused with the remedy named: issue it again. No amount
of pushing produces a password the broker has never heard of.
Found while making the builder a module, which is the first thing to hold one.
A board reads through interfaces and holds nothing. Everything it needs
is already answered — as text, for people, which is not something a page
can read.
`--json` rather than a serving API, because nothing needs one yet:
whatever serves a board runs the command, and the constraint holds either
way — the board never touches a context's store. An API is the larger
thing and should wait until something asks for it.
Both forms are gathered from the same reads before either says anything,
so they answer the same questions rather than being two implementations
that can drift. That was not true of the first version: the JSON printed
after the text, because the branch was too late.
Four properties, each asserted and each confirmed to fail when removed:
- refused and failed stay distinct all the way out. They are fixed in
different places, so one word for both sends half a page's readers to
the wrong one — and how much DID apply is carried, since "three of
eight" and "none of eight" are different machines
- a machine that never spoke carries no time at all, rather than a zero
one that any page would format as a date in 1970
- nothing is null. A page distinguishing "no machines are wrong" from
"this field is missing" has to handle both, and null is the one that
gets forgotten
- no field is named like a secret. Everything here comes from records
that hold no readable one, but a shape a page is built against is
exactly where one would eventually be added for convenience
Found by testing removal, which is the half nobody tests.
A grant was emitted for every secret the mesh held, whether or not the
machine still asked for it. So a consumer that was unassigned kept
appearing in its provider's manifest — and the provisioner's rule about
removing what nobody asks for can only fire if the mesh stops asking. The
login would have stayed live for ever, and nothing would have said so.
Skipped where the declaration is built rather than where grants are
gathered, so the rule holds whoever gathers them. No credential file is
written for a withdrawn consumer either, or the provisioner would find a
file its manifest does not mention and have to guess what that means.
The secret itself is deliberately kept. It is sealed and unusable to the
mesh, and a machine that comes back gets what it had — what withdraws the
login is the manifest, which is the thing that reconciles.
Two things, both found by trying to write a real postgres module and
discovering it could not be said.
A database has a superuser password, a broker an administrator, a
registry an account. None of them is *for* anybody — they are not the
credential a consumer is given, and the mechanism that hands those out
has a consumer in the middle of it. So a module may declare what it needs
and where to put it, and the mesh generates one per node, seals it, and
reads it no more than it reads any other.
Per node, deliberately: a module running on three machines has three
passwords. One in the manifest instead would put the same secret on every
machine that ever runs it, in a file anybody can read, for ever. Made
once and kept, or a running database would be handed a password it was
not started with; remade when the machine's sealing key changes, like
everything else sealed here.
A need declared and not made is refused rather than skipped, because a
module whose own credential is silently absent starts, fails to
authenticate, and the reason is three layers from the machine reporting
it.
And the provisioner can watch. That is what lets it be a module rather
than a binary somebody places: run once, it needs invoking after every
declaration by a timer or a unit wired to a file; watching, it is an
ordinary long-running service the host already supervises. It polls
rather than watching the filesystem, because the host writes atomically —
the file is replaced, so a watch on the path stops seeing anything after
the first replacement, and a watcher that silently stops working is worse
than a poll. Credentials are compared by digest and never held: this runs
for as long as the machine is up.
A node reports back after applying a declaration: it worked, some of it
failed, or the whole thing was refused. A refusal or a failure moved
last_seen and the reason went to a log line — so "which machine is not
doing what it was told" had no answer the next morning, which is the
question a mesh exists to answer.
Refused and failed are kept as different things, because they are
different situations with different remedies: refused means the machine
is exactly as it was and what is wrong is in what was sent; failed means
it is in a state nobody declared and what is wrong is on the machine. One
word for both would make the record say less than the node did.
One row per node, replaced. The question is the machine's current state —
"this failed an hour ago and then succeeded" is not a machine anybody
needs to look at, and a table of every report would bury the ones that
matter under the ones that do not.
`status` now answers three questions in the order somebody asks them: is
anything broken, is anything not answering, is anything out of date. The
first has consequences now, the third is a plan for later, and a status
leading with the third would bury the first. A machine that has never
spoken is reported as quiet rather than as broken — new, switched off and
unreachable are not the same as tried and could not.
The mapping from a report to an outcome had no test at all, which the
injection caught: it is the code deciding which of those situations a
machine is in. It has four now, including that a partial report never
becomes the account of what the machine holds — the fault that destroyed
a substrate once.
A build result was answered to whoever asked and kept nowhere. So "when
did this last build", "why did it fail" and "which machine built what is
running" had no answer, and a build nobody was waiting for was reported
into the void — which is the same as not reporting it.
Failures are recorded too, and that is the point rather than a detail: a
failed build that leaves no trace is indistinguishable from one nobody
asked for, and the difference is the whole of whether somebody should be
looking at something. A build that never learned what it was building
keeps the repository, because that is what a person goes and looks at.
Recording is idempotent on the correlation id, because a result can
arrive twice — as the answer to whoever asked, and on the exchange when
nobody was. Two rows would show one build as two, and which is real is
not answerable afterwards.
The serving control plane now binds `built` as well, so results from
builds it did not ask for are kept. It refuses them loudly when it has
nowhere to put them rather than dropping them, so the broker's own
counters show something arriving that nothing handles.
`builds [<module>]` reads it: what happened lately across the mesh, or
what has happened to one module — the first asked after something goes
wrong, the second when deciding whether to trust something.
What was published is kept with the build, so a digest traces back to
what made it without holding the manifest twice in a place that can
disagree with the first.
A build is work, not state. Everything else the control plane sends a
node is a declaration — this is what you should be — reconciled forever.
A build happens once and is finished. Putting it in a declaration would
mean rebuilding on every reconcile, or a declaration carrying "and I
already did this", which is state about an event rather than about a
machine.
So it travels on its own queue and the answer comes back correlated. One
queue, so several build machines share the work and each request is done
exactly once — which a per-machine routing key would not give.
mesh-builder is the program a build machine runs. Not the control plane,
which must not run commands on a machine; not the host, which would then
need a container runtime and git everywhere to do something almost no
machine will ever do. It holds its own broker credential and nothing
else.
Three properties that are decisions:
- a request is acknowledged only once the answer is away, so a builder
that dies mid-build leaves the work for another machine rather than
losing it with nobody ever hearing why
- one build at a time. Five at once against one runtime finishes all five
slower than it would have finished the first, and the queue is what
shares work between machines
- a failure is a RESULT. A build that fails silently is
indistinguishable from a builder that is not running, and those want
different responses
And `module list` is a catalogue: what exists, at which version, built
from which commit or handed over by hand or shipped with the control
plane, whether it is behind its source, and which machines run it. All of
that was recorded from the first build and none of it was shown, so "is
this current?" could only be answered by reading the database.
Proven against a real broker, registry and store: the mesh asked, a
builder consumed, built, published, answered; the manifest was recorded
with its commit; the source moved and the catalogue said "behind";
rebuilding caught it up with a new digest because the content changed.
The enrolment request is a struct in each repository. A node now reports
a third key — the one its secrets are sealed to — and that wiring had
unit tests on each side and had never been run across the join. A field
renamed on one side fails silently: enrolment succeeds, the key is
absent, and the node looks joined until the first thing sealed to it
cannot be opened, by which point nobody is looking at enrolment.
So the host's suite writes a real request and this one reads it, the same
way the declaration check already runs in the other direction. Both skip
with a reason when the neighbour is not checked out.
It does more than compare shapes: it seals something to the key that
arrived and opens it with the private half the host kept. Confirmed to
fail three ways — a renamed field, a value that is not a key, and a key
that is present, correctly named and simply somebody else's. Only the
last needs the sealing step, and it is the one a shape check would pass.
Also `inventory.ForTest`, because the check lives beside the link and a
second copy of the throwaway-database helper would be a second thing to
keep true.
HAL keeps env vars in the registry, encrypted at rest. Its own tooling
records what that bought and what it did not. `secret_locate` matches by
value rather than by name — because the same password sits in
mesh_provisions, in module_env, in each node's .env in plain text, and
inside every connection string composed from it, and its documentation
says those URL copies "are often the only copies actually in use". And a
query against the encrypted column returns zero rows and proves nothing,
so auditing moved to the decrypted copies on the nodes.
Two faults there, and encryption at rest addresses neither: the control
plane can read what it stores, so a copy of the database is a copy of
every credential; and one secret has many homes with nothing tracking
them.
So here the mesh generates a password, seals it to each end with keys
those nodes generated, stores both blobs, and discards the plaintext. It
cannot read what it holds. Neither can the broker relaying it. And
nothing is composed centrally — a connection string is assembled on the
machine that needs one — so no copy is ever minted in a shape nothing
tracks. `Compromise of a node is compromise of that node` (ADR 0004) is
now true of secrets, not only of identity.
Two files rather than one, because the mesh cannot compose a document
containing a value it discarded: `binds` carries the readable facts,
`secrets` carries the credential alone. The readable half stays readable
in the declaration; the secret half changes only when the secret does,
which makes restart-on precise. The provider gets a directory, one file
per consumer, for the same reason.
It is made once and kept — regenerating per declaration would restart
both ends on every push, and the password a provider was told to create
would never be the one its consumer was given. It is remade when either
end's sealing key changes, and both ends learn the new one in the same
push, so there is no window where half the mesh holds a dead credential.
Two tests found passing for the wrong reason, both caught because their
injection came back clean:
- the provider's copy was asserted non-empty, which reads the same
whichever column is selected. It now opens the blob with the
provider's own key.
- RotateSecret deleted and re-created; the re-create was dead, because
the next read makes one anyway. Removed, and a second path to the same
act is how two ends come to disagree.
And one real fault: three places built a declaration, and the one behind
`--json` predated credentials, so it silently produced a declaration
missing them — a difference between what `plan` showed and what anything
reading `--json` got. There is one path now.
Two different things were both written `requires`. A shell, a display
server and a private network have to be on the machine that needs them.
A database does not — it runs somewhere and is reached over the network.
Both were answered the same way, so requiring a database installed
PostgreSQL on every machine that ran a web application.
What a module provides now carries a scope, the same idea claims already
use, written short in the ordinary case:
"provides": ["shell"]
"provides": [{"name": "database", "scope": "mesh"}]
A mesh-scoped requirement is answered by finding the node already running
it — never by installing it here. Choosing a machine to put a database on
is a decision with consequences, and nothing resolving a web application
should make it silently. With nothing anywhere it refuses and says which
module to assign; with two it refuses and says how to choose.
Choosing is `pin <node> <provision> <from>`, kept per node because that
is the granularity the choice has. A pin at a machine that does not
provide it refuses rather than falling back — a fallback would quietly
move somebody's data. One provider does not overrule a pin either.
Resolving a node now needs to know what the others offer, and working
that out needs them resolved, so it is two passes: the first answers only
what each node offers, the second answers everything. Nothing is ever
declared from the first.
A node's plan says what it takes from elsewhere. It is the only part of a
set that stops working when a different machine goes away, and nothing
else in that output would have said so. It is also where a credential
will hang once there is a mechanism for handing one back.
One test found passing for the wrong reason: it read pins through a join
on the provider, which hides a dangling row whether or not it was cleaned
up. It counts rows now, and bites when the cascade is removed.