0056 is 'the authority is the control plane, not a database'. A citation
pointing at the wrong decision is worse than none: it reads as corroboration.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A store connection string carries a password, and the control plane took it
from MESH_STORE_<CONTEXT> — an environment variable, which is readable in
`docker inspect`, in the process's own /proc entry, and in whatever composed
it. Every other module in the catalogue is given secret material as a file the
mesh sealed to the machine and the host wrote.
That difference is what stopped the control plane from being an ordinary module
(novox/hq ADR 0067). A manifest can put a sealed value into a file's `content`
with ${secret:…}; it has no substitution into a container's `env` at all. So a
control-plane manifest could be written with the password in it, or without the
setting — neither honest. The fix is not to change the manifest format but to
let the control plane read what everything else reads: a file.
MESH_STORE_<CONTEXT>_FILE names one. Exactly one of the two may be set; both is
refused rather than settled by precedence, because whichever won, the other
would still read as the setting in force and the process would be writing to a
store nobody expects. Trailing whitespace is trimmed — a file written by a
person or by a filled-in placeholder ends in a newline, and a newline inside a
URL is rejected several layers from anything that could explain it. Leading
whitespace is left, being a mangled value rather than a habit.
The no-leak property is kept and extended: a file that cannot be read names its
path, never its contents, and the parse failure now names whichever source was
used because a variable name and a path are not the secret.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The sibling of node public-domain, and the worse one: a placement is three facts
declared together, so an invocation that said none of them took all three away —
the endpoint every other machine dials, the site, and the hub. A mesh whose hub
was placed that way has no paths left, at the moment somebody was trying to look
at it.
--nothing keeps the real case (a machine that roams and opens every path itself)
sayable, by name.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
`status --json` emitted no JSON at all when a single node was unresolvable. A blocked
node is not on the private network, and a mesh whose hub is that node has no hub — which
came back through the reading as a refusal, so `status` printed nothing and `status
--json` put multi-line prose on stderr and not one byte on stdout. A machine-readable
interface that stops being machine-readable exactly when something is wrong is one nobody
can build an alarm on.
Why a machine cannot be worked out is read as data now, per machine, through the same
whoResolves the private network is built from — so this and the network agree about who
could not be resolved rather than deciding it twice. The private network failing to
compute is kept as a note beside it instead of ending the read: it is almost always a
consequence of those same refusals, and every question that does not depend on it is
still answered.
It reaches all three ways of saying it, from the one reading: the text form leads with it
because a machine here is in none of the answers below, the JSON carries `unresolved`
(always a list, never null) and `network`, and the page has a section of its own.
That also closes a silent success. A machine that resolves to nothing has nothing
computed for it, so there is nothing to compare it against and nothing it can be behind —
it appeared in no answer at all, and `status` reported a mesh where nothing could be sent
anywhere as "all doing what they were told".
statusAsJSON takes the whole reading now rather than a growing argument list, which is
what let an answer be added to the text form and forgotten here. The two are one
function's output in two shapes and must not be able to differ about what was asked.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
`node public-domain <name>` cleared the domain. It reads like a question — it is exactly
what anybody types to find out what the answer is — and it silently took every routed
name the node had. There is no output that makes up for that: by the time it prints, the
fact is gone, and the mesh cannot tell a person what a domain used to be.
The bare form reports now. Clearing is still a real thing to want — a machine that stops
facing the outside composes no names, and lab-versus-production is this one setting
(novox/hq ADR 0056) — so it keeps a way to be said, by name: `--clear`. A domain and
`--clear` together are refused rather than one of them silently winning.
The other `node` subcommands were checked. `add`, `list` and `show` write nothing they
were not asked to, so there is nothing to make consistent with.
`overlay place <node>` with no flags has the same shape — it clears the endpoint, the
site and the hub flag — and is deliberately left alone here. It is a verb rather than a
question and every caller passes flags, so the fix is a different judgement and belongs
in its own change.
The usage text gains the three forms, and `module forget`'s new flag, neither of which
it named before.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
An assignment was verified by resolving the node it was made on. The verify that matters
is resolution over all of them: a module offering a mesh-scoped provision stops offering
it the moment its own node stops resolving, so an assignment could be reported as fine
while it took that provision away from every consumer elsewhere. Those consumers were
then told "nothing in this mesh provides it", naming as the remedy a module that was
already assigned — a wrong answer about a machine nobody had touched.
That is novox/hq 04-ISSUES/017's shape exactly: an action succeeds into a state its own
verify rejects, and it does so because the action's own test is not the test the verify
uses. 017's remedy was to make them the same test, and this makes them the same test.
The assignment is still kept, and that is the other half of the decision. Assignment is
not an ordering: a consumer assigned before its provider does not resolve for as long as
it takes to assign the provider, and refusing the first half of a pair would make the
order somebody types two commands in part of the mesh's rules. So `assign` and `unassign`
now name every OTHER machine that cannot be worked out as things stand, in the mesh's own
words, beside whatever they already said about this one. It reports the state and never
claims causation — saying "this assignment broke laptop" would mean resolving the whole
mesh twice and would still be a guess about which of several changes did it.
It costs a resolution per machine. Assignment is a person typing a command, and being
told which machines this just blocked is worth more than the milliseconds.
This is what a four-node raise read as "bumping a module's version broke provider
recognition". It was neither the version nor the provider: nothing in this codebase reads
the version column, every lookup is keyed on the module name alone, and a test in the
previous commit now says so. It was one machine's set of assignments, and nothing said so.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
`module forget` cascaded. The settings, the module's own secrets and the ports the mesh
chose all name the module by a foreign key that cascades, so removing the row took all
three and reported "forgotten" — an action succeeding into a state its own verify would
reject (novox/hq 04-ISSUES/017). A sealed secret is not recoverable afterwards, because
the mesh discarded the plaintext when it made it.
It now reads what it would destroy, names each thing one at a time, and refuses.
`--and-what-it-holds` is how somebody says they mean it, and the removal then reports
what went — this being the only record that any of it ever existed.
Reported as "operator settings do not persist, because re-registering a module
cascade-deletes them". Half of that is wrong, and the test now says so out loud: the
upsert is on the name, so `module add` at a new version leaves the settings, the secrets
and the ports exactly where they were. The command that destroyed them was `forget`, and
a wrong belief about which command destroys data is expensive in both directions — it
sends people looking for a fault that is not there, and leaves the real one unexamined.
Checked by internal/inventory/forget_test.go, which writes all three, re-registers the
module at a new version, reads them back, and only then tries to forget it.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The co-located fix could not reach a grant assembled for a consumer on another
machine: ContributionsFrom never sees a port map, so the proxy was told the
workload's software port and dialled a number that machine never published. The
consumer's own assignments are fetched where the grant is built and applied
there. The same fault as 038, one node over.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A whole-mesh push refused outright the moment any single node failed to resolve:
every node was composed first, and one entry in `refusals` returned "nothing was
sent" before the send loop ran. So on the ADR 0056 lab, one module on the anchor
requiring a provision nobody had assigned a provider for — `nothing provides
"acme-ca", wanted by route-proxy` — stopped every OTHER machine from being sent
anything. Nothing converged anywhere, and the machines that went unconverged were
the ones with nothing wrong with them. The failure and the punishment were on
different machines.
This is the rule a92c11b established one level down, where an un-hostable module
stopped taking down the healthy modules beside it, applied one level up: the blast
radius of a fault is the thing that has it. A node whose declaration composes is
sent; a node whose does not is named, with its reason, and the push still ends
non-zero — skipping is not succeeding, and a command that exits cleanly having
missed a machine is a command that lies. The message now says how many machines
WERE sent, because "nothing was sent" was the claim that had become untrue.
sendTo keeps its all-or-nothing rule, and its comment now says that push
deliberately does not share it: a rotation reaching the consumer and refusing on
the provider leaves one end holding a credential the other has never heard of,
which is a real coupling between two named machines. A whole-mesh push has no such
coupling and never did.
Composing is split out of pushCommand so the rule can be tested without a broker
and a database.
novox/hq ADR 0056
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A public route used to carry its whole hostname as a literal in the module
manifest, so running the same catalogue against a different domain meant
overriding that literal on every routed module, per node. The mesh was, in
effect, holding a map of names to services: the one thing it should never hold,
because the subdomain is the operator's choice and the domain is the node's.
Compose instead. A route contribution carries a `label` (the subdomain); a node
carries its `public_domain` as node-level configuration; the mesh joins
`<label>.<public-domain>` and grants exactly that, interpreting neither half.
Held as a node property beside the node's other node-level facts (endpoint,
site, overlay address), not in a module's settings — the ADR calls it
node-level, and the settings table is keyed per module.
Additive, so an unmigrated catalogue keeps working: a contribution that still
carries a full `name` and no `label` passes through unchanged, and the catalogue
can migrate module by module. A labelled contribution on a node with no public
domain composes nothing, reading downstream as a route that named no host.
And propagate: each granted route name is published into internal resolution
mesh-wide, mapped to the node that serves it, alongside the `<node>.internal`
names every container already gets. So a container — and an internal ACME
validator, which cannot complete a challenge for a name it cannot reach —
resolves a routed name to the proxy that serves it. Name-agnostic throughout:
the mesh propagates whatever names it was told to serve and knows nothing about
what they mean.
novox/hq 02-DECISIONS/0056
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A module a person assigns to a machine that cannot host it — its declared
capability has no detector there, as fail2ban does on a host with no firewall —
made Resolve refuse the entire node, so a whole-node push refused to send the
healthy modules beside it too. One module on the wrong machine took down every
other module on that node.
Assign already keeps such an assignment on purpose (it is what a person meant,
and acts.go says so), so the fix is on the resolve/push side: a directly-assigned
module the machine cannot host is left out of the closure and reported as
un-applied on the Resolution, rather than refusing the set. The healthy modules
still resolve, declare, and converge. A module that is *required* by something
running here and cannot be hosted still refuses — that set is genuinely
incoherent — so the distinction is who wanted it.
assign, plan and push now name the un-applied module and the missing capability,
via a shared WrongMachine message, so it is neither silently dropped nor fatal.
Reconciled two tests that encoded the old whole-node refusal for directly-assigned
un-hostable modules; added coverage for the healthy-modules-still-converge case
and the required-un-hostable-still-refuses distinction.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The refreshable-grant refresh token no longer rides a custom at-rest envelope that a
module opens with a node private key. A module is never given a node's private sealing
key, so that path could not exist -- the gap Phase C hit.
Instead the refresh token is a credential sealed to the MANAGER holder with the same
anonymous box (secrets.Seal / crypto_box_seal) every credential uses, stored as one
sealed blob, and delivered by the existing host-unseal-and-mount: the host opens it with
the node's real key and mounts the cleartext at the manager module's bound path, exactly
as a consumer's db password is delivered.
- refresh_grant now stores { sealed, manager_key }, dropping the AtRest token/wrapped_key
columns; internal/secrets/atrest.go is retired (nothing else used it).
- the licence records its manager as (node, module); KeyFor delivers the refresh token to
the manager holder and the access token to consumers, disambiguated by module so the two
can co-locate. Accept and the reseal skip the manager holder.
- the manager holder is delivered the node's PUBLIC sealing key in its bound facts, so the
module can re-seal a rotated refresh token with no private key of its own; the
declaration tolerates its empty pre-adoption secret rather than refusing.
- SubmitRefresh / set-grant take a sealed blob, never a refresh token in the clear.
The invariant holds unchanged: the control plane never reads the refresh token, and no node
but the manager holds it. A committed cross-language test proves the TypeScript module seal
opens under Go box.OpenAnonymous (the host's Unseal) -- both are NaCl crypto_box_seal.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Phase C of model-access (ADR 0050). Refresh above calls an in-process
VendorRefresher, which would open the at-rest envelope inside the control
plane's own process. Anthropic must not: its refresh runs on the manager
node. So add SubmitRefresh, the companion that publishes a refresh a
manager node already performed -- it is given only the new access token in
the clear (sealed per holder, as any accepted key) and an opaque re-sealed
refresh envelope (stored unopened). The refresh token in the clear never
crosses this boundary. The reseal-and-publish half is extracted and shared
with Refresh, so the sealing logic is one implementation.
CLI: licence grant (print the opaque envelope), set-grant (store a
module-produced envelope -- adoption), submit-refresh (access token +
optional rotated envelope). Tests defend that the manager alone opens the
refresh token and the control plane never holds it in the clear.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The ADR 0050 carve-out, built generic and vendor-neutral. A refreshable-grant
licence records one manager node; that node holds the refresh token encrypted at
rest, access tokens are still sealed per holder, and the refresh token is never in
a holder's delivery. Bounded on the three stated axes: refreshable-grant vendors
only, the refresh token only, the manager node only. Anthropic's actual OAuth
refresh stays a Phase-C plug-in behind a clean seam.
- New at-rest crypto (secrets.SealAtRest/OpenAtRest): envelope encryption distinct
from the per-holder anonymous-box seal. The refresh token is under a symmetric
data key (secretbox); the data key is wrapped to the manager node's public
sealing key. The database alone holds ciphertext and a wrapped key with no
private half to open either — only the manager node reads it back.
- Refreshable-grant adapter dispatch: anthropic is now refreshable-grant,
anthropic-api-key the static-key second case. The adapter implements the
Refresher seam by delegating to an injected VendorRefresher (the Phase-C plug,
none shipped). static-key is untouched. The type assertion to Refresher is what
gates the carve-out to refreshable-grant vendors.
- Refresh lease/rotate/publish flow (Licences.Refresh): a transaction-scoped
advisory lock is the single-refresher lease; the new access token comes from the
vendor refresh, is sealed per holder (secrets.Seal, as Accept does) and delivered
on the next push — doc 13's reseal-and-publish half, all-or-nothing. The refresh
token stays put, re-encrypted at rest only if the vendor rotated it.
- Manager and refresh_grant schema: consolidated into migrations/0001 and carried
by a new incremental 0003 (the dual-write rule).
- 17 new tests, including the four security checks: KeyFor never carries the
refresh token, a static key has no manager and cannot be refreshed, the at-rest
token needs the manager's key, and a refresh delivers a new sealed access token.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
ADR 0050 Phase A. Rename the licence's `provider` field to `vendor` — the
inventory already uses "provider" for which node answers a brokered provision,
and one word must not carry two facts — and route the licence layer's sealing
and delivery through a per-vendor adapter selected by that field.
The rename touches the Go struct/params/SQL in internal/licences, the operator
CLI, and the schema: 0001 (the consolidated schema) now creates the column as
`vendor`; a new guarded 0002 renames it on a database that predates the change,
and is a no-op on a fresh one.
The adapter (internal/licences/adapters) has a `shape` and the two verbs a
static-key vendor needs — accept (the generic anonymous-box seal) and deliver
(the sealed blob unchanged). refresh/identity/usage are named as optional
capability interfaces so the refreshable-grant seam exists before its code.
A registry maps vendor→shape (anthropic→static-key for now, with a Phase-B
TODO to swap it to refreshable-grant); an unknown vendor is refused clearly.
Behaviour is unchanged from the operator's view except the field name.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A module may declare a short `slug`; the mesh derives mesh_<node>_<slug|name> and
refuses at assignment (naming the slug as the remedy) when it would still overflow —
identityLimit is now 20, an S3 access key's, the tightest of the backends a login
reaches (04-ISSUES/010). The slug rides the grant so the provider derives the same
login the consumer does, even across nodes. CheckIdentity is now wired, in grantsFor.
Also, the minted secret shrinks to 40 chars (30 bytes) from 43: an S3 secret key is
8-40, the same fit-the-tightest-backend rule on the credential's other half.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
LavinMQ refuses a non-administrator declaring a queue with a dead-letter
exchange, so a scoped module cannot make its own. EnsureModuleQueue declares
<node>.<module>.events with its DLX as the mesh, and 'module issue' does so
for a consuming module — the runtime then passively checks it rather than
declaring. Verified against a real broker: the scoped account binds and
consumes the pre-declared queue.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
'module issue <module> --node <m>' looks up the module's emits/consumes from
the catalogue, ensures the bus exchanges exist, creates its scoped account
(CreateModuleAccount), and seals an amqps {url,fingerprint} to the node as the
module's broker own-secret — the same delivery as 'builder issue', now generic.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The report carries the digest of the declaration it applied (mesh-host
8211d8b), and the mesh stores it beside the outcome. `reported` rows in
the status JSON now say `current`: whether the machine's last word
names the declaration last sent.
Not derivable from the timestamps beside it, which is why they were
not enough: an apply begun under the previous declaration reports
after the next send — newer, and still about the old words. The lab
lost exactly that race between one test's closing push and the next
test's opening one.
Empty digests — every host from before reports carried one — read as
not current, which errs toward waiting rather than toward asserting on
files that are not there yet.
"Not waiting" says the declaration is current, not that the machine
finished applying it: the sent digest is recorded at send. So a test
that pushed, saw waiting clear, and asked the machine what it was
running found containers that did not exist yet — the certificate fix
made compositions stable, and the settling that used to fail first had
been hiding the gap behind it.
The mesh already held the missing half: every machine's last report,
with its time. It just was not in the JSON. `reported` now sets each
machine's last word beside when the current declaration went to it, and
"has it caught up" becomes a comparison of two timestamps the mesh
recorded itself — a report newer than the send means the machine acted
on what was sent; older means it is still working, which waiting alone
cannot distinguish.
Confining allocation to the push path took `plan` with it, and `plan`
belongs on the other side: it is a person asking what a push would do to
one named machine, so the port it shows and the secret it seals must be
the ones a push would use. Both are kept once chosen, so showing
numbers a later push would replace answers a question nobody asked.
Caught by the lab: composing the real modules stopped producing
postgres's sealed superuser, because nothing had minted it and the
read-only path correctly declined to.
The line is not question versus command. It is a person asking once
about one machine, against the mesh asking continuously about all of
them — the second is what hung, and the second is what reads.
object-store.json and minio.json described the same thing: same image,
same provision at the same scope, same provisioner. Not two
implementations a person could choose between — one module written
twice. Assigning both to a node would have collided on `s3-bucket`.
It exists because it was written first, to pair with photos.json for the
README's worked edge, and minio.json was the fuller version of the same
module written later. Nobody removed the first.
The pair test keeps its point and now reads the surviving one. Checked
across the rest: this was the only duplicate.
`status` hung. It composes a declaration for every node to answer *is
this machine running what I would send it*, and composing one assigns
each module a machine port — so the question wrote to the database, and
wrote to the same rows as the machine it was asking about.
`port_assignment` is unique on (node, machine). Two transactions
inserting the same port do not race, they queue: the second waits on the
index until the first commits. A status polled every two seconds while a
node applies is two writers on those rows, and the poll stopped
returning rather than returning something wrong — which is the better
failure of the two, and still a failure.
The latent version of this was there before anything polled: two
compositions running at once could both allocate.
So allocation belongs to the send path alone. The mesh chooses a port
when it commits to sending one; every other caller reads what was
chosen. A module with nothing assigned has never been sent, which is
precisely what "waiting" means — the read needs no number to be right
about that, and inventing one would make the answer worse.
Named rather than passed as a bare bool: at three call sites, `true` and
`false` say nothing about which of these two things is meant.
Checked by the lab, which now polls status throughout an apply.
The lab caught this: a module declaring a port and running no container
had its rule set opened on 20000 while its service sat on 9101. The
firewall reported success and blocked the thing it was told to admit,
which is the precise failure the filtering comment warns about, arrived
at from the other side.
Assignment was applied to every declared port. But a container's mapping
is the thing that translates, and where there is none the software binds
what it binds — the mesh choosing a number does not move the service, it
only makes the mesh wrong about where it is.
The declaration side already knew this: publishedOn rewrites container
ports and nothing else. Filtering did not, so the two disagreed about
the same fact. MachineSide is now the one derivation both follow.
It also fixes a second case nobody had hit yet: a mapping the manifest
wrote itself, like the mail system's 7080:80. That is passed through
untouched when composing, so assigning it a machine port would have
opened a rule on a port the container does not publish. Either side of
such a mapping now names it, and the host side is the answer — a module
may read `listens` as what its software binds or as what the machine
exposes, and both readings want the same number.
Recorded either way, assigned or not: the map means where this module's
port is on this machine, and every reader needs that answer regardless
of who chose it.
Tests bite — making it always assignable reproduces the lab failure.
novox/hq ADR 0038. A module cannot choose a port: it is written once and
assigned anywhere, so any number it picks is a guess about a machine it
has never seen. A database module met the mesh's own store on 5432 and
was told, by a container runtime three layers down, that the port was
already allocated.
The number used to appear three times in every module — the rule set,
what a consumer is told, and what the runtime publishes — agreeing only
because one person wrote all three. Now it appears once, in `listens`,
and the other two are derived: the container publishes `20000:5432`, the
consumer is told 20000, and the rule set opens 20000.
An assignment is made once and kept, as a credential is. A port that
moved on every declaration would restart both ends each time and hand a
consumer a number that was true when it was read.
Ports the protocol fixes — mail on 25, submission on 587, DNS on 53 —
say so, and are then claims: one holder per machine, and the second is
refused by name at assignment. That is the mechanism the mesh already
has for what is singular on a machine, pointed at ports.
A mapping written the long way is left exactly as it is. Some things
must be pinned by hand, and quietly overruling somebody who wrote both
halves would be worse than not offering the short form.
Still open, and known: the substrate is not a module, so the mesh has
never heard of its own store and cannot yet assign around it. That is
what 028 will still be about after this.
novox/hq 04-ISSUES/022. A credential was keyed by provision, consumer
node and provider node, so "who is asking" was answered by naming a
host. The node this mesh exists to take over runs eight modules against
one database server.
The symptom had two halves and only one was loud. The provider refused,
naming the modules and explaining they would share one credential, which
reads as a decision rather than a limit. The consumer did not refuse: it
resolved cleanly, wrote one module's credential file and left the others
absent — a service that starts and cannot authenticate, with nothing
saying why. That is 021 again on a different axis.
Three modules wanting one database produced one need, carrying whichever
module mentioned it first, because the resolution walk is a work-list
over names. The fan-out now happens in one place, after the walk. The
record path already did this correctly and said why: a consumer here is
a module on a machine. It is the same rule.
Downstream: the secret's key gains the consuming module, the grant file
is named after both halves, needs are matched by provision and module
rather than provision alone, and the provisioners name the role and the
access key after the module. The refusal in ContributionsTo is gone
because there is nothing left to refuse.
Worth stating plainly: without that refusal, gitea's login would have
opened keycloak's database. From the provisioner's side it created
exactly what it was asked to create.
Existing secrets are discarded rather than backfilled. They cannot say
which module they were for, and a secret is remade and delivered to both
ends on the next push — so this costs one rotation and invents nothing.
Also guards the role name against PostgreSQL's 63-byte truncation, which
is a notice rather than an error and would reintroduce exactly this
collision at a length nobody tests.
Three faults injected — the fan-out removed, needs matched by name
alone, the grant file named after the machine — each caught.
novox/hq ADR 0035: one implementation, several surfaces, and a surface
holds no decisions. The act of assigning — including that an assignment
which does not resolve is kept and still refused — moved into acts.go,
and the command line now calls it too. Two surfaces, one refusal, in the
same words.
It will not run without --issuer, and refuses at start rather than per
request so it is found by whoever ran it rather than by whoever finds
it. There is no flag that removes the check.
The authenticator is honest about what it is: no token can be verified
until an identity provider exists, because that is a module and none is
running, so every request is refused and told that the command line
still works. A surface that functioned without authentication would be
one somebody left running — and the board this stands behind is
published on a public name.
Four refusals, four tests. The last one first asserted "not 200", which
passed because a request with no database fails at the store anyway — it
proved nothing about whether the input was checked. It now asserts the
specific refusal, and bites when the check is removed.
The command refused every real invocation. Go's flag package stops
parsing at the first non-flag argument, so with the positionals first —
the order that reads correctly — `--from -` stayed among them and the
count check rejected it.
The host's own parser carries a note about this exact fault, and the
version it describes is worse: there a flag somebody passed was silently
ignored and the command succeeded anyway. This one at least refused.
The tests did not catch it because every case in them was a rejection.
The command was broken in the only way that matters — it refused what it
is for — and the suite was green. The lab found it at the first call.
Two tests now: the helper, and the command itself with a --from naming a
file that is not there, so the complaint must be about the file rather
than about usage. The second exists because injecting against the first
stayed silent: testing the helper alone left the command free to ignore
it entirely.
The entry point for adopting something already running, and the half
that was missing. The store has carried the distinction since the
beginning — a module secret records whether it was `made` or `accepted`,
and refuses to invent a replacement for the second — and
AcceptSecretForModule existed, with exactly one caller: the broker
account issued to a build machine. Nothing else could write one.
Without it every module secret is generated, which against a database
that already exists puts 32 random bytes where a working credential was.
The machine applies it, reports success, and whatever reads it fails to
authenticate somewhere else entirely, with the mesh insisting the secret
was delivered — which it was.
The value is read from a file or from standard input, never from an
argument: a value on the command line is in the shell's history and in
the process list. Same path a model-access key already takes, and no new
dependency — the first version reached for x/term and the existing one
needed nothing.
Sealed on the way in, plaintext discarded, and not printed back. The
only difference from a generated secret is where the value came from.
Two rules with a test each, and the second is the one that would have
been got wrong: only the line ending is removed, never surrounding
space. Trimming both ends is the obvious thing and would deliver a
password chosen with a leading space as a different password, silently.
Both were briefly untested for different reasons — the trimming lived
where no test could reach it, and then a -run filter matched neither
test. Extracted, and injected against the whole suite.
It sat beside `secrets` — where a *provision's* credential lands on a consumer.
Both were name-to-path, both held something secret, and the names
distinguished them not at all. Reaching for the wrong one parsed cleanly and
failed somewhere else entirely, which is the shape of fault this whole design
exists to prevent, sitting in the manifest format.
The axis that separates them is not how secret they are — both are — but
whose. `secrets` is keyed by the provision it is for and belongs to a
relationship with another machine. `own-secrets` is keyed by a name the module
chose and belongs to nobody else.
A manifest using the old name is told the new one rather than refused with
"unknown field": whoever wrote it knew what they meant, and the mesh knows what
it is called now. An invented key is still refused as one rather than guessed
at.
Found by auditing the 19 manifest fields for whether any could be mistaken for
another. This was the only pair that could — and while checking it, a second
instance of the same collision turned up one layer down: `Manifest.Needs` and
`Resolution.Needs` were different concepts sharing a name in Go. The rename
separates those too.
Working out what a machine should be reaches the identity context for its
certificate and the licence context for its model access. Both were opened —
and waited on — inside functions called for every node in a push. Two machines
hid it. Fifty would be fifty connect-and-wait cycles for data that does not
change while the push runs.
So a command holds what it has open, and passes it. Each context is opened on
first use rather than up front, because most commands need one and paying to
reach three would be the same waste from the other side.
The contexts stay separate, which is the point: this is one struct holding
three connections to three databases, not one connection to a shared one. No
context reaches another's store, and each still holds only its own credential
(novox/hq ADR 0008).
A pure move again — the gate is green before and after, and no test changed.
2,769 lines and 59 functions, holding command parsing, store opening,
resolution, the board, rotation, licences, builds and status rendering.
Nothing in it was wrong. It grew because appending was always the cheapest next
step, and no single edit was the one that should have been a new file.
That is exactly how novox/hq ADR 0001 records `hal/sdk` reaching 155 files and
34,636 lines — "containing code from every context", with each addition
avoiding a cycle and none of them the mistake. This is the same shape at 8% of
the size, which is why it is worth doing now rather than noting.
Eight files, along boundaries that already existed: what a machine is; the
private network; the catalogue; working out what one machine should be; sending
it; builds; the three questions; and reaching each context's store. main.go
keeps what a main is for — parsing arguments and dispatching.
A pure move. No behaviour changed, no test changed, and the gate is green
before and after — which is the only thing that makes a refactor this size
safe to do in one commit.
Services are named under the machine they run on — postgres.novox.internal,
plex.ace.internal. The first label is the service and the rest is the node, so
what has to resolve is anything under a node's name. What routes it once it
arrives is a proxy's concern and stays separate.
A hosts file cannot do that. It answers exact names, and a wildcard there would
mean writing down every service in advance — which is the enumeration the
arrangement exists to avoid. novox/hq 08-connectivity named this exact case as
the trigger for needing a resolver rather than a file, and it is the first
thing to meet it.
The mesh writes the data and runs no daemon. A resolver is third-party
software, and third-party software runs on the mesh rather than being of it
(ADR 0001): the mesh has no business shipping one, choosing which one, or
knowing its configuration language. What only the mesh can know is which
machines exist and where they are. A module that runs a resolver requires what
this provides and reads one file, so swapping the daemon changes that module
and nothing here.
Separate from names rather than part of them: a machine with no container
runtime can still have a hosts file, and folding them together would take exact
names away from a machine that cannot run a daemon in order to give it a
wildcard it cannot use either.
A machine with no address is left out. A wildcard pointing at nothing is worse
than no wildcard — every name under it resolves and then hangs, where an
unresolvable name fails at once and says which name it was.
Internal names are written to the machine's hosts file, which serves the
machine and not what the machine runs: a container gets its own hosts file
holding only its own hostname. So every name the mesh wrote was invisible to
the majority of things that need one — and on the machine it always worked,
which is exactly what made it easy to miss.
It was hit for real in the lab, and worked around by resolving the address on
the machine and passing it in. That workaround is now removed, and its absence
is the assertion.
A file rather than a resolver, which is the decision the mesh already made
about names and this extends rather than overturns: it works on every runtime,
needs no package and has no failure mode of its own. The stated trigger for a
resolver — names that are not one-per-node, service names, wildcards — is
still not met.
Given by the mesh, not chosen by a module: a module that listed the machines
would go stale the day one joins, and one that did not would be a module whose
containers cannot reach anything by name. A container that named its own keeps
them and gets the mesh's beside them.
Only containers, and not the ones on the machine's own network: a runtime
refuses to write a hosts file for those, and a file or a service given the
field is a declaration the host refuses outright — so getting it wrong breaks
the whole machine for something that was never about names.
Adding "not running what the mesh would send it" to `status` and not to the
board would have left two answers to one question with a person in front of
each — which is the single thing this page's design forbids, introduced by the
change that was supposed to make the question answerable.
The published JSON carries it as well, so the page, the command and anything
built against either say the same thing from the same read. Additive, because
that shape is hard to change once anything is built against it.
Never told stays separate from out of date on the page as it is everywhere
else: same remedy, and nobody has ever asked that machine to be anything.
It meant "failed or refused". So a machine that applied cleanly and whose
declaration has since changed was not behind — and novox/hq ADR 0010's
question, did my change go out?, was answerable exactly for the machines that
broke. For every machine that worked, the answer was silence whether the change
had gone out or not, which is the thing replacing a pipeline was supposed not
to cost.
The mesh now records a digest of what it last sent each machine. A digest
rather than the declaration: it can compute what a machine should be at any
moment, and keeping a copy would be a second account of it able to disagree
with the first. What cannot be recomputed is what was actually sent.
Recorded after the send, not before — a digest kept for something that failed
to send would make the machine look current for a declaration it never
received.
Never told stays separate from out of date. The remedy is the same push and the
situations are not alike: nobody has ever asked that machine to be anything.
And a machine the mesh could not work out is not reported as waiting, because
saying so would invent a comparison — that is `plan`'s answer to give.
`status` says it and `push --behind` sends it, or the flag would know something
the person reading the status does not.
novox/hq ADR 0009: a capability's presence gates an assignment and its detail
carries a value — seat: card1-DP-1, an architecture, an amount of memory. So
'can this run here' and 'what should it be configured as' are one fact read two
ways, and the mesh was keeping the first read and discarding the second.
The reason an absent capability is absent went the same way, which is the case
a person most needs: 'this machine has no container runtime' is the answer and
'docker is not installed' is why, and only the machine knows why.
`node show` says it back. Never reported and reported nothing stay different
things there — one machine has not run the host, the other ran it and can do
nothing, and those send a person to different places.
novox/hq 03-DESIGN/01-to-be/11-a-board.md, built. The board being replaced is
one service reading every context's database directly — ADR 0008 violated by
the one component with a reason to violate it. The cost is not hypothetical: a
boundary nothing may cross can move, and one thing crossing it is enough to
freeze it. A board that reads the provisioning tables breaks when provisioning
changes them, and the change then gets weighed against the board.
So the three questions are read once, by one function, for all three ways of
saying them — a person's status, its JSON, and this page. Three
implementations of "which machine is not doing what it was told" would be three
chances to disagree.
Refused and failed stay distinct all the way to the page: refused means the
machine is exactly as it was and what is wrong is in what was sent; failed
means it is in a state nobody declared. Different places to fix, so one word
for both would send half the readers to the wrong one.
It stores nothing, changes nothing, and every action it might offer already
exists as a command. A board that cannot reach the mesh says so rather than
rendering an empty page — an empty page says "nothing is wrong" in the one
situation where nobody can know that.
One test earns its place twice: a machine's own words are the whole reason the
page is useful and the one thing on it nobody in this repository wrote, so they
are shown and are not markup.
A status that names what is wrong and not what to do about it makes somebody go
and find the command — and the command is the whole point of having noticed.
The hint existed on one of the two paths that print this.
Its own store, its own test database, the same shape every other context has.
Five properties: a key with nobody to seal it to is refused rather than kept
readably; a key is sealed once per holder and the blobs differ because they are
sealed to different machines; a holder recorded afterwards has none and the
existing ones keep theirs; releasing a consumer takes its key; and a licence
nobody recorded is refused by name.
The last was the only one whose message mattered and whose message was not
checked — the database's own foreign-key error is true and mentions a
constraint, which sends somebody to read a schema instead of typing the name
they meant.
Partial sealing now says how far it got. The person holding the key is the only
one who can finish, and running it again knowing what it will do is different
from running it hoping.
novox/hq ADR 0010 replaced a pipeline with a comparison, and named the risk:
losing the question "did my change go out?". The mesh could already answer
which modules are behind their source — and then a person read that list and
retyped each repository, which is a person being the loop, and the loop is the
thing the pipeline was doing before it was taken away.
The mirror of `push --behind`, with the same argument and the same refusal to
combine the two forms: naming a repository and asking which need building are
different requests.
One failing does not stop the others, for the same reason one broken module no
longer blocks a machine's whole declaration: a mesh where one bad repository
holds back nine good ones is a mesh where nobody dares add the tenth.
Each is built from its own recorded ref rather than the commit the mesh
happened to notice — pinning to that would quietly turn a tracked branch into
a pin.
novox/hq ADR 0024, gaps 1 and 2. The user's stated requirement, and the first
thing here that no machine can answer: a hosted model is on nobody's node and
is reached over the public internet, so the rule that refuses two ends sharing
no private network must not apply to it.
A licence is a named thing and the name is the operator's — *the personal
account*, *the organisation's* — because the whole point is saying which one a
given consumer uses, and an anonymous credential hanging off a provider cannot
be said. Many to many, so deliberately not a claim: two machines sharing an
account is ordinary rather than a collision.
Gap 2 is the missing verb, *accept*: take a value somebody supplied, seal it to
each holder, discard the plaintext. With the consequence stated rather than
hidden — a holder recorded after the key was supplied has no key and the mesh
cannot make one, so it is refused by name with the remedy, not silently handed
an empty file.
Refusal is felt, as the record warns: a mesh holding three ways to reach a
model refuses every consumer that has not chosen. So the refusal names the
candidates and the exact command. Being right is not the same as being usable.
Gaps 3 and 4 — a consumer that is not a machine, and switching as a reaction
rather than a declaration — remain gaps. Half-building them would put a
conditional in the declaration language, which is what ADR 0024 says plainly to
avoid.
Its own context, with its own store and its own credential: a licence is a
different aggregate from anything inventory owns, and it refers to nodes by
name because that is what crossing a context boundary may carry.
novox/hq 08-connectivity §3, built. The mirror of a database grant: there the
consumer supplies a name and receives credentials; here it supplies a target
and receives a name. Nothing new in the vocabulary — a route is a provision
like any other.
One field was missing and it is the one that matters for anything reaching
back: a contribution now carries where the mesh says that machine is. A reverse
proxy is told to send traffic to a consumer and has to open a connection, so
without it every provider implementing a provision would have to know how the
mesh names machines — a convention leaking into every module.
The proxy itself is an example, not part of the control plane: the contract is
the file, not this program. It replaces its table whole rather than merging,
because the file is the whole truth about who has a route and merging would
keep serving a name whose module was unassigned — the stale-route fault
08-connectivity lists as open, reintroduced one level down. A name it does not
serve is refused by saying which it does: a route withdrawn and a name that
never existed are different things.
The invariant novox/hq ADR 0001 records as unowned, and it was measurably
false in HAL: a provision documented as never rotating minted a new password on
every adoption and updated only the provider's row. Consumers on three nodes
held dead credentials for two days while the mesh reported success. Nothing
enumerated who held the old one.
Three things make that impossible here. The holders are a set the mesh can name
— each pair has its own credential, so rotating one consumer touches one role
and the affected list is a query rather than an assumption. Both ends are
pushed by this command rather than a later one, because leaving the sending to
whoever remembered is the fault exactly. And it is all-or-nothing: if any
affected machine cannot be resolved, nothing is sent and the old credential
keeps working, which is a mesh that has not rotated rather than one that has
half-rotated.
The window is stated rather than hidden: a role's password changes on the
provider and the file changes on the consumer, and they cannot be simultaneous.
The provisioner now takes its superuser password from the file the mesh wrote,
which is how the mesh delivers one. Passing it through the environment needed a
person in the middle of the one path that exists so there is not one — and put
a superuser password where `docker inspect` prints it.
The builder hashed the certificate and compared a bare digest against a
fingerprint written as `sha256:` followed by 64 hex characters. It could never
match — and it failed as "this is not the broker this builder was told about",
which is the one thing this check exists to report truthfully. A check that
cries wolf on every correct broker is worse than no check, because the first
thing anybody does is remove it.
The error now prints what was expected beside what arrived, the way the host's
has always done: without both, the message describes a mismatch nobody can
confirm.
And the pin check has its own test, driven against a real TLS handshake — it
accepts the certificate whose fingerprint the mesh wrote and refuses another.
A pin only ever exercised through a live broker is a pin nothing tests.
The credential was a URL and nothing else, so the builder verified the broker
the ordinary way — against public roots. A mesh's broker presents a certificate
of the mesh's own, which is in no trust store anywhere, so the connection could
only ever succeed against a broker somebody else vouches for. It failed at TLS
with an error about an unknown authority rather than about a missing pin, and
the container sat there running: up, credential on disk, connected to nothing.
So the sealed credential now carries the URL and the broker's fingerprint —
the same two facts a node's token carries, for the same reason, delivered out
of band relative to the thing being trusted. The builder pins it: the standard
chain check is replaced rather than removed, and what replaces it is stricter,
accepting one certificate instead of every certificate a public authority
would sign.
A file holding only a URL still works, for a builder somebody runs by hand
against a broker with an ordinary certificate.
A rule nobody derives is a rule somebody keeps in step by hand, and five HAL
manifests carry a `scope:` key that reads as a restriction and restricts
nothing. Both halves are closed here.
Manifests are parsed strictly. An unknown key is refused, which is the
discipline the host's declaration parser has always had; `scope:` survived
because nothing rejected it.
A module says what it listens on and who may reach it, and saying from where is
required — a rule with no source is open, and must say so rather than appear to
restrict something. The mesh gathers every assigned module's ports, widens
where two overlap, names every module that wanted each one, and renders one
nftables file per node. What no module declared is closed.
Three things it deliberately does not do: it writes no forward policy, because
what a machine routes is the container runtime's business and dropping there
stops every container on the node; it never flushes the whole ruleset, only
its own table; and it carries no command to load itself, because the link may
not carry an action. A service declares `restart-on` the file instead, which is
the shape that rule leaves.
Also fixes a fault the lab found: certificateFor asked where every node is
without the catalogue, so nothing resolved, every machine looked like it was on
no private network, and every certificate the mesh was asked for was refused
with a reason that was not true. Asking that question without the catalogue is
now refused rather than answered wrongly.