The other half of ADR 0038, and what 04-ISSUES/028 was actually about.
A module can now avoid colliding with another module; until this it
could not avoid colliding with the mesh itself.
The substrate is not a module. A node raises it from the bundle it
carries before any mesh exists, so the control plane had never heard of
the store, the broker, or its own container — and handed a database
module 5432, which the store already had.
So the machine says. The host records what each resource binds,
distinguishing what it carried from what the mesh sent — a distinction
that already existed so the two never remove each other — and reports
the carried ones. The node states and this context writes, which is the
shape of every message between them.
What the declaration binds, not what is open. A machine's open ports are
a moving target, and assigning around them would mean a port that was
free when it was asked for and taken when it was used.
Replaced whole each time rather than merged: a machine that gave a port
back must be believed about that too, and a set that only grows keeps a
port reserved for something no longer there.
Tested against a real database, and the tests bite — removing the check
hands the module 20000, which the machine had said it holds.
novox/hq ADR 0038. A module cannot choose a port: it is written once and
assigned anywhere, so any number it picks is a guess about a machine it
has never seen. A database module met the mesh's own store on 5432 and
was told, by a container runtime three layers down, that the port was
already allocated.
The number used to appear three times in every module — the rule set,
what a consumer is told, and what the runtime publishes — agreeing only
because one person wrote all three. Now it appears once, in `listens`,
and the other two are derived: the container publishes `20000:5432`, the
consumer is told 20000, and the rule set opens 20000.
An assignment is made once and kept, as a credential is. A port that
moved on every declaration would restart both ends each time and hand a
consumer a number that was true when it was read.
Ports the protocol fixes — mail on 25, submission on 587, DNS on 53 —
say so, and are then claims: one holder per machine, and the second is
refused by name at assignment. That is the mechanism the mesh already
has for what is singular on a machine, pointed at ports.
A mapping written the long way is left exactly as it is. Some things
must be pinned by hand, and quietly overruling somebody who wrote both
halves would be worse than not offering the short form.
Still open, and known: the substrate is not a module, so the mesh has
never heard of its own store and cannot yet assign around it. That is
what 028 will still be about after this.
novox/hq 04-ISSUES/025. Every image reference in every example module
was sixty-four zeros — eighteen of them across five modules. Each
parsed, resolved, and composed into a declaration a host accepts, and
none could ever have started: the machine reaches `docker pull` and
stops. That is why those modules were written and not running, and no
check saw it because every check passed.
The host validates the shape of a reference and nothing more, which is
correct: verifying a digest exists means reaching a registry, and that
is the one thing a host must never have to do. So the last place that
could catch this is the wrong place to try.
The guard therefore sits where a declaration is composed, not where a
manifest is parsed. A file in a repository is allowed to await a pin —
the design already says the manifest in a repository names artifacts
while the manifest the mesh holds names digests, and the bundle works
exactly that way. What must never happen is a placeholder reaching a
machine, and composing is the last moment before one does.
Twelve third-party images resolved to real digests without pulling
anything, which is also the mechanism the open issue needs. Two
discoveries came free: mailu publishes to ghcr rather than Docker Hub,
so seven references named repositories that do not exist at all; and it
renamed roundcube to webmail, so that one would have failed even with
the right registry.
What stays a placeholder is the mesh's own provisioner images, which
genuinely have no digest until built and pushed — the bundle's problem,
legitimately unresolved here. The stand-in consumer now stands in with
a real image rather than an invented one.
I wrote that a test asserts the control plane's placeholder expression
and the host's still agree. None does, and none in this repository could
— a unit test here can only assert what this repository already
believes.
That is precisely the thing this project refuses to tolerate: a stated
rule with no way to check it, which costs more than no rule because
people believe it. Written by me, today, in the same file that closes a
gap of the same kind.
What actually proves it is the lab, and the comment now says so.
A contribution that is not a credential grant — a module offering
something to another on its own machine — has nobody to be identified
to, and was carrying an empty `as`. A field that is always present and
usually empty teaches a reader to ignore it, including when it is not.
novox/hq 04-ISSUES/023. A consumer was given its password, the address,
the port and where its credential lives, and still could not connect —
the user name was invented by the provisioner and recorded nowhere, and
the rest sat in a JSON binding that a program reading KEY=value cannot
use.
Both halves have the same cause: the mesh knew something and did not say
it.
**Who a consumer is, said once.** The provisioner used to derive
mesh_<node>_<module> and that string existed nowhere else — not in the
control plane, not in the binding, and above all not at the consumer,
which has to present it. Now the mesh derives it once and sends it to
both ends, so they agree by construction rather than by two conventions
that were the same on the day they were written. The provisioners refuse
to invent one if the mesh says nothing, because falling back to a name
of their own would create a role the consumer would never guess and
everything would report success.
**Bound values reach the file that needs them.** ${bound:provision:key}
is the symmetric twin of the sealed placeholder, and simpler: these
values are not secret, so the control plane fills them in before sending
and the host gains no field and learns no format. It stays
name-agnostic — at, as and from are true of any provision, and every
other key comes from what the provider said it serves.
The asymmetry it removes was backwards. The secret is the hard case,
because the mesh must not be able to read it, and the secret was the
part that already arrived.
Keycloak and Gitea now produce complete connections, asserted from the
manifests on disk rather than from fixtures: every part filled, no
placeholder surviving as a value, and the password still a hole only the
host can close. Three faults injected, each caught.
A module may not declare an action: the link may not carry a command to
run, and that bound is what limits a compromised control plane to shapes
it cannot turn into arbitrary code (novox/hq ADR 0005). The host
enforces it, correctly and in the right place.
But a module's resources reach a machine over the link, so a manifest
carrying an action was accepted here, stored, resolved, planned and
pushed — and refused on the machine, in the host's log, with nothing
connecting it back to the manifest that caused it.
The rule held. It was just unusable, which is the same shape as the
network shape earlier today: the refusal was right, arrived far from its
cause, and nobody was reading the log.
The refusal names the rule and what to do instead, because "you may not"
with no alternative is where a module author stops.
Found while checking a claim I had written in the coverage document —
that a module cannot declare one. It could; it just could not deliver
it. The document is corrected.
novox/hq 04-ISSUES/022. A credential was keyed by provision, consumer
node and provider node, so "who is asking" was answered by naming a
host. The node this mesh exists to take over runs eight modules against
one database server.
The symptom had two halves and only one was loud. The provider refused,
naming the modules and explaining they would share one credential, which
reads as a decision rather than a limit. The consumer did not refuse: it
resolved cleanly, wrote one module's credential file and left the others
absent — a service that starts and cannot authenticate, with nothing
saying why. That is 021 again on a different axis.
Three modules wanting one database produced one need, carrying whichever
module mentioned it first, because the resolution walk is a work-list
over names. The fan-out now happens in one place, after the walk. The
record path already did this correctly and said why: a consumer here is
a module on a machine. It is the same rule.
Downstream: the secret's key gains the consuming module, the grant file
is named after both halves, needs are matched by provision and module
rather than provision alone, and the provisioners name the role and the
access key after the module. The refusal in ContributionsTo is gone
because there is nothing left to refuse.
Worth stating plainly: without that refusal, gitea's login would have
opened keycloak's database. From the provisioner's side it created
exactly what it was asked to create.
Existing secrets are discarded rather than backfilled. They cannot say
which module they were for, and a secret is remade and delivered to both
ends on the next push — so this costs one rotation and invents nothing.
Also guards the role name against PostgreSQL's 63-byte truncation, which
is a notice rather than an error and would reintroduce exactly this
collision at a length nobody tests.
Three faults injected — the fan-out removed, needs matched by name
alone, the grant file named after the machine — each caught.
The gap that stopped keycloak and gitea from starting. A granted
credential arrives as a file whose entire content is the password, which
is what a program reading a password file wants — and most programs do
not read one. They read KEY=value, or a JSON document with the token at
an attribute inside it. A module in that position could be handed the
bare value or nothing, and both are useless.
The host has been able to do this all along: content with ${secret:name}
in it, sealed values beside it, substitution on the machine, which is
the only place both halves exist. Nothing filled the values in, so the
hole could be written and never closed and the host refused the file.
That refusal was correct and the feature was unreachable.
A module reaches its own secrets and the credentials it was granted —
both things it wrote in its own manifest — and nothing else. Naming
another module's is refused: two modules on one machine are as separate
as two on different machines, and letting one read the other's
credential by guessing a name would end that to save writing a file.
Filling runs after settings, which is the whole reason it sits where it
does. A setting is how a placeholder gets into a JSON document in the
first place — the desktop client that reads its token from an attribute,
not an environment variable. Before the merge that file's content is
"{}" and asks for nothing.
Tested through Declaration rather than through the helper. Three times
in this repository a test asserted on a helper while the code calling it
was wrong, and each time the injected fault stayed silent. Three faults
injected here — the call removed, the call moved before settings, and
the module boundary widened — each caught by the test meant for it.
novox/hq 04-ISSUES/021. Two modules where one provided what the other
required, on one node, resolved cleanly with zero needs: no credential
was made, the consumer's secret file was never written, and whatever
read it would fail somewhere else entirely. Nothing was refused and
nothing was reported.
The world a node resolves against is every OTHER node, so a provider on
the same machine never became a Needed, and the credential loop walks
Needs. Every step reasonable, the sum a silent gap.
It survived because everything proven until now was cross-machine —
the interesting case for a mesh and the rare one in practice. The first
module to want a database on its own machine was the first real one.
The assumption underneath was that a local consumer needs no credential,
which holds for a process reaching a unix socket where the system can
vouch for the caller. It does not hold for containers, which is how
nearly everything here runs: the consumer reaches the provider over TCP
from its own container and the database asks for a password exactly as
it would from another machine. **The machine stops being a trust
boundary once both ends are containers.**
A brokered provision answered here is now a need naming this node, and
carries what the provider serves — which a local provider never
contributes through the world. A name nothing grants is unchanged: a
shell answered here is answered, and nothing more is owed. Both
directions tested, both injections bite.
Work breakdown 1.2. Two sessions run on the control-plane node — the
node's own and the mesh's (novox/hq ADR 0026) — so a machine stopped
being a usable answer to "whose licence is this".
14-model-access.md called per-module-per-machine "a step toward it and
not it", and that is true of a worker: many run on one machine from one
module, so the pair cannot name them apart. It is not true of a session.
The two sessions are two modules — the same mechanism started in
different context roots, and a context root is what a module delivers —
so (node, module) tells them apart and nothing needed adding.
Checked rather than argued: different licences on one machine, each with
its own key, and a session on no licence is not handed the other's.
The third test exists because a fault injection stayed silent. The first
two put the sessions on different licences, so the licence alone
disambiguates and the module argument is never load-bearing — removing
it from the query changed nothing and everything still passed. Two
sessions on the SAME licence is the case that needs the pair to be the
identity: releasing one must leave the other, and a machine-shaped
answer takes both.
A migration refused when the record and the files disagree is the
property that makes a schema trustworthy months later. The test asserted
it without naming the decision, so an audit of which decisions are
defended could not see it. novox/hq ADR 0017.
Provisions were named after roles: provides "database", requires
"database". Nothing distinguished engines, so a module written against
PostgreSQL could be matched to a provider of SQL Server, resolve as
satisfied, deploy, and fail on its first query — with nothing
connecting that error back to a match made elsewhere by something that
believed it had done its job.
The failure is in the direction that hides. Refusing on ambiguity
exists precisely so this does not happen, and the generic name walked
around it: with one provider of each name nothing is ambiguous, so
nothing is asked.
How it got in: every resolver test had exactly one provider per name,
so no mismatch was expressible and none was caught. The fixtures agreed
with the design — the same fault as the imagined test output in
04-ISSUES/005, at the level of a name.
Refused rather than documented, because the old naming *was* the
documented convention. Providing database/db/sql/sql-database is now a
parse error naming what to write instead.
The rule is about coupling, not specificity everywhere: route and
resolver stay role-named, because a consumer genuinely cannot tell
which proxy answered. novox/hq ADR 0027.
It sat beside `secrets` — where a *provision's* credential lands on a consumer.
Both were name-to-path, both held something secret, and the names
distinguished them not at all. Reaching for the wrong one parsed cleanly and
failed somewhere else entirely, which is the shape of fault this whole design
exists to prevent, sitting in the manifest format.
The axis that separates them is not how secret they are — both are — but
whose. `secrets` is keyed by the provision it is for and belongs to a
relationship with another machine. `own-secrets` is keyed by a name the module
chose and belongs to nobody else.
A manifest using the old name is told the new one rather than refused with
"unknown field": whoever wrote it knew what they meant, and the mesh knows what
it is called now. An invented key is still refused as one rather than guessed
at.
Found by auditing the 19 manifest fields for whether any could be mistaken for
another. This was the only pair that could — and while checking it, a second
instance of the same collision turned up one layer down: `Manifest.Needs` and
`Resolution.Needs` were different concepts sharing a name in Go. The rename
separates those too.
The host does not sort, so the order written here is the order a machine
applies. Selection walks outward from what was assigned, which puts a consumer
before the thing it pulled in — and a service that reads a file another module
writes then starts before the file exists.
It fails, and the next reconcile fixes it. That is the worst shape a fault can
take: what gets remembered is that it works, and nobody looks again. It is
04-ISSUES/013 one level up from where that was found — there, the mesh's own
computed files came after a module's resources; here, a whole module comes
after the one that needed it.
Nothing had hit it because no module until now both required something with
resources of its own and had a resource depending on it. Writing the resolver
module was what made it reachable, and it would have shown up as dnsmasq
failing once on every fresh machine and working ever after.
Unrelated modules keep the order selection gave them — assigned first, then
what they pulled in. That order is meaningful, and reshuffling it would make
every declaration's diff unreadable for no gain.
Two modules requiring each other are both applied rather than refused: a cycle
is not a machine that cannot work, and refusing would make a cooperating pair
impossible to assign.
Three manifests and the rule that keeps them apart. Serving and asking are
genuinely different roles, and systemd-resolved can only do the second — it
cannot answer a wildcard, it routes the mesh's suffix to something that can. A
module that treated them as one role could not work, which is the mistake worth
naming rather than discovering.
So `the-dns-port` and `the-resolver-configuration` are two claims. A machine
gets one of each, and two of either is refused by the mesh rather than fought
over on the machine — which is what ADR 0009's table meant by listing resolvers
beside the seat and pid 1. That table names the resource `/etc/resolv.conf`,
which is what it is; a claim is a name in the catalogue's own form, and the
catalogue refuses the path as one.
Neither module knows anything about the machine it is on, which is what lets
them be static manifests: they name `mesh0` and `127.0.0.54`, both chosen by
the mesh, rather than an address only that machine has. Not 127.0.0.1 and not
127.0.0.53 — taking either would be a module claiming something it did not say
it claims.
A service can now reflect a file another module put on the machine, written
`<module>.<id>`. The resolver has to restart when the mesh rewrites the names;
without it, it would serve the names it started with for ever, with every
machine that joined afterwards unreachable and every check passing.
Services are named under the machine they run on — postgres.novox.internal,
plex.ace.internal. The first label is the service and the rest is the node, so
what has to resolve is anything under a node's name. What routes it once it
arrives is a proxy's concern and stays separate.
A hosts file cannot do that. It answers exact names, and a wildcard there would
mean writing down every service in advance — which is the enumeration the
arrangement exists to avoid. novox/hq 08-connectivity named this exact case as
the trigger for needing a resolver rather than a file, and it is the first
thing to meet it.
The mesh writes the data and runs no daemon. A resolver is third-party
software, and third-party software runs on the mesh rather than being of it
(ADR 0001): the mesh has no business shipping one, choosing which one, or
knowing its configuration language. What only the mesh can know is which
machines exist and where they are. A module that runs a resolver requires what
this provides and reads one file, so swapping the daemon changes that module
and nothing here.
Separate from names rather than part of them: a machine with no container
runtime can still have a hosts file, and folding them together would take exact
names away from a machine that cannot run a daemon in order to give it a
wildcard it cannot use either.
A machine with no address is left out. A wildcard pointing at nothing is worse
than no wildcard — every name under it resolves and then hangs, where an
unresolvable name fails at once and says which name it was.
Internal names are written to the machine's hosts file, which serves the
machine and not what the machine runs: a container gets its own hosts file
holding only its own hostname. So every name the mesh wrote was invisible to
the majority of things that need one — and on the machine it always worked,
which is exactly what made it easy to miss.
It was hit for real in the lab, and worked around by resolving the address on
the machine and passing it in. That workaround is now removed, and its absence
is the assertion.
A file rather than a resolver, which is the decision the mesh already made
about names and this extends rather than overturns: it works on every runtime,
needs no package and has no failure mode of its own. The stated trigger for a
resolver — names that are not one-per-node, service names, wildcards — is
still not met.
Given by the mesh, not chosen by a module: a module that listed the machines
would go stale the day one joins, and one that did not would be a module whose
containers cannot reach anything by name. A container that named its own keeps
them and gets the mesh's beside them.
Only containers, and not the ones on the machine's own network: a runtime
refuses to write a hosts file for those, and a file or a service given the
field is a declaration the host refuses outright — so getting it wrong breaks
the whole machine for something that was never about names.
The machine that most needed a firewall was the one that could not have one. A
hub is dialled by every node at other sites and needs its port open; a machine
that is not a hub dials out and needs nothing open. They are the same module,
and `listens` in a manifest is one answer for every machine that runs it — so
the machine a static answer gets wrong is the one facing the public internet.
A generator can now say what it opens, in a second interface rather than a
method on every generator: most have nothing to say here, and requiring an
empty method of each would be a cost paid everywhere for one caller.
The port is the one in the endpoint, which is where the interface takes its
ListenPort from. One source, so a rule set cannot open a port the interface is
not on. Open to everywhere and deliberately: a node at another site is not on
the private network until this port lets it on, so restricting it to the mesh
would be a rule that can never be satisfied by the thing it exists for.
And a generator that cannot say is refused rather than read as silence. Closing
a port on the evidence of a failure to look is how a machine is severed by a
fault somewhere else — and the machine it would sever is the hub, whose only
route to being fixed is the network it just closed.
It meant "failed or refused". So a machine that applied cleanly and whose
declaration has since changed was not behind — and novox/hq ADR 0010's
question, did my change go out?, was answerable exactly for the machines that
broke. For every machine that worked, the answer was silence whether the change
had gone out or not, which is the thing replacing a pipeline was supposed not
to cost.
The mesh now records a digest of what it last sent each machine. A digest
rather than the declaration: it can compute what a machine should be at any
moment, and keeping a copy would be a second account of it able to disagree
with the first. What cannot be recomputed is what was actually sent.
Recorded after the send, not before — a digest kept for something that failed
to send would make the machine look current for a declaration it never
received.
Never told stays separate from out of date. The remedy is the same push and the
situations are not alike: nobody has ever asked that machine to be anything.
And a machine the mesh could not work out is not reported as waiting, because
saying so would invent a comparison — that is `plan`'s answer to give.
`status` says it and `push --behind` sends it, or the flag would know something
the person reading the status does not.
novox/hq ADR 0009: a capability's presence gates an assignment and its detail
carries a value — seat: card1-DP-1, an architecture, an amount of memory. So
'can this run here' and 'what should it be configured as' are one fact read two
ways, and the mesh was keeping the first read and discarding the second.
The reason an absent capability is absent went the same way, which is the case
a person most needs: 'this machine has no container runtime' is the answer and
'docker is not installed' is why, and only the machine knows why.
`node show` says it back. Never reported and reported nothing stay different
things there — one machine has not run the host, the other ran it and can do
nothing, and those send a person to different places.
The pass that answers *what does this node offer* takes a failed resolution to
mean it learned nothing about that node. So refusing an unanswerable
requirement there made the machine disappear — and every other machine was then
told, wrongly, that the two of them shared no private network.
A wrong answer about a machine nobody asked about, caused by a fault on a
third. The lab found it: one module needing a licence that had not been added
yet made two unrelated machines look disconnected.
The second pass still refuses it, where the question is actually being asked.
Its own store, its own test database, the same shape every other context has.
Five properties: a key with nobody to seal it to is refused rather than kept
readably; a key is sealed once per holder and the blobs differ because they are
sealed to different machines; a holder recorded afterwards has none and the
existing ones keep theirs; releasing a consumer takes its key; and a licence
nobody recorded is refused by name.
The last was the only one whose message mattered and whose message was not
checked — the database's own foreign-key error is true and mentions a
constraint, which sends somebody to read a schema instead of typing the name
they meant.
Partial sealing now says how far it got. The person holding the key is the only
one who can finish, and running it again knowing what it will do is different
from running it hoping.
novox/hq ADR 0024, gaps 1 and 2. The user's stated requirement, and the first
thing here that no machine can answer: a hosted model is on nobody's node and
is reached over the public internet, so the rule that refuses two ends sharing
no private network must not apply to it.
A licence is a named thing and the name is the operator's — *the personal
account*, *the organisation's* — because the whole point is saying which one a
given consumer uses, and an anonymous credential hanging off a provider cannot
be said. Many to many, so deliberately not a claim: two machines sharing an
account is ordinary rather than a collision.
Gap 2 is the missing verb, *accept*: take a value somebody supplied, seal it to
each holder, discard the plaintext. With the consequence stated rather than
hidden — a holder recorded after the key was supplied has no key and the mesh
cannot make one, so it is refused by name with the remedy, not silently handed
an empty file.
Refusal is felt, as the record warns: a mesh holding three ways to reach a
model refuses every consumer that has not chosen. So the refusal names the
candidates and the exact command. Being right is not the same as being usable.
Gaps 3 and 4 — a consumer that is not a machine, and switching as a reaction
rather than a declaration — remain gaps. Half-building them would put a
conditional in the declaration language, which is what ADR 0024 says plainly to
avoid.
Its own context, with its own store and its own credential: a licence is a
different aggregate from anything inventory owns, and it refers to nodes by
name because that is what crossing a context boundary may carry.
novox/hq 08-connectivity §3, built. The mirror of a database grant: there the
consumer supplies a name and receives credentials; here it supplies a target
and receives a name. Nothing new in the vocabulary — a route is a provision
like any other.
One field was missing and it is the one that matters for anything reaching
back: a contribution now carries where the mesh says that machine is. A reverse
proxy is told to send traffic to a consumer and has to open a connection, so
without it every provider implementing a provision would have to know how the
mesh names machines — a convention leaking into every module.
The proxy itself is an example, not part of the control plane: the contract is
the file, not this program. It replaces its table whole rather than merging,
because the file is the whole truth about who has a route and merging would
keep serving a name whose module was unassigned — the stale-route fault
08-connectivity lists as open, reintroduced one level down. A name it does not
serve is refused by saying which it does: a route withdrawn and a name that
never existed are different things.
The invariant novox/hq ADR 0001 records as unowned, and it was measurably
false in HAL: a provision documented as never rotating minted a new password on
every adoption and updated only the provider's row. Consumers on three nodes
held dead credentials for two days while the mesh reported success. Nothing
enumerated who held the old one.
Three things make that impossible here. The holders are a set the mesh can name
— each pair has its own credential, so rotating one consumer touches one role
and the affected list is a query rather than an assumption. Both ends are
pushed by this command rather than a later one, because leaving the sending to
whoever remembered is the fault exactly. And it is all-or-nothing: if any
affected machine cannot be resolved, nothing is sent and the old credential
keeps working, which is a mesh that has not rotated rather than one that has
half-rotated.
The window is stated rather than hidden: a role's password changes on the
provider and the file changes on the consumer, and they cannot be simultaneous.
The provisioner now takes its superuser password from the file the mesh wrote,
which is how the mesh delivers one. Passing it through the environment needed a
person in the middle of the one path that exists so there is not one — and put
a superuser password where `docker inspect` prints it.
A binding was skipped when the provider turned out to be on the same node,
reasoning that a file saying "it is on this node" is a fact nobody needs. That
is right about the location and wrong about everything beside it: a binding
also carries what the provider said a consumer must know, which is the port,
and a consumer cannot invent that.
A build machine sharing a node with the registry it pushes to sat in a loop
saying it could not read its own binding. Nothing was wrong with the machine,
the module, the credential or the provision — the file was never written, and
the absence looked exactly like a mistake in the module.
The original intent is kept where it was right: a provision whose provider said
nothing a consumer must know is still not written. A shell is answered here and
there is nothing to say about it. A registry is answered here and the port is
still unguessable.
The address is this machine's name on the private network, or loopback when it
has none — a machine off the network still reaches itself, and a name nothing
resolves is worse than an address that always works.
The host does not sort — order is stated (novox/hq ADR 0005) — so the order the
mesh writes down is the order a machine applies. Certificates, credentials,
bound files and the rule set were appended after a module's own resources, so a
service or container that depends on one was applied before it existed.
It failed and the next reconcile fixed it, which is why nothing caught it. A
fault that repairs itself on the second attempt is worse than one that does
not: what gets remembered is that it works.
Nothing the mesh computes depends on a module's resources, so putting all of it
first is unconditionally right. Merged after the computed-resources branch,
which replaces a module's resources wholesale and would otherwise discard them.
nftables matches ip and ip6 separately and one set holding both is a syntax
error, so the file would not load: the service reports a configuration fault
and the machine filters nothing. Also a make target for the builder image,
which the lab now stocks.
Two kinds live in module_secret and they behaved identically, which is right
for one of them. A made secret is the mesh's: when a node regenerates its
sealing key the mesh makes another and nothing is lost, because nothing else
ever knew the old one.
An accepted secret is not. A broker account's password exists because the
broker was told about it. Regenerating one puts 32 random bytes where a working
credential was — and the machine applies it, reports success, and the program
reading it fails to authenticate somewhere else entirely, with the mesh
insisting the secret was delivered, which it was.
The row now records where the value came from, and a rejoined machine asking
for an accepted one is refused with the remedy named: issue it again. No amount
of pushing produces a password the broker has never heard of.
Found while making the builder a module, which is the first thing to hold one.
A rule nobody derives is a rule somebody keeps in step by hand, and five HAL
manifests carry a `scope:` key that reads as a restriction and restricts
nothing. Both halves are closed here.
Manifests are parsed strictly. An unknown key is refused, which is the
discipline the host's declaration parser has always had; `scope:` survived
because nothing rejected it.
A module says what it listens on and who may reach it, and saying from where is
required — a rule with no source is open, and must say so rather than appear to
restrict something. The mesh gathers every assigned module's ports, widens
where two overlap, names every module that wanted each one, and renders one
nftables file per node. What no module declared is closed.
Three things it deliberately does not do: it writes no forward policy, because
what a machine routes is the container runtime's business and dropping there
stops every container on the node; it never flushes the whole ruleset, only
its own table; and it carries no command to load itself, because the link may
not carry an action. A service declares `restart-on` the file instead, which is
the shape that rule leaves.
Also fixes a fault the lab found: certificateFor asked where every node is
without the catalogue, so nothing resolved, every machine looked like it was on
no private network, and every certificate the mesh was asked for was refused
with a reason that was not true. Asking that question without the catalogue is
now refused rather than answered wrongly.
08-connectivity keeps two authorities apart on purpose: a public one for
names the outside world reaches, and the mesh's own for names only the
mesh knows. Nothing implemented the second, so anything between machines
was plaintext or trust-on-first-use — which the design refuses everywhere
else.
A node now generates a fourth key at enrolment and reports the public
half. A fourth, because a key used for two purposes is one rotation away
from breaking the other: the identity key signs messages to the mesh and
would do for TLS, and reusing it would mean rotating a node's identity
every time its certificate is replaced.
**Nothing secret travels and nothing is sealed.** A certificate authority
says "this name belongs to the holder of this key", so the mesh signs a
public half it cannot use, and the certificate it issues is public. A
module asks for one and is given the certificate and, if it wants,
the mesh's own — the private key is a path to a file the machine already
has, the same arrangement the private network's key uses.
Asserted by verifying rather than inspecting, because a certificate that
parses and does not chain fails at the moment something connects:
- what the mesh issues verifies against the mesh, for the name asked for
- the name is in the subject alternative names, since a certificate
carrying it only in the common name is refused by every modern client
- it certifies the key the node generated and no other
- another mesh's certificate does not verify, which is the whole point of
two authorities being separate
- the authority cannot sign another authority — one that could is one
that can be delegated without anybody deciding to
- two control planes starting together agree on one authority, or a mesh
has certificates half its machines refuse
Certificates last ten years, which is a choice: a short life needs
something to renew it, and a renewal that fails silently is a mesh that
stops trusting itself on a date nobody wrote down. What makes one
replaceable is that the mesh reissues on demand, not that it expires.
A board reads through interfaces and holds nothing. Everything it needs
is already answered — as text, for people, which is not something a page
can read.
`--json` rather than a serving API, because nothing needs one yet:
whatever serves a board runs the command, and the constraint holds either
way — the board never touches a context's store. An API is the larger
thing and should wait until something asks for it.
Both forms are gathered from the same reads before either says anything,
so they answer the same questions rather than being two implementations
that can drift. That was not true of the first version: the JSON printed
after the text, because the branch was too late.
Four properties, each asserted and each confirmed to fail when removed:
- refused and failed stay distinct all the way out. They are fixed in
different places, so one word for both sends half a page's readers to
the wrong one — and how much DID apply is carried, since "three of
eight" and "none of eight" are different machines
- a machine that never spoke carries no time at all, rather than a zero
one that any page would format as a date in 1970
- nothing is null. A page distinguishing "no machines are wrong" from
"this field is missing" has to handle both, and null is the one that
gets forgotten
- no field is named like a secret. Everything here comes from records
that hold no readable one, but a shape a page is built against is
exactly where one would eventually be added for convenience
Found by testing removal, which is the half nobody tests.
A grant was emitted for every secret the mesh held, whether or not the
machine still asked for it. So a consumer that was unassigned kept
appearing in its provider's manifest — and the provisioner's rule about
removing what nobody asks for can only fire if the mesh stops asking. The
login would have stayed live for ever, and nothing would have said so.
Skipped where the declaration is built rather than where grants are
gathered, so the rule holds whoever gathers them. No credential file is
written for a withdrawn consumer either, or the provisioner would find a
file its manifest does not mention and have to guess what that means.
The secret itself is deliberately kept. It is sealed and unusable to the
mesh, and a machine that comes back gets what it had — what withdraws the
login is the manifest, which is the thing that reconciles.
A module usually runs software somebody else built: a database module
ships configuration and a provisioner and does not build a database. It
could name the upstream reference directly, and then every machine needs
a route to a public registry and the reference is a tag somebody else can
move — which is what pinning exists to prevent.
So an artifact may be `upstream`: pulled by the reference the module
names, pushed into the mesh's own registry, and pinned by the digest that
registry assigns. This is what the bootstrap already does by hand; it is
now something a module can say.
Refused: an upstream reference with no tag or digest, because what gets
mirrored would be whatever `latest` means today and a module pinned to
that is not pinned. And the rule that a build reads only its own
repository does not apply to it — applying it anyway refused every
reference with a registry host in it, which the test caught.
Written by trying to write a real postgres module and finding it could
not be said. It can now: two directories, two containers pinned by
digest, a superuser password sealed to the machine, and the grants
manifest — six resources from one assignment, all accepted by the host's
own parser.
That exercise also found my manifest wrong rather than the host: a
container declared `restart-on`, which is a service field, and the host
refused it by name. It is right to. A container whose own definition
changes is recreated, and a file it mounts is read by the process inside,
which is that image's business.
Two things, both found by trying to write a real postgres module and
discovering it could not be said.
A database has a superuser password, a broker an administrator, a
registry an account. None of them is *for* anybody — they are not the
credential a consumer is given, and the mechanism that hands those out
has a consumer in the middle of it. So a module may declare what it needs
and where to put it, and the mesh generates one per node, seals it, and
reads it no more than it reads any other.
Per node, deliberately: a module running on three machines has three
passwords. One in the manifest instead would put the same secret on every
machine that ever runs it, in a file anybody can read, for ever. Made
once and kept, or a running database would be handed a password it was
not started with; remade when the machine's sealing key changes, like
everything else sealed here.
A need declared and not made is refused rather than skipped, because a
module whose own credential is silently absent starts, fails to
authenticate, and the reason is three layers from the machine reporting
it.
And the provisioner can watch. That is what lets it be a module rather
than a binary somebody places: run once, it needs invoking after every
declaration by a timer or a unit wired to a file; watching, it is an
ordinary long-running service the host already supervises. It polls
rather than watching the filesystem, because the host writes atomically —
the file is replaced, so a watch on the path stops seeing anything after
the first replacement, and a watcher that silently stops working is worse
than a poll. Credentials are compared by digest and never held: this runs
for as long as the machine is up.
A node reports back after applying a declaration: it worked, some of it
failed, or the whole thing was refused. A refusal or a failure moved
last_seen and the reason went to a log line — so "which machine is not
doing what it was told" had no answer the next morning, which is the
question a mesh exists to answer.
Refused and failed are kept as different things, because they are
different situations with different remedies: refused means the machine
is exactly as it was and what is wrong is in what was sent; failed means
it is in a state nobody declared and what is wrong is on the machine. One
word for both would make the record say less than the node did.
One row per node, replaced. The question is the machine's current state —
"this failed an hour ago and then succeeded" is not a machine anybody
needs to look at, and a table of every report would bury the ones that
matter under the ones that do not.
`status` now answers three questions in the order somebody asks them: is
anything broken, is anything not answering, is anything out of date. The
first has consequences now, the third is a plan for later, and a status
leading with the third would bury the first. A machine that has never
spoken is reported as quiet rather than as broken — new, switched off and
unreachable are not the same as tried and could not.
The mapping from a report to an outcome had no test at all, which the
injection caught: it is the code deciding which of those situations a
machine is in. It has four now, including that a partial report never
becomes the account of what the machine holds — the fault that destroyed
a substrate once.
The builder was documented as holding its own broker credential and
nothing else, and nothing issued one — so in practice it used whatever it
was handed, which was the broker's administrative account. A program
documented as holding its own credential and given somebody else's is
worse than one with no story at all.
`builder issue <name>` creates an account that may read the build queue
and write to the mesh exchange. Not a node account: a build machine is
not a node, and a node's queue carries its declarations.
Two faults found by running it, both about the answer path:
- the reply queue was left for the broker to name, and the account was
scoped to `amq.gen-*` — one broker's convention. The builder built,
could not answer, and the connection closed. Reply queues are named
here now, deterministically.
- the answer then went via the DEFAULT exchange, where permission is
granted per exchange rather than per queue. A builder allowed to use it
could publish into any node's queue, which is the privilege a build
machine most obviously should not have. Answers go through the mesh
exchange, which it already may use, and an asker binds its reply queue
to the same key and filters by correlation.
Verified against a real broker: a builder cannot consume a node's queue
and cannot publish to the default exchange. That check nearly reported
the opposite — an unconfirmed publish is asynchronous, so the refusal
arrives as a channel close afterwards and a naive test sees success. With
publisher confirms it is immediate. A negative security assertion made
against an asynchronous call is not an assertion.
Redelivery was observed working while fixing this: builders that died
before answering left their work on the queue, and the next builder did
all of it.
Also: the queue and exchange names exist in both `broker` and `link`,
because `link` imports `broker`. A test in an external package keeps them
agreeing — a builder scoped to a queue nothing publishes to takes no work
and says nothing about why.
A build result was answered to whoever asked and kept nowhere. So "when
did this last build", "why did it fail" and "which machine built what is
running" had no answer, and a build nobody was waiting for was reported
into the void — which is the same as not reporting it.
Failures are recorded too, and that is the point rather than a detail: a
failed build that leaves no trace is indistinguishable from one nobody
asked for, and the difference is the whole of whether somebody should be
looking at something. A build that never learned what it was building
keeps the repository, because that is what a person goes and looks at.
Recording is idempotent on the correlation id, because a result can
arrive twice — as the answer to whoever asked, and on the exchange when
nobody was. Two rows would show one build as two, and which is real is
not answerable afterwards.
The serving control plane now binds `built` as well, so results from
builds it did not ask for are kept. It refuses them loudly when it has
nowhere to put them rather than dropping them, so the broker's own
counters show something arriving that nothing handles.
`builds [<module>]` reads it: what happened lately across the mesh, or
what has happened to one module — the first asked after something goes
wrong, the second when deciding whether to trust something.
What was published is kept with the build, so a digest traces back to
what made it without holding the manifest twice in a place that can
disagree with the first.
The last of the four gaps ADR 0024 names. Everything the mesh handles
today it generated itself, sealed to both ends, and discarded. An API key
for a hosted service comes from a person, and carrying it needs a verb
the mesh did not have.
Accept seals it on the way in and keeps no plaintext — the same storage
and the same property as a generated one, only a different origin. That
is the whole difference from the arrangement being replaced, where an
operator-supplied key sits in a column the control plane can read, which
makes a copy of the database a copy of every account the mesh touches.
The consequence is deliberate: the mesh cannot show it back. Somebody who
loses the key gets a new one from wherever it came from. There is no
reveal and there cannot be one, because a mesh that can reveal a secret
is a mesh that holds it — asserted as a test, because it is a property
somebody will eventually ask to break.
An empty value is refused. A credential that exists, authenticates
nowhere and looks exactly like a working one is the failure this whole
mechanism is arranged to prevent.
A build is work, not state. Everything else the control plane sends a
node is a declaration — this is what you should be — reconciled forever.
A build happens once and is finished. Putting it in a declaration would
mean rebuilding on every reconcile, or a declaration carrying "and I
already did this", which is state about an event rather than about a
machine.
So it travels on its own queue and the answer comes back correlated. One
queue, so several build machines share the work and each request is done
exactly once — which a per-machine routing key would not give.
mesh-builder is the program a build machine runs. Not the control plane,
which must not run commands on a machine; not the host, which would then
need a container runtime and git everywhere to do something almost no
machine will ever do. It holds its own broker credential and nothing
else.
Three properties that are decisions:
- a request is acknowledged only once the answer is away, so a builder
that dies mid-build leaves the work for another machine rather than
losing it with nobody ever hearing why
- one build at a time. Five at once against one runtime finishes all five
slower than it would have finished the first, and the queue is what
shares work between machines
- a failure is a RESULT. A build that fails silently is
indistinguishable from a builder that is not running, and those want
different responses
And `module list` is a catalogue: what exists, at which version, built
from which commit or handed over by hand or shipped with the control
plane, whether it is behind its source, and which machines run it. All of
that was recorded from the first build and none of it was shown, so "is
this current?" could only be answered by reading the database.
Proven against a real broker, registry and store: the mesh asked, a
builder consumed, built, published, answered; the manifest was recorded
with its commit; the source moved and the catalogue said "behind";
rebuilding caught it up with a new digest because the content changed.
One store, and it is the registry the bootstrap already pulls from. An
OCI registry is a content-addressed blob store that also understands
images: PUT a blob and it is retrievable at /v2/<name>/blobs/sha256:… for
ever, by digest. An archive is a content-addressed blob.
A second store beside it was considered and is the right answer for
objects that are mutable, need per-reader access, or are not build output
— somebody's uploads, a backup, a thing with a lifecycle. None of that
describes a digest-pinned archive, and running a second service to hold
one kind of immutable blob is two things to run, two to back up, and two
ways for an artifact to be missing. Overturnable by reading: the manifest
carries a URL and a digest, and neither says what served it.
`build <repository>` clones, reads module.json, builds what it declares,
publishes, and records the manifest with the commit it came from. It is a
command rather than something the control plane does on its own, because
building runs things on a machine and what the control plane may send a
machine is bounded by the declaration language. This is the shape the
builder module takes when it is given work over the broker.
Proven end to end on a real repository and a real registry: a shell
module with a package, a user and a dotfile archive built, published,
fetched back at the digest it declared, rebuilt to the same digest, and
its manifest accepted by the host's own parser — including `user` and
`archive`, which did not exist this morning.
A tag is never accepted as a pin, and a blob already stored is not sent
again — it is named by its content, so re-uploading asks the registry to
store what it already has under the name it already has.
It runs on a node, not in the control plane. Building needs a container
runtime and a working tree, and the control plane deliberately cannot run
commands on a machine — what it may send is bounded by the declaration
language, and "run this build" is not in it. So the builder is something
a node runs as a module, given work over the broker like anything else.
The alternative, the control plane holding a docker socket, would make it
the one component that can do anything anywhere, which is the property
the whole design is arranged to avoid.
A module repository has one file at its root, module.json, saying what it
is and what it builds. A convention somebody can look for beats a setting
somebody has to find.
Properties that are decisions rather than details:
- a fresh clone every time. A build reusing a working tree can succeed
because of something a previous build left behind, and that is a build
nobody can reproduce.
- archives are packed deterministically — sorted, and carrying no
timestamps, uid, gid or original names. Two builds of one commit must
produce one digest, or nothing downstream can tell "this changed" from
"this was built again", and every rebuild looks like a change to every
machine holding it.
- nothing is published until everything is built. Half a module in the
store under a digest the mesh never records is reachable,
unreferenced, and indistinguishable from something in use.
The reproducibility test was passing for the wrong reason: both builds
landed in the same second, so a packer carrying timestamps would still
have agreed. It now stamps the two trees a year apart, and a timestamp in
the header breaks it.
One line is honest about not being independently tested: the sort before
packing is belt and braces over filepath.Walk's documented lexical order,
and no injection can distinguish it.
document
The manifest in a repository names artifacts; the manifest the mesh holds
names digests. Keeping them the same file would mean a repository
carrying a digest — wrong the moment anybody edits anything, and pinning
a value nobody could have checked.
So a resource says `"artifact": "server"`, and resolving a build rewrites
it to the image reference or the archive's source and digest, removing
the build-time word entirely. The host has never heard of an artifact and
its strict decoder would refuse one, at the worst moment.
A module that builds nothing is ordinary and needs no build section —
most of what a person installs is configuration, and a field that exists
to be left blank is a field nobody fills in correctly.
Refusals worth having:
- an artifact declared and not produced blames THE BUILD, not the
resource. Both are failures and the remedies are in different places;
telling somebody to fix the wrong one costs an afternoon. Found by
injection: the first version's message could not be told apart from
the resource-level one, so the check was not actually tested.
- a build reads its own repository and nothing else. An input path
leaving it makes what gets built depend on whatever happens to be on
the machine building it.
- two artifacts with one name, because a resource naming it could mean
either.
resolve.go had grown to 796 lines doing four jobs: working out what a
machine should run, applying settings, collecting contributions, and
placing credentials. They answer different questions — the first is
"what", the rest are "what does that look like as resources" — and one
file doing both is how a thing starts becoming the kernel everything
imports.
Prompted by looking at why HAL's shared library became unmaintainable.
Measured while here, and the shape is the inverse of that one: the large
packages import nothing internal, and only inventory and link compose. A
change to module resolution cannot reach connectivity, because
connectivity does not import it.
Found by raising a mesh end to end. The broker account a joining node
authenticates as is named after the node, and exists before that machine
has been told anything — so the node has to know its name before the mesh
can tell it. Without it, enrolment fails at the broker with an empty
username, which says nothing about why.
Not a secret, and the issuer already knows it. The wire-format test now
covers it, so a rename on either side fails in both repositories rather
than at enrolment on a real machine.
The enrolment request is a struct in each repository. A node now reports
a third key — the one its secrets are sealed to — and that wiring had
unit tests on each side and had never been run across the join. A field
renamed on one side fails silently: enrolment succeeds, the key is
absent, and the node looks joined until the first thing sealed to it
cannot be opened, by which point nobody is looking at enrolment.
So the host's suite writes a real request and this one reads it, the same
way the declaration check already runs in the other direction. Both skip
with a reason when the neighbour is not checked out.
It does more than compare shapes: it seals something to the key that
arrived and opens it with the private half the host kept. Confirmed to
fail three ways — a renamed field, a value that is not a key, and a key
that is present, correctly named and simply somebody else's. Only the
last needs the sealing step, and it is the one a shape check would pass.
Also `inventory.ForTest`, because the check lives beside the link and a
second copy of the throwaway-database helper would be a second thing to
keep true.
Contributions were node-local, so a mesh-scoped provider — the one case
that most needs them — never heard from its consumers. A database was
given a password and no idea what to create it for.
Cross-node consumers now reach the provider's `receives` file, merged in
with the ones on its own machine: from the provider's side they are the
same thing, and a provider that had to read two lists would read one of
them. Each names the file its credential is in rather than carrying it,
because the mesh discarded the value and could not put it there. The
readable half therefore stays readable.
And examples/postgres-provisioner, which is the last step: it reads what
the host wrote and makes PostgreSQL accept it. Explicitly not part of the
control plane — the control plane decides and never touches a machine.
This runs on the machine and touches it, and a real one ships with the
module that ships PostgreSQL. It lives here because this is where the
contract is defined, written as something that runs so it can be read.
It reconciles rather than applying a change, because it is never told
what changed. Three things that follow, and each is a fault somebody has
shipped:
- the password is set every time, not only on creation, or a rotation
reports success and changes nothing
- what it made and nobody asks for any more is revoked, or a departed
consumer keeps a working login for ever
- what it did not make is left alone, or it cannot be run on a database
that predates it
Proven in the lab against a real PostgreSQL, each assertion confirmed to
fail with the behaviour removed. The suite is in mesh-lab, which also
records the two ways the test itself was wrong first.
HAL keeps env vars in the registry, encrypted at rest. Its own tooling
records what that bought and what it did not. `secret_locate` matches by
value rather than by name — because the same password sits in
mesh_provisions, in module_env, in each node's .env in plain text, and
inside every connection string composed from it, and its documentation
says those URL copies "are often the only copies actually in use". And a
query against the encrypted column returns zero rows and proves nothing,
so auditing moved to the decrypted copies on the nodes.
Two faults there, and encryption at rest addresses neither: the control
plane can read what it stores, so a copy of the database is a copy of
every credential; and one secret has many homes with nothing tracking
them.
So here the mesh generates a password, seals it to each end with keys
those nodes generated, stores both blobs, and discards the plaintext. It
cannot read what it holds. Neither can the broker relaying it. And
nothing is composed centrally — a connection string is assembled on the
machine that needs one — so no copy is ever minted in a shape nothing
tracks. `Compromise of a node is compromise of that node` (ADR 0004) is
now true of secrets, not only of identity.
Two files rather than one, because the mesh cannot compose a document
containing a value it discarded: `binds` carries the readable facts,
`secrets` carries the credential alone. The readable half stays readable
in the declaration; the secret half changes only when the secret does,
which makes restart-on precise. The provider gets a directory, one file
per consumer, for the same reason.
It is made once and kept — regenerating per declaration would restart
both ends on every push, and the password a provider was told to create
would never be the one its consumer was given. It is remade when either
end's sealing key changes, and both ends learn the new one in the same
push, so there is no window where half the mesh holds a dead credential.
Two tests found passing for the wrong reason, both caught because their
injection came back clean:
- the provider's copy was asserted non-empty, which reads the same
whichever column is selected. It now opens the blob with the
provider's own key.
- RotateSecret deleted and re-created; the re-create was dead, because
the next read makes one anyway. Removed, and a second path to the same
act is how two ends come to disagree.
And one real fault: three places built a declaration, and the one behind
`--json` predated credentials, so it silently produced a declaration
missing them — a difference between what `plan` showed and what anything
reading `--json` got. There is one path now.
Knowing that a machine needs the anchor's database is useless to the
program that needs it unless the program is told. It knew; nothing was
written anywhere it could read.
Two fields, mirroring contributes/receives in the other direction:
serves: {database: {port: 5432, driver: postgres}} on the provider
binds: {database: /etc/app/database.json} on the consumer
The provider says what a consumer needs to know; the mesh adds the half
only it has — which machine, and what that machine is called on the
private network. The file says, in itself, that it carries no credential
and why. A missing field looks like a bug; a stated absence looks like a
boundary.
Binding something answered on this machine writes nothing. A file saying
"it is on this node" is a fact nobody needs and one more thing to keep
true.
And two machines that share no private network are refused rather than
wired together. An app here and a database there with no path between
them is a mesh that reports itself configured and does not work — the
failure surfaces as a connection timing out, which is the slowest place
to find it. This is checkable now only because the network became
something a machine is given rather than something it has by having an
address.
One fault, found by running it: working out who is on the private network
resolved the mesh, and resolving the mesh asks who is on the private
network. It hung for two minutes. The comment above the function said not
to do that and the function did it anyway; it now resolves each node
locally, which is the right answer to the question regardless — whether a
machine is on the network depends on what it was assigned, not on what it
takes from others.