The fourteen undeclared mounts are declared. More to the point, a
manifest that does not declare one is now refused: they were right by
coincidence, and a checklist nothing enforces is a checklist that is
true until the next commit.
Refused in the control plane, because the machine cannot tell the
difference — asked to mount a path that does not exist, it makes the
directory, which is a thing it is perfectly able to do.
A module now says its port once, in `listens`, and the container's
mapping, the rule set and what a consumer is told are all derived from
one assignment. The three hand-written copies that agreed only because
one person wrote them are gone.
The half that made this an issue rather than an inconvenience was that
the substrate is not a module: nothing in the mesh had heard of its own
store, so it handed a database module the port the store already had.
The machine now says what it carries, and the mesh assigns around it.
Ports the protocol fixes became claims, which needed no new mechanism —
the mesh already had one for what is singular on a machine.
Left open, and unchanged by any of this: whether a module should publish
to the machine at all. Assignment makes publishing safe without making
it necessary.
Jochen's call, and the right one: a module cannot choose a port well,
because it is written once and assigned anywhere. Any number it picks is
a guess about a machine it has never seen, and two modules guessing the
same number is not a mistake either of them made.
Writing it up turned up something the issue had missed. The same number
appears three times in every module — the rule set, what a consumer is
told, and what the runtime publishes — and nothing checks that they
agree. They agree today because one person wrote all three. A module
whose `serves` said one thing and whose container published another
would resolve, compose, apply, and hand every consumer a port that
answers nothing.
So the decision is one source with the other two derived, and an
assignment made once and kept, as a credential is.
The part that needed thought is ports that cannot move — mail on 25,
submission on 587. Those become claims, which is what the mesh already
has for what is singular on a machine. Two modules wanting 25 is the
same shape as two wanting the seat, and gets refused by name at
assignment rather than by a container runtime at apply. That makes this
mostly a matter of pointing an existing mechanism at ports.
Left open: whether a module should publish to the machine at all.
Assignment makes publishing safe without making it necessary.
Found by fixing 027 and pushing again. The declaration is now accepted
and the database container still cannot start: the mesh's own store
holds 5432 on that machine, and the module publishes 5432.
Nothing catches it because the substrate is not a module. It arrives
from the bundle before there is a mesh to ask, so the control plane has
never heard of the store and does not know it holds a port. Resolution
can compare modules with each other and cannot compare one against what
the mesh is built on.
Nor does it compare modules with each other. A port is exclusive on a
machine in exactly the way a claim is, and the mesh has a mechanism for
that which ports do not use.
It has been met before: the end-to-end test that exercises a real
database publishes 5433 rather than 5432, inline, with nothing saying
why. That is how a constraint becomes folklore.
The open question is bigger than the bug. Whether a module should
publish to the machine at all decides how a consumer reaches it, and
changes what `serves` means.
Found by the forge failing to start. I had put `restart-on` on nine
containers so they would pick up a rotated credential; it belongs to a
service, and the host refused the whole declaration.
Removing it fixes the modules and leaves the reason I reached for it.
The mechanism is written against exactly this, in the host's own words:
a running service does not re-read its configuration, so replace the
file, find it already running, do nothing, and the machine keeps
behaving as before while every check passes. Every word of that applies
to a container, and nearly everything the mesh runs is one.
The cost is concrete. Rotation replaces the file and tells the provider
to accept the new credential. A provider reconciles, so it takes it. A
consumer is usually a container, so it does not — and the two ends hold
different passwords, which is the fault ADR 0001 records costing two
days. The test that proves rotation works uses a consumer that reads the
file on each attempt, so it does not meet this.
Two things a fix has to keep: it stays declared state rather than a
command, because the link may not carry an action; and where an env-file
changed, the honest verb is recreate rather than restart, because a
container's environment is fixed at creation.
Three corrections, two of them to things I wrote today.
026 is the serious one. Four modules mounted fourteen host paths nothing
declared — the mail spool, the databases, the object store's data. The
runtime creates those as root, so owner and mode go unapplied, and the
rule that keeps a directory holding data the mesh did not put there is
written in terms of declared directories. It reached the configuration
and missed the data. The cause was carrying compose files across: a
container shape that can express one gets filled in like one.
025 claimed nothing turns a tag into a digest. That is false, and the
answer was designed and built before I wrote it. A module names an
artifact, not an image, and `kind: upstream` mirrors somebody else's
image into the mesh's own registry, pinned by the digest it lands with.
The two-document split the issue described as the shape of a fix is the
design. Pinning twelve images by hand was treating the symptom, and left
them pointing at a public registry rather than the mesh's.
And the image store was written up as something the mesh does. It is an
ordinary module — considered for the substrate and removed, because the
test is whether the control plane needs it before its first instruction,
not whether it can grant itself one. So somebody's own registry is the
same module as the mesh's.
Recorded against phase 3, because the phase note said running them
needed images stocked and provisioners built — and missed that not one
of them named an image that exists. Sixty-four zeros where a digest
belongs, eighteen times, parsing and resolving perfectly.
The forge now runs: on a database another module provides, with a
password it did not choose and a connection string it could not have
written itself. First of these descriptions to be started rather than
planned, and it exercises the whole of the credential work.
What remains is the mechanism rather than the data. Nothing turns a tag
into a digest as part of the mesh's own work, so it was done by hand —
which is what the issue says a person should not be asked to do. Asking
a registry takes a second and pulls nothing, so the main argument for
leaving it undone is gone.
The refusal landed, and the examples pin images that exist. What is
still missing is the part that makes it unnecessary: nothing in the mesh
turns a tag into a digest, so it was done by hand — which is precisely
what the issue says a person should not be asked to do.
Recorded because the resolution mechanism turned out to be trivial:
asking a registry what a tag points at takes about a second and pulls
nothing. That removes the main argument for leaving this open.
Also records the two faults that fell out of pinning for real. The mail
system named seven repositories that do not exist, because it publishes
to a different registry than the manifest assumed, and one of the seven
had been renamed upstream. Nothing checking only the shape of a
reference could have found either.
The code and its tests went in hours before the record was touched, so
an issue that read `located` had been fixed all along. That is the exact
failure the frontmatter exists to prevent: status is meant to be
answerable from the record rather than by reading the code.
Closed with the commit that did it, and cross-referenced to 022 and 023,
which came out of the same mistaken instinct — treating the machine as
a boundary, then as an identity, then finding a consumer had a password
and no name to present with it.
The module descriptions sit in `examples/` inside the control plane, and
that name has been doing harm: everything there reads as a sketch, and
one shipped naming a container image nothing builds. A directory called
the catalogue would have made "does this work" the obvious question.
The shape of the answer turns on one measurement. Of the 126 modules in
the system being replaced, 47 are software in their own right — the
largest is 182 source files, and a speech-capture module carries a whole
daemon. Another 44 ship helper scripts. Only 35 are a description and
nothing else.
So a catalogue cannot be a folder of manifests, because two thirds of
modules are programs. That splits them four ways, and only two of the
four belong in a catalogue: things the world made that we describe, and
packages with some files. What the mesh is made of stays in the
repositories that build it. What we wrote keeps its description beside
its code, in the same commit, because nothing else can stop the two
drifting.
The mesh's list of modules is a table, not a repository, and it already
records where each module came from and at which commit. Nothing needs
inventing for modules from anywhere; a repository of ours is just the
source we curate.
The check that a description is valid should move to a command on the
control plane's binary. Today a test reaches into the control plane's
internals to parse manifests, and another reads its build file to check
images exist — two jobs tangled. A command would also give the same
check to somebody describing their own application, which is the case
that matters most and has none.
Left open: how a provisioner's image gets published and pinned, and
whether thirty-five install-a-package modules deserve to be modules at
all.
The cause was one line. Machines get a systemd-networkd unit with a
static address, so networkd finishes and reports the link configured.
The registry ran `ip addr add` inline, which leaves networkd waiting to
configure something it was never told about — and
systemd-networkd-wait-online has an infinite timeout.
So network-online.target was never reached and everything ordered after
it never started. On these machines that is Docker, so `docker load`
blocked on a socket whose daemon was queued behind a target that would
never come, and three bounded timeouts stacked to thirty-five minutes.
These machines have no DHCP by design, so that wait was never going to
end.
The hypothesis in this record was wrong and the record now says so.
Stocking had just been changed, so stocking looked guilty; stocking
takes 34 seconds and always did, timed directly before changing
anything.
Fixed with two things that made it cost hours instead of minutes: an
image placement now waits for the runtime and refuses after 120s naming
what systemd is waiting on, and the end-to-end test passes onProgress —
the raise reported every step and the test discarded it, which is why
thirty-five minutes and four minutes of silence looked the same.
The suite then ran to completion, 23 of 24, the one failure a check of
its own flagging a path as a credential because `/` is in the base64
alphabet.
Also recorded: a redirected log lags, because Node block-buffers stdout
to a file. Read as a stall twice, the second time right after the real
fix — where a buffering artifact argues the fix did not work.
Seen twice today. Once mid-run: thirteen passes, then the process ended
with no summary, no failure and no receipt. Once from the start: the
first test ran 35 minutes against a measured 4.5 and was still running
when it was stopped.
Ruled out rather than assumed: not memory (84 GiB free, no OOM), not the
daemon (the stalled machine answered `incus exec` immediately), and not
the changes under test — the anchor VM had no host log and no
containers, so the run never reached placing the host.
What changed just before is that the rebuild went from two artifacts to
six, and every one of them is pushed into the scenario's registry, which
is the step the second stall sat in. Recorded as what changed, not as
the diagnosis.
The reason this is an issue and not a slow test: the suite prints
nothing between starting a scenario and finishing its first test, so
four minutes and thirty-five look identical from outside, and the only
recourse is to guess. That is how a workstation was left unbootable in
August. And a run that ends silently after thirteen passes is a run
somebody may believe.
023 is fixed, so the design faults are gone and one concrete thing is
left: the realm provisioner does not exist. Its manifest named an image
nothing builds and no program backs, which has been removed — a manifest
describing a program nobody wrote is the same mistake as the credential
files that could never be read.
Keycloak's manifest now says what is true today, and the gap is loud: it
no longer claims to provide oidc-client, so a consumer asking for one is
refused by name at plan time instead of resolving cleanly and waiting
for a client nothing will create.
The provisioner should be written against a real Keycloak in the lab
rather than from the API documentation. The object store's took three
corrections that only a running server produced.
Tool servers — 56 modules, over half — were written up as the biggest
missing thing. They are expressible with what exists, and the first
framing was wrong in a way worth keeping: a module provides `tools` and
the session requires them does not work, because a requirement has one
answer and 56 modules offering tools would be 56 answers.
Turned around it fits exactly. The session provides `tool-host`; every
module offering tools requires it and contributes where its tools are.
Many-to-one is what `contributes` has always been, and the session
receives all of them in one file. Verified by resolving it rather than
by reading the code.
It only became possible today: until 022, several modules on one node
requiring the same thing was refused outright. Worth noting because it
means the credential fix bought more than credentials.
What remains is a decision about what a tool server is, which is work
rather than a missing shape.
The entry stays in the list rather than being deleted — a checklist that
quietly loses its biggest item reads as though nobody looked.
Both halves had one cause: the mesh knew something and did not say it.
Who a consumer is now comes from one derivation, sent to the provider in
its grant and to the consumer in its binding, so the two agree by
construction. The provisioners use the name they are given and refuse to
invent one, because a name of their own would create a login the
consumer could never guess while everything reported success.
Bound values reach the file that needs them through the symmetric twin
of the sealed placeholder — simpler, because they are not secret, so the
control plane fills them in and the host gains nothing.
The lab run meant to prove this failed in a way that looked like the fix
being wrong: rotation could not authenticate against a real database.
The cause was the suite rebuilding the control plane's image and not the
provisioner's, so an image built that minute ran against a provisioner
built the day before. That is 005's family and is recorded with the
issue, because the misleading part is worth more than the fix.
All three modules have manifests, all three parse, resolve and plan, and
none of them can start. Worth writing down before it reads as progress
or as failure, because it is neither.
The vocabulary held. Nothing in 3.1–3.3 needed a new shape — including
the mail system's several containers on a private network, which was the
one expected to break it. That was the question this phase was designed
to answer.
What did not hold was underneath: 022, now fixed, and 023, open. Both
are about credentials rather than about what a module can say.
The third fault was in the manifests, not the design: a secret declared
at a path named .env and read as one, when a sealed file holds a
password and nothing else. That is what a manifest checked only by a
parser buys, and it is why there are now two tests reading the manifests
on disk.
023 is the whole of what remains before the identity provider runs.
I wrote that a module cannot declare an action. It could — the parser
accepted one, and the refusal only came on the machine. The claim was
wrong in the direction that matters: it read as "the design prevents
this", when what prevented it was a check at the far end that nobody
would connect back to the manifest.
Health checks are still the gap most worth closing, but the shape of the
answer is different from what I wrote. An action is not available to a
module at all, so a health check needs a way to say ask this and expect
that without saying run this — closer to a listens entry than to an
action.
Also records the finding itself, because it is a recurring shape here
and not a one-off: a rule enforced only at the far end is enforced and
unusable.
The playbook offered `env-file` and `${secret:name}` as alternatives,
and that reading is what produced the bug every example module shipped
with: own-secrets pointing at a path named `.env`, mounted as env-file,
holding a bare password. The container starts with no password set —
which is a service running on the wrong credential, not a failure.
They are not alternatives. A sealed file holds a password and nothing
else, so env-file points at a file the module declares whose content
leaves a hole, and the host fills it on the machine. A provisioner is
the exception, because it reads a password file.
Written out as the three lines a module needs, with the failure it
prevents named, since the abstract version was already there and was
read the other way.
Amends the credentials page, which said "every pair has its own
credential" and meant two machines. Built that way, it was wrong in a
way that only shows on a real node: a machine running several services
against one database server had one credential between them, so the
provider refused to plan at all and the consuming node quietly gave the
first module a credential and the rest nothing.
The page already argues the case against itself — one credential with
many holders is the first of the three faults it was written to remove.
It just drew the boundary at the machine.
Two modules on one node are as separate as two on different nodes, and
one login opening both is what this page exists to prevent. It is also
what makes withdrawal possible: one role per machine cannot say that
this module has lost its login and the others still have theirs.
022 turned out to have a silent half worth recording: the provider
refuses loudly and names the modules, which reads as a decision, while
the consuming node does not refuse at all. Three modules wanting one
database produce one need, so two of them get no credential file and
each starts and fails to authenticate with nothing saying why.
023 is what remained after fixing it. A consumer now gets its own
password, in whatever shape its configuration wants, and still cannot
connect: the user name is invented by the provisioner and recorded
nowhere in the mesh, and the host and port sit in a JSON binding that an
application reading KEY=value cannot use.
The asymmetry is backwards and the coverage document now says so. The
secret is the hard case, because the mesh must not be able to read it,
and the secret is the part that arrives. The host and port are ordinary
facts the mesh holds in the clear, and they are the ones stuck.
Keycloak, Gitea, Mailu and MinIO all parse and resolve and none of them
can start. This is what stands between the module set and a running one.
Found while checking whether the module vocabulary covers real use
cases. A node running three modules that all want a database cannot be
planned at all:
anchor has 3 modules asking for "postgres-database" and they would
share one credential: gitea, keycloak, umami
The refusal is right about what it says and wrong about what it implies.
They would share one credential, and sharing is worse than refusing —
but the arrangement being refused is the ordinary one, and the node this
mesh exists to take over runs eight modules against one database server.
The cause is the key: a credential is keyed by provision, consumer node
and provider node, so `consumer` is a machine. The provisioner inherits
it and names the role `mesh_<node>`. The refusal is not a check that
caught something; it is the only honest thing that function can do with
a key that cannot tell two consumers apart.
It is the same mistake as 021 with a different face. There the machine
was treated as a trust boundary; here it is treated as an identity, as
though "who is asking" is answered by naming a host. Two modules on one
node are as separate as two on different nodes.
Worth stating plainly: without the refusal, gitea's login would have
opened keycloak's database, and nothing would have said so — from the
provisioner's side it created exactly what it was asked to create.
Not a local fix. It crosses the control plane, the grant file naming and
every provisioner that names something after a consumer.
Every manifest in the system being replaced was read and every key
counted, then set against what the new one can express. Three findings
worth more than the table.
**The most-used key was already covered and I expected a gap.**
Depending on another module — 65 manifests, the commonest thing any of
them says — is a requirement naming a module, which already means that
module rather than anything providing the name.
**The largest real gap is tool servers: 56 modules, over half.** A
module can already run one; what is missing is anything saying it offers
tools. That is plausibly a provision rather than new vocabulary, which
would need nothing added — not yet decided, and recorded as undecided.
**The gap most worth closing is health, at seven modules.** The mesh
knows a container is running, which is not whether it answers, and this
project has paid for that distinction twice. An action with a verify is
exactly the right shape and may not arrive over the link, so a module
cannot declare one.
Two things are missing deliberately and say so: stage hooks, because the
link may not carry an action and a module needing setup ships a program;
and flavours, retired in favour of claims.
Config merging is missing and should stay missing. A mechanism that
understands TOML gets asked for YAML, then INI, which is how the thing
being replaced became unholdable.
Also records what the survey found that is not about coverage: manifests
that had stopped matching what was actually brokered, one fact derived
in two places giving two answers, and a live listing returning
credentials in plaintext.
Written after porting the first real workload end to end. Every step
exists because skipping it cost something, and the ratio is recorded
because it is the lesson: six attempts, one real bug, and the mesh was
right every time.
The rule worth carrying out of it: read the host's log before
theorising. A declaration that was sent and not applied says so there
and nowhere else — it took an hour to look, and the answer was one line.
The conversion's detail is operational and names machines, so it lives
in the mesh's knowledge base rather than in this repository:
`migration/where-service-data-lives` for where every service's data
actually sits, and `troubleshooting/db-password-frozen-at-first-init`
for the lockout. This document says the rule; those say the specifics.
The lockout is the finding worth carrying here, because it is worse than
the one this plan was already guarding against and it is likelier. A
database image consumes its password variable only when its data
directory is empty. Everything keeps data on a persistent directory, so
the role holds whatever password it was created with for ever;
regenerate the variable and the application moves on while the database
does not, permanently, because nothing reconciles it.
Eight modules are in that state today and work only because nobody has
regenerated their credential since their data directory was created.
It was already documented in the knowledge base and my survey had missed
it — found by searching, which is the argument for the knowledge base
existing.
Pinned 2.5.0 rather than latest, on the suspicion that its draft
profiles extension was involved. Identical failure, so that is ruled out
and recorded — two of the three guesses in this issue have now been
tested and both were wrong, which is the useful half.
The scenario keeps the pin regardless; it should have had one from the
start.
Against a real ACME server the proxy orders, the challenge is answered
at the name on port 80 through the proxy itself, the authorisation goes
valid, finalisation is accepted, and the authority issues a certificate.
The client then posts to an empty URL to collect it, and never does.
Read from the authority's own log rather than inferred. Across one run
it issued two certificates and accepted finalise three times: the client
reaches issuance every attempt and fails at the same step after it.
Ruled out and recorded, so nobody repeats it: the directory is complete;
the authority's API certificate covers the address; the challenge path
works. A hand-written server config was suspected and was wrong —
replacing it with the server's own default, changing only the challenge
port, gives the identical error.
Filed rather than pursued because what remains is interop between two
libraries against a server that exists to be a test server, and may say
nothing about a real authority. What the mesh needed to show, it showed:
a routed name gets a certificate ordered from a configured authority,
and an unrouted one gets nothing — that second assertion passes.
Phase 1 closes with this one item partly open. Two of its four tasks
needed no code at all, the network shape was built, and the next thing
to learn comes from moving a module rather than a fourth lab run.
An earlier paragraph implied a secret becomes unrecoverable once
accepted. It does not. It is sealed to the node, which holds the private
half and writes the plaintext into the module's own file at 0600 — the
value is there, on the machine, as an ordinary file.
What does not exist is a way to ask the mesh what a secret is. That is
the property worth having and it is narrower than what was written.
The reason to capture the old system's environment first is simply that
adoption means supplying those values, not that they become
unrecoverable.
Nothing is rotated during the conversion. A service keeps the password
it is already using, because minting a new one is how a running service
stops being able to reach its own database mid-migration.
The mesh has both paths already: generate-and-seal for a new module,
accept-and-seal for an adopted one. Adoption needs the second, and it is
built.
Rotation becomes a separate act afterwards, once everything works — the
machinery is proven, and it is a thing to do deliberately rather than as
a side effect of moving a service between systems.
Records the step that has to come first and is easy to miss: read the
current environment out of the old system while it can still be read.
Once accepted, the mesh cannot show a secret back, and once the old
system is gone neither can that. A password nobody wrote down is a
service nobody can adopt.
Assumed throughout and stated nowhere — the wrong way round for the most
consequential fact about this component.
A board reachable only over the private network would sit inside the
boundary 0004 already calls the security boundary, and a login there
would guard a room whose door is inside the building. This one faces the
internet, so its login is a perimeter rather than defence in depth.
Which makes the identity provider the mesh's outermost gate. The board
presents the control plane, and the control plane's networked surfaces
can change the mesh (0035) — so whoever that provider admits can assign
modules, from anywhere. Written flatly because it is easy to arrive at
one reasonable step at a time and then be surprised by.
What follows is not the board's own design: who may log in is a decision
about the mesh rather than about an application; a public name needs a
certificate from an authority the world trusts, which is why that work
exists; and the provider going wrong in the permissive direction is a
mesh-wide exposure with no local symptom.
The command line is unaffected and is why this is tolerable — it
authenticates through nothing and answers to the machine's own login, so
the mesh stays operable by somebody standing at it whatever happens to
the gate. That is the property to protect if the rest is ever traded
away.
Bootstrap stopped when the control plane started — a mesh that runs and
cannot be used by anybody not standing at the machine, since the
networked surfaces need an identity provider and no module has been
assigned yet. It now runs through the provider and the first login.
The obstacle was not incidental. The mesh has never held a readable
secret: Make generates and seals, keeping no readable copy. An initial
administrator's credential is the first value a person must read.
Generating it and printing it once was the convenient option and is
refused. It would give the control plane a plaintext secret for the
first time — briefly, and to one terminal, but the capability would then
exist, and an exception made for one case does not stay one. The next
awkward credential gets printed too, and "a copy of the database is a
copy of nothing" stops being checkable by reading the code.
So the operator supplies it, on standard input, not echoed — the path
that already exists for a model-access key. What is created is an
account in the identity provider, not a user of the mesh; there is still
no user model.
Unattended bootstrap remains possible and the value still comes from
outside: automation supplying it is the operator supplying it. What is
refused is the mesh inventing one, so an unattended bootstrap with
nothing provided yields a mesh with no administrator — correct rather
than broken.
The mesh is operated from a command line and must be operable from a
browser and from a model's tools, without becoming three systems. The
pattern is already in the code and was unnamed: `board` serves HTTP by
calling the same functions the CLI calls, holding nothing.
Takes the decision 0034 said had to be taken deliberately rather than
arrive with a feature: the HTTP surface is not read-only, so a browser
login now carries authority over the mesh.
Names the dependency by protocol — an OAuth2 identity provider — as the
mesh does for AMQP, S3 and OCI. Keycloak is what fills the role; what
the control plane knows is that it validates a token, and replacing the
provider is a migration rather than a redesign.
Says what this must not become, because it is the failure the project
was started over: a kernel every module imports, 155 files of code from
every context. Shared surfaces are not a shared library. Three adapters
calling the same functions is not the same as logic leaving the context
that owns it.
And records the loop it creates. The networked surfaces depend on a
module the control plane assigns, so when identity is down nobody can
authenticate — including whoever is trying to fix it. The way out is the
command line, which authenticates through nothing and is available to
the account that owns the machine. Hence the rule: no capability exists
only behind an authenticated surface, because that is a capability which
disappears exactly when identity does.
Supersedes 0032, which decided the right thing and described it wrongly.
The decision is unchanged: the account that installed the host owns the
mesh, and there is no user model.
What was wrong was inventing "a surface that delegates authentication"
for the board. It is a web application with a login, in the way every
web application has a login. That is a fact about an application, not a
property of the mesh.
The cost was not cosmetic. It made the identity module look like part of
the mesh's authority — something the mesh depends on to know who anybody
is — when the mesh knows nothing about people at all and one of the
applications running on it happens to have a login.
Keeps the line that is worth writing down, and states it more plainly:
signing in to an application must not become authority over the mesh.
Today it cannot, because the board reads and does not act. The moment it
can assign a module, whoever it lets in has mesh authority — and it
would arrive as a feature rather than as a decision. So a surface that
can change the mesh is a change to who owns the mesh, and is taken as
one. Not forbidden; just not something that turns up in a pull request
titled "add assign button".
Third correction to one table today, found the same way as the other
two: by asking whether both halves of the test were answered, or only
the easy one.
0006 admits the registry because "it cannot grant itself a repository" —
true, and the second half. Nothing established that the control plane
needs one in order to run. Counted rather than argued: the bundle raises
twelve resources and no registry is among them. The registry arrives
afterwards as an ordinary module, which is exactly what the lab asserts.
0006 half-said this already, calling it "substrate by role and ordinary
by delivery, provisioned once there is a control plane to do it". A
member provisioned by the thing it supposedly precedes is not a member;
that phrase was carrying a contradiction rather than resolving one.
The registry is a closer call than the object store and the difference
is worth keeping: the control plane never touches an object store at
all, but it genuinely uses the registry. So the registry is a real
dependency of the mesh operating and not of the control plane starting —
and it is the second that the word means.
The substrate is now exactly what the bundle raises, which is the
strongest form the list can take: checkable by counting rather than by
reading an argument, and the two cannot drift.
The finding is not about substrates. A test with two conditions is a
test only when both are asked.
Answers what 0031 left open, and a question it did not ask — who owns
the mesh at all. There was no answer, and the absence was invisible
because every operation so far has been run by the person sitting at the
machine, so nothing had to say whether that was the design or the
circumstance.
The account that installed the host owns the mesh on that node. No user
model, no roles, nothing to administer. It follows from 0004 rather than
adding to it: there is no authorisation between nodes because every node
is the operator's own, so a user model inside that boundary would guard
nothing — anyone it could stop could read the node's key off the disk.
The board is different, and the difference is the network. A surface
reachable by a browser has to know who is asking, because those people
are not by construction people with a shell on the machine. So it
delegates to an OAuth provider, which is a module.
That does not make identity substrate. A surface delegating
authentication is not the control plane delegating it: the control plane
runs, applies declarations and reaches nodes with no identity provider
in existence. Only the board needs one.
Records the cost plainly: anybody with a shell on a node has full
authority there, and there is no way to give somebody authority over one
node without giving them a login on it.
Closes the last open question about what the substrate contains. 0006
left an identity provider conditional — substrate only if the control
plane delegated authentication — and said the decision had not been
taken. It is now: it delegates to nothing.
The conditional was never about machines. A node proves itself with a
keypair it generated over a broker account issued at enrolment, and
declarations are verified by signature; none of that involves an
identity provider. It was only ever about whether a person signing in to
a mesh surface would be authenticated by something else.
So the substrate is three — a relational store, a message bus, an image
registry — and with 0028 having removed the object store, no member is
conditional and every one is there for the same reason.
It does not settle how a person signs in to a surface, deliberately.
What is settled is that whatever answers that is not something which
must exist before the mesh does, so it can be decided late or replaced —
which being substrate would have prevented.
The conversion method, recorded because it decides everything else and
was not written down.
The old control plane is stopped — provisioning, coordinator, syncs, the
pipeline, anything that decides or writes. The workloads it was managing
keep running, because nothing is managing them. The new mesh then takes
ownership one module at a time.
Nothing is ever unassigned in the old system. Unassigning is how it
removes things and removing is how data is lost; it is asked to stop
having opinions, never to take anything away.
Disabled rather than merely stopped, which is the part easy to get
wrong: those units are enabled, so a stop lasts until the next reboot. A
reboot mid-conversion would bring the old control plane back to
regenerate managed files underneath the new one — the one situation
where two systems really would fight over a machine.
A service left running with nothing managing it is the safe state: it
has its data, its configuration is on disk, and nothing will change
either. The risk in a conversion is in the managing, not the running.
Also records why taking ownership piecemeal is safe: the new host's
orphan removal is per-origin, so it only removes what it recorded
itself. Services it was never told about are not orphans to it.
0030, found by asking what the conversion actually needs rather than by
reviewing anything. The host deleted a directory and everything under it
when it stopped being declared — which happens when a module is
unassigned, or when a manifest is edited to move a data folder, which is
the exact operation this plan needs. A database's files, a mail spool.
The report said "removed".
A directory still holding something is now kept and said so. No flag and
nothing to remember: emptiness is the test, and it works because the
removal order was already right — the mesh's own contents are gone by
the time the directory is reached, so what remains is by definition
something nobody declared.
The plan now says data outranks its own ordering: copy, read back
through the service that owns it, and only then point anything at the
new location. Never move and then check.
And it records where this starts — the node holding all the production
data — with what that costs stated rather than argued with. Everything
proven so far was proven on machines that could be destroyed and raised
again. A scenario proves the mechanism, not the state on that machine.
Recorded because it is load-bearing and was not written down: moving
from the current system to this one is a person at a command line, not a
migration program.
What that removes is larger than what it adds. Nothing in this plan
needs an importer, a translation layer, a compatibility shim, or a way
of keeping two systems agreeing while both are live — each of which
somebody would otherwise reasonably build, use once, and maintain for a
year.
It also settles what "safe" means for the system being retired: a fix to
it must be safe on its own, because there is no careful rollout to
sequence it into. A change needing three steps in the right order is a
change that will be half-applied. That reversed a certificate default I
had chosen this morning.
Ordering needed no change for the third time running — resources apply
in the order declared and nothing sorts them — and is now asserted,
because sorting them for any sensible reason would have passed every
other test.
Separates ordering from readiness, which the task had run together: a
container started is not a container ready. Nothing waits, and what
needs something usable retries. That is deliberate and more robust than
start ordering, since a dependency can restart long after apply.
The network was the first thing in Phase 1 that genuinely needed
building, and the first that needed a decision: 0029 records why a shape
rather than an action, and the vocabulary is nine.
A session as a licence consumer needed no change either: the two
sessions are two modules, so the existing (node, module) binding already
names them apart. 14-model-access.md's "a step toward it and not it" is
true of a worker and not of a session, and the difference is that there
is one session per node rather than many per machine.
Records what stays open: the worker half of that gap is real and
unaffected, and belongs with 0003, which is unbuilt.
Two tasks in a row that were already possible. Both were written from
the design rather than from the code — the review's own finding arriving
in the plan it produced. The remaining Phase 1 items should be checked
against the code before being started rather than after.
An object-store provision, proven against a real store with seven
assertions.
The finding is worth more than the task: the control plane
special-cases nothing. provides, requires, contributes and grants are
name-agnostic, so asking for a bucket needed no change to the mesh at
all. What was missing was a provider and the last step on the machine —
"add an object-store provision" was never mesh work, and the breakdown
now says so rather than leaving the next person to rediscover it.
Named s3-bucket by 0027: the coupling is to the API, not the product,
because swapping one store for another does not break a consumer. A
database is the other case and names its engine.
Records the assertion a database does not need, because it is the one
that will be forgotten when somebody writes the next provider: one store
holds every bucket behind one endpoint, so isolation is a policy rather
than a property, and a policy granting everything passes every test that
only checks a consumer can reach its own bucket.
**0024 accepted.** Model access was decided, built, and proven in the
lab, and two design documents rest on it; only the status had never
moved. The gate is green again.
**The work breakdown rewritten.** It planned a decomposition of the
existing system in place — extract contexts, declared features, shrink
the shared library. That is not the work. A replacement is being built
beside it, and only the old Phase 0 survived contact with reality, so
the one document meant to say what happens next was describing a system
being retired.
Now ordered by what "modules move across one at a time until the old
registry is off" actually requires:
- Phase 0 is marked done against the twenty-two lab assertions, **and
carries its own limitation**: every module exercised was written to
test the mechanism, so the vocabulary was shaped by its own fixtures.
- Phase 1 is the vocabulary gaps found by asking what real modules
need — an object-store provision, a session as a licence consumer, a
network shape with ordering, public certificate issuance.
- Phase 2 is one module, then a week of running it, because the point of
going first is to find what Phase 1 missed.
- Phase 3 picks modules that each prove something the first did not; the
mail system is last because it is the one that may send work back into
the declaration language.
- Phase 4 is switching the registry off, named as a phase so it is not
mistaken for the goal.
Keeps the rules of engagement unchanged — they were about how work is
done, not what it is — with one addition: stop and ask before anything
that touches a machine outside the lab.
Adds a section on keeping the list true, since the document it replaces
was wrong for weeks and nothing said so. A claim here is counted, not
reasoned, and a phase is done when the lab says so.
First pass of a design review, done by reading documents against code
and against a raised mesh rather than against each other. Every error
below was invisible to a proofread.
**Statuses were stale, and nothing checked them.** Ten to-be documents
said `designed` while naming working, lab-proven code — several with a
*What was built* or *Raised, and observed* section. Added a
`status-vs-code` check: naming a file is a claim that the file
implements this, so a document that points at one has stopped being
merely designed. It failed on all ten before it passed, per the rule
this folder sets for its own checks.
**The bundle carries three images, not two.** 07 reasoned about which
substrate services go in and overlooked that the control plane is in
there too — it is what the substrate exists to start, and there is
nothing to fetch it with yet. Counted, not deduced.
**The bootstrap uses four shapes, not six.** It listed `file` and
`directory`, which substrate-first-node.lock never asks for. The claim
that mattered — nothing is blocked on the host — was true either way,
which is why the wrong count survived.
**The eight capabilities were documented nowhere.** Implemented in
internal/profile/detectors.go and enumerated in no document, including
the one about the host that detects them. A vocabulary modules write
against, readable only by reading the code. Now written down, with the
seat/graphical-session distinction that is wrong in both directions if
collapsed.
**MinIO swept out of the to-be layer** per 0028.
The gate now fails on one thing left deliberately: ADR 0024 is
`proposed` while two documents rest on it and the feature it decides is
built and lab-proven. Accepting a decision is not mine to do.
**0027 — provisions.** A module written against PostgreSQL could be
matched to a provider of SQL Server, resolve as satisfied, and fail on
its first query. The name said the role, so nothing distinguished
engines. Refusing on ambiguity could not help: with one provider of
each name nothing is ambiguous. Enforced at parse rather than
documented, because the old naming was the documentation.
**0028 — the substrate.** 0006 admits an object store on the grounds
that it cannot grant itself a bucket. That answers the second half of
the test and assumes the first: the control plane does not need one.
Verified — no S3 client in mesh-control, and internal/builder/registry.go
records the deliberate choice to put artifacts in the OCI registry as
content-addressed blobs. The row was inherited from the system being
replaced, where an object store distributed module tarballs, and was
never re-tested against the definition above it.
So an object store is an ordinary module, and a mesh with nothing
needing one runs none. Migrating it is module work, not substrate work.
0028 also states what 0006 left unsaid: a substrate service and a
module of the same product are different instances. The substrate is
raised from the bundle before any mesh exists, so it is not in the
module graph — a workload depending on it would depend on something the
graph cannot see, cannot rotate a credential for, and cannot move, and
would put workload data in the store the control plane keeps its own
state in.
Both records were found by reading code against design rather than
design against itself, which is the review that should have happened
sooner.
Answers the question 15 raised: a board showing many sessions leaves
one-per-node untouched, because each is still one conversation. Only
concurrent conversations with the same session would touch 0004.
Records soulstream and herdr as the prior art to draw from, and marks
it explicitly off the provisioning path so it stays a note rather than
becoming the work.
Settles the question 15 left open: the mesh session holds its own
memory in the mesh root, rather than assembling a view over the node
sessions. Memory follows the rule the rest of the design already uses —
the context root is the whole of what makes one session a different
agent, and memory is part of what makes it that agent.
The control-plane node is what makes this load-bearing rather than
tidy. Two sessions share that machine; if memory belonged to the
machine instead of the root they would share it too, and the mesh's
recollection would be indistinguishable from that node's own — the
collision 0026 exists to avoid, arriving through the back door.
Also corrects an error made writing it up: memory is NOT declared
state. The engram and tools are — the mesh says what they are and the
host writes them (0011). Memory is written by the session itself and
declared by nobody, so a mechanism that regenerates the root wholesale
would erase it on the next heartbeat, silently, while reporting
success. The root is not uniformly managed and which parts are has to
be explicit.
A session for the mesh itself, addressed as the mesh, differing from a
node's in exactly three things: the context it starts in, its engram,
and its licence binding. Not a new kind of agent — the same mechanism
pointed at a different root. Two implementations of one mechanism drift,
and the vocabulary collision 0001 exists to undo began exactly that way.
It runs on the control-plane node, and the reasoning is easy to get
backwards: not "the important agent on the important machine", but that
this node is already the one place excepted from "compromise of a node
is compromise of that node". Placed anywhere else it would create a
second such place.
It is an addition to per-node messaging and never a replacement. 0001
holds that losing the control plane costs change, not operation — and a
mesh whose only conversational surface lived there would lose the
ability to ask anything while every machine kept running perfectly.
Writing it up exposed that the node session's setup was never designed
at all. 0004 gives behaviour and stops: nothing said how a session
starts, where its context lives, or how a broker message becomes a
prompt. That gap was invisible until something had to be built *like* a
node session. 15-the-agent-session.md covers both as one mechanism.
It also makes "a consumer that is not a machine" undeferrable. The
control-plane node now hosts two sessions that must hold different
licences, and a per-machine binding cannot express that at all. Noted in
14-model-access.md against the gap it was already recorded as.
Also completes the to-be index, which stopped at 10 and omitted four
documents. Pre-existing broken ADR references in the older rows are left
alone rather than guessed at.
Decides the question 006 narrowed to. An agent reads this repository
directly and the search consults it, so these documents surface beside
ordinary results instead of only when somebody already suspects they
exist.
A scheduled sync into the mesh's memory was the option that works with
what exists today, and lost on the ground this repository can least
afford: it makes a second copy, and the copy that is searched quietly
stops matching the copy that is edited. A design record that has
silently diverged from the reasoning it claims to carry is worse than
one that cannot be found — the first misleads, the second merely fails.
Amends what 0019 promised rather than satisfying it: these documents
will not be indexed, they will be read. The commitment that survives is
the one that mattered — that a searcher finds them without already
suspecting they exist.
Gated on an agent that does not exist yet, so 006 stays open on the
build with a decided shape. What closes it is a check that fails today
by design: search the mesh's memory for a phrase that appears only in a
design document here, and require it back.
**004 — certificate issuance.** The resolver declared no authority at
all, so the client fell to its built-in production default: there was no
setting set wrongly, there was no setting. It is now a node property
defaulting to staging, which answers the first open question. Staging by
default rather than production-with-an-override, because the alternative
leaves the safe path depending on remembering to opt out of it — 005's
lesson, in a second place. The rollout is ordered and the order is the
dangerous part; recorded, not performed.
**008 — node rescue.** Read back from running nodes as the report asked,
and one of its own claims was wrong in a way that matters: the health
timer does exist and does fire. It simply never calls the rescue script.
A trigger that exists and does not do what the script claims survives a
halfway check, which makes it worse than the absence the report
described. Resolved by making the documentation true, not by
implementing rescue — the replacement host already supervises recovery,
and wiring unattended restart into the fleet being retired is a
deliberate decision rather than a tidy-up. Two "self-healing" claims
narrowed to what they actually do.
**006 — deliberately not closed.** Re-checked today: the indexing still
does not exist. What is gone is the reason it was an issue — the claim
is no longer load-bearing, because the README names the gap and the
decision's reasoning never invoked indexing. A signpost now points here
from the knowledge base, and was measured rather than assumed: it is
reachable, it is not surfacing. Closing it while the indexing does not
exist would be this repository's own named failure, one folder from
where it names it.
Retired in favour of the lab rather than repaired — that answers the
first open question. The second finding is the one that generalises:
"nothing runs it, and nothing reports that nothing runs it" is not a
fact about that harness, it is a fact about any suite too expensive to
run on every push. The replacement inherited the fault it was replacing.
Records the three rules that now hold, and what the fix taught twice:
the remedy rebuilt the symptom inside itself, and the code that counts
results passed every test while reading nothing.
A service is reached at <service>.<node>.internal, so what resolves is anything
under a node's name. The mesh writes the data and runs no daemon; two roles,
two claims, because systemd-resolved cannot serve a wildcard at all.
Both prohibitions were found by a machine rather than by reasoning: an address
systemd already held, and reading resolv.conf for upstreams that now point at
itself.
Twice in one file, a statement about a machine that read as reasoned and was
wrong — and the module's unit tests all passed while the daemon could not
start. That is what a unit test is: it confirms the assertion was made, never
that it is true of any machine.
003 in prose rather than in a manifest key.
The field was called needs, beside secrets, and both were name-to-path holding
something secret. What separates them is whose, not how secret — so that is
what the name says now.
012 named its own closing condition — a scenario with four images coming up —
and the scenario now stocks seven and has raised cleanly many times at the
memory the wrong diagnosis had raised.
001 is answered by the host reading the package database back after installing.
002 was NOT answered and was present here too, so it is a fix rather than a
note: a stale index is now named instead of reported as a failed install.
The connectivity design still said a hub cannot be filtered — a gap recorded in
the morning and closed in the afternoon, left standing as though it were
current. Worse than a stale date: it would send somebody away from something
that works.
`restart-on` was described nowhere, including the part added today that lets a
service reflect a file another module put on the machine. A rule the host
enforces and no document mentions is a rule nobody can rely on.
And nine of fifteen design documents claimed an `updated:` older than their last
change, some by a week. That field is what cross-cutting views are generated
from, so it is not decoration.
A resolver takes over /etc/resolv.conf, which is a singular resource — ADR 0009
lists it in the table beside the seat and pid 1. So choosing between resolved,
dnsmasq and unbound is assigning a module, per machine, and the mesh refuses
two rather than letting them fight over the file.
Recorded because it was treated as an open question two days after being
decided, which is the argument for that table being a table.
Found by a container failing to resolve a name every machine could: a container
gets its own hosts file holding only its own hostname, and on the machine it
always worked, which is what made it easy to miss.
Declared containers are given the names. A container somebody starts by hand is
not the mesh's to configure — which is a second, different reason to want a
resolver, recorded beside the first rather than folded into it.
Asked whether a machine that drops off needs re-adopting: it does not, nothing
expires, and the only thing that forces re-enrolment is losing its own key.
The gap was the twenty or thirty seconds after a resume in which a node
believes it is in a mesh it has left — recovering on its own, which made it a
quality gap rather than a fault, and still a machine waiting to be told
something it already knew.
It meant failed-or-refused, so the question this record says must not be lost
was answerable only for the machines that broke. Out of date, never told, and
not worked out are kept apart: the remedy is the same push and they read
differently to whoever is looking.
Their subject matter has been built and proven for days and their frontmatter
still said code: [] — which is what the cross-cutting view is generated from,
so it was claiming nothing existed for the substrate, the node lifecycle and
delivery.
One reading answered three ways, holding nothing and touching no context's
store — which is the constraint the whole document is about, and the thing the
board being replaced gets wrong.
Rotation and the provisioner contract; model access as a provision answered by
a record, with ADR 0024's other two gaps left as gaps; exposure, which closes
the open question about revoking a route; and the delivery loop, which closes
the gap ADR 0010 left when it replaced a pipeline with a comparison.
The broker's fingerprint travels with its credential, and the machine's
filesystem does not travel at all — it runs in a container, which is the
arrangement working rather than a limitation to route around.
Found by the firewall: every packet filtered as declared, and the machine
reported as not doing what it was told, because the unit that loaded the rules
had finished. Stated as a gap rather than worked around silently.
Otherwise it succeeds into a state its verify rejects, and the host's report is
accurate and names nothing. Recorded where the vocabulary is described, because
it is a rule about writing an action rather than about one action.
A hub needs its overlay port open and a node that is not a hub does not, and
they are the same module — so listens, a static manifest field, cannot express
it while the overlay module's resources are computed per node. Written down
rather than left as an oversight for whoever first puts a firewall on a hub.
Issue 014: the node's serving key was stored in the host's own encoding, so
every check that reads the file passed and no server could start. Same shape as
013 — two halves of one mechanism designed separately, each correct about its
own half. Where a file exists so a third party can read it, the format is the
interface.
Issue 003 is answered in both halves: manifests are parsed strictly, and a
module says what it listens on and from where rather than carrying a key
nothing reads. The design records what was built and how each part is checked.
Issue 013 is new, found by reading while writing the first module that has
both a computed file and a service that needs it. The file arrived second.
It failed, then the next reconcile fixed it, which is why nothing caught it.
Two things changed at once: a fourth image in the scenario, and scenario
machines raised from 1 GiB to 2 GiB. The bootstrap then failed every
time, and the image was blamed.
Removing the image did not fix it. Removing the memory increase did —
nine assertions pass again on three images with the machines back at
1 GiB. Three machines at 2 GiB on a host doing other work contend enough
that the store container does not come up at all.
The ordinary lesson, and it still caught me: two changes together, the
failure attributed to the plausible one, and an issue written recording
the wrong cause. What found it was reverting to the exact last-known-good
state rather than reverting the suspicious change.
What remains untested is whether a fourth image alone is fine. Probably.
Nothing has measured it, and the honest state of this issue is that what
it was opened about was never demonstrated.
substrate
Adding a fourth image to the two-machine scenario makes the bootstrap
fail every time, with the store's readiness check producing no output at
all — which says the container was not running rather than that the
database was slow. Three images pass nine assertions; four never get past
the store.
More memory did not change it, so memory is not the cause; the change is
kept because the reasoning holds on its own. Disk is the most likely
explanation and nothing has measured it.
It blocks proving the mesh runs its own artifact store, since the
registry module needs a registry image to mirror. The module is written
and accepted; what is unproven is a machine assigned it serving another.
The first fix continued past every failure, and the next lab run failed
at the bootstrap: the store did not answer in three minutes and then said
"the database system is shutting down". Carrying on past the readiness
gate had started the broker and the control plane against a machine that
was not ready, and on a small machine that is how a database still
initialising has its memory taken away.
An action is the only shape whose purpose is to make something true
BEFORE the next thing needs it, which is why it is the only one with a
verify. So a failed action stops what follows and nothing else does —
which fixes both this and the hostage problem the issue was opened for.
Found in the lab. A machine with one impossible module applied nothing at
all on every later push, and the mesh said "failed" without saying the
rest was never attempted.
Recorded with the evidence, including that the behaviour's test cited a
record which does not decide it: ADR 0010 argues about pipelines against
reconcilers and says nothing about whether one resource failing should
stop the next being attempted.
Fixed in mesh-host: everything is attempted, every failure reported.
Three additions, all written by trying to write a real database module
and finding out what could not be said.
A module may mirror an image it did not write. Naming an upstream
reference directly needs every machine to reach a public registry and
pins to a tag somebody else can move.
A module may need a secret of its own — a superuser password is not FOR
anybody, so the mechanism that hands credentials to consumers cannot
express it. Per node, so three machines have three passwords.
And the provisioner watches, which is what lets it be a module rather
than a binary somebody places. It polls rather than watching the
filesystem, because the host writes atomically and a watch on a replaced
path silently stops working.
One assignment now gets a working database provider: two directories, two
pinned containers, a sealed password and the grants manifest.
A page nobody had thought to ask for turns out to be the one a person
opens first: what is not doing what it was told. Recorded with the order
that matters — broken, then quiet, then out of date — because a page
leading with the last would bury the first.
And refused stays distinct from failed all the way to the page. They are
fixed in different places, so one word for both sends half its readers to
the wrong one.
Two additions to the module-repository design, both from building it.
A build machine has its own credential and it is not a node's: read the
build queue, write the mesh exchange, nothing else. A node's queue
carries that node's declarations.
The answer goes through the exchange and never the default one, because
permission there is per exchange rather than per queue — anything allowed
to use it can publish into any node's queue. The price is that every
asker sees every result and filters by correlation, which is cheap
against a builder never needing that permission.
And every result is kept, failures included, because one that leaves no
trace is indistinguishable from a build nobody asked for. That is what a
builds view reads; the board page is corrected to say so.
A build is work, not state, and that is why it does not travel as a
declaration: as one it would either rebuild on every reconcile or carry
"and I already did this", which is state about an event rather than about
a machine. So it has its own queue and the answer comes back correlated.
Three properties recorded because they are decisions: acknowledge only
once the answer is away, one build at a time, and a failure is a result
rather than silence.
And the board page is corrected. A build result today is answered to
whoever asked and kept nowhere, so a builds view has nothing to read. A
record of past builds is the missing piece, not the builder.
Designed with no reference to what came before, which was asked for. The
system this replaces has features — several deployable units inside one
module — and they are deliberately absent.
That closes something ADR 0001 has been carrying as an open prerequisite.
It lists "named features with per-node opt-in" as required, or "every
independently deployable unit becomes a module again and the count
returns". The premise was right and the remedy already exists in another
form: several modules, assignment per node, and a module with
requirements and no files of its own. `networking` is exactly that. The
count does not return because what made it return — a module is
expensive, so put several things in one — is gone. A module here is a
manifest and usually nothing else.
The manifest in a repository names artifacts; the manifest the mesh holds
names digests. Two documents, because a digest is not knowable until
something is built and a repository carrying one is wrong the moment
anybody edits anything.
The builder runs on a node. Building needs a container runtime and a
working tree, and what the control plane may send a machine is bounded by
the declaration language. A control plane holding a container socket
would be the one component that can do anything anywhere.
And the host's vocabulary grew from six shapes to eight — user and
archive — with the reasoning for each and for the refusals that came with
them. The count is asserted by a test precisely because every addition
widens what a compromised control plane can express.
Read from the board that exists. Eight sections; four are about work and
workers and are held back with that domain. The other four are the mesh
itself, and everything behind the main one already exists here — it is a
reader, not a second source of truth.
The constraint is the point of writing this down now. The existing board
is one service that reads every context's database, because that is the
shortest path to a page showing all of them at once. That is ADR 0008
violated by the one component with a reason to violate it, and the cost
is the same one the shared library has: a boundary nothing may cross is a
boundary that can move, and one thing crossing it is enough to freeze it.
So a board reads through interfaces and stores nothing. If a question is
slow, the answer belongs in the context that owns it, where everything
else asking gets it too.
A new requirement, and it is mostly a shape the mesh already has. A
module that needs to think requires model-access; several vendors and a
locally-run model are several modules providing it; choosing is assigning
the one you want. A model the mesh runs itself needs nothing new at all —
it is a mesh-scoped provision on the node with the hardware, credential
included.
A licence is a named thing because the whole point is saying which one a
given consumer uses, and the names are the operator's. Many to many, so
not a claim: two machines sharing an account is ordinary, not a
collision.
Four gaps, written as gaps rather than design:
- a provider that is on no node, reached over the public internet, which
the reachability rule must not refuse
- a secret the mesh is GIVEN rather than mints. Every credential it
handles today it generated and discarded; an API key arrives from a
person, and accepting one must still discard the plaintext
- a consumer that is not a machine. Which licence a worker uses is a
binding to an agent, and the provisions model has no consumer identity
other than a node
- switching on exhaustion is a reaction to something observed, not a
declaration. It belongs with observability, changing a binding — saying
so is what stops the declaration language growing a conditional
The existing auto-refresh and switching is not being replaced because it
was wrong. It is being rebuilt because it lives somewhere that cannot
express the rest.
First end-to-end raise. A machine with a container runtime applied the
bundle its host carries and ended with a store, databases, schemas, a
broker holding a certificate it generated itself, and the control plane
serving. Then it took a token, checked the broker against the pinned
fingerprint, generated three keypairs and enrolled — the first node being
a node whose mesh is not up yet, observed rather than argued.
And a credential crossed. Declared the provider of a database for a
second node and pushed to over the broker, the machine ended with the
password in one file at mode 0600, and that password appears nowhere in
the declaration that crossed the broker, nowhere in the control plane's
database, and nowhere in what the node reported back. That is the whole
secrets argument, measured.
One fault, in the joining: the token did not say what the mesh calls the
machine, so enrolment needed a flag its own help said it did not, and
failed at the broker with an empty username. It is the fifth thing a
token carries now — the node cannot work its own name out, because the
broker account it authenticates as is named after it and exists before
the mesh has told it anything.
A password nothing was told to create authenticates nowhere. The mesh
generates one, seals it to both ends and cannot read it — so it cannot
tell the software to accept it either. Something on the providing machine
reads what arrived and makes it true.
That something belongs to the module, not to the mesh. The control plane
decides and never touches a machine; a provisioner runs on the machine
and touches it. What the mesh owns is the contract: a manifest of who
asked and where each credential is, and one file per consumer holding it.
It reconciles and is never told what changed, which forces three things
that are each a fault somebody has shipped: set the password every time
or a rotation changes nothing; remove what nobody asks for or a departed
consumer keeps a login for ever; leave alone what it did not make or it
cannot be run on anything that predates it.
Saying where the mesh stops is the point. It decides, delivers, and can
prove what it delivered; the last inch belongs to whoever knows what
`create role` means.
Written after looking at how the existing mesh does it, so this is a
reaction to a measurement rather than a preference.
There, credentials sit in a column encrypted at rest. Its own tooling
records what that bought: the tool for finding a secret matches by value
rather than by name, because the same password is in three tables, in
each node's environment file in plain text, and inside every connection
string composed from it — copies its documentation calls the ones usually
in use. And a query against the encrypted column returns zero rows and
proves nothing, so auditing moved to the decrypted copies.
Encryption at rest addresses neither fault. The control plane can read
what it stores, so a copy of its database is a copy of everything. And
composition is what mints the untracked copies.
So the value is sealed to the node that will use it before it is stored,
with a key that node generated. Nothing central is composed. What it
costs is auditing by value, which was never real anyway; what stays
answerable is which node holds what, which is what rotation asks.
What remains is a provisioner. The mesh generates the secret and tells
both ends; nothing yet acts on the telling.
Which turned out to be the useful way to cut it. A provider says what a
consumer needs in order to use it; a consumer says where it wants to be
told; the mesh adds which machine and what that machine is called on the
private network. So an app on one node reaches its database on another,
by a name the mesh also created.
The file says it carries no credential and why, because a missing field
looks like a bug and a stated absence looks like a boundary.
What remains is the secret itself, and the shape it will arrive in now
exists.
Also: two machines wired together across no private network is refused,
and that only became checkable when the network stopped being something a
machine has by virtue of holding an address.
0009 distinguishes presence from instantiation — what the edge hands
over. It never distinguished where the thing on the other end is, and
that turned out to be the half doing the damage: a shell and a database
were both written `requires`, so requiring a database installed one on
every machine that used one.
A provided name now carries a scope, as a claim already does. Scope
belongs to the name rather than to each provider, or one requirement
means two things depending on which module answers it.
A requirement answered from the mesh is never satisfied locally. Nothing
provides it, and it says which module to assign somewhere; two do, and it
says how to choose. Choosing is recorded per node, because two machines
may reasonably use two different databases.
And knowing which node answers is the first half of handing a credential
back — you cannot be given a database's password before it is settled
whose database it is.
0009 already said a consumer supplies a target and receives a name. What
it did not say is that those are two separate mechanisms.
Contribution — publish me at this name, on this port — now exists.
Binding — and hand me back a credential — does not, and is the larger
half: a secret has to exist, be stored, reach one node and not the
others, and rotate with every holder informed. That is the invariant set
found violated three ways at once, so it is not something to add in
passing.
The absence had a measured cost. Exactly two modules opened a direct
connection to the control plane's database, and they are the reason every
node permanently holds a credential to it. Both were doing by hand what
this edge is for. Neither needed a new kind of thing.
Two records, from building it.
0009 has a section titled "there are no domain modules", and `networking`
now exists. It is not a contradiction and it reads as one, so the
difference is written down: what was refused contains WireGuard and a
proxy and is assigned where half of it is unwanted. What exists contains
nothing — requirements and a name — so there is no half. Every artifact
it leads to is still an ordinary module assigned on its own terms.
With the cost stated, because it is real: adding a second implementation
turns a settled question into an open one for everyone using the bundle,
not only for whoever wanted the alternative. That is the refusing rule
applied consistently, and the alternative is a default, which is the
flavor field returning under a better name.
08-connectivity gains why the network stopped being code beside the
module system: a machine was on the private network because it had an
address, and there was no way to keep one off. A manifest can now say its
resources are computed, which is what a peer list needs.
And three modules rather than one, because WireGuard is one VPN of
several. Naming a module after the job and putting one implementation
inside it is flavor wearing a generic name — the second VPN has nowhere
to go.
Recorded while building the seat detector. A capability is a named fact about a
machine: its presence gates an assignment and its detail can carry a value, so
"can this run here" and "what should it be configured as" are the same fact
read two ways. A verdict has always had a detail beside its yes or no, so
panel: oled needs no new concept.
Two things that keep the set honest, both worth writing down before anyone adds
the fiftieth capability. It must be detected and the detector must say how it
knows -- so nobody can add one they cannot check, which is the whole of issue
007. And detectors ship inside the host, which is one static binary, so adding
a capability means shipping a new host everywhere. That argues for a small
general vocabulary rather than a specific one.
Three decisions, all Jochen's, and the first is the one that unlocked it.
Exclusivity is not a property of a module. It is a property of a singular
resource the module takes over. Two shells compete for nothing and any number
may be installed; two display servers both want the seat. So a module declares
what it CLAIMS, and two modules claiming the same thing cannot both be assigned
within that claim's scope.
Not "xorg conflicts with wayland". Pairwise exclusion has a property that only
shows up later: adding a third display server means editing xorg and wayland to
know about it. Every new module requires changing modules nobody who wrote it
owns, and the edits grow as the square of the count. With a claim the third one
says what it claims and nothing else changes anywhere.
Claims have a scope -- node, site, mesh -- which is not new. The mesh already
enforces exactly one hub with a unique index. Scope is that idea said once
rather than hard-coded per case.
And some conflicts need no claim at all: two modules declaring the same file or
binding the same port are visible from what they declare. A claim is only
written for the abstract ones.
A requirement with several answers is refused, never guessed. One candidate is
assigned silently because there was no choice to make; none is refused naming
what is missing; several is refused naming them. That is what makes a solver
unnecessary -- counting candidates has no surprising behaviour, and a solver
can be added later without changing a single manifest.
Flavor is retired. It was carrying three unrelated meanings: variants of a
thing, a subset of a module a node installs, and whatever the current system
does, which earned two knowledge-base entries about going wrong. A word with
three meanings cannot be reasoned about. What it reached for is two ordinary
things -- different modules providing the same thing, and one module with a
setting.
All on the first three machines to actually run it, and all invisible from the
mesh's own state: the graph was right, the files were right, the services were
up, every node reported success, and the network did not work.
A running interface does not re-read its configuration, so a node joining left
every existing node carrying a network that no longer existed. A hub sharing a
site with a spoke was emitted twice, which WireGuard refuses. Two nodes at one
site that neither can be dialled were peered directly, so nobody opened the
path and the more specific route blackholed -- this document's own warning
arriving in its implementation. And Docker sets the FORWARD policy to DROP, so
a hub with forwarding enabled still carried nothing between its spokes.
The last one is the sharpest: the substrate at tier 1 silently breaks the
network at tier 2, and nothing in either tier's state says so.
None of these is reachable by reasoning, and each was found within minutes of a
real machine trying it. That is the argument for the lab in one line.
The store records where each resource came from and each origin removes only
its own. Verified on the scenario that caused it -- eleven resources raised,
enrolled, sent the same two-resource declaration, and the store, broker and
control plane were all still running. A later declaration dropping a resource
still removed it, so removal by omission survived the fix.
Two more faults found while fixing it, both the same shape. A report published
to a routing key nobody bound vanishes: the broker accepts it, finds no queue,
drops it, and tells the publisher nothing -- so nodes announced what they had
applied into a void. And publishReport was discarding its error, so a node that
could not tell the mesh looked exactly like one that had.
Reports are mandatory now, so an unroutable one comes back and is said out
loud, and the binding covers every key a node may publish.
Found in the lab, doing the ordinary thing: raise a first node, enrol it, send
it a declaration. Both declared resources applied correctly and every container
on the machine was removed -- the store, the broker, and the control plane that
had sent the message. The link died mid-sentence because the broker carrying it
had just been torn down by what it carried.
Nothing is behaving incorrectly. Apply removes what the store holds and the
declaration does not name, which is what reconciliation means. The fault is
that the carried bundle and mesh declarations share one store, so the host
cannot tell what this machine raised for itself before there was a mesh from
what the mesh told it to have.
It is invisible until those two meet, which happens exactly once per mesh: on
the first node, after enrolment, the moment the control plane first speaks.
The report says what is not the answer, including the tempting one -- having
the control plane send the substrate back. It cannot: it was never told what
the bundle contained, and the bundle exists precisely because there was no
control plane to ask.
Two things this record never said, both asked directly.
Connecting to the mesh is one outbound AMQP connection from the node to the
broker, held open. There is no second connection and nothing is ever dialled at
a node. Being in the mesh means that connection is up.
Two different things ride on it and conflating them is what made this murky. An
AMQP account, which the mesh issues per node at enrolment, answers whether the
connection is accepted at all -- per node rather than shared, because a shared
one lets any node consume another's queue, which is the shared-credential fault
this record exists to remove reappearing at the transport.
The node's own keypair answers which node is speaking, on every message. It is
not made redundant by the account: with only an account the control plane knows
who is speaking because the broker says so, and that is the same transitive
authority this record already refuses in the other direction. A compromised
broker could attribute reports to whichever node it liked.
So a node holds two things after enrolment -- a credential the mesh issued for
reaching the broker, and a key it generated that the mesh only sees the public
half of. Both are its own, neither reaches anything else.
I have been treating "what a node presents to prove it is that node" as an
undecided design question for weeks, and blocking on it. It was decided.
08-connectivity says of the overlay keys: each node generates its own keypair,
the private key never leaves the machine, the public key is published to the
mesh -- and says explicitly that this IS ADR 0004's "a node holds its own
identity", applied. Nobody had applied it to the thing 0004 is actually about.
What caused it was a word. The lifecycle said a joining node receives its own
durable identity, which reads as the mesh issuing something, and then the
question is what. The mesh issues nothing. A node arrives holding its identity;
what it receives is being known. That line now says what happens: it presents
the one-time secret and its own public key, which the mesh records.
The rule above it then holds literally rather than aspirationally. The mesh
stores a public key, so a copy of the mesh's database grants nothing, and
compromise of a node really is compromise of only that node.
Also recorded, since it was asked directly: same principle as SSH, own key, not
the machine's SSH host key. Host keys are regenerated by reinstalls and image
clones, which would silently un-enrol a node; their lifecycle belongs to sshd
rather than the mesh; and a partial host has no SSH daemon at all, so an
identity scheme resting on one excludes a supported kind of node.
The good half of that idea is kept: the mesh knows every node, so it can
distribute host keys the way it distributes authorised keys, and node-to-node
SSH stops depending on trust-on-first-use.
Correcting an overstatement from the previous commit, where I had written that
a node IS a conversation. It is not. A node is a machine inside the mesh, and
the session is one of the things running on it -- like the host, like any
workload.
That also dissolves the conflict I flagged as unresolved rather than needing
anyone to decide it. 0001 says a node does not authenticate to a model
provider, agents do. Still true: the session authenticates, and the session is
not the machine. The node does not think, something on the node does. I had
manufactured the contradiction by promoting a feature into an identity.
0001's summary row is corrected the same way, and says explicitly that neither
the node's session nor a hired worker makes the node itself a thinking thing --
both run on a machine, which is what leaves that line untouched.
Moving this out of 0003 and out of its vocabulary. I had spent three attempts
fitting the node's own session into the agent-as-employee record, each time
bending hired, draining, reassigned and retired to cover something none of them
describe. 0003 is back to its original text.
It belongs in 0004, under what a node is, because that is what it is -- not a
program installed on a node but part of the node. It holds one session
permanently, anything in the mesh can message it, and it remembers across
callers and across weeks. Its system prompt is the engram, which is recorded
here for the first time despite running on every node.
Also recorded: it has its own narrower tool list, so it can go and look rather
than only report about itself; there is no authorisation between nodes, because
every node is the operator's own; and how a node passes a question on is its
own business rather than a protocol field.
Switched off it still answers, and that is the point of having an off state
rather than an absent one. A node with nothing there is a silence somebody has
to diagnose. A node that says it is switched off is not. Same rule the host
follows about a service that does not exist.
0001's summary is corrected too: it had one row for "agents", which is the
conflation being complained about. Two rows now. A node's own session and a
hired worker are built from the same parts and run on entirely different terms.
Left standing and NOT resolved here: 0001 says a node does not authenticate to
a model provider, agents do. A node that holds a session does. That is a real
conflict between what is recorded and what runs, and it needs deciding rather
than a fourth reconciliation from me.