Commit Graph
503 Commits
Author SHA1 Message Date
jschoubben ba0d01788e 0062: a host may be episodic; 0060's Android gap closed
0060 named the gap and did not close it: everywhere else an init runs the
launcher at boot, and Android grants neither an init to register with nor
anything worth supervising, because a supervisor would be killed alongside what
it supervises.

Closed by narrowing what is required rather than building something. A host is
resident or episodic, and both are hosts. Being killed by the platform is
disconnection, which 0036 already made ordinary -- and every mechanism an
episodic host needs already exists because it was built for laptops that close.

A partial host can join a mesh and cannot be the first node, since every
bootstrap step is a shape it refuses. Its bundle says so.

Two consequences that are easy to miss: last-heard-from means much less on an
episodic host, so a healthy phone reads as a dead server unless the reader
knows which kind it is; and a declaration may take a long time to land, which
makes 0058's outstanding-versus-failed distinction load-bearing.

Still open, and in that order: what an Android node is FOR, and only then how
it is started.
2026-08-28 01:24:07 +02:00
jschoubben f1b1cd9aa0 Review: three ADRs no longer said what we had concluded
A sweep for claims overtaken by the last few days. Annotated rather than
rewritten, following the pattern already in 0049 -- what changed and why is the
useful part, and an accepted record should not quietly become something else.

0057's init section was wrong on all three of its claims. It said the host
needs FOUR things from an init; 0061 reduced that to one. It said every machine
the mesh targets already has systemd; Alpine does not, and it is the intended
first node. It said there is no second init to abstract over; there is now, and
the answer is still not an abstraction -- it is a four-line file per system.
What survives is the part that was always right: an init is not a dependency in
0041's sense, because it is not installed, it is what the machine already is.

0048 named Docker as the container runtime. It is now docker or podman,
detected rather than chosen -- because adoption keeps what a machine already
has, so naming one contradicted a rule already decided. That row is the only
one of the five that names two, and the record now says why.

0060 claimed the bundle is portable across operating systems. Its mechanism is;
its contents are not -- package names, unit names, service names all differ, so
an Arch host embeds an Arch bundle. That was my error, and it is the exact
confusion behind the question that found it.

The design layer had the same drift: 07 and 09 said "Docker" where they meant a
container runtime, 09 said systemd restarts the host after an upgrade when the
launcher does, and both install snippets assumed Arch. They now show Alpine and
Arch side by side, which makes the point better than prose did -- step 1
differs per system, step 2 never does.

Checked and NOT changed: 0047's "the vocabulary grows by one shape" is a claim
about the rate, not the count, and is still true. 0037 lists docker among tools
the host manages, which it does. 0041 says nothing about either.
2026-08-28 00:43:47 +02:00
jschoubben 66df0eb53e 0061: the launcher supervises; init is asked for one thing
The record said an init is asked for two things -- start at boot and restart on
exit -- which was half a change. It moved the give-up logic out of unit files
and left the restart in one, so init still decided when the host came back.

The launcher no longer execs the host. It supervises it, so restarting is ours
too, and init is asked only to run it at boot. There is an OpenRC script beside
the systemd unit now.

Records the cost honestly: not exec'ing means the launcher must trap the
shutdown signal and pass it down, because a supervisor that exits while its
child runs leaves the host to be killed rather than to stop.

And records what the implementation found: the counter counts consecutive
FAILURES, not starts. Counting starts meant a host that upgraded itself three
times rolled itself back, having worked perfectly every time -- because a clean
exit IS the upgrade path. That is now the second time a clean exit has been
mishandled, so it is called out as the thing to check.
2026-08-28 00:38:18 +02:00
jschoubben c557f99cba Record what testing podman actually showed
0060 said the container runtime was a separate decision. It is now made, and
the reasoning is worth keeping because it is the opposite answer to the same
question one paragraph earlier.

Abstracting service managers is lossy -- systemd and OpenRC are different
models and LoadState has no equivalent. Container runtimes converged on one CLI
deliberately, so almost nothing is lost: checked against podman 6.1.0, run,
rm -f and docker's own template syntax for state and labels all work unchanged.
Only the probe differs. So: a two-entry lookup, not an interface.

The difference that is NOT in the CLI is the one that would have shipped
silently. Podman accepts --restart unless-stopped, records it, and has no
daemon to act on it -- containers do not return after a reboot unless
podman-restart.service is enabled, which by default it is not. Every command
reports success and the effect does not happen.

That belongs in the declaration rather than the host: a node using podman is
told to enable the unit. Which is what made the service shape's missing 'boot'
field visible, and it is now built.
2026-08-27 23:59:05 +02:00
jschoubben e1ad39b500 Per-OS hosts, and an init asked for only start and restart
0060 -- the host is built per operating system. systemd and pacman are the Arch
host's implementation, not abstractions the mesh has to grow. They are not
independent choices: a machine has pacman because it is Arch, and the package
manager, service manager and packaging format arrive together as one decision
somebody made at install time.

Rejected abstracting them, and the reason is correctness rather than effort.
The service applier reads LoadState to tell "not installed" apart from
"stopped", which is what stops it reporting absence as success. An interface
spanning systemd and OpenRC degrades to what both express, and the lowest
common denominator is exactly where that fault lives.

Almost all of it is shared -- the vocabulary, store, apply loop, read-back
discipline, refusal model, bundle and link are portable. Two appliers differ.
And delivery was already per-OS, since a .pkg.tar.zst is an Arch artifact, so
this is the seam that already existed.

Android is the interesting case rather than Debian: no service manager, no
package installation, usually no root. Such a host implements file, directory
and action and refuses the rest -- the same refusal a host already gives an
unknown type, with a different reason. Those three are the portable floor.

The container runtime is deliberately left open: it is not an OS split, since
Arch runs docker or podman.

0061 -- the init is asked for start-at-boot and restart-on-exit, and nothing
else. Both are expressible in OpenRC, runit, s6 and an Android init.rc.
Counting failed starts and rolling back moves into a launcher, because that is
the one piece which must work when the host does not, and a script with a
counter can be tested where OnFailure= can only be hoped for. Supersedes 0059,
keeping its reasoning in full.

The checker found all six places citing 0059 and refused the commit until they
named the replacement.
2026-08-27 23:46:11 +02:00
jschoubben dcc4b8339c Say who consumes the broker and who writes the registry
Left implicit by the previous commit, which said the owning context writes
without saying what does the consuming.

The control plane is the consumer, and there is one of it. Seven contexts but
one deployable, so it is one process dispatching internally rather than seven
consumers racing -- which matters because the as-is records two consumers
accidentally sharing a queue and silently splitting the traffic, each getting
half of what it expected. With one consumer that cannot arise.

The broker is also the buffer while the control plane is down: nodes keep
publishing, messages queue, the control plane drains them on return. That is
what makes a single control plane tolerable -- an outage delays the mesh's
knowledge rather than losing it.

One consequence named because it will otherwise be discovered: an unbounded
queue grows until the broker's disk is full, and the broker is the component
every node depends on. The bound is per queue and undecided -- dropping the
oldest health report is obviously right, dropping the oldest declaration
acknowledgement is not.
2026-08-27 22:20:03 +02:00
jschoubben 19997d56c3 Approve 0057-0059, with four corrections from review
Not approved as drafted -- four things came out of checking them against each
other, and one was a bug that would have broken every upgrade.

The bug: 0059 specified Restart=on-failure while 0057 has the host restart onto
a new binary by exiting CLEANLY. on-failure does not restart a process that
exited zero, so every upgraded node would have been left stopped, having
successfully upgraded. Found by reading the two records against each other
rather than by either alone. Now Restart=always in all three places that
mention it.

The host cannot run in a container, and the reason is decisive rather than
stylistic: step 0 of the substrate bootstrap installs the container runtime, so
a host inside a container would need the thing it exists to install. It would
also break 0041 -- copy it onto a machine and run it stops being true when the
machine must already have a runtime. Everything above tier 0 is a container;
the host is not. That split is the tier boundary, not an inconsistency.

systemd is named rather than abstracted. An init is not a dependency in 0041's
sense: 0041 is about what must be installed before the host works, and an init
is not installed, it is what the machine already is. The unit file is the only
systemd-specific artefact and it belongs to the package, so a machine with a
different supervisor ships a different package.

The mesh is a watchdog, and my first draft was half an answer. Recovery must be
local -- nothing dials a node, and a host that cannot start cannot report. But
detection is the mesh's, and a local supervisor structurally cannot do it: it
sees one process failing and cannot tell a broken machine from a broken
release. Only something watching every node can, and that distinction decides
whether the response is "fix this machine" or "stop shipping this version". So
a host rollout is staged -- a few nodes, wait for heartbeats, continue or stop
on silence. Local rollback still needed, because the canary nodes break and
because a node offline during the rollout gets the declaration later with no
batch around it.

The first declaration is the overlay and nothing else. Forced, because a node's
address and peers are assigned rather than chosen. But also the way back in: a
node reachable over the overlay can be fixed by hand if a later declaration
breaks it, and a large first declaration risks a node that is broken and
unreachable at once.

Also stated plainly, because it reads as a contradiction: nodes reach each
other over the overlay and every node consumes from the broker; what 0039
forbids is an inbound CONTROL surface, not reachability.

And in 06: no node holds a credential to any control-plane store, for reads or
writes. Four ADRs already say this separately and none of them said it in one
place. Nodes state over the broker; the owning context writes. With a note that
most high-frequency writes are observability's, not the registry's -- routing
logs into the registry would be the shared-schema mistake arriving through a
door marked performance.
2026-08-27 22:12:12 +02:00
jschoubben 605c9fd441 Changes are pushed, not polled; and a stuck host rolls itself back
Two corrections and one new decision, all from Jochen catching things.

Pushed, not polled. I described updates as landing "on the next reconcile",
which reads as polling and is not the design. A declaration arrives as a
message on a link that is already open; the host applies it then. Polling over
an existing connection would be slower to land AND constant traffic to learn
nothing.

The timer is for drift and nothing else, and it cannot be replaced by an event
for a definitional reason: drift is change the mesh did not make -- somebody
edited a managed file, a distribution upgrade replaced a config -- so nothing
will ever publish a message about it. Only looking finds it.

Separated the heartbeat from the reconcile timer, which I had been conflating.
They point in opposite directions and answer different questions: the timer
looks at the machine and asks whether it still matches; the heartbeat reports
upward and is what makes silence mean something. A node with nothing to do
sends nothing, and without a heartbeat that is indistinguishable from a node
that stopped.

0059 -- a host that cannot start is rolled back by the service manager. I had
left this open on the grounds that recovery meant the host judging its own
health. That objection does not survive being asked properly: a keepalive is
something else judging the host. The watchdog must be local, because nothing
dials a node and a host that cannot start cannot report -- so it is the service
manager, which is already there.

The failure it prevents is sharper than "the node is down": a host that will
not start looks exactly like a machine somebody switched off, which is the one
condition this design has deliberately decided not to alarm on. So a bad
release reaches every node, each goes quiet, and the mesh reports a fleet of
sleeping laptops.

Confirmed means started and completed one reconcile -- deliberately not "the
link is up", or a laptop on a train would roll itself back. The rollback is a
script shipped by the package, not a host subcommand, because a binary that
will not start cannot be its own recovery. It rolls back once: a second failure
means the machine is the problem, not the binary.

Also refined the records checker, which produced a false positive: a proposed
record may extend another proposed one, because decisions are drafted in chains
and the alternative is marking things accepted to satisfy a check. An accepted
document resting on a proposed record still fails, and that was verified.

0057, 0058 and 0059 are all proposed.
2026-08-27 22:04:26 +02:00
jschoubben aeea2a9f9a Resolve the host lifecycle's open items, and say how the host is delivered
The upgrade question turned out to be a delivery question, so 0058 answers
both.

Today's third silo runs once per node and sends each one a command to install
and start. That is where the as-is records a package install that 404ed from
every mirror while the job went green, an image pull failure that did not fail
the deploy, and a verify stage that was built and never scheduled because it
was missing from a list.

The shape underneath all of those is that the thing reporting success was not
the thing doing the work. Meanwhile ADR 0037 has given every node a component
that applies state, reads back and reports -- so two mechanisms now change a
node and only one checks its work.

0058: a pipeline ends when the declaration is updated. Deploy stops sending
commands to nodes and becomes one write. The host applies it on its next
reconcile, and the host cannot report success it did not verify. The verify
stage disappears as a stage, which is the point -- verification stops being a
step that can be left off a list.

A pipeline result now means "the declaration is updated, and here is which
nodes have applied it". It does not wait for every node, because a node may be
legitimately switched off for a week. Outstanding is reported separately from
failed, since conflating them is how the old system produced a stall with no
error anywhere.

The host is delivered by exactly this path and needs no new resource type: a
`file` writes the package manager's config pointing at the mesh's repository, a
`package` names the version. Added a step I had missed -- before exiting for a
restart, the host runs the new binary once. A package can install something
that does not execute here, and that turns "the node never came back" into "the
apply failed and said why".

Six open items resolved: re-enrolment is decided when the token is issued and
revokes the previous identity; the mesh keeps a recovery copy of what each node
reports it owns, which un-strands the orphans; last-contact is reported with no
threshold, because a laptop off for three weeks is doing nothing wrong;
adoption always completes but a failed line makes a node ineligible for
assignment; a briefing is a structured document whose outcome is computed from
its lines; and the token is printed once and carried by hand, which is the
property that makes it worth anything.

Still open and named: automatic rollback of a host version that will not start.

0057 and 0058 are both proposed.
2026-08-27 21:53:38 +02:00
jschoubben 2204b01909 Design the node lifecycle end to end
The host was described as a component and never as something that runs for
years on a machine somebody else also uses. 09 covers every state a machine can
be in and every transition between them.

Four states: unmanaged, hosted, enrolled, disconnected. Only the last two are
nodes, and they are the same node in two situations. `hosted` -- the host
installed but never told which mesh it belongs to -- had no name before and is
where a machine sits between the two adoption commands.

Things that were unclear and now are not:

The first node walks the same path in an unusual order: reconcile from the
bundle, the control plane it just raised issues a token, enrol against it. Its
specialness lasts two commands. A side effect worth having -- enrolment is
exercised on node one, rather than being written and first used on node two.

Enrolment reports profile and inventory BEFORE the control plane decides
anything. The profile is the input to that decision, not a diagnostic; the
control plane cannot decide what a machine should run without knowing what it
can run.

Rebooting mid-apply is safe by construction. The store records each resource
after it worked, so a host that dies half way through comes back and applies
the rest. The rule that stops the host lying about what it did also makes it
crash-safe.

Retiring splits in two. Graceful is a final empty declaration. A node that is
gone will reconcile its last declaration forever -- the honest consequence of
making disconnection ordinary. The answer is not to make the host expire but
that the node holds nothing that outlives revocation: every grant is a per-node
credential revoked at the provider. A lost node keeps running and stops being
able to reach anything. Said plainly rather than implying the mesh can switch a
machine off, which it cannot and should not.

Losing the store is quiet and permanent, so it gets its own section. The host
re-enrols and re-applies fine; what does not come back is removal, because
resources it no longer has a record of become unowned and sit there
indefinitely.

Also corrects 0057, which said the mesh must not upgrade the host at all. That
conflated two acts. Replacing the binary is safe -- Unix keeps the running
inode. Stopping the unit is not. So the host may apply a package naming itself,
and restarts by finishing its apply and exiting cleanly, letting the supervisor
start it on the new binary. It never asks the service manager to restart it.
That makes a fleet-wide host upgrade an ordinary declaration, which the first
draft gave up on.

0057 remains proposed.
2026-08-27 21:16:44 +02:00
jschoubben 3ab11c96ef Say what the host process is: a root service, installed as a package
The design described what the host does and never what it is at runtime. The
words daemon, long-running, interval, poll and heartbeat appeared nowhere in it
or in the relevant decisions. What exists is a command that runs and exits;
what the design needs is a process holding a link. Nobody had written down that
those differ, so several questions had no answer.

0057 settles them. It runs on every node -- the host is what makes a machine
managed, so a machine without one is not a node. Root, because no useful subset
of the job is unprivileged. A systemd unit, because something must survive a
reboot to hold the link.

It never manages its own unit. The temptation is obvious and it ends with a
host stopping itself half way through an apply, leaving a machine with nothing
running to fix it. The installation owns the host; the host owns everything
else.

Installed as a package, with a tarball as the floor. The package carries the
unit file, the state directory and an upgrade path, which a bare binary does
not. But the mesh's package repository is hosted on the mesh, so any route that
needs the mesh to install the thing that joins the mesh is a circle -- the
tarball is the path that must never acquire a dependency.

Reconciles on start, on a declaration, on a timer and on reconnect. The timer
is the one easy to leave out, and without it `owned` reports what the host
applied rather than what is there -- ADR 0035 violated by omission.

The records checker caught this commit on its first attempt: 05 listed 0057 in
its frontmatter while 0057 is still proposed, and a to-be document may not rest
on an unaccepted record. The section now says so in the body instead.
2026-08-27 21:06:00 +02:00
jschoubben 2330d74c1b The host's vocabulary is complete; 05 and 07 said otherwise
All six shapes are built. 07 still said the last three did not exist, and 05
still described stage 2 as having built three of six.

Records what the lab still cannot do, because that is now the only thing
between here and an end-to-end substrate bootstrap: a sealed scenario cannot
fetch an image and its machines carry no container runtime, so package,
container and action were verified against a real machine instead.
2026-08-27 20:36:58 +02:00
jschoubben 03874f3fe2 Add a structural check over HQ's own records
Nothing in this repository was verified by anything but reading, which is how
a superseded decision stayed live in the constitution and in the to-be README
at the same time. Both were found by a person looking, and nothing stopped a
third.

Five checks: links resolve; `decisions:`/`extends:` name records that exist and
are accepted; a governing document citing a superseded record must name its
replacement in the same paragraph; supersession is symmetric; filename number
matches heading number.

Each was made to fail before it was made to pass. The live-citation check was
verified against a reconstruction of the actual incident -- the to-be README
citing ADR 0017 as live guidance -- and reports it with file and line.

It found one thing nobody had noticed: ADR 0018 never declared that it
superseded 0011, though 0011 has named 0018 as its superseder since August.
Fixed.

Deliberately not checked, and said so in the README: 02-DECISIONS and
01-RESEARCH may cite superseded records freely, because a decision record
discusses history and research records what was observed. 00-as-is may rest on
one, per 0056. Flagging those would put noise on correct documents, and a check
that cries wolf gets suppressed -- which costs more than not having it.

Two bugs found by running it: the frontmatter reader iterated an inline list as
characters, and the as-is exemption was missing entirely.
2026-08-27 20:18:35 +02:00
jschoubben e1f4c7d9e0 Approve 0054-0056, apply them, and fix the two smaller findings
0003 is now superseded by 0056. Nothing is left proposed.

Applied:
- 06 corrected from ten contexts to seven plus the api, each row now stating
  why it passes the more-than-one-node test. work, knowledge and stream are
  named as mesh-hosted rather than dropped; `ai` folds into config; `record`
  is deferred explicitly rather than listed. Its frontmatter now cites 0055.
- how-we-build §4 amended per 0054, and the derived page republished by
  playbook 05.

The sync found the drift the playbook exists to catch: the published §4 and
the source did not say the same thing. The source said "four accidents, not
four boundaries"; the published page said "one intent expressed four times",
and only the published page carried the scope caveat. Same rule, two texts,
already diverging. Verified the republish by reading back -- the new rule is
present and the old section's body returns nothing -- rather than trusting the
success message.

The two smaller findings:
- 0051 separated the transport identity from the declaring authority. It said
  the token carries "an address" and "the identity to expect" without saying
  what the node dials. It dials the broker, so pinning only that would make the
  control plane's authority transitive and let a compromised broker forge
  declarations -- which, since the host applies whatever the link delivers, is
  the whole machine. The token now carries four things, and declarations are
  signed and verified per declaration. Cost recorded: rotating the signing
  identity is fleet-wide.
- 0026 no longer restates 0022's rule about generated views. 0022's own words
  are "prose does not restate status; one place, and two is one too many",
  which is what 0026 was doing to it.
2026-08-27 02:21:34 +02:00
jschoubben f49d177a31 Draft three records for the contradictions the review found
0054 -- things that change together share an authority, not a package.
The constitution instructs agents to group "how a node is reachable" into one
module, citing superseded ADR 0017; ADR 0044 says there is no networking thing
to install. Since the constitution is injected where work is decided, the
superseded rule is the one actually steering work. The observation behind it
was right -- research 005 measured that reachability is the only place modules
genuinely change together -- but the conclusion was wrong: tight coupling means
a shared authority, not one artifact. wireguard and traefik deploy to different
node sets, so the merged module would be assigned where half is unwanted.
Requires amending how-we-build and republishing the derived page.

0055 -- the control plane is the node-coordinating contexts.
Three context lists were in circulation (0015 says nine, 06 says ten, the
README said eight) and none was decided. Research 006 said explicitly that the
change "belongs in a new record -- not written here", and the design used the
list anyway. Reconciling them shows `stream` and `ai` were dropped with no
reasoning at all. Applying 06's own test -- needs to know about more than one
node -- gives seven contexts plus the api, with work, knowledge and stream as
hosted applications and `ai` folded into config as an ordinary grant. The
record defers rather than lists.

The cost is stated rather than reassured away: a board composing across the
boundary reads more than one interface. That was raised before as "only moves
the problem up a layer", and the answer is that 0045 already requires surfaces
to read interfaces rather than stores -- what changes is the count, not the
kind of work.

0056 -- the authority is the control plane, not a database.
Every clause of 0003 has been decided against in four separate records and it
is still accepted and cited as live. The error underneath is the same category
error 0054 corrects: "source of truth" named a storage location when it meant
an authority, and once the store is the answer, shared schemas follow. The
half that was right -- the repository defines what exists, the mesh defines
what runs where -- survives untouched. Best consequence: the cache mode
disappears, so a node that has not heard from the mesh is no longer
indistinguishable from one that has.

0003 is left accepted until 0056 is.
2026-08-27 01:35:37 +02:00
jschoubben ef5dd0751b Approve 0049-0053; drop a to-be item superseded by ADR 0044
The 'domain grouping' item cited ADR 0017 as live guidance. 0044 superseded
it -- there is no domain module to group into, so there is no domain list to
settle.
2026-08-27 01:00:29 +02:00
jschoubben ccbbfa9c8a One node runs the control plane, and nothing takes over
Closes the two open questions in 06 and 08, which turned out to be one
question: how many control planes run, and what happens when the hub is down.
Both were drifting toward redundancy by default -- a standby plane, a second
hub, an election to pick between them. That is not one feature but a property
every layer must then honour, and each layer gets it wrong independently.

Not wanted, and not needed. A handful of machines with one node hosting the
registry is not a distributed system.

The argument for why this is sound rather than merely cheap is that the design
already tolerates it by construction. ADR 0036 makes reachability state rather
than class; the host reconciles from its own store (0043) and never needed to
ask anybody to hold the state it was last given. So the control plane being
down is not a new failure mode -- it is every node in the ordinary disconnected
situation at once. What is lost is change, not operation.

No node holds a contended role: the control plane is assigned like any other
module, and the overlay hub is declared (0050). No promotion, no quorum, no
fencing, no split brain, no replicated store, and no "which node is
authoritative" recurring at every layer.

Two consequences stated plainly rather than buried. The control-plane node is a
single point of failure -- deliberate, and said out loud so it stays
deliberate. And recovery is restore rather than failover, which makes backup
the availability story rather than hygiene.

The sharpest one is the clock: the control plane owns certificate issuance
(0049), so an outage outlasting a renewal window expires every public name.
That bounds how long recovery may take, and nothing measures it today.
2026-08-27 00:55:10 +02:00
jschoubben 4e80820e2f Design connectivity in full: overlay, resolution, exposure, filtering, certificates
Written as one document because the five are one design. They share inputs,
they must agree, and every one of them today is computed in a different place
by a different module from a different copy of the same facts.

The through-line is that none of the five can be answered by a machine alone,
so all five are decided centrally and delivered as `file` resources. That costs
no new host vocabulary and removes both remaining direct database connections
from nodes -- wireguard and traefik are the only two, and both are connectivity.

Three decisions fall out, all proposed:

0050 -- reachability is declared, not inferred from an address. The RFC1918
regex is wrong for carrier-grade NAT (100.64/10 tests as public, so an endpoint
is written to an address nothing can reach), wrong for IPv6, and wrong for a
routable address behind a closed firewall. The lab needing TEST-NET-3 to
satisfy the regex is the same bug from the other side. Also kills hub election
by address prefix, which fails silently and makes renumbering an outage.

0051 -- the enrolment token carries where the mesh is and how to recognise it.
Closes two circles with one mechanism: verifying the mesh needed the CA, and
obtaining the CA meant trusting whoever handed it over; and a node had to reach
the mesh before it could resolve any mesh name. An address plus a fingerprint,
carried out of band, resolves both -- and closes the CA question 0049 deferred.

0052 -- a filter rule names its source. `scope:` is declared in five manifests,
is part of no rule type, and is referenced by no code, so those manifests
appear to restrict ports and restrict nothing. Removed rather than implemented;
the general fix is refusing unknown keys, which the host already does and
manifests do not.

Also corrects two claims in 0049 asserting wireguard was already handled.
Research 006 says both modules still reach upward; neither is.
2026-08-27 00:36:49 +02:00
jschoubben 8d9282d86b Resolve the ingress gap: a route is a grant
ADR 0048 named ingress as an unclosed hole -- nothing said what terminates
TLS, how a public name reaches a container, or which tier owned it. Resolving
it needed no new concepts, which is why it survived: nobody had applied the
rules already written to it.

Ingress is not substrate. The control plane does not need a route to start,
and no node needs one to reach it -- the node dials out and has no listening
control surface. It grants itself a route afterwards, like a bucket.

A route is an instantiation edge under ADR 0044. The direction mirrors a
database -- the consumer supplies a target and receives a name rather than
credentials -- but it is the same edge.

The substantive finding is that exposure is three facts at two scopes: name
resolution and certificate issuance need to know which node is publicly
reachable, and only the proxy mapping is a single machine's business. That is
why it belongs to the connectivity context, and why Traefik doing all three on
the node is wrong.

Which matters beyond tidiness: research 006 counted traefik as one of two
modules opening a direct Postgres connection, reading nodes and mesh_ca. That
violates 0037, 0045 and 0039 at once, and is why every node permanently holds
a credential to the control plane's database. Deriving the config centrally and
delivering it as `file` resources removes it, costs zero new host vocabulary,
and closes the set 0039 identified -- wireguard was the other.

Left open deliberately: the mesh's internal CA is the other thing traefik
reads, and it belongs to the link's mutual authority, not to exposure.
Conflating the two is what made the gap hard to see.

Also fixes an inconsistency from the previous commit: 06 still claimed the
virtual host was raised from the bundle.

Proposed, not accepted -- for review.
2026-08-27 00:22:52 +02:00
jschoubben 4d19e93900 Name the substrate's actual products
The design layer described every service by role and never once by name:
Postgres appeared in zero design documents. That was over-application of the
research rule "never identify the mesh it observed", which is about node names
and domains, not software.

Two things were actually broken by it. substrate.lock pins images by digest and
a digest belongs to a named image, so the bundle could not be written from the
design. And a reader could not tell a settled choice from an unexamined one --
"a relational store" reads identically either way.

ADR 0048 names them: PostgreSQL, LavinMQ, MinIO, an OCI registry, Docker. The
argument for each is continuity, which is a real argument -- replacing a
substrate service migrates the mesh's own state. Role and product are now both
written, because the design depends on the protocol while the installer needs
the product.

Also separates two questions the substrate doc had merged: being substrate and
being in the bundle. Only Postgres must precede the control plane; the rest are
substrate by role and ordinary by delivery. Whether the bus joins it is left
open, because it turns on the control plane's internal shape.

Names the forge as Gitea, and records ingress/Traefik as an unclosed gap rather
than a naming one -- nothing says what terminates TLS or which tier owns it.

Fixes a miscount: the host's bootstrap vocabulary is six shapes, not five.
2026-08-27 00:11:38 +02:00
jschoubben c631cbd07c The bootstrap starts a step earlier than recorded
Asked whether postgres has to be installed, and the answer exposed a missing
step. The store is a container, so something must run containers before anything
else happens — and a container runtime is a PACKAGE, not a container.

Step 0 is where several threads meet. It is what the host's capability detection
already reports, and the first use of that report by something other than a
person. It is adopted rather than installed when the machine already has a
runtime with configuration somebody chose. And it is a package, needing the
machine's own package manager and a network, both of which ADR 0046 permits.

So the host's bootstrap vocabulary is six shapes: package, container, file,
directory, service, action. Stage 2 built three of them.

The node host design now names which three remain and why the lab cannot yet
exercise them — a sealed scenario fetches nothing and its machines carry no
container runtime, which is lab-installation work rather than a constraint on
the design, because production machines have a network.
2026-08-26 23:58:19 +02:00
jschoubben 93470f6162 ADR 0047 — the bundle may carry actions the link may not
The bootstrap's sharpest open question, and the framing was wrong. "State on
this machine" was being read as the filesystem and the service manager. A
service running on this machine IS part of this machine — writing a file and
creating a database in a local store differ in mechanism, not in scope.

The real question was underneath: must the host learn what a database is? It
must not. Giving it a `database` resource type means tier 0 knows Postgres, then
a bucket, then a virtual host — the host acquiring the substrate's vocabulary
one service at a time, which is what ADR 0037 exists to stop.

So the bundle declares an ACTION and the host runs it and verifies it. What a
database means stays with the module that provides one; the host knows only how
to run a declared action against something local and check the result. Its
vocabulary grows by one shape rather than by one resource type per service.

Actions are permitted in the bundle and forbidden over the link, and the
asymmetry is deliberate. A bundle arrives WITH the binary: anyone able to put a
hostile action in it could equally have put it in the host itself, so refusing
actions there buys nothing and costs the bootstrap. The link is a separate
party, reachable separately, and an action there is the unbounded blast radius
ADR 0039 refuses. That decision stands unchanged.

And ongoing provisioning is not the host's at all — the control plane does it
once a mesh exists — so the asymmetry costs nothing.

Which dissolves the earlier worry about one mechanism with a tier boundary
inside it: there are two mechanisms, with different actors, scopes and trust
models, and that is the answer rather than a compromise.

Named rather than hidden: this is the escape hatch research 011 warned about,
arbitrary code in the place hardest to remove later. It is bounded by being
bundle-only and by every action having to declare how it verifies itself, and
that boundary is the whole defence.
2026-08-26 23:56:39 +02:00
jschoubben 5b3d0ebd4f ADR 0046 — the installer fetches what it pins
The blocking question was where a container image comes from, and the version
that blocked assumed the machine might have no network. That assumption came
from the LAB: a scenario is a closed address space by design, which is what lets
two scenarios hold the same addresses without meeting. Production is not sealed
— a machine being adopted has a network, and one that does not is a machine
where very little works anyway.

So substrate.lock carries references, not payload: an image name and a digest,
fetched at apply time. A first node pulls from upstream because no mesh registry
exists yet; every node after that pulls from the mesh's own. The lab is the
exception and places images itself, the way it already places the host binary —
a property of a test environment, and letting it dictate the production design
would be the tail wagging the dog.

Pinned by DIGEST rather than tag. Reproducibility comes from pinning the
identity of a thing, not from carrying its bytes, which is what makes fetching
acceptable rather than a compromise.

ADR 0041 survives untouched, which was the point. "Copy it onto a machine and
run it" stays literally true — one binary, a few megabytes, which then fetches
what it was told to. Carrying images would have quietly redefined the property
that decision rests on.

Costs accepted and named: an apply can now fail because something is
unreachable, which a self-contained artifact could not, so it must fail legibly
— naming what it could not fetch and from where. And the lab needs a way to
place images into a machine that also has no container runtime, both of which
are lab-installation concerns and neither solved here.

Research 012's build-time-versus-apply-time reframing narrows accordingly: it
still holds for what a tailored installer contains, and no longer has to hold
for images.
2026-08-26 23:52:46 +02:00
jschoubben 0531d6fc38 ADRs 0044 and 0045 — the module design, closed; 011 graduates
011 opened asking what a graph deletes and found the graph already existed. The
work became design, worked through twenty cases and one provider in full. Two
decisions close it.

0044 — a module declares presence, instantiation and exclusion. Two kinds of
edge because a game wanting a database is not a game wanting postgres to exist:
one creates something per consumer, carries credentials back, can be revoked,
and leaves the provider holding state. Names are concrete unless providers are
genuinely substitutable — `terminal` passes, `database` fails, and the adapter
is what creates an interface. Where there is no contract there is a tag, which
describes and does not bind. Exclusion is a third relation and is not derivable.
A node provides names too, which makes capability checking stop being a separate
mechanism and makes the host's detection an input to resolution. Constraints,
never placement. Scope decides which provider and the binding is written down
and sticky, in a place that follows the scope. And there are three entities, not
two — the assignment carries what belongs to neither end, which is what
node-agnostic modules ran out of.

0045 — a context owns its store, exclusively. No shared writes and no read roles
on another context's store, because reading couples you to its layout just as
firmly and invisibly. The unit is the CONTEXT, not the process: a board showing
the mesh's own data is the mesh showing its own data. Asking or subscribing is
derived from ADR 0036 rather than chosen. And it is the first clear list of what
the design removes: grant kinds, table ownership, cross-context migration
ordering, and a class of permission modelling.

0017 is superseded rather than narrowed — its text unchanged, its status
changed. Folders assert relationships where edges record them, and the domain
module goes with it.

Left explicitly undecided in both: what a resolver delegates rather than
reimplements, how many instances a module should have, and what a provider hands
back.
2026-08-26 23:41:18 +02:00
jschoubben 60aea14935 Define the substrate, and answer 006's four-or-five conditionally
Same gap as the control plane: load-bearing and unpinned.

The substrate is what the control plane CONSUMES AND CANNOT GRANT ITSELF. Every
module needing a database asks provisioning for one; the control plane needs one
too and cannot ask itself, because it is not running yet. That circularity is
not an awkwardness to work around — it is the definition, and anything on the
wrong side of it must be raised by the bundle the host carries.

Which answers 006's open question in the honest form rather than with a number.
The identity provider is substrate only if the control plane DELEGATES
authentication — then it cannot serve anybody before the provider exists and
cannot grant itself a client. If it authenticates natively, the provider is an
ordinary hosted service. So the count follows from a decision not yet taken, and
asserting four was asserting that decision.

The test also rules out the tempting wrong answer: an identity provider, a mail
server and an analytics service are all infrastructure by any ordinary reading,
and none are substrate, because the control plane starts and runs without them.
Important is not the test.

Records why the bundle is pinned by hand — it is applied when no mesh exists, so
nothing can resolve a version or ask a registry — and why it must be
self-contained, which makes it an artifact built on a machine with a network for
a machine that may have none.
2026-08-26 23:39:11 +02:00
jschoubben 148395ca54 Define the control plane, which was used 79 times and defined nowhere
Nineteen files, seventy-nine mentions, no definition. That is how-we-build §5
failing on this repository's own vocabulary — ubiquitous language is checked,
not assumed.

The definition, and it is not arbitrary: the control plane is everything that
needs to know about MORE THAN ONE NODE. It follows from ADR 0037, which has the
host applying rather than deciding precisely because deciding needs knowledge
the machine does not have. So the line falls exactly there — writing a file is
the host's, choosing which nodes run the store is the control plane's, and
anything a single machine could answer alone does not belong here at all.

That last consequence is worth having: putting a single-machine concern in tier
2 is a mistake the tier rule will NOT catch, because the dependency direction
stays correct.

Also states what it is not — not the thing that changes machines, not a surface,
not the substrate, and not privileged on a node beyond what the declaration
vocabulary allows. And the property that makes tier 2 unlike the others: it is
itself a consumer, with the same requirements as any module, which is the
circularity the bundle exists to resolve rather than hide.

Scoped deliberately: this defines the term and does not design the contexts
inside it. Ten is the skeleton's claim rather than a settled list, and research
006 still asks whether the record belongs here or in the substrate.
2026-08-26 23:33:20 +02:00
jschoubben a4ab3e15c2 011: one interface, many contexts — and the constraint that hides in it
The objection is right: if every context runs its own service with its own
interface, the board is coupled to N interfaces instead of N schemas, something
has to compose them, and composition is logic — which tier 3 says a surface does
not hold. That moves the problem up a layer rather than solving it.

The skeleton already answers it, and the previous entry talked past it. `work`
and `knowledge` are not separate services; they are contexts INSIDE the control
plane, alongside the record, inventory and delivery — and `api` is listed there
as the one interface every surface speaks to.

So the board speaks to one interface. Behind it the contexts stay separate,
integrating through the record, but they are one tier, one repository, one
deployable — and coupling within a tier is not what the tier rule forbids. The
problem does move up a layer, and the layer it moves to already exists and has
this as its job.

The caveat is load-bearing and now recorded as an open question: this holds only
while the contexts are not separate deployables. The moment one becomes its own
service with its own interface, the board is back to N clients, something must
compose them, and the composition has nowhere to live that tier 3 permits. That
is a real constraint on how far the control plane may be split, and it is worth
knowing before splitting rather than after.
2026-08-26 23:29:56 +02:00
jschoubben a4a25ca7e3 011: one surface over several contexts is normal
The board visualises the mesh, the work engine, the knowledge base and more, and
the alternative — a web application per context — is worse for everyone using
it. Composing several sources into one view is what a surface IS, so this is not
a compromise with the ownership rule.

What changes is only where it reads from: each context's interface rather than
each context's store. Most of that already exists — 56 of 126 modules carry a
tool surface, more than carry a service.

And the unified board is what keeps those interfaces honest. A view that cannot
be built from a context's interface proves the interface inadequate, discovered
where it is cheap to notice rather than the first time something else needs the
same data and quietly reaches for the store instead.

If composing many calls proves too slow, the answer is a projection the board
owns and keeps current from events, not access to somebody else's tables.
2026-08-26 23:27:42 +02:00
jschoubben fa62c7f0e4 011: correct the rule — contexts, not processes
An earlier version argued a dashboard reading a dozen stores was caught by
exclusive ownership, because a dashboard is a surface and surfaces speak to an
interface. Wrong, and it drew the line in the wrong place.

The mesh's own board showing nodes, modules and deployments is not a separate
context reaching across a boundary — it is the mesh showing its own data.
Requiring it to go through an interface to reach facts its own context owns is
ceremony.

The rule is that a CONTEXT is granted what it exclusively owns. Everything
inside it — service, surface, tools — reads that store freely. What is forbidden
is a different context reading it.

Which is what the consumer count already showed: the problem was never surfaces,
it was three other contexts keeping their tables in the mesh's database.
2026-08-26 23:25:46 +02:00
jschoubben e71d532c2e 011: request or subscription is derived, not chosen
Asked what the distinction actually is, and the SQL half needed correcting
first: under exclusive ownership SQL runs against your own database and nothing
else, whatever transport a query might travel over. Both options are the mesh's
own channel and both ride the broker, so the transport is not the distinction.

The distinction is where the answer lives when you need it. A request asks at
the moment and waits — always current, costs a round trip, cannot answer when
the other side is down. A subscription keeps a local copy — instant, works
offline, as current as the last event received, and you must handle what you
missed.

What decides is not taste. ADR 0036 makes disconnection an ordinary situation
rather than an exception, so anything that must keep working while disconnected
CANNOT use a request: there is nobody to ask. And the converse — anything where
a stale answer is worse than no answer cannot use a subscription. A display can
lag; a decision about whether a grant is still valid cannot.

So an apparently open question turns out to be derived from a decision already
taken. What stays open is narrower: what a consumer does about the events it
missed while disconnected — replay from a point, ask once for a full picture and
resume, or rebuild. The question every projection has.

Also recorded: separate databases are required in the new design, and the shared
registry is a leftover rather than a pattern.
2026-08-26 23:23:39 +02:00
jschoubben 7e83723b7b 011: rewrite the question table, which had gone stale silently
Several edits to the overview matched nothing and returned success, so the
question table still carried answers superseded two or three exchanges ago —
"when two modules provide one name, who chooses" was still open in the table
while answered in the file it pointed at, and nothing recorded the instantiation
edge, instance counts, grants, bootstrap provisioning, the tool audience, or the
registry consumer check.

That is the fault this repository catalogues, committed by the thing cataloguing
it: a string replacement that found no match, reported nothing, and left the
document claiming a state it did not have. Rewritten from what the documents
actually say rather than patched again.

Nine questions settled, fourteen live, and the split is now visible instead of
implied.
2026-08-26 23:17:25 +02:00
jschoubben afcc355744 011: checked the registry's real consumers, and the question was the wrong shape
The exclusive-ownership rule turned on whether every reader of the mesh registry
could be served another way. Eighteen consumers open a direct connection. Four
groups, and only one is work.

The owner and its machinery keep reading, because they own it. The node appliers
are already resolved — ADR 0037 stops the host querying the mesh database,
decided for tier reasons with nothing to do with this.

The bulk are FOREIGN TENANTS. The work engine holds ten of its own tables in the
registry's database, the knowledge base two, pipeline logs one. Thirteen foreign
tables across three contexts, which is how-we-build §4's shared schema counted.

So the question was the wrong shape: the problem is not readers needing a new
route to data, it is tenants needing to move out. Tasks, agents and teams have
nothing to do with nodes and modules and are co-located by history. Give that
context its own database and its dependency on the registry shrinks to one
table.

A handful of genuine cross-context reads remain, small enough to enumerate
rather than estimate. The rule holds.

Left open: whether those reads want an interface or events. Asking which nodes
exist at the moment you need to know is a request; reacting when a node appears
is a subscription, and some consumers want both.
2026-08-26 23:16:13 +02:00
jschoubben 6b1aab6a1e 011: the dashboard case, and why exclusive ownership is the tier rule
Raised as the hardest test of the rule: a board showing nodes, modules,
pipelines, agents and tasks wants to read a dozen stores, and under exclusive
ownership it can read none of them.

It survives, and not by luck. The board is a SURFACE, and surfaces already may
not do this — the skeleton puts `api/` in the control plane as the one interface
every surface speaks to, and tier 3 as thin, no logic. A board reading stores
directly is a surface reaching past the context that owns the data, which the
tier rule forbids for reasons that have nothing to do with databases.

So it is not a counter-example; it is an instance the rule catches. And the two
rules turn out to be one rule seen from two sides: exclusive ownership is the
tier rule expressed in terms of storage.

The general shape for anything needing to see across many things: consume the
record and own your own view. A reporting context builds a projection from
events and reads its own store, never anybody else's.

The cost said plainly rather than buried: a projection is more work than a join,
and it lags. A board queries the mesh's own database directly today — ordinary,
working — and this rule makes that a migration rather than a preference. The
reason to pay it is §4's already-measured cost, not elegance.
2026-08-26 23:10:36 +02:00
jschoubben aa767d17a8 011: a module is granted only what it exclusively owns
Reconsidered by the operator — maybe shared databases should not be allowed at
all — and the stricter version is better and goes further than the schemas it
replaces.

No shared writes, and no read-only role on another module's database either.
Reading another context's tables couples you to its layout exactly as firmly as
writing them does, and the coupling is harder to see because nothing breaks
until the owner changes a column.

That is how-we-build §4 taken at its word rather than at its letter. The
permissive version — a per-consumer schema, revocable, with cross-context joins
possible but deliberate — kept the letter and left the temptation. A boundary
that is merely inconvenient to cross is a boundary that gets crossed.

The cost is cross-module reporting, and it is the point rather than a
regrettable side effect: anything wanting to know what several modules hold
consumes their events or calls their interface. That is §4's whole argument, and
the mesh already has both mechanisms. What gets harder is precisely the thing
that was making work belonging to one context keep having to be implemented in
another.

And it is the first clear instance of what this effort has been hunting — what
the design DELETES rather than adds. Grant kinds collapse to one: an exclusive
resource. With them go the question of who owns which table, the guessing at
revocation time, cross-module migration ordering, and a class of permission
modelling a shared store would otherwise need.

One thing it does not answer, recorded because it could make the rule
unworkable: the mesh's own registry is read directly by many things today, and
under this rule they consume events or call tools instead. Achievable in
principle. Whether EVERY current consumer can be served that way is unchecked,
and should be before this becomes a decision.
2026-08-26 23:09:44 +02:00
jschoubben 4c8515507a 011: where a binding lives follows the scope, and a grant is not always a whole resource
Two questions asked directly, and the second collides with a rule in force.

One module on two nodes sharing a database corrects something stated flatly: the
binding is not "recorded on the assignment". Where it is written down FOLLOWS
THE SCOPE. A shared grant belongs to the module and every assignment references
the same one — which is the answer for two nodes wanting one database between
them. A per-instance grant belongs to the assignment. Same relation, two homes,
and which home is what makes two instances share something or not.

Several modules adding their own tables to one database is three needs wearing
one sentence, and a provider offers KINDS of grant rather than one: a database
for a consumer whose tables are nobody else's business, a read-only role for one
that needs to see what another holds, and a SCHEMA within a shared database for
the case actually asked about.

Loose tables in a shared database is what how-we-build §4 warns against in as
many words — several domains sharing one forty-five-table schema, which is why
work belonging to one context keeps having to be implemented in another. Not a
style objection; the observed cost, already paid.

A per-consumer schema keeps what the request wants and drops what §4 objects to.
Same database, same connection, same backup, and a cross-schema read remains
physically possible when genuinely needed. What it adds is ownership: migrations
touch one namespace, two modules cannot collide over a table name, and revoking
drops the schema rather than guessing which tables belonged to whom.

So the fault §4 names is still possible and no longer accidental — a
cross-context join becomes something somebody deliberately writes rather than
the path of least resistance. And revocation becomes answerable, which the
whole-database version never was.
2026-08-26 23:08:32 +02:00
jschoubben fa9889536c 011: tools have a different audience, migrations cross the edge, provisioning is early
Three additions, and the third kills an assumption.

Tools are the most common content in the catalogue — 56 of 126 modules, more
than carry a service — and they survive the split without fitting either half. A
tool is not an artifact and not node state; it is a contract the mesh publishes
on a module's behalf, and what consumes it is an AGENT rather than another
module. That is a second audience the design has not described. Whether it is
one relation with two audiences or two relations is cheap to decide now and
expensive later.

A migration belongs to the CONSUMER and runs on the PROVIDER. A game's
migrations run against the database the store granted it: owned by the consumer,
hosted inside something it does not control, ordered after the provisioning edge
because there is nothing to migrate until the grant exists, and scoped to that
grant. Ownership crosses the edge, which nothing in provides and requires
expresses — and it gives a consumer's own install an internal order, provisioned
then migrated then started, that depends on an edge rather than on its contents.

And provisioning is EARLY, not late. The assumption worth killing is that it is
something the control plane does for consumers once a mesh is running. The
mesh's own registry database is provisioned before there is a mesh, and so is
its virtual host on the broker: the store runs from the carried bundle, a
database is created in it, the mesh's own schema is applied, and only then does
a control plane exist. Steps two and three happen before there is a mesh to do
them, so provisioning is part of the bootstrap and part of what the bundle has
to express.

Which strains ADR 0043. The host applies declared state ON THIS MACHINE, and a
database inside a running store is not a file or a unit. At bootstrap it is at
least local — the store is on the same machine. Afterwards a consumer on one
node provisioned from a store on another is the ordinary case and reaching it is
not the host's job. The same operation is local at bootstrap and remote later,
which is either two mechanisms or one with a tier boundary crossing inside it.
Currently the sharpest unresolved thing in the effort.
2026-08-26 23:02:52 +02:00
jschoubben a3c7e7e1f1 011: providing is a facet, and the assignment is a third thing
Any hosted service can be a factory — an identity provider grants clients, an
analytics service grants a tracking identity, a mail server grants mailboxes, an
application platform grants a project that is several of those at once.
Providing is a FACET a module may have, not a kind of module it is, which is the
same conclusion this effort reached about services and applications arriving
from the other direction. So `provider` stops being a category too.

Two relational stores from different vendors both grant "a database" and are the
sharpest possible test of the substitutability rule. They fail it completely —
different protocol, dialect, driver, client library compiled into the consumer —
so `database` stays a tag, now with two real providers rather than a thought
experiment.

The assignment is a third entity, recorded because the operator tried the
alternative: modules were once node-agnostic and it did not survive. Several of
a provider's properties belong to neither end — where its state lives, how it is
reached, tuning derived from the machine's hardware, which instance serves a
given consumer. Not the catalogue, because they differ per node; not the node,
because they are about this module. A design with only modules and nodes has
nowhere to put them, which is what node-agnostic ran out of. The current system
already stores environment values per module AND per node, arriving the same
way.

Which answers the question asked directly: two nodes both run a store, so which
serves a consumer? Neither obvious answer. Not the consumer naming a node — that
is placement in the consumer's manifest, a game edited because a database moved.
Not the consumer not caring — for presence it genuinely does not, for
instantiation it cares permanently.

What the consumer knows is the SCOPE of its own need: one instance shared across
every instance of itself, or one each. That decides, and needs no node named.
Then the mesh binds, and the binding is recorded on the assignment and is
sticky — a resolver that re-derives which store serves a consumer will one day
derive a different answer and relocate a database.
2026-08-26 22:56:36 +02:00
jschoubben 13c6068874 011: the provider shape generalises, and two things differ inside it
The broker has all nine properties the store has. So do the object store and the
image registry. A substrate service is a SERVICE PLUS A FACTORY, there are four
of them, and the pattern generalises past the substrate: anything granting
something per consumer has this shape.

Two differences matter more than the similarity.

The broker cannot be managed over the broker. ADR 0001 makes it the channel
every node takes work from and ADR 0039 makes it the security boundary, so the
module providing it is also the way modules are managed — a declaration cannot
be delivered to it over itself. Nothing else has that property; the store is
consumed by the control plane but is not how the control plane REACHES anything.
This is what the carried bundle exists for: the broker is raised from what the
host carries because there is no other way to raise it. A constraint on one
module, not a general rule, and a schema with no way to say so hides it.

And two modules of identical shape want opposite instance counts. The broker is
one per mesh by decision. The store cannot be, because a node that must keep
working while disconnected cannot depend on a database elsewhere. Which settles
what cases.md left open: how many instances is NOT derivable from what a module
is. It is a per-module decision, it has to be declared, and nothing in provides,
requires or excludes says it.

Revocation differs in consequence too. Dropping a database leaves data until
something removes it — a leak, recoverable. Dropping a virtual host loses
whatever was undelivered — silent, and not. Same relation, different blast
radius, which argues for the provider deciding what revocation means rather than
the mesh applying one rule.

File renamed: it was never really about postgres.
2026-08-26 22:54:49 +02:00
jschoubben f160b28a71 011: postgres worked through, and "one kind of edge" was wrong
The tidy version said a module provides names and requires names and that is the
only edge. Working postgres through completely disproves it.

A small game wanting to store data does not require postgres to EXIST. It
requires postgres to MAKE IT A DATABASE and hand back credentials. Those are
different relations in every way that matters: one creates something per
consumer, carries a payload back, can be revoked, and leaves the provider
holding state about who was granted what. The other creates nothing.

So: two kinds of edge, one graph. Instantiation implies presence; presence does
not imply instantiation. The current system already had exactly this split —
`dependencies` for presence, `requires: provision:` for instantiation, with the
resolver deriving one from the other. analysis.md called that derivation a
convenience. It is not: it is the correct relationship between two genuinely
different relations, and the design had collapsed them.

Postgres also turns out to be nine things, not one. A container. Persistent
state where moving nodes is a migration rather than a reschedule. Configuration
partly derived from the machine's hardware. A tool surface. A provisioner. Its
own bookkeeping about what it granted, which is not the data it stores. An
exposure decision per node it runs on. Credentials it generates, which means a
provisioning edge carries a secret. And health that is not "the container is up".

Four questions the worked example makes concrete rather than abstract. WHICH
postgres, when there are two — a consumer of `terminal` does not care and a
consumer of a database cares permanently. How many instances a module should
have, which cannot be a global rule because one-per-mesh is wrong for a store a
disconnected node needs and one-per-node is wrong for the mesh's own registry.
What happens to a grant when its consumer is removed, where dropping is data
loss and keeping is a leak. And whether a declaration is composed PER NODE from
what that node reported — because tuning follows hardware the control plane
cannot know, and the alternative is the host deciding, which ADR 0037 forbids.
2026-08-26 22:53:59 +02:00
jschoubben c9c2dfe686 011: what a feature is, and what it splits into
The operator wants features gone, and 006 left it open. Measured, and the answer
is that nothing replaces them because they were never one concept.

A feature is a kind of content a module carries, detected from its directory:
twenty-one of them, each with a handler owning six stages — build, publish,
install, configure, start, verify.

The structural finding: EVERY handler implements EVERY stage. `configs` writes
files onto a node, has nothing to build, and has a build stage. `npm` publishes
to a registry, has nothing to start, and has a start stage. One interface spans
build-time and apply-time, so every kind of content must implement both halves
and most do nothing in one — and a stage that does nothing looks exactly like a
stage that failed to do anything.

They split four ways, across three tiers. Artifacts built once per version and
published, where no node is involved — delivery. Resources that are desired
state on a machine, which is what ADR 0043 already describes and the host already
does — tier 0. Actions run once against something that is not this machine, like
a migration against a database on another node — delivery, and seeds go
entirely. And checks: the prerequisites are REQUIREMENTS IN DISGUISE, a module
saying what must be true before it can be installed, which is what an edge in
the graph says; the verifiers are the read-back the host already performs.

So `feature` is one word for four things spanning three tiers, which is why the
pipeline is hard to reason about.

One property must survive the split, and it is the thing the current design got
right: content is DETECTED, relationships are DECLARED. A module that says it
has migrations and has none is a fault nobody sees until it matters — but what
it requires and provides is not visible in a directory and has to be said.
2026-08-26 22:51:19 +02:00
jschoubben ae099482a9 011: twenty cases, and two axes nothing covers
Before settling a schema, what a module can actually be. Twenty kinds of thing,
with the hard ones at the end because they are the point.

The ordinary nine are unsurprising: a supervised service, a system package with
configuration, an application a person launches, a command-line tool, a library
that never runs, a one-shot task, a scheduled one, an adapter, and a standalone
application whose only difference is where its source lives.

The eleven that break a naive schema are where the work is. Something that is a
service AND an application — a git forge is consumed as a remote and operated
through a web interface, and neither reading is wrong. Something that provides
and consumes, because provider and consumer are ends of edges rather than kinds
of module. Something the mesh installs that then becomes a node CAPABILITY,
which means a node's provides-list is partly derived from what is installed on
it and not only detected. Something that must be adopted rather than installed.
Something that is a set rather than a thing. Something with exactly one instance
for the whole mesh, where assigning it twice is not redundancy but two meshes.
Something that is not software at all — a firewall policy, a DNS record, pure
desired state, which fits the host's declaration model exactly and an installable
package model not at all. An agent. The host itself, which is not a module and
needs a schema that can say so. And the things the mesh depends on and does not
control, which are why a node can be perfectly configured and still not work.

Nine axes come out of it. Two are covered by nothing anyone has proposed: HOW
MANY INSTANCES a thing may have, and WHETHER TWO CAN COEXIST — `excludes` covers
part of the second and nothing covers the first.

And one question the cases sharpen: is "runs" a property or a kind? The axes say
property — one schema with a field saying how it runs, `never` included. The
alternative is several kinds of module with different schemas, which is the
taxonomy this effort already rejected once for services and applications.
2026-08-26 22:48:20 +02:00
jschoubben f55ecc1a47 011: an abstract name needs providers that are actually substitutable
Two corrections from the operator, and the first improves the design rather than
narrowing it.

`database` is not an edge. The test it fails, and the test the proposal was
missing: can a consumer be switched from one provider to another WITHOUT
CHANGING? A module speaking Postgres does not speak MongoDB or SQL Server —
different wire protocol, dialect, driver — so a consumer declaring `requires:
database` and handed any of them breaks. The name promises what no provider can
deliver, and the resolver would report a requirement satisfied that is not.

`terminal` passes: anything that runs a command in a terminal works and the
consumer never learns which it got.

So the ADAPTER is what creates an interface. `ai-assistant` is legitimate exactly
because adapters normalise what is behind it. Without one there is no interface,
there is a category — and a category is a TAG. Tags describe, edges bind, and
keeping them apart is what stops the catalogue acquiring a second kind of
relationship that looks like a dependency and is not, which is what a folder
named after a domain already was.

And the domain module goes. A `networking` module gathering a firewall, a
resolver and a proxy under one name came from an older shape and does not fit —
there is no such thing to install. There is core infrastructure: concrete
modules named individually, not flavourable, with no grouping module standing in
front of them.

Fixed three places where the revision left the old rule standing, including an
example manifest still requiring `database` — the kind of contradiction that
would have been read as the design rather than as a leftover.
2026-08-26 22:44:37 +02:00
jschoubben e20a09ae80 011: one kind of edge
The design, rather than an account of what exists. A module provides names and
requires names, and that single relation absorbs three things this effort had
listed separately: requiring another module is requiring a concrete name,
requiring a resource is requiring an abstract one, and an interface is simply a
name with more than one provider. Nothing has to declare that it is an
interface — it either has one provider or several.

The move that does the most work: a NODE provides names too. Its profile is a
set of them — display-server, container-runtime, an architecture — so a module
requiring a display server is satisfied by the node exactly as one requiring a
database is satisfied by another module. One resolution instead of two, and a
graphical application cannot land on a node without a display server for the
same reason, through the same code, that it cannot land without its libraries.

Which makes the host's capability detection an input to resolution rather than
something a person reads. It was built to be read; it turns out to be a
provides-list.

`excludes` is the one genuinely new relation, because it is not derivable: two
modules that both provide message-bus look interchangeable when installing both
would break the machine.

Constraints are not placement. They say what must be true of a node, never which
node — which is the mistake the measurement found in the current catalogue,
where a module pins its database to a named node so a second node cannot provide
it without editing the consumer.

What it deletes, for the design: the module/resource distinction, the interface
as a kind of thing, capability checking as a separate mechanism, domain grouping
— folders assert relationships where edges record them, so a domain becomes a
query over the graph rather than a directory somebody keeps true — and possibly
tiers, if a tier is just a computed level.

What it does not delete, stated so it is not discovered later: a resolver still
has to exist, with version constraints and conflicts, and the design owes an
answer on what it delegates rather than reimplements.
2026-08-26 22:39:17 +02:00
jschoubben c3a2984b3e 011: measured, and the premise was wrong — the graph is not missing
The effort was opened to ask whether the catalogue's missing structure is a
graph. It is not missing. 126 manifests, 103 edges, no cycles, nothing dangling,
deepest chain of five — and a resolver in the SDK that topologically sorts them,
already called by the tool loader at startup, the installer when syncing modules
onto a node, and the delivery coordinator when expanding what a change affects.

It already does something this effort assumed would need designing: a
requirement on another module's provision is treated as an implicit edge to the
module that provides it. So "ordering by the graph", which ADR 0043 makes the
control plane's job, is a thing to call rather than a thing to build.

The one place the graph is wrong, it is wrong about the substrate. A module
needing a database declares `provider: postgres` inside `provisions:` — which is
what a module OFFERS — so the resolver, which reads `dependencies:` and
`requires:`, never sees it. Three edges are invisible this way, and they are the
mesh's own database, the mesh's own broker, and the work engine's database.

The consequence is measurable: computing what a working mesh needs from the
declared graph gives registry -> sdk -> mesh -> meshware. Four modules, four
levels, no database. Arithmetically correct and obviously wrong, for exactly one
reason — a field that means "depends on" is not read as one. That is
04-ISSUES/003 in a new form: not a key nothing reads, but a key read as
something other than what it means.

Two latent defects, both contrary to ADR 0008 and both in the component ADR 0043
makes responsible for ordering a host will apply without question: a cycle warns
and falls back to input order, and a dependency that does not exist warns and
continues. Neither has fired, because the catalogue currently has no cycles and
nothing dangling, which is why nobody has noticed.

And placement is decided in the catalogue: a provision pins itself to a named
node in the manifest. Which node runs what is an inventory decision — tier 2 by
the skeleton's own test — so a second node cannot provide the mesh's database
without editing the module that consumes it.

What the graph would DELETE is currently nothing. What it would add is three
declarations that no manifest uses today: excludes, a required node capability,
and an interface with adapters. Whether they would be used is not measured, and
zero usage is equally consistent with nobody needing them and nobody being able
to express them.
2026-08-26 22:35:39 +02:00
jschoubben 278f7427ed 012: the briefing carries an outcome, derived from its lines
Proposed by the operator: state plainly whether adoption succeeded, partly
succeeded or failed, with a severity per line.

Taken with one change — the overall is DERIVED as the worst mark present, never
written alongside. Two fields maintained independently drift, and a briefing
reading "full success" while carrying a failed line is exactly the fault this
record keeps cataloguing. An outcome computed from its lines cannot disagree
with them.

Four marks: ok, kept, unknown, failed. "unknown" is not a shade of success —
adoption will meet configuration it cannot parse and state it cannot read, and
folding those into "fine" is the same move as reporting an installed package as
a capability.

And adding severity reopens something the earlier rule did not cover. "Flags
inform, they do not block" was decided about CONFLICTS, where the mesh chose
deliberately and the machine still works. A failure is not "we chose" but "we
could not". Treating both the same makes a node where something the mesh needed
never happened indistinguishable from one where a log level differed.
2026-08-26 21:55:08 +02:00
jschoubben 0106318bcb 012: on conflict, keep the machine's configuration
Reversed by the operator, and both directions are recorded because the reasoning
for each is the useful part.

What is already on the machine stays, the conflict is flagged, adoption
completes. This buys non-destructiveness by construction: the class that made
the opposite rule dangerous — a storage driver against the filesystem it is
actually on, a data directory pointing at a mount that exists — cannot arise,
because nothing tied to the machine's physical reality is overwritten.

It exposes the mirror. The mesh's configuration is not only preference; some of
it is what a module needs to function. Keeping the machine's version there
produces a module that is installed and does not work, which is 04-ISSUES/007
arriving from a direction that issue did not anticipate. And a fleet where every
node kept its own settings is one where a module works on one node and fails on
another with nothing able to say why.

So neither direction is right as a blanket, and the question is not whose
configuration wins. It is whether the module REQUIRES the setting or merely
PREFERS it — required contradictions cannot be kept without breaking the module,
preferences should always yield to what is there.

That is a property of the module's declaration rather than of the adoption
algorithm, which makes it one more thing the graph would carry. Until modules
can say which of their settings are load-bearing, adoption is defaulting in the
dark, and the default chosen is the one that does not break the machine it is
adopting.
2026-08-26 21:40:25 +02:00
jschoubben bcb18c7329 012: on conflict, install the mesh's version
Decided by the operator. Where the existing configuration and the mesh's
disagree, the mesh's version is installed, the conflict is flagged, and it is
reconciled afterwards — the mesh's configuration is known to work, the machine's
is not, and a half-adopted machine is a state nobody understands.

So adoption always completes and flags inform rather than block, which also
settles what 'adopted with open questions' prevents: nothing. The node is a
node. The original is kept, so nothing is unrecoverable.

One class left open rather than folded in, because it is the one place the
oldest rule in this record argues the other way. 'Known to work' is true of the
mesh's configuration in isolation, not on this machine. Most disagreements are
preference and overwriting them is right. A few are tied to what is physically
present — a storage driver against the filesystem it is actually on, a data
directory pointing at a mount that exists — and installing ours there does not
discard a preference, it can make existing data unreadable. Restoring the
configuration file afterwards does not undo that.

The default is settled. The exception is not 'there is a conflict' but 'applying
ours would destroy something a configuration backup cannot restore', and
identifying that class is open.
2026-08-26 21:37:04 +02:00
jschoubben 60736199a4 012: keep the original, and flag what cannot be decided
Two additions from the operator, and the second answers a question this effort
had open with two bad answers.

Nothing is taken over without keeping what was there. Adoption happens on
machines somebody is already using, and the configuration being taken over is
configuration somebody chose. This is a never rule rather than a courtesy, and
it earns that by the same incident the mesh's strongest rule carries: the worst
loss in this record came from a tool acting on a path it did not own. Adoption
is that act made deliberate, which makes the safeguard obligatory.

And adoption produces a briefing, not just a result. It meets things a script
cannot decide — a runtime configured one way against a mesh wanting another, a
package pinned for a reason, local settings the mesh has no opinion about.
Silently winning is wrong in both directions and refusing outright makes a
machine in use unadoptable. So conflicts are FLAGGED: what it found, what it
took over, what it could not resolve, written to be read by a person or an agent
as the first thing a session on that node has to work with.

That is the declaration parser's principle at a larger scale — name every
problem at once, to somebody who can act on it.

The question it turns on is recorded rather than assumed away: are flags
advisory or blocking? A briefing nobody opens is worse than a failure, because
the machine is in service and the record says it went well — 04-ISSUES/003
again. Working position: the node is usable and the mesh KNOWS it has unresolved
adoption questions, as a state something can ask about rather than a document in
a log directory. What that state prevents is undecided.
2026-08-26 21:33:09 +02:00
jschoubben ddb8091f68 Research 012 — the minimum viable node, and adopting what is already there
Building tier 0 reached a wall that looked like a packaging problem and is not.
The host can be told to run a container or install a package; both need a file,
and asking where the host gets it produced a bad trilemma — carry everything,
download at apply time, or push the files in first. Downloading fails on the
first node, which cannot fetch the image registry from the image registry it is
trying to start.

The reframing came from the operator: the machine is not offline, and what
matters is WHEN the fetching happens. Move it from apply time to build time —
build the installer on a machine with a network, tailored to the target, apply
it on a target that then needs nothing. The same move the lab already made for
its router image.

Which makes the question not where artifacts come from but what is missing from
THIS machine, and that needs two things answered: the closure for a one-node
mesh, and how a machine already in use becomes one.

Adoption is the second half, and it is sharper than it sounds. Having a package
installed is not owning it: a container runtime found already present carries
settings somebody chose, and noticing the binary exists discovers none of them.

It was also the original path — 00-as-is/05 records adoption of a pre-existing
machine's configuration as the original mechanism, since made legacy and
explicitly out of scope for the lab. It returns for a different reason than it
was dropped for.

Two collisions recorded rather than discovered later. ADR 0004 has managed files
generated and never edited, and adoption needs a one-time import before that
rule starts applying — three states, and the middle one is new. And ADR 0043
says the host never touches what it did not create, which is exactly what
adoption does; that rule needs a companion rather than an exception.

Eight open questions, including whether 'tier' is just a coarse view of a graph
level, whether owning a package means owning its version, and what cannot be
precomputed at all — because tailoring moves the cost of building from source
rather than removing it.
2026-08-26 21:30:43 +02:00
jschoubben 64913ed0d3 Merge pull request 'The approval is the checkpoint, and what a declaration is' (#10) from design/approval-is-the-checkpoint into main 2026-08-26 20:37:43 +02:00