Commit Graph
79 Commits
Author SHA1 Message Date
jschoubben cabe873503 Playbook 06 — one feature, one branch, one MR per repo
Records the branching-and-merging workflow for code changes across the mesh
repos, written against a failure it names: branches and MRs opened per unit of
thought, treated as done when opened not merged, and named differently per repo,
so they pile up unmerged — one session left sixteen to consolidate by hand. The
rule is one feat/<slug> shared across every repo a feature touches, isolated in
.work/<slug>/<repo> worktrees off main, pushed and opened as one MR per repo only
when the whole feature is done, then merged promptly. Adds the ground-rule
pointer in AGENTS.md and the row in the process overview.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 03:40:39 +02:00
jschoubben 36936dc7a3 Merge pull request 'Issue 009 (fixed) + Issue 008 (resolved via ADR 0053) — module-runtime config & provider contract' (#21) from worktree-issue-provider-seal-key into main 2026-09-05 03:02:34 +02:00
jschoubben bfa08742f8 Consolidate hq: ADRs 0044-0052 merged in, and 0017/0049-0052 accepted
Brings the independent ADR branches (0044-0052) onto one branch so hq lands as a
single MR, and ratifies the five that were still proposed — 0017, and 0049-0052,
which are implemented and green in the lab. With 0053/0054 already accepted here, the
whole ADR chain 0044-0054 is accepted on this branch.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 03:01:06 +02:00
jschoubben b5a74178c3 Merge remote-tracking branches 'origin/design/adr-0044-sdk-boundary', 'origin/design/adr-0045-what-a-module-is', 'origin/design/adr-0046-events', 'origin/design/adr-0047-event-wire-shape', 'origin/worktree-adr-0048-module-broker-account', 'origin/worktree-adr-reachability-dns-firewall', 'origin/worktree-adr-config-is-the-assignments' and 'origin/worktree-adr-module-runtime' into worktree-issue-provider-seal-key 2026-09-05 03:00:24 +02:00
jschoubben 1a83ed28fa Merge pull request '011: measure the graph against every facet a module carries' (#11) from research/011-facet-coverage into main 2026-09-05 02:53:20 +02:00
jschoubben 984194e765 Accept ADR 0054 (slug) and resolve issue 010
ADR 0054 accepted with option E (a declared slug). Issue 010 resolved: the login fits
via the slug, and the minted secret shrinks to 40 chars for S3's secret-key limit —
both halves of an S3 credential now fit the tightest backend.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 02:52:03 +02:00
jschoubben af5c939d16 ADR 0054: add option E (a declared slug) and recommend it over A+B
Implementing A (bound the identity at 20) revealed the readable budget is node+module
<= 14 chars — so tight that the catalogue's own test names (workstation+keycloak, 25)
compact to an opaque hash. B's fallback would fire for the common case, not the rare
overflow, inverting A+B into mostly-opaque identities. Option E — an optional short
slug a module/node declares, preferred over the cleaned name — is the escape hatch B
wanted to be without the opacity: legible because a person chose it, and it makes an
early refusal palatable (refuse on the slug field, not the machine's name). B dropped;
E recommended over a bound of 20, composing with C later if needed.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 02:30:11 +02:00
jschoubben 89b65dd4c0 ADR 0054 (proposed) — a consumer's identity is bounded by the tightest backend
Sketches the options for issue 010: the mesh's identityLimit (63, postgres's) is not
the shortest among the backends the derived name reaches — S3's is 20 — so CheckIdentity
lets an over-long access key through and minio fails at provision time. Options: bound
by the true minimum and refuse at assignment (recommended, with a compact fallback held
in reserve), per-interface bounds, or a provider-generated identity (rejected — breaks
"the mesh says the identity once"). Links issue 010 to it.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 02:17:46 +02:00
jschoubben 75887d0bc5 Issue 010 — the mesh's derived login does not fit every backend's identity rules
Found doing the per-backend provider e2e: redis and postgres accept the mesh's `as`
(mesh_<node>_<module>) verbatim, but minio's S3 access key is capped at 20 chars and
`as` is 22, so the provisioner cannot create the service account. `as` is doing two
jobs — a stable identity the two ends agree on, and a literal identifier a backend
must accept — and those are not always the same string.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 01:58:11 +02:00
jschoubben e5f4af8cf2 Accept ADR 0053 and resolve issue 008 — provider contract implemented and proven
ADR 0053 accepted; adds the scope boundary the umami rework surfaced (credential
provisions vs data provisions — analytics' generated siteId return is left to a
separate decision) and records the lab proof. Issue 008 marked resolved: the sdk
harness and the four adapters are reworked, the symmetric seal removed, and
provider-uses-mesh-credential is green.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 00:27:56 +02:00
jschoubben e27f8b1b7d ADR 0053 (proposed) — a provider creates the credential the mesh minted, and seals nothing
Issue 008's trace confirmed the premise in control-plane code: the mesh already
mints one password per consumer/provider pair and delivers the provider its copy
(SecretFor/SecretsFrom/grantsFor -> Grant.Sealed; the receives contribution carries
As + Secret). The provisioner's symmetric seal is an orphaned, contradictory second
model. ADR 0053 corrects the provider contract in one place (the sdk harness):
providers create the resource with the mesh-supplied login and password and drop
seal/key/return entirely. Reframe 008 as contract-first (every provider, not four).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 00:06:50 +02:00
jschoubben 332c334767 Issue 008 — sharpen: the provisioner's seal model is orphaned, not just undelivered
A cross-repo trace showed nothing writes the provisioner's grant-request files,
nothing reads its sealed credentials, and no consumer unseals — while the mesh
already mints and delivers provider/consumer credentials asymmetrically with no
shared key. The fix is to drop the symmetric seal and have providers consume the
mesh-minted password, a breaking provider-contract change that wants an ADR.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 23:40:20 +02:00
jschoubben 47b0ca5c90 Issue 009 — settings change does not restart a container-hosted runtime
Found rolling the runtime out to the catalogue: a module's runtime reads its
settings-merged config file once at start, but a container is only recreated on a
spec change, and file content is not part of the spec. So updating settings
re-renders the file and nothing re-reads it — ADR 0051's "on the fly" holds only
for config set before first start. Services have restart-on; containers do not.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 23:06:48 +02:00
jschoubben d1aa4254c9 Issue 008 — a provider runtime has no seal key the mesh can deliver
Found building the module-runtime vertical slice: a provider's provisioner
requires a seal key it has no way to receive, and the consumer no way to obtain
the matching one. The runtime cannot come up as delivered.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 22:23:24 +02:00
jschoubben f63eca13b3 ADR 0052 — a module runs its code as its own process, with its own account
The runtime-model gap the review found. A module with tools or events runs one
container — the tool runtime carrying its code — holding the one scoped account
ADR 0048 gave it. A node-wide runtime can't: it would hold the union of every
module's permissions, the isolation 0048 draws. So per-module: one module, one
process, one account. Tools served per key (serve.<tool>) so a caller names a
tool and only its module answers (superseding a shared tools.invoke); events in
the same process under the same account; the runtime image is the tool runtime
plus the module's code (the audit-logger's shape, made the rule). A plain
service module runs no such process. A provider's provisioner is a runtime too —
which is why a provisioner that emits must carry a broker credential or not emit.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 21:24:46 +02:00
jschoubben e898a3ec44 ADR 0051 — a module's configuration is its assignment's, not its manifest
A module is assigned to a node (there is no mesh assignment; 'mesh' is a scope).
The manifest is what the module IS, plus defaults; the configurable values are
settings, carried by the assignment — per-node or mesh-wide, applied at
resolution, changeable live (what a meshboard edits). Extends settings from a
config file's content to the manifest fields marked settable: foremost
listens.from (postgres from:mesh by default, from:anywhere per node — the
firewall follows), and a provider's own config (a registrar's zone/domain/
ingress). Static config in a manifest is config in the wrong place: it cannot
vary per node and cannot change without a rebuild.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 21:05:55 +02:00
jschoubben 9eb5576682 04-ISSUES/003 — fixed: the firewall scope is enforced now
The manifest refuses unknown keys (DisallowUnknownFields), 'from' is the field
that scopes a port and it is rendered to nftables (AsNftables), and the firewall
module applies the rule set. The chain from a declared scope to a dropped packet
is closed. Amended-design: ADR 0050.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 20:56:44 +02:00
jschoubben 5118258ee3 ADR 0049 and 0050 — public DNS and the firewall, two reachability decisions
0049: a public name is provisioned like any capability — a module requires
public-dns and contributes its host; a neutral interface answered by
registrar-scoped providers (cloudflare-dns, route53-dns) that create/remove
the record pointing the name at the mesh's public ingress. Pairs with route
(the proxy) and a public cert (the proxy's ACME).

0050: answers the firewall question. The firewall is NOT a provider like the
proxy — it is a machine's own filter, derived by the host as the sum of what
its modules declare they listen on, with 'from' the whole of public-vs-internal.
Enforced both ways, unknown keys refused — closing 04-ISSUES/003. A public
service is exposed through the proxy (listens from:mesh + requires route), not
by opening its own port.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 20:45:58 +02:00
jschoubben cfa3ad1930 ADR 0048 — ratified: status accepted 2026-09-04 01:22:47 +02:00
jschoubben b622de4fe6 ADR 0048 — a module's broker account is scoped by its emits and consumes
Events (0046) and their wire (0047) left open how a module reaches the
broker. The code has no generic module broker-account: only node and
builder scopes exist, so emits/consumes are enforced by nothing — a
manifest declaring a scope the broker does not draw (04-ISSUES/003).

Decides: on assign, a module gets a broker account whose permissions ARE
the manifest — read on mesh.events + its own queue bound to consumes;
write to mesh.events under module.<self>.* only; nothing else. Consuming
'#' is a deliberate, auditable grant. The account is what makes the
declaration a rule the broker enforces, not a comment.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 01:19:15 +02:00
jschoubben 820ce8fc8c ADR 0047 — the shape of an event on the wire
The wire contract ADR 0046 left open: two topic exchanges (mesh.events,
mesh.rpc, kept apart so # is a clean audit); the routing key as the event
type namespaced by origin (module.*, mesh.*, node.*); metadata in AMQP
headers (required x-event-id/x-source/x-node/x-time/content-type; optional
x-causation-id/x-schema; unknown x- headers ignored) with the body only the
payload; persistent messages; per-consumer durable dead-lettered queues
with prefetch; at-least-once with idempotent consumers (no false exactly-
once). The precedent is ADR 0043 for declarations.

Supersedes the sdk's first cut (metadata in body -> headers); that and the
queue config are code to align in mesh-sdk and mesh-tools.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-03 23:47:58 +02:00
jschoubben 4bb6dd26ed ADR 0046 — events are a relationship, provisioning's lighter sibling
A module emits and consumes events, both declared (emits/consumes),
parallel to provides/requires. Events are 1:many, broadcast, credential-
free — no provisioner, just the broker's topic routing — so most inter-
module reaction should be an event, not a provision. Every event carries
source/node/time so it is auditable; the audit logger is just a module
consuming '#', no privilege. A consumes for an event nothing emits is a
dangling edge and refused, like requires. One per-node runtime serves
tools, provisioning and events alike.

Extends ADR 0045; builds on ADR 0001 (the broker) and 0044 (emit/on are
stable sdk surface; the binding and runtime are not).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-03 23:42:25 +02:00
jschoubben b2cb481681 ADR 0045 — what a module is
A module is one self-contained piece of software the mesh installs and
manages; the software is its identity, and capabilities/seats/provisions
are the relationships between modules, not what a module is. Records the
three relationships (shared seat, exclusive seat, provide/require), that
interfaces are mesh-owned and providers adapt to them, and the naming
rule: draw the interface at the consumer's real coupling — neutral where
the coupling is thin (analytics), protocol-scoped where the consumer
speaks a protocol (postgres/mssql/mongodb), never false genericity.

Supersedes 0017 (domain grouping — wrong axis), refines 0002, generalises
0027's protocol-not-product rule.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-03 22:50:52 +02:00
jschoubben 50d61398b6 ADR 0044 — what the SDK holds, and what it refuses
Supersedes ADR 0030's 'types, not behaviour' line for mesh-sdk. The
boundary is change-frequency, not kind: the SDK holds the stable spine
(tool-serving harness, messaging/event framework, contracts, core
primitives) and refuses per-module clients, per-module tool code, and
anything volatile — because those are what turned hal/sdk into constant
maintenance and made every edit rebuild every module.

States the rule (frequent AND cascading is the disease), why the root
cause was intra-module feature-sharing leaking into inter-module
coupling, and where per-module shared code lives instead (in the
module — a shared file, or a module-local sdk for the few large ones).
Updates repos.md's canonical mesh-sdk description to match; leaves 0030
untouched (immutable).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-03 21:41:08 +02:00
jschoubben b245f5e254 011: measure the graph against every facet a module carries
The five declarations cover relations; a module is more than its
relations. Add the facet-by-facet coverage table so the effort cannot
conclude while tools, verification and contributions are unplaced —
and weigh each candidate gap rather than adopting it: contributions
probably dissolve into declared resources, mandatory verifiers risk
trivial ones, and the tool surface is the one facet with no home.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-01 23:14:03 +02:00
jschoubben 64913ed0d3 Merge pull request 'The approval is the checkpoint, and what a declaration is' (#10) from design/approval-is-the-checkpoint into main 2026-08-26 20:37:43 +02:00
jschoubben 00d5ba8376 0042 and 0043, as approved
0042 becomes the operator's own rule and stops there: every merge into the main
branch is notified and approved. Notified means proposed and said out loud, not
performed and mentioned; approved means a person says yes to THAT merge. Who
performs it is not the thing worth constraining, which is what makes an agent
merging its own work unremarkable — the checkpoint already happened.

0043 answers the question it did not: where the ordered list comes from. By hand
today, in substrate.lock, because the first node has no control plane to derive
anything from. Afterwards the control plane derives it from module assignments,
resolved configuration, and what each module declares it needs — ordered by the
dependency graph, which is research 011.

So the record is complete on the consumer side and deliberately silent on the
producer side, and that is a legitimate order to settle them in: the host must
refuse what it does not understand whoever wrote it.

One consequence that only appeared when the question was asked: if the graph
turns out not to determine a total order, that is 011's problem and not the
host's. The host is still handed a list and still applies it as given. Recorded
because it is the seam where a future difficulty would otherwise try to migrate
into tier 0.

Both were marked accepted before they had been read. Approved now, so the field
is true — which it was not when it was written.
2026-08-26 20:37:13 +02:00
jschoubben bea052753e ADR 0043 — what a declaration is
Stage 2 could not start without it. Three constraints already bound the shape
and between them they decide most of it.

JSON, because the standard library carries it and carries no YAML, and a YAML
declaration would put a third-party parser inside the one binary whose whole
argument is that it needs nothing — to gain authoring comfort in a document
generated by a machine and read by a machine.

An ordered list, because ordering is a DECISION. A host deriving order from
declared dependencies would be deciding the thing most likely to differ between
what the control plane intended and what the machine does. The control plane
knows what depends on what; it says so by saying when.

Unknown is refused, never skipped — an unknown version, type or field refuses
the whole declaration. A host that skipped what it did not understand would
apply most of a declaration and report success, which is 04-ISSUES/003 with the
declaration on the other side of the wire.

Complete for what the host OWNS, and only that. It removes what it previously
applied and is no longer declared, which it knows from the store rather than by
inference, and never removes what it did not create — a converger that treats
'not declared' as 'must not exist' deletes what the mesh never put there.

Two consequences arriving earlier than the build order suggested: the store is
load-bearing at stage 2, because nothing can be removed without knowing what was
applied. And a closed address space bounds the first vocabulary to what needs no
network, because a scenario has no route to a package repository.
2026-08-26 02:04:41 +02:00
jschoubben 23232e019a ADR 0042 — the approval is the checkpoint, not the second pair of hands
§2 said "never merge your own", written for people. Applied to an agent it
produced a contradiction that surfaced immediately: an agent asked to merge
cannot merge, because it authored what it is being asked to merge. So every
merge here was either performed by the thing that wrote it, or not performed.

Rejected the literal reading, because an operator clicking merge dozens of times
without reading is not a checkpoint — it is the SHAPE of one, which is worse,
since the record then claims a review that did not happen. Rejected dropping the
rule, because the failure it prevents is not one an agent is less prone to.

So the rule names what the checkpoint actually is: a person deciding, not a
person clicking. Work may be merged by whoever wrote it once a human has
explicitly approved that merge. What "explicit" excludes is the half that can
rot, so it is enumerated: a standing permission cited forever, an instruction to
do the work read as approval to merge it, silence, and the author's own
judgement that it is ready.

This narrows rather than relaxes. The obligation moves from who performs the
merge to whether a person decided — a higher bar in the case the old wording
permits, where a reviewer merges someone else's work without reading it.

Synced, and verified by reading the rule back out of the live page rather than
by trusting the publish.
2026-08-26 01:08:34 +02:00
jschoubben 7ef1c561e0 Merge pull request 'Tier 0: the questions answered, the decisions taken, and the design' (#9) from design/close-the-record into main 2026-08-26 00:56:07 +02:00
jschoubben 92e8c74ce4 ADR 0041 and the build handoff for the node host
Building tier 0 forced the question "the one binary installed by hand" had been
carrying unexamined. A TypeScript host needs a runtime present before it runs,
so the thing installed by hand becomes two — and the second must be installed by
the means the host exists to replace.

So the host is a statically linked binary that requires nothing present, written
in Go. Rejected: a runtime installed first, which breaks the property the tier
rests on; and bundling the runtime into the executable, which carries ninety
megabytes to preserve a language choice and puts a young feature at the bottom
of the stack.

The argument that decided it is architectural rather than about taste. 0037
means the host never queries the mesh database and 0039 means it only receives
declarations, so the host shares NO code with any other tier — not a client, not
a schema, not the SDK. The language boundary falls exactly on a boundary that
already exists, and a second language usually costs duplicated logic where here
there is none to duplicate.

§8 gains a scope: it said "TypeScript throughout" when everything was a service
or a surface, and is now scoped to those with tier 0 named. Another sync owed.

Playbook 04 steps 2 and 4: repos.md records mesh-host as existing, the design
takes code: [mesh-host] and status: in-progress.
2026-08-26 00:16:07 +02:00
jschoubben b9facf9375 Design the node host
Playbook 02 step 3, on four recorded decisions. Tier 0 has one job — apply
declared state on this machine — and the six absorbed concerns are instances of
it, not additions to it.

Specifies the six parts and what each owns, and the two properties that make
apply trustworthy rather than merely present: every applier reads back, because
setting a value is not evidence the value took; and what was applied is recorded
after it works, never before, because a failed apply leaves the machine wherever
it reached and nothing must claim otherwise.

Build order is staged so each stage is verifiable in the lab before the next
exists. Stage 1 is profile and inventory — no control plane, no declarations, no
network — and it is deliberately the smallest useful thing, because `place:` has
nothing to place and the lab therefore raises empty machines. Stage 1 ends that,
and every later stage is tested by a lab that already works.

Stage 2 is the one that could invalidate the tier boundary: whether one host can
raise the substrate alone is Move 1's assumption and has never been proved.

Every decision the design rests on is given the test that asserts it, per 0034 —
including the dependency-direction lint, which is what makes "the host never
queries the mesh database" a rule rather than an intention.

Six things left open and named, including the one that host-size.md could not
measure: zero dependencies, but still six vocabularies.
2026-08-26 00:10:24 +02:00
jschoubben 6e4fc5d69b ADR 0040 and the constitution sync — absorb, then publish
The sync came due for three accepted rules. Reading the target before
overwriting it found the source and the enforced copy had diverged unrecorded,
and that a literal republish would have DELETED rules the mesh enforces: the
live page's §4 carried SOLID, layering, TypeScript strictness and DRY/YAGNI,
which appear in no decision record anywhere and have been checked against for
six weeks.

Removal was not the safe alternative either. The orchestrator reads "when
absent, no constitution is injected (backward-compatible)" — so deleting the
page would not fail, it would silently inject nothing, and every design meeting
would run unchecked. Three unenforced rules would have become all of them.

So: absorb first. The code-quality rules land as §8 rather than §4, because
appending renumbers nothing and every existing citation stays valid. They are
marked as inherited — every other rule states the incident behind it, these
state nothing because nothing was written down, and importing them silently
would have claimed a provenance the document does not have.

The review bar is resolved to a person who is not the proposer. The live page
required two node operators; there is one, so the rule was never met and could
not be — a rule that cannot be satisfied is not a high standard, it is one
everything silently violates.

Then published, and the read-back earned its place in the playbook: the FIRST
publish reported success and changed nothing. New revision, new title, body
unapplied — a malformed argument dropped silently. §5 demonstrating itself
during its own publication.
2026-08-26 00:05:50 +02:00
jschoubben 34e7a4780a Accept 0035 — a report is read from the system
The principle is obvious; the obligations it imposes are not, and those are what
was actually being decided.

The applier records what it did, including facts it never reads itself, purely
so something else can read them back. It records them AFTER the thing works, so
a half-finished run ends up with less metadata rather than optimistic metadata —
rejecting the simpler alternative of tagging at creation with a status field,
because a failed raise deliberately leaves wreckage standing and wreckage tagged
at creation claims things that never happened. And it binds anything that
reports on the mesh, not only the drawing.

§5 carries the rule unqualified now. The constitution sync is owed for this and
for 0034 and 0018, and has not been run.
2026-08-25 23:40:21 +02:00
jschoubben 66a89cefbe Accept 0039 — a node owns no password, only an identity
Settled in the operator's own words: nodes should not own passwords, only an
identity when communicating to the broker. Recorded that way at the top of the
record, because it is the whole decision in one line and the rest is why.

What this obliges, in order of newness: enrolment is the one mechanism that does
not exist. Per-node broker users, virtual hosts and per-queue permissions are
broker configuration. Mutual authority is certificates on a connection already
open. And the boundary must fail legibly, which is the requirement the debugging
objection earned.
2026-08-25 23:37:52 +02:00
jschoubben 9ee0c63c7a 0039: the link already exists, and most of the cost is already paid
Written first as though the link were a thing to build. It is not. ADR 0001
already has it — every node connects outbound to a single broker, nothing ever
connects to a node, each node declaring an exchange named for itself and
consuming from its own queue. Already outbound-only, already per-node addressed,
already the one channel everything arrives through.

So this record is not proposing a channel. It proposes that the channel carry
per-node identity instead of one shared credential. The same as-is records the
fault it fixes, for the broker rather than the database: the broker's credential
is mesh-wide, rotating it is a mesh-wide operation, and doing it wrong has taken
the broker down.

That reframes the overhead objection, which was fair against what the record
said and not against what it means. Of six properties, four are already true,
one is broker configuration — users, vhosts and per-queue permissions the broker
already implements — and exactly one is new machinery: enrolment. Meanwhile 0037
subtracts, since a node under it holds no database credential at all. Three
hand-carried shared secrets become one identity that grants only identity.

Adds the option that was actually being weighed and was missing: accept the
exposure as the cost of simplicity. Rejected because the simplicity IS the
unfixability — the credential cannot be rotated precisely because everything
holds the same one.

And adds a requirement from the debugging objection, which was the strongest
part of it: it must fail legibly. A boundary that refuses a node without saying
why is worse than the password it replaced, because a wrong password at least
announces itself. That is §5 applied to a security mechanism.
2026-08-25 23:29:07 +02:00
jschoubben 902739acb6 Research 011 — the module graph
The proposal to split modules into provisioning services and applications was
worked through and abandoned, for a reason worth keeping: it cannot be filed
consistently. A git forge is consumed as a service and operated through a web
interface; an analytics service grants tracking identity and is a dashboard.

The operator's correction is the sharper form — what runs on the machine is a
supervised container, not something a user started. That is a fact about HOW a
thing runs, not about what kind of thing it is. So it is a facet, and 0002
survives: everything is a module.

What the catalogue is missing is not a taxonomy but a graph. Grouping asserts
relationships; a graph records them. Five declarations, of which two exist:
requires/provides a resource (yes), requires/excludes another module (no),
requires a node capability (no). Plus interface modules that carry no
implementation, with adapters providing them.

Recorded because it matters: this is a package manager's model, and pacman
already has all of it — depends, conflicts, and provides as virtual packages,
which is exactly the interface/adapter idea. Arriving there independently is
evidence for the shape. It is also a warning about what not to reimplement.

Working position on capabilities, to be tested: intrinsic ones (hardware,
architecture, network position) are detected and never installed, and a module
requiring one it lacks is impossible rather than unresolved. Provided ones (a
display server, a container runtime) are not a separate kind of thing — they are
modules that provide a capability, so "may the mesh install a capability" is not
policy, it is dependency resolution. Issue 007 then bears directly: an installed
package is not a capability.

Also captured: the operator's assessment that the machinery around a module —
scheduled tasks, hooks, migrations, config and env — is worth keeping, seeds are
not, and the integration is wrong enough to need a major refactor. Research 005
found supporting evidence from another direction, that the densest apparent
coupling in the catalogue is manifest boilerplate churn.

The first open question is the one that decides whether this is progress: what
does the graph DELETE? If modules gain declarations and lose nothing, it is
motion.
2026-08-25 23:19:57 +02:00
jschoubben f5073b9b00 Accept 0034 and 0018; 0011 is superseded
0034: a test defends a decision. §5 carried it marked "proposed, pending
review"; the marker is removed and the rule now stands unqualified. The lab was
already built to it, which is the inversion §5 exists to catch — closed now
rather than left standing.

0018: the mesh creates no symlinks. §3 said the installer owns the links today
and the INTENT was that the mesh creates none. It is no longer an intent, so the
wording says so, and ADR 0011 becomes superseded rather than edited — its
reasoning is why the rule exists at all, and the incident behind it is the
reason anyone believes either record. The links the installer still reconciles
are a migration, not a permission.

The constitution sync (§6 step 4, playbook 05) is NOT done. An unsynced rule is
a rule the mesh does not enforce, whatever this document says — and publishing
it changes what every design meeting is checked against, so it wants saying out
loud rather than doing quietly.
2026-08-25 23:08:54 +02:00
jschoubben 4866007c04 ADR 0038 accepted; ADR 0039 (proposed) — the link is the security boundary
0038 accepted: no two modes. One behaviour, two sources of declaration.

0039 fills the gap 0038 named. Proposed rather than measured — it designs a
boundary that does not exist — but what it replaces IS measured, and that is the
argument.

Today adopt.sh asks the operator to paste in the postgres password and the
object store password, the same ones on every node, and they are not discarded
after adoption: wireguard and traefik open a pg connection on every reconcile.
So every node permanently holds a credential to the control plane's database,
and 00-as-is/06 records that nothing rotates it. Compromise of any node is
compromise of the mesh's store, with no way back.

Four properties make the link a boundary rather than a pipe: it is outbound and
node-initiated, so a node has no listening control surface — which the topology
already requires, since most nodes have no forwarded port. A node holds its own
identity and nothing else, so compromise of a node is compromise of that node.
Authority is mutual, because a host that applies whatever the link delivers must
know the mesh from something impersonating it. And what may be pushed is bounded
by FORM — declarations of known shape, never a command to run.

That last property is stated with its limit rather than oversold: it bounds
form, not impact. A compromised control plane can declare harmful state and the
host will apply it faithfully. What it buys is a describable blast radius.

Joining uses a one-time short-lived enrolment token, useless once used and
useless after a while, in place of hand-carried shared secrets.

Named rather than hidden: the first node's identity is self-issued and becomes
the root of trust, which is the one place 0038's "no special first node" does
not fully hold. Rotation becomes possible and is still not designed. And what
may expire is constrained by 0036 — an identity needing refresh would make a
laptop fail for being a laptop.
2026-08-25 22:27:31 +02:00
jschoubben 72b22830f3 ADRs 0036, 0037, 0038 — what a node is, what the host does, how one joins
0036 (accepted): a node is a managed machine, and disconnection is a situation.
The open question posed a class distinction — full nodes and lesser presences.
There is none. Reachability is state, not kind, which promotes the host's local
store from a component to a requirement: it is what makes disconnection ordinary
rather than exceptional. The reduced contract the question reached for is real
but it is capability, and that belongs in the profile.

0037 (accepted): the host applies, it does not decide. Measured rather than
argued — the absorption is smaller than the machinery that already applies
state, and eight of ten adapters carry no dependency to move. The two that do
open a Postgres connection to the control plane, which inside tier 0 is the one
thing the tier rule exists to forbid. So each concern splits: deciding needs
every other node and stays in tier 2; applying needs root and locality and goes
to tier 0. The host carries ONE concern, of which the six are instances.

0038 (proposed): a node joins by linking first. The operator's two-modes
proposal, adopted as intent and corrected as structure. Two modes is two code
paths where the first runs once per mesh and rots — and the mesh already has
that fault in its worst form, as three hand-run shell scripts. Instead: one
behaviour, two sources of declaration. The first node is not a different kind of
node, it is a node whose mesh is not up yet, and its specialness is temporary
and self-erasing.

0038 also shrinks the migration 0037 called expensive: a joining node never
needs mesh-wide state, because the hard part of the overlay is only needed to
compute the WHOLE mesh. It needs one peer. The rest arrives.

Left open and said so: what may be pushed over the link and how a joining node
proves it is entitled to join, and whether one host can raise the substrate
alone.
2026-08-25 10:34:45 +02:00
jschoubben 42bce02bba 006: answer the host-size question by measuring it
The skeleton's biggest unproven claim was that absorbing six concerns makes a
binary whose whole argument is having no dependencies carry six of them.

Measured against origin/main, and the question turns out to ask about the wrong
axis. By size the absorption is SMALLER than the machinery that already applies
state on a node — 2755 lines of adapters against 3059 lines of meshware,
env-sync and config-sync. The host is not a new large thing; it already exists,
spread across three core modules.

The real risk is direction, and it is two modules wide rather than six concerns
wide. Eight of ten adapters already receive derived state and only apply it, so
absorbing them moves code that has no dependency to move. Two — wireguard and
traefik — open a Postgres connection to the control plane and compute their own
configuration, which inside tier 0 would be an upward dependency and is exactly
what the tier rule forbids.

And the split has already been happening without being named: dnsmasq-app needs
the same node data as wireguard and does not query for it, because hand-
duplicated state went wrong and someone derived it centrally instead. Eight of
ten adapters are on the far side of that migration.

So the absorption is not a move, it is a split: deciding stays in tier 2,
applying goes to tier 0. The claim survives with its scope corrected — the host
carries ONE concern, apply declared state on this machine, of which the six are
instances.

Stated open rather than glossed: the two unsplit modules are the two hardest,
six concerns is still six vocabularies even at zero dependencies, and what the
host must carry versus find is issue 007 and unresolved.

Question B also recorded as answered by the operator — a node is a managed
machine, and a disconnected node is still a node in a different situation. The
question posed a class distinction; there is none, and what varies is state.
2026-08-25 02:20:32 +02:00
jschoubben 4bf7a35568 Close the record on the lab
Playbook 02 and 04 were followed for the substance — decisions before design,
design before build — and skipped for the bookkeeping. This closes that.

004 graduates. Its one open item was "not yet stood up"; the lab is stood up,
and the substitution the effort turned on is now enforced by the validator
before anything is raised rather than left as a thing to remember. Its
certificate conclusion has a home in 01-end-to-end-testing and is designed but
not built — implementation is a third axis, and an effort graduates on its
conclusions.

One item leaves 004 without a home and is recorded rather than lost: the reverse
proxy does not set caServer, so it defaults to the production endpoint.

The two lab designs read `designed` while running in production of a sort, so
they become `in-progress`.

And the lab gets an as-is document, which it did not have. It records what runs
including the parts nobody would choose again: that `place:` is refused and the
lab therefore raises EMPTY MACHINES, that the drawing shipped with no design
document behind it, that a router is tagged as a machine for a reason found by a
bug, and that the integration suite raises two of five scenarios while both
faults found so far lived in the three it does not.

006 stays active, deliberately. Two of its open questions ARE the tier 0 design
— whether absorbing six concerns makes the host too large, and whether an
unprivileged node earns a place in the inventory. Playbook 04 is explicit that
an open question is a reason to research, not to build around.
2026-08-25 01:55:52 +02:00
jschoubben b225b07625 ADR 0035 (proposed): a picture is read from what runs
Drawing a scenario forced a choice that looks cosmetic and is not. A diagram
built from the declaration and captioned "as raised" answers "is what is running
what I asked for?" with the request, which always agrees with itself.

So: a picture captioned as raised reads only the running system, and where the
hypervisor does not hold a fact the picture needs, the raise records it on the
resource. With the rule that makes the recording worth anything — a behavioural
tag is written after the behaviour works, never at creation, because a failed
raise leaves wreckage standing and a picture of wreckage must not badge what the
wreckage was supposed to be.

It earned itself on the first comparison: every VM showed no addresses, because
a container's interface carries the device's name and a VM names its own. The
two pictures disagreed, so a whole class of machine silently losing its
addresses was visible in seconds.

§5 of how-we-build gains the general form, marked proposed. The constitution
sync is deliberately NOT done — a rule the mesh enforces before a second person
agreed to it is what §6 exists to prevent.
2026-08-24 22:54:30 +02:00
jschoubben 8efa063f21 The snapshot question is answered by a test
The lifecycle design asked whether a scenario snapshot needs the machines
stopped. The integration test answered it on its first run: no, but they
must be flushed.

A snapshot captures disk and not memory, so a write still in the guest's
page cache is absent from it — not stale, absent. A file written seconds
before a snapshot did not survive the restore.

Flushing first buys write-durability. It does not buy
application-consistency: anything mid-transaction is still captured
mid-transaction, and that limit is now stated rather than left implied.
2026-08-24 22:26:47 +02:00
jschoubben a2495e4d8e ADR 0034 (proposed): a test defends a decision
how-we-build 5 already says that if a document states a rule about the
mesh, it says how the rule is verified — an unenforced rule being
indistinguishable from a wrong one, and costing more because people
believe it. That has never been applied to decisions, and a decision
record states the same kind of claim.

The gap was found by review: the lab reached 2,128 lines with 1,072
untested and no stated rule broken, because there is no testing posture in
how-we-build at all. Every decision the lab embodies was verified by hand
and none of it survives the terminal it ran in — which is
04-ISSUES/005 in miniature, coverage assumed rather than checked.

Rejected a coverage percentage: it measures how much code a test touched,
not whether anything important is defended, and would have been satisfied
by testing the parser harder while the hypervisor integration stayed
unasserted.

Rejected test-driven development as a hard rule, and not because it is
wrong in general. Half this implementation was discovery — that the
hypervisor CLI reads a definition from stdin and hangs, that it assigns a
MAC without recording it, that a stock image's networking flushes a static
address. A test written first against undiscovered behaviour asserts a
guess.

So: structure and logic tested first, behaviour against a real system
tested alongside, mocking the boundary forbidden, and a blocking gate as
the definition of done. A test names the decision it defends, which is what
makes the pairing checkable — a decision without one can be found rather
than noticed.

Stated as proposed rather than adopted: 6 requires review by someone who
is not the proposer. Records 0001-0033 predate it and are not retroactively
invalid, but each should acquire a test or an explicit note that it cannot
have one, and until then the rule is aspirational for them — which is the
state 5 warns about, recorded rather than hidden.
2026-08-24 22:22:19 +02:00
jschoubben eab4598494 ADR 0033: a router is scenery, not a node
ADR 0016 makes a lab node a virtual machine, and its reasoning is
fidelity: a node boots a stock image and runs the real install, so it has
to be a real machine or the thing under test is not the thing that ships.

That reasoning does not reach a router. Nothing under test runs on one, it
holds no identity, the mesh never installs anything on it, and no assertion
is ever made about its internals. It exists so packets behave the way they
behave in the world, which is the definition of scenery.

So a router is a system container. What it must reproduce is kernel
behaviour — translation, connection tracking, filtering, forwarding — and a
container has the same kernel.

Verified before deciding rather than assumed. In a plain unprivileged
container: ip_forward and ipv6 forwarding both settable, nftables
masquerade accepted and listed back, and the conntrack timeouts that
mapping_ttl depends on both writable. No privileged mode, no nesting, no
capability grants.

Rejected letting the hypervisor provide NAT, on a stronger ground than
speed: it makes the lab provide what the declaration is supposed to own,
and it cannot express a mapping that expires, a gateway that refuses to
forward, or policy between siblings. The model would shrink to fit the
tool.

The distinction is now load-bearing and has to stay legible: node means
something under test, scenery means something that makes the test real. If
the mesh ever installs anything on a router, it has become a node and this
record no longer covers it.
2026-08-24 01:28:08 +02:00
jschoubben 132f6a0626 Merge pull request 'The scenario model, the lifecycle, and what the lab actually costs' (#6) from design/scenario-underlay-detail into main 2026-08-24 00:40:54 +02:00
jschoubben e88b448145 The fix is real: 76x, verified. And how the lab installs on a clean machine
Snapshot 9.9s -> 0.13s. Restore 10.4s -> 0.80s. Three snapshots sharing
1.36 GB instead of costing 4.8 GB. The projected four-machine reset cycle
falls from ~90s, unbounded at worst, to ~15s dominated by a boot that
cannot be avoided. ADR 0029's inner-loop argument holds with copy-on-write
and did not without it.

The consistency matters as much as the speed: three consecutive snapshots
took 0.13, 0.12 and 0.13 seconds, against a dir second snapshot that never
finished.

One honest counter-observation recorded: launching onto the fresh
copy-on-write pool was slower, 20.2s against 14.3s, because the image had
to be unpacked into a pool that had never seen it. Paid once per pool, and
dwarfed by what snapshotting saves, but it went the other way.

Doing the measurement produced the answer to how the lab installs on a
clean machine, because both failure modes appeared while doing it.

Installed is not available: the daemon was present with units disabled and
no group. Issue 007.

Available is not adequate, and this is worse: with the storage tooling
absent everything worked and snapshots were seventy-six times slower.
Nothing failed, nothing warned. That is a variant the mesh has not
catalogued — its usual failure is reported success and did nothing; this is
reported success and did it seventy-six times slower, which no error
surface catches because nothing is wrong.

So the lab verifies CAPABILITY, never installation, and refuses to run
degraded rather than warning — a warning about a slow inner loop is read
once and ignored forever. Prerequisites may arrive from a mesh module or
from the lab's own bootstrap, and the second path is required rather than
convenient: a lab installable only by a mesh cannot host the development
of the mesh that installs it.

The lab is the second thing installed by hand, after the node host, and for
the same reason: something has to be first, and pretending otherwise
produces a circularity papered over by a script nobody exercises.
2026-08-24 00:14:57 +02:00
jschoubben 98bcd5cc49 Measure the lab's inner loop — it is too slow, for a fixable reason
The lifecycle design closed on a question that was measurable rather than
arguable, so it was measured. One virtual machine on a workstation with
hardware virtualisation and NVMe.

Raising: the launch call returns in 3.4s, the machine is actually usable
after 14.3s. The gap is a design constraint — raise must wait for the
second number, because reporting the first would be transport reported as
effect, which is the mesh's own recurring failure.

Snapshot: 9.9s and 1.6 GB for a 1.5 GB instance. A dir snapshot is a full
copy; nothing is shared. Restore: 10.4s, usable again after 20.1s.

The second snapshot exceeded two minutes and never completed. That is the
more troubling number: snapshot cost here is not merely high, it is
unpredictable, and a loop with a variable multi-minute step is one nobody
trusts.

Projected to a four-machine scenario, a reset-and-rerun cycle is about a
minute and a half at best and unbounded at worst, before any of the mesh's
own work begins. That is too slow for an inner loop, and ADR 0029's whole
argument — that making the bootstrap path the inner loop turns the
least-exercised code into the most-exercised — holds only while resetting
is cheap.

The cause is not virtual machines. Hardware virtualisation is present and
machines boot in fourteen seconds. It is that the daemon offers exactly one
storage driver, dir, which has no copy-on-write and therefore no cheap
snapshot. The btrfs kernel module is available; btrfs-progs is simply not
installed, which is the entire reason the driver is absent.

The copy-on-write comparison was deliberately NOT run, because running it
would mean installing a package by hand — which the rules forbid and which
would have made the measurement unreproducible. So the honest statement is
that the current configuration is too slow and the likely fix is known but
unverified, rather than that btrfs fixes it.
2026-08-24 00:07:22 +02:00
jschoubben a253afe020 Scenario lifecycle, and how two scenarios coexist
ADR 0032: a scenario is a closed address space. Every segment materialises
as its own isolated link belonging to one instance, so two scenarios raised
from the same declaration hold the same addresses and never meet. The
declaration keeps its literal addresses and they mean what they say —
allocating from a pool would have made them a fiction, so a scenario
reproducing a specific topology would stop reproducing it.

The constraint that follows shapes everything: the lab never reaches into a
scenario over IP. It talks to machines through the virtualisation layer's
own channel. If it reached them by address, the workstation would need a
route into each scenario, and two carrying the same prefix would give it
two routes to one destination — failing not with an error but by one
scenario's traffic arriving in another.

That also makes reachability an honest question. Can this machine reach
that one is asked from INSIDE, by executing on the first, rather than
probed from a workstation that is not on the network and whose opinion
would be a different question with a misleadingly similar answer.

The lifecycle itself: six verbs, of which raise and destroy are enough to
be useful and the rest are what make repetition cheap. Raising is
convergent rather than incremental, because a lab behaving differently
from the thing it tests teaches the wrong habit.

A failed raise leaves the wreckage standing. Tearing down on failure
destroys the only evidence, which is backwards — a scenario that failed to
raise is more interesting than one that succeeded.

Snapshots are whole-scenario. Per-machine would be cheaper and wrong: the
mesh keeps state spanning nodes, so restoring one machine while its peers
move on produces a mesh that has never existed, and faults found there
would be artefacts of the lab.

Closes the declaration's open question about running several scenarios at
once.
2026-08-23 23:53:00 +02:00