Merge branch 'issue/021-provider-port-published-on-loopback' into design/bootstrap-is-a-pivot

# Conflicts:
#	03-DESIGN/01-to-be/04-lab-installation.md
This commit is contained in:
2026-09-11 00:45:59 +02:00
188 changed files with 14552 additions and 2968 deletions
@@ -1,11 +1,12 @@
---
topic: the mesh
status: accepted
date: 2026-08-22
deciders: jochen
reconstructed: false
---
# 15. The mesh brokers capabilities; nodes host; agents think
# 1. The mesh brokers capabilities; nodes host; agents think
## Context
@@ -69,6 +70,52 @@ invariants were found violated simultaneously (see Consequences).
**The mesh brokers capabilities. Nodes are places where work runs. Agents are personas
that think and act.** Everything else supports one of those three.
### What this is, plainly — and what "mesh" does not mean
*Written 2026-08-29, from working through connectivity and asking whether the word still fits.*
Four layers. Naming them honestly is worth more than the word on the tin:
| | |
|---|---|
| **machines are linked by a private network** | and every machine reaches every other over it |
| **one node holds knowledge of all of them** | the control plane, and only it |
| **modules are how anything is built and delivered** | this *is* the CI/CD, not something beside it ([ADR 0010](0010-delivery.md)) |
| **every node runs a session you can message** | a feature of the node, remembering across callers; any node can message any other ([ADR 0004](0004-a-node-and-how-it-joins.md)) |
| **workers are hired onto nodes to do tasks** | employees, with a lifecycle — a different thing from the row above ([ADR 0003](0003-agents-are-persistent-employees.md)) |
**The last two rows are not the same thing and the vocabulary of one does not describe the other.**
A node's session comes with the machine: nobody hires it, it holds no tasks, it is never
reassigned, and it goes when the node leaves. A worker is an employee — named, hired, drained,
retired, movable. They are built from the same parts and run on entirely different terms, and
collapsing them is how the employee vocabulary ends up stretched over something it does not fit.
**Neither makes the node itself a thinking thing.** A node is a machine; both of these run *on*
one, which is why *a node does not authenticate to a model provider, agents do* is unaffected by
either.
**Where the value is** is that both can reach across the whole set: a shell, a service, a file, or
simply a question to another node. Not machines that can be configured centrally, which is
ordinary, but a set of machines that can be worked across as though they were one.
**This is not a mesh in the peer-to-peer sense and will not become one.** The word describes what
machines can reach, not how they are governed:
| | a mesh? |
|---|---|
| what a machine can reach | **yes** — genuinely any to any |
| how the traffic travels | no — anything crossing sites transits the hub |
| who decides | no. One node, declared |
**And *master* overstates it in the other direction.** A master implies the others need it in
order to function. They do not: every node holds what it was last told and runs from that copy
*always* — not as a fallback, as the only mode it has. So the control plane being gone is every
node in the ordinary disconnected situation at once, and **what is lost is change, not
operation.**
The accurate phrase is **one authority, no failover**, and both halves are deliberate
([ADR 0006](0006-the-substrate-and-the-control-plane.md)).
### Nodes and agents are decoupled
A node is a place where an agent can run — that is the entire relationship. There is no
@@ -182,7 +229,7 @@ existing pipeline. Nothing here requires a flag day, and nothing here is cheap.
- `modules/hal/sdk/src/feature-handlers/index.ts` — `FEATURE_HANDLERS`, the fixed handler
array that makes a feature a singleton per module
- `modules/postgres/tools/index.ts` — the adoption path that rotates a shared credential
- Mediahuis `papa-hq`, ADR 0009 *Composable, independently-shippable modules* — the
- Mediahuis `papa-hq`, ADR 0020 *Composable, independently-shippable modules* — the
constraints that make a unit independently shippable, applicable unchanged to features
- impire.io / soulstream — *the record* as integration substrate, personas over services,
and "cheap awareness and expensive thinking"
@@ -1,73 +0,0 @@
---
status: accepted
date: 2026-03-14
deciders: jochen
reconstructed: true
---
# 2. Everything is a module, and one manifest describes all of them
> Reconstructed after the fact from the evidence cited below.
## Context
The mesh carries several kinds of thing: containerised services with data and ports, pure
capability providers with no service at all, and bare markers whose only content is that a
node has them. Before this decision these were separate concepts with separate handling —
the earlier vocabulary was *capabilities*, and services were installed by a different path
than tools.
Every distinct kind of thing needs its own install path, its own change detection, its own
place in the delivery pipeline, and its own documentation. Three kinds means three of each,
and every new feature has to be built three times or, more commonly, once — leaving two kinds
quietly unsupported.
## Considered options
1. **Separate concepts per kind** — a service registry, a tool registry, a node feature flag
list. Rejected: it is what existed, and the cost was paid in every cross-cutting change.
2. **One manifest, kind inferred from directory contents.** Chosen.
3. **One manifest with an explicit `type:` field on every module.** Partly adopted — a service
still declares itself — but the general rule became inference, because a declared list and
the directory it describes drift, and the directory is the one that is true.
## Decision
Everything the mesh installs is a **module**: a directory with a manifest. The manifest
declares identity, environment variables, what the module provides, what it requires, and how
it is exposed. What kind of module it is follows from what the directory contains:
| Contains | Is |
|---|---|
| a compose definition | a service |
| a tools directory | a capability provider |
| a daemon or unit directory | a long-running process |
| a configs directory | a source of managed files |
| nothing but a manifest | a flag — presence is the whole content |
A module may be several of these at once. Each is a **feature**, and the delivery pipeline
addresses features, not modules.
The mesh's own components are modules on exactly these terms. They get no privileged install
path, no separate registry, and no exemption from the pipeline.
## Consequences
- One mechanism to learn, one to document, one to fix. A pipeline improvement reaches
everything the mesh carries.
- Dogfooding stops being a discipline and becomes structural: if the mesh's own components
need an exception, the machinery is unfinished, and that is visible immediately.
- Feature detection from directory contents means a directory rename silently changes what a
module *is*. This has bitten repeatedly — a hook named for a feature the module does not
have is skipped without complaint.
- The manifest becomes load-bearing and grows. It is now the largest single point of
coupling in the mesh.
## References
- `Rename capabilities → modules across the entire codebase`, 2026-03-14.
- `Merge fail2ban, ufw, firewall apps into modules`, 2026-03-15 — the first modules to arrive
by conversion rather than by creation.
- Knowledge base: `modules`, `modules/manifest-reference`, `conventions/modules`.
- The rename-breaks-detection shape: `troubleshooting/hooks-named-for-missing-feature`,
`troubleshooting/health-check-tools-index-false-positive`.
@@ -1,11 +1,12 @@
---
topic: the mesh
status: accepted
date: 2026-02-25
deciders: jochen
reconstructed: true
---
# 1. Nodes communicate over a message broker, not over HTTP
# 2. Nodes communicate over a message broker, not over HTTP
> Reconstructed after the fact from the evidence cited below. The decision was taken in
> implementation, not in a record; this document states what was decided and why, not a
@@ -1,11 +1,12 @@
---
topic: the mesh
status: accepted
date: 2026-07-12
deciders: jochen
reconstructed: true
---
# 12. An agent is a persistent employee, not an instance of a pool
# 3. An agent is a persistent employee, not an instance of a pool
> Reconstructed after the fact from the evidence cited below.
@@ -59,7 +60,7 @@ itself is an agent of a kind exempt from the hiring lifecycle.
or another agent is hired — both deliberate acts.
- The transition was not free. Lifecycle columns had to reach every query that selects an
agent, and the ones that were missed failed at the moment of hiring rather than at startup.
- This is the decision [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) generalises:
- This is the decision [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) generalises:
one kind of participant, differing only in modality.
## References
@@ -1,65 +0,0 @@
---
status: accepted
date: 2026-04-02
deciders: jochen
reconstructed: true
---
# 3. The mesh database is the source of truth; the repository is node-agnostic
> Reconstructed after the fact from the evidence cited below.
## Context
Two things must be known to run the mesh: **what exists** — which modules there are, what each
declares, how each is built — and **what runs where** — which node hosts which module, with
which settings, at which version.
The repository is the natural home of the first. It was initially also the home of the second:
per-node directories held that node's configuration, and adopting a machine meant committing
its files. That has three costs. A node cannot be changed without a commit, so runtime state
and source share a review cadence they do not share a rhythm with. Two nodes cannot be
reconciled, because nothing holds both. And the repository becomes an inventory of the
installation, which is exactly the content that cannot be made public.
## Considered options
1. **Per-node directories in the repository.** Rejected — it is what existed. Every binding
change is a commit and a deploy, and the repository accumulates an inventory of one
particular mesh.
2. **Configuration files distributed to nodes and edited there.** Rejected. There is then no
authority: two nodes disagreeing have no arbiter, and drift is invisible until something
breaks.
3. **A mesh database as the single authority, cached locally for resilience.** Chosen.
## Decision
A single database holds every binding: which node hosts which module, at which selection, with
which environment overrides, plus mesh-level settings that all nodes read. The runtime loads
its configuration from that database at startup and falls back to a local cache when the
database is unreachable.
**The repository defines what exists. The database defines what runs where.** No node-to-module
mapping is ever committed.
A node is therefore not described anywhere in source. Bringing one into the mesh is a database
operation.
## Consequences
- The repository becomes node-agnostic, and can be published without disclosing an
installation. This repository's public stance rests on that property.
- A binding changes without a commit, a build, or a deploy.
- The local cache means a node survives losing the database, but a node running from cache is
running from a snapshot with no indication of its age. Divergence is silent by construction.
- The database is the hardest dependency in the mesh. It is also a module, provisioned like
any other, which makes its bootstrap circular — resolved by the first-node initialisation
script, and the reason such a script exists.
- Nothing on a node is authoritative. That is what makes the next decision necessary.
## References
- `Phase 3: rename core modules to hal/ namespace`, 2026-04-02, and the mesh configuration
tables that landed with it.
- Knowledge base: `mesh` — "The repo is node-agnostic. It contains no per-node assignments."
- The stale-cache shape: `troubleshooting/installed-version-and-deployments-are-stale`.
@@ -0,0 +1,263 @@
---
topic: the tiers
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 4. A node, and how it joins
*Consolidated 2026-08-28 from four records.*
## What a node is
**A managed machine inside the mesh.** Not a device that is merely known about, not an
unprivileged something. If the mesh does not manage it, it is not a node — it is a client, a
peer, or a thing on the network, and those want their own names rather than a weakened version of
this one.
**A disconnected node is still a node, in a different situation.** Reachability is **state, not
class**. A machine switched off, roaming, or behind a connection that has dropped has not become
a lesser kind of thing; it has a last-known state and a pending set of declarations.
The distinction people reach for is real, but it is **capability** — what this machine can be
asked to do — and that belongs in the host's profile rather than in the definition of a node.
**This is the rule that does the most work elsewhere.** A single control plane is tolerable
because its absence is every node in the ordinary disconnected situation at once. An episodic
host on a phone is that situation more often. Neither needed a new mechanism.
### A node runs one agent session
*Written 2026-08-29. It runs on every node today and appeared in no record, which is how something
deliberate comes to look accidental.*
**A node is a machine. The session is a feature of it** — one of the things running there, like the
host, like any workload. The node does not think; something on the node does. Which is why
[ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)'s *a node does not authenticate to a model
provider, agents do* holds unchanged: the session authenticates, and it is not the machine.
**It is permanent, and it remembers.** Anything in the mesh can send it a message; it replies; and
what it was asked ten minutes ago is still there next week, alongside what everything else asked in
between — the same way both sides of any conversation remember it.
Its system prompt is the node's **engram** — what makes one node's replies recognisably its own
rather than generic.
**It has its own tools**, and fewer than a session a person is driving directly. So a question can
be answered by going and looking: *what is in our forge*, not only *what is your battery*.
**Messages travel the broker like everything else** ([ADR 0002](0002-nodes-communicate-over-a-broker.md)).
There is no second transport and nothing is dialled.
**Any node can message any node, and this is the one part of the system that is genuinely a mesh**
— symmetric, with no centre. A node that is asked something it does not have can ask another, and
how it passes the question on is its own business: it may say who wants to know, or simply ask. A
person relaying a question makes the same choice, and it follows from the engram rather than from a
message format.
**There is no authorisation between nodes.** Every node is the operator's own, so a message from
one is a message from them, and asking a node something is asking a colleague rather than
presenting credentials. Stated once so it is not discovered later: **the mesh boundary is therefore
the security boundary** — anything inside can reach what any node can reach, which is what puts the
whole perimeter on the token and the overlay
([ADR 0007](0007-connectivity.md)).
**It can be switched off, and switched off it still answers.** A node whose session is disabled
replies saying so, at once, with no model involved — the queue is still read, and the state is the
reply. That is deliberate and it is the same rule the host follows about a service that does not
exist: **absence must never be indistinguishable from a failure to answer.** A node with nothing
there is a silence somebody has to go and diagnose; a node that says *I am switched off* is not.
**One per node, always, and it cannot be moved to another machine.** Two and nothing decides which
replies; none and the node is mute; moved, and one machine is answering as another.
**It is not an employee** ([ADR 0003](0003-agents-are-persistent-employees.md)). Nobody hires it,
it holds no tasks, it drains nothing and it is never reassigned — that vocabulary was written for
workers and does not describe this. It exists because the node does, and it is gone when the node
leaves.
## How it joins
**The host has one behaviour and two sources of declaration.** What differs between the first
node and the fiftieth is not what the host does but where the declaration comes from — and, as
above, that is a situation rather than a class.
| | declaration comes from |
|---|---|
| no mesh reachable | the pinned bundle the host carries |
| mesh reachable | the control plane, over the link |
**The first node is not a different kind of node.** It is a node whose mesh is not up yet. It
applies the bundle it carries, the control plane comes up on top of it, and from that moment it
takes declarations like everything else. **Its specialness is temporary and self-erasing**, which
is what the hand-run bootstrap scripts never were.
**A joining node does the minimum to be reachable and nothing else** — an identity, an address,
and one peer to reach. It does **not** compute the overlay: the whole peer set is derived
centrally and pushed down.
That is also why the migration is smaller than it looked. The hard part of the overlay — every
node's key, address, site and reachability — is only needed to compute the *whole* mesh, and a
joining node needs one peer.
## The link is the security boundary
**Everything reaching a node arrives one way**, and four properties make that a boundary rather
than a pipe.
**It is outbound and node-initiated.** The node dials the control plane; nothing dials a node. Not
only defensive — most nodes sit behind a household connection with no forwarded port, so an
inbound control channel would work for one node and not the rest, and the difference would be
invisible until it mattered. **A node has no listening control surface at all.**
**A node holds its own identity and nothing else.** No shared secret, no credential to anything it
does not own. **Compromise of a node is compromise of that node** — which the current arrangement
does not have, because every node permanently holds the same database and object-store
credentials, and there is no mechanism that rotates one and informs everything holding it.
### What that identity is: a keypair the node generates
*Written 2026-08-29. This is the same rule as the sentence above, and it had been treated as an
open question for weeks because of a word.*
**The node generates a keypair. The private half never leaves the machine. The mesh records the
public half.** Ed25519, the same as the control plane's signing key, in the other direction:
the mesh proves itself to a node by signing, and a node proves itself to the mesh by signing.
**This was never open.** [`08-connectivity.md`](../03-DESIGN/01-to-be/08-connectivity.md) already
says it of the overlay keys, in these words: *each node generates its own keypair, the private key
never leaves the machine, the public key is published to the mesh* — and adds that this **is**
ADR 0004's *a node holds its own identity*, applied. What was missing was applying it to the thing
this record is about.
**The word that caused it:** the lifecycle says a joining node *receives* its own durable identity,
which reads as the mesh issuing something, and then the question becomes *issuing what*. It does
not issue anything. The node arrives holding its identity; what it receives is **being known**.
Enrolment is the moment the mesh writes down a public key it will believe, and the one-time secret
is what buys the right to have it written down.
**Everything above then holds literally.** Nothing is stored that could be stolen and replayed: the
mesh's copy is a public key, so a copy of the mesh's database grants nothing. *Compromise of a node
is compromise of that node* becomes true rather than aspirational, because the only secret on a
machine is the one that identifies it.
### What "connecting to the mesh" is, concretely
*Written 2026-08-29, because it was asked and this record had never said it.*
**One outbound AMQP connection from the node to the broker, held open.** That is all of it. There
is no second connection and nothing is ever dialled *at* a node. Being in the mesh, operationally,
means that connection is up; being disconnected means it is not
([`09-the-node-lifecycle.md`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md)).
**Two different things ride on it, and conflating them is what made this confusing:**
| | what it answers | who issues it |
|---|---|---|
| **an AMQP account** | may this connection be accepted at all | **the mesh, at enrolment** |
| **the node's keypair** | which node is speaking, on every message | **the node**, above |
**The account is the mesh's to issue**, and per node. The broker has to authenticate somebody
before a connection exists, and a shared account would let any node consume another's queue —
which is the shared-credential fault this record exists to remove, reappearing at the transport.
So enrolment creates that node's account and hands it over, and it is rotatable without touching
the node's identity.
**The keypair is not made redundant by it.** With only an account, the control plane would know
which node is speaking *because the broker says so* — and that is the same transitive authority
this record refuses in the other direction. A compromised broker could then attribute reports to
whichever node it liked, and the control plane would act on them. Signing is what removes the
broker from the question in both directions.
**So a node holds two things after enrolment**: a credential the mesh issued for reaching the
broker, and a key it generated itself that the mesh only ever sees the public half of. Both are
its own, neither reaches anything else, and *compromise of a node is compromise of that node*
still holds.
### Its own key, not the machine's SSH host key
Reusing the host key is the obvious economy and it is refused, for reasons that are operational
rather than fastidious:
- **It is regenerated by ordinary events.** A reinstall, an image cloned, `ssh-keygen -A` on a
rebuild — each silently un-enrols the node, and the failure appears as an authentication problem
with no cause anybody changed.
- **It is managed by something else.** Its lifecycle belongs to the machine's SSH daemon, and an
identity the mesh depends on should not rotate on a schedule the mesh does not know about.
- **Not every node has one.** A partial host has no SSH daemon
([ADR 0005](0005-the-node-host.md)), and an identity scheme that excludes a supported kind of
node is not one.
**The mesh should still know the host key** — it knows every node, so it can distribute host keys
the same way it distributes authorised keys
([ADR 0006](0006-the-substrate-and-the-control-plane.md)), and node-to-node SSH stops depending on
trust-on-first-use. That is the good half of the idea, kept.
**Authority is mutual.** The node proves it may join, and the control plane proves it is the
mesh. One-way is not enough: the host applies whatever the link delivers, so a node that cannot
tell the mesh from something impersonating it will apply that something's declarations.
**What may be pushed is bounded by form, not by trust.** Declarations of known shape, never a
command to run. Stated honestly, **this bounds form and not impact**: a compromised control plane
can declare harmful state and the host will apply it faithfully, because that is what it is for.
What the property buys is that the blast radius is *describable* — exactly what the declaration
language can express, which can be reviewed. An arbitrary command channel has no such bound.
## The enrolment token carries the mesh
Mutual authority needs the node to verify something before it trusts anything, and that is a
circle: verifying the mesh needs the mesh's certificate authority, and obtaining one means
trusting whoever hands it over. There is a second circle beside it — a node must reach the mesh
before the mesh has configured it, so it can resolve no mesh name.
**Both are the same shape: a node needs a fact about the mesh before it has any trustworthy way
to obtain one.** So that fact arrives by a path other than the mesh.
**The token carries five things**, and it is the only thing a joining node needs:
| | |
|---|---|
| **who it is** | the name the mesh calls this machine |
| **where** | the broker's **address**, not a name — there is no resolution yet, and this is why none is needed |
| **what it is connecting to** | the fingerprint of the broker's certificate |
| **who it will believe** | the control plane's signing identity |
| **the right to join** | a one-time secret, useless once used and useless after it expires |
*The first row was added 2026-08-30, from raising a mesh end to end for the first time.* It reads
like an oversight and is not: **the node cannot work its own name out.** The name is the mesh's,
chosen when the record was created, and the broker account the node authenticates as is named
after it — so it must be known *before* the mesh can tell the node anything. It is not a secret,
and whoever issues the token already has it.
Without it, enrolment fails at the broker with an empty username and a message about credentials,
which points at everything except the cause. **A missing fact that surfaces as an authentication
error is worse than one that surfaces as a missing fact.**
**Carried out of band**, by the person adopting the machine. That is what breaks both circles:
its authenticity comes from the channel it travelled, not from anything the node can check
afterwards. **Trust on first use, with the first use moved out of band** — the difference between
a pin and a guess.
**The endpoint and the authority are two identities.** A node connects to the broker and takes
instruction from the control plane behind it. Pinning only the broker would make the control
plane's authority *transitive*, and a compromised broker could then forge declarations — which,
since the host applies whatever the link delivers, is the whole machine. So the transport is
verified once at connect, and **each declaration is verified by its signature, every time**.
**What this settles:** the mesh's certificate authority is not a bootstrap concern — it certifies
internal names once a node is a member. Nothing needs name resolution before the link. And
nothing is placed on disk beforehand except the token, which is the first moment *a node holds
only its own identity* becomes true rather than aspirational.
## Consequences
- **Declarations must be signed**, and the host must tell *this is not from the mesh I joined*
apart from *this is malformed*. Rotating the signing identity is a fleet-wide operation with an
overlapping rollover, and that is the cost of not trusting the broker.
- **The token becomes security-critical**, because it carries the pin. Tampering substitutes the
mesh — which is strictly better than the alternative, where there is nothing to tamper with and
the node trusts the first answer unconditionally.
- **A rejoining node is ordinary.** There is no long-lived secret to recover, so a node that lost
its identity gets a new token.
@@ -1,67 +0,0 @@
---
status: accepted
date: 2026-04-06
deciders: jochen
reconstructed: true
---
# 5. Capabilities are provisioned on declaration, not configured by hand
> Reconstructed after the fact from the evidence cited below.
## Context
Most modules need something another module holds — a database, a cache, a bucket, a message
vhost, an identity client. Wiring that by hand means creating the resource, creating a user,
generating a credential, putting it in the consumer's configuration, and repeating all of it
on every node the consumer runs on.
Every step is a place to make a mistake that surfaces much later, and the credential ends up
written somewhere it can be read.
## Considered options
1. **Manual setup, documented.** Rejected. Documentation of a manual procedure is a
description of the mistakes people will make.
2. **A shared credential per resource type**, distributed to all consumers. Rejected: no
isolation, and rotation becomes a mesh-wide outage.
3. **Declared requirements, satisfied by the provider module.** Chosen.
## Decision
A module declares what it **provides** and what it **requires**. A requirement names the
provider, the resource type, optionally a name and a target node, and a mapping from the
resource's connection fields to the consumer's environment variables.
The mesh satisfies it: a provisioner belonging to the provider creates the resource and its
credential, records the grant, and writes the mapped values as database overrides. The
synchroniser from [ADR 0004](0004-managed-files-are-generated-never-edited.md) then
materialises them. Neither the credential nor the topology is ever written by hand.
A requirement may name a provider on another node. The grant records consumer and provider
nodes separately, so cross-node wiring is the same declaration.
## Consequences
- **Provisioning becomes a core concern of the mesh, not plumbing.** A module asks for a
capability; where it lives is the mesh's problem. This is the property
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) later builds the whole domain model
around.
- Credentials are never authored, so they are never authored badly, and they are never in the
repository.
- Each consumer gets its own credential, so revocation is per-consumer.
- Rotation is where this bites. A shared secret rotated for a new consumer invalidates the
peers holding the old one, and this has taken the mesh down. The declaration model makes
granting easy and says nothing about fan-out.
- A module with no requirements skips the stage entirely, which is correct and also means the
absence of provisioning is indistinguishable from provisioning that did not run.
## References
- `Remove shell/ helper library; split brain into independent workspaces`, 2026-04-06 — the
provisioner daemon becomes its own component.
- `Coordinator refactor: centralize pipeline orchestration`, 2026-04-04 — the provision-then-
environment-then-start sequence becomes the coordinator's.
- Knowledge base: `provisioning`, `provisioning/requires`.
- The rotation failure: `troubleshooting/provision-rotation-invalidates-peers`,
`troubleshooting/provision-adoption-rotates-live-credential`.
+155
View File
@@ -0,0 +1,155 @@
---
topic: the tiers
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 5. The node host
*Consolidated 2026-08-28 from eight records. Tier 0 is one component and was decided over a
week; the reasoning is kept, the fragmentation is not.*
## It applies; it does not decide
**The host makes a machine match what it was told, and never works out what that should be.**
This is the line the whole tier rests on, and it is not about privilege — it is about what a
single machine can *know*. Deciding needs knowledge the machine does not have: which nodes should
run a store, which peers belong in an overlay, whether a node has been unreachable for a week.
Anything needing a second node is the control plane's.
The practical form: **the host never queries the mesh's database and holds no credential to it.**
Two modules in the current mesh do, and they are the reason every node permanently carries a
database credential.
## It depends on nothing that must be installed first
**A statically linked binary. Copy it onto a machine and run it — that is the whole
installation.** Written in Go, because the job is system-level and because a runtime that must be
installed first would make the host depend on the thing it exists to install.
**What it needs from the machine is not a dependency in this sense.** An init is not installed;
it is what the machine already is. A package manager is the distribution. Those are what a
machine *is*, not what must be put on it before the host works.
## It is built per operating system
**`systemd` and `pacman` are the Arch host's implementation, not abstractions the mesh grows.**
They are not independent choices: a machine has pacman *because* it is Arch, and the package
manager, service manager and packaging format arrive together as one decision somebody made at
install time.
```
mesh-host-arch pacman · systemctl · a container runtime
mesh-host-alpine apk · rc-service
mesh-host-android neither — a partial host
```
**Abstracting them was rejected on correctness, not effort.** The service applier reads systemd's
`LoadState` to tell *not installed* apart from *stopped* — which is what stops it reporting
absence as success — and OpenRC has no equivalent. An interface spanning both must drop it, and
the lowest common denominator is exactly where that fault lives.
**Almost all of it is shared.** The declaration vocabulary, the store, the apply loop, the
read-back discipline, the refusal model and the link are portable. Two appliers differ.
**A host that cannot implement a shape refuses it.** Android has no package manager it may drive
and no init it may register with, so it implements `file`, `directory` and `action` and refuses
the rest — the same refusal an unknown type gets, with a different reason. Those three are the
portable floor, and they are what makes a partial host a real thing rather than a broken one.
## It is a root service, and it never manages its own unit
**Root**, because no useful part of the job is unprivileged: it writes under `/etc`, installs
packages, manages units and runs containers.
**It cannot run in a container**, and the reason is decisive rather than stylistic: installing the
container runtime is a step of the bootstrap, so a host inside a container would need the thing
it exists to install. Everything above tier 0 is a container; the host is not. That split is the
tier boundary made concrete.
**The installation owns the host; the host owns everything else.** It manages `service` resources
and its own unit is one — the temptation is obvious and it ends with a host stopping itself half
way through an apply, leaving a machine with nothing running to fix it.
## An init is asked for one thing
**Start this at boot.** That is all, and every init can express it — systemd, OpenRC, runit, s6.
**Everything else is a launcher the host ships**, which supervises it: restart it when it exits,
count consecutive failures, roll back after too many, halt after that. Policy in a unit file can
only be read and hoped for; a script with a counter can be tested, and this is the one piece that
must work on a machine where the host does not.
**The launcher does not exec the host, it supervises it** — so restarting is ours rather than the
init's. The cost is signals: a supervisor that exits while its child runs leaves the host to be
*killed* rather than to *stop*, and an apply interrupted that way is the half-configured machine
this design is about. So it traps the shutdown signal, passes it down, and waits.
**A clean exit is the upgrade path**, and it is the easiest thing to get wrong — twice now. The
host stands aside for a new binary by exiting zero, so anything supervising must restart on a
zero exit and must not count it as a failure.
**Recovery is local, and detection is the mesh's.** Nothing dials a node and a host that cannot
start cannot report, so the node must recover itself. But a local supervisor sees one process
failing and cannot tell a broken machine from a broken release — only something watching every
node can, which is why a host rollout is staged and stops when nodes go quiet.
## A host may be episodic
**Resident or episodic, and both are hosts.** A phone has no init to register with and nothing
worth supervising, because a supervisor would be killed alongside what it supervises. So it runs
when the platform allows and is killed when the platform wants the memory — **and that is
disconnection**, which is already an ordinary situation.
It needs no keep-alive and no new mechanism: the store is already authoritative while
disconnected, reconcile already happens on start, and *last heard from* is already reported
rather than alarmed on. An episodic host cannot be the first node, because every bootstrap step
is a shape it refuses.
## What a declaration is
**An ordered list of resources the host owns.** JSON, because Go's standard library carries a
JSON parser and no YAML, and the one binary whose argument is that it needs nothing must not
gain a parser to buy authoring comfort in a machine-written document.
**Ordered, because ordering is a decision.** The host does not sort and does not resolve
dependencies — that would be deciding, and deciding the thing most likely to differ between what
the control plane intended and what the machine does.
**Every resource has a stable identity** — a name the control plane keeps across declarations, not
a position and not a hash of content. It is what lets the store say *this is the same resource I
applied last time*, which is what makes removal possible at all.
**Unknown is refused, whole.** A field the host does not know is something the control plane
believes it asked for. A declaration naming one is rejected entirely, naming every problem at
once — a host that applied the parts it understood would leave a machine that looks configured
and is not.
**Six shapes:** `file`, `directory`, `service`, `package`, `container`, `action`. Every addition
widens what a compromised control plane can express, so the list is a security artefact and grows
deliberately.
### The bundle may carry actions; the link may not
An `action` runs a command, and the host never learns what it means. It is needed because the
bootstrap creates a database before there is any mesh to ask for one, and the host must not learn
what a database is.
**Permitted from the bundle, refused from the link**, and the asymmetry is the whole point: a
bundle arrives *with* the binary, so anyone able to put a hostile action there could have put it
in the host itself — refusing it buys nothing and costs the bootstrap. The link is a separate
party, reachable separately, and an action there is an unbounded blast radius.
**An action must carry its own verification**, which is also its idempotency check. The host does
not know what a database is, so *is it already there* is a question only the declaration can ask.
## Consequences
- **The migration is smaller than it looks.** A joining node never needs mesh-wide state — it
needs an identity, an address and one peer, and the rest arrives as declarations.
- **What is applied is recorded after it works, never before.** A failed apply leaves the machine
in whatever state it reached, and nothing must claim otherwise.
- **A second operating system is additive**: two appliers and a four-line init file.
@@ -0,0 +1,279 @@
---
topic: the tiers
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 6. The substrate and the control plane
*Consolidated 2026-08-28 from six records. Extended 2026-08-29, by building it: the language, and
what must be running before the control plane starts — which this record had left not established
and could not have settled the way it was asking.*
## The control plane is what needs to know about more than one node
That is the whole test, and it follows from the host applying rather than deciding: **deciding
needs knowledge a single machine does not have.**
| question | whose |
|---|---|
| write this file, with this content, with this mode | the **host** |
| which nodes should run the store | the **control plane** |
| is this unit running | the **host** |
| which peers belong in this node's overlay | the **control plane** |
| has this node been unreachable for a week | the **control plane** — nobody else is watching |
**Anything a single machine could answer alone is not the control plane's.**
### Seven contexts and one interface
**inventory, config, connectivity, provisioning, delivery, observability, identity** — plus
`api`, the one interface every surface speaks to. Each earns its place by the test above rather
than by being ours.
**`work`, `knowledge` and `stream` are mesh-hosted applications, not control plane.** A task does
not need to know a node exists. *Being ours does not make something infrastructure.*
**`identity` owns SSH access.** *Written 2026-08-29, on noticing it was assumed everywhere and
stated nowhere.* SSH appears three times across this design and every time as something that
*uses* the overlay — "the way back in", "every node reaches every other: SSH, services, ordinary
traffic" — while nothing said who hands out the keys. Nobody else could: the mesh is the only
thing that knows which humans and agents exist and which nodes they may reach, which is
`identity`'s definition. The node end already works, since an `authorized_keys` file is a file.
It is three questions wearing one name, and only two of them are the mesh's:
| | |
|---|---|
| **humans** | their key, on the nodes they are allowed on |
| **agents** | the same, with a lifetime — and revocation that has to actually work |
| **an agent reaching another node** | **this is the point of the mesh, not an exception to it** — see below |
**There is no such thing as node-to-node SSH here, and that is a clarification rather than a
restriction.** The actor is always an **agent**; a node is only where it happens to be running —
[ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md): *a node is a place where an agent can
run, that is the entire relationship.* An agent hired onto one node reaching another to do work is
the capability the whole arrangement exists to provide.
**The credential is the agent's, never the node's.** It lives in the agent's own credential
directory ([ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)), so a node's
`authorized_keys` lists **agents** and never nodes. Three things follow, and they are why this
shape is better rather than merely allowed:
- [ADR 0004](0004-a-node-and-how-it-joins.md)'s *a node holds its own identity and nothing else*
stays true — no node holds a key that reaches another node;
- a compromised node costs the credentials of the agents that were on it, not a way into
everything;
- **who may reach what is a mesh-wide fact**, which is exactly why it is `identity`'s and not
something arranged locally.
**And it does not conflict with the host having no inbound control surface.** That rule is about
how a node's *declared state* changes: over the broker, never by being dialled. An agent with a
shell is not the mesh reconfiguring a machine — it is what a person with a terminal has always
been, and this design already depends on it working
([ADR 0007](0007-connectivity.md): the overlay is *the way back in*). What such a session can leave
behind is drift, and drift is what reconciliation is for.
**Where the record lives is deliberately open.** Contexts integrate through it, which makes it
load-bearing, and putting it in the substrate risks recreating the circularity the tiers just
removed. Listing it as an eighth context would settle by naming what has not been settled by
arguing.
### One node runs it, and nothing takes over
**Declared, never elected.** No promotion, no quorum, no fencing, no split brain — none of it
built, so none of it can be subtly wrong.
#### The option that would make it a real mesh, and why not
*Written 2026-08-29. It had been rejected by never being written down, which is the weakest way
to reject anything.*
A genuine peer-to-peer mesh means **no node is special**, and that has a concrete price:
- every node holds the **whole inventory**, so there is a replication process between them;
- replication needs a writer, so one node is elected **master**, and something promotes a new one
when it drops — Redis Sentinel and its whole family of problems;
- and it still would not deliver what the name promises, because **application databases are not
replicated.** A workload's store lives where the workload lives.
That last point is the one that settles it. To make the mesh genuinely peer-to-peer we would have
to become **a replicated database system for everything running on it** — not for our own
inventory, for every consumer's data too. That is a product, and a much larger one than the thing
it would be supporting.
**So there are three central roles, not one**, and it is worth seeing them separately because
only the third costs operation:
| | its loss costs |
|---|---|
| **the control plane** | nothing can be *changed*. Nothing stops running |
| **the broker** | nothing can be told anything, or report anything |
| **the hub** | nodes in different places **cannot reach each other** ([ADR 0007](0007-connectivity.md)) |
**Whether these are one node is not decided here.** All three must be dialable by every node, which
pushes toward one; nothing says they must be.
**That is sound rather than merely cheap**, because the design already tolerates its absence by
construction: a node reconciles from its own store and never needed to ask anybody to hold the
state it was last given. **The control plane being down is not a new failure mode — it is every
node in the ordinary disconnected situation at once.** What is lost is *change*, not *operation*.
The honest half: this node is a single point of failure, recovery is **restore rather than
failover** — which makes backup the availability mechanism rather than hygiene — and
**certificate renewal is the clock.** An outage outlasting a renewal window expires every public
name, which turns an inconvenience into an outage on a timer. Nothing measures that today.
## The authority is the control plane, not a database
**There is no single mesh database.** Each context owns its store exclusively, and *the mesh
database* names a thing that will not exist.
**No node reads any of them** — not for writes, not for reads. A node is *told* what to own, over
the link, in a bounded vocabulary; it **states** what it applied, and the owning context writes.
The difference is the security boundary: something that can write cannot be prevented from
writing anything.
**A node runs from its own store always, not as a fallback.** The current arrangement's nastiest
property is that *a node running from cache looks identical to a node running from the database*,
with no age on the cache and nothing reporting divergence. Under this there is no second mode to
be mistaken for the first.
**What survives from the original decision:** the repository defines what exists, the mesh defines
what runs where, and no node-to-module mapping is ever committed. That is what makes the
repositories node-agnostic and why anything about the mesh can be published at all.
**The error underneath was a category error**: *source of truth* named a storage location when it
meant an **authority**. Once the store is the answer, *which database* becomes the question, and
shared schemas follow.
## The substrate is what the control plane consumes and cannot grant itself
Every module needing a database asks provisioning for one. The control plane needs a database too
and cannot ask itself, because it is not running yet. **That circularity is the definition**, and
anything on the wrong side of it is raised from the bundle the host carries.
| role | product | |
|---|---|---|
| relational store | **PostgreSQL** | its own state lives there |
| message bus | **LavinMQ** | it cannot grant itself a virtual host — and **precedes it**, below |
| object store | **MinIO** | it cannot grant itself a bucket |
| image registry | **an OCI registry** | it cannot grant itself a repository |
| identity provider | — | **conditional**: substrate only if the control plane delegates authentication, which is undecided |
**The role and the product are both written.** The role is what the argument turns on; the product
is what gets installed and pinned, and a design that names only the role does not record that the
choice was made. **The dependency is on the protocol** — AMQP, S3, OCI — which is what keeps
naming them safe. The store is the exception: the provisioning model uses databases, roles and
schemas as PostgreSQL means them.
**A container runtime is detected, not chosen** — docker or podman, because a machine that
already has one keeps it. Only the version probe differs between them; the behavioural difference
(podman has no daemon, so containers do not return after a reboot unless a unit is enabled)
belongs in the declaration rather than the host.
**Being substrate and being in the bundle are different questions.** PostgreSQL and LavinMQ must
precede the control plane; the object store and the registry are substrate by role and ordinary by
delivery, provisioned once there is a control plane to do it.
### Why the broker precedes it too
Written after the fact, because this record first left it *not established* and framed it as
turning on whether the control plane's own contexts talk to each other over the bus.
**They do not** — they are one process and dispatch internally. Under that framing the broker is
provisioned like anything else and the bundle stays at one image.
**The framing cannot answer the question.** What decides it is not how the contexts reach each
other. It is how the control plane reaches a *node* — and that is settled above: never except over
the link, and the link is AMQP ([ADR 0002](0002-nodes-communicate-over-a-broker.md),
[ADR 0004](0004-a-node-and-how-it-joins.md)). So:
```
the bundle raises the control plane
the control plane provisions the broker ← by telling a host to run it
telling a host happens over the link
the link is the broker
```
**And it is not avoided by the first node being local.** `enrol` dials the broker at the address in
its token, which is the first node's own third step. A machine that raised the mesh still joins it
the ordinary way, and that was deliberate — its specialness lasts two commands. Making it join by
some other route would buy a smaller bundle by giving up the property the design was built to have.
The broker precedes the control plane for the same reason PostgreSQL does: **the control plane
cannot grant itself the thing it would need in order to grant it.**
**What it costs.** Two images rather than one, against the wish above to keep the bundle at roughly
one so a person can read it — two is still readable, four would not be. And three things that are
not images, each an action the bundle declares and the host runs, the way the database already is:
a virtual host, a credential on it, and **a certificate**. That last is the awkward one: a token
pins the fingerprint a host must expect *before it sends anything*, so the broker needs a
certificate at a moment when there is no mesh to issue one and no public name to obtain one for.
Self-signed and pinned is the shape that fits; how it is later replaced by the certificates in
[ADR 0007](0007-connectivity.md) is not decided here.
## The control plane is written in Go
The same language as the host, so tiers 0 and 2 are one language and not two.
The reason that decides it is not familiarity. **Its image is pinned by digest in the bundle**,
which means it is fetched and run on a machine where no mesh exists yet — nothing to check it
against, nothing watching, and a person expected to have read the bundle and believed it. A
statically linked binary makes that image the program and nothing else: no interpreter, no package
tree, no transitive dependency that arrived because something needed a date library. Everything
under that line is something somebody would have to audit, on the one image the whole mesh is
raised from.
A second reason, smaller and still real: the control plane runs a reconcile loop of its own
([ADR 0010](0010-delivery.md) — artifacts against source, as the host reconciles machine state
against declarations). Two loops of the same shape are cheaper to hold in one head when they are
also the same language.
**The option rejected** is TypeScript, matching the lab and the surfaces that will speak to this.
The argument for it is that tier 3 is web and CLI, so a TypeScript control plane would share types
with its callers rather than generating a contract. True, and it does not reach far enough:
`mesh-sdk` is *contracts shared across tiers* and **tier 0 is Go**, so the contracts cross a
language boundary whatever tier 2 is written in. The choice is between generating them for one
consumer or for two.
**What it costs, plainly:** the control plane can import nothing that exists today, and a person
moving between tier 2 and tier 3 changes language. Neither is recovered later — the language is the
most expensive thing in this record to reverse.
## The installer fetches what it pins
`substrate.lock` carries **references, not payload** — an image name and a **digest**, fetched at
apply time. A tag moves; a digest does not, and reproducibility comes from pinning the identity of
a thing rather than carrying its bytes.
The assumption that a machine might have no network came from the lab and was wrong: a machine
being adopted has one, and the sealed case is the lab.
**The lab places images by raising a registry inside the scenario**, which is what a real node
pulls from anyway — so it tests the real path rather than a stand-in for it. The digests that
registry serves are its own, and that satisfies this rule: what is required is a reference that
is **exact and cannot move**, and one it assigned is both. Assuming an upstream digest had to be
preserved is what made this look impossible for a while
([04-ISSUES/009](../04-ISSUES/009-a-digest-pinned-image-cannot-be-placed-in-the-lab/00-report.md)).
**Its contents are per operating system** even though its mechanism is not — package names, unit
names and service names all differ, so an Arch host embeds an Arch bundle.
## Consequences
- **The bundle stays small and reviewable.** A list of pinned references is something a person can
read; a bundle containing images is not.
- **An apply can fail because something is unreachable**, which a self-contained artifact could
not. That must fail *legibly*, naming what could not be fetched and from where.
- **Cross-context reporting is harder, and that is the point.** Anything wanting to see across
contexts consumes their events or calls their interfaces.
- **A queue with no limit grows until the broker's disk is full**, and the broker is what every
node depends on. The bound is per queue and is not decided.
- **The bundle carries two images and four actions**, and the substrate bootstrap grows a step.
- **Nothing in the first node's path is special-cased.** Enrolment is walked on node one.
- **The broker's certificate at bootstrap has no answer yet**, and is named as unfinished rather
than assumed. It is the first thing that will be wanted when the link is built.
- **The language cannot be revisited cheaply.** It is the one line here close to irreversible.
+191
View File
@@ -0,0 +1,191 @@
---
topic: the tiers
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 7. Connectivity
*Consolidated 2026-08-28 from three records. Overlay, resolution, exposure, filtering and
certificates are one design.*
## Why it is control-plane work
Apply the test — *everything that needs to know about more than one node* — and not one of the
five can be answered by a machine on its own:
| | needs to know |
|---|---|
| **overlay** — who peers with whom | every node, and which can be dialled |
| **resolution** — which name is which node | every node |
| **exposure** — which public name reaches which container | which node is publicly reachable |
| **filtering** — which port is open, to whom | what is assigned here, and the overlay's shape |
| **certificates** — who may present which name | which name belongs to which node |
That is exactly what the current arrangement gets wrong, by computing all five on the node from a
direct database connection. Two modules do this, and they are the only two left holding a
credential to the control plane's database.
**The shape of the fix, once for all five:** the connectivity context computes the configuration;
it arrives over the link as `file` resources; the service reads files and knows nothing about the
mesh. **This costs no new host vocabulary.**
## A route is a grant
**Ingress is not substrate.** The control plane does not need a route to start — it listens
locally — and no node needs one to reach it, because the node dials out and has no listening
control surface. It grants itself a route afterwards, the way it grants itself a bucket.
The strongest objection deserves stating: the `api` is the one interface every surface speaks to,
so eventually it *does* want a public name. But **wanting one later is not needing one to
start**, and that distinction is the entire substrate test.
**A module that must be reachable declares it needs a route; the proxy provides one.** Ordinary
instantiation, with the direction mirrored — the consumer supplies a target and receives a name.
**Exposure is three facts at two scopes**, which is why it cannot live on the node:
| the fact | scope |
|---|---|
| the public name resolves to an address | **mesh** — which node is publicly reachable |
| a certificate valid for that name exists | **mesh** — issued once, used on one node |
| the proxy maps that name to that container | **node** |
**A node without a public address is proxied by one that has**, across the overlay. Most nodes sit
behind a connection with no forwarded port, so exposure cannot assume the workload's node is
reachable.
## Reachability is declared, not inferred
The overlay's peer graph is computed from whether a node can be dialled, and that was inferred
from a regular expression over the address. **The address is evidence of reachability; it is not
the fact**, and the gap has already cost:
| address | the regex says | actually |
|---|---|---|
| `100.64.0.0/10` — carrier-grade NAT | **public** | **not reachable.** An endpoint is written to an address nothing can reach |
| any IPv6 address | public | the test is v4 shapes only |
| a routable address behind a closed firewall | public | not reachable |
| a documentation range standing in for a public segment | private | reachable — this is the lab bug |
**A test environment having to choose its addresses to satisfy a regex is the regex telling us it
is not a fact.**
So: **an endpoint, or none** — declared. And **the hub is declared, never derived from an address
prefix**, because an election decided by the first four characters of an address fails silently,
cannot be queried, and makes a renumbering an outage.
**The address remains evidence and stops being the fact.** Where an observed endpoint disagrees
with a declared one, the disagreement is a **reportable condition**, not a silent correction.
**What does not change** is the lesson underneath: role does not imply reachability — a
home-hosted node is a server that cannot be dialled. This keeps that and stops encoding it as a
pattern match.
### Some nodes must be reachable, and this had not been said
*Written 2026-08-29, on being asked and finding no answer.*
Everything above treats reachability as a **fact to record** — which node can be dialled, so that
exposure and certificate issuance can be placed. It never said the converse, and the converse is a
hard requirement:
| role | dialled by | so it needs |
|---|---|---|
| **the node running the broker** | every node, outbound ([ADR 0002](0002-nodes-communicate-over-a-broker.md)) | to be reachable from wherever nodes are, at a **stable address** |
| **the hub** | every node not co-located with its peer | the same |
**Across the internet, "reachable from wherever nodes are" means publicly reachable.** For a mesh
confined to one network it does not — the requirement is about the nodes that exist, not about the
public internet.
**Stable is the sharper half.** A token carries the broker's *address, not a name*, because there
is no resolution before joining ([ADR 0004](0004-a-node-and-how-it-joins.md)). A broker node whose
address moves invalidates every token issued for it, and a node that was disconnected across the
change cannot get back.
**A mesh whose nodes are all behind NAT cannot be raised.** That is a real precondition and it
belongs with the others rather than being discovered.
### The link stays on the underlay, and that is a repair channel
The obvious objection is that all traffic should run over the overlay. Nearly all of it does — SSH,
services, node to node — and the exception is each node's own outbound link to the broker.
**At join time it is forced**: a node has no overlay yet, so it cannot use one to ask for one.
**Afterwards it is a choice**, and the reason is that the link is how a broken node is fixed. A
repair channel carried over the thing being repaired is not a repair channel: a node whose only
path home was the overlay is gone the moment an overlay declaration is wrong.
**What it does not cost is confidentiality.** The link is already authenticated and encrypted
against a pinned fingerprint ([ADR 0004](0004-a-node-and-how-it-joins.md)), so moving it onto the
overlay would not protect traffic that is unprotected today.
## A filter rule names its source
`scope: public` is declared in five manifests, is part of no rule type, and is **referenced by no
code**. So five manifests appear to restrict a port and restrict nothing — on the modules most
worth restricting.
**A rule names its source. `from:` is the only way to scope one, and a rule without one is open**
— which it must say plainly rather than appear to deny.
**`scope:` is removed rather than implemented**, because giving it meaning would leave two ways to
express one thing. And the general fix is that **an unknown key is refused**: the host's
declaration parser already works this way, and manifests are the layer where that discipline is
missing. `scope:` survived because nothing rejected it, and it spread by copying to five
manifests.
## Order, and what it costs
**The link runs on the underlay and never on the overlay.** The overlay is configured by the mesh,
so a link requiring it could never be established on a new node.
**The first declaration is the overlay and nothing else** — because a node's address and peers are
*assigned* so it cannot come earlier, and because it is the way back in. A node reachable over the
overlay can be fixed by hand if a later declaration breaks the machine; **a large first
declaration risks a node that is broken and unreachable at once.**
**Reachable is not the same as having a control surface.** Every node reaches every other over the
overlay — SSH, services, ordinary traffic — and every node consumes from the broker. What is
forbidden is a listening thing that accepts instructions and changes the machine.
## Consequences
- **The last two direct database connections leave the nodes**, and with them the database
credential every node carries.
- **The `/etc/hosts` floor goes**, along with the bootstrap circularity it patched.
- **Two certificate authorities stay separate on purpose**: a public one for public names, the
mesh's own for internal ones. A single-CA lab would hide any bug living in the split.
- **What happens when the hub is down**: nothing takes over. Non-co-located paths stop; co-located
peers and every assigned workload keep running.
## Open — the link over the overlay, with a fallback
*Raised 2026-08-29 and deliberately left open, because the honest gain is smaller than it looks
and it is a decision rather than a derivation.*
The proposal: a node prefers the overlay for its link and drops to the underlay when the overlay
is not working — so ordinary operation is private and the underlay stays as the way back.
**Two things it would have to get right:**
- **The trigger cannot be "is the overlay up".** A WireGuard interface has no link state; once
configured it is up whether or not the far end exists. So there is no flag to read, and failing
over means *try, fail, time out, retry elsewhere*.
- **Running on the fallback has to be visible.** A node that quietly drops to the underlay is a
node whose overlay is broken with nothing to say so, and it will stay broken because everything
still works. That is this repository's recurring fault — a failure that reads as success — and a
fallback is the easiest place in the design to reintroduce it.
**What stops it being an obvious win:** if the underlay path must stay available for the fallback,
the broker stays exposed on it. So the exposure is unchanged and the traffic was already encrypted
— the gain is which network carries bytes, not what an attacker can reach.
**The version that would buy something is overlay-only**, with the broker firewalled to the overlay
in steady state, accepting that a node whose overlay breaks needs hands-on recovery. That is a real
trade: it exchanges the automatic way back in for a closed port.
Not decided either way here.
@@ -0,0 +1,114 @@
---
topic: the tiers
status: accepted
date: 2026-08-26
deciders: jochen
reconstructed: false
extends: 0009-modules-and-the-graph.md
---
# 8. A context owns its store, exclusively
## Context
[ADR 0009](0009-modules-and-the-graph.md) settles what a module
declares. This settles what a grant may be, and it is the half that **removes** things.
`how-we-build` §4 already says *contexts integrate through the record, never through a shared
schema*, and states the cost: several domains share one forty-five-table schema, which is why
work belonging to one context keeps having to be implemented in another.
That was written as a principle. Counted, it is thirteen foreign tables belonging to three
separate contexts, living in the mesh's own registry database.
## Considered options
1. **A schema per consumer inside a shared database.** Namespaced, revocable by dropping the
schema, with a cross-context join possible but deliberate. Rejected: it keeps the letter of
§4 and leaves the temptation in place, and a boundary that is merely inconvenient to cross
gets crossed.
2. **Read-only roles on another context's store.** Rejected for the same reason and one worse:
reading another context's tables couples you to its layout exactly as firmly as writing them,
and the coupling is invisible until the owner changes a column.
3. **Exclusive ownership.** Chosen.
## Decision
> **A context is granted only what it exclusively owns.**
No shared writes. No read-only role on another context's store. If you need what another context
holds, you ask it or you subscribe to it.
**The unit is the context, not the process.** Everything inside a context — its service, its
surface, its tools — reads its own store freely. A board showing the mesh's own nodes and
modules is the mesh showing its own data, not a boundary crossing. What is forbidden is a
*different* context reading it.
### Asking or subscribing is derived, not chosen
[ADR 0004](0004-a-node-and-how-it-joins.md) makes disconnection an ordinary situation. So:
- **Anything that must keep working while disconnected cannot ask** — there is nobody to ask. It
keeps a local copy, which means subscribing.
- **Anything where a stale answer is worse than none cannot subscribe.** A display may lag; a
decision about whether a grant is still valid may not.
Neither is a query against another store, whatever transport it travels over.
### What that is, concretely
*Written 2026-08-29, on building the first one — the rule above was clear and what to type was not.*
**One PostgreSQL database per context, named for the context.** A separate database rather than a
separate schema is the whole point: a cross-schema join is a qualified name away, and a
cross-database join needs a foreign data wrapper somebody has to install and explain.
**And one credential per context, held only by it.** There is no mesh-wide connection setting and
no way to ask for one, so reaching another context's store is not a matter of restraint — a process
has no address for it and nothing to present. That is also how this rule is *checked*: what a
context can reach is the list of variables the declaration running it grants, and it is read there
rather than audited in code.
**Contexts that do not exist yet do not get a database.** The bootstrap creates the ones there are.
**How a context added later gets its database is open**, and it is a real question: by then there is
a control plane, but a control plane holding a credential that can create databases is holding
rather more than the thing it exclusively owns.
## What this removes
The first clear list of what the design deletes rather than adds:
- **Grant kinds.** There is one: an exclusive resource. No schema grants, no read roles, no
rules about who may see what inside a shared thing.
- **The question of who owns which table**, and the guessing at revocation time. Removing a
consumer drops what it was granted, whole.
- **Cross-context migration ordering.** Two contexts migrating one database must be ordered
against each other. Exclusive ownership means a context's migrations are ordered only against
itself.
- **A class of permission modelling** a shared store would otherwise need.
## Consequences
- **Cross-context reporting is harder, and that is the point.** Anything wanting to see across
contexts consumes their events or calls their interfaces. That is §4's argument, and the cost
it names is the one already paid.
- **A single surface over several contexts still works** — that is what a surface is. It reads
interfaces, not stores. This holds while the contexts sit behind **one** interface; splitting
a context into its own deployable costs that, and the composition would have nowhere to live
that tier 3 permits. **A real constraint on how far the control plane may be split.**
- **Three contexts must move out of the registry database**, taking thirteen tables with them.
Their dependency on the registry then shrinks to almost nothing — one of them needs a single
table.
- **The node appliers were already handled.** [ADR 0005](0005-the-node-host.md)
stopped the host querying the mesh database for tier reasons unrelated to this, and it removes
most of the remaining direct readers as a side effect.
- **What a consumer does about events missed while disconnected is not decided** — replay from a
point, ask once and resume, or rebuild. The question every projection has.
## References
- [`how-we-build.md`](../00-META/how-we-build.md) §4 — the rule this makes enforceable.
- [Research 011](../01-RESEARCH/011-the-module-graph/worked-provider.md) — the count, the worked
provider, and the dashboard case.
- [ADR 0004](0004-a-node-and-how-it-joins.md) — why asking or subscribing is derived.
@@ -1,71 +0,0 @@
---
status: accepted
date: 2026-06-05
deciders: jochen
reconstructed: true
---
# 8. A step that fails must fail the job
> Reconstructed after the fact from the evidence cited below.
## Context
The mesh's expensive faults are not crashes. They are the operations that reported success and
did nothing: an artifact that partially downloaded and was extracted anyway, a package that
404ed from every mirror while the job went green, a hook that never ran because it was named
for a feature the module does not declare, a deploy that reported the transport succeeded
rather than that the effect happened.
Each of these was found long after it happened, by someone investigating an unrelated symptom.
The cost is not the failure; it is the interval between the failure and anyone learning of it,
during which decisions are made on the assumption that the thing worked.
## Considered options
1. **Continue on error and report at the end.** Rejected — it is largely what existed. A
summary nobody reads is not a report, and later steps run against the state the failed step
should have produced.
2. **Continue on error, and let health checks catch the divergence.** Rejected. It converts a
precise, located failure into a vague one discovered elsewhere, and requires a health check
for every possible partial state.
3. **Fail the step, fail the job, say which step.** Chosen.
## Decision
A step that fails stops the sequence it is part of, and the failure is surfaced where the work
was requested — not only in a log.
Concretely, and these are the forms it takes:
- A scripted sequence gates each step on the previous one. A directory change that fails must
stop the commands that assumed it.
- An artifact that does not fully download is not extracted.
- A stage reports the **effect** it achieved, not that it dispatched a message. "Started" must
mean the thing is running, not that a command returned.
- A template that cannot resolve a variable is not written half-rendered.
**Prefer failing to lying.** A green result that is not true costs more than a red one.
## Consequences
- Failures are noisier and land earlier, on the person who caused them.
- Some jobs that used to complete now stop. In every case examined so far, that job was
producing a partial result that something downstream trusted.
- This is a rule the mesh has adopted repeatedly rather than once, because each instance is
written in a different place — a shell hook, a download path, a deploy stage. It is not
enforced by a mechanism, and cannot currently be checked in general. New instances are still
being found; the package-install case remains open as
[`04-ISSUES/001`](../04-ISSUES/001-failed-package-install-reports-success/00-report.md).
## References
- `fix(installer): fail loudly when feature artifact download fails` (#244), 2026-06-05.
- `A flavor template with an unresolved variable is written to disk instead of failing`
(#710), 2026-08-08.
- Knowledge base: `troubleshooting/deploy-reports-transport-not-effect`,
`troubleshooting/service-started-is-not-ready`,
`troubleshooting/green-pipeline-means-transport-not-effect`,
`troubleshooting/silent-failures-and-stale-state`.
- The core value it became: [`00-META/mission.md`](../00-META/mission.md), "Failure must
be loud."
+441
View File
@@ -0,0 +1,441 @@
---
topic: what runs on it
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 9. Modules and the graph
*Consolidated 2026-08-28 from six records.*
## Everything is a module
One kind of thing, one manifest describing all of them. A database, a web application, a window
manager and a firewall rule set are all modules — not because they are alike, but because
**anything else means a second kind of thing with its own rules, and then a third.**
**A module is the unit of delivery**: assignable to a node, versionable, replaceable on its own.
## There are no domain modules
An earlier decision grouped modules by domain — four things constituting *how a node is
reachable* becoming one `networking` module. **That was wrong, and the correction is worth
keeping** because the observation behind it was right.
The measurement holds: reachability is the **only** place in the catalogue where modules
genuinely change together under one intent. What did not hold is the conclusion. Tight coupling
means they share an **authority** — one place that decides for all of them — and not that they
should be one artifact. `wireguard` and the proxy are deployed to different sets of nodes, so a
module containing both would be assigned where half of it is unwanted.
> **Coherence is a context. Delivery is a module.**
**Folders assert relationships; edges record them.** What grouping was for — finding things,
seeing what belongs together — is a tag and a query, neither of which anybody has to keep true by
hand.
### What a domain module turns out to be, and why it is not the one refused above
*Written 2026-08-29, from building it. The heading above reads as a contradiction of what now
exists and is not one — but only if the difference is stated, so it is stated here.*
**What was refused contains things. What exists contains nothing.**
| | `networking` as refused | `networking` as built |
|---|---|---|
| what is in it | WireGuard, a proxy, a firewall — artifacts | nothing at all |
| what it says | *these ship together* | *I want a private network and names* |
| what is assigned | one module, half of it unwanted | whatever answers each requirement, each on its own |
The objection above is untouched by this and still correct: a module holding WireGuard and a
proxy is assigned where half of it is unwanted. **A module holding nothing cannot be, because
there is no half.** It is requirements and a name, and every artifact it leads to is still an
ordinary module assigned on its own terms.
**Why it is worth having.** Most people want the network working and do not want to choose a VPN.
`assign networking` finds one answer to each requirement and takes it without asking, because
with one candidate there was never a question — the rule below about refusing does the work.
Somebody who does care assigns the VPN they want, and *that is the whole of choosing*: there is no
flavor field, no variant syntax, and no second verb. **Picking an implementation is assigning a
module.**
**What it costs, stated because it is real.** Adding a second implementation to the catalogue
turns a settled question into an open one for **everyone using the bundle**, not only for whoever
wanted the alternative. Every node assigned `networking` refuses until somebody says which. That
is [the refusing rule](#a-requirement-with-several-answers-is-refused-never-guessed) applied
consistently, and the alternative is a default — which is the flavor field returning under a
better name. The cost is one assignment per node, and the message names the candidates.
**A consequence that had to be found by running it.** A bundle can drag an implementation in
through a requirement nobody looked at. Choosing a different VPN still installed WireGuard,
because the names module needed addresses only WireGuard hands out, and nobody was told. Two VPNs
on one machine is not always wrong — a machine may run one for another purpose — but being **the**
network the mesh runs over is singular, so that is a claim, and the collision is refused by name.
**The general rule: what a bundle pulls in is only as safe as the claims on what it pulls in
from.**
## Three edges
| edge | means | declared? | satisfied |
|---|---|---|---|
| **presence** | that thing must exist and be reachable here | yes | at provisioning |
| **instantiation** | that thing makes something for me and hands back credentials — a database, a bucket, a route | yes | at provisioning, and again whenever it must be |
| **build** | I was compiled against that artifact | **no — read from imports** | **at build, once** |
**Instantiation implies presence; presence does not imply instantiation.**
**A route is an instantiation edge**, and it is worth noticing because the direction is the mirror
of a database: the consumer supplies a target and receives a *name*, rather than supplying nothing
and receiving credentials. Same edge.
**Provider stops being a category.** Any hosted thing can be a factory — an identity provider
grants clients, a mail server grants mailboxes. It is a facet, not a kind.
**A module may also declare what it claims**, because some things cannot coexist and that is a
fact about the module rather than about a particular node. What that means precisely is below.
### Where the answer to a requirement is allowed to live
*Written 2026-08-29, from building it. The table above distinguishes **presence** from
**instantiation** and this is the half of that distinction nobody had noticed was missing: not
what the edge hands over, but **where the thing on the other end is.***
Two different things were both being written as a requirement:
| | *a shell*, *a display server*, *a private network* | *a database*, *an object store*, *an identity provider* |
|---|---|---|
| where the answer lives | **this machine** | **somewhere in the mesh** |
| how it is answered | install another module here | find the node already running it |
| what is missing if absent | a module to assign here | **a decision about where**, which is nobody's to make silently |
Answering the second like the first installs a database on every machine that uses one, which is
what it did.
**So a provided name carries a scope**, the same idea a claim already has, and written short in
the ordinary case so the few that are not node-scoped stand out rather than drowning. Scope is a
property of **the name, not of each provider**: two modules disagreeing about whether a database
is local would make one requirement mean two things depending on which happened to answer it, so
that is refused.
**A requirement answered from the mesh is never satisfied by installing it here.** Nothing, and
the mesh refuses and says which module to assign somewhere. Two, and it refuses and says how to
choose — the same rule as everywhere else, for the same reason: picking is guessing, and the wrong
guess puts somebody's data on a machine they did not choose.
**Choosing is recorded per node**, because that is the granularity the choice actually has — two
machines may reasonably use two different databases and a mesh-wide answer could not say so. A
choice pointing at a machine that does not provide the thing is refused rather than quietly
replaced by one that does, and a single available provider does not override a choice either.
**Both are the same rule: the mesh does not overrule a person, and it does not move data without
being told to.**
**What this is a prerequisite for.** Knowing *which node* answers is the first half of handing a
credential back — you cannot be given a database's password before it is settled whose database it
is. So a node's resolution now records what it takes from elsewhere, which is both the only part
of its set that stops working when a *different* machine goes away, and the place a credential
will hang.
### An edge has two directions, and only one of them is built
*Written 2026-08-29, from building it. The row above already says a consumer **supplies a target
and receives a name**; what it did not say is that those are two separate mechanisms, and that
having one without the other is what forced two modules outside the system entirely.*
| direction | the consumer says | who needs it |
|---|---|---|
| **contribution** | *publish me at this name, on this port* | the proxy, the DNS server, a firewall |
| **binding** | *and give me back a credential to it* | the database, the object store, the identity provider |
**Contribution is built.** A module declares what it contributes to a requirement; the control
plane collects every contribution on a node and writes them to a path the provider named, as a
file, in the mesh's own shape. **Contributing to something is requiring it** — asking to be
published means a publisher must exist, and a module that had to say both would eventually say
one, with the failure appearing as a machine where nothing serves the route.
**The control plane does not know what a reverse proxy is**, and does not write one's
configuration. It delivers the facts; the module turns them into whatever it runs. That boundary
is what makes swapping the proxy cost nothing in any module that publishes through it, and it is
[the same separation](0001-mesh-brokers-nodes-host-agents-think.md) that keeps third-party
software *on* the mesh rather than *of* it. It also costs the host nothing: a received file is a
file, which was checked by putting the control plane's output through the host's own parser rather
than by asserting it.
**Binding is built except for the secret**, and that turned out to be the useful way to cut it.
A provider says what a consumer needs in order to use it — a port, a driver, a realm — and a
consumer says where it wants to be told. The mesh adds the half only it has: **which machine, and
what that machine is called on the private network.** So an application on one node is handed the
address of its database on another, as a file, and reaches it by a name the mesh also created.
**The file states that it carries no credential, and why.** A missing field looks like a bug; a
stated absence looks like a boundary, and somebody wiring this up should not spend an afternoon
looking for a password that was never going to be there.
### And the secret, which is delivered without ever being held
*Written 2026-08-30, after looking at how the existing mesh does it. The design here is a reaction
to a measurement, not a preference.*
**The obvious arrangement is a credentials column, encrypted at rest.** It exists, and its own
tooling records what it bought:
| | |
|---|---|
| the tool for finding a secret matches **by value**, not by name | because one password is in the provisions table, the environment table, each node's environment file in plain text, and **inside every connection string composed from it** — copies its documentation calls *"often the only copies actually in use"* |
| a query against the encrypted column **returns zero rows and proves nothing** | so auditing moved to the decrypted copies on the machines |
**Two faults, and encryption at rest addresses neither.** The control plane can read what it
stores, so a copy of its database is a copy of every credential in the mesh. And one secret has
many homes with nothing tracking them — **composition is what mints the untracked ones**, because
building a connection string centrally creates a new secret-bearing value no rotation path knows
about.
**So the value is sealed to the node that will use it before it is stored.** With a key that node
generated and whose private half the mesh has never seen — a third key beside the identity and the
overlay, for the same reason those are two rather than one. What is stored is unusable by whoever
holds it, the mesh included, and the broker relays a blob it cannot read. This is what makes
[ADR 0004](0004-a-node-and-how-it-joins.md)'s *compromise of a node is compromise of that node*
true of secrets rather than true of identity and quietly false of everything that matters.
**And nothing is composed centrally.** A connection string is assembled on the machine that needs
one, if at all. The mesh delivers parts.
**What it costs, stated because it is real:** the mesh cannot audit by value. That is the right
trade rather than an oversight — a query over an encrypted column could not either, so the audit
was never real. What *is* answerable is which node holds what, which is the question rotation
actually asks.
**A consequence that shapes the mechanism.** The mesh discarded the plaintext, so it cannot
compose a file containing it. The credential is therefore **its own file**, holding the value and
nothing else, beside the readable one. That is better than the alternative it was forced into:
the readable half stays readable in the declaration, and the secret half changes only when the
secret does, so a service reloading on it reloads for a real reason.
**Rotation is generating a new one**, because reading the old one back is not possible. Both ends
are re-sealed and reach their machines in the same push — which removes the window where half the
mesh holds a dead credential, the failure
[recorded in ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) as consumers on three nodes
holding one for two days.
### The provisioner, which is where the mesh stops
*Written 2026-08-30, from building one and running it against a real database.*
**A password nothing was told to create authenticates nowhere.** The mesh generates one, seals it
to both ends and cannot read it — so it cannot tell the software to start accepting it either.
Something on the providing machine reads what arrived and makes it true. That is a provisioner.
**It belongs to the module, not to the mesh**, and the boundary is the same one that keeps
third-party software running *on* the mesh rather than being *of* it
([ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)). The control plane decides and never
touches a machine. **What the mesh owns is the contract**, which is two files the host writes from
an ordinary declaration:
| | |
|---|---|
| the manifest | every consumer, what it asked for, and **where** its credential is |
| one file per consumer | that credential, alone in it |
Two files because the mesh discarded the value and cannot compose a document containing it. As
before, the constraint produces the better shape: the readable half stays readable and auditable
in the declaration, and the secret half changes only when the secret does.
**It reconciles; it is never told what changed.** It runs after every declaration and must reach
the same state from wherever it starts. Three consequences, and each of them is a fault that has
been shipped somewhere:
- **the password is set every time, not only on creation** — otherwise the role already exists,
nothing happens, and a rotation reports success while changing nothing
- **what it made and nobody asks for any more is removed** — otherwise a consumer that left keeps
a working login for ever and nothing ever says so. This is the same rule the host follows about
[removing what it declared and no longer declares](../04-ISSUES/010-the-first-declaration-destroys-the-substrate/00-report.md)
- **what it did not make is left alone** — otherwise it cannot be run on a system that predates
it, which is every system anybody would want to adopt
**A missing credential is refused rather than worked around.** A role created without one is a
login nothing can use, and nothing would report it until something tried to connect.
**This is where the mesh stops**, and saying so is the point of the section. It decides, delivers
and can prove what it delivered; the last inch belongs to whoever knows what `create role` means.
**One check that only became possible now.** Two machines wired together across no private network
is a mesh that reports itself configured and does not work, and the failure surfaces as a
connection timing out — the slowest place to find anything. It is refused, and it is only
*checkable* because the network became [something a machine is
given](#what-a-domain-module-turns-out-to-be-and-why-it-is-not-the-one-refused-above) rather than
something it has by virtue of holding an address.
**What the absence cost, measured.** Exactly two modules opened a direct connection to the control
plane's database — the proxy and the VPN — and they are the reason every node permanently holds a
credential to it. Both were doing by hand what this edge is for. The VPN's half is closed by being
[a module whose files are computed](../03-DESIGN/01-to-be/08-connectivity.md); the proxy's is
closed by contribution. **Neither needed a new kind of thing, and both had been outside the model
for as long as there was one.**
### Why the build edge is a different kind
It is fixed inside an artifact rather than negotiated when something runs, and **its only remedy
is a rebuild** — nothing can re-provision it.
It is also **derived rather than declared**, and the asymmetry is deliberate: a runtime edge is an
*intention* somebody has about how the mesh should be wired, and only a person can state it. A
build edge is a *fact about code that already exists*, and a declared list of dependencies drifts
from the imports it describes.
**An artifact is out of date when its source moved, or when anything it was built against moved.**
So what is recorded is a commit *and the identity of every artifact it was built against*, which
is what makes the rebuild set computable and *is this current?* answerable without building.
**The graph measures design quality, not just build order.** A module with many inbound build
edges is one whose every change is expensive — and that is readable before anything is built. The
current shared library is exactly that, and nobody could see it because nothing drew the edges.
## Provisioning is declared, never configured by hand
A module declares what it **provides** and what it **requires**. The mesh satisfies it: a
provisioner belonging to the provider creates the resource and its credential, records the grant,
and the values are derived onto the consumer. **Neither the credential nor the topology is ever
written by hand.** A requirement may name a provider on another node, so cross-node wiring is the
same declaration.
## What a module claims, and why it is not a list of rivals
*Written 2026-08-29, replacing pairwise exclusion.*
**Exclusivity is not a property of a module. It is a property of a singular resource the module
takes over.** Two shells do not compete for anything and any number may be installed. Two display
servers both want the seat, and only one may have it.
> **A module declares what it *claims*. Two modules claiming the same thing cannot both be
> assigned within that claim's scope.**
**Not "xorg conflicts with wayland".** Pairwise exclusion has a property that only shows up later:
adding a third display server means **editing xorg and wayland to know about it**. Every new
module requires changing modules nobody who wrote it owns, and the edits grow as the square of
the count. With a claim, the third one says `claims: the seat` and nothing else changes anywhere.
**The new module is the only thing that has to know anything** — which is the difference between
a catalogue that grows and one that calcifies.
The pattern is common enough to be worth listing, because seeing it is most of understanding it:
| these coexist | these claim one thing |
|---|---|
| shells — bash, zsh, fish | display servers — xorg, wayland (*the seat*) |
| editors — vim, emacs, helix | init — systemd, openrc (*pid 1*) |
| language runtimes | container runtime — docker, podman |
| terminal emulators | reverse proxies — nginx, caddy, traefik (*ports 80/443*) |
| browsers | time — chrony, timesyncd, ntpd (*the clock*) |
| | resolvers — resolved, dnsmasq, unbound (*`/etc/resolv.conf`*) |
| | network management — NetworkManager, networkd, netctl |
| | mail — postfix, exim, msmtp (*port 25*) |
| | audio — pipewire, pulseaudio (*the device*) |
**A claim has a scope**, because not everything singular is singular per machine:
| scope | example |
|---|---|
| **node** | the seat, pid 1, port 443 |
| **site** | a DHCP server on a segment |
| **mesh** | the hub, the control plane |
The last is not new — the mesh already enforces exactly one hub with a unique index
([ADR 0007](0007-connectivity.md)). Scope is that idea, said once rather than hard-coded per case.
**Some conflicts need no claim at all.** Two modules declaring the same file, or binding the same
port, are visible from *what they declare* — the mesh already holds every resource of every
declaration. So a claim is only written for the abstract ones, where nothing in the declaration
reveals the clash. That keeps the manifest small, which is worth protecting.
## A requirement with several answers is refused, never guessed
A module requiring *a shell* may be satisfied by three. The mesh does not pick.
| candidates | what happens |
|---|---|
| exactly one | assigned, silently — there was no choice to make |
| none | refused, naming what is missing |
| several | **refused, naming them**, and a person chooses |
**This is what makes a solver unnecessary.** Counting candidates is a few lines and has no
surprising behaviour; a solver that picks has to be understood before its answer can be trusted,
and it is understood by whoever is debugging it at the time. Nothing here is lost by waiting —
a solver can be added later without changing a single manifest, and the reverse is not true.
**Requiring a module and requiring a capability are different fields**, because the remedies
differ and the message should say which:
- *i3 needs xorg, which is not assigned here* — assign it.
- *this machine has no seat* — wrong machine; nothing can be installed to fix it.
### A capability may carry a value, and that is not a new idea
A capability is a named fact about a machine, **detected and never assumed**. Its presence gates
an assignment; its detail can also carry a value — `seat: card1-DP-1`, `panel: oled`, an
architecture, an amount of memory. Nothing new is needed for that: a verdict has always had a
detail beside its yes or no.
So *can this run here* and *what should it be configured as* are answered by the same fact, read
two ways. A module that must not be assigned without an OLED panel and one that dims itself
differently on one are reading the same line.
**What keeps the set from sprawling is the cost of adding one.** A capability must be detected,
and the detector must say how it knows — so nobody can add one they cannot check, which is the
whole of [04-ISSUES/007](../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md):
an installed package was treated as a capability and a node was assigned work it could not do.
**And detectors ship inside the host**, which is one statically linked binary. Adding a capability
means shipping a new host to every node that needs it. That is a real cost and it argues for
keeping the vocabulary small and general — `seat`, not `has-nvidia-with-two-outputs`.
## "Flavor" is retired
It was carrying three unrelated meanings — variants of a thing, a subset of one module a node
installs, and whatever the current system does, which earned two knowledge-base entries about
going wrong. **A word with three meanings cannot be reasoned about**, and every attempt to design
around it produced a rule that was right for one meaning and wrong for the others.
What it was reaching for is two ordinary things:
- **Different modules that provide the same thing.** `zsh` and `fish` both provide *a shell*. They
are two modules, not one module with a switch: they share a name and nothing else — different
packages, different configuration, different everything.
- **One module with a setting.** A monitoring module that is an agent here and a server there is
one module, configured. Nothing varies but a value.
If something is neither, it is probably two modules.
**A third thing it was reaching for, added 2026-08-29:** *I want this working and I do not care
which one.* That is a module with requirements and no files —
[a domain module](#what-a-domain-module-turns-out-to-be-and-why-it-is-not-the-one-refused-above) —
and it is what makes "different modules that provide the same thing" bearable for somebody who
does not want to know there is a choice.
## The core library is the mesh's domain
One module everything may depend on. It holds **what is true of the mesh regardless of which
context you are in**: a module, a node, an assignment.
The test: *would this still mean the same thing in a context that had never heard of the one it
came from?* A node would. A pipeline stage would not — that is delivery's.
**Types ship with the module that owns them**, not here. A consumer needing `inventory`'s types
depends on `inventory` — one narrow, visible edge — rather than everything depending on a hub
where the relationship cannot be seen. **A library everything depends on is expensive to change
whether it holds types or code; the fan-in is what makes it expensive**, which is why *types, not
behaviour* was the wrong guard.
**It stays small on its own.** A domain model changes when what the mesh *is* changes, which is
rare. A drawer labelled *shared* changes whenever anybody writes something reusable, which is
constantly — and *who else might want this* always answers yes, which is how the current one grew.
## Consequences
- **Fewer things will be shared, and some code will be written twice.** That is the trade: the
current library exists because sharing felt free. Two similar functions in two modules is often
the better answer.
- **The check is a measurement rather than a prohibition.** Inbound build edges say when something
is becoming a hub, while it is happening rather than after.
- **Reading build edges needs a language-aware tool per language**, which is the real cost and the
reason declaring them looks tempting. It is still wrong.
+161
View File
@@ -0,0 +1,161 @@
---
topic: what runs on it
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 10. Delivery
*Consolidated 2026-08-28 from five records.*
## Delivery is a comparison, not a pipeline
**The control plane holds what source exists and what has been built from it, and builds the
difference.** A change becomes a build because source is **ahead of artifacts** — answerable at
any moment — rather than because a message arrived.
**An event makes it fast. Nothing makes it necessary.** A missed notification costs latency and
cannot cost correctness.
That is the same shape the host uses on a machine, one layer up:
| | reconciles | against |
|---|---|---|
| the control plane | artifacts | source |
| the host | machine state | declarations |
**This is not the current coordinator repaired.** That is a state machine over stages; the value
of it here is as a catalogue of the ways this fails, and it has been used for exactly that.
### The module system is the CI/CD
*Written 2026-08-29, because this was the intention throughout and was never stated in one line.*
**There is no pipeline product beside the mesh, and there is not going to be one.** A module
declares what it is ([ADR 0009](0009-modules-and-the-graph.md)); the control plane notices its
source is ahead of its artifacts and builds it; the graph says what else that invalidates; the
node that should run it is told. Build, test, publish and deploy are the same reconciliation seen
at four points, not four stages wired together.
**Which is why the module system is the core of the setup rather than one component of it.** Every
other layer is carried by it: the substrate is modules the bundle raises before there is a mesh,
the control plane is a module, and an application is a module with a different manifest. A thing
that cannot be expressed as a module cannot be delivered at all — that is a real constraint, and it
is the one keeping a second delivery mechanism from growing beside this one.
**What disappears is the pipeline as a state machine** — no stage list something can be omitted
from, which is how a verify stage was built and never scheduled, and no run to lose.
### Currency is the whole input closure
**An artifact is out of date when its source moved, or anything it was built against moved.** So a
shared library changing invalidates everything with a transitive build edge to it, in dependency
order, because a module cannot be built against a new library until it exists.
**The module graph is a prerequisite of this, not an enabler of it.** Without it there is no
rebuild set and no ordering, and this cannot be implemented.
## An artifact is build output, never a source tree
Compiled and bundled with its dependency graph inlined. **A deploy is extract-and-run and touches
no network.**
The consequence is the whole cost of the decision: **anything not in the build output does not
ship.** Every file kind had to be brought into that rule separately, and each was discovered by
something silently not happening after a deploy — migrations reading a source layout,
provisioning scripts reading a source layout, selection files never packaged at all.
## Three silos, and the third is not a stage
The cardinality observation holds and is what the split is for:
| silo | runs | ends with |
|---|---|---|
| **build** | once per module | a self-contained artifact |
| **publish** | once per module | that artifact addressable — an image by digest, a package in the mesh's repository |
| **deploy** | **once, not once per node** | the affected nodes' **declarations updated** |
**Deploy stops sending commands to nodes.** It changes what the control plane says each node
should be, which is one write. What happens on the machines is the host's ordinary reconcile.
**Why this fixes the failure class rather than patching it.** Every recorded fault shares one
shape: *the thing that reported success was not the thing that did the work.* A coordinator
dispatching a command can only report on dispatch. Under this the reporter **is** the applier —
which already refuses to record a resource until it read it back, and already fails the whole
apply on one failed step.
**The verify stage disappears as a stage**, which is the strongest evidence for the shape:
verification stops being a step that can be omitted from a list and becomes a property of applying
at all.
**There is no fan-out**, so the defect class that came from the build node having passed through
two silos while others had not cannot arise.
## A step that fails must fail the job
A step that fails and lets the job continue **reports success for work that did not happen**.
Absence of an error is not evidence of an effect.
This is the mesh's most consistent failure shape, and it is not incidental — it is what stage
reporting measured. Documented instances: a service reported started when the container command
merely returned; an image pull failure that did not fail the deploy; a package install that 404'd
from every mirror while the job went green; a node left on old code after a failed download with
a version marker that had already advanced.
## The verdict is tiered
An artifact may not be declared until something has judged it fit. **Two tiers, because one gate
would be both slow and unreliable:**
| | judged by | when |
|---|---|---|
| the module's own tests | the build | **always** — this is most of it |
| the lab | a raised scenario | when an assertion genuinely needs a mesh |
A lab scenario takes tens of seconds and can fail for reasons that have nothing to do with the
artifact, and a shared-library change produces a cascade of dozens. One expensive
non-deterministic gate fails in both directions: a flaky run marks a good artifact unfit, a lucky
one marks a bad artifact fit, and **neither failure looks like itself.**
**A run that failed environmentally is not a verdict.** A machine that would not boot says nothing
about the artifact, and recording it as *unfit* is the same untruth as recording a dispatch as a
deploy.
## What a result means
> **The declaration is updated, and here is which nodes have applied it.**
A pipeline does not wait for every node — one may be legitimately switched off for a week, and a
delivery mechanism that blocks on a sleeping laptop is one nobody will use.
```
delivered declaration updated for 5 nodes
applied 3 of 5
outstanding 2 — last seen 4 days ago, 20 minutes ago
```
**Outstanding is not failure**, and conflating them is how the old system produced a stall with no
error anywhere.
## What must exist first
1. **The module graph, with build edges.** No graph, no rebuild set and no ordering.
2. **A recorded input closure per artifact**, so currency is answerable without building.
3. **Something that notices a reconciler is not converging.** Below.
## Open, and the first is the real risk
- **A loop that will not converge is harder to debug than a job that failed.** A failed job stops
and names its step; a reconciler retries forever. Without something that notices *this has been
trying for an hour*, the failure is **silence** — the fault this removes, reintroduced in a new
place.
- **The run identity people use is lost.** *Did my change go out?* is answerable today by opening
a pipeline. Something must replace that or this is worse to live with, whatever its properties.
- **Does a fit artifact declare itself?** If it does, merging to main deploys to production —
which may be wanted and is far too large a property to acquire by omission.
- **Rebuild storms are mostly behaviourally empty.** Reproducible builds would stop a cascade at
the first module whose output did not move; without them one commit redeploys the fleet for no
change in behaviour.
- **Detection stays the fragile input for latency**, though no longer for correctness.
@@ -1,17 +1,18 @@
---
topic: building it
status: accepted
date: 2026-04-03
deciders: jochen
reconstructed: true
---
# 4. Managed files are generated onto nodes and never edited there
# 11. Managed files are generated onto nodes and never edited there
> Reconstructed after the fact from the evidence cited below.
## Context
[ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) put every binding in the mesh
[ADR 0006](0006-the-substrate-and-the-control-plane.md) put every binding in the mesh
database. But the things that consume those bindings — environment files, service
definitions, daemon configuration, firewall rules — are files on a node's disk, because that
is what the software reading them requires.
@@ -26,7 +27,7 @@ directions at once.
1. **Bidirectional sync** — a node's edits flow back to the database. Rejected, and removed.
Two writers and no arbiter: whichever synced last wins, and neither is authority.
2. **Files are authoritative; the database is a cache of them.** Rejected — it inverts
ADR 0003 and returns to state that cannot be reconciled across nodes.
ADR 0013 and returns to state that cannot be reconciled across nodes.
3. **Strictly one-directional: the database is written, files are generated.** Chosen.
## Decision
@@ -1,60 +0,0 @@
---
status: superseded
superseded-by: 02-DECISIONS/0018-the-mesh-creates-no-symlinks.md
date: 2026-07-10
deciders: jochen
reconstructed: true
---
# 11. The installer owns linking; nothing else creates a symlink
> Reconstructed after the fact from the evidence cited below. The incident that earned the rule
> predates the record, and its date is not established here.
## Context
A service's definition lives in the module catalogue; its runtime directory and persistent data
live outside it. The mesh connects the two by linking the definition into the runtime location
— deliberately, so that runtime state and source stay separate while the running service reads
a current definition.
A link is also the easiest thing in the world to create by hand while fixing something, and a
container engine resolves a bind mount through it. A hand-made link pointed a volume somewhere
it should not have, and **production data was lost**.
## Considered options
1. **Copy instead of linking.** Rejected. A copy goes stale silently, which trades data loss
for a service running a definition nobody can find.
2. **Allow links, document the hazard.** Rejected. The hazard is not knowable at the moment of
the mistake — the link looks right and the resolution happens inside the container engine.
3. **One component owns linking; everyone else is forbidden.** Chosen.
## Decision
The installer creates and repairs every link the mesh needs. It reconciles them: a missing
link is created, a stale one is repointed, and a real file found where a link belongs is
adopted into the node's override location and replaced.
**Nothing else creates a symlink** — not a hook, not a fix, not an agent, not a person
debugging. The prohibition is absolute because the judgement required to make a safe exception
is exactly the judgement that was not available at the moment it mattered.
## Consequences
- The class of failure is closed, at the cost of a rule that reads as arbitrary to anyone who
has not seen the incident. That is why it is recorded here rather than only asserted.
- Links become reconcilable state rather than incidental filesystem facts.
- The rule is stated for humans and agents and is enforced by convention, not mechanism. A
check does not exist.
- The rule as written governs the mechanism rather than removing it. A link made by the
installer resolves the same way as one made by hand, so the hazard is narrowed and not
closed. [ADR 0018](0018-the-mesh-creates-no-symlinks.md) proposes widening this to "nothing
links, the installer included"; until that is accepted, this record governs.
## References
- Recorded as a non-negotiable in the governed constitution page, §2: *"Symlinks to repos or
service directories have caused production data loss via Docker volume path resolution. The
installer handles all linking. Never create symlinks manually."*
- Knowledge base: `services` — the reconciliation behaviour, including adoption of real files.
@@ -1,16 +1,16 @@
---
topic: building it
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
extends: 0011-the-installer-owns-linking.md
---
# 18. The mesh creates no symlinks — a derived file is a copy
# 12. The mesh creates no symlinks — a derived file is a copy
## Context
[ADR 0011](0011-the-installer-owns-linking.md) responded to production data loss — a hand-made
[ADR 0012](0012-the-mesh-creates-no-symlinks.md) responded to production data loss — a hand-made
link, resolved through a container engine's volume handling, pointing a mount somewhere it
should not have — by centralising linking in the installer and forbidding it everywhere else.
@@ -25,7 +25,7 @@ Two things have changed since, and together they remove the argument that kept i
silently while the catalogue moves on, so a link was the cheap way to guarantee the running
node reads a current definition. That argument assumes the node's copy is unmanaged.
**It is not.** [ADR 0004](0004-managed-files-are-generated-never-edited.md) established that
**It is not.** [ADR 0011](0011-managed-files-are-generated-never-edited.md) established that
everything on a node's disk is derived from the mesh and regenerated when its inputs change,
and the installer already **reconciles** links rather than assuming them — repointing stale
ones, adopting real files it finds where a link belongs. Reconciling content is the same
@@ -34,12 +34,12 @@ operation as reconciling a pointer, plus a comparison.
So the mesh already has the machinery that makes a copy safe, and is using a link to solve a
problem that machinery solves better. Worse, a link is conceptually the wrong shape: it makes
the node's runtime state a *pointer into source*, which is the one thing
[ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) and ADR 0004 exist to prevent.
[ADR 0006](0006-the-substrate-and-the-control-plane.md) and ADR 0011 exist to prevent.
State is derived onto nodes; it does not reach back.
## Considered options
1. **Keep ADR 0011 as the final position** — centralised linking, forbidden elsewhere.
1. **Keep ADR 0019 as the final position** — centralised linking, forbidden elsewhere.
Rejected as the status quo. It governs the mechanism rather than removing it, and the
failure it was written for remains reachable by any code path the installer trusts.
2. **Keep links but harden them** — canonicalise before mounting, refuse a link that escapes
@@ -54,12 +54,12 @@ State is derived onto nodes; it does not reach back.
**The mesh creates no symlinks.** A file a node needs is placed on that node as a real file,
derived from the mesh and reconciled by the installer like every other managed file
([ADR 0004](0004-managed-files-are-generated-never-edited.md)).
([ADR 0011](0011-managed-files-are-generated-never-edited.md)).
The prohibition in ADR 0011 stands and widens: it ceases to be "only the installer may link"
The prohibition in ADR 0019 stands and widens: it ceases to be "only the installer may link"
and becomes "nothing links, the installer included".
When this is accepted, ADR 0011 becomes superseded rather than edited — its reasoning is why
When this is accepted, ADR 0019 becomes superseded rather than edited — its reasoning is why
the rule exists at all, and the incident behind it is the reason anyone believes either record.
## Consequences
@@ -73,7 +73,7 @@ the rule exists at all, and the incident behind it is the reason anyone believes
cost, and it is the whole cost: today a link cannot be stale, and a copy can. The answer has
to be detection — the installer comparing what is on disk against what the mesh says should
be — and it must be loud, because a silently stale definition is exactly the failure shape
this mesh keeps producing ([ADR 0008](0008-a-failed-step-fails-the-job.md)).
this mesh keeps producing ([ADR 0010](0010-delivery.md)).
- Reconciliation gets more expensive: comparing content rather than checking a pointer's
target, on every module, on every node.
- Disk usage rises, trivially, and is not a consideration.
@@ -90,12 +90,12 @@ the rule exists at all, and the incident behind it is the reason anyone believes
- **Migration order.** Converting a node's links is a change to how its services resolve their
own definitions, which is not a change to make everywhere at once.
Until those are answered this record stays `proposed`, and ADR 0011 remains the governing rule.
Until those are answered this record stays `proposed`, and ADR 0019 remains the governing rule.
## References
- [ADR 0011](0011-the-installer-owns-linking.md) — the incident, and the rule this widens.
- [ADR 0004](0004-managed-files-are-generated-never-edited.md) — the machinery that makes a
- [ADR 0012](0012-the-mesh-creates-no-symlinks.md) — the incident, and the rule this widens.
- [ADR 0011](0011-managed-files-are-generated-never-edited.md) — the machinery that makes a
copy safe.
- [`03-DESIGN/00-as-is/05-runtime-and-installation.md`](../03-DESIGN/00-as-is/05-runtime-and-installation.md)
— what the installer does today, including reconciliation and adoption.
@@ -1,63 +0,0 @@
---
status: accepted
date: 2026-08-04
deciders: jochen
reconstructed: true
---
# 13. An artifact is build output, never a source tree
> Reconstructed after the fact from the evidence cited below.
## Context
A module is built once and deployed to every node assigned to it. What travels between those
two events is the artifact.
For a long time the artifact was a filtered copy of the module's source directory. Deploying it
therefore meant resolving and installing its dependencies **on the target node** — which
requires the target to reach a package registry, at deploy time, for every node, every deploy.
A node with no route to the registry could not deploy code that had already been built
successfully.
## Considered options
1. **Ship source, install dependencies on the target.** Rejected — it is what existed. Deploy
becomes a network operation with a failure mode per node, and the code that runs is
assembled independently on each one.
2. **Ship source plus its resolved dependency tree.** Rejected: large, slow, and it ships the
dependency resolution's platform assumptions along with it.
3. **Ship a self-contained build output; a failed bundle fails the build.** Chosen.
## Decision
The artifact is the module's **build output directory** — compiled and bundled, with its
dependency graph inlined. Deploy is extract-and-run and touches no network.
A build that cannot produce a self-contained output **fails**. It does not fall back to
shipping a dependency tree, because a fallback that works is a fallback that is never fixed —
an application of [ADR 0008](0008-a-failed-step-fails-the-job.md).
## Consequences
- A node can deploy without reaching a registry. What was built is what runs, identically, on
every node.
- Deploys are faster and their failure modes are local.
- **Everything not in the build output does not ship.** This is the decision's whole cost, and
it was paid several times before it was understood: migrations that read the source layout,
provisioning scripts that read the source layout, selection files never packaged at all. Each
worked in development, where the source is present, and silently did nothing after deploy.
- Any file a module needs at runtime must be deliberately placed into the build output. The
rule "the artifact is `dist/`" has to be applied to every file kind, not just compiled code,
and that generalisation was the expensive part.
- Bundling has its own failure modes that a compiler will not catch — a bundler can exit
successfully and produce output that cannot load.
## References
- `build: bundle artifacts so a deploy is extract-and-run` (#673), 2026-08-04.
- The consequences, in order: `Provision migrations and seeds read the source layout, not the
artifact` (#699), `Local migrations read the source layout too` (#700), both 2026-08-07.
- Knowledge base: `pipeline/artifacts-are-build-output`, `pipeline/bundling`,
`troubleshooting/shell-migrations-never-packaged`, `troubleshooting/flavors-never-packaged`,
`troubleshooting/esbuild-silent-tla-breakage`.
@@ -1,11 +1,12 @@
---
topic: building it
status: accepted
date: 2026-05-14
deciders: jochen
reconstructed: true
---
# 6. Schema and state changes are numbered migrations, in the same language as the code
# 13. Schema and state changes are numbered migrations, in the same language as the code
> Reconstructed after the fact from the evidence cited below.
@@ -1,78 +0,0 @@
---
status: accepted
date: 2026-08-04
deciders: jochen
reconstructed: true
---
# 14. Build, publish and deploy are three silos with different cardinality
> Reconstructed after the fact from the evidence cited below.
## Context
Delivery had been treated as one pipeline that a module passes through. It is not: its stages
run a different number of times.
- Compiling happens **once per module feature**, on the build node.
- Packaging and uploading happens **once per module feature**, on the build node.
- Installing, configuring, starting and verifying happens **once per module feature per node**.
Conflating them is what made earlier versions slow and hard to reason about. Work that should
happen once was being repeated per node, and the fan-out point was implicit rather than a
boundary anything could observe.
The split had been declared before it was real. Packaging still happened inside the build,
which meant the boundary existed in the documentation and not in the code.
## Considered options
1. **One pipeline, stages that know their own cardinality.** Rejected — it is what existed.
Cardinality is then a property of each stage's implementation, and nothing can reason about
the pipeline as a whole.
2. **Two silos: build-and-publish, then deploy.** Rejected. It leaves packaging inside build,
so build must know every module, every feature, and how each composes its artifact —
exactly the coupling the split exists to remove. A failed upload then retries by re-sending
a stale package instead of re-packaging.
3. **Three silos, with an explicit handover between each.** Chosen.
## Decision
Delivery is three silos, and the boundaries are real:
| Silo | Runs | Where |
|---|---|---|
| **build** | once per module feature | the build node |
| **publish** | once per module feature | the build node |
| **deploy** | once per module feature **per node** | every assigned node |
Commands and events are addressed **per feature**, not per module.
Build compiles and hands over a **staged tree** — not a package. Publish applies the module's
packaging rules, packages that tree, and uploads it. Publishing to a package registry *is*
publishing, so a module whose artifact is a package publishes in the publish silo, not the
build one.
Modules are resolved into dependency **levels**, and a level completes before the next begins,
so a module always builds against its dependencies' freshly published versions.
## Consequences
- Work that should happen once happens once. The fan-out point is explicit and observable.
- A failed upload retries by re-packaging, because packaging belongs to the stage that
uploads.
- The handover is a staged tree in a known location rather than the build's working directory,
which is reference-counted and cannot be assumed to still exist when a later stage runs.
- The build node is now the only node that has already passed through two silos when the
fan-out happens. Anything tracking a node's stage must account for **both** pre-fan-out
stages; code that knew only about the first parked the build node forever while every other
node deployed cleanly.
- A recovery mechanism that knows a subset of the stages it guards is worse than none — it
reports success over a stall it cannot see.
## References
- `publish owns packaging — the silos were not actually split` (#677), 2026-08-04.
- Knowledge base: `pipeline/three-silos` — including the note that the older architecture
documents claimed otherwise and were stale until 2026-08-06.
- The build-node stage-tracking failure was observed on pipeline #5557.
@@ -1,11 +1,12 @@
---
topic: building it
status: accepted
date: 2026-06-04
deciders: jochen
reconstructed: true
---
# 7. No workspace — each module is a standalone package consuming published dependencies
# 14. No workspace — each module is a standalone package consuming published dependencies
> Reconstructed after the fact from the evidence cited below.
@@ -47,7 +48,7 @@ its dependencies' freshly published versions.
- Development and the pipeline resolve imports identically. The divergence is gone by
construction rather than by discipline.
- A module in its own repository is not a special case. It builds exactly as a module in the
monorepo does — which is what makes [ADR 0010](0010-applications-live-in-their-own-repository.md)
monorepo does — which is what makes [ADR 0015](0015-applications-live-in-their-own-repository.md)
cheap.
- A cross-package change costs a publish-and-consume round trip. This is the real price, paid
on every shared-library change.
@@ -1,11 +1,12 @@
---
topic: building it
status: accepted
date: 2026-07-10
deciders: jochen
reconstructed: true
---
# 10. Applications live in their own repository; the monorepo is for the mesh
# 15. Applications live in their own repository; the monorepo is for the mesh
> Reconstructed after the fact from the evidence cited below.
@@ -46,7 +47,7 @@ reject it.
- An application's cadence is its own. It is not reviewed as mesh code and does not queue
behind mesh work.
- The separation is safe **only because** the pipeline and provisioning are identical either
side of it — which [ADR 0007](0007-no-npm-workspace.md) is what makes true. Without
side of it — which [ADR 0014](0014-no-npm-workspace.md) is what makes true. Without
standalone packages this decision would fork the build.
- The monorepo stops being an inventory of the installation, which is a precondition for
publishing anything about it.
@@ -59,6 +60,6 @@ reject it.
- The rule is stated in the governed constitution page authored 2026-07-10, §3, as a
convention violation reviewers must reject.
- [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) decision 4 extends this from *new*
- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) decision 4 extends this from *new*
applications to the modules already in the monorepo.
- Knowledge base: `troubleshooting/unregistered-module-source`.
@@ -1,138 +0,0 @@
---
status: accepted
date: 2026-08-22
deciders: jochen
reconstructed: false
---
# 16. A lab node is a virtual machine running the real install
## Context
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) makes a local mesh a prerequisite
rather than a convenience: *"everything that manifests between nodes is discoverable only in
production, which is where every fault of 2026-08-22 was found."*
Two things were measured while establishing what exists
([`01-RESEARCH/002-local-mesh`](../01-RESEARCH/002-local-mesh/analysis.md)):
- **There is no local mesh.** The dev tooling starts providers through the host's own init
system and reads credentials from host paths (`modules/hal/developer/tools/dev-env.ts:146-178`).
It borrows the machine because there is nowhere else to put a mesh.
- **The one containerised node in the repository has been unable to build since 2026-06-04**,
when the npm workspace it depends on was removed. Nothing runs it, so nothing reported it.
So the question is not how to improve a local mesh. It is what a node *is* when it is not a
physical machine. Every subsequent question — how faithful is faithful enough, what may be
mocked, which failures remain reachable — follows from that one answer.
The hardware available is not a constraint: 125 GB of memory with 71 free, 24 threads, and
hardware virtualisation present.
## Considered Options
1. **An application container.** Rejected. **A node's job is to run containers**, so modelling
a node as one inverts the thing being modelled: module service stacks then require nested
containers through a privileged daemon, or a shared socket that makes isolation between
nodes cosmetic. Init is not PID 1, so units and timers need workarounds. Cheapest to start
and the least like a node.
2. **A system container.** Rejected, after first being recommended. It is genuinely good —
real init, properly nested containers, roughly a second to boot, cheap snapshots — and it
is the only option that makes a twenty-node run affordable. It was rejected because **the
scale requirement that justified it was invented rather than required**: the stated goal is
to run the real mesh, which is four nodes, on one computer. And a system container still
forces the question a virtual machine dissolves — *how faithful must a node be?* — which
then has to be answered again for every capability under test.
3. **`systemd-nspawn`.** Rejected. Already present, so nothing to install, but too primitive:
no storage pools, no snapshot management, no network management, no virtual machines.
Snapshots are what make the loop fast, so the saving is not worth what it costs.
4. **A virtual machine.** **Adopted.** A bare Arch Linux machine that the real install script
turns into a node.
## Decision
**A node in the mesh development lab is a virtual machine.** It boots a stock Linux image,
runs the real install, and becomes a node. It is not a model of a node, so no question arises
about how good the model is.
The environment is called **the lab**.
Three things follow directly and are decided here:
### The lab is driven by `incus`
Chosen for what it manages, not for what it is: virtual machines, their snapshots, and the
bridges between them, through one interface. It also manages system containers, so if a run
ever genuinely needs twenty nodes, that is a change of instance type rather than a rewrite.
Declared in `modules/hal/developer/module.yml`, so it installs the way every other package
does.
### The simulated public segment uses TEST-NET-3
`203.0.113.0/24`, reserved by RFC 5737, never routable.
This is not cosmetic. WireGuard decides per pair whether to write an `Endpoint` by testing the
peer's underlay address against an RFC1918 regex
(`modules/wireguard/hooks/index.ts:225-240`). A simulated public segment addressed from
private space makes the hub test as unreachable, so no spoke writes an endpoint for it,
nothing can initiate, **and the mesh silently never forms** — appearing as a WireGuard fault
rather than an addressing mistake.
The production LAN subnet and the entire overlay address plan are reproduced unchanged.
### The lab issues its own certificates
Public names are certified by an ACME server inside the lab; `.internal` names keep the mesh
CA. **The lab keeps production's two-authority split rather than collapsing it**, because a
single-authority lab would hide any fault living in that split.
This also makes the lab's port forward load-bearing: an HTTP-01 challenge must reach a
published-but-NATed node on port 80, so a broken forward becomes a reproducible certificate
failure rather than a mystery.
## Consequences
**The fidelity question disappears, and with it a class of argument.** There is no "how real
is this node" to litigate per capability, because the node is real. What remains not-real is a
short, enumerable list: the model provider, the public internet, and the public certificate
authority.
**The install becomes the thing under test.** A container-shaped lab would have had to skip
the bootstrap entirely. Here it runs, so it is exercised on every fresh lab.
**Reproducing the network is mostly a data problem.** The bootstrap performs no network
configuration at all; WireGuard, DNS, routing and internal TLS are generated by module hooks
from mesh-DB rows. The lab therefore exercises the same code production runs rather than a
reimplementation ([`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md)).
**Scale runs get expensive, and this is the real cost.** Four virtual machines are
comfortable; twenty are not, on a workstation. Faults that only appear at scale — a fan-out
reaching most consumers rather than all, a cascade that stalls with many modules — stay hard
to reproduce. The mitigation is that the same tooling runs system containers, so a scale run
remains possible at lower fidelity if one is ever genuinely needed.
**Boot is slower, and it does not matter.** Ten to twenty seconds against roughly one. A run
includes a full delivery — build, publish, install, migrate — measured in minutes, so boot
time is noise.
**One change is required before the lab can issue certificates.** The reverse proxy sets no
`caServer`, so it defaults to the public authority's *production* endpoint
(`modules/traefik/docker-compose.yml:17-19`). It must become configurable, defaulting to
production so real nodes are unaffected. Worth noting on its own: aiming at production rather
than staging means every certificate experiment on a real node consumes issuance quota.
## References
- [`01-RESEARCH/002-local-mesh`](../01-RESEARCH/002-local-mesh/analysis.md) — what exists, and
the four host couplings that only obstruct a container-shaped node
- [`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md) — the topology
being reproduced and the endpoint constraint
- [`03-DESIGN/01-end-to-end-testing.md`](../03-DESIGN/01-to-be/01-end-to-end-testing.md) — what the lab
is for
- `modules/wireguard/hooks/index.ts:206-240` — the endpoint rule, and the incident comments
recording what it cost to get right
- RFC 5737 — reserved documentation address blocks
+84
View File
@@ -0,0 +1,84 @@
---
topic: building it
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 16. The lab
*Consolidated 2026-08-28 from five records. The lab is one design and was split across five
decisions taken over three days; the reasoning is kept, the fragmentation is not.*
The environment a change is run against before it reaches real machines.
## A node in the lab is a virtual machine
It boots a stock Linux image, runs the real install, and becomes a node. **It is not a model of
a node**, so no question arises about how good the model is — which is the whole reason for
paying the cost of virtual machines rather than containers.
The lab is driven by **incus**, and a scenario is raised from a declaration.
## A router is scenery, and is therefore a container
**Nothing under test runs on a router.** It is not a participant, holds no identity, has nothing
installed on it by the mesh, and no assertion is ever made about its internals. It exists so that
packets between machines behave the way they behave in the world.
The fidelity argument that makes a node a virtual machine does not reach it: what a router *is*
does not matter, only what it *does to traffic*. So a router is a system container, and the lab
is cheaper for it.
## A scenario declares the underlay, and only the underlay
**What a hosting provider and a home router would have provided**, before any of our software
touched the machine:
- which segments exist, and their address ranges
- which machine sits on which segment, at which address
- what NAT sits between them, and which ports are forwarded through it
- which machines are detached, and may be attached or detached during a run
**A scenario declares nothing about the overlay** — no overlay addresses, no hub, no peering, no
names, no certificates. Those are the mesh's job, and a scenario that supplied them would be
testing itself.
> A scenario provides what a hosting provider and a home router would provide, and nothing our
> software is responsible for.
## A scenario is a closed address space
Every segment materialises as its own isolated link belonging to one scenario instance. **Two
scenarios raised from the same declaration hold the same addresses and never meet**, because
nothing joins their links. The declaration therefore keeps its literal addresses and they mean
exactly what they say.
**The consequence that constrains everything else: the lab never reaches into a scenario over
IP.** It talks to a machine through the virtualisation layer's own channel — the way one would
use a console rather than the network. That is what makes two identical scenarios able to run at
once, and it is why placing anything inside a machine is a hypervisor operation rather than a
network one.
## Two scenario classes, and the first has no pipeline
| | **bootstrap** | **full** |
|---|---|---|
| contains | machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules |
| verdict from | what the host reports about the state it reconciled | a delivery result ending in verification |
| exercises | tiers 0 and 1 | tiers 2 and above, and modules |
**The bootstrap class comes first**, because it is what develops the node host, and because a
full scenario needs tiers that do not exist yet. A lab that could only raise the larger class
would be a lab nobody could use until everything else was built.
## Consequences
- **The lab tests the real code path**, not a reimplementation of it. The network a scenario
produces is generated by the same code production runs.
- **Isolation is what makes it usable in parallel**, and it costs the ability to reach in over
IP. Everything the lab puts inside a machine — a binary, an image, a file — goes through the
hypervisor.
- **A sealed scenario cannot fetch anything**, which is a real limit rather than an inconvenience:
it is why images have to be placed and why a container runtime has to be in the base image.
@@ -1,11 +1,12 @@
---
topic: checking it
status: accepted
date: 2026-08-24
deciders: jochen
reconstructed: false
---
# 34. A test defends a decision
# 17. A test defends a decision
## Context
@@ -1,97 +0,0 @@
---
status: proposed
date: 2026-08-23
deciders: jochen
reconstructed: false
extends: 0015-mesh-brokers-nodes-host-agents-think.md
---
# 17. Modules outside the platform core are grouped by domain, not by single function
## Context
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) recomposes the platform's own modules
into bounded contexts named after their aggregates, and sends the rest out of the monorepo on
the grounds that they run *on* the mesh rather than being *of* it.
That leaves the larger half unaddressed. Around three quarters of the catalogue are modules
that are neither part of the mesh's domain nor standalone applications: a firewall, a VPN, an
SSH daemon and a resolver; a file manager, a media player and a system monitor; a set of
media-library services. Today each is its own module, because one module is the unit of *one
piece of software*, and no other grouping exists.
The result is that the catalogue's shape records what was installed, not what anything is for.
Four modules that together constitute "how a node is reachable" have no relationship the mesh
can see: they cannot be assigned, versioned, reasoned about or replaced as one thing, and a
change to how the mesh handles connectivity has to be made four times.
This is the same failure ADR 0015 names for the core — *boundaries drawn by deployment accident
rather than by domain* — appearing outside it.
## Considered options
1. **Leave them as they are.** Rejected. The core gets domain boundaries and everything else
keeps accident boundaries, so the catalogue becomes harder to read after the refactor than
before it.
2. **One module per piece of software, with a tag or category field.** Rejected. A label is not
a boundary: it does not change what can be assigned, versioned or replaced as a unit, and it
drifts from the thing it labels.
3. **Group them into domain modules, each owning the software that serves one purpose.**
Proposed here.
4. **Extend ADR 0015's contexts to cover everything.** Rejected. Those contexts are named for
the mesh's own aggregates; a media library is not an aggregate of the mesh, and forcing it
into that model repeats the metaphor-naming mistake ADR 0015 exists to correct.
## Decision
*Proposed — the principle is settled; the domain list is not. See "Open" below.*
Modules that are not part of the platform core are grouped into **domain modules**. A domain
is named for the concern it serves, and owns the software that serves it. The unit stops being
one piece of software and becomes one purpose.
This extends ADR 0015 rather than replacing it. The eight bounded contexts for the mesh's own
domain stand unchanged. This decision covers what ADR 0015 leaves outside them.
Naming follows the same rule as the core: **name the domain for what it does, not for what it
is made of**. Connectivity, not a VPN implementation.
## Consequences
- A domain becomes assignable, versionable and replaceable as one thing. Changing how nodes
reach each other is a change to one module.
- The catalogue's shape starts describing purpose. A reader can tell what a mesh is *for* from
its module list.
- Swapping an implementation stops being a module replacement, with the data-volume and
provisioning consequences that carries, and becomes a change inside a domain.
- The count drops sharply, which is a symptom of the improvement rather than the point of it.
- **Grouping conceals.** A domain module hides which implementation is in use, and every
operational question — which port, which unit, which credential — gains an indirection.
- The migration is not free and has no obvious increments: a domain is only useful once
everything belonging to it has moved.
- Some modules genuinely serve one purpose and are already correctly sized. Grouping for its
own sake would be the same error in the other direction.
## Open
**The domain list is not settled and this record does not invent one.** What is decided is the
principle; what is not decided is the set. Candidate groupings are visible in the catalogue —
connectivity and reachability, node presentation and desktop, media libraries, observation and
metrics, storage and data services — but naming them here would be reconstructing a decision
that has not been taken.
Settling the list is a research effort, not an act of this record. Until it concludes, this
ADR stays `proposed`. That effort is
[`01-RESEARCH/005-domain-grouping`](../01-RESEARCH/005-domain-grouping/00-overview.md), and its
first measurement already narrows this record's scope: co-change analysis supports grouping for
reachability, argues against it for the provisioned infrastructure providers, and finds no
signal either way for the fifty modules that never change alongside anything.
## References
- [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) — the core decomposition this
extends, and its rule about naming a context after its aggregate.
- [ADR 0010](0010-applications-live-in-their-own-repository.md) — standalone applications are
already out of scope here; they are not domains and do not group.
- [`03-DESIGN/00-as-is/10-module-catalogue.md`](../03-DESIGN/00-as-is/10-module-catalogue.md)
— the catalogue's current shape, which is the evidence for the problem.
@@ -1,18 +1,19 @@
---
topic: checking it
status: accepted
date: 2026-08-24
deciders: jochen
reconstructed: false
extends: 0034-a-test-defends-a-decision.md
extends: 0017-a-test-defends-a-decision.md
---
# 35. A picture of a system is read from the system, never from what asked for it
# 18. A picture of a system is read from the system, never from what asked for it
## Context
A scenario declaration is a file. A raised scenario is a set of machines, links and rulesets.
The two are supposed to correspond, and the entire value of the lab rests on noticing when
they do not — [ADR 0034](0034-a-test-defends-a-decision.md) says a claim nothing checks is a
they do not — [ADR 0017](0017-a-test-defends-a-decision.md) says a claim nothing checks is a
claim that will quietly stop being true.
Drawing a scenario makes that concrete, and forces a choice that looks cosmetic and is not.
@@ -91,8 +92,8 @@ difference read off directly.
## References
- [ADR 0034](0034-a-test-defends-a-decision.md) — a claim nothing checks stops being true.
- [ADR 0031](0031-the-lab-provides-the-underlay.md) — why the lab must not supply what the
- [ADR 0017](0017-a-test-defends-a-decision.md) — a claim nothing checks stops being true.
- [ADR 0016](0016-the-lab.md) — why the lab must not supply what the
mesh is responsible for; the same instinct, applied to facts rather than to configuration.
- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — the fault in production
form.
@@ -0,0 +1,132 @@
---
topic: how we work
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 19. How this repository works
*Consolidated 2026-08-28 from ten records that were one decision seen from ten angles. The
reasoning is kept; the fragmentation is not.*
## The repository
**`novox/hq` is Novox's headquarters, and it is public.**
Company-scoped, not the mesh's. Today almost everything in it is about the mesh, because the
mesh is what Novox is building — a fact about the present rather than a definition. A second
product would live here too.
**Public** means written for a reader who is not its author and has no access to the mesh it
describes. Nothing here may contain routable addresses, real domain names, node names, absolute
paths, usernames or credentials. The test: *would this paragraph still teach a stranger running
an entirely different mesh?*
**Separate from the code** because the cadence differs — a decision changes when thinking
changes, not when code changes — and because a public repository cannot be a private one's
subdirectory.
**The naming rule:** a repository belonging to a product carries that product's prefix. A
company-scoped one does not. So this is `hq` and the mesh's are `mesh-*`.
| Repository | Tier | Holds |
|---|---|---|
| `novox/mesh-host` | 0 | the node host — the one binary installed by hand |
| `novox/mesh-substrate` | 1 | the pinned tier-1 services, as declarations |
| `novox/mesh-control` | 2 | the control plane and its contexts |
| `novox/mesh-surfaces` | 3 | tools, web, cli — thin, no logic |
| `novox/mesh-sdk` | — | the mesh's own domain ([ADR 0009](0009-modules-and-the-graph.md)) |
| `novox/mesh-lab` | — | the lab: scenario lifecycle, networking, placement |
| `novox/hq` | — | this one |
**The product is `Novox Mesh`**, shortened to `mesh` in internal use — repository names, the
module namespace, environment variables, paths. **`Nox` is an identity of Novox**, an agent
participant within the mesh's own model, not a second system.
## The folders, and why they are numbered
**The numbering is the flow.** Research produces a decision; the decision authorises a design.
Following the numbers walks the process in the order it happens.
| | |
|---|---|
| `01-RESEARCH` | an open question, while it is open |
| `02-DECISIONS` | what was decided, and why |
| `03-DESIGN` | what is being built |
| `04-ISSUES` | something wrong at the level of design or governance |
**`03-DESIGN` has two layers and they are never mixed.** `00-as-is/` describes the mesh that
exists, written from the implementation and the operational record. `01-to-be/` describes the one
being built toward. Every document says which it is. A statement about the future does not belong
in an as-is document, and an as-is document is never edited to describe an intention.
**`04-ISSUES` is for design-level faults** — a rule enforced by nothing, a stated invariant that
is false, a failure the design permits to be silent. Not an operational ticket queue.
## What a decision record is, and is not
**If a decision is worth recording, it is worth a record. If it is not worth a record, it is not
recorded.**
That bar has been read too generously. A *finding* is not a decision. A bug is not a decision.
**A record is warranted when there is a genuine fork**: a direction reversed, an alternative
seriously considered and likely to be proposed again, or something contested that needs to stay
settled. Everything else belongs in the design document, where the reasoning is read.
**There is no ledger** — no separate document summarising, ranking or tracking decisions. A
chronological view is generated from frontmatter, which is what a ledger was actually for.
**A number identifies a record and never changes.** It is not a position, and it cannot be
both — a position moves when the set changes, and an identity that moves is not one.
That is not a preference. Records are referenced from **outside** this repository: code
comments, commit messages, the knowledge base. Renumbering once cost 96 references across two
code repositories, and nothing in either would have failed to compile — the comments would
simply have pointed at the wrong reasoning, which is worse than a broken link because nothing
reports it.
**So the reading order lives in a generated index**, from each record's `topic:` — what the mesh
is, then its tiers from the bottom up, then what runs on them and how it gets there, then how it
is built, how it is checked, and how we work.
**And the index is written, not only generated on demand.** A reader looking at the folder on a
forge sees the folder, not a command. The objection to a written index is that it drifts, and
that is answered by **checking** it rather than by refusing to write one — which is §5's own
rule: a rule states how it is checked. A record with no topic, or a topic nobody defined, fails
the same check, because the quiet failure is a record that vanishes from the order rather than
appearing in the wrong place.
**The design layer is what you read.** These records explain *why* a thing is as it is. They are
not a description of the system, and needing to read them to understand it would mean the design
documents had failed.
## Status, and views over it
**Every document carries its state in YAML frontmatter** — research overviews, design documents,
decision records, issue reports.
**There are no central status files.** Every cross-cutting view — a status matrix, a decision
index, an open-issue list — is generated from frontmatter when asked for, never written to disk.
Two places holding one fact drift, and the written one wins by being closer to hand.
**Prose does not restate status.** One place, and two is one too many.
## Workflows are playbooks
Every workflow is a playbook in [`00-META/process/`](../00-META/process/): trigger, who runs it,
steps, outputs. People and agents follow the same ones, and **agents do not act outside them**.
Each is wrapped by a thin skill that defers to the playbook as authoritative and adds only the
mechanical scaffolding — so the process has one definition rather than a document and an
implementation that disagree.
## Consequences
- **A reader has one place per thing.** The design layer describes the system; these records say
why; the playbooks say how work is done.
- **Records will accumulate more slowly**, because the bar is a fork rather than a finding. This
record is itself the correction: ten records became one because they were one decision.
- **The public rule constrains everything written here**, permanently and at every commit. It is
the reason research describes real observations without identifying the mesh it observed.
@@ -1,69 +0,0 @@
---
status: accepted
date: 2026-08-22
deciders: jochen
reconstructed: false
supersedes: none
---
# 19. HQ is its own repository, and it is public
## Context
The mesh's reasoning — mission, research, design, decisions — began inside the code
repository, under a folder there. The objection to moving it out was specific and good: the
mesh already has an operational memory and a structured archive, and adding a third store
repeats the mistake that consolidation was meant to fix.
## Considered options
1. **Keep it in the code repository.** Rejected, but the objection it rests on is correct and
is answered rather than dismissed — see Consequences.
2. **Put it in the structured archive**, alongside the governed documents. Rejected: the
archive is not reviewable as a diff, and a design argument is exactly the thing that needs
line-by-line review and a branch.
3. **Its own repository.** Chosen.
## Decision
HQ is its own repository, and it is **public** — written for a reader who is not its author
and has no access to the mesh it describes.
Three reasons it is separate:
- **The cadence differs.** A decision changes when thinking changes, not when code changes.
Tying documents to a code branch merges them on the code's schedule.
- **The reviewers differ.** A design argument is not reviewed the way an implementation is,
and should not queue behind a build.
- **The scope is wider than one repository.**
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) sends most modules out of the
monorepo; documentation governing several repositories cannot live inside one of them.
Being public is not incidental. It is enforceable only because
[ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) made the code repository
node-agnostic: there is no per-node content to leak. Nothing here may carry routable
addresses, real domain names, hosting providers, node names, absolute paths, usernames,
credentials, or operational detail useful only to an attacker.
The test is whether a paragraph would still teach a stranger running an entirely different
mesh.
## Consequences
- A document and the code it describes can no longer land in one commit. Keeping them honest
is a discipline rather than a mechanism — which is why decisions are recorded as they are
taken, and why a document stating a rule must say how the rule is checked.
- Research must state evidence without identifying the mesh it observed. The shape of a
finding survives anonymisation; the instance does not travel.
- **The objection is answered by indexing, not by location** — the claim being that these
documents remain searchable beside everything else, one source with many surfaces.
**That indexing does not exist.** Checked 2026-08-23, it returns nothing. Until it does, the
objection stands unanswered and this repository is the third knowledge store it was argued
not to be. Recorded as
[`04-ISSUES/006`](../04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md).
## References
- Supersedes the earlier position that documentation lives inside the code repository under a
folder there. That position was never recorded separately and has no record of its own.
- [`README.md`](../README.md) — the public-repository rule in full.
@@ -1,67 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 20. Design is written in two layers: what is, and what is intended
## Context
HQ held only the intended mesh. Every reader had to already know the running system in order
to understand what the decisions were about, and a statement about current behaviour had
nowhere to live except inside a document describing an intention.
The consequence was invisible until looked for: an as-is claim inside a to-be document is
indistinguishable from the intention around it, so the document silently stops being true as
the system moves — and nobody can tell which half went stale.
The mesh also has a large body of shipped behaviour that nobody would choose again. It is not
design in the sense of "what we decided"; it is design in the sense of "what is there", and it
is exactly the part a person changing the system most needs.
## Considered options
1. **One layer, describing the target.** Rejected — the status quo. The running system goes
undocumented and the target document accumulates unmarked claims about it.
2. **One layer, describing what runs, with intentions only in decision records.** Rejected:
a decision record is an argument, not a specification, and a multi-part intention has
nowhere coherent to live.
3. **Two layers, declared per document, never mixed.** Chosen.
## Decision
`03-DESIGN` holds two layers, and every document declares which it is:
| Layer | Describes | Written from |
|---|---|---|
| `00-as-is/` | The mesh that exists | The implementation and the operational record |
| `01-to-be/` | The mesh being built toward | Decision records |
An as-is document **records what is, not what should be** — including behaviour nobody would
choose again. A layer that keeps only the good decisions is a brochure.
When a to-be design ships it **does not move**. Its as-is counterpart is written or updated,
the to-be document's status becomes `implemented`, and both stand: one describing what runs,
the other recording what was intended. Deleting the intention loses the reasoning.
Where implementation and intention disagree, the as-is document records the implementation and
says they disagree.
## Consequences
- A reader can tell, from the folder and from one frontmatter field, whether they are reading
a description or a plan. That distinction was previously unavailable at any price.
- Correcting an as-is document requires evidence from the implementation, not agreement — and
needs no decision record, because nothing was decided.
- Two documents must be kept current per subsystem instead of one. This is the cost, and it is
paid on every ship.
- Something that shipped differently from its design becomes a visible divergence rather than
a silently wrong document, and may deserve an issue.
## References
- [`03-DESIGN/README.md`](../03-DESIGN/README.md) — the layer contract and frontmatter schema.
- [`03-DESIGN/00-as-is/`](../03-DESIGN/00-as-is/) — the first eleven as-is documents, written
2026-08-23 from the monorepo and the operational memory.
@@ -1,11 +1,12 @@
---
topic: how we work
status: accepted
date: 2026-07-10
deciders: jochen
reconstructed: true
---
# 9. The mesh is governed by a constitution, injected where work is decided
# 20. The mesh is governed by a constitution, injected where work is decided
> Reconstructed after the fact from the evidence cited below.
@@ -1,15 +1,16 @@
---
topic: how we work
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 25. HQ is the source of the mesh constitution
# 21. HQ is the source of the mesh constitution
## Context
[ADR 0009](0009-the-mesh-is-governed-by-a-constitution.md) established a canonical rule set,
[ADR 0020](0020-the-mesh-is-governed-by-a-constitution.md) established a canonical rule set,
injected into every eligible design session and checked before output is accepted. It lives in
the knowledge base, where the orchestrator reads it.
@@ -39,7 +40,7 @@ directly.
Publishing is a playbook step, not a manual act, and it ends with **reading the page back and
verifying the change is present**. A publish that reported success and did nothing is exactly
the failure class this mesh keeps producing
([ADR 0008](0008-a-failed-step-fails-the-job.md)).
([ADR 0010](0010-delivery.md)).
Section numbering is stable, because the orchestrator and the review fragments cite sections by
number.
@@ -47,7 +48,7 @@ number.
## Consequences
- One source, many surfaces — the same argument HQ's separation already rests on
([ADR 0019](0019-hq-is-its-own-repository.md)), applied to the rules themselves.
([ADR 0019](0019-how-this-repository-works.md)), applied to the rules themselves.
- Each rule keeps the incident that earned it, in a place that is reviewed as a diff.
- An edit to the derived page survives until the next sync and then vanishes. The playbook says
so, and nothing mechanically prevents it.
@@ -63,5 +64,5 @@ number.
- [`00-META/process/05-constitution-sync.md`](../00-META/process/05-constitution-sync.md) —
the sync, including the read-back.
- [ADR 0009](0009-the-mesh-is-governed-by-a-constitution.md) — the governed page and why it
- [ADR 0020](0020-the-mesh-is-governed-by-a-constitution.md) — the governed page and why it
exists.
@@ -1,57 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 21. Every workflow is a playbook, and agents operate through them
## Context
HQ stated a knowledge flow — research becomes design — and nowhere stated how anything moves
along it. What graduation required, who wrote the decision, what closed an effort, what
happened when something shipped: all of it was convention held in one person's head.
A large share of the work here is done by agents. An unwritten convention is not available to
an agent at all, so each one either invents a procedure or asks. Both produce a repository
whose shape depends on who last touched it.
## Considered options
1. **Convention, learned by reading existing documents.** Rejected — the status quo. It
transmits shape but not rules, and it transmits the mistakes along with the pattern.
2. **One long contributing document.** Rejected: it is read once, and the step someone needs
is never the step they are reading.
3. **A playbook per workflow, each with trigger, steps and outputs, wrapped by a thin skill.**
Chosen.
## Decision
Every workflow is a playbook in [`00-META/process/`](../00-META/process/): trigger, who runs
it, steps, outputs. Engineers and agents follow the same playbooks, and **agents must not act
outside them**.
Each playbook is wrapped by a thin skill that defers to it as authoritative and adds only
mechanical scaffolding — next free number, frontmatter block, where the file goes. The
playbook holds the reasoning; the skill holds the steps. When they disagree, the playbook
wins.
## Consequences
- An agent arriving with no context can act correctly, because the procedure is retrievable
rather than remembered.
- The playbooks are themselves reviewable. A bad rule can be found and changed, which is not
true of a convention.
- Duplication between playbook and skill is real, and is managed by making the skill thin and
naming the playbook as authoritative in the skill's first lines. Nothing prevents them
drifting; the constraint is that only one carries reasoning.
- A workflow with no playbook is a workflow agents will get wrong. Adding one is part of
adding the workflow.
## References
- [`00-META/process/00-overview.md`](../00-META/process/00-overview.md) — the five playbooks
and the flow they implement.
- Modelled on the process layer in the sibling HQ repository for the PAPA platform, which
arrived at the same shape and the same thin-skill split.
@@ -1,54 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 22. Status lives in frontmatter; cross-cutting views are generated
## Context
Status was carried in prose — a bold line near the top of a document saying what state it was
in — and indexes were maintained by hand. The decision-record index had already drifted from
the folder it described **after a single addition**, which is about as short a demonstration
as the failure mode offers.
A hand-maintained index is a copy of something the filesystem already knows. It is correct
only for as long as everyone remembers it exists, and its being wrong is silent.
## Considered options
1. **Prose status plus hand-maintained indexes.** Rejected — the status quo, already
demonstrably broken.
2. **A central status file.** Rejected. It centralises the drift rather than removing it: the
file and the documents disagree, and the file is the one people read.
3. **Machine-readable frontmatter per document; every cross-cutting view generated on
demand.** Chosen.
## Decision
Every document carries its state in YAML frontmatter — research overviews, design documents,
decision records, issue reports — with a schema stated in the section README.
**There are no central status files.** Every cross-cutting view — a status matrix, the
decision-record index, the open-issue list — is generated from frontmatter when asked for, and
never written to disk.
Prose does not restate status. One place, and two is one too many.
## Consequences
- A view cannot drift from what it describes, because it does not persist.
- Status becomes queryable. Inconsistencies — a closed effort with nothing in `became:`, an
`implemented` design with no owning repository — are findable mechanically, and the
generator reports them as flags rather than silently rendering around them.
- Frontmatter must be valid and paths in it must resolve, which is now something to check.
- A reader browsing the repository on a forge sees no index. That is the trade: the index is
correct and absent rather than present and wrong.
## References
- [`.claude/skills/hq-status/SKILL.md`](../.claude/skills/hq-status/SKILL.md) — the
generator, including the inconsistencies it flags.
- [`02-DECISIONS/README.md`](README.md) — the hand-written index that drifted, and its removal.
@@ -1,15 +1,16 @@
---
topic: how we work
status: accepted
date: 2026-08-26
deciders: jochen
reconstructed: false
---
# 40. The constitution absorbs what is already enforced
# 22. The constitution absorbs what is already enforced
## Context
[ADR 0025](0025-hq-is-the-source-of-the-constitution.md) makes this repository the source
[ADR 0021](0021-hq-is-the-source-of-the-constitution.md) makes this repository the source
and the knowledge base a derived copy, and playbook
[05](../00-META/process/05-constitution-sync.md) publishes the copy whenever a rule changes.
@@ -97,8 +98,8 @@ trusting the second success message either.
## References
- [ADR 0025](0025-hq-is-the-source-of-the-constitution.md) — source and copy.
- [ADR 0009](0009-the-mesh-is-governed-by-a-constitution.md) — why the copy is injected at all.
- [ADR 0021](0021-hq-is-the-source-of-the-constitution.md) — source and copy.
- [ADR 0020](0020-the-mesh-is-governed-by-a-constitution.md) — why the copy is injected at all.
- [Playbook 05](../00-META/process/05-constitution-sync.md) — the sync this record interrupts.
- [ADR 0034](0034-a-test-defends-a-decision.md), [ADR 0035](0035-a-picture-is-read-from-what-runs.md),
[ADR 0018](0018-the-mesh-creates-no-symlinks.md) — the three rules whose sync surfaced this.
- [ADR 0017](0017-a-test-defends-a-decision.md), [ADR 0018](0018-a-picture-is-read-from-what-runs.md),
[ADR 0012](0012-the-mesh-creates-no-symlinks.md) — the three rules whose sync surfaced this.
@@ -1,11 +1,12 @@
---
topic: how we work
status: accepted
date: 2026-08-26
deciders: jochen
reconstructed: false
---
# 42. The approval is the checkpoint, not the second pair of hands
# 23. The approval is the checkpoint, not the second pair of hands
## Context
@@ -73,5 +74,5 @@ the previous sync reported success and changed nothing.
## References
- [`how-we-build.md`](../00-META/how-we-build.md) §2 — the rule, now carrying this.
- [ADR 0008](0008-a-failed-step-fails-the-job.md) — the standard a checkpoint is held to: a
- [ADR 0010](0010-delivery.md) — the standard a checkpoint is held to: a
step that reports success without doing anything is the fault, not the shortcut.
@@ -1,60 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 23. Issues have a front door, separate from the operational memory
## Context
Findings that were nobody's task accumulated in a table inside the decision ledger — a package
install reporting success while installing nothing, a manifest key read by no code, an
end-to-end harness dead for months. They were measured, true, and unowned: a table row cannot
be assigned, diagnosed or closed.
The mesh already has an operational memory holding roughly a hundred and thirty entries,
indexed on symptoms. The obvious move — put these there — is wrong, and the reason is the
distinction worth recording.
## Considered options
1. **Leave them in the ledger.** Rejected: a ledger records decisions taken, and these are
the opposite — questions nobody has answered.
2. **Put them in the operational memory.** Rejected. That store answers *how do I fix this
occurrence*; these are *why does the design allow this at all*. Filing them there makes
them findable by symptom and unfindable as open questions, and nothing there has a state
that can be closed.
3. **A numbered issue folder in HQ, deliberately narrow.** Chosen.
## Decision
`04-ISSUES` is the front door for something wrong at the level of **design or governance**:
a rule enforced by nothing, a stated invariant that is false, a failure the design permits to
be silent, or a symptom whose owner cannot be found without the whole mesh in view.
One numbered folder per issue: the report with the symptom as observed and the evidence, and
a diagnosis document carrying the trail, dated, including what was ruled out.
**This is not a second copy of the operational memory.** An issue here is a question HQ must
*answer*; an entry there is an incident someone must *clear*. An issue whose answer is a
general lesson belongs in both — and the operational memory is searched first, because if the
answer is already there this was never an issue.
## Consequences
- A finding gets a number, a state and an owner, and closing it is a visible act.
- The symptom-to-component trail accumulates in a place where the whole mesh is visible, which
is where cross-component diagnosis has to happen.
- The boundary needs judgement on every report, and will sometimes be got wrong. Filing too
narrowly loses a finding; filing too widely rebuilds the operational memory here, which is
the outcome HQ's separation was argued against
([ADR 0019](0019-hq-is-its-own-repository.md)).
- Six issues opened on creation, all previously unowned observations.
## References
- [`04-ISSUES/README.md`](../04-ISSUES/README.md) — the boundary table and the frontmatter
schema.
- [`00-META/process/03-issues.md`](../00-META/process/03-issues.md) — the playbook.
@@ -0,0 +1,95 @@
---
topic: what runs on it
status: accepted
date: 2026-08-30
deciders: jochen
reconstructed: false
---
# 24. Model access is a provision, and a licence is a thing with a name
## Context
Everything in this mesh that thinks needs a model, and there is more than one way to reach one:
| | |
|---|---|
| **hosted services** | several vendors, each with its own account, quota and key |
| **models the mesh runs itself** | open-weight models on a node with the hardware for them |
And the choice is **per consumer, deliberately**: a workstation's own session on one account, a
laptop on the organisation's, two hired workers on the mesh's local model because their work does
not justify paid tokens. Those are three different answers to one requirement, held at once, in
one mesh.
**The existing system has the hard half of this already** — automatic licence refresh and
switching between accounts when one is exhausted — and it works. It is not being replaced because
it was wrong; it is being rebuilt because it lives in a place that cannot express the rest.
## Decision
**Model access is a provision.** A module that needs to think declares `requires: model-access`;
anthropic, openai, grok and a locally-run model are four modules that provide it. That is
[ADR 0009](0009-modules-and-the-graph.md)'s mechanism unchanged, and it buys the things that
mechanism already buys: several implementations of one job, a refusal when more than one could
answer, and choosing by assigning the one you want.
**A locally-run model needs nothing new at all.** It is a mesh-scoped provision on the node with
the hardware — the same shape as a database, including the credential.
### A licence is a named thing, and the name is the operator's
Not an anonymous credential hanging off a provider. *The personal account*, *the organisation's
account* — those are names a person uses, and the mesh has to use them too, because the whole
point is saying **which one** a given consumer uses.
**Many to many.** One provider has several licences; one licence serves several consumers. So it
is **not a claim** — claims are for things only one holder may have, and two machines sharing an
account is the ordinary case rather than a collision.
### Four things this needs that the mesh does not have
Written as gaps rather than as design, because each is a real piece of work and pretending
otherwise is how a plan becomes a surprise.
**1. A provider that is not on a node.** A mesh-scoped provision today is answered by *the machine
running it*, and the reachability rule refuses two ends that share no private network. A hosted
service is on nobody's machine and is reached over the public internet. That is a third scope —
answered by a record rather than by a node — and the reachability rule must not apply to it.
**2. A secret the mesh is given rather than one it mints.** Every credential the mesh handles
today it generated itself, sealed to both ends, and discarded. An API key arrives from a person.
The missing verb is *accept*: take a value, seal it to each holder, and **discard the plaintext**
— because a mesh that keeps operator-supplied keys readably is the arrangement this project
[measured and rejected](0009-modules-and-the-graph.md).
**3. A consumer that is not a machine.** *This worker uses that licence* is a binding to an agent,
not to a node. [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) already says an agent
holds credentials and that delivery follows its node bindings and modality — so what is delivered
still lands on a machine, and what is **chosen** is chosen per agent. The provisions model has no
consumer identity other than a node.
**4. Switching is a reaction, not a declaration.** Everything here is desired state, reconciled by
comparison. A licence that hits its limit and must be swapped is a response to something observed,
and it cannot be expressed as a declaration without the declaration meaning *whatever is working
right now* — which is not a thing anybody declared. **It belongs with observability, changing a
binding**, and the binding is then declared as usual. Saying this plainly is what stops the
declaration language growing a conditional.
## Consequences
- **The refusing rule applies here and will be felt.** A mesh holding three ways to reach a model
refuses every consumer that has not said which — which is correct and is a great deal of
saying-which the first time. The remedy is one assignment per consumer, and the message names
the candidates.
- **A licence outliving its holder is a live credential nobody is watching.** The same rule the
provisioner follows applies: what the mesh granted and no longer grants is withdrawn.
- **Nothing here makes a node authenticate to a model provider.** ADR 0001 holds: an agent does.
What changes is that the mesh can now say *which agent, which licence, which node it lands on*,
which is the fact ADR 0001 records as missing.
## References
- [ADR 0009](0009-modules-and-the-graph.md) — provisions, scope, choosing, and sealed credentials
- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — agents hold credentials, not nodes;
`hal/ai` as the context owning provider grants and rotation
@@ -1,76 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 24. The folder numbering is the flow, and decision records run oldest first
## Context
Two orderings were wrong in ways that only show up when someone new reads the repository.
**The folders.** Decisions lived in an unnumbered folder that sorted after the numbered ones,
so the repository's most load-bearing content read as an annex.
The sibling HQ repository for the PAPA platform had already solved this and solved it
crookedly: its design folder existed from its initial commit, and when its decision folder was
finally promoted it took the **next free number** rather than its place in the sequence. That
repository now reads `01 research → 03 decision → 02 design`. A decision precedes the design
it authorises and is numbered after it. By the time this was visible, the design folder was too
settled to renumber.
**The records.** Fourteen decisions had been taken in implementation and never written down —
the broker, the module abstraction, the mesh database, the artifact, the three silos and the
rest. Meanwhile two records existed, holding numbers 0001 and 0002, for decisions taken last.
## Considered options
1. **Match the sibling repository exactly**, inheriting its ordering. Rejected: structural
parity is worth something, but not the cost of copying a scar the other repository would
not choose again.
2. **Leave the folder unnumbered.** Rejected — the annex problem, and it leaves an unexplained
gap for anyone arriving from the sibling repository.
3. **Number by position in the flow, and renumber the records chronologically.** Chosen, on
the grounds that this repository was four commits old and nothing outside it cited a
number. That is the only window in which either renumbering is free.
## Decision
**The numbering is the flow.** Research produces a decision; the decision authorises a design.
So `01-RESEARCH`, `02-DECISIONS`, `03-DESIGN`, `04-ISSUES`. Following the folder numbers walks
the process in the order it happens.
**Decision records are a chronological ledger.** They run oldest first. The fourteen decisions
already taken in implementation were back-filled as records 0001–0014, each dated from the
history, each carrying `reconstructed: true` and saying so in its first lines, and each citing
the commit, pull request or knowledge-base entry it was recovered from. The two existing
records moved to 0015 and 0016.
A reconstructed record is not a transcript. Where the deliberation is not recoverable it states
what the alternatives were and why the chosen one won on the evidence available — not a
discussion that did not happen. Where a date is not establishable it says so.
The foundational folder is `00-META`, matching the sibling repository.
## Consequences
- The repository reads in process order, and the gap at `03` that a reader coming from the
sibling repository would notice is explained by this record.
- The design documents can cite reasoning instead of asserting rules, because the reasoning now
exists.
- Structural divergence from the sibling repository, deliberately, in exactly one place. It is
recorded here so that the difference reads as a choice rather than an accident.
- **Record numbers are now stable and renumbering is over.** This decision spends the one
window that existed; a future record takes the next free number regardless of its date.
- Reconstructed records carry a standing risk: they are the most confident-sounding documents
in the repository and the least directly witnessed. The `reconstructed` flag exists so that
is never invisible.
## References
- The sibling repository's restructure of 2026-07-13 moved its decision folder in a single
commit of twelve renames with no content change, alongside the same status-into-frontmatter
and playbook changes made here.
- [`02-DECISIONS/README.md`](README.md) — the format, and the note on reconstructed records.
@@ -0,0 +1,105 @@
---
topic: how we work
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0019-how-this-repository-works.md
---
# 25. The design record is read where it is written, never copied to be found
## Context
**These documents cannot be found by searching the mesh's memory, and never could.** Checked on
2026-08-23 and again on 2026-08-31, against both the symptom-indexed store and the structured
archive, using a decision record's full title and a distinctive phrase from a design document: no
result, no partial match, no stale copy.
That matters because of what was promised. The objection to giving this material its own
repository was that the mesh already has a knowledge store, and a second one repeats the mistake
that store was created to fix. **The answer offered was indexing rather than location** — that
these documents would be returned beside everything else in a search, so where they were authored
became a separate question. The indexing was never built.
**The claim has since stopped being load-bearing**, which is why this is a decision rather than an
incident. [`README.md`](../README.md) names the gap in the place the claim used to sit, and
[ADR 0019](0019-how-this-repository-works.md)'s reasoning rests on cadence, reviewers and scope —
none of which depend on being searchable from elsewhere. What remained was an unbuilt capability
and an open question, recorded as
[`04-ISSUES/006`](../04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md).
**A signpost was added on 2026-08-31 and measured.** One entry in the mesh's memory naming what
lives here and when to come looking. A search for *design records, decisions, repository* returns
it; a search phrased the way somebody actually asks — *why is the mesh built this way* — returns
nothing, because the store matches terms and not meaning. **Reachable is not the same as
surfacing**, and the measurement is what established which one a signpost buys.
## Considered Options
1. **A one-way sync into the mesh's memory.** A job reads this repository on a schedule and writes
the documents into the searchable store. It works with what exists today and needs nothing
built first. **Rejected**, because it creates a second copy of every document, and the failure
mode of a derived copy is the one this repository is least able to tolerate: *the copy that is
searched quietly stops matching the copy that is edited*, and the enforced one wins. A design
record that has silently diverged from the reasoning it claims to carry is worse than one that
cannot be found — the first misleads, the second merely fails.
2. **Leave the signpost and close nothing.** Honest, free, and it keeps the gap visible.
**Rejected as an end state**, though it is what stands until the option below exists. It
answers only for a reader who already suspects these documents exist, which is precisely not
the reader the mesh's memory is designed for.
3. **An agent reads this repository directly, and the search consults it.** Nothing is copied.
**Adopted.**
## Decision
**The design record is read where it is written.** Retrieval is an agent reading this repository,
not a copy living in a second store — and a search of the mesh's memory consults that agent, so
what it knows appears beside ordinary results rather than only when it is asked.
Both halves are the decision. The first alone is merely a reader, and would leave this repository
reachable but not surfacing — the state measured above. **The second half is what discharges the
promise** that these documents are returned beside everything else.
**There is no copy, and that is the point.** No sync, no schedule, no reconciliation, and nothing
that can drift, because there is only ever one of each document. It is also always current,
including for work that is not yet committed.
**The direction of reading is one-way and stays that way.** The agent reads this repository and
answers from it. Nothing flows back: this repository is public, the mesh is not, and a return path
would be how installation-specific detail arrives into documents that must not carry it
([`README.md`](../README.md)).
## Consequences
**This repository stops being a fourth knowledge system, properly.** The original objection was
about adding a knowledge *system*. An agent with read access adds no store at all — which answers
the objection more completely than the indexing that was promised, rather than merely as well.
**ADR 0019's promise is amended, not satisfied.** It said these documents would be *indexed*. They
will not be. They will be *read*, and the search will ask. The commitment that survives is the one
that mattered — that a searcher finds them without already suspecting they exist — and the
mechanism behind it is different from the one named.
**It is gated on an agent that does not exist yet.** Until it does, the signpost is what stands,
and this repository is reachable rather than surfacing. That is a known and stated gap, not a
silent one — and the gap is now a build task with a decided shape rather than an open question.
**The search must degrade honestly.** When the agent cannot be reached, a search has to say that
this material was not consulted. A result set that silently omits it looks identical to one where
nothing matched, and *silence and success must never look alike*
([ADR 0004](0004-a-node-and-how-it-joins.md)) — the rule this repository has now paid for twice.
**A rule states how it is checked, and this one is checkable.** The check is the measurement that
produced this record: search the mesh's memory for a phrase that appears only in a design document
here, and require it back. That check fails today, deliberately, and passing it is what closes
`04-ISSUES/006`.
## References
- [`04-ISSUES/006`](../04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md) —
the gap, the two measurements, and why closing it early was refused
- [ADR 0019](0019-how-this-repository-works.md) — the promise this amends
- [`README.md`](../README.md) — the objection, and the gap named where the claim used to sit
@@ -1,89 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 26. Every decision is a record; there is no ledger
## Context
HQ carried a decision ledger at its root: a chronological table of forty-one numbered
decisions, each with who decided and a pointer to where the reasoning lived. It was created
deliberately, to make decisions findable and to give a home to decisions too small to warrant
a document.
By the time the decision records were back-filled
([ADR 0024](0024-the-numbering-is-the-flow.md)) the ledger had become three things at once,
and only one of them was still needed.
Classified, its forty-one entries were: ten restating a record, eleven restating design
documents, fifteen describing how this repository works — with the reasoning in a README rather
than anywhere citable — three small rules with no home at all, and two superseded stubs.
So the ledger was mostly a copy. Worse, it was a **hand-maintained index**, which
[ADR 0022](0022-status-lives-in-frontmatter.md) had just finished rejecting for the
decision-record index on the grounds that it had drifted after a single addition. Keeping one
copy of that pattern while removing another is not a position.
It had also produced a naming collision that a directory listing makes plain: `DECISIONS.md`
beside `02-DECISIONS/`, holding different things.
## Considered options
1. **Keep the ledger.** Rejected. It duplicates the records, restates status, and is the exact
hand-maintained index this repository decided against elsewhere.
2. **Keep it, renamed, for small decisions only.** Rejected, and this is the option worth
arguing with — it is genuinely useful to record a decision without writing a document. But a
decision small enough to be one table row is almost always a **rule** rather than a
decision, and a rule belongs in [`how-we-build.md`](../00-META/how-we-build.md) where it is
enforced and where its reasoning is kept. That is where the three orphans went.
3. **Every decision is a record; nothing else.** Chosen. This is how the sibling HQ repository
for the PAPA platform works, and it has no ledger of any kind.
## Decision
**If a decision is worth recording, it is worth a record. If it is not worth a record, it is
not recorded.**
`02-DECISIONS` holds every decision. There is no ledger, no index file, and no central status
of any kind. The chronological view — decisions in the order they were taken — is *generated*
from record frontmatter, which is what the ledger was actually for.
Content that was only in the ledger was rehomed rather than dropped:
| Was | Went to |
|---|---|
| Decisions about how this repository works | Records [0019](0019-hq-is-its-own-repository.md)–[0025](0025-hq-is-the-source-of-the-constitution.md) |
| Small rules with no record | [`how-we-build.md`](../00-META/how-we-build.md) — the package rule, and two already there |
| Lab decisions not stated in the design | [`03-DESIGN/01-to-be/01-end-to-end-testing.md`](../03-DESIGN/01-to-be/01-end-to-end-testing.md) |
| "Deliberately not decided" | The research effort and design document each question belongs to |
| Unowned observations | [`04-ISSUES`](../04-ISSUES/) ([ADR 0023](0023-issues-have-a-front-door.md)) |
## Consequences
- One place to look, and nothing to keep in sync. The collision between the ledger and the
record folder is gone.
- Structural parity with the sibling repository on decisions, which
[ADR 0024](0024-the-numbering-is-the-flow.md) deliberately broke on folder numbering. The
divergence is now exactly one thing, and it is the one thing that was argued for.
- **Writing a record is now the only way to record a decision, and a record is more work than
a table row.** The real risk is that a small decision goes unrecorded because nobody wanted
to write a document. The mitigation is that a small decision is usually a rule, and
`how-we-build.md` takes rules cheaply — but this is a cost, not a solved problem, and it is
the thing to watch.
- The chronological view now depends on the generator existing and being run. It did not
before.
- Two superseded ledger stubs had no record of their own. The position that documentation
lives inside the code repository is now recorded only as superseded context in
[ADR 0019](0019-hq-is-its-own-repository.md); the system-container position is explained in
[ADR 0016](0016-a-lab-node-is-a-virtual-machine.md). Neither is lost.
## References
- The sibling PAPA HQ repository: root holds only agent instructions and a README; every
decision is a numbered record, and its graduation playbook has no path for an unrecorded
decision.
- [ADR 0022](0022-status-lives-in-frontmatter.md) — the hand-maintained-index argument this
applies consistently.
@@ -0,0 +1,130 @@
---
topic: what runs on it
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0004-a-node-and-how-it-joins.md
---
# 26. The mesh has a session of its own, and it is the node session's mechanism
## Context
[ADR 0004](0004-a-node-and-how-it-joins.md) gives every node a session: one per node, permanent,
remembering across callers, its system prompt the node's engram, reachable over the broker like
everything else. **Any node can message any node**, and that is called the one part of the system
that is genuinely a mesh — symmetric, with no centre.
**There is no way to address the mesh itself.** A question that spans machines — *what is running
across all of this*, *which nodes are behind*, *why is it built this way* — has to be put to some
node, which then asks the others. That works, and it makes a mesh-wide question **nobody's
question**: every node answers it as a foreigner, from a position where the whole is not in view.
**Three things independently arrived at the same missing piece.**
[ADR 0025](0025-the-design-record-is-read-not-copied.md), taken hours before this one, commits to
an agent that reads the design repository directly and answers into search. That agent has to
exist, run somewhere, and be askable — and nothing in the record says what it is or where it
lives.
[`14-model-access.md`](../03-DESIGN/01-to-be/14-model-access.md) records, as a gap deliberately
not half-built: *this worker uses that licence is a binding to an agent, not to a node* — and the
provisions model has no consumer identity other than a node. A session that must be assigned a
licence is exactly that consumer, and node sessions are already one.
**And ADR 0004 never said how a session is set up.** It describes behaviour and stops: nothing
states how a session starts, where its context lives, how the engram reaches it, or how a message
off the broker becomes a prompt. There is no design document for it. That gap was invisible until
something had to be built *like* a node session, because describing a second instance of a
mechanism requires the mechanism to have been described once.
## Considered Options
1. **No mesh session; keep relaying through a node.** Costs nothing and works today. **Rejected.**
It leaves mesh-wide questions belonging to nobody, and it does not survive contact with
ADR 0025 — that agent still needs a home, so the thing gets built anyway, unnamed, as an
attachment to whichever node happened to host it.
2. **A new kind of agent, built separately.** Purpose-built for the whole mesh. **Rejected.** It
would hold a session, a memory, a licence and broker plumbing — every one of which the node
session already has. Two implementations of one mechanism drift, and the vocabulary collision
that [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) exists to undo began exactly this
way: two things that were nearly the same, built twice, until neither word meant one thing.
3. **The same mechanism, started in a different context.** **Adopted.**
## Decision
**The mesh has one session, addressed as the mesh, and it is a node session in every respect but
three.**
| | |
|---|---|
| **the context it starts in** | the mesh's, not a machine's — this is the whole of what makes it different |
| **its engram** | the mesh's system prompt, as a node's engram is that node's |
| **its licence binding** | assigned in its own right, not inherited from the machine it runs on |
Everything else is unchanged and deliberately so: it is permanent, it remembers, it is reachable
over the broker, it holds its own tools, and switched off it still answers *I am switched off*
rather than falling silent.
**It runs on the node that holds the control plane** — not for convenience, but because that node
is already the one place excepted from *compromise of a node is compromise of that node*
(ADR 0004). An agent able to reach everything, placed anywhere else, creates a **second** such
place. Putting it where the authority already sits concentrates nothing new.
**It is an addition to per-node messaging and never a replacement.** Every node remains directly
addressable. This is not a preference: ADR 0001 holds that losing the control plane costs *change,
not operation*, and a mesh whose only conversational surface lives on that node would lose the
ability to ask anything while every machine kept running perfectly. **The front door may not be
the single point.**
**It is not an employee** ([ADR 0003](0003-agents-are-persistent-employees.md)). Nobody hires it,
it holds no task queue, it is never drained or reassigned. What it does with work that belongs
somewhere else is **dispatch it** — to node sessions, or to workers — which is what a node session
already does when asked something it does not have.
**It is ADR 0025's reader.** The agent that reads the design repository and answers into search is
this session, not a second one. One agent, one memory, one place to reach; two would both need
that repository and would eventually disagree about what it says.
**"One per node" is about address, not about process count.** ADR 0004's rule — *two and nothing
decides which replies* — forbids ambiguity in who answers when a **node** is addressed. The mesh
session answers when the **mesh** is addressed. The control-plane node therefore hosts two
sessions and no ambiguity, and stating this here is what stops it reading as a contradiction
later.
## Consequences
**The node session's setup must now be designed, and it never was.** This decision is expressed as
*the same as a node session, elsewhere*, which is only meaningful once that mechanism is written
down. The design document covering both is the immediate consequence of this record, not a
follow-up to it.
**A consumer that is not a machine stops being deferrable.** The licence binding above is the gap
`14-model-access.md` names, and it now has two consumers rather than a hypothetical one. Until it
exists, a session's model access can only be expressed as *this module on this machine*, which
cannot say *this node's session uses the personal licence and the mesh's uses the company one* —
the thing the binding is for.
**Symmetry is preserved, and it is worth being precise about why.** ADR 0004's claim is about what
a node can reach, and it is untouched: node-to-node messaging is unchanged, nothing is routed
through the mesh session, and it is a participant rather than a hop. What arrives is a
participant that happens to be the one a person usually addresses.
**Availability degrades to inconvenience rather than to silence** — but only because of the
addition rule above. If that rule is ever relaxed, this consequence inverts, and it inverts
quietly: everything keeps working and nobody can ask about it.
**The surface a person uses is not decided here.** That a board is a good place to talk to it is
likely and is not this record's business; the session is reachable over the broker like everything
else, and what puts a text box in front of it is a separate choice.
## References
- [ADR 0004](0004-a-node-and-how-it-joins.md) — the node session this extends
- [ADR 0025](0025-the-design-record-is-read-not-copied.md) — the reader this session is
- [ADR 0003](0003-agents-are-persistent-employees.md) — the vocabulary this is not
- [`03-DESIGN/01-to-be/14-model-access.md`](../03-DESIGN/01-to-be/14-model-access.md) — *a
consumer that is not a machine*, the gap this makes concrete
@@ -0,0 +1,109 @@
---
topic: what runs on it
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0009-modules-and-the-graph.md
---
# 27. A provision names what the consumer is coupled to, not the role it plays
## Context
Provisions are named after roles. The catalogue and every test fixture built so far say:
```
provides: database
requires: database
```
**Nothing distinguishes one engine from another.** A module requiring `database` is satisfied by
any module providing `database`, so a module written against PostgreSQL can be matched to a
provider of Microsoft SQL Server, resolve as satisfied, deploy, and fail on its first query.
**The mesh runs several engines** — PostgreSQL, Microsoft SQL Server, MariaDB, and others behind
products that expose their own. This is not a hypothetical collision.
**The failure is in the direction that hides.** Resolution *succeeds*. Nothing is refused, nothing
is logged, and the breakage surfaces later as an error inside an application, on a machine, with
nothing connecting it back to a match made elsewhere by something that thought it had done its
job. **A wrong answer delivered confidently costs more than a refusal**, and the whole point of
refusing on ambiguity ([ADR 0009](0009-modules-and-the-graph.md)) was to not do this.
**How it got in:** every test written for the resolver had exactly one provider of each name, so no
mismatch was expressible and none was caught. The fixtures agreed with the design. That is the same
fault as [`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)'s
imagined output and [`019`](../04-ISSUES/019-a-comment-asserting-a-fact-about-a-machine/00-report.md)'s
unchecked comment, at the level of a name rather than a line.
## Considered Options
1. **Keep role names; let the operator assign correctly.** The mesh would refuse ambiguity when two
providers exist, so a person picks. **Rejected.** It makes correctness depend on somebody
knowing that the module they are assigning speaks a particular dialect — which is exactly the
knowledge the provisioning model exists to remove. And with one provider of each name, nothing
is ambiguous and nothing is asked.
2. **A role name plus a `flavour:` or `engine:` qualifier**, matched as a second field.
**Rejected.** Two fields that must agree is a constraint the resolver has to enforce and a
manifest author has to remember, to express something one field already can. The name is the
contract; splitting it invites a requirement that names a role and forgets the qualifier, which
then matches everything again.
3. **The name says what the consumer is coupled to.** **Adopted.**
## Decision
**A provision is named for the thing a consumer's code is written against.**
```
provides: postgres-database
requires: postgres-database
```
**The test is whether the consumer can tell the difference.** If swapping the provider would break
the consumer, the name must say which provider — because a match that breaks the consumer is not a
match. If the consumer genuinely cannot tell, a role name is correct and better.
| provision | | why |
|---|---|---|
| `postgres-database`, `mssql-database` | **specific** | applications are written against a dialect; a swap breaks them |
| `route` | **role** | the consumer wants its name reachable and does not care what proxies it |
| `resolver` | **role** | the consumer wants names to resolve |
| `artifact-store` | **role** | the consumer fetches by digest over a protocol, and nothing else |
**`database` is not a provision and may not be provided.** There is no context in which an
application talks to a generic database: it talks to PostgreSQL or it talks to SQL Server. A name
that cannot be true of any real consumer should not be expressible.
**This is about coupling, not about products.** Two providers of `postgres-database` — a container
on this node and a managed instance elsewhere — are interchangeable and *should* both match. What
may not be interchangeable is what the consumer's queries are written in.
## Consequences
**Every manifest that names a database changes.** Doing this now costs a rename across a handful of
examples. Doing it after modules are migrated costs it across all of them, plus every deployment
that resolved against the old name.
**Wrong requirements now fail loudly, and at the right moment.** A module requiring
`postgres-database` where only `mssql-database` is provided is unsatisfiable, so it is **refused at
resolution** with both names visible — rather than deployed and broken later. This is the property
that was lost, restored.
**Generic role names are still right, and the rule says when.** This does not push specificity
everywhere; it puts it exactly where a consumer is coupled. Naming `route` after a particular proxy
would be the same error in the other direction, and would prevent a swap that genuinely changes
nothing.
**It is checked, not merely stated** ([`00-META/how-we-build.md`](../00-META/how-we-build.md) §5).
A manifest providing a name known to be engine-generic is refused, naming what to say instead.
Without that, this record is a convention, and a convention is what the previous naming was.
## References
- [ADR 0009](0009-modules-and-the-graph.md) — provisions, and refusing on ambiguity
- [`03-DESIGN/01-to-be/07-the-substrate.md`](../03-DESIGN/01-to-be/07-the-substrate.md) — *the
provisioning model uses databases, roles and schemas as PostgreSQL means them*, which is this
record's point made about the substrate before it was made about modules
@@ -1,112 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 27. The product is Novox Mesh; Nox is an identity, not a second system
## Context
The name `HAL` was never chosen. This began as a dotfiles repository, the first commits in
February 2026 adopt dotfiles and per-node overrides, and the name arrived with the code — as
recorded in
[`03-DESIGN/00-as-is/10-module-catalogue.md`](../03-DESIGN/00-as-is/10-module-catalogue.md),
most of the current shape is inherited from that origin rather than designed for a mesh. The
name is part of the inheritance.
Three things make it worth changing rather than living with.
**It is borrowed, and borrowed badly.** HAL is the canonical *untrustworthy* machine
intelligence. For infrastructure whose entire proposition is that it manages your machines,
heals itself, and is trusted with credentials, that is an unhelpful flag to fly, and it is not
a name anyone owns.
**There is a name available that is owned.** The company is Novox. A product of Novox should
carry that lineage rather than a film reference.
**A platform and a persona are different things, and one name was doing both.** `HAL` named the
mesh *and*, implicitly, the thing an operator talks to. Those are separate concerns — the
platform is what runs; the persona is who answers.
## Considered options
1. **Keep `HAL`.** Rejected. Every reason to keep it is sunk cost, and the sunk cost is at its
smallest today.
2. **Rename everything to a single new name covering platform and persona.** Rejected: it
repeats the conflation that made `HAL` ambiguous.
3. **Separate the two: a product name and an identity.** Chosen.
## Decision
**The product is `Novox Mesh`**, shortened to `mesh` in internal use — repository names, the
module namespace, environment variables, paths.
**`Nox` is an identity of Novox**, and specifically an **agent identity within the mesh's own
model** — a named participant, exactly as
[ADR 0012](0012-agents-are-persistent-employees.md) defines one. Not a separate product, not a
separate runtime, not a privileged path.
**Nox is the agent of the mesh, not of a node.** This is the part that carries weight:
- **Every node keeps its own identity.** That already exists and stays — a node is a named
participant with its own character, and addressing one directly remains possible and normal.
- **Nox is scoped to the whole mesh.** It is what the mesh is called when the mesh itself
speaks, rather than one machine within it.
- **Nox addresses node identities.** Asking Nox for something that lives on one node is Nox
talking to that node, not a human choosing a machine.
- **A human mostly talks to Nox.** It is the front door.
That last point makes Nox the concrete form of the vision in
[`00-META/mission.md`](../00-META/mission.md): *an agent states an intent and the mesh carries
it out — no console to open, no runbook to follow, no remembering which node holds which
thing.* Nox is who that intent is stated to. The mission described the behaviour; this names
the thing that has it.
Nox holds no private channel. Whatever it can do, it does through the same surfaces every other
agent uses — which is not a naming detail: a persona with its own path would be the one part of
the mesh with no human checkpoint, and the skeleton already rules that out.
`HAL` is retired.
**Timing is the substance of this decision, not an aside.** The skeleton in
[research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) is not built. Renaming
before it exists costs a search and replace across research documents. Renaming after costs the
same class of migration as everything else this repository is trying to avoid, and would
therefore not happen.
## Consequences
- **The as-is layer keeps `HAL`.** It describes what runs, and what runs is called HAL. The
to-be layer uses `mesh`. The rename is part of the migration, and the two-layer split
([ADR 0020](0020-design-is-written-in-two-layers.md)) is what makes holding both names
coherent rather than confusing.
- **Records 0001–0026 keep `HAL`.** They are immutable and they say what was decided when it
was decided. No record is edited for a name.
- Tier 2 cannot be `mesh-mesh`. The control plane is **`mesh-control`**; `mesh-broker` was
rejected because the substrate already contains a message broker.
- The namespace, environment variable prefix, service names and on-disk paths all change. In
the existing system that is a migration and is not attempted here.
- **`mesh` is a generic word**, and it already means something specific in infrastructure — a
service mesh is a different thing. Recorded as a known trade rather than an oversight: the
full name `Novox Mesh` is distinctive, and the short form is internal.
- The persona has a name before it has behaviour. That is the right order — it is an identity in
a system that already has a model of identities, so it needs no new machinery to exist.
- **Except in one respect, and it is a real gap.**
[ADR 0012](0012-agents-are-persistent-employees.md) binds every agent to a home node, one to
one, with a workspace on that machine. A mesh-scoped agent has no home node by definition, so
the model does not currently have a shape for Nox. Extending it — an agent whose scope is the
mesh rather than a machine — is a decision of its own and is not taken here.
- Two levels of identity now exist where there was one: the node, and the mesh. The distinction
has to stay visible in every surface, or "ask Nox" and "ask a node" collapse into each other
and it stops being clear who is answering.
## References
- [ADR 0012](0012-agents-are-persistent-employees.md) — what an identity is in this system, and
why `Nox` needs no separate mechanism.
- [ADR 0020](0020-design-is-written-in-two-layers.md) — why the as-is and to-be layers can
legitimately use different names for the same system.
- The dotfiles origin, and the naming inheritance it explains:
[`03-DESIGN/00-as-is/10-module-catalogue.md`](../03-DESIGN/00-as-is/10-module-catalogue.md).
-89
View File
@@ -1,89 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 28. HQ is company-scoped; the mesh is its first product
## Context
This repository was `hal-hq` — one product's headquarters, named for the product. Then the
product was renamed ([ADR 0027](0027-the-product-is-novox-mesh.md)), which forced the question
of what the repository is actually the headquarters *of*.
Two facts settled it, and both were checked rather than assumed.
**Novox already delivers other things.** The company's forge organisation holds live projects
beside the mesh, and they are registered as build sources — meaning the mesh already builds and
deploys them. They are not hypothetical future products; they exist and ship today.
**They are tenants, not peers.** They run *on* the mesh. Every one of them is developed,
delivered and hosted by it. So the mesh is not one product among several — it is the ground the
others stand on.
That distinction decides the scope. If the mesh were a product beside others, a per-product HQ
would be right. Because it is the substrate the company operates on, a decision about the mesh
is a decision about how the company works.
## Considered options
1. **`mesh-hq` — one HQ per product.** The safe choice, and the reversible one: a second
product creates its own HQ and shared practice graduates upward later. Rejected, knowingly,
because it models the mesh as a peer of things that are actually its tenants.
2. **A company HQ *and* a product HQ, from the start.** Rejected as ceremony — two repositories
for one operator, and the constitution's own YAGNI rule says not to.
3. **One company-scoped HQ, `novox/hq`, with the mesh as its first product.** Chosen.
## Decision
The repository is **`novox/hq`** — Novox's headquarters, not the mesh's.
It holds the reasoning behind what Novox builds. Today almost all of that is the mesh, because
the mesh is what Novox is building. That is a fact about the present, not a definition of the
repository.
**The scope of each document is fixed now, so the eventual split is mechanical rather than
archaeological:**
| Scope | Documents | Moves if products separate? |
|---|---|---|
| **Company** | [`how-we-build.md`](../00-META/how-we-build.md), [`process/`](../00-META/process/), [`repos.md`](../00-META/repos.md), this record and [0019](0019-hq-is-its-own-repository.md)–[0027](0027-the-product-is-novox-mesh.md) | No — they stay at the top |
| **Product (mesh)** | [`mission.md`](../00-META/mission.md), [`context.md`](../00-META/context.md), [`effect.md`](../00-META/effect.md), `01-RESEARCH`, `03-DESIGN`, `04-ISSUES`, records 0001–0018 | Yes — into a product section |
The folders are **not** restructured now. One product's content under a company name is
correct while there is one product's worth of it, and nesting before there is anything to nest
is the ceremony option 2 was rejected for.
## Consequences
- Engineering practice has a home that does not belong to the mesh. `how-we-build.md` — never
write to production directly, migrations for schema changes, runtime evidence for behavioural
criteria — is true of any Novox project, and its being in a mesh repository was always a
slight mislabelling.
- The constitution derived from it ([ADR 0025](0025-hq-is-the-source-of-the-constitution.md))
can legitimately govern work outside the mesh. Under a product HQ it could not have, without
either duplicating or reaching across repositories.
- **The bet is not entirely forward-looking, and that is worth being honest about.** Novox
already has work that is *not* a mesh tenant — client engagements and at least one product
that is developed outside it. So the company genuinely has a scope wider than the mesh
**today**, which strengthens the case for a company HQ and simultaneously means the split in
the table above is closer than "some day". The table is not a precaution; it is a plan whose
trigger already half-exists.
- What has *not* happened yet is any of that work needing the constitution. That is the actual
trigger ([ADR 0025](0025-hq-is-the-source-of-the-constitution.md)): the moment something
outside the mesh must be governed by the same rules, product-level content moves down a level
and this repository becomes what its name already claims.
- A new repository was created rather than the old one transferred, because the forge predates
the transfer API. The original was verified to contain nothing the new one lacks — every ref
an ancestor, no tags, issues, pull requests, releases or wiki content — and then removed.
- The mesh's own documents now live one conceptual level below the repository they are in. A
reader arriving at `01-RESEARCH` should understand it as the mesh's research, not Novox's.
Nothing in the folder names says so, and that is the cost of not restructuring.
## References
- [ADR 0027](0027-the-product-is-novox-mesh.md) — the product name that forced the question.
- [ADR 0019](0019-hq-is-its-own-repository.md) — why HQ is a repository at all. Unchanged; only
its scope moves.
@@ -0,0 +1,118 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0006-the-substrate-and-the-control-plane.md
---
# 28. The substrate supplies the control plane and nothing else
*Corrects one row of [ADR 0006](0006-the-substrate-and-the-control-plane.md) and makes explicit
something it left unsaid. The rest of that record stands.*
## Context
ADR 0006 defines the substrate by a circularity: **what the control plane needs in order to run,
and cannot ask itself for, because it is not running yet.** Two questions, and both must be
answered *yes* for something to be substrate.
Its membership table admits the object store on this line:
| role | product | |
|---|---|---|
| object store | **MinIO** | it cannot grant itself a bucket |
**That answers the second question and assumes the first.** It is true that a control plane cannot
grant itself a bucket. Nothing establishes that it needs one.
**It does not.** Verified 2026-08-31 against `mesh-control`: no S3 client, no bucket, no object
storage of any kind outside comments. Artifacts reach nodes as content-addressed blobs in the OCI
registry, and the code records the decision and its reasoning:
> One store, and it is the registry the bootstrap already pulls from. An OCI registry is a
> content-addressed blob store that happens to also understand images… The alternative considered
> was a second store beside it — S3-shaped, buckets, signed URLs. It is the right answer for
> objects that are *mutable*, or need per-reader access, or are not build output. None of that
> describes a digest-pinned archive, and standing up a second service to hold one kind of
> immutable blob means two things to run, two things to back up and two ways for an artifact to be
> missing.
**The row is inherited from the system being replaced**, where an object store distributed module
tarballs. Here nothing does, and the row was never re-tested against the definition it sits under.
**A second thing ADR 0006 never says:** whether a substrate service and a module of the same
product are the same instance. It says the substrate is *not the control plane* and *not a place
for logic*, and stops. The question is not idle — an application wanting a database, on a mesh
whose substrate is already running PostgreSQL, has an obvious wrong answer available.
## Considered Options
1. **Leave the object store as substrate, unused.** Harmless-looking. **Rejected.** A membership
list that includes something nothing needs is a list that has stopped being derived from its
test, and the next member is admitted by precedent instead of argument. It also mandates that
every mesh run a service no mesh uses.
2. **Applications share the substrate's instances.** One PostgreSQL, one of everything.
**Rejected**, below.
3. **The substrate is exactly what the control plane consumes; everything else is a module.**
**Adopted.**
## Decision
**The object store is not substrate.** It fails the first half of the test: the control plane does
not need one. An object store is an ordinary module, required through the module graph like
anything else, and a module wanting one depends on a module providing one.
**The substrate has four members, not five**: a relational store, a message bus, an image registry,
and conditionally an identity provider. The registry stays — the control plane genuinely cannot
deliver an artifact without somewhere to put it.
**A substrate service and a module of the same product are different instances, and are not
shared.** The mesh's own PostgreSQL and a PostgreSQL a workload was given are two servers, two
containers, two lifecycles.
Three reasons, and the first is the one that matters:
**The substrate is not in the module graph.** It is raised from the pinned bundle the host carries,
before any mesh exists to declare it. A workload depending on it would depend on something the
graph cannot see, cannot rotate a credential for, and cannot move — which is every property the
provisioning model exists to provide.
**It would put workload data in the control plane's own store.** The mesh's contexts own their
stores exclusively ([ADR 0008](0008-a-context-owns-its-store.md)). An application sharing that
server can exhaust it, lock it, or fill its disk, and the failure is the control plane going down
— which is the one failure that makes every other one harder to fix.
**They are bounded differently.** The substrate is sized, backed up and upgraded as part of
bootstrapping a mesh. A workload's database follows the workload — moved with it, destroyed with
it, restored with it.
## Consequences
**Migrating an object store is ordinary module work**, not substrate work. It was previously going
to be done as part of completing the substrate, which would have been the wrong shape and would
have coupled every mesh to a service the mesh does not use.
**A mesh with no workload needing one runs no object store at all.** That is the correct outcome
and was not previously available.
**Two PostgreSQL containers on a node that hosts both is expected**, not duplication to be
optimised away. Anyone tidying them together should find this record first.
**"Substrate by role and ordinary by delivery" loses one of its two members.** ADR 0006 uses that
phrase of the object store and the registry — things that are substrate but provisioned once a
control plane exists. It now describes the registry alone.
**The definition is applied, not just stated.** Both halves of the circularity test are asked of
each member, and *cannot grant itself one* is not sufficient on its own — it is true of almost any
service, which is what made it possible to admit a member on that half alone.
## References
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — the definition, and the table this
corrects one row of
- [ADR 0008](0008-a-context-owns-its-store.md) — a context owns its store exclusively
- `mesh-control internal/builder/registry.go` — where artifacts go, and why not S3
@@ -0,0 +1,112 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0005-the-node-host.md
---
# 29. A network is a shape, because an action cannot be undone
## Context
**A module of several containers has no way to let them reach each other by name.** A container
declaration carries a `network` field, and it only ever *joins* one that already exists — it was
added so the control plane could reach the store and the broker on the machine it was raised on.
Nothing in the vocabulary **creates** a network.
Without one, containers on a machine share the runtime's default bridge, which gives addresses and
no name resolution between them. So a module that is several containers can only be written by
publishing ports onto the machine and pointing its own parts at the host — which puts a module's
private wiring on the machine's own address space, where anything else on the machine can reach it
and any other module can collide with it.
**This is the gap a mail system meets and nothing else so far does**
([`00-work-breakdown.md`](../03-DESIGN/01-to-be/00-work-breakdown.md) 3.3). It is being taken now
rather than then, because 3.3 is the task most likely to send work back into the declaration
language and the least useful place to discover it.
**Adding a shape is not a small change, and the host says so** — the vocabulary is asserted
against a stated number, with the reason written into the failure: *every addition widens what a
compromised control plane can express, so a change here is a decision.* The host applies what it
is told; the only bound on a hostile control plane is what the language can say
([ADR 0004](0004-a-node-and-how-it-joins.md)).
## Considered Options
1. **An `action` that creates the network.** The vocabulary already has one, the bundle already
uses seven of them, and `docker network create` with a `verify` is exactly the shape an action
takes. Nothing would need adding. **Rejected**, on removal:
> An action has no footprint the host can undo — it ran, and whatever it did belongs to
> whatever it acted on.
A network made this way **leaks when the module is unassigned**, and the mesh cannot tell: the
record says an action ran, and there is nothing to reverse. Unassigning a module would leave a
network behind on every machine it was ever on, and the only way to find them would be to go
and look. *A resource the mesh can create and never clean up is one it should not create.*
There is a second reason, and it is the one that generalises: an action is opaque. **The mesh
cannot tell what an action did**, so a network created by one is not a thing the mesh knows
about — it cannot be reported, counted, or reasoned about, and a module could not require one.
2. **Publish ports on the machine instead.** No new shape, and it works today. **Rejected.** It
makes a module's internal wiring part of the machine's address space: two modules that each
want a database on a fixed port collide, and anything else on the machine can reach what was
meant to be private. It also makes the module's manifest depend on what else is installed,
which is the thing provisioning exists to remove.
3. **`network` as a ninth shape.** **Adopted.**
## Decision
**`network` joins the vocabulary, and the vocabulary is nine shapes.**
```
{"id": "internal", "type": "network", "name": "mail"}
```
**A name and nothing else.** Not a driver, a subnet, an address range or a gateway: every one of
those is a thing a module would have to know about the machine it lands on, and a module that
names a subnet is a module that collides with whatever else chose the same one. The runtime picks;
the mesh names.
**It is created if absent and removed when no longer declared** — an ordinary shape, with the same
lifecycle as a directory. That is the whole reason it is a shape.
**Declared before the containers that join it.** Resources are applied in the order the module
wrote them, and orphans are removed in **reverse** — so a network written first is created first
and removed last, after the containers attached to it are gone. This is not a new rule; it is the
existing one, and it happens to be exactly right here. A network written *after* its containers
would fail to remove while they still hold it, and that failure is reported rather than silent.
**What it does not do:** it does not reach across machines. A network is one machine's, like
everything else the host applies. Modules on different machines reach each other over the private
network the mesh already provides ([ADR 0007](0007-connectivity.md)), and a shape that tried to
span machines would be a second overlay with a worse contract.
## Consequences
**The vocabulary is nine, and the count moves with a record.** The test that asserts it names this
one, so the next person to change it finds the argument rather than a number to edit.
**A compromised control plane can now create and destroy networks on a machine.** Stated plainly
because that is the cost, and the bound is the point: it can create a named network and remove
one, and it can do neither to anything it did not declare. It cannot inspect, attach to, or
reroute what is already there — those would be different shapes, and are not being added.
**A multi-container module becomes expressible**, which unblocks 3.3 and, less obviously, makes
several smaller modules simpler: anything that is a service plus a sidecar currently has to
publish a port to talk to itself.
**Nothing is required to use it.** A module of one container declares no network and joins none,
exactly as now. The substrate keeps using `host`, which is a runtime-provided network and not one
the mesh creates.
## References
- [ADR 0005](0005-the-node-host.md) — the host's vocabulary, and why each shape is a decision
- [ADR 0004](0004-a-node-and-how-it-joins.md) — what may be pushed is bounded by form, not by trust
- [`03-DESIGN/01-to-be/00-work-breakdown.md`](../03-DESIGN/01-to-be/00-work-breakdown.md) — 1.3,
and the mail system at 3.3 that this is for
@@ -1,101 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
extends: 0016-a-lab-node-is-a-virtual-machine.md
---
# 29. The lab's first scenario has no pipeline, and the lab comes first
## Context
[The lab design](../03-DESIGN/01-to-be/01-end-to-end-testing.md) opens with *"what is under
test is a module; the mesh is the harness"*, and everything follows from that: a scenario has
its own forge, its own coordinator, and its own delivery cascade ending in verify. The verdict
*is* a pipeline result.
That is the right design for testing a module against the mesh that exists. It is unusable for
the thing now being built.
**The new mesh has no coordinator.** Tier 0 is a host binary and tier 1 is a pinned bundle
([research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md)). A scenario that
requires a forge, a coordinator, a cascade and a meshware daemon cannot exercise them, because
all four are tier 2 and do not exist yet.
**And the sequencing was written backwards.** [Research 009](../01-RESEARCH/009-migration/00-overview.md)
placed the lab at phase B, as verification of tiers already built. But tier 0 is the component
that takes over a machine's packages, services and network — it cannot be developed against a
machine anyone needs. It needs somewhere disposable to exist **before** it is written, not
after.
## Considered options
1. **Develop tiers 0 and 1 against a real machine; add the lab afterwards.** Rejected twice
over. Developing something that reformats a machine, against a machine that is in use, is
how a machine is lost. And it would leave the bootstrap path exercised only when performed
for real — which is precisely the property that makes the current first-node script the
least-tested code in the system.
2. **Build the full lab first.** Impossible, not merely unwise: the full scenario needs a
coordinator, a forge and a delivery cascade, all of which are tier 2. It cannot precede the
tiers it is meant to test.
3. **Two scenario classes, the smaller one first, the larger a superset.** Chosen.
## Decision
The lab has **two scenario classes**, and the first has no pipeline in it at all.
| | **Bootstrap scenario** | **Full scenario** |
|---|---|---|
| Contains | one or more virtual machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules |
| Verdict from | what the host reports about the state it reconciled | a pipeline result ending in verify |
| Exercises | tiers 0 and 1 | tiers 2 and above, and modules |
| Exists to | develop the mesh | test what runs on it |
The bootstrap scenario is a **strict subset** of the full one — the same virtualisation, the
same networking, the same scenario lifecycle, simply stopping before a control plane exists.
Nothing forks, which is the same rule the existing design already holds itself to.
**The lab is built first**, ahead of tier 0, and [research 009](../01-RESEARCH/009-migration/00-overview.md)
is resequenced accordingly. It is the environment everything else is developed inside.
Of the runner's two candidate jobs, this settles their order: **scenario lifecycle is needed
immediately** — something must materialise, snapshot and destroy a mesh before anything else
can be written. **Assertion execution comes later**, with the full scenario, because a
bootstrap scenario's assertions are about the state a single host reconciled and are small
enough to state directly.
## Consequences
- **The hardest path to test becomes the one exercised most.** Raising a node from nothing is
currently a script that runs when a node is created and is otherwise never touched. Under
this decision it is the inner development loop for every change to tiers 0 and 1.
- The first thing built is small: virtualisation, a network, a way to place a binary, and a way
to snapshot and reset. No forge, no coordinator, no pipeline, no modules.
- The full scenario becomes reachable by *addition* rather than by rework, because it differs
only in what is placed inside the machines.
- The lab acquires a second audience. It was designed for a module author and now also serves
whoever is building the mesh itself — which is the same "one runner, two callers" argument
the design already makes, extended one step.
- **A stale claim in the design is corrected.** It argues that scenarios are *"affordable with
system containers and would not be with virtual machines — the unit choice is what makes the
gate possible at all."* [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) superseded that:
a lab node is a virtual machine, and the scale argument for system containers was found to
have been invented rather than required. The design text did not follow the decision. It does
now.
- The lab's home is `novox/mesh-lab`, recorded in
[ADR 0030](0030-the-repository-structure.md) — written after this record, because this one
needed a repository that no decision had yet named.
- The bootstrap scenario's fidelity is its whole value, and also its risk: if it diverges from
how a real node is raised, it certifies something that does not happen. That is the same
hazard the existing design names for the full scenario, and the same answer applies —
nothing new drives it, and what runs is the real thing.
## References
- [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) — a lab node is a virtual machine.
- [Research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) — the tiers, and the
observation this rests on: a scenario needing only tiers 0 and 1 is one machine and a pinned
bundle, which is also exactly the bootstrap path.
- [Research 009](../01-RESEARCH/009-migration/00-overview.md) — the migration sequence this
reorders.
@@ -0,0 +1,96 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0005-the-node-host.md
---
# 30. Data outlives the mesh that declared it
## Context
**The conversion runs on live services holding real data**, and starts on the node that holds all
of it. Identity, mail, everything. The requirement stated plainly: a data directory may be
*moved*, and may never be *lost*.
**The host deleted them.** A directory that stopped being declared was an orphan, and an orphan
directory was removed with `os.RemoveAll` — everything under it — while the report said
`removed`. A module unassigned took its database's files with it, and nothing anywhere said what
had been in there.
Reproduced before it was fixed: assign a module, let a service write into its directory, unassign
the module, and the file is gone.
**A directory stops being declared for ordinary reasons**, which is what makes this sharp rather
than theoretical. A module unassigned from a node. A manifest edited to move a data folder — the
exact operation the conversion needs. A resource renamed. A typo. **Every one of those is a normal
day's work, and every one of them was destructive.**
**The removal order was already right, and that is what makes a fix possible.** Everything the
mesh puts inside a directory is itself a declared resource, and orphans are removed in reverse
declaration order — so by the time a directory is reached, what the mesh wrote there is already
gone. Anything still present was put there by something else.
## Considered Options
1. **A `keep` flag on the directory.** A module declares which of its directories hold data, and
the host leaves those. **Rejected.** It is safe only when somebody remembered, and the failure
of forgetting is total and silent. A rule that protects data only when it was asked to is not
a rule about data, it is a rule about attentiveness — and this is the one place in the system
where being wrong does not recover.
2. **Never remove a directory.** Simple and unarguably safe. **Rejected**, narrowly: every module
ever assigned would leave its directories behind for ever, and a machine that accumulates
things nobody can account for is one where nobody can tell what is still in use. The clean-up
that is genuinely the mesh's is worth keeping.
3. **Remove a directory only when it is empty.** **Adopted.**
## Decision
**A directory that still holds anything is kept, and the mesh says so.** An empty one is removed.
**This is the host's existing line applied to the one shape where getting it wrong is
unrecoverable** — *it removes what it made and leaves what it merely configured*
([ADR 0005](0005-the-node-host.md)). An empty directory is what the host made. A full one is not.
**No flag, no declaration, nothing to remember.** Emptiness is the test, and it is derived from
the removal order rather than asserted: the mesh's own contents are gone by then, so what remains
is by definition something nobody declared.
**It is reported, not silent.** The outcome is `kept`, naming how many items are inside and saying
they are for a person to deal with. A directory quietly left behind is how a machine accumulates
things nobody can account for — which is the objection to option 2, and it is answered by saying
so rather than by deleting.
**Files are unchanged.** A declared file is the mesh's own — it wrote it, it owns it, and losing a
configuration file is not the failure this is about. The distinction is deliberate: **directories
hold what other things produced; files are what the mesh itself put there.**
## Consequences
**Moving a data directory is now safe by default.** The manifest changes, the old path stops being
declared, and the data stays where it is until somebody has looked at it. That was the operation
most likely to destroy something during the conversion, and it is now the operation that does the
least.
**Unassigning a module leaves its data.** Correct, and it means unassignment is no longer a way to
clean up — removing data is a person's act, done knowingly. Given what unassignment did before,
that is the trade being made and it is the right way round.
**A machine can accumulate directories nobody removed.** Accepted, and mitigated by saying so
every time rather than by a periodic sweep. A sweep would be the deletion this record exists to
prevent, on a timer, with nobody watching.
**It is not a backup, and must not be mistaken for one.** This stops the mesh destroying data. It
does nothing about a disk, a mistaken `rm`, or a service corrupting its own store. The conversion
still needs backups taken and **restored** before anything is moved — a backup nobody has restored
is a belief, not a copy.
## References
- [ADR 0005](0005-the-node-host.md) — the host removes what it made
- [`03-DESIGN/01-to-be/00-work-breakdown.md`](../03-DESIGN/01-to-be/00-work-breakdown.md) — the
conversion this was found by planning
@@ -1,99 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 30. The repository structure, and the rule that names them
## Context
The tiers are settled ([research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md))
and the product is named ([ADR 0027](0027-the-product-is-novox-mesh.md)), but the repositories
themselves were only ever sketched in research. Two consequences had already appeared.
[ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) makes the lab phase 0 of the entire
migration and **could not say where it lives**, because no record named a repository.
And the research contradicted an accepted record: it listed `mesh-hq` for this repository, while
[ADR 0028](0028-hq-is-company-scoped.md) had decided `novox/hq` and explicitly rejected that
name. A design resting on research is a design resting on something that can change without a
decision.
There is also an implied naming rule that has never been written down. ADR 0027 says
repository names take `mesh`; ADR 0028 gives this repository no prefix at all. Both are right,
for a reason neither states.
## Considered options
Only the naming rule had genuine alternatives; the tier repositories follow from the tiers.
1. **No prefix — `novox/host`, `novox/control`.** The organisation already says Novox, so the
prefix reads as stutter. Rejected once it was established that Novox delivers more than the
mesh: with several products the prefix is not stutter, it is the product namespace doing
real work, and the forge has no nested groups to do it instead.
2. **An organisation per product — `novox-mesh/host`.** Puts the product boundary where the
forge's only real grouping primitive lives, so permissions and teams attach to it. Rejected
for now as premature: no per-product access boundary exists yet, and it costs `novox-`
repeated across every organisation.
3. **Product-prefixed repositories in the company organisation.** Chosen.
## Decision
**The naming rule:** a repository that belongs to a product carries that product's prefix. A
repository that is company-scoped does not.
That is why this one is `hq` and the mesh's are `mesh-*`. Both records were already correct;
the rule connecting them is stated here.
**The repositories:**
| Repository | Tier | Holds |
|---|---|---|
| `novox/mesh-host` | 0 | the node host — the one binary installed by hand |
| `novox/mesh-substrate` | 1 | the four pinned services, as declarations |
| `novox/mesh-control` | 2 | the control plane and its contexts |
| `novox/mesh-surfaces` | 3 | tools, web, cli — thin, no logic |
| `novox/mesh-sdk` | — | contracts shared across tiers: types, not behaviour |
| `novox/mesh-lab` | — | the lab: scenario lifecycle, networking, placement |
| `novox/hq` | — | this repository. Company-scoped ([ADR 0028](0028-hq-is-company-scoped.md)) |
**The lab is its own repository.** Its lifecycle differs from everything else in the list: it
is never shipped to a node, it outlives any single tier, and it drives virtualisation on a
workstation — which nothing else in the mesh does. Putting it inside the host would couple
development tooling to a shipped component; putting it inside the control plane would make the
bootstrap scenario depend on a tier that does not exist when it is needed.
**Tier 4 is deliberately not decided here.** Whether the catalogue is one repository, one per
domain, or one per application remains open from
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) and is blocked on
[research 005](../01-RESEARCH/005-domain-grouping/00-overview.md): how many repositories hold
domains cannot be answered before what the domains are. Recording the gap is the point —
`mesh-catalog` appears in the research sketch and is **not** decided by this record.
## Consequences
- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) can name its target. Phase 0 has
a home, which was the immediate blocker.
- The research sketch stops being load-bearing. It remains what it is — a sketch — and the
design layer can now cite a record instead.
- **Seven repositories where there is currently one**, for a mesh that today lives in a single
monorepo. That is the cost, and it is not small: seven release cadences, seven sets of
dependencies, and cross-repository changes that were previously one commit. The offsetting
argument is the tier rule — a boundary that only points downward is enforceable across
repositories and merely conventional inside one.
- The prefix will read as redundant for as long as the mesh is the only product with
repositories. That is accepted deliberately: the alternative is renaming everything at the
moment a second product appears, which is the class of migration this project is trying to
stop performing.
- Nothing is created yet. This records what the repositories *are*; creating them is part of
phase 0 and after.
## References
- [ADR 0027](0027-the-product-is-novox-mesh.md) — the product name the prefix comes from.
- [ADR 0028](0028-hq-is-company-scoped.md) — why this repository has no prefix.
- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the lab, and why it is first.
- [Research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) — the tiers, and the
sketch this supersedes as a source.
@@ -0,0 +1,72 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0006-the-substrate-and-the-control-plane.md
---
# 31. The control plane authenticates nobody, so identity is a module
## Context
[ADR 0006](0006-the-substrate-and-the-control-plane.md) left one member of the substrate
conditional, and said exactly why:
| role | product | |
|---|---|---|
| identity provider | — | **conditional**: substrate only if the control plane delegates authentication, which is undecided |
[`07-the-substrate.md`](../03-DESIGN/01-to-be/07-the-substrate.md) carried it as an open question —
*whether identity is the fifth* — noting it followed from a decision nobody had taken.
**The decision is taken: the control plane does not delegate authentication.** There is no mesh
identity provider.
**Nothing in the mesh's own machinery ever needed one.** A node proves itself with a keypair it
generated, over a broker account issued at enrolment
([ADR 0004](0004-a-node-and-how-it-joins.md)). Declarations are verified by signature. None of
that touches an identity provider, and the conditional was never about machines — it was only ever
about whether a *person* signing in to a mesh surface would be authenticated by something else.
## Decision
**Identity is a module**, like the mail system and the forge. It runs *on* the mesh, not *of* it
([ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)) — a provider other modules require,
which is the ordinary shape and needs nothing new to express.
**So the substrate is three, and no longer conditional**: a relational store, a message bus, and
an image registry. Together with
[ADR 0028](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md), which removed the
object store, the list is settled and every member is there for the same reason — the control
plane needs it and cannot ask itself for it.
**A mesh that wants no identity provider runs none.** That is now expressible, and was not while
it sat in the substrate as a maybe.
## Consequences
**The last open question about substrate membership is closed.** Both halves of ADR 0006's test
now have an answer for every candidate, and the answer for identity is *the control plane does not
need it*.
**It does not settle how a person signs in to a mesh surface**, and that is deliberately left
open. What is settled is that whatever answers it is not part of what must exist before the mesh
does — so it can be decided late, changed, or replaced, which is precisely what being substrate
would have prevented.
**It becomes a real test of the module graph.** An identity provider is a module that *other
modules require* — the object store already consumes it — so it exercises the provider chain more
seriously than anything ported so far, where the provider was written alongside its consumer.
**Ordering follows from it rather than from preference.** Anything requiring identity has to move
after it, which is a dependency the graph can state rather than something a person has to
remember.
## References
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — the conditional this closes
- [ADR 0028](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md) — the other member
removed, and the test applied properly
- [ADR 0004](0004-a-node-and-how-it-joins.md) — how a node proves itself, which needs none of this
@@ -1,87 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 31. The lab provides the underlay; the mesh builds the overlay
## Context
[Research 004](../01-RESEARCH/004-lab-network/analysis.md) worked out the topology a lab has to
reproduce: a routable segment using documentation addresses, a household segment behind NAT, a
router that forwards exactly one port so a *published-but-NATed* node is real, and a machine
that can attach to either segment or detach entirely.
It also records what makes that topology **mean** something, and this is where a boundary has
to be drawn. Hub election is by convention rather than by flag — the hub is the node whose
profile is server and whose overlay address begins `10.10.0.1`. Direct peering depends on two
nodes sharing a site. Names resolve from mesh configuration on each node.
Those are all facts the *mesh* establishes. The question is whether a scenario declares them.
It is tempting to say yes, because a scenario that hands you a working overlay is a scenario
you can start testing against immediately.
## Considered options
1. **The lab configures the overlay too** — assign the overlay addresses, elect the hub, write
the peer configuration, seed the names. Rejected, and the reason is the whole point of the
lab: **a lab that builds the overlay certifies its own work.** If the mesh's peering logic
is broken, a scenario that pre-built the peering still comes up green. The most valuable
thing the lab can test is precisely the part this would replace.
2. **The lab provides nothing but bare machines** — no addressing, no segments, no NAT. Also
rejected. Then the scenario cannot reproduce *published but behind NAT*, which research 004
identifies as the case that only exists in production today, and the lab loses its reason to
use virtual machines at all.
3. **The lab provides the underlay; the mesh builds the overlay.** Chosen.
## Decision
**A scenario declares the underlay** — the facts a machine would have before any of our
software touched it:
- which segments exist, and their address ranges
- which machine sits on which segment, at which address
- what NAT sits between them, and which ports are forwarded through it
- which machines are detached, and can be attached or detached during a run
**A scenario declares nothing about the overlay** — no overlay addresses, no hub, no peering,
no names, no certificates. Those are the mesh's job, and a scenario that supplied them would be
testing itself.
The rule stated in one line: **a scenario provides what a hosting provider and a home router
would provide, and nothing our software is responsible for.**
## Consequences
- **The overlay becomes a thing under test rather than a fixture.** Whether peers form,
whether the hub is elected, whether a NATed node's endpoint is learned — all of it is
observed rather than arranged. That is the class of fault research 004 says is discoverable
only in production today.
- The lab stays small, and stays honest. It needs to know about virtualisation, bridges,
addresses and NAT. It never needs to know what a mesh node is.
- A scenario cannot assert "the overlay came up" as a precondition, because it is an outcome.
A bootstrap scenario that wants a working overlay has to wait for one and check, which is
the correct shape.
- **The address ranges are load-bearing, not cosmetic.** The routable segment uses RFC 5737
documentation space specifically because the mesh's own code decides *public versus private*
by matching the address — a private range there makes the hub test as unreachable, and the
mesh silently never forms. Research 004 calls this the single most important fact in the
document, and the declaration format has to make getting it wrong hard.
- The router is a machine the lab materialises without being asked, because NAT requires
somewhere to run. That is an implicit machine in an otherwise explicit declaration, and it is
worth knowing about rather than discovering.
- **Host capability profiles are detected, not declared** — a consequence of
[research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md), and consistent here: a
scenario does not say what a machine is allowed to do, it provides a machine. Which leaves an
open question: a lab machine is always privileged, so the `user` and `edge` profiles have no
scenario that exercises them yet.
## References
- [Research 004](../01-RESEARCH/004-lab-network/analysis.md) — the topology, the documentation
ranges, and the hub-election and peering conventions this deliberately does not touch.
- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the two scenario classes this
declaration has to serve without forking.
@@ -1,71 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 32. A scenario is an isolated address space, and the lab never reaches into it over IP
## Context
A scenario declares literal addresses —
[the declaration](../03-DESIGN/01-to-be/02-scenario-declaration.md) is full of them, and it has
to be, because reproducing *published but behind NAT* means saying which address the world sees.
That raises a question the declaration left open: **two scenarios at once.** Several agents
working means several scenarios, and the lab design already calls that a requirement. But two
scenarios built from the same declaration want the same addresses, and there are only three
documentation ranges in existence.
## Considered options
1. **Allocate addresses from a pool at raise time**, rewriting the declaration's literals.
Rejected. It makes the addresses in a declaration a fiction, so a scenario reproducing a
specific topology no longer reproduces it; it breaks the RFC-range validation, since
allocated addresses would have to come from somewhere real; and the numbers a person reads
in the file stop being the numbers they will see in a capture.
2. **One scenario at a time.** Rejected — it is the requirement, not an inconvenience. A gate
an agent has to queue for is a gate that gets bypassed.
3. **Give each scenario its own network stack, so the addresses do not collide.** Chosen.
## Decision
**A scenario is a closed address space.** Every segment materialises as its own isolated link,
belonging to one scenario instance. Two scenarios raised from the same declaration hold the same
addresses and never meet, because nothing joins their links.
The declaration therefore keeps its literal addresses, and they mean exactly what they say.
**The consequence that constrains everything else: the lab never reaches into a scenario over
IP.** It talks to a machine through the virtualisation layer's own channel — the same way one
executes a command in a container without the container being routable.
That is not a preference. If the lab reached machines by address, the workstation running it
would need a route into each scenario, and two scenarios carrying the same prefix would give it
two routes to the same destination. Concurrency would be impossible, and it would fail in the
worst available way: not with an error, but by one scenario's traffic arriving in another.
## Consequences
- Scenarios are concurrent by construction, with no allocation, no bookkeeping and no limit
beyond the machine's capacity.
- The three documentation ranges stop being a scarce resource. Every scenario may use all of
them, because no two scenarios share a link.
- **The lab cannot use IP to check anything**, which is more of a constraint than it first
appears: *"can this machine reach that one"* has to be asked **from inside the scenario**, by
executing on a machine, rather than probed from outside. That is the honest way to ask it
anyway — reachability from the workstation is not the question.
- A scenario is a unit that can be paused, snapshotted and destroyed whole, because nothing
outside holds a reference into it.
- The lab needs a scenario **instance** identity distinct from the scenario name in the
declaration: the declaration is a kind, and several instances of one kind may exist.
- **The workstation is not on the scenario's network, so it is not a node in it.** Anything a
developer wants to reach — a web interface, a database — needs an explicit, deliberate
forward out of the scenario, which is a feature rather than a gap: nothing leaks by default.
## References
- [ADR 0031](0031-the-lab-provides-the-underlay.md) — the declaration whose literal addresses
this preserves.
- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the lifecycle jobs this shapes.
@@ -0,0 +1,79 @@
---
topic: how we work
status: superseded
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0031-the-control-plane-authenticates-nobody.md
superseded-by: 02-DECISIONS/0034-the-local-account-owns-the-mesh.md
---
# 32. The local account owns the mesh; a surface delegates to a module
## Context
[ADR 0031](0031-the-control-plane-authenticates-nobody.md) settled that the control plane
authenticates nobody, and deliberately left one thing open: **how a person signing in to a mesh
surface is authenticated.** This answers it, and answers a question 0031 did not ask — *who owns
the mesh at all.*
**There was no answer, and the absence was invisible** because every operation so far has been run
by the person sitting at the machine. Nothing had to say whether that was the design or the
circumstance.
## Decision
**The account that installed the host owns the mesh on that node.** Authority is a local login,
and there is nothing else to hold.
**No mesh user model.** No accounts, no roles, no grants, nothing to administer. A person with a
shell on a node can do anything the mesh can do there, because that is already true and pretending
otherwise would be a boundary that does not exist.
**This follows from what was already decided rather than adding to it.**
[ADR 0004](0004-a-node-and-how-it-joins.md) says there is no authorisation between nodes — every
node is the operator's own, so a message from one is a message from them, and *the mesh boundary
is therefore the security boundary*. A user model inside that boundary would guard nothing: anyone
who could be stopped by it could equally read the node's key off the disk.
**The board is different, and the difference is the network.** A surface reachable by a browser
has to know who is asking, because the people reaching it are not, by construction, people with a
shell on the machine. **So the board delegates to an OAuth provider** — which is a module.
## What this does not change
**The identity provider is still not substrate** (ADR 0031). A *surface* delegating
authentication is not *the control plane* delegating it. The control plane runs, applies
declarations and reaches nodes with no identity provider in existence; only the board needs one,
and only to decide whose browser it is talking to.
The test is unchanged and still answers no: *does the control plane need it in order to run?*
## Consequences
**The board depends on a module, and says so.** An ordinary edge in the graph, which means the
board cannot come up before the provider it authenticates against — stated as a dependency rather
than discovered as an outage.
**Moving the identity provider takes the board with it.** During that module's own conversion the
board is unavailable, and that is acceptable: it is a surface, nothing depends on it, and a brief
interruption is the trade already accepted everywhere else. Nothing that keeps a service serving
goes through it.
**Anyone with a shell on a node has full authority there.** Written down rather than left implied,
because it is the sentence that decides who gets an account on a machine. The protection is the
machine's own login, and the overlay that keeps the machine unreachable from outside
([ADR 0007](0007-connectivity.md)).
**A node cannot be operated by somebody without a login on it.** Deliberate, and the cost of
having no user model: there is no way to give a person authority over one node without giving them
a shell there. If that is ever wanted, it is a new decision and not a gap in this one.
## References
- [ADR 0031](0031-the-control-plane-authenticates-nobody.md) — the control plane authenticates
nobody; this answers what it left open
- [ADR 0004](0004-a-node-and-how-it-joins.md) — no authorisation between nodes, and why the mesh
boundary is the security boundary
- [`03-DESIGN/01-to-be/11-a-board.md`](../03-DESIGN/01-to-be/11-a-board.md) — the surface this is
about
@@ -1,81 +0,0 @@
---
status: accepted
date: 2026-08-24
deciders: jochen
reconstructed: false
extends: 0016-a-lab-node-is-a-virtual-machine.md
---
# 33. A router is scenery, not a node — so it is a container
## Context
[ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) settles that **a lab node is a virtual
machine**, and its reasoning is fidelity: a node boots a stock image and runs the real install,
so it has to be a real machine or the thing under test is not the thing that ships.
A scenario also needs routers. NAT, port forwarding, policy between segments and mapping
expiry are all things a router does, and until one is materialised a multi-segment scenario
raises isolated islands
([03-DESIGN/01-to-be/02-scenario-declaration.md](../03-DESIGN/01-to-be/02-scenario-declaration.md)).
The declaration already implies them: a gateway is *the one implicit machine in an otherwise
explicit declaration*.
The question is whether ADR 0016 binds those too.
## Considered options
1. **A router is a node, so it is a virtual machine.** Consistent, and pays for a consistency
nobody needs. A router boots in roughly ten seconds against a container's one; a
six-segment scenario wanting three routers spends thirty seconds per raise on scenery.
2. **The hypervisor provides NAT** — bridges with translation switched on, and its own
forwarding primitives. Rejected on a stronger ground than speed: it makes the *lab* provide
what the declaration is supposed to own, and it cannot express a mapping that expires, a
gateway that refuses to forward, or policy between siblings. The model would shrink to fit
the tool.
3. **A router is scenery, and scenery is a container.** Chosen.
## Decision
**ADR 0016 binds nodes. A router is not a node.**
Nothing under test runs on a router. It is not a participant, it holds no identity, the mesh
never installs anything on it, and no assertion is ever made about its internals. It exists so
that packets between machines behave the way they behave in the world — which is the definition
of scenery.
So a router is a **system container**, and the fidelity argument does not reach it: what a
router must reproduce is kernel behaviour — translation, connection tracking, filtering,
forwarding — and a container has the same kernel.
**Verified before deciding, not assumed.** In a plain unprivileged container:
| Needed for | Works |
|---|---|
| routing at all | `net.ipv4.ip_forward`, `net.ipv6.conf.all.forwarding` |
| `nat:` | nftables masquerade, rules accepted and listed back |
| `mapping_ttl:` | `nf_conntrack_udp_timeout`, `nf_conntrack_tcp_timeout_established` |
No privileged mode, no nesting, no capability grants.
## Consequences
- A raise stops paying a boot per router. Scenery costs about a second where a node costs ten,
and a scenario's cost tracks the machines actually under test.
- **The distinction is now load-bearing and has to stay legible.** *Node* means something under
test; *scenery* means something that makes the test real. If anything is ever installed on a
router by the mesh, it has become a node and this decision no longer covers it.
- Routers and nodes are different kinds of thing in the lab's own model, which is a small extra
concept — justified by it being true, rather than by the saving.
- A container shares the host kernel, so a scenario cannot reproduce a router running a
*different* kernel from the workstation. Nothing currently wants that; if something does, that
router becomes a virtual machine and this record needs revisiting rather than bending.
- The gateway stays implicit in the declaration. A scenario declares `gateway:` on a segment and
never names the machine that serves it — which is right, because it is not a machine the
scenario has anything to say about.
## References
- [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) — what a lab *node* is, unchanged.
- [Research 004](../01-RESEARCH/004-lab-network/analysis.md) — the topology needing a router,
and why *published but behind NAT* only exists in production today.
@@ -0,0 +1,91 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0028-the-substrate-supplies-the-control-plane-and-nothing-else.md
---
# 33. The substrate is a store and a broker
## Context
Third correction to one table in one day, all found the same way: by asking whether **both** halves
of the substrate test were actually answered for a given member, or only the second.
The test ([ADR 0006](0006-the-substrate-and-the-control-plane.md)) is *what the control plane needs
in order to run, and cannot ask itself for, because it is not running yet.* ADR 0006 admits the
image registry on this line:
| role | product | |
|---|---|---|
| image registry | **an OCI registry** | it cannot grant itself a repository |
**That is the second half again.** It is true that a control plane cannot grant itself a
repository. Nothing establishes that it needs one *in order to run*.
**Counted rather than argued.** `substrate-first-node.lock` — the only bundle there is, and what a
first node actually becomes — raises twelve resources, and no registry is among them:
```
container runtime · the store · one database per context · the schemas
· the broker's certificate · the broker · the control plane
```
The registry arrives afterwards, as an ordinary module the mesh assigns. That is what the lab
asserts, in those words: *the mesh runs its own artifact store.*
**ADR 0006 half-said this already**, calling the registry *substrate by role and ordinary by
delivery, provisioned once there is a control plane to do it.* A member that is provisioned by the
thing it supposedly precedes is not a member; the phrase was carrying a contradiction rather than
resolving one.
**The registry is a closer call than the object store, and the difference is worth keeping.** The
control plane never touches an object store at all — no client, no bucket, ever
([ADR 0028](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md)). It genuinely
*uses* the registry: the builder pushes to it, hosts pull from it, and nothing reaches a machine
without it. **So the registry is a real dependency of the mesh operating, and not of the control
plane starting** — and it is the second that the word substrate means.
## Decision
**The substrate is two things: a relational store and a message bus.** Both are in the bundle,
both must exist before the control plane's first instruction, and neither can be asked for.
**The registry is an ordinary module.** The mesh cannot deliver anything without one, and it
installs one the way it installs everything else. The first node's chicken-and-egg is already
solved and needs nothing from this list: it fetches upstream images directly, then runs a registry
of the mesh's own.
**The test is applied to both columns, every time.** *Cannot grant itself one* is true of almost
any service and settles nothing on its own. It is what admitted the object store, and then the
registry, and both were removed by asking the other question.
## Consequences
**The substrate is now exactly what the bundle raises**, which is the strongest form this list can
take: it can be checked by counting rather than by reading an argument. A member that is not in
the bundle is not substrate, and the two statements cannot drift apart.
**A mesh that builds nothing still needs a registry** — to receive anything at all — but it needs
it as a module, on its own schedule, replaceable. That was already true and was obscured by the
list.
**The word may now be doing too little work.** "Substrate" for *a database and a broker* is a term
of art for two things everybody can name. Renaming is not taken here and is worth considering
separately; what this record fixes is the membership, not the vocabulary.
**Three removals from one table in one day is itself the finding.** Each member was admitted on the
half of the test that is easy to answer, and the design read plausibly throughout. The rule that
comes out of it is not about substrates: **a test with two conditions is a test only when both are
asked.**
## References
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — the definition, and the table this
corrects a second row of
- [ADR 0028](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md) — the object
store, removed for the same reason
- [ADR 0031](0031-the-control-plane-authenticates-nobody.md) — identity, which was conditional and
is now a module
@@ -0,0 +1,81 @@
---
topic: how we work
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
supersedes: 02-DECISIONS/0032-the-local-account-owns-the-mesh.md
---
# 34. The local account owns the mesh, and a web application's login is not that
*Supersedes [ADR 0032](0032-the-local-account-owns-the-mesh.md), which decided the right thing and
described it wrongly. The decision below is unchanged; what it said about the board was an
invention.*
## Context
ADR 0032 answered *who owns the mesh* — the account that installed the host — and then framed the
board as **a surface that delegates authentication**, a category it made up for the occasion. It
does not need one.
**The board is a web application.** It has a login, provided by the identity module, in the way
every web application has a login. That is a fact about an application, not a property of the
mesh, and giving it a name in the mesh's vocabulary implied a relationship that is not there.
The cost of the invented category was not cosmetic. It made the identity module look like part of
the mesh's own authority — something the mesh *depends on* to know who anybody is — when the truth
is that the mesh knows nothing about people at all, and one of the applications running on it has
a login.
## Decision
**The account that installed the host owns the mesh on that node.** Authority is a local login.
There is nothing else to hold, no user model, no roles, and nothing to administer.
**This follows from what was already decided.**
[ADR 0004](0004-a-node-and-how-it-joins.md) says there is no authorisation between nodes — every
node is the operator's own, so a message from one is a message from them, and *the mesh boundary
is the security boundary.* A user model inside that boundary would guard nothing: anyone it could
stop could read the node's key off the disk.
**A web application's login is its own business.** The board authenticates its users through the
identity module. So might anything else the mesh runs. **None of that is mesh authority**, and the
mesh does not learn who anybody is from it.
## The line this draws, which is the reason to write it down
**Signing in to an application must not, on its own, become authority over the mesh.**
Today it cannot: the board reads and does not act
([`11-a-board.md`](../03-DESIGN/01-to-be/11-a-board.md) — *not the way to change things*). Looking
at a page tells you what is true and changes nothing.
**The moment the board can assign a module, whoever it lets in has mesh authority** — and it would
arrive as a feature rather than as a decision. That is the failure this record exists to make
visible, because it is the kind that is only obvious afterwards.
So: **a surface that can change the mesh is a change to who owns the mesh**, and is taken as one.
Not forbidden — wanting to manage nodes from a browser is reasonable — but not something that
turns up in a pull request titled *add assign button*.
## Consequences
**The identity module is not special.** Not substrate ([ADR 0031](0031-the-control-plane-authenticates-nobody.md)),
not part of the mesh's authority, and nothing about the mesh stops working when it is down. Some
applications cannot be logged into, which is what it means for an application's login provider to
be unavailable.
**Anyone with a shell on a node has full authority there.** Unchanged from ADR 0032, and still the
sentence that decides who gets an account on a machine. The protection is the machine's own login
and the overlay that keeps it unreachable from outside ([ADR 0007](0007-connectivity.md)).
**A node cannot be operated by somebody without a login on it.** The cost of having no user model.
If that is ever wanted, the paragraph above says what it costs.
## References
- [ADR 0032](0032-the-local-account-owns-the-mesh.md) — superseded; same decision, invented category
- [ADR 0004](0004-a-node-and-how-it-joins.md) — the mesh boundary is the security boundary
- [ADR 0031](0031-the-control-plane-authenticates-nobody.md) — the control plane authenticates
nobody
@@ -0,0 +1,115 @@
---
topic: what runs on it
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0034-the-local-account-owns-the-mesh.md
---
# 35. One implementation, several surfaces, and what that costs
## Context
The mesh is operated from a command line today. It needs to be operable from a browser and from a
model's tools as well, and the three must not be three different systems.
**The pattern is already in the code and unnamed.** `board` serves HTTP by calling the same
functions the CLI calls; it holds nothing and decides nothing. What follows makes that the rule
rather than a property of one command.
**The board is a presentation layer over the control plane.** Not an application beside it holding
a database credential — the thing that shows what the control plane knows, and asks it to do what
a person asked for.
## Decision
**The logic lives once, in the context that owns it. A surface is an adapter with no decisions in
it.**
| surface | for |
|---|---|
| **command line** | a person on a machine, and the recovery path below |
| **HTTP** | the board, and anything else that speaks to the mesh over a network |
| **model tools** | an agent asking the mesh to do something |
**Every surface refuses identically, because the refusal is not in the surface.** An assignment
that cannot be satisfied is refused by the same resolution whichever way it arrived. The moment a
surface can accept something another would reject, the mesh has two answers to one question and
people learn which to trust.
**Reading and doing are both exposed.** The HTTP surface is not read-only: managing the mesh from
a browser is the point. This takes the decision
[ADR 0034](0034-the-local-account-owns-the-mesh.md) said had to be taken deliberately —
**a browser login now carries authority over the mesh** — and takes it knowingly rather than
letting it arrive with a feature.
**The networked surfaces authenticate through an OAuth2 identity provider.** Named by protocol
rather than by product, like every other dependency the mesh takes — AMQP for the bus, S3 for an
object store, OCI for the registry
([ADR 0006](0006-the-substrate-and-the-control-plane.md)). What fills the role today is a module
running Keycloak; what the control plane knows is that it validates a token against a provider
speaking OAuth2, and replacing that provider is a migration rather than a redesign.
**The command line does not authenticate at all**: it is already behind the machine's own login,
which is what owns the mesh (ADR 0034).
## What this is not: a kernel every module imports
**The shared library is the failure this project was started over**, and the difference has to be
stated or it will be rebuilt. The old one is 155 files and 34,636 lines *containing code from
every context* — work-domain logic sitting in the kernel every module imports, each piece landing
there to avoid a cycle between two modules that both needed it.
**Shared surfaces are not a shared library.** What is shared here is that three adapters call the
same functions. Those functions stay in the context that owns them — provisioning's logic in
provisioning, identity's in identity — and no module imports another's. A surface may call many
contexts; a context still may not reach into another's store
([ADR 0008](0008-a-context-owns-its-store.md)).
The test, when something is about to be put "somewhere shared": *does this belong to a context, or
does it only belong to the surface?* If it belongs to a context it goes there, even if two
surfaces want it.
## The loop this creates, and the way out
**The control plane's networked surfaces will depend on a module the control plane assigns.**
An identity provider is an ordinary module ([ADR 0031](0031-the-control-plane-authenticates-nobody.md)). When
it is down, or being migrated, or misconfigured, the HTTP and tool surfaces cannot authenticate
anybody — including the person trying to fix it.
**The command line is the way out, and it is why local ownership matters more rather than less.**
It authenticates through nothing, needs no network, and is available on the machine to the account
that owns the mesh. **A mesh must always be operable by somebody standing at it.**
So the rule: **no capability exists only behind an authenticated surface.** Anything the board can
do, the command line can do. That is not a courtesy to CLI users; it is the recovery path, and a
capability that exists only over HTTP is one that disappears exactly when identity does.
## Consequences
**Identity is still not substrate**, and the test still answers no: the control plane runs, applies
declarations and reaches nodes with no identity provider in existence
([ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md)). What is unavailable without it is two
surfaces, not the mesh.
**Whoever the identity provider admits has authority over the mesh.** That is now a real perimeter
with real consequences, where before it guarded a page that only read. Who may log in, and to
which realm, becomes a decision about the mesh rather than about an application.
**A surface must not grow an opinion.** The likely erosion is a validation added to the board
because it was quicker there — and then the CLI accepts something the board rejects, or worse the
reverse. Adapters hold no decisions.
**Three surfaces over one implementation is a cost paid three times if it is not one
implementation.** The reason to write this down now is that the second surface is the cheapest
moment to get it right, and the third is where the drift usually starts.
## References
- [ADR 0034](0034-the-local-account-owns-the-mesh.md) — the local account owns the mesh, and the
line this record deliberately crosses
- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — the shared library this must not
become
- [ADR 0008](0008-a-context-owns-its-store.md) — a context owns its store, which a surface does
not change
@@ -1,75 +0,0 @@
---
status: accepted
date: 2026-08-25
deciders: jochen
reconstructed: false
---
# 36. A node is a managed machine, and disconnection is a situation
## Context
[Research 006](../01-RESEARCH/006-mesh-from-scratch/00-overview.md) left open: *"does an
unprivileged node earn a place in the inventory, or only a presence? Decides whether 'node'
means one thing or two."*
The question came from requirement 6 — *Arch Linux only for now; ideally any device, including
phones, on lighter terms* — and from the observation that some machines cannot be fully
managed. A phone will not run the host. A laptop is absent for days.
The question assumed the answer was a **class**: full nodes and lesser ones, with the
inventory recording the first and merely acknowledging the second.
## Considered options
1. **Two classes — nodes and presences.** An unprivileged device gets a lighter record and a
reduced contract. Rejected: it makes "node" mean two things, so every context that reasons
about nodes acquires a branch, and the branch is invisible in the type. The mesh already has
one instance of this shape and it is the one this repository keeps writing issues about —
a declared thing that is only sometimes honoured.
2. **One class, membership by capability.** Everything is a node; what it can do is a property.
Chosen.
## Decision
**A node is a managed machine inside the mesh.** Not a device that is merely known about, not
an unprivileged something. If the mesh does not manage it, it is not a node — it is a client, a
peer, or a thing on the network, and those want their own names rather than a weakened version
of this one.
**A disconnected node is still a node, in a different situation.** Reachability is state, not
class. A node that is switched off, roaming, or behind a connection that has dropped has not
become a lesser kind of thing; it has a last-known state and a pending set of declarations.
The distinction the original question reached for is real, but it is **capability**, not kind —
what this machine can be asked to do — and that belongs in the host's profile, not in the
definition of a node.
## Consequences
- **The inventory has one shape.** No branch, no second record type, no context that must ask
which kind it is holding.
- **Local state is structural, not a convenience.** If disconnection is an ordinary situation
rather than an exception, the host's store is authoritative while disconnected by design —
it is what makes the situation ordinary. This promotes `store/` from a component to a
requirement.
- **Absence is not failure.** A node that has not been seen is in a state, and the mesh must be
able to say which. Anything that treats unreachable as broken will be wrong most of the time
about a laptop.
- **Devices that cannot be managed do not become nodes by being lenient about the word.** A
phone that cannot run the host is not a node under this record. Whether the mesh should reach
such devices at all, and as what, is not decided here and needs its own record if it is
wanted.
- **The reduced-contract idea is not lost, it is relocated.** What a given node can be asked to
do is its profile — the host's capability detection — and varies per machine without varying
what a node is.
## References
- [Research 006](../01-RESEARCH/006-mesh-from-scratch/00-overview.md) — the open question, and
requirement 6 that raised it.
- [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) — *nodes host*; this says what a
node is.
- [Issue 007](../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md) —
capability as something detected rather than assumed, which is where the reduced contract
now lives.
@@ -0,0 +1,99 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0035-one-implementation-several-surfaces.md
---
# 36. Bootstrap ends at a usable mesh, and the first credential comes from a person
## Context
Bootstrap currently ends when the control plane starts
([`07-the-substrate.md`](../03-DESIGN/01-to-be/07-the-substrate.md)). That is a mesh that runs and
cannot yet be used by anybody who is not standing at the machine: the networked surfaces need an
OAuth2 identity provider ([ADR 0035](0035-one-implementation-several-surfaces.md)), the provider is
a module, and no module has been assigned.
**So bootstrap should go further** — through the identity provider and the first login — and stop
at a mesh somebody can actually use.
**One thing in the way, and it is not incidental.** The mesh has never held a readable secret. The
sealing code says what it does and why:
> Make generates a secret and seals it to both ends, **keeping no readable copy.**
An initial administrator's credential is the first value a **person must read**. Everything else
the mesh generates is something no human ever sees, and everything a human provides is something
the mesh immediately stops being able to read.
## Considered Options
1. **The mesh generates it and prints it once**, to the terminal of whoever ran the bootstrap.
Convenient, and needs no prompt. **Rejected.** It would give the control plane a plaintext
secret for the first time — briefly, and only to one terminal, but the capability would then
exist. *An exception made for one case does not stay one*: the next credential that is awkward
to supply gets printed too, and the property that a copy of the mesh's database is a copy of
nothing stops being checkable by reading the code.
2. **No password: a one-time link that lets the operator set their own.** The nicest to use.
**Rejected for now** — it needs a mechanism that does not exist, and the thing it improves is
one prompt, once, on a new mesh.
3. **The operator supplies it.** **Adopted.**
## Decision
**Bootstrap runs to a usable mesh**: the substrate, the control plane, the identity provider as an
ordinary module, its realm and client provisioned, an administrator able to log in, and the
networked surfaces available.
**The administrator's credential is supplied by the person doing the bootstrap**, on standard
input and not echoed — the path that already exists for a model-access key. The mesh seals it and
cannot read it afterwards.
**What is created is an account in the identity provider, not a user of the mesh.** The mesh still
has no user model and gains none here ([ADR 0034](0034-the-local-account-owns-the-mesh.md)). What
this produces is the first login for the applications that have one.
**The provisioning is ordinary.** A realm, a client and a first account are what an identity
module's provisioner makes from what the mesh granted it — the same shape as a database and a
bucket, which are built and proven.
**The surfaces arrive when their dependency does.** The command API is not started with the
control plane and then broken until identity exists; it becomes available once it can authenticate,
the way anything else waits for a provider.
## Consequences
**The mesh still never holds a readable secret**, and that sentence needs no exception clause.
That is the whole reason for the prompt.
**An unattended bootstrap is still possible, and the value still comes from outside.** Automation
supplying the credential is the operator supplying it. What is refused is the *mesh inventing*
one — so an unattended bootstrap with no credential provided produces a mesh with no
administrator, which is correct rather than broken.
**Bootstrap gains an interactive step**, and it is the only one. Worth stating because a bootstrap
that cannot run without a person is a real constraint on how a node is stood up, and this is
deliberate rather than an oversight.
**The identity provider is still not substrate.** It is assigned by the control plane, so it comes
after it, and a thing that comes after cannot be a thing that must exist before
([ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md)). Bootstrap running through it does not
move it: bootstrap is a sequence, the substrate is a dependency.
**And the recovery path is unchanged.** When the identity provider is broken later — which is the
failure that matters, not the one at first start — the command line still works, because it
authenticates through nothing (ADR 0035).
## References
- [ADR 0035](0035-one-implementation-several-surfaces.md) — the surfaces, and why the command line
must keep working
- [ADR 0034](0034-the-local-account-owns-the-mesh.md) — the local account owns the mesh; this adds
no user model
- [ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md) — what must exist before the control
plane, which this does not change
@@ -1,97 +0,0 @@
---
status: accepted
date: 2026-08-25
deciders: jochen
reconstructed: false
---
# 37. The host applies; it does not decide
## Context
The skeleton absorbs overlay membership, packet filtering, package management, service
supervision, the container runtime and filesystem management into tier 0, and
[research 006](../01-RESEARCH/006-mesh-from-scratch/00-overview.md) called this *"the
skeleton's biggest unproven claim. A binary whose whole argument is that it has no dependencies
now carries six concerns."*
That claim has now been measured against the monorepo's `main`:
[`host-size.md`](../01-RESEARCH/006-mesh-from-scratch/host-size.md).
The measurement says the question asked about the wrong axis.
## Considered options
1. **Absorb the six concerns as they are.** What the skeleton literally proposes. Rejected on
evidence: two of the ten modules implementing them open a direct connection to the control
plane's database and compute their own configuration. Absorbing those unchanged puts a
Postgres client and knowledge of the mesh schema inside tier 0 — an upward dependency, and
the tier rule is the whole of the bootstrap argument.
2. **Leave them as modules.** Keeps the tier rule trivially, and keeps the fault that prompted
the skeleton: four modules constituting *how a node is reachable* with no relationship the
mesh can see, so one intent is expressed four times
([research 005](../01-RESEARCH/005-domain-grouping/analysis.md) finding 4 measures this and
finds it is the only place in the catalogue where the shape genuinely occurs).
3. **Split each concern: decide centrally, apply locally.** Chosen.
## Decision
**The host has one concern: apply declared state on this machine.** The six are not six
concerns it carries; they are instances of the one.
Each divides:
- **Deciding** — what this node's overlay, names, exposure, filtering, packages and services
*should be*. This needs every other node, and belongs to the control plane.
- **Applying** — putting that on the machine. This needs root and locality, and belongs to the
host.
**The host never queries the mesh database.** A host that reads the control plane's schema is
tier 0 depending on tier 2, and the tiers stop being a bootstrap answer the moment that is
permitted once.
## Why the evidence supports it
**Size was the wrong worry.** The ten modules total 2 755 lines. The machinery that already
applies state on a node — `meshware`, `env-sync`, `config-sync` — is 3 059. Everything being
absorbed is smaller than what already exists to apply it. The host is not a new large thing; it
already exists, spread across three core modules and unnamed.
**Eight of the ten are already pure appliers.** They receive derived state and put it on the
machine. Absorbing them moves code that has no dependency to move.
**The split has already been happening, unnamed.** `dnsmasq-app` needs the same mesh-wide data
as `wireguard` and does not query for it. Its own comments record why: the values were
*"duplicated by hand on all four nodes"* until someone derived them centrally, after a rename
meant editing four override rows nobody knew about. That is this decision, reached once by
fixing a bug.
## Consequences
- **Two modules must be split before they can be absorbed**, and they are the two hardest.
`wireguard` needs every node's key, address, site and endpoint reachability; `traefik` needs
certificates and every node's exposed names. The measurement says the design is right; it
does not say the migration is cheap, and this record does not claim it is.
- **The overlay and firewall modules stop existing** as the skeleton says — but the reason is
now sharper than "they are host concerns". The host holds membership and applies filtering;
the control plane decides policy; swappable backends stay modules.
- **Six vocabularies remain.** Zero dependencies, but the host must still know what a WireGuard
peer, an nftables rule, a package, a unit, a container and a dataset *are*. That surface is
the residue of the original worry and is not measured by anything here.
- **A dependency-direction lint is now load-bearing**, not a nicety. This record is a rule
about direction, and per this repository's own standard a rule states how it is checked: an
upward import fails the build. A tier rule enforced by intention is the same as no tier rule.
- **What the host carries versus what it finds is still open.**
[Issue 007](../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md) — the
host manages `wg`, `nft`, `pacman`, `docker`; it does not contain them, and *installed* is
not the same as *usable*.
## References
- [`host-size.md`](../01-RESEARCH/006-mesh-from-scratch/host-size.md) — the measurement.
- [Research 005](../01-RESEARCH/005-domain-grouping/analysis.md) — reachability as the only
measured co-change cluster in the catalogue.
- [ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) — what the control plane decides
from.
- [ADR 0030](0030-the-repository-structure.md) — `mesh-host` as tier 0.
- [ADR 0008](0008-a-failed-step-fails-the-job.md) — the standard the direction lint is held to.
+103
View File
@@ -0,0 +1,103 @@
---
topic: building it
status: proposed
date: 2026-09-01
deciders: jochen
reconstructed: false
rests-on: 02-DECISIONS/0009-modules-and-the-graph.md
---
# 37. Where a module lives
## The question
The mesh's own module descriptions currently sit in `examples/` inside the control plane, beside
the small programs that hand out logins. That was fine while there were three of them. It is
wrong now, and the name is doing active harm: everything in `examples/` reads as a sketch, and one
of them shipped naming a container image that nothing in the repository builds. A directory called
*the catalogue* would have made *does this actually work* the obvious question to ask of it.
So: **one repository holding the modules we ship?** And if so, where does everything that is not
ours go?
## What a module actually is, counted
The system being replaced has **126 modules** on its main branch. The shape of them is the whole
argument, so it is measured rather than assumed:
| | count | what it is |
|---|---|---|
| **the module is software** | 47 | its own source tree lives inside the module — a daemon, a service, a library |
| **helper scripts only** | 44 | no application of its own; scripts it runs at install time or offers to an agent |
| **a description and nothing else** | 35 | a package to install and some files to write |
**Two thirds of modules contain code.** The largest is a shared library of 182 source files. A
speech-capture module carries a complete daemon — audio capture, mixing, transcription, a model
runner. Treating a module as *a description of something else* is true of barely a quarter of them.
That kills the simplest answer. A catalogue cannot be "a folder of manifests" when most modules
are programs.
## The four kinds, which want different homes
**1. What the mesh is made of.** The control plane, the host, the shared library, the board.
These are not modules that happen to be ours; they are the mesh, expressed as modules so it can
install itself. They belong in the repositories that build them, which already exist.
**2. Something the world made, that we describe.** A forge, a mail system, an identity provider,
a media server. Nobody upstream ships a description; somebody has to write one, and it is the same
description for everybody who runs it. **This is what a catalogue is for.** It is also where the
small programs that create accounts belong, because such a program is part of describing that
service, not part of the mesh.
**3. Something we wrote, that runs somewhere.** An application, a site, a side project. The
description belongs **with the code, at the root of its own repository**, because the two change in
the same commit. A repository that gains an environment variable and a description that gains it
elsewhere will drift, and there is no mechanism that could stop it. This is already how it works
and it should stay that way.
**4. A package and some files.** A tool, a font, a shell. Thirty-five of these, and each is a few
lines. The catalogue.
## The proposal
**A `mesh-catalog` repository** holding kinds 2 and 4: descriptions of software we did not write,
and the programs that provision it. Not kind 1, which is the mesh itself. Not kind 3, which lives
with its own code.
**The mesh's list of modules is not this repository.** It is a table in the control plane, filled
by adding a description to a running mesh. The catalogue is a *source* to add from — one of
several, and the mesh already records which: every module carries where it came from, the branch
followed there, and the commit its description was read at. **Nothing needs inventing to support
modules from anywhere**; a repository of our own is simply the source we curate.
**A description is checked by the tool, not by a test that imports the tool.** Today a test in the
control plane parses the example manifests by reaching into the control plane's internals, and
another reads the control plane's own build file to check every image a module names can be built.
Two jobs tangled. A `module check` command on the control plane's binary would let the catalogue
hold data validated from outside, and would give the same check to somebody describing their own
application in their own repository — which is the case that matters most and currently has no
check at all.
## What this costs, and the argument against
**It is early.** Ten modules exist, four of them ours. Moving ten files is a morning; moving a
hundred is a week — but the hundred is not here yet, and splitting now adds a second repository to
release across before there is anything to release.
The counter is that the tangle is already producing faults rather than merely threatening to. A
manifest naming an unbuildable image, and a test reading a build file two directories up, are both
symptoms of one repository doing two jobs. And the moment the first module is adopted on a real
machine, the descriptions stop being examples and become the thing deployments come from. **That
is the moment this becomes urgent, and it is close.**
## What it does not settle
**Where a provisioning program's image is published**, and how a description pins it. A description
names an image by digest; the image is built from the catalogue; the catalogue must therefore both
produce an image and refer to it, which is the same knot the bootstrap has and solves by writing
the digest down after building.
**Whether kind 4 deserves a module at all.** Thirty-five descriptions that say *install this and
write these files* may be better as one module with settings than as thirty-five modules. Left
open deliberately; it is a question about the shape of the catalogue, not about whether to have one.
@@ -1,106 +0,0 @@
---
status: accepted
date: 2026-08-25
deciders: jochen
reconstructed: false
extends: 0037-the-host-applies-it-does-not-decide.md
---
# 38. A node joins by linking first, and the mesh finishes the job
## Context
[ADR 0037](0037-the-host-applies-it-does-not-decide.md) settles that the host applies and the
control plane decides. That leaves the case where there is no control plane to decide: the
first node, which must raise a mesh from nothing, and the second, which must join one.
Raised by the operator: *"shouldn't the host have two modes — one for the initial node, setting
up the mesh, so we know the full state; then when adopting a second node, we enter the mesh
early and let our first node take over the mesh-related work? The host should only set up the
bare minimum for the other nodes in the mesh to complete adoption."*
The instinct is right and it is the resolution of the gap ADR 0037 leaves open. The framing
needs one correction, and the correction comes from what the mesh already does.
**Today there are three hand-run shell paths**: `install.d/adopt.sh` (183 lines),
a separate first-node bootstrap, and `install.d/rescue.sh`. The skeleton names the cause —
*"the first node is raised by a special script that exists only because of the circularity"* —
and Move 1 exists to remove it. Three paths that do nearly the same thing, maintained
separately, run by hand, outside anything that checks them.
**That is the two-mode problem, already at its worst.** A decision that gives the host two
modes risks rebuilding `adopt.sh` and `bootstrap.sh` inside the binary, where they will drift
in exactly the same way and be harder to see.
## Considered options
1. **Two modes — genesis and join.** What was proposed. Rejected as a *structure* while adopted
as an *intent*: two modes is two code paths, the first is exercised once per mesh and the
second constantly, so the rarely-run one rots. The current three scripts are the evidence.
2. **One path, and the first node is special-cased inside it.** The conditional moves rather
than disappearing, and now it is scattered instead of named.
3. **One behaviour, two sources of declaration.** Chosen.
## Decision
**The host has one behaviour: apply the declaration it has.** What differs between the first
node and the fiftieth is not what the host *does* but **where the declaration comes from** —
and, exactly as in [ADR 0036](0036-a-node-is-a-managed-machine.md), that is a situation rather
than a class.
| Situation | Declaration comes from |
|---|---|
| no mesh reachable | the pinned bundle the host carries (`substrate.lock`) |
| mesh reachable | the control plane, over the link |
**The first node is not a different kind of node.** It is a node whose mesh is not up *yet*. It
applies the bundle it carries, the control plane comes up on top of it, and from that moment it
takes declarations like everything else. Its specialness is temporary and self-erasing, which
is the property `adopt.sh` and the bootstrap script do not have.
**A joining node does the minimum to be reachable, and nothing else.** It establishes identity
and a route to the control plane — the `link` — and then stops deciding. Everything after that
arrives as declarations.
**The minimum is deliberately small:** an identity, an address, and one peer to reach. A
joining node does **not** compute the overlay. It needs a single peer to reach the mesh; the
full peer set is derived centrally and pushed down, like everything else.
## Why this resolves what 0037 left open
ADR 0037 records that `wireguard` and `traefik` are the two modules that must be split before
they can be absorbed, and that they are the hardest because they need mesh-wide state.
**A joining node never needs that state.** The hard part of the overlay — every node's key,
address, site and endpoint reachability — is only needed to compute the *whole* mesh, which is
the control plane's job. The node needs one peer. The rest arrives.
So the migration ADR 0037 calls expensive is smaller than it looked, and this record is what
makes it smaller.
## Consequences
- **Adoption stops being a script.** The three hand-run paths collapse into the host: joining
is establishing a link, and rescue is a node whose local state is discarded so the mesh can
re-derive it. Whether rescue is fully covered by this is not decided here.
- **The bundle is a fallback, not a mode.** It is what a host applies when nothing better is
available, which also covers a node that has been disconnected for a long time — ADR 0036's
ordinary situation.
- **The rarely-run path is now the common one.** The first node exercises the same code every
other node exercises constantly. That is the whole reason for choosing this over two modes.
- **The link becomes the security boundary.** Everything a node applies arrives through it, so
what may be pushed, and how a joining node proves it is entitled to join, is its own
question — taken up by [ADR 0039](0039-the-link-is-the-security-boundary.md).
- **The bundle must be able to raise the substrate alone.** Whether one host can bring up the
four pinned services with no mesh present is Move 1 of the skeleton and remains unproven.
This record depends on it and does not establish it.
## References
- [ADR 0037](0037-the-host-applies-it-does-not-decide.md) — the split this completes.
- [ADR 0036](0036-a-node-is-a-managed-machine.md) — situation rather than class, applied here
to the first node.
- [Research 006, Move 1](../01-RESEARCH/006-mesh-from-scratch/skeleton.md) — the pinned bundle,
and the special script it exists to remove.
- [`00-as-is/05-runtime-and-installation.md`](../03-DESIGN/00-as-is/05-runtime-and-installation.md)
— how a node comes into being today.
@@ -0,0 +1,82 @@
---
topic: what runs on it
status: proposed
date: 2026-09-01
deciders: jochen
reconstructed: false
rests-on: 02-DECISIONS/0009-modules-and-the-graph.md
---
# 38. The mesh assigns the port, and a module does not care
## The problem, as met
A database module cannot start on a machine that runs the control plane. The mesh keeps its own
store there and holds 5432; the module publishes 5432. Nothing notices until a container runtime
three layers down says `port is already allocated`
([`028`](../04-ISSUES/028-two-things-want-one-port-and-nothing-says-so/00-report.md)).
A module cannot fix this by choosing better, because **a module cannot know what else is on the
machine.** It is written once and assigned anywhere. Any number it picks is a guess about a
machine it has never seen, and two modules guessing the same number is not a mistake either of
them made.
## The number is written three times, and nothing makes them agree
Every module says its port in three places:
| where | for | example |
|---|---|---|
| `listens` | the rule set that lets traffic in | `{port: 5432, from: mesh}` |
| `serves` | what a consumer must know to connect | `{port: 5432}` |
| a container's `ports` | what the runtime publishes | `"5432:5432"` |
They agree today because one person wrote all three. Nothing checks it. A module whose `serves`
said 5432 and whose container published 5433 would resolve, compose, apply, and hand every
consumer a port that answers nothing.
## The decision
**The mesh assigns the machine-side port, and the module says only what it needs.** A module
declares that a container port must be reachable and what it is for. Which number the machine uses
is the mesh's to choose, because the mesh is the only thing that knows what else is there.
**One source, and the other two are derived.** `serves` carries the assigned port so a consumer is
told where to connect without the module having written it down; the rule set is computed from the
same assignment. Three copies become one fact.
**An assignment is made once and kept**, exactly as a credential is. A port that moved on every
push would restart both ends each time and would hand consumers a number that was true when it was
read.
## Some ports cannot move, and that is a claim
Mail is 25, submission is 587, IMAP over TLS is 993. A mail system on a strange port is not a mail
system. So a module may say a port is **fixed by the protocol** rather than assigned.
**A fixed port is exactly a claim** — the thing the mesh already has for what is singular on a
machine: one seat, one display server, one artifact store. Two modules wanting 25 on one machine is
the same shape as two wanting the seat, and gets the same answer: the second is refused, by name,
when it is assigned rather than when it is applied.
That is why this does not need a new mechanism so much as it needs the existing one pointed at
ports.
## What follows
- **A module becomes portable in a way it was not.** Two databases on one machine stop being a
collision and become two assignments.
- **The substrate has to be visible.** The mesh cannot assign around its own store while it has
never heard of it. What the bundle holds must be written down somewhere the assignment can read
— which the bundle does not say today.
- **A refusal can be useful.** *25 is held by the mail system on this machine* is a sentence a
person can act on. `port is already allocated` is not.
- **`serves` stops being written by hand**, which is a small vocabulary change with a large
consequence: what a consumer is told is now derived from what actually happened.
## What this does not settle
**Whether a module should publish to the machine at all.** Assignment makes publishing safe; it
does not make it necessary. Consumers could instead reach a provider on the module's own network by
name, with nothing published — which would make the question moot for anything inside the mesh, and
would still leave it for anything reached from outside.
@@ -1,190 +0,0 @@
---
status: accepted
date: 2026-08-25
deciders: jochen
reconstructed: false
extends: 0038-a-node-joins-by-linking-first.md
---
# 39. The link is the security boundary
## Context
[ADR 0038](0038-a-node-joins-by-linking-first.md) makes the link the one channel a node takes
declarations from, and names the gap it leaves: *"everything a node applies arrives through it,
so what may be pushed, and how a joining node proves it is entitled to join, is now a question
worth its own record."*
This is that record. It is a design decision about a boundary that does not exist yet — but
what it replaces is measured, and that is the argument.
Settled as: **a node owns no password. It owns an identity, and that identity is what it
presents to the broker.**
### What adoption does today
`install.d/adopt.sh` asks the operator to paste credentials in by hand:
```
The meshware module needs registry database and minio credentials.
REGISTRY_DB_PASSWORD=<postgres password from novox>
REGISTRY_MINIO_PASSWORD=<minio password from novox>
```
plus an `NPM_TOKEN` for the private registry. These are not adoption-time credentials that are
then discarded: `wireguard` and `traefik` open a `pg` connection on every reconcile
([ADR 0037](0037-the-host-applies-it-does-not-decide.md)).
**So every node permanently holds a credential to the control plane's database, and to the
object store.** They are the same credentials on every node. There is no rotation —
[`00-as-is/06`](../03-DESIGN/00-as-is/06-configuration-and-secrets.md) records that *"there is
no mechanism that rotates one and informs everything holding it. Where a rotation has been
done, it has been done by hand, and doing it wrong has taken services down."*
Compromise of any node is therefore compromise of the mesh's database, and there is no
mechanism to recover from it.
### The link is not new
Written first as though the link were a thing to build. It is not.
[ADR 0001](0001-nodes-communicate-over-a-broker.md) already has it: *every node connects
outbound to a single broker; nothing ever connects to a node*, each node declaring an exchange
named for itself and consuming from its own queue
([`00-as-is/01`](../03-DESIGN/00-as-is/01-mesh-and-transport.md)).
That is already outbound-only, already per-node addressed, and already the one channel
everything arrives through. **This record is not proposing a channel. It is proposing that the
channel carry per-node identity instead of one shared credential.**
The same as-is records the fault, for the broker rather than the database: *"the broker is a
single point of failure and a single point of trust. Its credential is mesh-wide, so rotating
it is a mesh-wide operation, and doing it wrong has taken the broker down."*
## Considered options
1. **Keep shared credentials, scope them per node.** Least change: give each node its own
database role. Rejected — it makes the blast radius smaller without changing its shape, and
it keeps tier 0 speaking the control plane's schema, which ADR 0037 forbids for reasons that
are not about security at all.
2. **Accept the exposure as the cost of simplicity.** A shared credential is one thing to
understand and nothing to build, and the objection to replacing it is real: mutual
authentication fails opaquely, and a node that cannot link is harder to debug than a node
with a wrong password. Rejected on the ground that the simplicity is what makes it
unrotatable — the credential cannot be changed *because* everything holds the same one, so
the arrangement's convenience and its unfixability are the same property.
3. **Mutual authority on a node-initiated link, with the node holding nothing but its own
identity.** Chosen.
## Decision
**The link is the only way anything reaches a node**, and four properties make it a boundary
rather than a pipe.
### It is outbound and node-initiated
The node dials the control plane. Nothing dials a node. This is not only defensive — it is what
the topology already requires: most nodes sit behind a household connection with no forwarded
port ([research 004](../01-RESEARCH/004-lab-network/00-overview.md)), so an inbound control
channel would work for the hosted node and not for the rest, and the difference would be
invisible until it mattered.
A node therefore has **no listening control surface at all**.
### A node holds its own identity and nothing else
No shared secret, no credential to anything it does not own. A node's identity authenticates it
to the control plane and grants access to nothing else.
**Compromise of a node is compromise of that node.** That is the property today's arrangement
does not have, and it is the main reason for this record.
### Authority is mutual
The node proves it may join, and **the control plane proves it is the mesh**. One-way is not
enough here: the host applies whatever the link delivers, so a node that cannot tell the mesh
from something impersonating it will apply that something's declarations. Given ADR 0038, an
attacker who can answer a joining node's first call owns the machine.
### What may be pushed is bounded by form, not by trust
The control plane may push **declarations of known shape** and nothing else. It may not push a
command to run. The host's vocabulary is finite, versioned and auditable, and anything outside
it is refused rather than best-effort interpreted.
**Stated honestly: this bounds form, not impact.** A compromised control plane can declare
harmful state — a malicious package, an open firewall — and the host will apply it faithfully,
because that is what it is for. What the property buys is that the blast radius is describable:
it is exactly what the declaration language can express, which can be reviewed. An arbitrary
command channel has no such bound. This is a real limit and not a defence-in-depth story.
### Joining is a deliberate, bounded act
A joining node presents a **one-time, short-lived enrolment token** issued by the mesh for that
purpose, and exchanges it for its own durable identity. The token grants exactly one thing:
the right to become a node. It is not a credential to any service, it does not persist after
exchange, and it expires whether used or not.
This replaces hand-carried shared secrets with a thing that is useless once used and useless
after a while.
## What this actually costs
The objection to weigh is overhead, and it is smaller than it looks because most of it is
already running.
| Property | Where it comes from |
|---|---|
| outbound, node-initiated | already true — ADR 0001 |
| per-node addressing | already true — per-node exchange and queue |
| per-node credential | a broker user per node; the broker already has users, virtual hosts and per-queue permissions |
| mutual authority | transport-level certificates on a connection that already exists |
| bounded by form | already true — three message shapes and only three |
| **enrolment** | **the one genuinely new mechanism** |
And ADR 0037 subtracts rather than adds: under it a node holds **no** database credential at
all, so this record replaces three hand-carried shared secrets with one per-node identity that
grants only identity.
**It must fail legibly.** A boundary that refuses a node without saying why is worse than the
credential it replaced, because a wrong password at least announces itself. A node that cannot
link must report which side rejected it and on what grounds, in terms someone can act on. This
is `how-we-build` §5 applied to a security mechanism: a refusal that proves only that something
went wrong is transport reported as effect.
## Consequences
- **ADR 0037 removes a standing exposure as a side effect.** Its rule — the host never queries
the mesh database — was chosen for tier discipline. It also removes the reason every node
holds the database password. Worth recording because the two arguments are independent and
both hold.
- **Rotation becomes possible and is still not designed.** Per-node identities can be revoked
individually, which is what makes rotation tractable at all. The mechanism —
what rotates, on what trigger, and how holders learn — is **not decided here** and remains
the open weakness `00-as-is/06` records.
- **The enrolment token has to come from somewhere.** Issuing it is a control-plane operation
and the first node has no control plane, so the first node's identity is self-issued and
becomes the root of trust when the mesh comes up. **That is a real asymmetry** — the one
place ADR 0038's "no special first node" does not fully hold — and it is named here rather
than hidden.
- **A declaration vocabulary is now a security artefact, not only a design one.** Every
addition widens what a compromised control plane can express. That is a reason to keep it
small and a reason for additions to be reviewed as such.
- **Offline nodes need identities that survive disconnection.** Per
[ADR 0036](0036-a-node-is-a-managed-machine.md) disconnection is ordinary, so an identity
that must be refreshed to remain valid would make a laptop fail for being a laptop. What
expires and what does not is **not decided here**.
- **This is a boundary that does not exist yet.** Nothing in the current mesh implements any of
it, and the migration from shared credentials to per-node identity touches every node and the
substrate. No estimate is offered.
## References
- [ADR 0038](0038-a-node-joins-by-linking-first.md) — the link, and the gap this fills.
- [ADR 0037](0037-the-host-applies-it-does-not-decide.md) — why the host stops holding database
credentials at all.
- [ADR 0036](0036-a-node-is-a-managed-machine.md) — disconnection as ordinary, which constrains
what may expire.
- [`00-as-is/06-configuration-and-secrets.md`](../03-DESIGN/00-as-is/06-configuration-and-secrets.md)
— secrets today, and the absence of rotation.
- [Research 004](../01-RESEARCH/004-lab-network/00-overview.md) — why most nodes cannot accept
an inbound connection.
@@ -0,0 +1,106 @@
---
topic: building it
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
---
# 39. What the SDK holds, and what it refuses
_Reconciliation note (2026-09-05): supersedes the earlier "repository structure" decision, which the consolidation folded; no standalone record remains to point at, so body references to it now point at the nearest surviving record, [ADR 0015](0015-applications-live-in-their-own-repository.md)._
## Context
The earlier "repository structure" decision (folded in consolidation; see the reconciliation note
above, and [ADR 0015](0015-applications-live-in-their-own-repository.md) as the nearest survivor)
named `mesh-sdk` "contracts shared across tiers:
types, not behaviour." That line is superseded here, because it draws the boundary in the wrong
place. The boundary that matters is not *types versus behaviour* — it is **how often the thing
changes**.
The current SDK is the cautionary tale, and its failure is precise. `hal/sdk` holds all the
code, including a per-module API client for every service (`clients/plex.ts`, `clients/gitea.ts`,
…) and a per-module tool implementation for each (`tools/plex.ts`, …). Every module depends on
the SDK, so **every edit to any of that per-module code rebuilds every module** — the cascade.
The SDK is under constant maintenance precisely because it became the place all the volatile
per-module logic accumulated.
The root cause is worth stating exactly, because the fix follows from it: the pressure was never
to share a client *between* modules. It was to share a client between one module's *own features*
— plex's tools, its health check and its hooks all wanted the same `PlexClient` — and the only
place to share code across a module's features was the global SDK. So **intra-module sharing
leaked out as inter-module coupling.**
## Decision
The SDK holds the **stable spine** that modules build against, and earns its place by rarely
changing. The test for membership is change-frequency, not kind.
### What it holds
- The **tool-serving harness** — the worker and registration mechanism, and the tool-definition
type. *How* a tool is declared and served is settled; it does not change when an individual
tool does.
- The **messaging and event framework** — the broker client, the event consumer, the envelope.
- The **contracts** — the manifest, declaration, provision and link shapes.
- **Core primitives** — sealing and crypto, semver, the shared resolution helpers.
These change rarely and deliberately. When one of them does change, a rebuild of everything is
the *correct* outcome, because the contract every module shares has genuinely changed.
### What it must not hold — the more important half
- **A module's API client.** A Plex client, a Gitea client, a MinIO client belong in their
module. They change when that service's API or the module's use of it changes, which is often,
and which has nothing to do with any other module.
- **A module's tool implementations.** Same reason, same place: in the module.
- **Anything volatile** — anything that changes when one service's features change.
The rule, stated so it can be applied without re-deriving it:
> If editing a thing recompiles unrelated modules **and** it changes often, it does not belong
> in the SDK.
Both conditions are load-bearing. A rare change that cascades is fine — that is a contract, and
the cascade is correct. A frequent change that stays local is fine — that is a module minding its
own business. Only **frequent *and* cascading** is the disease, and per-module clients and tools
are its carriers.
### Where per-module shared code lives instead
Code shared among a module's *own* features lives **in the module**. The default is the plainest
thing that works: an ordinary shared file the features import — `plex/client.ts`, imported by
`plex/tools/`. Within one module, features are files importing sibling files; no package
boundary, no ceremony.
A **module-local SDK** (a sub-package with its own version) is warranted only for the few modules
whose shared surface is large enough to version on its own. It is the exception, not the shape.
Either form gives the property the global SDK could not: editing a module's shared code rebuilds
**that module and nothing else**.
## Consequences
- The cascade becomes **structurally impossible for module logic**. There is no longer an edge
from one module's internals to another, so the only thing that can rebuild everything is a real
change to a shared contract in the SDK — which is rare, and when it happens, is right.
- The SDK is small and stable **by construction**, not by discipline. Its size is no longer a
thing anyone has to police.
- **Converting a module from the current system is partly a de-coupling, not just a move.** Its
client and its tools are pulled *out* of the shared SDK and *into* the module. A conversion
that copied `clients/plex.ts` into the SDK's replacement would rebuild the exact mistake.
- The host still does not import the SDK. It depends on nothing
([ADR 0005](0005-the-node-host.md)) and **mirrors** the contracts rather than
importing them, exactly as its apply-shapes table already does deliberately. The SDK is shared
by the tiers that *can* share code; the host is not one of them.
## References
- The earlier "repository structure" decision — named the repositories; its `mesh-sdk`
description ("types, not behaviour") is superseded by this record (folded in consolidation;
nearest survivor [ADR 0015](0015-applications-live-in-their-own-repository.md)).
- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — the decomposition this serves: code
belongs to the boundary that owns it.
- [ADR 0005](0005-the-node-host.md) — why the host mirrors the contracts instead of
importing the SDK.
+101
View File
@@ -0,0 +1,101 @@
---
topic: what runs on it
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
extends: 0009-modules-and-the-graph.md
---
# 40. What a module is
_Reconciliation note (2026-09-05): supersedes the earlier "grouped by domain" decision, which the consolidation folded into how-we-build.md; no standalone record remains to point at._
## Context
[ADR 0009](0009-modules-and-the-graph.md) settled that everything is a module, but never said what a
module *is* beyond "a directory the mesh processes." That gap let the catalogue's breadth read as a
smell: a module can carry a container, a built image, tools, a provisioner, migrations, health,
config, seat claims, requires and provides — so much that the unit seemed ill-defined.
The earlier "grouped by domain" decision (folded in consolidation; see the reconciliation note
above) tried to organise modules by domain, which is the wrong axis. This record states what a module is, drawn from the cases that
stress-tested it: the shell, i3-vs-sway, umami, and "database."
## Decision
**A module is one self-contained piece of software the mesh installs and manages** — everything
needed to make that one thing real and integrable: what runs, the seats it claims, what it provides
to other modules, what it requires from them, and what operates it.
The **software is the module's identity.** Capabilities, seats and provisioned resources are the
**relationships *between* modules**, not what a module is — and that is what binds a module into one
thing. umami is bound by *being umami*: its container runs umami, its provisioner creates umami sites,
its tools query umami, its `requires` gets umami a database. Every feature serves the one software.
### The three relationships
1. **Shared seat** — several modules fulfil a capability and coexist; one may be default. bash, zsh
and fish all join `shell`.
2. **Exclusive seat** — modules contend for a single slot; one holds it. i3 (needs x11) and sway
(needs wayland) contend for `display-session`.
3. **Provide / require** — a provider ships the **provisioner** that creates instances of the
resource it offers and returns sealed credentials; a consumer requires it and the mesh wires the
credential in. Symmetric: umami requires a database *and* provides analytics.
### Interfaces are mesh-owned; providers adapt to them
The mesh **defines the interface** for a capability — the provider-neutral contract of what a
consumer receives and how it integrates. Both sides conform: a provider's provisioner **adapts** its
software's real API to the mesh contract; a consumer depends on the **interface**, never on a
provider. Swap one provider for another and the consumer does not change.
### The naming rule — draw the interface at the consumer's real coupling
Name a `provides`/`requires` at the **widest boundary across which the consumer genuinely does not
care which implementation serves it**:
- Where the consumer's coupling is thin — an analytics embed snippet and dashboard, opaque to it —
the mesh defines a neutral interface (`analytics`) and providers (umami, amumi) adapt. Swappable
across vendors.
- Where the consumer **speaks a protocol** — a database's wire protocol and query dialect — the
interface *is* the protocol: `postgres-database`, `mssql-database`, `mongodb-database`. Swappable
only among protocol-compatible implementations, **never across**, because the application cannot
cross it either. "database" is not a capability; the protocol is.
- **Never false genericity.** A name must not promise a swap the contract cannot deliver
([research 005](../01-RESEARCH/005-domain-grouping/analysis.md)).
This is [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md)'s rule made general — "names the protocol, not
the product; a database names the engine because the app targets it" — with the reason stated: the
contract sits where the coupling is.
### What is not a module
- A **library** (built against, never deployed — [ADR 0039](0039-what-the-sdk-holds-and-refuses.md)).
- A **control-plane context** (the mesh itself — [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)).
A **swappable machine mechanism** (a firewall — ufw, nftables) *is* a module implementing a
capability. The host hardcodes no firewall, supervisor, package manager or runtime; it owns only the
generic apply primitives and platform detection, so it runs where none of those exist — an Android
phone has no ufw, systemd, pacman or Docker.
## Consequences
- **Supersedes the earlier "grouped by domain" decision** (folded in consolidation; see the
reconciliation note above). Modules are
organised by their relationships (seats, provisions), not grouped into domain folders.
- **Refines [ADR 0009](0009-modules-and-the-graph.md).** Everything the mesh runs and integrates is
a module — but a module is defined by the *software it delivers*, not by being a bucket of features.
- The target is a **self-fulfilling mesh**: declared wants bound to swappable modules, provisioners
wiring credentials, nothing hardcoded. The control plane's whole job is the binding.
- Converting a module from the old system includes pulling its per-module code out of the shared SDK
([ADR 0039](0039-what-the-sdk-holds-and-refuses.md)) and shipping its provisioner as an adapter to a
mesh interface — a de-coupling, not just a move.
## References
- [ADR 0009](0009-modules-and-the-graph.md) — everything is a module; this says what one is.
- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — contexts are the mesh, not modules.
- The earlier "grouped by domain" decision — superseded (folded in consolidation; see the note above).
- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — protocol-not-product, generalised here.
- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — per-module code lives in the module.
- [research 011](../01-RESEARCH/011-the-module-graph/00-overview.md) — the graph of these relationships.
@@ -0,0 +1,85 @@
---
topic: what runs on it
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
extends: 0040-what-a-module-is.md
---
# 41. Events are a relationship, the lighter sibling of provisioning
## Context
[ADR 0040](0040-what-a-module-is.md) names two relationships between modules — seats and
provide/require (provisioning). A third is latent in the mesh and worth making first-class: the
broker every node already runs ([ADR 0002](0002-nodes-communicate-over-a-broker.md)) can carry a
module's activity as **events**, which any other module reacts to. A logger that writes an audit
trail, a module that acts when another module acts, observability — all of it is one mechanism, and
today it is ambient rather than declared.
## Decision
**A module emits events and consumes events, and both are declared** — parallel to `provides` /
`requires`, so the mesh knows the event graph the same way it knows the provisioning graph.
### Events are provisioning's lighter sibling
| | provisioning | events |
|---|---|---|
| shape | **1:1**, a provider creates a resource *for* one consumer | **1:many**, a module emits, any number listen |
| credential | yes — sealed, per consumer | none — it is broadcast |
| machinery | a provisioner (the reconcile adapter) | nothing but the broker's topic routing |
| declared as | `provides` / `requires` | `emits` / `consumes` |
Because an event is broadcast and credential-free, there is no provisioner and no per-consumer
setup — only a subscription. That is why it is the *lighter* relationship, and why most
inter-module reaction should be an event, not a provision.
### An event carries what an audit needs
Every event carries its **type** (a dotted topic key, so listeners match by prefix), its **source**
module, the **node** it came from, and the **time**. A body follows. The metadata is not optional:
a reaction may only need the body, but an audit trail needs to know who did what, where and when,
and an event that cannot answer that is not auditable.
### The audit logger is just a consumer of everything
A logger that records the whole mesh's activity is **not a privileged component** — it is an
ordinary module that consumes `#` (every event) and writes them down. It holds no special access;
it only listens widely. That it falls out of the model with no new machinery is the check that the
model is right.
### `consumes` is validated like `requires`
A `consumes` for an event that **nothing** `emits` is a dangling edge, and the mesh refuses it
before deploy — the same rule that catches a `requires` for a resource nothing provides
([research 011](../01-RESEARCH/011-the-module-graph/00-overview.md)). A listener waiting for an
event that can never arrive is a silent failure, and this repository's whole discipline is against
silent failure.
### One runtime serves all three
The per-node module runtime that serves a module's tools also wires its `consumes` (subscribe,
dispatch to the handler) and lets its code `emit`. Tools are *invoked* (request/reply), resources
are *provisioned* (1:1, credentialed), events are *emitted and consumed* (1:many, broadcast) —
three relationships, one broker, one runtime, all declared on the manifest.
## Consequences
- The mesh gains a declared **event graph** alongside the provisioning graph — visible, validated,
reasoned over.
- **Reaction becomes the default coordination**: a module acts on another's event without either
knowing the other, and without a credentialed link. Coupling drops.
- An **audit trail** is a module, not a platform feature — and can be swapped, extended or run more
than once (a file logger and a queryable one) with no change to anything that emits.
- The runtime must dispatch a module's event handlers as well as its tools; that generalisation is
small (both arrive by importing the module's entrypoint) but it is real work.
## References
- [ADR 0002](0002-nodes-communicate-over-a-broker.md) — the broker events ride.
- [ADR 0040](0040-what-a-module-is.md) — the relationships this extends.
- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — `emit`/`on` are stable sdk surface; the
broker binding and the runtime are not.
- [research 011](../01-RESEARCH/011-the-module-graph/00-overview.md) — the graph these edges join.
@@ -1,78 +0,0 @@
---
status: accepted
date: 2026-08-26
deciders: jochen
reconstructed: false
extends: 0037-the-host-applies-it-does-not-decide.md
---
# 41. The host depends on nothing that must be installed first
## Context
[ADR 0030](0030-the-repository-structure.md) calls tier 0 *"the one binary installed by hand"*,
and [research 006](../01-RESEARCH/006-mesh-from-scratch/00-overview.md) states the property the
whole tier rests on: *"a binary whose whole argument is that it has no dependencies"*.
Building it forced the question that phrase had been carrying unexamined. Everything else in
the mesh is TypeScript, and `how-we-build` §8 says so. A TypeScript host needs a runtime present
before it can run — so the thing installed by hand becomes **two** things, and the second must
be installed by the means the host exists to replace.
## Considered options
1. **TypeScript, with a runtime installed first.** Simplest, and matches every other
repository. Rejected: it breaks the property the tier is built on. A host that cannot run
until something else has been installed by hand is not the bottom of the stack.
2. **TypeScript, bundled as a single executable.** Preserves the language and produces one
file. Rejected on two grounds: it carries roughly ninety megabytes of runtime to preserve a
language choice, and single-executable bundling is a young feature — tier 0 is the worst
place in the system to discover its edges.
3. **A statically linked binary in a language built for it.** Chosen; Go.
## Decision
**The host is a single statically linked binary that requires nothing to be present.** Copy it
onto a machine and run it. That is the whole installation.
**It is written in Go.** The job is system-level — run commands, write files, speak to the
firewall, the overlay, the service manager and the package manager — which is what Go's
ecosystem is for, and it cross-compiles to every architecture the mesh might reach, including
the lighter devices requirement 6 anticipates.
**The second language costs less here than anywhere else it could appear**, and the reason is
architectural rather than convenient. [ADR 0037](0037-the-host-applies-it-does-not-decide.md)
means the host never queries the mesh database.
[ADR 0039](0039-the-link-is-the-security-boundary.md) means it only ever receives declarations.
So the host shares **no code** with any other tier — not a client, not a schema, not the SDK.
It is joined to the mesh by a message contract and nothing else.
The language boundary therefore falls exactly on an architectural boundary that already exists.
A second language usually costs duplicated logic; here there is none to duplicate.
## Consequences
- **`how-we-build` §8 needs a scope.** It reads *"TypeScript throughout"*, which was true when
everything was a service or a surface. It is now scoped to those, with tier 0 named as the
exception and this record as the reason. That is a constitution change, and the sync it owes
is part of it.
- **Agents must write Go to work on the host.** A real cost, and the one genuine argument
against this. It is bounded by the host being the only thing in tier 0 — nothing else in the
mesh acquires a second language because of this.
- **The dependency-direction lint the design calls for gets easier, not harder.** A Go module
cannot accidentally import a TypeScript control-plane client; the boundary is enforced by
there being no path across it.
- **Cross-compilation replaces per-node builds.** The host is built once per architecture and
copied, rather than built on the machine it runs on — which is what makes *"copy it and run
it"* true rather than nearly true.
- **Two toolchains in the lab.** Scenarios that place a host need a Go build available, and the
lab is TypeScript. The binary is built before the scenario runs, not inside it.
- **This is reversible at a cost that will only grow.** It is being taken at the moment the
first line is written, which is the cheapest point it will ever be taken.
## References
- [ADR 0030](0030-the-repository-structure.md) — *the one binary installed by hand*.
- [ADR 0037](0037-the-host-applies-it-does-not-decide.md) — why the host shares no code.
- [ADR 0039](0039-the-link-is-the-security-boundary.md) — why it receives declarations only.
- [`05-the-node-host.md`](../03-DESIGN/01-to-be/05-the-node-host.md) — the design this serves.
@@ -0,0 +1,116 @@
---
topic: what runs on it
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
extends: 0041-events-are-a-relationship.md
---
# 42. The shape of an event on the wire
## Context
[ADR 0041](0041-events-are-a-relationship.md) made events a relationship — `emits`/`consumes`, the
graph, the audit logger. It did not say what an event *is* on the broker: the exchanges, the
routing keys, the headers, the queues and their configuration. That shape is a contract every
emitter and consumer conforms to, exactly as [ADR 0010](0010-delivery.md)
is for declarations — and it was being decided ad-hoc in code. This settles it, so the sdk and the
runtime implement one contract and a module never reinvents it.
## Decision
### Two exchanges, kept apart
- **`mesh.events`** — a durable topic exchange. Every event rides it: module, mesh and node.
- **`mesh.rpc`** — a durable topic exchange. Tool invocations (request/reply) ride it.
Kept separate because RPC is not an event: a `#` subscription on `mesh.events` is then a complete
audit of what happened, with none of the invocation traffic.
### The routing key is the event type, namespaced by origin
Dotted and hierarchical — `<origin>.<name>.<event…>` — with three reserved origins:
- `module.<module>.<event>` — `module.umami.site.created`
- `mesh.<context>.<event>` — `mesh.delivery.deployed`, `mesh.provisioning.granted`
- `node.<node>.<event>` — `node.anchor.joined`, `node.anchor.unreachable`
Topic matching gives a consumer `node.*.joined`, `module.umami.#`, or `#`. The origin roots are
reserved; everything after is the emitter's own namespace.
### Metadata in headers, payload in the body
An event's identity and provenance are AMQP **headers**, so a consumer — or the broker, or an
audit tool — reads who/when/what without parsing the body, and the body is only the domain payload.
**Required headers**
| header | meaning |
|---|---|
| `x-event-id` | a unique id — for dedup and audit (delivery is at-least-once, below) |
| `x-source` | the emitter: the module, context or node name |
| `x-node` | the node it was emitted from |
| `x-time` | emit time, RFC-3339 |
| `content-type` | `application/json` |
**Optional headers**
| header | meaning |
|---|---|
| `x-causation-id` | the event or command that caused this one — tracing |
| `x-schema` | a version of the body's shape, so a body evolves without silent misreads |
The routing key already carries the type; it is not duplicated as a header. An **unknown `x-`
header is ignored, not refused** — unlike a declaration, an event is observed by parties that need
not all understand every header, and refusing would couple every consumer to every emitter's
additions.
### Messages are persistent
Events are published persistent (delivery-mode 2). An audit trail that loses events on a broker
restart is not one, and the cost is disk the broker already spends on everything durable.
### Queues: one per consumer, durable, dead-lettered
- **A consumer's queue** is `<node>.<module>.events`, durable, bound to that module's consumed
patterns. Durable so a restart does not drop what arrived while it was down. **Manual ack** after
the handler succeeds — at-least-once.
- **Prefetch** bounds in-flight work (default 32) so one slow consumer does not pull the whole
backlog into memory.
- **A dead-letter exchange** `mesh.events.dead` receives a message rejected past a redelivery limit,
so a poison event is set aside for inspection rather than looping forever or vanishing silently.
- **The audit logger's queue** `<node>.audit-logger.events`, bound to `#`, is the same shape —
durable, persistent, dead-lettered — because completeness is its whole job.
- **RPC reply queues** are exclusive, auto-delete and server-named; **RPC serve queues**
`serve.<key>` are durable and shared, so several runtimes serving one tool key compete rather than
each answer.
### At-least-once, and consumers are idempotent
A handler may see an event twice — a redelivery after a crash between doing the work and acking.
Consumers must be idempotent, and `x-event-id` is what makes dedup possible. **Exactly-once is not
offered**: it is a promise no broker keeps honestly, and saying so is better than pretending.
## Consequences
- The event shape is a versioned, enforced contract, not conventions each module reinvents. The
sdk's `emit`/`on` and the runtime's AMQP binding implement it; a module never sees an exchange or
queue name.
- Metadata-in-headers means the body is exactly the domain payload, and a consumer that only wants
provenance never parses it.
- Adding a header or an origin root widens the contract and is reviewed as one — the discipline
[ADR 0010](0010-delivery.md) applies to the
declaration vocabulary.
- The sdk's first cut carried source/node/time in the *body*; this supersedes that — they move to
headers. That is code to align, in `mesh-sdk` (`emit`/`on`) and `mesh-tools` (the binding, queue
config, dead-letter).
## References
- [ADR 0041](0041-events-are-a-relationship.md) — events as a relationship; this is their wire shape.
- [ADR 0010](0010-delivery.md) — the precedent: a wire
contract, versioned, additions reviewed as security.
- [ADR 0002](0002-nodes-communicate-over-a-broker.md) — the broker.
- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — `emit`/`on` are stable sdk surface; the
binding, queue config and dead-letter are the runtime's, not the sdk's.
@@ -1,140 +0,0 @@
---
status: accepted
date: 2026-08-26
deciders: jochen
reconstructed: false
extends: 0037-the-host-applies-it-does-not-decide.md
---
# 43. A declaration is an ordered list of resources the host owns
## Context
[`05-the-node-host.md`](../03-DESIGN/01-to-be/05-the-node-host.md) leaves *what a declaration
is* open and calls it the first thing to settle in build. Stage 2 — applying with no mesh
present — cannot start without it.
Three constraints already bind it, and between them they decide most of the shape:
- **Data, not instructions**, with a finite, versioned vocabulary and anything outside it
refused rather than interpreted ([ADR 0039](0039-the-link-is-the-security-boundary.md)).
- **The host applies; it does not decide** ([ADR 0037](0037-the-host-applies-it-does-not-decide.md)).
- **The host depends on nothing** ([ADR 0041](0041-the-host-depends-on-nothing.md)).
## Decision
### JSON, because the host has no dependencies to spend
Go's standard library carries `encoding/json` and no YAML. A YAML declaration would put a
third-party parser inside the one binary whose entire argument is that it needs nothing — to
gain authoring comfort in a document that is, in the ordinary case, generated by a machine and
read by a machine.
The mesh's *authoring* formats stay YAML. What crosses the link is JSON.
### An ordered list, because ordering is a decision
A declaration states the order its resources are applied in. The host does not sort, does not
resolve dependencies, and does not decide what must come before what.
This follows from [ADR 0037](0037-the-host-applies-it-does-not-decide.md) more strictly than it
first appears. A host that derived ordering from declared dependencies would be **deciding**,
and it would be deciding the thing most likely to differ between what the control plane
intended and what the machine does. The control plane knows what depends on what; it says so by
saying when.
Consequence accepted: the control plane must order correctly, and a mis-ordered declaration
fails at the step that needed something not yet there — which is at least the *right* failure,
naming the resource rather than a mystery.
### Every resource has a stable identity
Not a position, not a hash of its content: a name the control plane keeps stable across
declarations. It is what lets the store say *this is the same resource I applied last time*,
which is what makes convergence possible at all.
### Unknown is refused, never skipped
An unknown declaration version, an unknown resource type, or an unknown field is a **refusal of
the whole declaration**. Not a warning, not a skip, not best-effort.
A host that skipped what it did not understand would apply most of a declaration and report
success — a node that looks configured and is not, which is
[04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) with the
declaration on the other side of the wire. Refusing whole also means an older host cannot be
handed a newer vocabulary and quietly do half of it.
### Complete for what the host owns, and only that
*Desired state* invites the question of removal, and the honest answer needs a boundary.
**The host removes what it previously applied and is no longer declared.** It knows what it
applied because it recorded it (`store`), so this is a fact it holds rather than an inference.
**The host never removes anything it did not create.** A machine has things on it that the mesh
did not put there, and a converger that treats *not declared* as *must not exist* deletes them.
The rule that prevents production data loss elsewhere in this repository is the same one:
[ADR 0018](0018-the-mesh-creates-no-symlinks.md) exists because a tool did something to a path
it did not own.
So: authoritative over its own footprint, inert everywhere else.
### Addressed, and checked when it can be
A declaration names who it is for. A host that has an identity refuses one addressed elsewhere.
A host that has no identity yet — the first node, applying the bundle it carries — has nothing
to check against and applies it.
## Where the list comes from
This record specifies what the host **accepts**. What produces a declaration is deliberately
not settled here, and the reason is worth stating rather than leaving as an omission.
**Today, and at stage 2: by hand.** `substrate.lock` is authored and pinned — a person writes
the resources and writes the order. That is the first node's path, where there is no control
plane to derive anything from.
**Afterwards: the control plane derives it**, from three things it already holds — which
modules are assigned to this node, what those modules' configuration resolves to, and what each
module declares it needs.
**And the order comes from the graph.** Each module expands to resources; the modules are
ordered by their declared dependencies on one another. That is
[research 011](../01-RESEARCH/011-the-module-graph/00-overview.md) — `requires`, `provides`,
`excludes` — and a declaration is the graph's output, flattened for one node.
So this record is complete on the consumer side and silent on the producer side, because the
producer does not exist and its shape is what 011 is investigating. The consumer can be settled
first because the host must refuse what it does not understand whoever wrote it.
**What this means for ordering.** [ADR 0037](0037-the-host-applies-it-does-not-decide.md) puts
the ordering decision in the control plane; 011 decides how the control plane makes it. If the
graph turns out not to determine a total order, that is 011's problem to solve and not the
host's — the host will still be handed a list, and will still apply it as given.
## Consequences
- **Ordering is now a control-plane responsibility**, and getting it wrong is a class of bug
that will appear. It is the correct place for it: the alternative puts a dependency solver in
tier 0 and a decision in the wrong tier.
- **The vocabulary is a security artefact.** Every type added widens what a compromised control
plane can express, so additions are reviewed as such rather than as features.
- **Removal is bounded but not free.** A resource dropped from a declaration is deleted on the
next apply, so removing a line is an act with an effect — which is the point, and is worth
saying out loud because it does not look like one.
- **The store becomes load-bearing at stage 2**, earlier than the build order suggests. Nothing
can be removed without knowing what was applied, so the record of applied resources arrives
with the first apply rather than with the link.
- **A closed address space bounds what the first types can be.** A scenario has no route to a
package repository, so a declaration whose resources must be fetched cannot be applied in the
lab at all. The first vocabulary is therefore what needs no network — files, directories,
service state — and packages and containers wait on *where `place:` gets its artifacts from*,
which is open in
[`02-scenario-declaration.md`](../03-DESIGN/01-to-be/02-scenario-declaration.md).
## References
- [ADR 0037](0037-the-host-applies-it-does-not-decide.md) — why ordering is not the host's.
- [ADR 0039](0039-the-link-is-the-security-boundary.md) — bounded by form.
- [ADR 0041](0041-the-host-depends-on-nothing.md) — why JSON.
- [ADR 0008](0008-a-failed-step-fails-the-job.md) — why a refusal is whole.
@@ -0,0 +1,101 @@
---
topic: what runs on it
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0041-events-are-a-relationship.md
---
# 43. A module's broker account is scoped by what it emits and consumes
## Context
[ADR 0041](0041-events-are-a-relationship.md) made events a relationship — `emits` and `consumes`
on the manifest. [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) gave them a wire shape — the
`mesh.events` exchange, the durable per-consumer queue, the reserved routing-key origins. Neither
said how a module *reaches* the broker: what account it holds, and what that account is allowed to
do.
As the code stands, there is no answer. The mesh can provision a **node** account (at enrolment)
and a **builder** account (scoped to the build queue), and it can *deliver* any module a sealed
own-secret at a declared path — but it has no way to provision a broker **account** for a general
module. A module that declares `own-secrets: {broker: …}` and nothing more receives thirty-two
random bytes, not a credential. So on the broker, `emits` and `consumes` are enforced by nothing: a
running module could bind any queue, consume any pattern, and publish under any origin, and the
manifest that says otherwise would be describing a boundary no code draws — the exact shape of fault
[04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) records, a scope
declared in manifests and read by nothing.
This settles it, so a module's place on the bus is a thing the broker enforces rather than a thing
the manifest merely claims.
## Decision
### A module gets a broker account when it is assigned, and its permissions are the manifest
When the mesh assigns a module to a node it provisions a broker account for that module on that node,
sealed to the node ([ADR 0004](0004-a-node-and-how-it-joins.md)) and delivered as the
module's `own-secrets` broker — `amqps://` with the mesh's fingerprint, the shape
[ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) already carries. The account's permissions are
derived from the manifest, and are exactly these:
- **What it consumes.** Read on `mesh.events`, and configure-and-read on its own queue
`<node>.<module>.events` bound to the patterns in `consumes`. It cannot bind or read another
module's queue. A module that consumes nothing gets no read on the events exchange at all.
- **What it emits.** Write to `mesh.events`, restricted to routing keys under its own origin,
`module.<name>.*`. It cannot publish as another module, and cannot publish under the reserved
`mesh.*` or `node.*` origins — those belong to the mesh and the host (ADR 0042). A module that
emits nothing gets no write.
- **Nothing else.** The events account reaches `mesh.events` and that module's own queue, and no
more. Tool serving and calling over `mesh.rpc` is a separate grant on the same principle — a
module serves the tool keys it declares and calls the ones it is bound to — and is scoped the same
way rather than folded in here.
### Consuming everything is a privilege, granted deliberately
`consumes: ["#"]` — the audit logger — is read across the whole bus: every module's events, the
mesh's, every node's. That is not a pattern like any other; it is the power to see everything, and
the account is where it becomes visible. The grant that lets one module read the entire bus is one
the mesh issues on purpose and can be audited — the answer to *who can read everything* is a row, not
a guess — rather than a breadth any manifest acquires by typing a single character. A `#` consume is
a reviewed grant, not a default one.
### The account is how the declaration is enforced
Because the account can do only what `emits` and `consumes` name, the broker itself refuses a module
that tries to consume a queue it did not declare or emit under an origin it does not own. That is what
makes an event relationship a rule and not a comment — the discipline that a stated rule says how it
is checked. A manifest that over-declares grants more than the module uses, which is visible and
reviewable; one that under-declares makes the module fail closed at the broker, which is the safe
direction to be wrong in.
## Consequences
- The control plane gains a **generic module broker-account**, derived from the manifest. The
builder stops being a special case: its access to the build queue becomes an ordinary expression of
what it consumes and serves, not a bespoke account method. One rule, and the builder is an instance
of it.
- The runtime reads its credential from a file (the broker own-secret), `amqps://` verified against
the mesh's fingerprint. The `guest` account is for raising the substrate, never for a module — a
module documented as holding its own credential and handed the broker's administrative one is worse
than one with no credential story at all.
- `emits` and `consumes` stop being advisory. They are the module's authority on the bus, so the
manifest is now a security boundary and is reviewed as one, the discipline
[ADR 0010](0010-delivery.md) applies to the declaration
vocabulary.
- *Who can read the whole bus* becomes an answerable question, because `#` is a grant and not an
accident.
## References
- [ADR 0041](0041-events-are-a-relationship.md) — events are a relationship; this scopes the account
by that relationship.
- [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) — the wire this account secures: the queue,
the origins, the `amqps` credential shape.
- [ADR 0004](0004-a-node-and-how-it-joins.md) — the link is the security boundary; a
module's account is sealed to its node the same way a node's is.
- [ADR 0010](0010-delivery.md) — a declaration is owned
and its additions reviewed; a module's broker permissions are that discipline applied to the bus.
- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — a scope declared
in manifests and enforced by no code: the fault this decision closes for events.
@@ -0,0 +1,94 @@
---
topic: what runs on it
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0027-a-provision-names-what-the-consumer-is-coupled-to.md
---
# 44. A public name is provisioned, not registered by hand
## Context
The mesh names and resolves its own machines internally: the overlay generates
`<service>.<node>.<suffix>` wildcards, dnsmasq answers them (`wildcard-resolution`), and the mesh
issues a certificate for each internal name. A service reachable at a *public* domain —
`plex.example.com`, not `plex.anchor.internal` — needs three things that machinery does not give it:
- a **public DNS record** at a registrar or DNS provider, so the name resolves on the internet;
- a **publicly-trusted certificate** for it, because the mesh's own authority is trusted by nobody
outside the mesh;
- and routing from that name to the module — which the reverse proxy already does: a module
`requires` the `route` capability and the proxy provides it, routing by the host it was asked for.
The routing exists. The public DNS record does not: the mesh has no way to make a name resolve on
the public internet, so today that is a step someone does by hand at a DNS provider, outside the
mesh, remembered nowhere. A public name is therefore the one part of reaching a service that the
declaration graph cannot grant or withdraw — which means it is created once and outlives whatever it
was for, the shape of drift this project exists to remove.
## Decision
### A public name is a capability, requested like any other
A module reachable at a public host declares `requires: ["public-dns"]` and contributes the hostname
it wants — beside `requires: ["route"]`, which exposes it through the proxy. The name is then
provisioned on declaration ([ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md)): created
when the module is assigned, removed when it is withdrawn, reconciled like every provision.
### The interface is neutral; the providers are the registrars
`public-dns` is drawn at the consumer's coupling: the consumer wants *a public name that resolves to
me*, and does not care whether Cloudflare, Route 53 or a registrar's own API puts the record there.
So the interface is neutral and the providers are provider-scoped — `cloudflare-dns`,
`route53-dns`, `porkbun-dns` — each implementing the one `public-dns` contract, the same way a
neutral database coupling is answered by `postgres-database` and `mssql-database`. A module names
`public-dns`; it never names a registrar.
### The record points at the mesh's public ingress, not at the node
What the name resolves to is the address the reverse proxy answers on, not the consuming machine's.
A public service is reachable only *through* the proxy — the proxy holds the `route` grant and routes
by host to the module — so the public name must resolve to the proxy. `public-dns` and `route` are
the two halves of one public exposure: the name, and what the name reaches.
### The record is a fact, not a secret
A DNS record is public by definition, so the grant returns the fully-qualified name and its TTL and
nothing sealed. The only secret is the provider's own API credential, which is the provider module's
own-secret and never leaves it — the module that wanted the name never sees it.
### Events
The provider emits `module.<provider>.record.created` and `module.<provider>.record.removed`
([ADR 0041](0041-events-are-a-relationship.md)), so *which names the mesh publishes, and where* is a
question answered from the event trail and the grants, not from a folder of records edited at a
provider.
### The public certificate is the proxy's, and is named here only to pair it
A public name without a publicly-trusted certificate is reachable and not trusted — the same pairing
the internal name and the mesh-issued certificate already have. Obtaining that certificate (ACME
against the now-resolving public name) is the reverse proxy's to do, and its mechanism is its own
decision; it is named here so the pairing is not forgotten, not resolved here.
## Consequences
- A public name is created and torn down with the module, so it cannot outlive it, and the mesh can
say which public names it publishes without anyone reading a registrar's dashboard.
- Adding a registrar is adding a provider that answers `public-dns`; the modules that want names do
not change.
- Public exposure of a service is a trio of separate, declared, enforced relationships: the firewall
opens the proxy's public port ([ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)),
`route` routes the host to the module, and `public-dns` makes the host resolve.
## References
- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — a capability is provisioned on
declaration; a public name is one.
- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the firewall, the other
half of the reachability question this was asked with.
- [ADR 0041](0041-events-are-a-relationship.md) — the provider's record events.
- [ADR 0040](0040-what-a-module-is.md) — a provider and its interface; the neutral-interface,
scoped-provider naming this follows.
@@ -0,0 +1,93 @@
---
topic: what runs on it
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0005-the-node-host.md
---
# 45. A machine's firewall is the sum of what its modules listen on
## Context
The reverse proxy is a *provider*: a module `requires` the `route` capability and a running proxy
provides it, routing traffic by name and reaching back to the consumer. A fair question follows —
is the firewall the same shape? Should a module *register* a port with a firewall provider the way
it requests a route?
It should not, and the difference is the point. A reverse proxy is a service another component
performs; a firewall is a property of the machine — a packet filter the host applies to itself.
Modelling it as a provider would invent a credential and a reach-back for something that has neither.
And the mesh already has the registration: a module declares `listens: [{ port, from }]` — the port
it accepts connections on, and from where. That *is* how a service says it wants a port open. What is
missing is not a model but enforcement. [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md)
records that a `scope:` key five manifests carry is read by no code: a manifest can appear to
restrict a port and restrict nothing — the exact fault
[how-we-build.md](../00-META/how-we-build.md) names, *an unenforced rule is indistinguishable from a
wrong one*, made worse because the declaration reads as a restriction.
## Decision
### The firewall is derived and host-applied, not a provider
A machine's firewall is the sum of what the modules assigned to it declare they listen on, computed
by the host and applied as one of its owned resources ([ADR 0005](0005-the-node-host.md):
the host applies, it does not decide; [ADR 0010](0010-delivery.md):
the declaration is owned resources). It is not a capability, not a per-consumer grant — opening a
port is a declarative fact about a machine, so it is computed and applied, not requested and
credentialed.
### `from` is the whole of public-versus-internal
The distinction the question is really about lives in `from`:
- `listens: [{ port: 5432, from: mesh }]` — open to the private overlay only.
- `listens: [{ port: 443, from: anywhere }]` — open to the public internet.
A module registers a port on the firewall by listening on it and saying from where. There is no
separate firewall capability, because the firewall is not a thing that reaches back or holds a
secret; it is the machine's own filter over the ports its modules named.
### The host enforces it both ways, and unknown keys are refused
A port a module listens on is opened to exactly the scope it named; a port nothing declares is
closed. And a key the firewall does not read — the `scope:` of issue 003 — is refused at the
manifest, not accepted and ignored, so a declaration that reads as a restriction is one. This is the
discipline [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) applied to the
broker account, applied here to the packet filter: the declaration is the enforcement, or it is a
comment.
### A public service is exposed through the proxy, not by opening its own port
Reaching the public internet is normally not `from: anywhere` on the service's own port. The service
listens `from: mesh` — only the proxy reaches it — and `requires: route`, so the sole machine with a
public opening is the one running the reverse proxy, and the service is exposed by name through it.
`from: anywhere` is the deliberate direct-exposure case, for a service that is its own front door.
## Consequences
- Issue 003 is closed: the firewall is computed from `listens` and enforced, so a declared scope is
real and an undeclared port is shut. Rejecting unknown manifest keys is the general fix, of which
the `scope:` key was one instance.
- The firewall and the reverse proxy stop being confused for one model: the firewall is the machine's
filter (host-derived from `listens.from`); `route` is a name-router (a provider); the public DNS
name is a third thing ([ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md)). A
public service uses all three.
- The modelling question is answered: a module registers a port by declaring `listens`, and reaches
the public internet by name through `route` + `public-dns` — never by the firewall being a
provider.
## References
- [ADR 0005](0005-the-node-host.md) — the host applies; the firewall is one of
the things it applies.
- [ADR 0010](0010-delivery.md) — the firewall is a derived
owned resource, not a grant.
- [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) — the same discipline:
a declaration is enforced, or it is a comment.
- [ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md) — the public name, the other
half of the reachability question this was asked with.
- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — the unenforced
`scope:` this closes.
@@ -0,0 +1,88 @@
---
topic: what runs on it
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0027-a-provision-names-what-the-consumer-is-coupled-to.md
---
# 46. A module's configuration is its assignment's, not its manifest's
## Context
A module is assigned to a node — `assign <node> <module>`, always to a machine; there is no
assignment to the mesh. "Mesh" is a *scope*, not a place: a `provides` or a `claim` scoped `mesh`
reaches the whole mesh, but the module still runs on a node. So the two kinds of thing a module can
carry are the manifest (what the module *is*) and, separately, what it should do *here* — which
differs by deployment and by node.
The mesh already has the second: **settings**. `settings set <module> [--node <node>]` — with a node
it is that machine's, without it the whole mesh's — layered over what the module declares and applied
at resolution, changeable without editing the module and without a rebuild. That is the surface a
meshboard would edit.
But settings today reach only a module's **config-file content** (a mergeable file the module owns).
Configuration that is not a file has been landing in the manifest instead, statically — a registrar's
zone and domain, the address public names point at, and, most sharply, `listens.from`. That last one
is the tell: whether a port is open to the private overlay or to the public internet is a
*per-node deployment choice* — the same database internal on one machine and public on another — and
a value fixed in the manifest is one value for every machine, so it cannot be. Static configuration in
the manifest is configuration in the wrong place: it cannot vary per node, and it cannot change
without a new module version.
## Decision
### The manifest is identity and defaults; the assignment's settings are the configuration
A module's manifest declares what it is — what it provides, requires and claims, the shape of its
resources — and, for anything configurable, a **default**. The values that make a running instance
*this* instance are settings, carried by the assignment: per-node, or mesh-wide when no node is named,
applied over the defaults at resolution. Change one and the next reconcile carries it; nothing is
edited on a machine and nothing is rebuilt.
### Settings drive the configurable fields the manifest marks, not only file content
Settings extend beyond a config file's content to the manifest fields a module declares settable —
foremost:
- **`listens.from`**: a module declares its safe default (`from: mesh`), and a per-node setting
raises or lowers it. postgres declares `listens: [{ port: 5432, from: mesh }]`; on the machine that
should expose it, a setting makes that port `from: anywhere`. Same module, different exposure, and
the firewall ([ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)) is computed
from the effective value, so the packet filter follows the setting.
- **A provider's own configuration**: a registrar's zone, domain and the ingress its names point at
([ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md)) are mesh-wide settings, not
manifest constants — one mesh's Cloudflare zone is not another's, and the module description is the
same for both.
### Unset is the default, and an unknown setting is refused
A field with no setting keeps the manifest's default, so a module runs correctly configured by nobody.
A setting that matches no settable field — like a config value that reaches no file today — is named,
not silently dropped, so a misspelled setting is found rather than believed (the discipline of
`UnusedSettings`, and of [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md):
a declaration is enforced or it is a comment).
## Consequences
- The postgres case works: one module, `from: mesh` by default, `from: anywhere` where a setting says
so — internal on ace, public on novox, changeable live.
- Provider modules stop carrying a mesh's specifics: `cloudflare-dns` describes *a Cloudflare
registrar*, and *which* zone and ingress is a setting, so the same module serves every mesh.
- Configuration becomes a thing a meshboard manages — set per node or mesh-wide, applied on the next
reconcile — rather than a manifest edit and a rebuild ([ADR 0011](0011-managed-files-are-generated-never-edited.md):
the way you change a managed thing is not by editing it).
- What a manifest may not do is grow a value that differs per machine; if it differs per machine it is
a setting, and the manifest holds only the default.
## References
- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — what is provisioned on
declaration; its per-instance values are the assignment's.
- [ADR 0011](0011-managed-files-are-generated-never-edited.md) — a managed thing is changed through
the mesh, not by editing it; settings are that, for configuration.
- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the firewall follows the
effective `listens.from`, so making `from` a setting makes exposure a setting.
- [ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md) — the provider whose zone and
ingress are settings, not manifest constants.
@@ -0,0 +1,87 @@
---
topic: what runs on it
status: accepted
date: 2026-09-04
deciders: jochen
reconstructed: false
extends: 0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md
---
# 47. A module runs its code as its own process, with its own account
## Context
A module is one self-contained thing ([ADR 0040](0040-what-a-module-is.md)), and it gets a broker
account scoped to what it emits and consumes ([ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md)).
The catalogue now gives modules **tools** and **events** — real code, in the module ([ADR 0039](0039-what-the-sdk-holds-and-refuses.md)) —
but nothing has said what *runs* that code. The audit-logger showed one shape and was treated as an
exception: a container running the tool runtime carrying the module's compiled code, holding the
module's own scoped account. Every module with tools or events needs the same, and the tempting
alternative does not work.
**A node-wide runtime that loaded every assigned module's code cannot hold a per-module account.** It
would run under one account with the union of every module's permissions — able to emit as any of
them and read any of their queues — which is exactly the isolation ADR 0043 exists to draw. So the
runtime is per-module, not per-node, and treating the audit-logger as special left the other
modules' code with nothing to run it: the conversion produced tools and events that, as it stands,
never execute.
## Decision
### A module with tools or events runs a process of its own
A module that has tools or events runs a **runtime process** — a container, the tool runtime carrying
that module's compiled code — assigned and started like the module it is, holding the single broker
account the mesh scoped to it (ADR 0043). One module, one process, one account.
### It serves its tools, each on its own key
A tool is served on its own key (`serve.<tool>`), and a caller invokes a named tool. Only the module
that serves it answers, and the module's account is scoped to exactly its tool keys — so one module
cannot answer another's calls, the isolation ADR 0043 gives events extended to tools. This supersedes
a single `tools.invoke` endpoint that dispatched by name: that shape assumed one runtime for the
whole node, and per-module runtimes competing on one key would each be handed calls for tools they do
not have.
### It runs its events in the same process, under the same account
Emitting under the module's own origin and consuming its own queue ([ADR 0042](0042-the-shape-of-an-event-on-the-wire.md))
happen in that same process, with that same account — not a second one to scope and seal. A module's
tool code, its event code and, for a provider, its provisioner are the one module's code and run as
the one module's process.
### The runtime image is the tool runtime plus the module's code
Built from the module's source like any module image — the audit-logger's shape, made the rule, not
the exception. The module declares a `container` for it carrying `MESH_BROKER_FILE` (its sealed
credential, ADR 0043) and its compiled code. A module with **neither** tools nor events runs no such
process: a plain service module — the plex *server*, dnsmasq the resolver — is its service and files
and nothing more. A module that is both a service and code declares both containers: the service, and
the runtime beside it.
## Consequences
- The catalogue's tools and events become runnable: each tools-or-events module gains a runtime
container with its scoped credential, and the audit-logger stops being special. Until this, the
converted modules held code with nothing to execute it.
- A process, and a small image, per tools-or-events module. That is the cost of ADR 0043's isolation:
one account per module means one process per module. It is paid deliberately — a shared runtime is
cheaper and cannot be scoped, and a mesh where any module can emit as any other is not one worth the
saving.
- `serve.<tool>` per key replaces the single `tools.invoke` dispatch. The sdk's serving and a module's
account scope both come to name tools individually.
- **A provider's provisioner is a runtime process too.** It already runs as its own container; its
events (`bucket.created`, `database.provisioned`) belong to *that* process and need the same
credential. So a provisioner that emits carries `MESH_BROKER_FILE` and its scoped account like any
runtime — or it does not emit. (This is the fix for provisioners that emit today with no broker
bound: the emit is a runtime's, and the provisioner is a runtime.)
## References
- [ADR 0040](0040-what-a-module-is.md) — a module is one self-contained thing; its code runs as one
process.
- [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) — the scoped account
this process holds, and the isolation that makes it per-module.
- [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) — the events this process runs, and the
`serve.<key>` queue tools now use.
- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — the code lives in the module; this runs it.
@@ -0,0 +1,131 @@
---
topic: what runs on it
status: accepted
date: 2026-09-05
deciders: jochen
reconstructed: false
---
# 48. A provider creates the credential the mesh minted, and seals nothing
## Context
A provider module stands up a per-consumer resource — a database, a cache bucket, an object
store user — and the consumer must end up holding a credential that authenticates against it.
Building the module runtime (ADR 0047: a module runs its own code as its own process under its
own account), the provider's provisioner was run for the first time as a delivered thing, and
it did not work. It reads a seal key from the environment that nothing sets, and it seals every
credential it produces to that key with a symmetric passphrase.
Tracing the credential's path turned up something larger than a missing key. **The provisioner
harness the whole catalogue is built on describes a credential flow the mesh does not have, and
duplicates — incorrectly — one it does.**
What the sdk's `runProvisioner` does today:
- reads request files named `*.grant.json` — which nothing in the mesh writes;
- calls an adapter whose `create` **generates its own password** and returns it;
- seals that password with a symmetric key (`$MESH_SEAL_KEY`) and writes a `*.credential`
file — which nothing in the mesh reads, and no consumer ever unseals.
What the mesh already does, and has wired end to end:
- The control plane mints one password per (consumer, provider) pair (`Inventory.SecretFor` →
`secrets.Make`) and seals it to **both** node keys asymmetrically — a copy the consumer's
host can open and a copy the provider's host can open. No shared symmetric key exists
anywhere, on purpose: a key both ends hold is a key the mesh would have to distribute, which
is the same problem one level down, and the control plane's own code refuses it.
- The provider is handed, at the path its `receives` names, one contribution per consumer:
the **login to create** (`As`, derived by the mesh so the two ends agree by construction),
the consumer's address and requested values, and a **`Secret` file** holding that consumer's
password sealed to the provider and unsealed onto the machine by its host.
- The consumer is handed the *same* password, as plaintext its own host wrote by unsealing its
copy and substituting it into a config file. The consumer never unseals anything itself and
holds no key.
So the password a provider's provisioner invents is not even the password the consumer was
given: a consumer authenticating with the mesh's password against a resource the provisioner
created with its own would simply fail. The symmetric seal is not an incomplete feature to
finish delivering a key for. It is a second, contradictory credential model bolted beside the
real one, and it cannot be made to work without building the very thing the mesh was designed
not to have.
This is a decision and not a patch because the harness is the **provider contract**. Every
provider — the four that exist and the many a real mesh grows — is built on
`runProvisioner(resource, adapter)`. Whatever it says a provider is, they all inherit; and
changing it later is one migration per provider. It is cheaper and more honest to settle what
a provider is now.
## Decision
**A provider is handed the credential; it does not make one, does not seal one, and does not
hand one back.** The provisioner's only job is to make the mesh's grants true in its own
software.
Concretely, for the sdk harness and the adapter contract:
- The harness reconciles the **contributions the mesh delivers** to the provider's `receives`
path — the list of consumers, each with its login name (`As`), address, requested values,
and the path to its unsealed password (`Secret`). It does not read `*.grant.json` and it
does not write `*.credential`.
- For each consumer present, the harness reads the password from that consumer's `Secret` file
and calls the adapter to bring the resource into being under the given login. For each
consumer no longer present — the mesh drops it from the contributions file when its consumer
goes away — the harness calls the adapter to withdraw it.
- The adapter shrinks to the per-software half and nothing else. It is given the login, the
password, and the values, and it makes the resource exist or removes it. It generates no
password, derives no name, seals nothing, and returns no credential:
roughly `create({ as, password, values })` and `remove({ as })`, both returning nothing.
- `$MESH_SEAL_KEY`, the symmetric `seal()`/`writeSealedCredential` path, and the `*.grant.json`
/ `*.credential` files are removed from the provisioning path entirely. The credential
reaches the consumer through the mesh's own asymmetric channel, which already crosses node
boundaries and holds no shared secret.
Identity stays the mesh's to say. The login the provider creates is the name the mesh derived
and gave the consumer to present; the provider never invents a name, because a name the
consumer cannot learn is a name it cannot authenticate with.
## Consequences
- A provider module becomes smaller and unable to be wrong in this way: with no password to
generate and no key to seal to, the class of bug where the two ends hold different secrets
cannot be written. A provider added after this inherits the corrected contract and has no
seal to reintroduce.
- The four current providers (redis, postgres, minio, umami) each lose their `generatePassword`
+ seal code and gain a `create` that takes the password it is given. Their teardown becomes
"withdraw the login named `As`".
- The symmetric `seal()`/`unseal()` primitive loses its only caller and leaves — checked, not
assumed: nothing else in the sdk or the catalogue called it, so it is removed with the
provisioner it belonged to.
- **How this is verified:** redis is assigned as a provider in the lab, the contributions and
the unsealed password the mesh would deliver are put in its `receives` path, and a client
authenticates as that consumer with the mesh's password and gets PONG — where a provider that
invented its own password answers WRONGPASS — with `$MESH_SEAL_KEY` set nowhere and no
`.credential` file written. Proven: `provider-uses-mesh-credential` is green.
**What this does not cover — credential provisions, not data provisions.** This decision is about a
provision whose credential is a *secret the mesh mints* — a login and password (redis, postgres,
minio). A provider that instead *generates* the thing the consumer needs, and that thing is not a
secret — umami's `analytics`, where the consumer wants back a `siteId` umami assigned — does not fit,
because a contract that returns nothing has no way to hand that data back. The seal-key fault was
never umami's (it sealed no password; it returned a public id), so removing the seal does not break
it further, and it still reconciles its sites off the mesh's contributions. But delivering
provider-generated data back to a consumer is a *return path* the mesh does not have and this
decision does not build — a separate shape, left to a separate decision.
- Teardown beyond "remove the login" — data an object store leaves behind when a consumer
leaves — is named by each provider's adapter, not by the harness, and is out of scope here
except to say the contract must leave room for it.
## References
- [04-ISSUES/032](../04-ISSUES/032-provider-runtime-has-no-seal-key/00-report.md) — the
observation and the cross-repo trace this decision rests on.
- ADR 0047 (the module runtime) — what first ran a provider's provisioner as a delivered
process and exposed this; link to be filled when 0047 lands on the trunk.
- Control-plane mechanisms this relies on already existing: `mesh-control` —
`internal/inventory/secrets.go` (`SecretFor`, `SecretsFrom`), `internal/secrets/seal.go`
(`Make`, the two-blob asymmetric sealing), `cmd/mesh-control/plan.go` (`grantsFor`, the
`Grant.Sealed = ForProvider` delivery), `internal/catalogue/declaration.go` (the `receives`
contribution: `As`, `At`, `Values`, `Secret`).
- The path being removed: `mesh-sdk` — `src/provisioner/index.ts` (`runProvisioner`, `sealKey`,
`writeSealedCredential`) and the symmetric `src/primitives/index.ts` `seal()`/`unseal()`.
@@ -0,0 +1,130 @@
---
topic: what runs on it
status: accepted
date: 2026-09-05
deciders: jochen
reconstructed: false
---
# 49. A consumer's identity is bounded by the tightest backend that must accept it
## Context
The mesh says who a consumer is, once, and hands the same name to the provider (to create) and the
consumer (to present), so the two ends agree by construction rather than by two conventions (the
principle behind `ConsumerIdentity`, 04-ISSUES/023). The name is `mesh_<node>_<module>`, cleaned to
lower-case letters, digits and underscore.
Proving the provider contract per backend (ADR 0048) turned up 04-ISSUES/034: redis and postgres
create that name verbatim, but **minio refuses it** — an S3 access key is capped at 20 characters,
and `mesh_anchor_bucketuser` is 22. The provisioner then retries for ever, per consumer, and the
consumer holding that same too-long name could never present it either.
Two things about the existing derivation decide most of this:
- **The charset is already right.** `[^a-z0-9_]` is deliberately conservative, and its own comment
says it reaches "a PostgreSQL role, a MinIO access key, an LDAP uid and a Keycloak client without
quoting." That much is true.
- **The length is wrong.** `CheckIdentity` refuses names over `identityLimit = 63`, commented as
"the shortest identifier limit among the systems these names reach: PostgreSQL's". It is not the
shortest — S3's 20 is shorter — so the guard that was meant to catch exactly this lets it through,
and the failure lands at provision time as a silent retry instead of at assignment as a refusal.
So this is a small wrong constant with a real cost attached: whatever bound we set, `mesh_` (5) plus
a node name plus `_` plus a module name has to fit inside it.
## The options
**A — Bound the identity by the true minimum, and refuse early.** Lower `identityLimit` to the real
shortest (20, S3's), so `CheckIdentity` refuses an over-long name *at assignment* with a clear
message, the way it already refuses over-63 names. The derivation does not change; long names are
simply rejected before anything is provisioned.
- *For:* smallest change; keeps "the mesh says the identity once, verbatim" intact; the failure
moves from a per-consumer provision-time retry to an up-front, legible refusal — which is what
`CheckIdentity` exists to do.
- *Against:* a hard budget. `mesh_` + node + `_` + module ≤ 20 means node + module ≤ 14 characters.
`anchor` + `bucketuser` (16) is already over. It pushes the constraint onto how machines and
modules are named, which is a real limitation on legible names.
**B — Keep the readable name when it fits, compact it when it does not.** Below the bound, the name
is `mesh_<node>_<module>` as today; over it, the mesh substitutes a deterministic short form (e.g.
`mesh_` + a truncated hash of node+module) — still one derivation, so both ends still agree.
- *For:* no naming constraint; short backends always satisfied; the common case stays legible.
- *Against:* some identities become opaque, and a provisioner tracing "whose login is this" loses
the answer for exactly the consumers that overflowed. The mesh now owns a fallback format and its
collision properties (a truncated hash is not free of collisions at 15 characters).
**C — Let each interface declare its identifier bounds, and derive within the tightest a consumer
reaches.** `s3-bucket` states `identifier: { max: 20 }`; `postgres-database` states 63; the mesh
derives a name that fits the **minimum** bound across the providers a given consumer is granted.
- *For:* the most precise — each provision gets exactly the room it has, and a database consumer
keeps long legible names while an S3 consumer gets a short one; the constraint lives where the
fact does (on the interface).
- *Against:* the most work, and a consumer of two interfaces with different bounds must satisfy the
smaller — so its name shortens for both, reintroducing B's opacity in a narrower case. It also
means one consumer can hold **different** identities per provision, which the "said once" model
currently forbids.
**D — Let the provider generate a backend-valid identity and hand it back (rejected).** minio mints
its own access key and returns it to the consumer. This is the data-provision return path this era
keeps meeting — but it directly contradicts 023 and ADR 0048: the identity would no longer be the
mesh's single derivation the two ends share, it would be a value one side invents and the other must
be told. Listed for completeness; not recommended.
**E — A module (and a node) may declare a short slug; the identity is built from it.** The identity
becomes `mesh_<node-slug|node-name>_<module-slug|module-name>`: where a slug is declared it is used,
otherwise the cleaned name. A slug is a deliberately short, operator-chosen identifier — `kc` for
keycloak, `wkstn` for a workstation. It is optional: short names (`anchor`, `redis`) need none.
- *For:* this is the escape hatch B wanted to be, without the opacity. The name stays legible — a
provisioner can read `mesh_wkstn_kc` and know who is asking — because a person chose it, not a
hash function. And it makes an early refusal *palatable*: if even the slug-built identity overflows,
the refusal points at the slug, a field made for exactly this, rather than at the machine's name.
Both ends still derive it from one declared thing, so they agree by construction.
- *Against:* a new optional manifest field, and someone must pick the slug — but only for names that
would otherwise overflow, and picking a short legible identifier is a better job than being handed
a hash.
## What implementing A revealed
A was tried first. At `identityLimit = 20`, the readable budget is `mesh_` (5) + node + `_` + module
≤ 20, i.e. **node + module ≤ 14 characters** — far tighter than it looked. The catalogue's own
existing tests use `workstation`+`keycloak` (25), which compacts to `mesh_dbbc02f8dde34d3`; common
mesh names (`home-server`, `the-build-node`, `workstation`) blow the budget with any module. So B's
compact fallback would fire for the *common* case, not the rare overflow — which inverts A+B: most
identities would be opaque hashes. A alone (hard refusal at 20) would refuse most realistic names.
This is what moved the recommendation to E: the problem is not the limit, it is that the *readable
name* is the wrong source when it is long, and a slug is a better source than either a hash or a ban.
## Recommendation
**E, over a per-consumer bound (start with the global minimum, 20).** Build the identity from an
optional slug, keep it when it fits, and refuse at assignment with "declare or shorten `<module>`'s
slug" when it does not — no hash, no lost legibility, and the fix is a first-class field. Set the
bound to the true minimum (20) now; it needs no per-interface machinery to unblock S3, and a module
that consumes S3 simply declares a short slug. Graduate to **C** (per-interface bounds) later if it
turns out that non-S3 consumers are paying for S3's limit often enough to mind — E and C compose:
slugs are the mechanism, per-interface bounds refine where the ceiling sits. **B is dropped**: a
declared slug is a strictly better escape hatch than an opaque hash. **D stays rejected.**
## Consequences (of E)
- A module manifest gains an optional `slug`; a node may carry one too. `ConsumerIdentity` prefers
the slug over the cleaned name for each half. `identityLimit` becomes 20 (the true minimum), and
`CheckIdentity` refuses at `module add` / assignment — now with a message naming the slug to set.
- The common case stays legible; only names that overflow the budget need a slug, and what they get
is a name a person chose, not a hash.
- Existing modules/nodes whose names overflow declare a slug once — a migration cost paid as a clear
refusal with an obvious remedy, not a silent hash or a silent truncation.
- minio (04-ISSUES/034) is unblocked: an S3 consumer declares a short slug and its access key fits.
- **How it is checked:** the minio grant e2e — a consumer whose (slugged) identity fits reaches its
bucket with the credential the mesh delivered — plus unit tests that a slug is preferred, that an
un-sluggable over-long identity is refused (naming the slug), and that two consumers never collide.
## References
- [04-ISSUES/034](../04-ISSUES/034-mesh-login-exceeds-s3-access-key-limit/00-report.md) — the
observation.
- ADR 0048 — a provider creates the credential the mesh minted; the identity it creates it under is
the one this decision bounds.
- `mesh-control` `internal/catalogue/identity.go` — `ConsumerIdentity`, `identityUnusable`,
`identityLimit`, `CheckIdentity` — where the constant and the check live.
@@ -0,0 +1,206 @@
---
topic: what runs on it
status: accepted
date: 2026-09-05
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0024-model-access-is-a-provision.md
---
# 50. Model access is vendor-agnostic, and a vendor is an adapter
## Context
[ADR 0024](0024-model-access-is-a-provision.md) settled that model access is a provision and that
a licence is a named thing an operator uses. What shipped, and runs, is a single vendor: the mesh's
"claude" feature. A read-only trace of that feature (2026-09-05, in the code workspace) was made to
answer whether the model-access provision is Anthropic-shaped or genuinely general. The finding is
that **the vendor-agnostic layer already largely exists**, and the Anthropic specifics are a thin
band around it that a per-vendor adapter can hold.
**What is already general, with evidence.** `mesh-control internal/licences` models
`licence(name, vendor, serves)` and `licence_holder(licence, node, module, sealed)`, and each
holder's credential is sealed per-holder through `internal/secrets`. `serves` carries the
non-secret facts (a base URL, a model) and is not vendor-specific. The `accept` verb
([ADR 0024](0024-model-access-is-a-provision.md), and
[`14-model-access`](../03-DESIGN/01-to-be/14-model-access.md)) already takes an operator-supplied
value, seals it to each holder and discards the plaintext. A model the mesh runs itself answers
`model-access` at node scope with no licence at all. None of that mentions Anthropic.
**What is Anthropic-specific.** The credential is not a static key: it is a subscription OAuth grant
— an hourly access token plus a refresh token. That shape drags four things behind it that a static
key does not need: **central rotation** (one manager node refreshes under a lease and publishes the
new token), **delivery that strips the refresh token** so a consuming node holds only an access token,
an **identity guard** that reads the credential to catch a mis-binding, and a **usage** reading with
Anthropic's own `utilization%` semantics. Most vendors are a single static key, which the sealed-key
model already handles and which needs none of these four.
**The tension at the centre of this.** A refreshable credential cannot be both *sealed so the mesh
cannot read it* and *rotated centrally*. Central rotation means some node in the mesh holds the
refresh token in readable form, because that is what refreshing requires. Per-holder sealing means no
node but the holder can read the credential. For a static key the two never meet — there is nothing to
rotate. For a refreshable grant they collide directly, and this record exists to say which gives way,
and by how much.
## Considered Options
1. **Keep Anthropic special-cased in the core.** Leave the three binding columns and the `claude_*`
schema, and add other vendors beside them the same way. **Rejected.** It is exactly what
[ADR 0024](0024-model-access-is-a-provision.md) ruled against: a module that names a vendor cannot
be moved onto another model without editing it, and moving it is the point. It also grows the core
by one band per vendor, when the bands are the same shape.
2. **One provision, and refuse to hold any refresh token — re-seal only.** Make every credential
purely sealed per-holder, including refreshable ones; let each holder refresh its own grant.
**Rejected.** It throws away the hard half [ADR 0024](0024-model-access-is-a-provision.md) says
already works — the lease, the single-refresher, the switch-on-exhaustion — and replaces it with N
nodes each holding a refresh token, which is the very thing today's delivery strips on the stated
ground that *a node never holds a refresh token*. A refresh token is the long-lived secret; spraying
it across every holder is strictly worse than keeping one copy on one node.
3. **One provision, and abandon central rotation entirely** for refreshable vendors — treat the grant
as opaque and let it expire. **Rejected.** For a subscription-seat vendor an expired access token is
a dead licence; without rotation the feature that works today stops working. This is option 2's cost
without option 2's autonomy.
4. **One vendor-blind provision, plus a per-vendor adapter, with a bounded carve-out for the
refreshable case.** **Adopted**, below.
## Decision
**`model-access` is one consumer-facing, vendor-blind provision.** A consumer declares
`requires: model-access`, and is coupled to *reaching a model* — a base URL, a model name, a key —
and not to which vendor answers. That is the coupling the name is drawn at
([ADR 0040](0040-what-a-module-is.md)'s rule: name the interface at the widest boundary across which
the consumer does not care which implementation serves it). Where a consumer were genuinely coupled to
a specific wire API it could not swap across, the same rule would split the name — but the consumers
that exist reach their model through a CLI or SDK that hides the vendor, so `model-access` is the true
coupling and stays one name. This extends [ADR 0024](0024-model-access-is-a-provision.md) and
[ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) without changing them.
**A vendor is an adapter, keyed by the licence's `vendor` field.** The lifecycle a licence needs is
vendor-specific and lives in a per-vendor adapter selected by `licence.vendor`, exactly as
`public-dns` is one neutral interface answered by registrar-scoped providers —
`cloudflare-dns`, `route53-dns` ([ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md)).
A consumer names `model-access` and never a vendor, the same way a module names `public-dns` and never
a registrar.
**The field is named `vendor`, not `provider`.** The inventory already uses "provider" for the
provider-pin — *which node answers a brokered provision*. Reusing it for *which company sells this
licence* would collide two unrelated facts on one word. `vendor` is the licence's, and is separate.
### The adapter's capabilities, all but one optional
An adapter declares:
- **`shape`** — `static-key` or `refreshable-grant`. This is the switch the carve-out below turns on.
- **`accept(value) → sealed`** — take an operator-supplied credential and seal it to the holders, the
`accept` verb [ADR 0024](0024-model-access-is-a-provision.md) already defines.
- **`refresh(licence)`** — refreshable-grant only: the lease / rotate / publish machinery.
- **`identity(credential) → account-id`** — the mis-binding guard, for a vendor whose credential
carries an identity worth checking.
- **`usage(licence) → normalised rows`** — the vendor's usage reading, mapped to the common shape below.
- **`deliver`** — the credential *value* only; the destination path is the consumer's, not the
adapter's.
**A static-key vendor implements almost nothing** — `shape: static-key`, `accept` is the generic
seal, `deliver` is the value, and `refresh`, `identity` and `usage` are absent or trivial. The
abstraction earns its keep by making the common vendor small, not the rare one clever.
### The carve-out — the one place the guarantee is relaxed, said plainly
The mesh's standing principle is that it cannot read what it stores: `accept` seals to the holders and
discards the plaintext ([ADR 0024](0024-model-access-is-a-provision.md)), and a provider seals nothing
because the credential travels the mesh's own asymmetric channel
([ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)). A `refreshable-grant`
credential cannot honour that principle and be centrally rotated at the same time, and central rotation
is the working half [ADR 0024](0024-model-access-is-a-provision.md) is explicit about keeping.
**So, for `refreshable-grant` vendors only:**
- the **manager node holds the refresh token encrypted at rest** — readable by that node, because
rotation requires it. This is the bounded exception.
- **access tokens are still sealed per-holder**, as every credential is; a holder reads its own and no
other node reads it.
- the **refresh token is stripped on delivery** — it never reaches a consuming node. *A node never
holds a refresh token* stays true for every node but the one manager.
**Static-key vendors keep the full guarantee.** There is no token to rotate, so there is nothing to
hold readably, so `accept` discards the plaintext and the carve-out never fires. The majority of
vendors are static-key, and the majority therefore lose nothing.
The exception is stated rather than hidden because a relaxed guarantee that is not written down is
indistinguishable from a broken one. It is bounded on three axes at once: **refreshable-grant vendors
only, the refresh token only, the manager node only.**
### The settled details this record also fixes
- **Usage is normalised to `(licence, consumer, period, metric, value)` plus the raw response as
jsonb.** The metric is vendor-defined — Anthropic's `utilization%` is one metric, a token count is
another — and no common unit is forced across vendors. The raw response is kept so a reading can be
re-derived if the normalisation is later found wrong.
- **Binding is explicit per consumer, and an unchosen consumer is refused — no implicit fallback.**
This is the direction the resolver already takes, and it is the safe one: a mesh with several ways
to reach a model refuses a consumer that has not said which, naming the candidates and the command,
rather than silently choosing one ([ADR 0024](0024-model-access-is-a-provision.md), and
[`14-model-access`](../03-DESIGN/01-to-be/14-model-access.md)).
- **Subscription-seat authentication lives entirely inside the adapter**, never in the generic core.
So does an interactive `/login` — an adapter-specific "adopt" origin for a credential a person must
produce in a browser; the generic `licence key <name>` covers the static-key case.
- **Anthropic is the first `refreshable-grant` adapter**, carrying the OAuth refresh, the usage
reading, the identity guard and access-token-only delivery. **`anthropic-api-key` is a
`static-key` adapter for the same vendor's plain API keys**, and is the early second case that
proves the abstraction is not a single vendor wearing a coat: it exercises the whole path with the
carve-out switched off.
## Consequences
- **The carve-out is the mesh's one deliberate relaxation of "it cannot read what it stores."** It is
bounded to refreshable-grant vendors, to the refresh token, and to the manager node; the static-key
majority keep the full guarantee unchanged. This is the open risk the analysis carried here, and it
is recorded as an exception rather than pretended away.
- **Anthropic collapses from special case to adapter.** The three binding columns
(`nodes.node_license`, `nodes.hal_claude_account`, `agents.claude_account`) become three ordinary
consumers of `model-access`; the `claude_*` schema becomes the generic licence tables plus one
adapter. What was hardcoded becomes data keyed by `vendor`.
- **Adding a vendor is adding an adapter, and a static-key vendor is nearly free.** The modules that
want a model do not change when a vendor is added — they named `model-access`, not a vendor.
- **The refresh-token concentration is now a stated property to defend, not an accident.** The manager
node is a place a long-lived secret lives readably, and losing it or compromising it is a bounded,
named blast radius rather than a surprise.
### How each claim here is checked
- **Vendor-blind provision, static-key path.** A lab scenario binds an `anthropic-api-key` licence to
a consumer; the consumer resolves, receives a key sealed to its node, and reaches a model — and the
key is **nowhere in the control plane's database** nor in anything that crossed the broker. This is
the `licence_holder` sealed-per-node check that [`14-model-access`](../03-DESIGN/01-to-be/14-model-access.md)
already runs, now asserted for a second vendor.
- **The carve-out is exactly as narrow as stated.** For a `refreshable-grant` licence, a test asserts
the refresh token exists (encrypted) **only on the manager node**, is **absent from every holder's
delivery**, and that the delivered credential is access-token-only — and that for a `static-key`
licence no refresh token is stored anywhere.
- **Adapter selection is keyed by `vendor`.** A scenario with two vendors on two licences verifies each
licence's lifecycle runs its own adapter, and that a consumer naming `model-access` never names a
vendor to get one.
- **Refuse-if-unchosen.** Already checked in [`14-model-access`](../03-DESIGN/01-to-be/14-model-access.md):
a consumer with more than one candidate is refused with the candidates and the command named.
## References
- [ADR 0024](0024-model-access-is-a-provision.md) — model access is a provision, a licence is a named
thing, and `accept`; this record generalises its single vendor and keeps its working central
rotation.
- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — a provision names the
coupling; `model-access` is drawn at the consumer's.
- [ADR 0040](0040-what-a-module-is.md) — the naming rule and the neutral-interface / scoped-provider
shape a vendor adapter follows.
- [ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md) — registrar-scoped `public-dns`
providers, the precedent a `vendor`-scoped adapter mirrors.
- [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md) — the mesh seals credentials
and holds no readable copy; the carve-out here is the bounded, named exception to that for a
refreshable grant.
- [`03-DESIGN/01-to-be/14-model-access.md`](../03-DESIGN/01-to-be/14-model-access.md) — the design this
record extends, amended to describe the adapter generalisation.
- The read-only vendor-agnostic analysis, 2026-09-05 (code workspace) — the inventory and the decisions
taken on the open questions this record encodes.
@@ -0,0 +1,166 @@
---
topic: what runs on it
status: accepted
date: 2026-09-05
deciders: jochen
reconstructed: false
extends: 0030-data-outlives-the-mesh-that-declared-it.md
---
# 51. Shared data is the operator's, and a module is granted access to it
## Context
**Eight modules declared one filesystem as eight private ones.** The media stack — a library
server, the acquisition managers for films, series, music and books, a subtitle fetcher and two
download clients — shares directories on one machine: the download clients write into
`/services/media/downloads` and the managers read it; the managers write into the libraries and
the library server reads them. That sharing is the entire point of the stack. Yet each module
declared every shared directory it touched as its own `directory` resource, with an owner and a
mode. `/services/media/downloads` was written seven times, as seven private directories that
happen to be the same path.
**The resolver refuses exactly that, and is right to.** Two modules declaring one path on one node
are refused by name, with no exemption for identical content and no merge — because two owners of
one path is the class of fault this repository keeps recording ([04-ISSUES/036](../04-ISSUES/036-six-modules-own-what-they-must-share/00-report.md)).
So the stack as written refuses its own only sensible assignment: all of it on one machine,
sharing one filesystem. It passed today only because no test co-resolves any two of the eight. The
first machine assigned two of them together is where the refusal would have surfaced.
**The vocabulary had one word for two intentions, and this was already seen.**
[04-ISSUES/026](../04-ISSUES/026-the-data-directories-are-not-declared/00-report.md) found the
same gap from the other side and named it precisely: two kinds of mount are spelled identically —
*the directory my data lives in*, which the mesh creates and owns, and *a facility I was granted*,
which already exists and the mesh only reaches. That issue deferred inventing a field to tell them
apart, because doing so is a design decision and it declined to make one to get a check green. This
is that decision.
**A `directory` resource is owned, on every axis.** The host creates it, sets its owner and mode,
and removes it when it is empty and no longer declared — it *removes what it made and leaves what
it merely configured* ([ADR 0005](0005-the-node-host.md), [ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md)).
A shared media library is none of that. It existed before the mesh, several modules read and write
it at once, and losing it is the one failure that does not recover. It is not any module's
resource; it is the operator's, and a module only needs to be let at it.
## Considered Options
1. **Make the stack one module with several containers**, the way the mail system already is. The
issue raises it directly: is a set of modules that must share a filesystem really one module?
**Rejected.** The eight are independently assignable and independently useful — a person may run
the download client without the library server, or the film manager without the music one — and
folding them into a single module to express a shared directory would make *what a module is*
turn on an incidental filesystem contract. It also does not generalise: the next pipeline of
modules handing files to each other on one machine (an ingest folder, a spool, a drop directory)
would face the same wall and the same wrong remedy.
2. **One module owns the directories and the rest `require` them.** **Rejected**, and this is the
heart of the decision. Nobody owns shared operator data. The library predates the mesh and
outlives any one module, so making the library server or a manager its owner means unassigning
that module orphans everyone else's access — and the owner would set the owner and mode of a
tree it did not create. Ownership is the wrong relationship to model, because the true owner is
not a module at all.
3. **A flag on a directory resource** — `external: true`, or an owner of `operator`. **Rejected.**
It overloads one shape with a boolean that inverts every one of its semantics: created becomes
*must already exist*, owned becomes *touch nothing*, removed-when-empty becomes *never removed*.
That is the *two-kinds-spelled-identically* trap of 04-ISSUES/026 reintroduced with a single
quiet field — a reviewer reading `type: directory` would have to check one boolean elsewhere to
know whether the mesh owns the thing at all.
4. **A distinct `accesses` declaration, separate from resources.** **Adopted.**
## Decision
**Shared, pre-existing data is operator-owned and external. The mesh does not create it, does not
set its owner or mode, does not reconcile it and does not remove it.** A media library, a download
spool, an ingest directory is the operator's, and the mesh is a guest in it.
**A module declares that it needs *access* to such a path, not that it owns a resource there.** The
manifest field is `accesses`: a list of `{path, mode}`, where mode is `read` or `read-write` and
absent narrows to `read` — the safe default, because the danger with an access is being given more
than was meant, not less. A module's own configuration and state directories stay owned
`directory` resources; only the shared, pre-existing paths become accesses.
**The host mounts an accessed path and owns nothing about it.** It reaches the machine as a new
declaration shape, `access`, distinct from `directory`. The host confirms the path is present and
does nothing else — no create, no chown, no mode, no removal.
**An accessed path absent at apply time is refused, clearly, not created.** The mesh does not own
it, so conjuring it would be a lie the host then acts on — and specifically the lie 04-ISSUES/026
records, where a bind mount whose source does not exist is made by the container runtime as root
with the wrong ownership. The host says the operator must provide the path instead.
**Several modules accessing one path is normal, and never refused.** The duplicate-path refusal is
about *ownership*, not *use*: it applies to resources a module owns and to those alone. An access
is not a resource and never enters the check, so the eight-module stack co-resolves. What stays
refused is genuine rivalry — two modules owning one path — and the new contradiction it exposes: a
path one module owns while another merely accesses it, because that asserts both that the mesh owns
the directory and that the operator does.
This is a decision and not a patch because it settles *what a module may say about a path it did
not make*, which every co-located file-handoff in the catalogue now and later depends on — and
because it draws the ownership line [ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md)
started: the host owns what it made and keeps what it merely configured, and this adds the third
case it did not have a word for — what it neither made nor configured, and must not touch.
### How each claim is checked
- **The stack co-resolves.** A control-plane unit test assigns two modules that declare access to
one path on one node and asserts no refusal — the exact case the resolver refuses when the same
path is owned. The mirror test, two modules *owning* one path, still refuses, so the sharing
vocabulary does not weaken the rule it sits beside.
- **Ownership and access cannot both be claimed of one path.** A unit test asserts the resolver
refuses a path one module owns and another accesses, naming both.
- **Absent is refused, not created.** A host unit test applies an access to a path that does not
exist and asserts a clear refusal that names the operator, and that nothing was created.
- **Present is confirmed and nothing moves.** A host unit test applies an access to an existing
directory and asserts the apply reports no change and disturbs nothing.
- **Undeclaring never removes.** A host unit test drops a previously declared access and asserts
the operator's directory and its contents are left exactly as they were — the data-loss failure
[ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md) exists to prevent, on a directory the
mesh never made.
- **Every media manifest is corrected.** No `/services/media/*` path is an owned `directory`
resource in any of the eight; each is an `accesses` entry, and each module's own config and state
directories remain owned. Checked by the control-plane manifest parser, which now understands
`accesses` and refuses a malformed one.
## Consequences
**A shared filesystem between co-located modules now has a vocabulary**, and it is not the media
stack's alone: any pipeline handing files to a neighbour on one machine — an ingest directory, a
spool, a drop folder — says *I access this operator path* rather than *I own this directory*, and
several of them may say it of one path.
**Unassigning a module that reached shared data leaves the data.** Correct, and the same trade
[ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md) made for owned directories: removing
data is a person's act, done knowingly, not a side effect of unassignment.
**The operator must provision the shared paths before the stack is applied**, and a machine that
lacks one is told plainly which. That is a real new obligation, and it is the right one: the mesh
cannot own what predates it, so it cannot create it either, and saying so at apply time beats a
directory conjured as root and a service that half-works.
**The host vocabulary grew by one shape**, which is a cost — every added shape widens what a
compromised control plane can express ([ADR 0005](0005-the-node-host.md)). It is a narrow one: an
`access` is confirmed by a stat and grants the host no new action. It earns its place by letting
the host refuse to create what it must not own, which no existing shape could say.
**A path can be both owned and accessed only by refusal.** If a future manifest declares one path
as an owned directory in one module and an access in another, the resolver refuses it rather than
guessing which is meant — the two assertions about who owns the data cannot both hold.
## References
- [04-ISSUES/036](../04-ISSUES/036-six-modules-own-what-they-must-share/00-report.md) — six (in
fact eight) modules own what they must share; the problem this resolves
- [04-ISSUES/026](../04-ISSUES/026-the-data-directories-are-not-declared/00-report.md) — the two
kinds of mount spelled identically, which deferred this field to a decision
- [ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md) — data outlives the mesh; the host
keeps what it did not make. This record adds the case it had no word for
- [ADR 0005](0005-the-node-host.md) — the host removes what it made and leaves what it merely
configured; the vocabulary is finite and every shape is a security decision
- [ADR 0040](0040-what-a-module-is.md) — what a module is; an access is a new thing a module may
say about the machine it lands on
- mesh-control `feat/shared-data-access`, mesh-catalog `feat/media-access-not-ownership`,
mesh-host `feat/mount-operator-owned` — the mechanism, the corrected manifests, and the host
shape
@@ -0,0 +1,198 @@
---
topic: what runs on it
status: accepted
date: 2026-09-05
deciders: jochen
reconstructed: false
extends: 0005-the-node-host.md
---
# 52. An init step is a container run once to completion, gating what follows
## Context
**A module can declare things that exist; it cannot declare a step that runs.** The host owns a
finite vocabulary of shapes — `file`, `directory`, `service`, `package`, `container`, `action` —
and every one but `action` describes *state*: a thing that should be present, with content or a
mode or an image, which the host reconciles toward ([ADR 0005](0005-the-node-host.md)). That is
right for what it covers. But a real class of modules needs, once, to *run their own code at a
point in their own lifecycle* — and the vocabulary has no word for it
([04-ISSUES/037](../04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md)).
**mosquitto is the sharp case, and it fails silently without this.** Its Dynamic Security plugin
loads at broker start and refuses to come up unless `dynamic-security.json` already holds an admin
client. Seeding that file is a step that must happen *after* the data directory exists and *before*
the broker container starts. The manifest can declare the directory, the config file and the broker
container; it cannot declare "seed this, once, before that container starts." Written as it is
today, the broker starts against an unseeded store and the plugin aborts — and the next reconcile
does not fix it, because nothing in the declaration ever seeds the file.
**It is not one module's defect.** The database providers need the same to run a first-boot
migration, an extension enable, or a health gate before they are announced ready; today that works
only where the *image* happens to seed itself from an environment variable, and anything the mesh
must run once against the server has no home. This is the timing face of the same gap
[04-ISSUES/035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md) records
from the content side: the manifest needed *the file to exist before first start*, and had only
*the file has this content, forever*.
**The obvious answer is the one that already went wrong.** An earlier mesh had exactly this as a
feature — event-driven hooks that ran custom code at phases of build, publish and deploy. It was
powerful and it was *complex to set up and flaky*, and that fragility, not the need, is the content
of the issue. Whatever this becomes must not rebuild that engine.
**The ground has shifted since that engine, in a way that makes a much smaller answer possible.** A
module with tools or events now runs a **process of its own** — a container carrying the module's
compiled code, holding the single broker account the mesh scoped to it, isolated from every other
module ([ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)). The
code that must run at first boot is *already in that image, under that account*. So the mesh does
not need a way to run a module's code — it has one. It needs a way to say **run this container to
completion, and start the one that depends on it only after it has.**
## Considered Options
1. **A host `run` shape — a command the host executes on the machine.** The direct reading of
"run code at a lifecycle phase." **Rejected.** The host's vocabulary is finite and every added
shape is a security decision, because it widens what a *compromised control plane* can express
([ADR 0005](0005-the-node-host.md)). A general "run this command as the host" is the largest
such widening there is: the blast radius is the whole machine, as root. The mesh already drew
this exact line for `action` — it runs a command, and so it is *permitted from the bundle and
refused from the link*, because the bundle arrives with the binary and the link is a separate
party with an unbounded reach. A new host-command shape usable by an ordinary module would be an
`action` from the link by another name, which is precisely what is refused.
2. **Per-phase lifecycle hooks on a module** — `pre-start`, `post-start`, `pre-remove`, and their
build/publish cousins, each naming code the mesh runs at that phase. The general answer, and the
old feature. **Rejected for now.** It is the flaky engine the issue warns against, and most of
its phases have no present need. Deciding the full set of phases, where each one's code runs, and
how each is made idempotent is a large design taken to buy capability nothing yet asks for. The
three blocked modules all need one phase — *before a container starts* — and a mechanism narrow
enough to be obviously correct beats a general one that is not.
3. **A distinct one-shot resource type** — a new shape, sibling to `container`, that names an image
and runs it once. **Rejected.** It grows the host vocabulary by a whole shape (a `Type`, a
struct, an applier, a place in every host's shape list) to express something a `container` almost
already is. A one-shot *is* a container — a pinned image, an account, volumes, an environment —
that happens to exit. Spending a new shape on the difference is the cost of option 1 in smaller
type, for a capability the existing shape can carry with one modifier.
4. **A modifier on the existing `container` shape: this container runs once, to completion, and the
host gates the apply on it.** **Adopted.** It reuses the shape the host already has, adds no new
host action, and leans on two guarantees the host already gives — *apply in declared order,
never sorted*, and *a failed step fails the apply* — to turn "before that container starts" into
an emergent property of ordering rather than a dependency graph the host must resolve.
## Decision
**A run-once step is an ordinary `container`, marked to run to completion.** The manifest sets
`run-once: true` on a container resource. Everything else about it is a container as before — a
digest-pinned image ([ADR 0006](0006-the-substrate-and-the-control-plane.md)), volumes, an
environment, and for a module's own code the same scoped account its runtime already holds
([ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)). The mesh adds
no way to run code; it marks a container the host must **run to completion and require to exit 0**,
rather than start and leave running.
**The gate is declaration order, not a named dependency.** The host applies a declaration in the
order it is given, does not sort, and does not resolve dependencies — ordering is a decision, and it
is the control plane's ([ADR 0005](0005-the-node-host.md),
[04-ISSUES/013](../04-ISSUES/013-a-file-arrives-after-the-service-that-needs-it/00-report.md)). A
run-once step is placed *before* the container that depends on it, and **a failed run-once step
halts the apply**, exactly as a failed `action` does — so everything the declaration places after
it, the broker included, is never reached until the step has completed. "Before the broker starts"
is therefore expressed by list position plus completion, and the host cross-references nothing.
**Completion is recorded, and a re-apply does not re-run it.** The host records what it applied only
after the fact, as the digest of the declaration that produced it
([ADR 0018](0018-a-picture-is-read-from-what-runs.md)) — a run-once step no differently. Because the
step leaves nothing running to inspect, that persisted digest, not a live container, is the marker
that it happened. On a later apply the host finds the digest already recorded for this exact
declaration and does nothing; it re-runs only when the declaration's digest has changed, and a step
that exited non-zero recorded nothing and so is retried next apply. This is the reconcilable,
idempotent discipline the state shapes get for free, made explicit for a step.
This is a decision and not a patch because it settles **what a module may say about running its own
code**, which the whole catalogue of providers — a seed before start, a first-boot migration, a
health gate — now and later depends on, and because it draws the line the issue asked for: the
narrowest sound mechanism that unblocks the three modules without rebuilding the hook engine whose
fragility is the warning.
### The security bound, stated plainly
**The host gains no new action and no new shape.** `run-once` is a boolean modifier on the
`container` shape that already exists. A run-once container is strictly *less* powerful than an
`action`: it cannot run an arbitrary host command, only a digest-pinned image under an account the
mesh scoped — which is exactly the capability `container` already grants from the link. A control
plane that is compromised can express nothing through `run-once` it could not already express by
declaring an ordinary `container`. The dangerous expansion of option 1 — a command the host runs on
the machine — is not made.
### How each claim is checked
- **A run-once step runs to completion and its exit 0 is required.** A host unit test applies a
run-once container whose image exits 0, asserts the host ran it to completion (not detached, not
left running) and reported it done; a sibling test applies one that exits non-zero and asserts
the apply fails, naming the step.
- **A failed run-once step gates what follows.** A host unit test places a run-once container that
exits non-zero before another container and asserts the second is never started and the apply is
reported gated — the mirror of the existing test that a failed action stops what follows.
- **It is not re-run once it has completed.** A host unit test applies a run-once step, then applies
the identical declaration again with the first run's record present, and asserts the second apply
runs nothing and reports the step unchanged.
- **A changed declaration re-runs it.** A host unit test applies a run-once step, then applies one
whose image or environment differs, and asserts it runs again — the digest moved, so the marker no
longer matches.
- **The vocabulary carries the field end to end.** A control-plane unit test resolves a module whose
manifest marks a container `run-once` and asserts the rendered host declaration carries the field,
in author order before the container it gates; the manifest parser refuses a `run-once` that is
not boolean.
- **mosquitto seeds before the broker.** mosquitto's manifest declares a run-once init container,
before the `server` (broker) container, that writes the admin client into `dynamic-security.json`
and exits — checked by the control-plane resolver, which now understands the field, and by the
ordering of the rendered declaration. The end-to-end proof that it runs exactly once, at the right
phase, and converges on re-apply is owed to a lab scenario ([04-ISSUES/037] open question), which
this record does not close.
## Consequences
- **The three blocked modules gain a home for their step.** mosquitto seeds its dynsec admin before
the broker; a provider that must migrate or health-gate at first boot declares a run-once step in
its own runtime image, under its own account, before the container that depends on it.
- **The general lifecycle hook is deferred, deliberately.** Only *before a container starts* is
bought here. `post-start`, `pre-remove` and the build/publish phases remain unbuilt, and the day
one is genuinely needed it is decided then, against a need, not speculatively — the same restraint
that kept this from being the old engine.
- **The seed-then-mutate file is safe if the step is written to be.** A run-once seed writes
`dynamic-security.json` only when it is absent and never reconciles it, so what the running plugin
grows in that file afterward is never wiped
([04-ISSUES/035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md)).
The host's marker guarantees the step is not re-run; the step's own code guarantees it does not
clobber on the pass it does run.
- **The host vocabulary did not grow, and that is the point.** The cost of a run-once step is one
boolean and a completion path in the container applier, not a new shape and not a new action. The
mesh expresses ordering and completion; the module runs its own code, where it already runs it.
- **A run-once step that never converges is a stuck apply, loudly.** A step that exits non-zero
every time halts the apply every time, and the container it gates never starts — which is the
correct failure, reported, rather than a broker that half-starts against an unseeded store and a
reconcile that reports success. It is failed forward, not failed silent.
## References
- [04-ISSUES/037](../04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md) — a
module cannot run its own code at a lifecycle phase; the gap this resolves, and the warning about
the old hook engine
- [04-ISSUES/035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md) — the
content face of the same gap: a file needed before first start, that the running program then
mutates
- [04-ISSUES/013](../04-ISSUES/013-a-file-arrives-after-the-service-that-needs-it/00-report.md) — the
order of a declaration is the control plane's, and the host applies it as given; the gate rests on
this
- [ADR 0005](0005-the-node-host.md) — the host's finite vocabulary, the ordered declaration it does
not sort, and `action` as the shape a command already is and why it is refused from the link
- [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) — a module runs
its code as its own process under its own account; a run-once step is that process, run to
completion
- [ADR 0018](0018-a-picture-is-read-from-what-runs.md) — what was applied is recorded after it works;
the completion marker is that record
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — a container is pinned by digest; a
run-once container no differently
- mesh-control `feat/lifecycle-run-once`, mesh-host `feat/apply-run-once`, mesh-catalog
`feat/mosquitto-bootstrap` — the vocabulary, the apply support, and mosquitto's seeded broker
@@ -0,0 +1,181 @@
---
topic: what runs on it
status: accepted
date: 2026-09-06
deciders: jochen
reconstructed: false
extends: 0052-a-step-that-runs-once-before-a-container.md
---
# 53. A scheduled step is a container run on a recurring schedule
## Context
**[ADR 0052](0052-a-step-that-runs-once-before-a-container.md) gave the mesh a step that runs *once*;
a real class of modules needs one that runs *again and again*.** kometa reconciles a media library
against its lists on a timer; a ticketing integration polls its source for new work every few minutes;
a backup, a cache warm, a metrics roll-up all recur. The host's vocabulary describes *state* — a file,
a directory, a container that should be running — and 0052 added *a step that happens once and is
done*. Neither says *this should happen every night at 3, forever*
([04-ISSUES/037](../04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md) named
the lifecycle gap; 0052 closed the once-before-start face of it and left the recurring face open).
**The modules that need it already run their own code.** As with run-once, the ground has shifted
since the old mesh's flaky hooks: a module with tools or events runs a **process of its own** — a
container carrying the module's compiled code under the single scoped account the mesh gave it
([ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)). The work that
must recur is already in that image, under that account. The mesh does not need a way to run a
module's code on a timer — it has the code and the account. It needs a way to say **run this
container again on this cadence**.
**The shape of the answer is already decided, one modifier over.** 0052 rejected a host `run` command,
per-phase hooks, and a distinct one-shot resource type, and adopted *a modifier on the `container`
shape the host already has*, because a one-shot is a container that happens to exit. A scheduled step
is the same container that happens to exit — run repeatedly. Reusing that shape keeps the host
vocabulary flat and inherits 0052's security bound whole. The only thing 0052's `run-once` does not
carry is *when to run it again*.
**One difference from run-once changes a rule, and it is the reason this is its own record.** A
run-once step **gates the apply**: it is placed before the container that depends on it, and a failure
halts everything after it, because "seed the store before the broker starts" is a correctness
precondition ([ADR 0052](0052-a-step-that-runs-once-before-a-container.md)). A scheduled step is the
opposite: it runs *after* the machine is up and converged, on its own clock, and a single failed run
is an ordinary operational event — the next run comes anyway. A scheduled step that halted the apply,
or that a failed run marked the node not-current over, would make a routine poll into a reason the
whole machine reads as broken. So the gating rule 0052 established is exactly the rule this record
must **not** inherit.
## Considered Options
1. **A host `cron`/`timer` shape — the host installs a system timer that runs a command.** The direct
reading. **Rejected**, for the reason 0052 rejected a host `run` shape: it widens what a
*compromised control plane* can express toward "run this command on the machine, forever," which is
the largest widening there is, and it is `action`-from-the-link by another name
([ADR 0005](0005-the-node-host.md)). A recurring command is worse than a one-off, because it
persists.
2. **Per-phase lifecycle hooks** — `on-schedule` joining `pre-start`/`post-start` as named code the
mesh runs. **Rejected for now**, as in 0052: it is the flaky hook engine the issue warns against,
and the three modules that need this need one thing — *run this container on a cadence* — which a
narrow modifier expresses without deciding a whole hook vocabulary.
3. **A distinct `scheduled` resource type**, sibling to `container`. **Rejected**, as 0052 rejected a
distinct one-shot type: it spends a whole new host shape (a `Type`, a struct, an applier, a place
in every host's shape list) on something a `container` already almost is — a scheduled task *is* a
container (pinned image, account, volumes, environment) that runs on a clock.
4. **A modifier on the existing `container` shape: `schedule`, a cron expression the host runs the
container on.** **Adopted.** It reuses the shape the host has, adds no new host action, and sits
beside `run-once` as its recurring twin — the same container, exited, run again.
## Decision
**A scheduled step is an ordinary `container`, marked with a `schedule`.** The manifest sets
`schedule: "<cron>"` on a container resource — a standard five-field cron expression. Everything else
about it is a container as before: a digest-pinned image
([ADR 0006](0006-the-substrate-and-the-control-plane.md)), volumes, an environment, and for a module's
own code the same scoped account its runtime already holds
([ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)). The mesh adds no
way to run code; it marks a container the host must **run on that cadence, each time to completion**,
rather than start once and leave running (a service) or run once and gate (a run-once step).
**A run is fired by the clock, not by the apply, and does not gate it.** Applying the declaration
installs the schedule; it does not run the step. The machine converges — reports applied and current —
as soon as the schedule is installed, exactly as it does for a service that is running. Thereafter the
host fires the container when the cron expression is due. This is the deliberate inversion of 0052:
a scheduled step is downstream of convergence, not a precondition of it.
**A failed run is recorded and the next run still comes; it never marks the node not-current.** A run
that exits non-zero is logged against the module — visible, auditable — but it does not fail the apply,
does not halt other resources, and does not flip the node's reported state. A poll that fails at 03:00
and succeeds at 03:05 is the system working, not a machine that needs attention. A step that fails
*every* time is a loud, repeating log entry, which is the correct signal for "this recurring job is
broken" — distinct from "this machine did not converge."
**Runs do not stack.** If a run is still going when the next is due, the host skips the due run rather
than starting a second copy, and logs the skip. A slow nightly job that occasionally overruns must not
spawn a growing pile of concurrent containers competing for the same account and volumes — the failure
mode that made the old timers dangerous.
**Each run is independent and idempotent by the module's own code.** The mesh guarantees only *the
container is run on the cadence*; that a run does the right thing when the previous one half-finished
is the module's contract, the same discipline a run-once seed owes
([ADR 0052](0052-a-step-that-runs-once-before-a-container.md),
[04-ISSUES/035](../04-ISSUES/035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md)).
This is a decision and not a patch because it settles **what a module may say about running its own
code on a cadence**, which every recurring provider job — a sync, a poll, a roll-up — now and later
depends on, and because it draws the line 0052 left: the recurring twin of run-once, with the gating
rule deliberately reversed so a routine job's failure is never a machine's failure.
### The security bound, stated plainly
**The host gains no new action and no new shape.** `schedule` is a string modifier on the `container`
shape that already exists. A scheduled container is strictly *less* powerful than an `action`: it runs
a digest-pinned image under an account the mesh scoped, which is exactly what `container` already
grants from the link, and it cannot run an arbitrary host command. A compromised control plane can
express nothing through `schedule` it could not already express by declaring a `container` — the cron
string only says *how often*, not *what*. The dangerous expansion of option 1 — a command the host
runs on the machine on a timer — is not made.
### How each claim is checked
- **A scheduled step runs when the schedule is due.** A host unit test installs a container with a
schedule that is due immediately (or advances a injected clock to when it is due) and asserts the
host ran it to completion; a sibling test with a schedule not yet due asserts it has not run.
- **Installing it does not run it, and the node is current without a run.** A host unit test applies a
scheduled container and asserts the apply reports current *before* any run has fired — the schedule
is state that is present, not a step that gated.
- **A failed run does not fail the apply or the node.** A host unit test fires a scheduled container
that exits non-zero and asserts the failure is recorded against the module, the apply is not failed,
and the node stays current — the mirror of the run-once test where a non-zero exit *does* halt.
- **Runs do not stack.** A host unit test fires a scheduled container whose run outlasts its next due
time and asserts the host skipped the due run and logged the skip, rather than starting a second
container.
- **The vocabulary carries the field end to end.** A control-plane unit test resolves a module whose
manifest sets `schedule` on a container and asserts the rendered host declaration carries the field;
the manifest parser refuses a `schedule` that is not a valid cron expression, and refuses a container
that is both `run-once` and `schedule` (a step is one or the other, never both).
- **A real module recurs in the lab.** A converted module declaring a scheduled step (kometa's library
sync, or a poller) is assigned in a lab scenario, and the scenario asserts the scheduled container
fires on its cadence and its account and volumes are the module's — the end-to-end proof this record
owes, as 0052 owed its run-once lab proof.
## Consequences
- **The recurring providers gain a home for their cadence.** kometa reconciles on its schedule; a
poller polls; a roll-up rolls up — each a scheduled step in its own runtime image, under its own
account, on the cron it declares.
- **The general lifecycle hook is still deferred.** Only *run once before* (0052) and *run on a
cadence* (this) are bought. `post-start`, `pre-remove` and the build/publish phases remain unbuilt,
decided when a real need arrives, not speculatively — the restraint that kept both from being the old
engine.
- **A recurring job's failure is loud but not fatal.** The node stays current while a scheduled step
fails and retries; a step that fails forever is a repeating log entry, not a machine marked broken.
This is the correct separation — a machine's convergence and a job's success are different questions —
and it is why this could not simply be `run-once` without the schedule.
- **The host vocabulary did not grow, again, and that is the point.** The cost of a scheduled step is
one string field and a cron loop in the container applier, not a new shape and not a new action. The
mesh expresses cadence; the module runs its own code, where it already runs it.
- **`run-once` and `schedule` are exclusive and complete for now.** A container runs once and gates, or
runs on a cadence and does not, or runs and stays up (a service). A manifest that asks for two of
these at once is refused, because the three are distinct answers to "how does this container run."
## References
- [ADR 0052](0052-a-step-that-runs-once-before-a-container.md) — the run-once step; this is its
recurring twin, reusing the `container`-modifier shape and inheriting its security bound, and
reversing its gating rule
- [04-ISSUES/037](../04-ISSUES/037-a-module-cannot-run-code-at-a-lifecycle-phase/00-report.md) — a
module cannot run its own code at a lifecycle phase; 0052 closed the once-before-start face, this
closes the recurring face
- [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) — a module runs
its code as its own process under its own account; a scheduled step is that process, run on a cadence
- [ADR 0005](0005-the-node-host.md) — the host's finite vocabulary, and why a host command (even on a
timer) is refused from the link
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — a container is pinned by digest; a
scheduled container no differently
- [ADR 0018](0018-a-picture-is-read-from-what-runs.md) — what was applied is recorded after it works;
the installed schedule is state, a fired run is an event
- mesh-control `feat/schedule-container`, mesh-host `feat/apply-schedule`, and the converted module
that first declares a scheduled step — the vocabulary, the apply support, and the recurring proof
@@ -0,0 +1,178 @@
---
topic: model access
status: accepted
date: 2026-09-06
deciders: jochen
reconstructed: false
extends: 0050-model-access-is-vendor-agnostic.md
---
# 54. Model usage is a vendor-neutral record, produced by the adapter, at two grains
## Context
**[ADR 0050](0050-model-access-is-vendor-agnostic.md) fixed the *shape* of a usage reading and left
its *home* open.** It decided that an adapter may expose `usage(licence) → normalised rows`, that a
row is `(licence, consumer, period, metric, value)` plus the raw vendor response as jsonb, and that the
metric is vendor-defined (Anthropic's `utilization%` is one metric, a token count another). It did not
say where those rows are stored, how they get produced on a cadence, or whether the only grain is the
whole licence — and the mesh's first real consumer needs all three answered.
**The predecessor mesh recorded usage at two grains, and both are wanted.** It polled the vendor for a
licence-level reading (an account's `utilization%`), and it also attributed **per-session** token and
cost — which model, how many input and output tokens, what it cost — parsed from the agent's own
transcript, including the case where a long session switched the account it billed against mid-way. The
licence-level reading answers "how close is this subscription to its cap"; the session-level reading
answers "what did this piece of work cost, and against which account." A mesh that kept only the first
could not bill a project or notice a runaway session; keeping only the second could not see a cap
approaching. Both are load-bearing and neither subsumes the other.
**A session is already a consumer, so the second grain needs no second vocabulary.**
[ADR 0026](0026-the-mesh-has-a-session-of-its-own.md) and design 15 establish that an agent session is
a thing in its own right, and that **model access is bound to the session, not to the machine** — a
session is a *consumer* of `model-access`, identified by `(node, module)` and its own id. That is
exactly the `consumer` column ADR 0050 already put in the usage row. So the two grains are not two
schemas; they are the same row at two consumer resolutions: the holding module for the licence grain,
the session for the finer one. The design's own still-open worker-naming gap (design 14) is the same
gap here and is left where it is — a session id distinguishes what `(node, module)` cannot.
**The pieces this needs already exist.** A periodic reading is a **scheduled step**
([ADR 0053](0053-a-step-that-runs-on-a-schedule.md)) — the adapter's `usage()` poll is a container the
mesh runs on a cadence, which is precisely what 0053 was built for. A durable audit of what happened is
**an event the audit trail records** ([ADR 0041](0041-events-are-a-relationship.md),
[ADR 0042](0042-the-shape-of-an-event-on-the-wire.md)); the `audit-logger` module already consumes every
event. And a queryable current picture is **a context store**, the same shape the licences themselves
live in ([ADR 0008](0008-a-context-owns-its-store.md)). Nothing new in kind is required; what is missing is
the decision to point them at usage.
## Considered Options
1. **Licence grain only — a single `utilization%` poll, nothing per session.** Rejected: it cannot
attribute cost to a piece of work or catch a session that is burning an account down, which is half
of why usage is recorded at all.
2. **A bespoke `sessions` schema mirroring the old mesh's `*_sessions` / `*_session_account_usage`
tables.** Rejected: it reintroduces a second, vendor-shaped vocabulary for something the mesh
already names — a session is a consumer, and its usage is a usage row. A parallel schema would drift
from the `model-access` vocabulary and force every reader to learn two.
3. **Store usage only as raw vendor blobs, normalise later.** Rejected as the *whole* answer (kept as a
fallback within the chosen one): a reader that must parse Anthropic's response shape to answer "what
did this cost" has the vendor coupling the whole feature exists to remove. The raw blob is kept
beside the normalised row (0050 already requires this), not instead of it.
4. **One vendor-neutral usage record at two consumer grains, produced by the adapter, recorded as
both an event and a queryable row.** Adopted.
## Decision
**Model usage is one vendor-neutral record — `(licence, consumer, period, metric, value)` plus the raw
response — recorded at two grains that differ only in the `consumer`.** At the **licence grain** the
consumer is the holding module and the metric is the vendor's own account reading (Anthropic:
`utilization%`). At the **session grain** the consumer is the agent session — `(node, module)` and its
session id ([ADR 0026](0026-the-mesh-has-a-session-of-its-own.md)) — and the metrics are the ones a
session bills: input tokens, output tokens, model, and cost. The row shape is 0050's, unchanged; the
grain is which consumer the row is *for*.
**The adapter is the only thing that knows the vendor, and it produces both grains.** Reading an
account's cap is the adapter's `usage(licence)` verb ([ADR 0050](0050-model-access-is-vendor-agnostic.md));
attributing a session's cost is the adapter reading that vendor's transcript or usage API and emitting
rows keyed to the session. The mesh defines the row and the plumbing; the adapter fills it from whatever
the vendor exposes, and a static-key vendor that exposes nothing simply produces no rows — usage is an
optional reading, not a requirement of holding a licence.
**A reading is taken on a schedule, not on a request.** The licence-grain poll is a **scheduled
container** ([ADR 0053](0053-a-step-that-runs-on-a-schedule.md)) the adapter runs on a cadence; the
session-grain rows are produced as sessions progress, from the transcript the session already writes.
Neither blocks anything: a poll that fails is a logged, retried scheduled run (0053's rule), and a
session whose cost cannot yet be attributed is a row not yet written, never a session refused.
**Usage is recorded two ways, for two audiences.** Each reading is **emitted as an event**
([ADR 0041](0041-events-are-a-relationship.md)) — an immutable "this was observed at this time" that the
`audit-logger` already records, so the history of what an account did is in the audit trail by default,
under nobody's special arrangement. And the **current** picture — the latest reading per
`(licence, consumer, period, metric)` — is upserted into a **usage context store**, so "how close is
this cap" and "what has this project spent this month" are a query, not a fold over the event log. The
event is the record of what happened; the store is the answer to what is true now.
**Usage is not a credential, and is recorded in the clear.** The one thing the mesh must not read is the
sealed key ([ADR 0050](0050-model-access-is-vendor-agnostic.md), the refresh-token carve-out aside). A
token count and a cost are not secrets; they are the operator's own operational facts, recorded openly
so they can be queried, audited, and charged against. This is the deliberate opposite of the credential
rule, and stating it prevents a later reader assuming usage inherits the key's secrecy and hiding it
from the person who is paying.
This is a decision and not a patch because it settles **where a model's usage lives and at what grain**,
which every reader — a bill, a cap alarm, a per-project report — depends on, and because it closes the
half [ADR 0050](0050-model-access-is-vendor-agnostic.md) explicitly left open, reusing the session
([ADR 0026](0026-the-mesh-has-a-session-of-its-own.md)), the schedule
([ADR 0053](0053-a-step-that-runs-on-a-schedule.md)), the event ([ADR 0041](0041-events-are-a-relationship.md))
and the store ([ADR 0008](0008-a-context-owns-its-store.md)) the mesh already has rather than inventing a
vocabulary beside them.
### How each claim is checked
- **A usage row is vendor-neutral and carries the raw beside it.** A unit test constructs an Anthropic
`utilization%` reading and a token/cost reading and asserts both render to
`(licence, consumer, period, metric, value)` with the vendor response preserved in the raw column;
a reader that answers "what did this cost" touches only the normalised columns.
- **The two grains differ only in the consumer.** A unit test records a licence-grain row (consumer =
the module) and a session-grain row (consumer = a session id) for one licence and asserts both are
the same shape and both are returned when the licence's usage is asked for, distinguishable by
consumer.
- **A poll is a scheduled run and its failure is not fatal.** The adapter's usage container declares a
`schedule` ([ADR 0053](0053-a-step-that-runs-on-a-schedule.md)); a host test (0053's) already proves a
scheduled step runs on cadence, does not gate, and logs rather than fails on a non-zero run — the poll
inherits this and adds nothing to check.
- **Each reading is an event the audit trail records.** An integration check asserts a usage reading
emits an event that the `audit-logger` receives (it consumes `#`), so the history is present without
the usage module and the audit module knowing about each other beyond the event.
- **The current picture is a query.** A store test upserts two readings for one
`(licence, consumer, period, metric)` and asserts the later replaces the earlier, so "what is true
now" is one row, while the event log keeps both.
- **A static-key vendor with no usage reading records nothing, and that is fine.** A test resolves a
`static-key` model-access consumer whose adapter has no `usage` verb and asserts the licence works and
no usage rows or poll are required — usage is optional, holding a licence is not conditioned on it.
- **Usage is readable in the clear; the key is not.** A test asserts a usage row is stored unsealed and
is returned to an ordinary query, while the licence key remains sealed and absent from the same
surfaces — the deliberate inversion of the credential rule.
## Consequences
- **A bill and a cap alarm are both queries.** "What did project X spend this month" reads the
session-grain rows; "how close is account Y to its cap" reads the latest licence-grain metric — both
from the usage store, neither a fold over events or a call to the vendor.
- **The session becomes the unit of cost, which is what it already is.** Because a session is the
consumer, attributing cost needs no new identity — and the worker-granularity gap
([design 14](../03-DESIGN/01-to-be/14-model-access.md)) surfaces here exactly as it does for access,
to be closed once, for both, when a session id is threaded through.
- **The adapter carries the vendor's usage quirks alone.** Anthropic's `utilization%`, its transcript
shape, a mid-session account switch — all live in the Anthropic adapter; the mesh, the store, and
every reader see only rows. A second vendor adds a second adapter and no new table.
- **Usage history is durable and tamper-evident by reuse, not by a new mechanism.** It rides the event
trail the mesh already keeps, so an operator who wants the whole history has it, and one who wants the
current number has the store — without usage owning either mechanism.
- **The refresh-token carve-out is untouched by this.** Usage is read *from* an authenticated adapter;
it neither holds nor exposes the credential, so the one place the mesh reads what it stores
([ADR 0050](0050-model-access-is-vendor-agnostic.md)) is not widened by recording what that credential
was spent on.
## References
- [ADR 0050](0050-model-access-is-vendor-agnostic.md) — model access is vendor-agnostic; fixes the
usage row shape and the `usage(licence)` adapter verb, and leaves its home open — which this closes
- [ADR 0026](0026-the-mesh-has-a-session-of-its-own.md) — the mesh has a session of its own; a session
is the consumer the finer grain attributes to
- [ADR 0053](0053-a-step-that-runs-on-a-schedule.md) — a scheduled step; the licence-grain poll is one
- [ADR 0041](0041-events-are-a-relationship.md) — events are a relationship; a usage reading is one, and
the audit-logger records it
- [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md) — the shape of an event on the wire; the form a
usage reading takes to reach the audit trail
- [ADR 0008](0008-a-context-owns-its-store.md) — a context owns its store; the current usage picture
lives in one
- [ADR 0024](0024-model-access-is-a-provision.md) — model access is a provision; usage is a reading of
what that provision was used for
- [03-DESIGN/01-to-be/14-model-access.md](../03-DESIGN/01-to-be/14-model-access.md) — the model-access
design; this fills its usage section and shares its open worker-naming gap
- [03-DESIGN/01-to-be/15-the-agent-session.md](../03-DESIGN/01-to-be/15-the-agent-session.md) — the agent
session; the consumer the session grain is keyed to
@@ -0,0 +1,146 @@
---
topic: model access
status: accepted
date: 2026-09-07
deciders: jochen
reconstructed: false
extends: 0050-model-access-is-vendor-agnostic.md
---
# 55. Model access is answered by a licence, or by a node that hosts the model
## Context
**[ADR 0024](0024-model-access-is-a-provision.md) made model access a provision, and
[ADR 0050](0050-model-access-is-vendor-agnostic.md) fixed what answers it: a record — a licence — that
an adapter turns into a sealed vendor credential.** A consumer requires `model-access`, is put on a
licence, and is delivered a key (a static API key, or an access token a manager refreshes). Every
answer so far has been a credential to reach a vendor's API across the internet.
**But a model need not come from a vendor. A node in the mesh can host one.** An operator with a GPU
runs Ollama or vLLM, which serves an OpenAI-compatible API on that node. A consumer that wants that
model does not need a vendor credential — it needs the model server's **endpoint**: the base URL and
the model name, and a key only if the server is configured to want one. This is model access answered
by a **node**, not by a record.
**The mesh already knows how a node answers a provision — it is the ordinary provider/consumer path.**
A provider `provides` a provision at a scope, `serves` its connection facts, and the mesh fills the
consumer's bound facts with the provider's `at`/`port` and each served fact — exactly how a Postgres
consumer learns where its database is. Nothing reserved `model-access` to records: a node offering
`provides: ["model-access"]` resolves through this path, and the resolver already **prefers a local
answer over a licence** — its own comment names the case, "a model the mesh runs itself." So the
capability exists; what is missing is the decision to use it, and the statement of what a node-answer
delivers and where it stops.
**A node-answer and a record-answer are the same provision with two shapes of answer.** This is the
same move [ADR 0050](0050-model-access-is-vendor-agnostic.md) already made for shapes within the
vendor path (static-key vs refreshable-grant): one provision, more than one way it is answered. A
consumer written against `model-access` should not care whether the model behind it is a vendor's or
the mesh's own — it asks for model access and is given what reaches a model.
## Considered Options
1. **A separate provision for the local case (`local-model`, `model-endpoint`).** A node answers that;
`model-access` stays record-only. Rejected: it splits "where my model comes from" into two
provisions a consumer must choose between in its manifest, when the mesh already models a
record-answer and a node-answer to **one** provision. A consumer would have to know, at authoring
time, whether its model will be a vendor's or the mesh's — the exact coupling the provision was
meant to remove. It is the safer implementation (see the limitation below) but the worse interface.
2. **Model access answered by either a licence or a node, under the one provision.** A consumer
requires `model-access`; the operator answers it with a licence (a vendor) or by assigning a
node that hosts a model. Adopted: one interface, and the answer is an operator's deployment choice,
not a consumer's authoring choice.
## Decision
**Model access is one provision answered two ways: by a licence (a record, turned into a sealed vendor
credential by an adapter — [ADR 0050](0050-model-access-is-vendor-agnostic.md)) or by a node that hosts
the model (a provider that serves an endpoint).** A consumer requires `model-access` and is delivered
whichever the operator assigned; it does not name the kind.
**A node-answer delivers an endpoint, not a credential.** The provider `provides: ["model-access"]`
and `serves` its connection facts — the port it listens on and the model it runs — and the mesh fills
the consumer's bound facts with the provider node's `at`, the served `port`, and the served `model`,
the same way every provider consumer learns where its provider is. The consumer assembles a base URL
(`http://<at>:<port>/v1`) and points an OpenAI-compatible client at it. If the local server wants a
key, the provider mints one the ordinary way (a per-consumer secret, sealed and host-unsealed); if it
does not — the common Ollama case — the consumer lists `model-access` under `binds` and **not** under
`secrets`, and no key is delivered. Secret delivery and fact delivery are already independent, so a
keyless endpoint is expressed by asking for the facts and not a secret.
**A node-answer uses no adapter.** The adapter registry ([ADR 0050](0050-model-access-is-vendor-agnostic.md))
is the vendor-credential machinery — accept-and-seal, refresh, usage. A node-hosted model has no vendor
secret to seal; its endpoint is served, and its key (if any) is minted like any provider's. The adapter
is consulted only for the record/vendor answer. So the vendor-agnostic decision is untouched, and the
node-answer adds no vendor logic anywhere.
**The resolver prefers a local answer.** When a node's own set answers `model-access` — a model the
mesh runs itself — a licence for it is not consulted. This is already the resolver's behaviour and is
made a decision here: a mesh that runs a model uses it, and a licence is the answer for a consumer that
has no local model, not a competitor to one that does.
**One limitation, stated so it is not found as a bug.** The local-preference above is exact for a
**node-scope** provider co-located with its consumer, and for any node-answer in a mesh that holds no
`model-access` licence. It is *not* yet exact for a **mesh-scope** model server — one node serving the
model to others — **while a licence for `model-access` also exists in the same mesh**: the record pass
that turns a licence into an answer keys on same-node satisfaction and would still demand the licence be
used, double-answering. Until that pass is taught to stand down when a brokered node need already
answers, a mesh-scope local model and a vendor licence must not both answer `model-access` in one mesh.
A node-scope local model has no such constraint. This is named because an unstated limitation is
indistinguishable from a bug, and costs more.
This is a decision and not a patch because it settles **what may answer model access** — a question
every model-access consumer's meaning depends on — and because it lets the mesh's own hosted models sit
behind the same provision as the vendors', which is what makes "the mesh can run its own model" a
deployment choice rather than a second interface to build against.
### How each claim is checked
- **A node answers model access without a licence, and is preferred over one.** A resolver test
assigns a consumer and a module that `provides: ["model-access"]` at node scope on the one node, with
a licence also present, and asserts no resolved need is answered by the record — the local model
answers and the licence is ignored. (This test exists; the decision adopts what it proves.)
- **The consumer is delivered an endpoint, not a credential.** A mesh bed assigns a model-server
provider and a consumer that binds `model-access` and does not list it under `secrets`, and asserts
the consumer's config carries `OPENAI_BASE_URL` built from the provider's served `at`/`port`, and
that no key file was delivered to its secret path.
- **The local endpoint is reachable through what the consumer was given.** The bed makes a request to
the base URL the consumer wrote and asserts the model server answers — the wiring, not a model's
output, is what is proven (the server may be a stub; a real model is not needed to prove the mesh
routed the consumer to it).
- **A node-answer consults no adapter.** A node-answered `model-access` need is resolved with the
vendor registry never read — asserted by the absence of any vendor on a node-answered need and the
ordinary served-facts delivery.
- **The scope limitation holds where stated.** The node-scope case is what the bed and the resolver
test exercise; the mesh-scope-plus-licence collision is recorded here and left for the resolver
change that reconciles the two answer passes, not worked around in a module.
## Consequences
- **The mesh can run its own model, and a consumer reaches it through the same `model-access` it uses
for a vendor.** One interface, two answers; a consumer moves between a vendor and a local model by an
operator reassigning its provision, not by a code change.
- **A local model is keyless by default and keyed by the ordinary path when it must be.** Nothing new
is invented for the local server's credential: it either has none, or mints one the way every
provider does.
- **The vendor path is untouched.** Adapters, the refresh carve-out, and usage
([ADR 0054](0054-model-usage-is-recorded-at-two-grains.md)) are the record answer's business; a
node-answer neither uses nor changes them. A local model that exposes usage would serve it as facts,
not as an adapter's usage verb.
- **The two answer passes meet in one place, and must be reconciled there.** The mesh-scope limitation
is the single point where a node-answer and a record-answer to the one provision can collide; it is
named, and its fix is a resolver change, not a per-module workaround.
## References
- [ADR 0024](0024-model-access-is-a-provision.md) — model access is a provision; this decides a node
may answer it, not only a record
- [ADR 0050](0050-model-access-is-vendor-agnostic.md) — model access is vendor-agnostic; the record
answer and its adapters, which the node answer sits beside and does not use
- [ADR 0054](0054-model-usage-is-recorded-at-two-grains.md) — model usage; a vendor's business on the
record path, served as facts (if at all) on the node path
- [ADR 0008](0008-a-context-owns-its-store.md) — a context owns its store; a licence is a record in one,
a node-answer needs none
- [03-DESIGN/01-to-be/14-model-access.md](../03-DESIGN/01-to-be/14-model-access.md) — the model-access
design, which this extends with the node answer
+89 -4
View File
@@ -15,7 +15,7 @@ The records run in the order the decisions were taken, oldest first.
**Every decision is a record.** There is no ledger and no index file — if a decision is worth
recording it is worth a record, and if it is not worth a record it is not recorded
([ADR 0026](0026-every-decision-is-a-record.md)). A "decision" small enough to be one line is
([ADR 0019](0019-how-this-repository-works.md)). A "decision" small enough to be one line is
almost always a **rule**, and a rule belongs in
[`00-META/how-we-build.md`](../00-META/how-we-build.md), where it is enforced and keeps the
incident that earned it.
@@ -62,6 +62,91 @@ rather than guessing.
## Index
The index is **generated, not maintained** — run the `hq-status` skill, which reads the
frontmatter of every record. A hand-written index drifts from the folder it describes, and
this one had already done so after a single addition.
**A number identifies a record and never changes.** Records are referenced from outside this
repository — code comments, commit messages — so a number that moves invalidates them silently.
Renumbering once cost 96 references across two code repositories, and that is why the numbers
are now fixed.
So the folder is in creation order, and **the reading order lives here.** It is generated from
each record's `topic:` and written, because a reader looking at the folder on a forge sees the
folder rather than a command. The objection to a written index is that it drifts — which is
answered by checking it rather than by refusing to write one:
```
python3 00-META/checks/index.py --write regenerate
python3 00-META/checks/index.py fail if stale
```
<!-- index:start -->
### What the mesh is
- **0001** — [The mesh brokers capabilities; nodes host; agents think](0001-mesh-brokers-nodes-host-agents-think.md)
- **0002** — [Nodes communicate over a message broker, not over HTTP](0002-nodes-communicate-over-a-broker.md)
- **0003** — [An agent is a persistent employee, not an instance of a pool](0003-agents-are-persistent-employees.md)
### Its tiers, from the bottom up
- **0004** — [A node, and how it joins](0004-a-node-and-how-it-joins.md)
- **0005** — [The node host](0005-the-node-host.md)
- **0006** — [The substrate and the control plane](0006-the-substrate-and-the-control-plane.md)
- **0007** — [Connectivity](0007-connectivity.md)
- **0008** — [A context owns its store, exclusively](0008-a-context-owns-its-store.md)
- **0028** — [The substrate supplies the control plane and nothing else](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md)
- **0029** — [A network is a shape, because an action cannot be undone](0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md)
- **0030** — [Data outlives the mesh that declared it](0030-data-outlives-the-mesh-that-declared-it.md)
- **0031** — [The control plane authenticates nobody, so identity is a module](0031-the-control-plane-authenticates-nobody.md)
- **0033** — [The substrate is a store and a broker](0033-the-substrate-is-a-store-and-a-broker.md)
- **0036** — [Bootstrap ends at a usable mesh, and the first credential comes from a person](0036-bootstrap-ends-at-a-usable-mesh.md)
### What runs on them, and how it gets there
- **0009** — [Modules and the graph](0009-modules-and-the-graph.md)
- **0010** — [Delivery](0010-delivery.md)
- **0024** — [Model access is a provision, and a licence is a thing with a name](0024-model-access-is-a-provision.md)
- **0026** — [The mesh has a session of its own, and it is the node session's mechanism](0026-the-mesh-has-a-session-of-its-own.md)
- **0027** — [A provision names what the consumer is coupled to, not the role it plays](0027-a-provision-names-what-the-consumer-is-coupled-to.md)
- **0035** — [One implementation, several surfaces, and what that costs](0035-one-implementation-several-surfaces.md)
- **0038** — [The mesh assigns the port, and a module does not care](0038-the-mesh-assigns-the-port.md) *(proposed)*
- **0040** — [What a module is](0040-what-a-module-is.md)
- **0041** — [Events are a relationship, the lighter sibling of provisioning](0041-events-are-a-relationship.md)
- **0042** — [The shape of an event on the wire](0042-the-shape-of-an-event-on-the-wire.md)
- **0043** — [A module's broker account is scoped by what it emits and consumes](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md)
- **0044** — [A public name is provisioned, not registered by hand](0044-a-public-name-is-provisioned-like-any-capability.md)
- **0045** — [A machine's firewall is the sum of what its modules listen on](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)
- **0046** — [A module's configuration is its assignment's, not its manifest's](0046-a-module-configuration-is-its-assignments-not-its-manifest.md)
- **0047** — [A module runs its code as its own process, with its own account](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)
- **0048** — [A provider creates the credential the mesh minted, and seals nothing](0048-a-provider-creates-the-credential-the-mesh-minted.md)
- **0049** — [A consumer's identity is bounded by the tightest backend that must accept it](0049-a-consumers-identity-fits-the-tightest-backend.md)
- **0050** — [Model access is vendor-agnostic, and a vendor is an adapter](0050-model-access-is-vendor-agnostic.md)
- **0051** — [Shared data is the operator's, and a module is granted access to it](0051-shared-data-is-the-operators.md)
- **0052** — [An init step is a container run once to completion, gating what follows](0052-a-step-that-runs-once-before-a-container.md)
### How it is built
- **0011** — [Managed files are generated onto nodes and never edited there](0011-managed-files-are-generated-never-edited.md)
- **0012** — [The mesh creates no symlinks — a derived file is a copy](0012-the-mesh-creates-no-symlinks.md)
- **0013** — [Schema and state changes are numbered migrations, in the same language as the code](0013-schema-changes-are-numbered-migrations.md)
- **0014** — [No workspace — each module is a standalone package consuming published dependencies](0014-no-npm-workspace.md)
- **0015** — [Applications live in their own repository; the monorepo is for the mesh](0015-applications-live-in-their-own-repository.md)
- **0016** — [The lab](0016-the-lab.md)
- **0037** — [Where a module lives](0037-where-a-module-lives.md) *(proposed)*
- **0039** — [What the SDK holds, and what it refuses](0039-what-the-sdk-holds-and-refuses.md)
### How it is checked
- **0017** — [A test defends a decision](0017-a-test-defends-a-decision.md)
- **0018** — [A picture of a system is read from the system, never from what asked for it](0018-a-picture-is-read-from-what-runs.md)
### How we work
- **0019** — [How this repository works](0019-how-this-repository-works.md)
- **0020** — [The mesh is governed by a constitution, injected where work is decided](0020-the-mesh-is-governed-by-a-constitution.md)
- **0021** — [HQ is the source of the mesh constitution](0021-hq-is-the-source-of-the-constitution.md)
- **0022** — [The constitution absorbs what is already enforced](0022-the-constitution-absorbs-what-is-enforced.md)
- **0023** — [The approval is the checkpoint, not the second pair of hands](0023-approval-is-the-checkpoint.md)
- **0025** — [The design record is read where it is written, never copied to be found](0025-the-design-record-is-read-not-copied.md)
- **0032** — [The local account owns the mesh; a surface delegates to a module](0032-the-local-account-owns-the-mesh.md) *(superseded)*
- **0034** — [The local account owns the mesh, and a web application's login is not that](0034-the-local-account-owns-the-mesh.md)
<!-- index:end -->