Compare commits

..
Author SHA1 Message Date
jschoubben 0b9fc90885 ADR 0188: a provider declares what it derives for each consumer, and the mesh tells both ends
Issue 124: a value the mesh's own rule produced reached neither end as a
statement. The object store's provisioner derived each consumer's bucket in
its own code; all three consumers transcribed the rule into their own
definitions, one of them wrong, and each of the three also named the machine
it happens to run on.

A served value may now name the consumer the mesh is serving. Design 27
amended; issue 124 resolved.
2026-10-02 21:24:31 +02:00
62 changed files with 206 additions and 3263 deletions
+1 -1
View File
@@ -108,7 +108,7 @@ term retired here may still appear there, and the mapping above is how to read i
memberships issue; its serving mode on loopback is what was called **the console**
([ADR 0175](../02-DECISIONS/0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md)).
Replaces **"console"** as the module's name; *console* remains the word for the person's end of it.
- **bundle** — the artifact a module's own code is built into — its tools, a seat's implementation, a daemon — in any language the mesh has a toolchain for, interpreted or compiled; never an image. One module may declare several ([ADR 0188](../02-DECISIONS/0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md)).
- **bundle** — the artifact a module's tools are built into, interpreted or compiled; never an image.
- **kept region** — a marked block in a managed file the mesh writes *into*, where the operator's own
lines survive every push and are given back when the module goes
([ADR 0174](../02-DECISIONS/0174-a-node-varies-a-module-through-settings-and-kept-regions-never-an-edit.md)).
@@ -1,34 +0,0 @@
---
status: graduated
initiated: 2026-10-03
touches: [the tool runtime, the catalogue's tool bundles, the controller's declaration composer, settings, own secrets, 03-DESIGN/01-to-be/38-building-the-operators-machine.md]
became: [02-DECISIONS/0192-a-tools-bundle-declares-what-it-is-given-and-the-runtime-hands-it-to-that-bundle-alone.md, 03-DESIGN/01-to-be/38-building-the-operators-machine.md]
---
# 020 — What a bundled tool is given
## What is being investigated
How a module's tools, once they are a bundle the node's runtime loads
([ADR 0175](../../02-DECISIONS/0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md),
[ADR 0188](../../02-DECISIONS/0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md)),
learn the things their container used to be handed: where the module's configuration file is, where
its token or password is, which port the service listens on, where a provision's address is written.
A container is given these as an environment and mounts, composed by the mesh per module per machine.
A bundle has no environment of its own: the runtime's process carries four words for every bundle it
loads, and nothing per module ([design 38](../../03-DESIGN/01-to-be/38-building-the-operators-machine.md) WP4).
## Why
Two holders moved on 2026-10-03 — the packet filter and the intrusion prevention — and both could,
because neither needs anything but a fixed path and root. Of the thirty-three modules whose tools
still run as containers on the runtime's image, thirty-one are not like that: their environment
names a configuration file, a credential file, a service address, a grants directory. Moving them
one by one without a rule for this would give the mesh thirty-one answers to one question. The
measurement and the options are in [01](01-what-the-containers-are-given.md).
## What it touches
The runtime (which hands a bundle what it is given), the composer (which resolves `${dir:…}` and
`${port:…}` for a container today and would for a bundle), the manifest (where a bundle would say
what it needs), and design 38, which records the gap and must say the rule once there is one.
@@ -1,86 +0,0 @@
# What the tool containers are given, measured
Counted 2026-10-03 in the catalogue, after the two holders moved.
| | |
|---|---|
| modules whose tools still run as a container on the runtime's image | 33 |
| tool containers among them (two modules run two) | 36 |
| modules whose container's environment carries only the bus credential | 1 (the intrusion prevention, now moved) |
| modules whose container's environment carries more | 32 — 31 still containers |
## What "more" is
Every value a container is given is one of five shapes. The reference kinds the composer resolves
in those values, over the 36 containers: a module directory (`${dir:…}`) in all 36, a mesh-chosen
port (`${port:…}`) in 12, a seat and an access grant once each.
1. **A file the mesh already places on the host, mounted in.** The module's configuration as
JSON (`…_CONFIG_FILE`), its own secret (`…_TOKEN_FILE`, `…_PASSWORD_FILE`, `MESH_BROKER_FILE`),
a provision's address and secret written for it. Every one is a path under one of the module's
directories — its mesh state, its state, its grants, what it has written — mounted at a path of
the container's choosing and named to the tool through the environment. **The file is on the
host already; only the name under which the tool finds it is the container's.**
2. **The service's address, with the port the mesh chose:** `http://127.0.0.1:${port:3000}`. The
port is the composer's; the rest is the manifest's constant.
3. **A provision's address as a constant string** (a database's URL on the module's own network
name), paired with a mounted secret file from shape 1.
4. **A directory of grants** (`MESH_RECEIVES`): shape 1 again, a directory rather than a file.
5. **Literals the image needs:** a time zone, a user id, a memory limit. These belong to the
service's container where one exists; a tool bundle needs none of them.
So the whole of what a bundled tool needs is: the paths of its module's directories on this
machine, the ports the mesh chose for its module here, and the constants its own manifest wrote.
Nothing a container had that a bundle cannot have; the mesh composes all three for the container
today, per module per machine.
## What the runtime already has for it
- The SDK's tool contributor is `(env) => tools`, and `collectTools(env)` takes the environment to
hand each contributor. The runtime calls it without one, so every contributor reads the process's
— the four words. The hook for a per-module environment exists and is unused.
- A launched bundle ([ADR 0188](../../02-DECISIONS/0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md))
is spawned with the runtime's environment; the launch takes an environment argument.
- The composer resolves `${dir:…}` and `${port:…}` for a container's `env` and `volumes`; the same
resolution over a bundle's declaration is the same code.
## Options
**A. The bundle declares its environment on its artifact, and the mesh composes it as a
container's.** The manifest's tools artifact gains `env`, resolved with the same references;
values that were mount targets become the host-side paths directly (`${dir:mesh-state}/config.json`
rather than `/run/config/config.json`). The controller composes one environment per bundle per
machine into the runtime's declaration; the runtime hands it to the bundle's contributor and to a
launched child, and to nothing else. *For:* the tool code does not change — it reads the same
names; the conversion of the thirty-one is a mechanical move of the container's `env` with the
mounts folded in; one rule, one place. *Against:* the runtime's process carries thirty-one
environments in its declaration, and a bundle's environment is visible to the other bundles in the
process unless the runtime keeps them apart, which it must — a tool that reads `process.env`
instead of the environment it was handed would see its neighbours' paths.
**B. The runtime derives the environment from the module's placed manifest.** No new field: the
runtime reads, for each module it serves, where that module's directories and ports are, and hands
a conventional set of words. *For:* nothing to declare. *Against:* a convention the tool code must
be rewritten to, thirty-one times; the runtime learns the composer's job; a module that names its
file `config.json` and one that names it `settings.json` need different words anyway.
**C. Tools read their module's files through the bus** — ask the controller. *Against:* a tool
that cannot start without the bus answering a question is a tool that fails in the one case the
tools exist for, and a secret crossing the bus to reach a file already on the machine is a
disclosure for nothing.
A is the one that keeps the tool code and the composer's vocabulary as they are, and names the one
thing the runtime must add: an environment per bundle, kept apart. The thing to decide beside it:
whether a bundle's environment may name a secret file at all, or whether secrets stay mounts in
spirit — a path the tool reads, never a value in the environment — which is what every container
does today and what A keeps if the rule says *paths, not values*.
## What a decision would have to say
- Where a bundle says what it is given (the artifact, option A), and that values are paths and
constants, never a secret's content.
- That the composer resolves it with the references it already has, per module per machine.
- That the runtime hands each bundle its own environment and nothing of another's, and how that is
checked: a test loading two bundles whose environments differ and asserting each sees only its own.
- That the thirty-one move in one mechanical change after the rule lands, each proven by its tools
answering from the runtime, and the registration gate then refuses the container shape for all.
@@ -1,48 +0,0 @@
---
status: graduated
initiated: 2026-10-03
touches: [the console, 03-DESIGN/01-to-be/34-the-console.md, the tool runtime, seats, assignments]
became: [02-DECISIONS/0195-the-meshs-tools-are-found-by-address-not-announced-whole.md, 03-DESIGN/01-to-be/34-the-console.md]
---
# 021 — Finding a tool in the mesh
## What was investigated
How an agent finds the one tool it needs among everything the mesh answers, and how a call names
exactly what it asks — a role the mesh holds once, a role every machine holds, or one assignment of a
module on one machine — rather than receiving the whole catalogue and a name that can mean several
things.
## Why
The operator's observation on 2026-10-03: *Claude should not see all tools at once; they should be
discoverable — and `postgres.list_databases` is wrong, asking one machine's postgres is not asking
another's.* Measured the same day from the console's own answer:
| | |
|---|---|
| tools announced to every session at its start | 228, in 110 KB |
| names (module or seat prefixes) | 43 |
| node seats' verbs, which require `node` | 22 |
| modules with tools on more than one machine | 4 — fail2ban, nftables (every machine), postgres, mssql (two each) |
| modules reported "not answering", most with no tools and several retired | 47 |
The two stateful modules on two machines are listed **once**, with `node` optional and *whichever
answers* when it is left out — though their two instances hold different databases. Design 34 §3 says
such a module is listed once per machine; the live console does not do that. The list is taken once
per session, so a tool that arrives later is invisible until the client reconnects. And only Claude
Code's own deferral of long tool lists keeps the 228 from the model's context; another MCP client
would receive them whole.
## Options
1. **Keep the flat list; rely on the client to defer it.** Rejected: a property of one client, and it
leaves the ambiguity and the stale list.
2. **One flat tool per assignment** (`ace_postgres_list_databases`). Removes the ambiguity, multiplies
the list, and runs into the API's tool-name limit (letters, digits, `_`, `-`, 64 characters).
3. **A small fixed set of tools that walk the mesh's own structure**, with the full address as an
argument: the mesh's seats; a machine's node seats and assignments; a search; a description; a
call. Chosen — see [ADR 0195](../../02-DECISIONS/0195-the-meshs-tools-are-found-by-address-not-announced-whole.md).
4. **MCP resources or prompts for discovery.** Clients support them unevenly, and an agent acts
through tools; a resource it cannot be relied on to read is not a discovery path.
@@ -1,41 +0,0 @@
---
status: graduated
initiated: 2026-10-03
touches: [the tool runtime, the per-module containers, the SDK, the bus grants, 03-DESIGN/01-to-be/38-building-the-operators-machine.md]
became: [02-DECISIONS/0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md, 03-DESIGN/01-to-be/38-building-the-operators-machine.md]
---
# 022 — Where a module's long-running code runs
## What was investigated
Twenty-three modules still run their own code in a container built on the runtime's image. Their tools
can move as bundles ([ADR 0192](../../02-DECISIONS/0192-a-tools-bundle-declares-what-it-is-given-and-the-runtime-hands-it-to-that-bundle-alone.md),
[ADR 0193](../../02-DECISIONS/0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md));
the rest of what those containers run cannot yet. This asks where that code goes and how it reaches
what its container handed it.
## What that code is, measured 2026-10-03
| | modules |
|---|---|
| subscribes to events on the bus | audit-logger (everything), mesh-catalog (two seat events), mesh-vault, records (`gitea.pull.merged`), and postgres, mongodb, mssql, redis, mosquitto logging their own lifecycle |
| provisioners: read the grants the mesh delivered as files, act on the backend, emit | 12 |
| a run-once preparation step | mesh-catalog |
| a command-line client of the backend | psql, mosquitto_ctrl, git (packages on every machine's system); mongosh, sqlcmd (not in its repositories) |
| a service reached by a container name | icecast, mailu-admin, minio, mongodb-server, mssql |
| a main of its own | anthropic-consumer, openai-consumer, route-adapter |
A provisioner needs nothing a launched bundle lacks: files named by its words, its backend, and an emit
that already travels through the runtime. The one thing missing is **a subscription** — events
delivered to the module's code, acknowledged when it has handled them.
## Options
1. **The runtime launches it and is its bus**: the stdio channel gains a subscription; the runtime
binds the module's durable consumer and delivers each event to the child, acknowledging when the
child answers. One bus connection per machine; any language. Chosen.
2. **A process per module with its own bus client and credential.** Every language's SDK would carry
a transport and every module a credential on disk — what ADR 0188 rejected for tools, for the same
reasons.
3. **Keep the containers for this code.** Leaves ADR 0188's rule broken for 23 modules indefinitely.
@@ -8,8 +8,6 @@ reconstructed: false
# 39. What the SDK holds, and what it refuses
> **The mechanism changed — 2026-10-02, by [ADR 0188](0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md).** The test of this record — frequent *and* cascading does not belong — stands and now applies to one SDK per language. Where *what it holds* names the broker client and the event consumer, read: the local protocol a tools bundle speaks to the node's runtime; the transport lives in the runtime and in no SDK, which is what keeps a bus change from rebuilding any module in any language.
_Reconciliation note (2026-09-05): supersedes the earlier "repository structure" decision, which the consolidation folded; no standalone record remains to point at, so body references to it now point at the nearest surviving record, [ADR 0015](0015-applications-live-in-their-own-repository.md)._
## Context
@@ -9,13 +9,6 @@ extends: 0007-connectivity.md
# 66. Public routing is name-agnostic, its names are resolved inside the mesh, and an internal authority can certify them
> **Narrowed, not replaced — 2026-10-03.** One clause of the decision below no longer holds: *publishing
> a granted name into internal resolution, mesh-wide*. A public name now resolves publicly, and only
> names under the mesh's own suffix get a private answer — [ADR 0191](0191-the-meshs-resolver-holds-only-the-meshs-own-names.md).
> Inside the mesh a route is reached and certified by its internal name
> ([ADR 0151](0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md)). The label, the
> node's public domain and their composition stand as decided here.
## Context
**[ADR 0007](0007-connectivity.md) and [connectivity §3](../03-DESIGN/01-to-be/08-connectivity.md)
@@ -9,14 +9,6 @@ extends: 0110-a-seat-is-a-module-assignment-from-a-closed-set.md
# 121. A system seat is named for its scope, and a module may define its own
> **Narrowed, not replaced — 2026-10-03.** *"`the-dns-port` → `node-dns-resolver`"* no longer holds:
> the serving role moves to mesh scope as `mesh-resolver`, one per mesh, and `node-dns-resolver` is
> retired ([ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)). The
> distinction this record kept — serving and asking are two roles, two seats — stands, and
> `node-resolver-config` is unchanged.
> **The mechanism changed — 2026-10-02, by [ADR 0190](0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md).** The naming rule stands. The build role this record made mesh-scoped — *the mesh's single build machine* — is node-scoped now: `node-build-agent`, one holder per machine, every holder taking from one work queue.
## Context
[ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md) made seats a closed set the
@@ -9,8 +9,6 @@ extends: 02-DECISIONS/0047-a-module-runs-its-code-as-its-own-process-with-its-ow
# 150. A module's own code runs as supervised processes under the module's one account
> **Widened — 2026-10-02, by [ADR 0188](0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md).** A module's long-lived process is a bundle in any language the mesh has a toolchain for, run as a unit the host writes; this record never said one language and never meant one, and 0188 says so as the rule.
> **The mechanism changed — 2026-10-02, by [ADR 0175](0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md).** For a module's *tools*, read that record: one runtime per node, the node's one account, bundles loaded from the memberships. This record still governs a module's long-lived processes — a daemon, a provisioner, a scheduled ingest — and the account invariant for them.
## Context
@@ -9,11 +9,6 @@ extends: 02-DECISIONS/0066-public-routing-is-name-agnostic.md
# 151. A route's internal name is composed under the node that serves it
> **Narrowed, not replaced — 2026-10-03.** *"The roster publishes it as itself, once, at the serving
> node's address"* no longer holds: a public name is never given a private answer, and resolves publicly
> ([ADR 0191](0191-the-meshs-resolver-holds-only-the-meshs-own-names.md)). The internal name this record
> composes is what that rests on, and stands.
## Context
A module that requires a route is given two names from one label: a public one, `<label>.<public
@@ -9,8 +9,6 @@ extends: 02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md
# 162. A merge produces a tiered plan the mesh keeps, and a module's dependencies are one relation in the catalogue
> **Progressive insight — 2026-10-02.** The context below says a dependent is *built by whichever build machine is running — the only one there could be*. That was a fact of the day, not of the decision: since [ADR 0190](0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md) a tier's asks are taken by every machine holding the build seat. The plan and its tiers are unchanged.
## Context
A merge on the forge reaches the controller as an event, and the controller asks the build
@@ -1,146 +0,0 @@
---
topic: what runs on it
status: proposed
date: 2026-10-01
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md
---
# 164. A setting is declared with its default, its meaning and what changing it costs
## Context
The operator asked for one thing for every module, with the container runtime as the first case: **one
consistent default configuration for every machine, overridable per assignment, and easy to change
later.** The four machines' runtime configurations were each written by hand and differ — one keeps
running containers through a daemon restart and one does not, their log rotation differs, and each
names its resolver and its trusted registries in its own words.
Most of this was already decided.
[ADR 0046](0046-a-module-configuration-is-its-assignments-not-its-manifest.md) said the definition is
identity and **defaults**, the assignment's settings are the configuration, *unset is the default*, and
*an unknown setting is refused*. [ADR 0112](0112-a-module-definition-names-no-node-mesh-or-path.md) made
a setting an operator requirement whose contract is "a type and, optionally, a default", answered by
"the assignment's, or the requirement's default, or unresolved". An assignment is a module on a node,
so the node layer of a module's settings already *is* the per-assignment override, and the mesh-wide
layer is the one consistent default a person changes once.
What was built is narrower than what was decided, measured in the controller on the day of deciding:
- **Nothing declares which keys are settable.** A mergeable file's content is its defaults, and every
key of it — and every key not in it — is accepted. Nothing tells a person, or the console, what can be
set, of what type, or what it means.
- **The refusal of an unknown setting is not there for most modules.** The stray-setting report returns
nothing at all for a module with any mergeable file, because such a file "takes any key"
([issue 173](../04-ISSUES/173-a-modules-settings-reach-every-fact-it-contributes/00-report.md) left
files that way on purpose). It reports rather than refuses where it does run.
- **A value in a file that is not JSON can have no default.** `${setting:<key>}`
([ADR 0155](0155-a-definition-names-no-installation-and-how-that-is-checked.md)) is refused when no
layer sets it — right for a mail domain, where a default is the very literal 0155 removes, and wrong
for a tunable like the resolver's upstreams, which the resolver module therefore carries as literals
in its file.
- **A setting reaches every mergeable file its module owns.** The layers are one flat map per module,
laid over each such file. Adding a setting to the resolver module for its own configuration put the
key into the container runtime's file as well — the resolver writes into that file too — and the
runtime refuses keys it does not know. The plan showed it before any push; the runtime's file was
then made to take no settings at all ([issue 198](../04-ISSUES/198-the-lans-dns-server-ran-outside-the-mesh-and-its-filter-closed-it/00-report.md)). Issue 173 stopped settings leaking into
contributions and served facts; between one module's own files the leak remains.
- **What a change costs is said per file, not per key.** A service names the files it is reloaded or
restarted on. The runtime re-reads its trusted registries on a reload and its `dns` key only when it
starts; the resolver module declared a reload, so on two machines the key was written, reloaded,
and never read, and every container got a public resolver for weeks while everything read as
current ([issue 110](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/01-resolution.md)).
## Considered Options
1. **Leave settings implicit; document each module's keys in its README.** Rejected: a key the mesh
does not know cannot be refused, typed, listed by the console or costed, and a README is a rule
enforced by nothing.
2. **A second mechanism for tunables beside settings** — defaults in a new block, settings untouched.
Rejected: two ways to state one person's value, and design 27 already retires six mechanisms
that grew that way.
3. **Settings declared in the definition, as 0112's operator requirement: a key, a type, a meaning,
optionally a default, and what a change costs.** Adopted.
## Decision
**A module declares every setting it takes.** Each declared setting has a name, a type, one sentence
of meaning, optionally a default, and what a change to it costs. The spelling is design 27's to settle
with the rest of the requirement form; this record decides the content.
**A setting with a default is a tunable; a setting without one is the operator's.** A tunable resolves
to its default when no layer sets it — wherever it is read, a mergeable file or `${setting:<key>}` in a
file of any format. A setting with no default is refused by name when nothing sets it, as 0155 decided;
0155's refusal is narrowed to exactly that case, not changed for it. Whether a value has a default is a
fact about the software (a log size does, a mail domain does not), and the definition states it once.
**The layers stay as they are, and every value says where it came from.** The definition's default,
then the mesh-wide layer, then the node's — later wins, objects merge, lists replace. One consistent
configuration for every machine is the default plus the mesh-wide layer; one machine that differs says
so in its own layer and nothing else. Asked for a module's configuration on a machine, the mesh lists
every declared setting with its effective value and its source: *default*, *mesh*, or *node*.
**Changing later is changing one of three places, and the plan shows its reach before anything moves.**
A new default ships with the module's next version and reaches every assignment that does not override
it; a mesh-wide setting reaches every assignment of the module; a node's reaches one. The plan of a
change names each assignment whose effective value moves.
**A declared setting says where it lands.** Each names the file or files of its module that read it,
and reaches no other: a module that owns two mergeable files no longer has one flat map laid over both.
A file that names no setting takes none.
**A declared setting is the only kind accepted.** Setting a key the module does not declare is refused
when it is set, naming the declared keys, rather than reported when the machine is planned. The mesh's
own words — where a port, a directory or an operator's data is placed, how far an endpoint reaches —
are the mesh's to validate as they are today, and no module declares them. A module
that declares no settings keeps today's behaviour until it does; a catalogue test lists those modules,
and the list shrinks to empty before the implicit form is removed — design 27's rule for every retired
mechanism.
**A setting says what it costs: nothing, a reload, or a restart.** When a file changes, the host
applies the strongest cost among the settings whose values moved in it, so a key the software reads
only at start can no longer be written and never read. A setting that reaches a container's environment
costs that container being recreated, which the host already does when a container's specification
changes; it needs no declaration. A service's `reload-on` and `restart-on` keep
naming the files that are not settings — a generated roster, a credential.
**The container runtime is the first module to declare its settings** and the model for the rest:
its log rotation, keeping containers through a daemon restart, and its resolver are tunables, and
its trusted registries are what the mesh tells it.
## Consequences
- The console can show a module's settings as a form: what can be set, of what type, its default,
and where the current value came from. That is the surface the operator wants for changing a
default later.
- `settings set` can refuse an unknown key, so ADR 0046's rule is enforced where it was only stated.
- The resolver's upstreams, the runtime's log rotation, and other literals a definition carries
because it could not give them a default become declared tunables.
- **What got harder:** every module that takes settings must list them, and a mergeable file no
longer silently accepts a key its author did not foresee. A person who needs one adds it to the
definition, which is a new module version, not a setting.
- Issue 173's open question — a consumer checks nothing against a contract — is unchanged; this record
is the operator half of design 27's contract, not the provider half.
- Not decided here: the spelling (design 27); whether a node's layer may be narrowed to a single key
rather than replaced whole, as `settings set` does today.
## How this is checked
| Rule | Checked by |
|---|---|
| Every setting a module takes is declared | A parser test refusing a setting declaration without a type or meaning; a catalogue test listing modules with mergeable files or `${setting:}` and no declarations, which must be empty before the implicit form is removed |
| A tunable resolves to its default; an operator value without one is refused | Resolution tests: an unset tunable in a JSON file and in a text file both take the default; an unset setting with no default is refused naming it (0155's existing test) |
| A setting reaches only the files it names | A resolution test: a module with two mergeable files and a setting declared for one; the other file's content is unchanged by it (the case of issue 198) |
| An undeclared key is refused when set | A controller test: `settings set` with an undeclared key fails naming the declared keys, and nothing is stored |
| Every effective value names its source | A test listing a module's configuration on a node with one key from each of default, mesh and node |
| A change's reach is shown before it moves | A plan test: changing a mesh-wide setting names every assignment whose effective value moves and no other |
| The strongest cost applies | A host test: a file where a reload-cost key and a restart-cost key both moved restarts; a file where only reload-cost keys moved reloads |
| Live | The container runtime's module lists its settings with their sources on every machine, and a mesh-wide change to its log rotation reaches all four at the next push |
## References
- [ADR 0046](0046-a-module-configuration-is-its-assignments-not-its-manifest.md), [ADR 0112](0112-a-module-definition-names-no-node-mesh-or-path.md), [ADR 0155](0155-a-definition-names-no-installation-and-how-that-is-checked.md), [ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md)
- [Design 27 — a module requires, the mesh resolves](../03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md)
- Issues [110](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md), [173](../04-ISSUES/173-a-modules-settings-reach-every-fact-it-contributes/00-report.md)
- mesh-controller `internal/catalogue/settings.go` (`settle`, `UnusedSettings`), `internal/catalogue/setting_into.go`
@@ -1,105 +0,0 @@
---
topic: what runs on it
status: proposed
date: 2026-10-01
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0161-what-deserves-a-seat.md
---
# 165. `container-runtime` is what a machine can run; that a runtime is running is its holder's health
## Context
A capability is a requirement a module places on a machine, detected by the host and renewed with
every report ([ADR 0161](0161-what-deserves-a-seat.md)). The host's `container-runtime` asks the
daemon for its version: *a running daemon, not an installed client*. It was made that way by
[issue 007](../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md), where an
installed package was believed to be a working service, and
[design 05](../03-DESIGN/01-to-be/05-the-node-host.md)'s table says the same: *a runtime is
running*. The installer's preflight borrows the same detector to wait for the runtime the
foundation bundle installs, so there is one answer to "is there a runtime here".
The mesh is now to have a module for the runtime itself — its packages, its configuration, its
service — on every machine ([ADR 0166](0166-the-container-runtime-is-a-node-seat-and-the-host-creates-containers-through-its-holder.md)).
That module cannot declare `container-runtime` as defined: it would require the very thing it
installs, the cycle [research 011](../01-RESEARCH/011-the-module-graph/cases.md)'s case 12 names
("something the mesh installs that then becomes a node capability"). The operator defined the word
for it: **`container-runtime` means the machine is able, at the kernel level, to install a runtime and
execute containers** — not that one is installed, and not that one is running.
The host already draws this line once. `seat` is hardware, a display server *could* run here;
`graphical-session` is state, one *is* running; the detector's own comment says "assignment needs the
first". A machine without a display has no seat however much software is installed, and a machine
with one has a seat before anything is.
Fifty-four catalogue modules declare `container-runtime` today, counted on the catalogue's main
branch on the day of deciding: every module that delivers a container. Each relies on the current
meaning to keep it off a machine with no running runtime.
## Considered Options
1. **Keep the meaning; let the runtime's module declare nothing.** Rejected: a module that installs
the runtime has requirements on the machine — the kernel features without which installing it is
pointless — and would state none of them. The cycle stays, only hidden.
2. **Two capabilities, "can run" and "is running".** Rejected: the second is made true by assigning a
module, so it is the module's state, not a fact of the machine; a capability the mesh itself
flips by its own assignment is case 12's cycle with an extra name.
3. **The capability is the kernel's; whether a runtime runs is the runtime module's health, and a
module that delivers a container needs the runtime's seat held.** Adopted.
## Decision
**`container-runtime` is detected from what the kernel offers**, as `seat` is: the namespaces a
container needs, a control-group hierarchy the runtime can manage, and an overlay filesystem the
running kernel has or can load. Present when all three are; absent naming the missing one. Nothing is
run and no runtime is asked. The verdict's detail names what was found, not a runtime's version.
**"A runtime is running and answers" is one probe, owned by the host and used twice:** by the
installer's preflight, which waits for the runtime the foundation installs, and as the runtime
module's health. It asks the daemon, as issue 007 requires. The preflight stops borrowing the
capability's detector, and there is still one answer to "is a runtime running here".
**The runtime's module declares `container-runtime`**, with `package-manager`, `service-manager` and
`privileged`, like any module that manages machine software.
**A module that delivers a container needs the runtime seat held on its machine**, and is refused
otherwise, naming the seat and the modules that could hold it — the refusal design 27 already lists
for an unheld seat. That requirement is derived from the container resource and needs no manifest
field ([ADR 0166](0166-the-container-runtime-is-a-node-seat-and-the-host-creates-containers-through-its-holder.md)).
The fifty-four existing declarations of the capability stay valid and become redundant; a catalogue
test lists them, and they retire when the list is empty.
**The order is fixed, not preferred.** The detector changes only once the seat requirement is
enforced. In between, a machine with the kernel and no running runtime would read as able to run
every containerised module, which is issue 007 again.
## Consequences
- Design 05's capability table changes its `container-runtime` row from *a runtime is running* to
*the kernel can run containers*, and names the runtime module's health as where "running" is now
asked.
- The node listing stops showing the runtime's version beside the capability. The version moves to
the runtime module's health and its seat's verbs.
- A fresh machine with no runtime reads as able to run one, so it can be assigned the runtime's
module, which is what makes the mesh able to install the runtime instead of the bootstrap alone.
- **What got harder:** "is this machine running containers" is no longer one glance at the profile;
it is the runtime seat's holder and its health. The node's listing should show both side by side.
## How this is checked
| Rule | Checked by |
|---|---|
| The capability is the kernel's | Host detector tests over a fixture `/proc` and `/sys`: all three present → present; each one missing → absent naming it; no runtime binary on the fixture machine changes nothing |
| One probe asks whether a runtime runs | A host test that the preflight and the runtime module's health call the same probe, and that the probe fails against a stopped daemon with an installed client (issue 007's shape) |
| A containerised module needs the runtime seat held | A resolution test: a module with a container resource on a machine whose runtime seat is unheld is refused, naming the seat and its candidate holders |
| The order holds | The host release that changes the detector is gated on the controller release that enforces the seat requirement — stated in both changes' descriptions and checked at review |
| Live | Every machine's profile shows `container-runtime` present with the kernel's features as its detail; a machine with no runtime installed reads present too |
## References
- [ADR 0161](0161-what-deserves-a-seat.md) — the profile renewed by every report; a capability that names a dialect
- [ADR 0166](0166-the-container-runtime-is-a-node-seat-and-the-host-creates-containers-through-its-holder.md) — the seat and its holder
- [Issue 007](../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md), [research 011](../01-RESEARCH/011-the-module-graph/cases.md) cases 12–13
- [Design 05 — the node host](../03-DESIGN/01-to-be/05-the-node-host.md)
- mesh-host `internal/profile/detectors.go` (the runtime detector), `internal/profile/seat.go` (the hardware/state split), `internal/bootstrap/preflight.go` (the preflight that borrows it)
@@ -1,161 +0,0 @@
---
topic: what runs on it
status: proposed
date: 2026-10-01
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0161-what-deserves-a-seat.md
---
# 166. The container runtime is a node seat, and the host creates containers through its holder
## Context
Every container the mesh runs on a machine is created by the host, which looks for a runtime
(`docker info`, then `podman info`) and drives that runtime's command line itself: run, inspect,
remove, exec. Research 012 called this "detected rather than declared": the host takes over whatever
runtime it finds. Nothing in the mesh owns the runtime. Its package came from the foundation bundle
or was already on the machine. Its configuration file was written by hand, differs on each of the
four machines, and is also written into by two modules that are not the runtime's
([issue 190](../04-ISSUES/190-the-runtimes-configuration-is-written-by-modules-that-are-not-the-runtime/00-report.md)).
Its service is declared by those same two.
The operator set the direction:
- a module for the runtime, on every machine, owning "what is needed to run containers here": its
packages, its configuration and its service;
- that module holds a node seat for the runtime, so a second runtime (podman) can later compete
for the seat;
- the host stops speaking to the runtime directly and uses the seat's holder. The host stays the one
that decides, and the holder becomes the one that executes;
- every container on the machine is in scope, not only the mesh's. A development environment started
by hand, or a test database a tool runs, is legitimate. The host already calls these *strays*: 3,
8 and 25 on three of the machines on the day of deciding;
- the runtime's events and verbs are subjects on the bus, and the mesh's own interface is built on
them. The third-party interface run until now was removed by hand.
[ADR 0161](0161-what-deserves-a-seat.md)'s test for a seat is whether the mesh's own code finds it by
name. Here it does: the host would look up the holder on its own machine. A singular role of a module
held once per machine is a `node-*` seat ([ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md)),
in the controller's seed.
The constraint that decides most of this record is a cycle. The bus runs in containers. On the
broker's machine, the broker's own container is created by the host. A holder's code served from a
container cannot create the container that runs it. On a first machine, before the controller exists,
nothing holds anything.
## Considered Options
1. **The host calls the holder's verbs over the bus.** Rejected: with the broker down, no machine can
create any container, including the broker's. The mesh would be unable to restart its own
transport.
2. **The holder picks a dialect that the host speaks itself, as with the uplink.** Rejected: the
host would still drive the runtime, and the module would drive it too for every other caller.
That is two programs speaking to one daemon, and they come to disagree about the same machine
(the installer's preflight already exists to avoid this).
3. **The holder's code runs as a supervised process on the machine and serves the seat's verbs
twice: locally to the host, on the bus to everyone else.** Adopted.
## Decision
**`node-container-runtime` is a seat of the mesh's own, node-scoped,** in the controller's seed under
this record. It delivers no provision; what it carries is its role's protocol: verbs its holder must
serve ([ADR 0132](0132-a-seat-carries-the-tools-its-holder-must-serve.md)) and events its holder emits
([ADR 0129](0129-a-seat-carries-the-protocol-of-its-role.md)). The runtime's module, `docker`, claims
it and is assigned to every machine. A podman module may claim it later; one machine runs one.
**The seat's verbs cover every container on the machine:** list, inspect, logs, stats, start, stop,
restart, create and remove. A mesh-held container is marked by the host's label and says which
assignment holds it. **A container the runtime runs can be root on the machine** — privileged, a host
path mounted, the host's network or process namespace, the runtime's own socket — so a caller other
than the host may not create one that is any of these; only a declaration the mesh composed may ask
for them. And the verbs that change anything are granted by name, never by a wildcard: a grant of
every tool (the console's today) reaches the reading verbs only. Issue 193 is what a verb that trusts
its caller costs. **Creating or removing a mesh-held container is the host's alone.** Any other
caller is refused naming the assignment, because the host would undo it at its next apply. Starting,
stopping or restarting one is allowed, and the answer says the host will restore what its
declaration says. A container the mesh does not hold is the caller's to do anything with.
**The seat's events are the runtime's own** — a container created, started, stopped, died, removed,
its health changed. They are emitted on the seat's subjects, so every holder emits the same events and
no reader depends on which runtime holds the seat. As
[ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)
decides, the subjects are issued by the controller, not composed by the module.
**The holder's code is a supervised process, not a container** ([ADR 0150](0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md)).
A runtime cannot be run by the thing it runs. The process serves the seat's verbs on the bus to the
console, to tools and to the mesh's interface. The same verbs are served on a local socket on the
machine, which only the host may use. **The host creates, inspects and removes its containers
through that socket and nothing else.** If the holder does not answer, the host creates nothing. It
says so in its report, naming the seat. It never falls back to the command line.
**A container needs the seat held on its machine.** An assignment that delivers a container on a
machine whose runtime seat is unheld is refused, naming the seat and its candidates
([ADR 0165](0165-container-runtime-is-what-a-machine-can-run-and-a-running-runtime-is-its-holders-health.md)).
Mounting the runtime's socket into a container is granted by the seat, not by the capability. The
socket's path is the holder's to state, because podman's is not docker's.
**The runtime module owns the runtime's configuration.** Its settings are declared with defaults
([ADR 0164](0164-a-setting-is-declared-with-its-default-its-meaning-and-what-changing-it-costs.md)):
the resolver containers use, the registries it trusts, log rotation, and keeping containers through a
daemon restart. The module is given the resolver's address and the mesh's registry as values; no other
module writes the runtime's file or declares its service.
**The first machine is bootstrapped with the holder, and adopted afterwards.** The foundation bundle
already installs the runtime's package and service. It also carries the holder's process, delivered as
a binary the way the host is ([ADR 0142](0142-the-mesh-delivers-its-own-components-as-binaries.md)).
When the runtime module is assigned, it adopts what the bundle made, as the store and broker modules
adopt theirs ([ADR 0078](0078-the-store-and-broker-are-modules.md)).
## Consequences
- **The migration on the running mesh has a fixed order:**
1. Each machine's hand-written configuration is read, because the module's defaults replace what
differs.
2. In one push per machine: the resolver module and the private network stop writing the
runtime's file ([issue 190](../04-ISSUES/190-the-runtimes-configuration-is-written-by-modules-that-are-not-the-runtime/00-report.md)),
and the runtime module is assigned and adopts the runtime, its file and its service. Split in
two, either the controller refuses two modules declaring one path, or a machine is left with
nothing setting `dns` and `live-restore`.
3. The controller seeds the seat and enforces the container requirement.
4. The host releases the version that uses the holder.
5. The host's command-line path is removed in the release after every machine's holder answers.
Until then, the host reports per machine which path it used.
- **Every container the host makes depends on the holder's process.** A crash-looping holder stops
new containers on its machine. Running containers are unaffected. The host's report names the cause.
- The process form of a module's own code must serve tools on the live mesh before this ships. Only the
showcase declares it, and [issue 117](../04-ISSUES/117-a-modules-own-code-is-a-container-and-a-process/01-diagnosis.md)
found the showcase's tools declared in a form nothing runs. The runtime module is the first whose
tools cannot fall back to a container.
- A user interface subscribing to events directly does not exist. Today a reader of events is a module
that consumes them. The mesh's container view is a module, or waits for that path.
- [ADR 0005](0005-the-node-host.md) ("a container runtime is detected, not chosen") and
[ADR 0006](0006-the-substrate-and-the-control-plane.md)'s matching line describe the mechanism this replaces: on
acceptance, each gets a dated note saying the runtime is now a seat's holder, as the decision
records' rule for a moved mechanism requires. Design 05 and design 26 are amended after acceptance.
- The operator's decision to remove the third-party interface by hand needs no mechanism. No
module-retires-module rule is introduced.
- **What got harder:** the host gains a dependency it did not have, and a first machine's bundle gains
a component. The direct path was simpler and is what makes a runtime a black box to the rest of the
mesh.
## How this is checked
| Rule | Checked by |
|---|---|
| The seat is the mesh's own, node-scoped, with its verbs and events | A catalogue test on the default seats; registration refuses a claimant that does not serve every verb (design 33's existing check) |
| Creating or removing a mesh-held container is the host's alone | A test of the runtime module's verbs: create or remove of a container carrying the host's label, from any caller but the host's socket, is refused naming the assignment; the same verbs on an unlabelled container succeed |
| No caller but the host creates a container that is root on the machine | A test of `create` from the bus: privileged, a host path, the host's namespaces and the runtime's socket are each refused; the same request on the host's socket is accepted. A broker test: a grant of every tool does not reach a changing verb |
| The host uses the holder and never the command line | A host test with a fake holder on the local socket: every container operation goes to it, and with the holder absent the apply creates nothing and reports the seat; after step 5, the host carries no command-line runtime code (checked by build: the package is gone) |
| A container needs the seat held | A resolution test refusing a containerised assignment on a machine with the seat unheld, naming the seat |
| Socket mounts are granted by the seat | A catalogue test: a module mounting the runtime's socket on a machine whose holder states a different path is refused |
| No other module writes the runtime's file | The existing collision check, once the private network's computed resources are inside it ([issue 190](../04-ISSUES/190-the-runtimes-configuration-is-written-by-modules-that-are-not-the-runtime/00-report.md)) |
| Live | `seats` lists `node-container-runtime` held on every machine; the node listing shows each machine's containers, strays included, from the seat's `list` verb; a container started by hand appears as an event on the bus |
## References
- [ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md), [ADR 0129](0129-a-seat-carries-the-protocol-of-its-role.md), [ADR 0132](0132-a-seat-carries-the-tools-its-holder-must-serve.md), [ADR 0159](0159-a-tool-call-names-the-machine-and-a-holder-serves-its-seats-verbs.md), [ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md), [ADR 0161](0161-what-deserves-a-seat.md)
- [ADR 0150](0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md), [ADR 0142](0142-the-mesh-delivers-its-own-components-as-binaries.md), [ADR 0078](0078-the-store-and-broker-are-modules.md), [ADR 0005](0005-the-node-host.md)
- [ADR 0164](0164-a-setting-is-declared-with-its-default-its-meaning-and-what-changing-it-costs.md), [ADR 0165](0165-container-runtime-is-what-a-machine-can-run-and-a-running-runtime-is-its-holders-health.md), [issue 190](../04-ISSUES/190-the-runtimes-configuration-is-written-by-modules-that-are-not-the-runtime/00-report.md)
- [Design 26 — the seats](../03-DESIGN/01-to-be/26-the-seats.md), [design 33 — the tools the mesh answers](../03-DESIGN/01-to-be/33-the-tools-the-mesh-answers.md)
- mesh-host `internal/apply/apply.go` (the runtime lookup and the command line it drives)
@@ -9,8 +9,6 @@ extends: 02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under
# 175. One tool runtime per node serves every module's tools, on the host side
> **The mechanism changed — 2026-10-02, by [ADR 0188](0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md).** Everything decided here stands: one runtime per node, host-side, every module's tools and every held seat's verbs on the memberships' subjects, root the module's concern, any node calls any tool, the console its serving mode. What moved is how the runtime brings a bundle to life. Decision 3 and the consequence *the node tools runtime needs an interpreter on the machine* read as though a bundle were always interpreted code the runtime imports; a tools bundle is now a process in any language that speaks MCP over stdio to the runtime, and importing a TypeScript bundle is the shortcut, not the contract.
## Context
A module's tools are code the module wrote, one function behind each verb, served on the subjects
@@ -1,131 +0,0 @@
---
topic: what runs on it
status: accepted
date: 2026-10-02
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md
---
# 188. A module's own code is bundles in any language, and a tools bundle speaks MCP to the runtime
## Context
[ADR 0175](0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md) put one
tool runtime on every node and said a module brings its tools as a bundle. The runtime that exists
is written in TypeScript and brings a bundle to life by **importing it into its own process**, which
only JavaScript can be. The SDK ([ADR 0039](0039-what-the-sdk-holds-and-refuses.md)) is one
TypeScript package. The builder knows three toolchains — TypeScript, Go, Python — and every one of
the 35 catalogue modules with tools wraps them in a container on the runtime's TypeScript image.
Nothing in the records says a module's code may be written in anything else, and nothing refuses a
module that wraps its own code in an image to get around that.
The operator's direction, stated on 2026-10-02 and repeated: *the SDK is the most important part;
we must not limit developers; tools can be written in any possible language — Rust, C, Go,
JavaScript. A service in Go or Rust as a systemd unit must be possible too. One module can deliver
all kinds of bundles: one for its tools, one for a seat's implementation, one for a daemon. Support
the bare minimum first, as a skeleton; a full implementation comes when the work requires it.*
Measured against that: the `bundle` artifact kind already names a language and the `process`
resource already runs a command from an unpacked bundle as a unit the host writes
([ADR 0150](0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md)), so a Go
daemon as a native service is possible today and one module in the catalogue does it. What is not
possible is a tool in any language but one, and what is not written is that any of this is the
rule.
## Considered Options
1. **One SDK, one language, as now.** Rejected: it limits who can write a module to one
ecosystem, which the operator declines, and it is what made every module's tools a container
on one image.
2. **A full bus client per language.** Each SDK speaks the bus itself; the runtime only
supervises. Rejected: a transport in every SDK is what ADR 0039 refuses, and a bus change
would then rebuild every module in every language — the cascade, multiplied.
3. **A tools bundle is a process the runtime launches and speaks a small local protocol to,
and that protocol is MCP over stdio.** Chosen. The runtime already speaks MCP outward (the
console); speaking it inward to a child process is the same vocabulary. Every language that
has an MCP server library can write a tools bundle today with no mesh SDK at all, and the
mesh's own SDK for a language is a thin convenience over it. The transport stays in the
runtime, so a bus change rebuilds nothing.
4. **A protocol of the mesh's own design.** Rejected: a second way to describe a tool, its
schema and its call, inventing what MCP already settled, for no gain.
## Decision
**1. A module's own code is bundles, in any language the mesh has a toolchain for, and never an
image.** A `bundle` names its language and what it is for. Images are for third-party software a
module installs — a database, a forge — never for code the module wrote. One module may declare
several bundles: its tools, its implementation of a seat's verbs, a daemon, a step. Each is built
alone and delivered alone, as [ADR 0156](0156-an-artifact-is-what-a-build-produces-and-the-store-is-named-for-its-scope.md)
already has it.
**2. A bundle the runtime serves is a process that speaks MCP over stdio.** The node's runtime
launches it as the bundle names it — an interpreter and a file, or a binary — with the runtime's
environment, asks `tools/list`, and answers each call on the bus by `tools/call`. A tool whose name
is `<seat>.<verb>` is the module's implementation of that seat's verb; any other name is the
module's own tool. Everything the runtime does with what it is told — subjects from the membership,
a held seat's verbs, the `tools` answer, a bundle that fails named and the others serving — stays as
[ADR 0175](0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md) and
[ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)
have it. A TypeScript bundle may still be imported into the runtime's own process; that is a
shortcut over the same contract, not a second contract, and a TypeScript bundle written against
the protocol is served the same way as any other.
**3. A bundle that is a service is a `process`**, run by the host as a unit, in whatever language it
is compiled from, exactly as the host's own bundle already is. Nothing new is decided here; it is
said so that it is the rule and not an example.
**4. One thin SDK per language, and the test of ADR 0039 applies to each.** An SDK for a language
holds the MCP-over-stdio loop, the tool-definition type and the few primitives a module's code
needs; it holds no transport, no module's client and nothing volatile. Where a language has a
sound MCP library, the SDK wraps it rather than re-implementing it. The languages are those that
make sense to write a module in; the first set is TypeScript, Go, Python, Rust and C, and the set
grows when a module needs one, not before.
**5. Skeleton first.** Each piece — a toolchain, a launcher, an SDK — exists at the bare minimum
that lets one bundle in that language be built, delivered and answer one tool on the live mesh.
Anything beyond that is added when a module needs it. A skeleton that is not proven by one bundle
answering is not a skeleton; it is a promise.
> **The mechanism changed — 2026-10-03, by ADR 0193.** §2's allowance that a TypeScript bundle may
> be imported into the runtime's own process is withdrawn: every served bundle is launched, and the
> build makes each served entrypoint executable. The rest of §2 stands.
## Consequences
- The runtime gains a launcher beside its loader. The loader, the memberships, the seats and the
failure handling built for ADR 0175 stand; the launcher is the one new step.
- The builder gains a toolchain per language, each at the skeleton: compile, pack, name the
entrypoint. Rust and C are new; a language that compiles to a binary says its operating system
as a Go bundle already does.
- An existing MCP server in any language is already a valid tools bundle. What the mesh adds is
the subjects, the seats and the memberships around it.
- The gate [to-be 38](../03-DESIGN/01-to-be/38-building-the-operators-machine.md) WP2 adds —
refusing a tools container built on the runtime's image — widens: a module whose own code is
an image artifact is refused at registration, naming this record.
- What got harder: a tools bundle is now a process per module on the node rather than code in
one process, so the runtime supervises children and restarts one that dies. The one-process
shape ADR 0175 counted on for the TypeScript shortcut remains available for it.
- ADR 0039's "what the SDK holds" now reads per language; its refusals are unchanged and are the
reason option 2 was rejected.
## How it is checked
| Rule | Checked by |
|---|---|
| A module's own code is never an image | the catalogue's registration check: a manifest with a `bundle` kind of own code *and* an image artifact built from the module's own directory is refused, naming this record |
| A tools bundle in a language other than TypeScript answers on the bus | the runtime's tests: a bundle written against the protocol in a second language, launched, its tool called over a real bus |
| A TypeScript bundle written against the protocol is served like any other | the same tests, with the TypeScript shortcut off |
| Each SDK is thin | each SDK's own README states what it holds under ADR 0039's test, and its size is in the mesh's records |
| Live | a tool in a compiled language answers from the node's runtime on one machine |
## References
- [ADR 0175](0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md),
[ADR 0039](0039-what-the-sdk-holds-and-refuses.md),
[ADR 0150](0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md),
[ADR 0156](0156-an-artifact-is-what-a-build-produces-and-the-store-is-named-for-its-scope.md),
[ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)
- [To-be 38](../03-DESIGN/01-to-be/38-building-the-operators-machine.md) — the work packages this
record widens
- The Model Context Protocol's stdio transport — the local protocol a tools bundle speaks
@@ -0,0 +1,130 @@
---
topic: the mesh
status: accepted
date: 2026-10-02
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0048-a-provider-creates-the-credential-the-mesh-minted.md
---
# 188. A provider declares what it derives for each consumer, and the mesh tells both ends
## Context
An arrangement between a consumer and a provider is delivered entirely by the mesh. Where the
provider is, which port it answers on, what name the consumer must present, where its password
is — each arrives as a fact the consumer reads from its binding, or as `${bound:…}` filled into a
file before the declaration leaves the control plane. The provider invents none of it and hands
none of it back ([ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)).
One kind of value escapes that. Where the **provider names the resource** — a bucket, a database,
a vhost — the name is derived from the consumer, per consumer, and the mesh has no way to carry
it. `serves` is a literal block in the provider's definition: the same values for every consumer.
A provisioner's contract takes a provision and returns nothing. So a value the mesh's own rule
produced reaches neither end as a statement; it is recomputed at one end and transcribed at the
other.
The object store is the instance ([issue 124](../04-ISSUES/124-a-consumer-cannot-be-told-what-its-provider-derived/00-report.md)).
Its provisioner normalises the login the mesh minted into a bucket name and creates, checks and
removes exactly that; the rule lives in twenty lines of the module's own TypeScript. Its three
consumers each write the answer into their own definition by hand. Two transcribed it correctly;
one named a predecessor's bucket, and would have authenticated successfully and been refused on
every object, which reads like a credential fault and is not one.
Even corrected, the transcriptions are wrong in a second way. Each is `mesh-<node>-<slug>`, so
each **names the machine the module happens to run on today** — a definition stating a fact about
one installation, which [ADR 0155](0155-a-definition-names-no-installation-and-how-that-is-checked.md)
forbids and whose check does not catch because the name is not a domain. Move any of the three to
another machine and its configuration points at a bucket its key cannot open.
The shape is not the object store's. A database provisioner that prefixed names, a queue provider
that scoped vhosts, any provider that derives a resource from who is asking: each forces the
consumer to reproduce somebody else's rule and keep it in agreement by hand.
## Decision
**1. A served value may name the consumer the mesh is serving.** A `serves` block, which is
literal today, may interpolate the mesh's own statement of who the consumer is:
- `${consumer:as}` — the identity the mesh minted for this consumer, exactly as the login it is
told to present ([ADR 0049](0049-a-consumers-identity-fits-the-tightest-backend.md));
- `${consumer:as:dns}` — the same identity written as a DNS label.
Nothing else. **The mesh learns no protocol here; it spells its own name in an alphabet it already
knows.** The identity is the mesh's, minted by the mesh, already capped at twenty characters
because of what an S3 access key accepts; `dns` is that same name with its separator written `-`
instead of `_`, which is the whole of the difference between the mesh's identifier alphabet and
the one buckets, vhosts and hostnames use. A provider that needs a prefix or a suffix writes it
around the placeholder, because a served value is a string.
The rejected alternative is **the provider returning values from provisioning** — the natural
channel, since the provider is what derived them. It is rejected for three reasons, in order of
weight. It inverts the delivery the mesh is built on: a grant would carry data the provider wrote
rather than only data the mesh minted, and [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)
removed exactly that second path once already. It makes a consumer's declaration incomplete until
its provider's reconcile loop has run, so a consumer could not be composed before a provider
answered — a bootstrap order the mesh does not have and does not want. And it puts the rule where
nothing can check it: a value that arrives from a running process cannot be refused at resolution,
only discovered wrong later, which is the failure this record exists to end.
**2. The mesh resolves it once, per consumer, and tells both ends from the one resolution.** At the
moment a consumer's declaration is composed, the mesh knows exactly who the consumer is. There, and
only there, the placeholders are filled. The result reaches:
- the **consumer**, as the served facts in its binding file and as `${bound:<provision>:<key>}` in
any file it writes — unchanged mechanisms, carrying one more key;
- the **provider**, as `serves` on that consumer's entry in its contributions file, so the
provisioner is *told* the name rather than recomputing it.
**The provider stops deriving in code and starts declaring.** One statement, filled once, delivered
to both ends: the two cannot disagree, because there is no second computation to disagree with.
**3. A served value stays settled before it is per-consumer.** Settings still compose into `serves`
([ADR 0174](0174-a-node-varies-a-module-through-settings-and-kept-regions-never-an-edit.md)), and the consumer
placeholders are filled after that, so an operator may set a prefix and the mesh still derives the
rest. A `${consumer:…}` naming a fact or an alphabet the mesh does not have is refused when the
definition is parsed, with what it may say.
**4. A consumer may no longer name the resource its provider derives.** With the value delivered,
a literal in a consumer's definition is not merely redundant — it is the one thing that can
disagree with what the provider will actually create. The three object-store consumers lose their
hand-written bucket names in this change.
## Consequences
- One more thing a definition may say, and one less thing a module may be wrong about. The
vocabulary grows by a placeholder; the catalogue loses three literals that named this
installation's control node.
- A provider's naming rule becomes readable in its definition instead of in its source. `minio`'s
`bucketFor` goes; the manifest says `"bucket": "${consumer:as:dns}"` and the provisioner uses
what it is given.
- A provider that already serves consumers keeps serving them: the derived value equals what the
code derived, so no bucket, database or login changes name. This is a change of **who says it**,
not of **what is said**.
- The mesh now holds a rule in another system's alphabet — one rule, `dns`, stated once. A second
alphabet is a decision, not an addition: the cost of each is that the mesh must be right about
somebody else's naming, and that cost is only worth paying where the mesh already mints the name.
## How this is checked
- A served value naming an unknown fact or alphabet is refused at parse, with the list of what it
may say — tested on both halves of the message.
- Resolving a consumer whose provider derives a value puts that value in the consumer's binding
file, in its `${bound:…}` substitutions, and in the provider's contributions entry for that
consumer — one test asserting the three agree, because agreeing is the whole point.
- Two consumers of one provider on one machine get two different derived values, and neither gets
the other's.
- A catalogue-wide test refuses a consumer definition that writes a literal where its provider
derives: the provider's `serves` names the key, so the catalogue can say which definitions
transcribe one.
- `dns` is checked against the identity the mesh actually mints, not against an invented string:
the test derives an identity with `ConsumerIdentity` and asserts the label it becomes.
## References
- [issue 124 — a consumer cannot be told a value its provider derived for it](../04-ISSUES/124-a-consumer-cannot-be-told-what-its-provider-derived/00-report.md)
- [ADR 0048 — a provider creates the credential the mesh minted, and seals nothing](0048-a-provider-creates-the-credential-the-mesh-minted.md)
- [ADR 0049 — a consumer's identity fits the tightest backend](0049-a-consumers-identity-fits-the-tightest-backend.md)
- [ADR 0174 — a node varies a module through settings and kept regions, never through an edit](0174-a-node-varies-a-module-through-settings-and-kept-regions-never-an-edit.md)
- [ADR 0155 — a definition names no installation, and how that is checked](0155-a-definition-names-no-installation-and-how-that-is-checked.md)
- [design 27 — a module requires, the mesh resolves](../03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md)
@@ -1,117 +0,0 @@
---
topic: the mesh
status: accepted
date: 2026-10-02
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md
---
# 190. A seat's work is shared by its holders, and building is the first such role
## Context
Work addressed to a role goes to the seat's `accept` subjects, on a per-seat work queue
([design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md) §1, [ADR 0041](0041-events-are-a-relationship.md)).
The holder's worker on that queue is already a queue group — *"even though the seat guarantees one
holder … the day somebody allows two holders for throughput, every message is processed twice with
nothing reporting it"* — and design 25 already says what a build queue shared by several machines is:
*a seat's `accept` subjects, on a work queue with a queue group of holders*. The mechanism was drawn.
Two things stopped it being used.
First, the build role is a **mesh-scoped** seat, `mesh-build-machine`, so there is one holder in the
whole mesh ([ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md):
*the mesh's single build machine*). Second, the worker is a push consumer with **one delivery in
flight** — set so after 2026-10-01, when a push consumer handing out many at once left twenty-six of
forty-three asks undelivered ([issue 175](../04-ISSUES/175-an-announcement-behind-a-long-build-comes-back/00-report.md)) —
and one in flight on a shared consumer is one build at a time across every holder there could be.
Measured on 2026-10-02: a change to code comments in the tool runtime rebuilt its thirty-five
dependent images, one after another, on one machine, for about half an hour, while three other
machines with a container runtime sat idle; the work the mesh wanted next waited behind it. The
operator's words: *this is our first occurrence of a mesh advantage* — and: *make sure the setup is
done generically, so if another module also requires mesh functionality it can re-use the pattern.*
## Considered Options
1. **A second build machine by configuration** — a concurrency setting on the one holder, or a second
holder admitted by hand. Rejected: a setting on one machine shares nothing, and a second holder
of a mesh-scoped seat contradicts what a mesh seat means.
2. **A build-specific dispatcher** — the controller choosing a machine per build and asking it by
name. Rejected: it reinvents the queue the bus already is, it makes the controller a scheduler,
and it is specific to building; the next role needing the same would build its own.
3. **A seat's work is shared by its holders, and the build role becomes node-scoped.** Chosen. It is
what the bus was drawn to do, it is one rule for every role rather than one for building, and
"the machines that are online and hold the seat" is exactly the set a queue group's members is.
## Decision
**1. Work asked of a seat is taken by whichever of its holders is idle.** Every holder of a seat with
`accepts` reads the seat's one work queue; a node-scoped seat held on several machines has several
holders, and an ask goes to one of them. The asker addresses the role — `mesh.seat.<seat>.accept.<verb>`
— and never a machine. The outcome, the role's own event, says which machine did the work (`on`), as a
build's already does ([ADR 0157](0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md)).
**2. A holder takes one ask at a time, when it is idle, by pulling.** The worker is a pull consumer:
a holder fetches one ask, works it, acknowledges, fetches the next. The server never hands an ask
to a busy holder, so a slow machine never holds work an idle one could take — the fault issue 175
found in push delivery is removed by the shape rather than by a limit, and the one-in-flight limit
that made the shared queue serial goes with it. A holder that dies mid-work leaves its ask to be
redelivered to another, as today.
**3. Work that must run on one particular machine is not a work queue.** That is a node seat's verb
asked of that machine ([design 33](../03-DESIGN/01-to-be/33-the-tools-the-mesh-answers.md) §4), and
nothing here changes it. A role's work queue is for work whose result is the same whichever holder
does it: a build is, because what comes out is published by digest to the mesh's store.
**4. This is one pattern, not one role's.** Any module that declares a node-scoped seat with
`accepts` gets decisions 1 and 2 with no further mechanism: the controller derives the queue and the
worker, the holders pull, the module's manifest says what every holding machine must have. The
build agent is the first; a module needing work done *somewhere on the mesh* — a scan, a
conversion, a fetch — declares a seat of its own the same way
([ADR 0126](0126-a-module-declares-its-own-seats.md)).
**5. Building is the first such role.** The build role is `node-build-agent`, scope node, with the
same `build` ask and the same `started`, `built` and `log.<id>` events as before. Its holder is the
`build-agent` module: the builder as it is — a container runtime, the artifact store and the package
registry resolved as provisions, a workspace, the bus credential — assignable to every machine that
has a container runtime. `mesh-build-machine` and the `builder` module are retired when the new
holder is assigned where the old one was. The tiered plan ([ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md))
is unchanged: a tier's asks go out together and are now worked together.
## Consequences
- A tier of thirty-five images is built by as many machines as hold the seat and are online. A
machine that is off builds nothing and blocks nothing.
- Every holding machine fetches base images from the store and pushes what it builds; the store is
reached as a provision, so this is what the provision was for. A machine with a slow link builds
slowly, and takes fewer asks for it, which is the point of pulling.
- A build's outcome carries which machine built it, so a build that fails on one machine and not
another is a fact the record shows, not a mystery.
- What got harder: a build's cache is per machine, so a cold machine pays the first pull of every
base it has never seen; the artifact store is now asked by several machines at once, and the
package registry likewise. Both are provisions and both are made for that.
- ADR 0121's *"the mesh's single build machine"* and ADR 0162's *"built by whichever build machine
is running — the only one there could be"* were true and are no longer; both records carry a note.
## How it is checked
| Rule | Checked by |
|---|---|
| Two holders of one seat each take one of two asks, and a third ask waits for the first to be idle | the controller's test over the work queue against a real bus: two machines bound to one worker, three asks |
| An ask is never delivered to a busy holder | the same test: the busy holder's ask count stays at one until it acknowledges |
| A holder that dies mid-work leaves its ask for another | the same test, one holder closed mid-ask |
| The asker names no machine | the controller's seat table: `node-build-agent` accepts `build` and the asking side publishes to the seat's accept subject, as the existing tests of `build` already require |
| Live | `builds` shows a tier's builds `on` more than one machine within one plan; `seats` shows `node-build-agent` held on every machine with a container runtime — *held on all four machines and a build taken by a workstation's agent, 2026-10-03* |
## References
- [Design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md) §1 and §5 — the work queue and the queue
group of holders this uses as drawn
- [Design 18](../03-DESIGN/01-to-be/18-building-a-module.md) — building a module, amended for
where a build runs
- [ADR 0041](0041-events-are-a-relationship.md), [ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md),
[ADR 0126](0126-a-module-declares-its-own-seats.md), [ADR 0157](0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md),
[ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md)
- [Issue 175](../04-ISSUES/175-an-announcement-behind-a-long-build-comes-back/00-report.md) — why the
worker had one in flight, and why pulling removes the cause rather than the symptom
@@ -1,120 +0,0 @@
---
topic: the tiers
status: accepted
date: 2026-10-03
deciders: jochen
reconstructed: false
supersedes-in-part:
- 0066-public-routing-is-name-agnostic.md
- 0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md
---
# 191. The mesh's resolver holds only the mesh's own names; a public name resolves publicly
> **Progressive insight — 2026-10-03.** The first implementation told the mesh's names from public
> ones by their spelling — a name ending in the mesh suffix — and this record said so: the Decision
> read *"only names under its own suffix"*, and the roster check *"every name the roster carries ends
> in the mesh suffix"*. The mesh needs no such test, nor any per-route name: domains are a node's. A
> node has **one internal domain**, `<node>.internal`, and every route on it is a name under that domain
> (ADR 0151), answered by one wildcard per node; a node has **one or more public domains**, which public
> DNS answers. So the mesh's resolver holds the nodes' internal domains and nothing else, and the roster
> carries the machines and no routed name. Both sentences now say that; what was decided — a public
> name is never given a private answer — is unchanged.
## Context
**[ADR 0066](0066-public-routing-is-name-agnostic.md) published every routed name into internal
resolution, mesh-wide, at the address of the node that serves it.** The reason was an internal
certificate authority in the lab: it validates by connecting to the name it certifies, and a routed
public name that nothing inside the mesh resolved could not be certified.
[ADR 0151](0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md) kept it:
*the roster publishes it as itself, once, at the serving node's address.*
**So every machine's resolver answered public names with private-network addresses.** On a
production mesh on 2026-10-03, each machine's hosts region carried 47 lines of the form
`<private address> <label>.<public domain>` — every public name of the control-node at its tunnel
address, every public name of the home server at its own. For the machines themselves this is merely
a detour: their traffic to a public name goes through the tunnel instead of the internet.
**For anything that is not a member it is an outage.** The home server's resolver also answers its
LAN — a listen address added as a setting on 2026-10-02. A phone on that LAN asked for the mail
server's public name, was given the control-node's tunnel address, and could not connect:
*couldn't connect to host, port: 10.10.0.1:143*. Every public name of the mesh failed the same way for
every non-member on that LAN — a phone, a television, a guest — while every check the mesh runs
reported success, because every check runs from a member.
**And the reason for publishing them is gone.** ADR 0151 gave every route an internal name,
`<label>.<serving node>.internal`, under the node's own name. It resolves inside the mesh without any
entry of its own, the proxy serves it, and the internal authority certifies it — the proxy has two
authorities since 2026-09-25: a public one for public names, the internal one for internal names.
Measured the same day: `drive.<control-node>.internal` resolves to the control-node's tunnel address
and answers 200 with a certificate that verifies against the internal root. Nothing the mesh runs
needs a public name to resolve to a private address. The one consumer that did — an internal
authority validating a public name — is the case the second authority removed.
The predecessor's resolver held exactly this and no more: an address per machine under `.internal`,
and everything else forwarded to public resolvers.
## Considered Options
**1. Keep publishing public names; stop the resolver answering the LAN.** Fixes the phone and
nothing else. The mesh would still hold a second, private answer for names the public DNS already
answers — two answers for one name, which disagree by design and are correct in different places.
And it forbids a reasonable setup: a home server's resolver serving its own LAN.
**2. Answer per source: private addresses to members, public ones to everyone else.** Split-horizon
by client. It is what a resolver serving two audiences would need *if* the private answer were worth
giving. It is not — option 3 shows nothing needs it — and it makes a name's address depend on who
asks, which is the hardest kind of fault to see from a member.
**3. The mesh's resolver holds only the mesh's own domain.** Names under the mesh suffix — machines,
and routes' internal names under them — resolve to private addresses. Every other name, including
every public name the mesh serves, is forwarded and resolves publicly. Chosen.
## Decision
**The mesh's resolver holds each node's internal domain and nothing else** — `<node>.internal` and
everything under it, at that node's private address. A machine's name,
and through it every `<label>.<node>.internal`, resolve to that machine's private address. **A public
name is never given a private answer by the mesh**: it resolves through public DNS to the public
address, from members and non-members alike.
This replaces ADR 0066's clause *"when the proxy is granted a name, the mesh publishes that name →
the node that serves it into internal resolution, mesh-wide"*, and ADR 0151's *"the roster publishes
it as itself, once, at the serving node's address."* Everything else in both stands: the label, the
node's public domain, the composition, and the internal name under the serving node.
**Inside the mesh, a route is reached by its internal name.** A container or a validator that must
reach a routed service inside the mesh uses `<label>.<node>.internal`; the internal authority
certifies that name, and a public authority certifies the public one. A mesh with no public
reachability — the lab — certifies its internal names and has no public names to resolve.
## Consequences
- **A resolver serving a LAN is safe.** What it adds to public resolution is the mesh's own domain,
which no public resolver answers.
- **A member reaches a public name over the internet, as anyone does.** A route the proxy restricts
to the private network is reached by its internal name, never by its public one — a public name
is, by this decision, public.
- **The internal authority certifies internal names only.** It was the only consumer of a public
name's private answer; the proxy's second authority already took that role away from it.
- **Public names leave every machine's hosts region** on the first push after the change.
Containers do not move with it: the roster is not part of a container's identity
([ADR 0148](0148-the-meshs-names-are-resolved-not-copied-into-containers.md)).
**How each is checked:**
- **The roster:** the controller's tests assert that the roster names the machines and nothing
else — a routed name in it, public or internal, fails the build.
- **On a machine:** asking the machine's resolver for a public name the mesh serves returns the
public address, and asking it for that route's internal name returns the private one. Asked from a
non-member on a LAN the resolver answers, the first must hold as well.
## References
- [ADR 0066 — public routing is name-agnostic](0066-public-routing-is-name-agnostic.md), whose
propagation clause this replaces.
- [ADR 0151 — a route's internal name is composed under the node that serves it](0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md),
which made the private answer unnecessary.
- [Connectivity design §2 and §5](../03-DESIGN/01-to-be/08-connectivity.md), amended alongside this
record.
@@ -1,126 +0,0 @@
---
topic: what runs on it
status: accepted
date: 2026-10-03
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md
---
# 192. A tools bundle declares what it is given, and the runtime hands it to that bundle alone
## Context
[ADR 0175](0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md) put one
runtime on every node serving every module's tools from a bundle, and
[ADR 0188](0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md)
said a module's own code is bundles and never an image. The two holders that moved first
([to-be 38](../03-DESIGN/01-to-be/38-building-the-operators-machine.md) WP4) needed nothing a
bundle does not have: a fixed path, and root. Research
[020](../01-RESEARCH/020-what-a-bundled-tool-is-given/00-overview.md) measured the rest before
they move: thirty-three modules still run their tools as a container on the runtime's image, and
thirty-one of them are handed, through the container's environment and mounts, things a bundle
has no way to receive — the module's configuration file, its own secret as a file, the service's
address with the port the mesh chose, a provision's address, a directory of grants. Every one of
those is a file the mesh already places on the machine or a value the controller already composes
for the container, per module per machine, from references the manifest writes: a placed
directory, a chosen port. And the SDK's tool contributor is a function of an environment that the
runtime calls without one, so every bundle reads the process's four words.
Without a rule, each of the thirty-one would answer the question its own way, and the runtime's
process would be the one place where every module's paths meet.
> **Progressive insight — 2026-10-03.** The context above calls the thirty-one remaining containers
> tool containers handed what a bundle cannot receive. Measured the same day while building this
> record: nine of them run only tools; three run a main of their own; twenty import, beside their
> tools, the module's own long-running code — event handlers that subscribe on the bus and
> provisioners that act on grants — under the module's own bus identity, and some reach their
> service by a container network name or need a package the image installed. That code is not a
> tool and is not this record's to move: under
> [ADR 0188](0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md)
> §3 it is a `process` bundle, and how it is given its credential, its words and its reach is the
> open question of [design 38](../03-DESIGN/01-to-be/38-building-the-operators-machine.md) WP4c.
> Decision 4 applies to these containers' tools; the containers themselves go when their other
> code has moved. The decision and its options stand.
## Considered Options
1. **The bundle declares its environment on its artifact, and the mesh composes it as a
container's.** Chosen. The tools artifact gains `env`: names to values, the values written with
the references the composer already resolves for a container — `${dir:…}`, `${port:…}` — and
what was a mount target becomes the host path itself. The controller composes one environment
per bundle per machine into the runtime's declaration. The runtime hands it to that bundle's
contributor, or to the child it launches, and to nothing else. The tool code reads the names it
read before.
2. **The runtime derives it from the module's placed manifest** — a conventional word per
directory and port, no new field. Rejected: a convention the thirty-one tools must be rewritten
to, the runtime learning the composer's job, and a module that names its file one way and a
module that names it another needing different words regardless.
3. **The tool asks the controller over the bus.** Rejected: a tool that cannot start until the
bus answers fails in the one case tools exist for, and a secret crossing the bus to reach a file
already on the machine is a disclosure for nothing.
4. **Leave each module to its own device.** Rejected by the measurement: thirty-one modules, one
question.
## Decision
**1. A tools bundle says what it is given, on its artifact.** `build.artifacts[].env` names the
words the bundle reads and their values. A value is a path or a constant, composed with the
references a container's environment may use; **never a secret's content.** A secret reaches a
tool the way it reaches a container: as a file the mesh places, whose path the environment names.
A bundle that declares no `env` is given nothing beyond the runtime's own words, which is what the
two holders that moved have.
**2. The mesh composes it, per bundle per machine, as it composes a container's.** The same
references, resolved the same way, to the host's own paths. The composed environment travels in
the node's declaration beside the bundle's archive; a change to it is a change to the bundle for
the purpose of `restart-on`.
**3. The runtime hands each bundle its own environment, and nothing of another's.** A bundle
imported into the runtime's process receives it as the argument its contributor is written to
take; a bundle launched as a child ([ADR 0188](0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md) §2)
receives it as the child's environment, over the runtime's own words. The runtime's process
environment is not where a module's words go, and a tool that reads the process's environment
rather than the one it was handed finds the runtime's four words and no module's.
**4. The remaining tool containers move in one change** after this is built, each proven by its
tools answering from the runtime, and the registration gate of
[to-be 38](../03-DESIGN/01-to-be/38-building-the-operators-machine.md) WP2 then refuses the
container shape for every module, as ADR 0188 already provides.
> **The mechanism changed — 2026-10-03, by ADR 0193.** Decision 3's imported path — the environment
> handed to an imported bundle's contributor — has nothing left to do: every served bundle is
> launched, and a launched bundle's environment is its own. The decision stands.
## Consequences
- The manifest gains one field on one artifact kind; the composer gains one more thing to resolve
with references it has; the runtime gains the hand-off and the separation. The thirty-one modules'
tool code does not change, and their conversion is the move of a container's `env` with its
mounts folded into host paths.
- A tool's inputs become legible in the manifest where its container hid them in mounts: what a
module's tools read is declared beside what the module writes.
- What got harder: the runtime must keep thirty-one environments apart in one process, and a
bundle's author must not reach for the process's environment. The separation is a rule the
runtime's test holds, not a property of the language.
- `MESH_BROKER_FILE` is not a bundle's to declare: the runtime speaks with the node's credential
([ADR 0175](0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md)), and
a module's own bus credential went with its container.
## How it is checked
| Rule | Checked by |
|---|---|
| A value in a bundle's `env` is a path or a constant, never a secret's content | the catalogue's manifest check refuses a `${secret:…}` reference in a bundle's `env`, naming this record |
| The composer resolves a bundle's `env` as a container's | the controller's composition test: one module, one bundle with `${dir:…}` and `${port:…}` in its `env`, the declaration carrying the host paths and the chosen port |
| Each bundle sees its own environment and no other's | the runtime's test: two bundles with different `env`, loaded in one runtime, each answering with its own words and none of the other's; the same for a launched bundle |
| A change to a bundle's environment restarts the runtime | the composition test above, with `restart-on` naming the bundle |
| Live | a module whose tools read a configuration file and a token file answers from the runtime on one machine with no container |
## References
- [ADR 0175](0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md),
[ADR 0188](0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md)
- Research [020](../01-RESEARCH/020-what-a-bundled-tool-is-given/00-overview.md) — the measurement
and the options
- [to-be 38](../03-DESIGN/01-to-be/38-building-the-operators-machine.md) — where the work is listed
@@ -1,88 +0,0 @@
---
topic: what runs on it
status: accepted
date: 2026-10-03
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md
---
# 193. Every bundle the runtime serves is launched, and the runtime knows no language
## Context
[ADR 0188](0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md)
§2 made a served bundle a process that speaks MCP over stdio, and kept one exception: a TypeScript
bundle may be imported into the runtime's own process, "a shortcut over the same contract". Every
module's tools today take the shortcut, and it is where the day's defects came from:
- [Issue 209](../04-ISSUES/209-a-bundles-own-sdk-copy-registers-into-a-registry-the-runtime-never-reads/00-report.md):
an imported bundle's own copy of the SDK registered into a registry the runtime never read; fixed
by a resolve hook that redirects every bundle's SDK import to the runtime's copy.
- [ADR 0192](0192-a-tools-bundle-declares-what-it-is-given-and-the-runtime-hands-it-to-that-bundle-alone.md):
thirty modules' environments in one process had to be kept apart by an SDK change and runtime
bookkeeping, where a process of its own has an environment of its own by construction.
- One faulty module can block or crash every other module's tools on its node.
And the shortcut ties the runtime to Node.js: only a JavaScript runtime can import JavaScript. The
operator's direction on 2026-10-03: *a module's tools are written in any language and the builder
builds them; the runtime runs them all and announces them; it should be fully language agnostic,
and node-tools can be rewritten in Go.*
## Considered Options
1. **Keep the shortcut.** Rejected: it is the cause of the three defects above, and it pins the
runtime's language.
2. **Launch every served bundle; the runtime knows how to start each language** (`node` for a
`.js`, exec for a binary). Rejected: the runtime would hold a table of interpreters, and a
runtime in Go would carry Node.js's knowledge for nothing.
3. **Launch every served bundle, and the build makes each served entrypoint executable.** Chosen.
A compiled language's binary is executable already; for an interpreted one the toolchain writes
a launcher beside the entrypoint — for TypeScript, a file that imports the entrypoint and serves
what it registered over stdio, using the bundle's own SDK. The runtime execs what it is given.
## Decision
**1. Every bundle the node's runtime serves is a child process speaking MCP over stdio.** The
in-process shortcut of ADR 0188 §2 is withdrawn. Everything else ADR 0188 §2 says — `tools/list`,
`tools/call`, `<seat>.<verb>` naming a seat's verb, the runtime serving each on the bus — stands.
**2. The runtime knows no language.** It is given an executable per served entrypoint and starts
it, with that bundle's environment ([ADR 0192](0192-a-tools-bundle-declares-what-it-is-given-and-the-runtime-hands-it-to-that-bundle-alone.md))
over its own words, and the module it serves it as. What makes an entrypoint executable is the
toolchain's business: a binary is one; an interpreted language's toolchain writes a launcher.
**3. A launched bundle is told the module it serves as**, so that what it lists unprefixed is that
module's own tools and a seat's verbs are always `<seat>.<verb>`, whichever it registered first.
**4. The runtime may be written in any language.** Nothing it does needs it to share a language
with a bundle; the mesh's runtime moves to Go, against this contract, once the contract is proven
in the runtime that exists.
## Consequences
- A process per served module per node. On the busiest machine that is a score of small children
where there was one process; a Go bundle costs a fraction of a Node.js one.
- The SDK resolve hook (issue 209) and the per-registration environment hand-off (ADR 0192 §3,
imported bundles) have nothing left to do and go; a launched bundle's environment is its own.
- A module's TypeScript tool code does not change: it registers as before, and the generated
launcher serves what it registered.
- A bundle that crashes or hangs takes only its own tools down, and is started again on its next
call, as ADR 0188 already provides for a launched bundle.
## How it is checked
| Rule | Checked by |
|---|---|
| Every served bundle is launched | the runtime's tests: a TypeScript bundle and a bundle in a second language, both launched, both answering over a real bus; a non-executable entrypoint is refused by name |
| The runtime knows no language | the runtime holds no interpreter: it execs the path it is given (code review; the Go runtime has no Node.js dependency at all) |
| A TypeScript served entrypoint is executable | the builder's test: a TypeScript bundle's served entrypoint has a launcher beside it, mode 0755 |
| A seat's verbs are named as the seat's whichever registers first | the SDK's test: a bundle registering its seat first and its own tools second lists `<seat>.<verb>` and its own tools unprefixed |
| Live | the packet filter's and intrusion prevention's seat verbs and the moved tools answer from launched bundles on every machine |
## References
- [ADR 0175](0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md),
[ADR 0188](0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md),
[ADR 0192](0192-a-tools-bundle-declares-what-it-is-given-and-the-runtime-hands-it-to-that-bundle-alone.md)
- [to-be 38](../03-DESIGN/01-to-be/38-building-the-operators-machine.md) WP4d
@@ -1,157 +0,0 @@
---
topic: the tiers
status: accepted
date: 2026-10-03
deciders: jochen
reconstructed: false
supersedes-in-part:
- 0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md
extends: 0191-the-meshs-resolver-holds-only-the-meshs-own-names.md
---
# 194. The mesh has one resolver, and every node asks it for the mesh's names
> **Narrowed, not replaced — 2026-10-03.** How a node asks is decided again by
> [ADR 0196](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md): every
> node and container asks `mesh-resolver` first and a public resolver only when it is silent. There is
> no `systemd-resolved` module and no runtime `dns` naming `mesh-resolver`, and step 2 of the migration
> reads as 0196 states it. Option 2 below was rejected for a laptop with its tunnel down resolving
> nothing; a public resolver listed second answers exactly then. The one resolver, its placement and
> the retirement of every per-node copy stand.
## Context
**Every node runs its own resolver and holds its own copy of the mesh's names.** On the production
mesh on 2026-10-03, each of the four nodes held `node-dns-resolver` with dnsmasq, fed on every push
with a zones file (one wildcard per node) and a region of `/etc/hosts` (the machines), and pointed
its own `/etc/resolv.conf` at itself. The controller computes the names once; four daemons then hold
four copies, each read in its own way.
**Every resolution fault found that day was a copy disagreeing with the truth, not the truth being
wrong:**
- **A copy read once.** dnsmasq reads `/etc/hosts` at start. After the controller stopped publishing
public names ([ADR 0191](0191-the-meshs-resolver-holds-only-the-meshs-own-names.md)), every node's
hosts file was right and every resolver still answered the mail server's public name with a
tunnel address, until each was restarted.
- **A copy beside other copies.** On the workstation, a name resolved to two addresses in rotation:
the mesh's region gave the tunnel address, and two lines the operator had written before the mesh
existed — one in `/etc/hosts`, one in a file the resolver also reads — gave the LAN address. A
comment beside one of them said to delete it once the mesh took over; nothing made that happen.
- **A copy that became somebody else's resolver.** The home server's resolver also answers its LAN
(a listen address added 2026-10-02), and the LAN's router hands that address out as the only DNS
server. Every phone and television on the LAN resolved through a mesh node's private copy, which is
how ADR 0191's outage reached them.
**And the overlay already has one centre.** Every node has exactly one tunnel peer — the anchor —
and routes the whole private range through it. Two nodes on the same LAN reach each other through
the anchor. So a name under `.internal` is only ever useful while the anchor is reachable: a resolver
anywhere else adds a copy without adding an answer anybody can use.
**What a node asks is already a separate role.** The connectivity design split *serving* (answers
the names) from *asking* (decides what the machine asks), because systemd-resolved cannot answer a
wildcard and can only route the mesh's suffix to something that can
([connectivity §2](../03-DESIGN/01-to-be/08-connectivity.md)). ADR 0121 kept them as two seats,
`node-dns-resolver` and `node-resolver-config`, both at node scope. No node runs systemd-resolved
today; each writes `/etc/resolv.conf` as a plain file pointing at its own dnsmasq.
## Considered Options
**1. Keep a resolver on every node, and make the copies more careful.** Restart on every file it
reads, own every file it reads, refuse to listen on a LAN. Each is a fix for one way a copy goes
stale, and the next way is not on the list yet. It keeps four answers to one question.
**2. One resolver for the mesh, and every node sends it every query.** The simplest asking side —
`resolv.conf` names the mesh's resolver and nothing else. Rejected: public resolution then depends on
the tunnel. A laptop whose tunnel is down could resolve nothing at all, and a public name would take
a detour through the anchor for no reason ADR 0191 left standing.
**3. One resolver for the mesh's names; each node asks it for those only.** The mesh's resolver holds
every node's internal domain. Each node's asking role routes the mesh's suffix to it and every other
name to public resolvers. Chosen.
## Decision
**The mesh has one resolver.** It is a module holding a new mesh-scoped seat, **`mesh-resolver`**
(capacity one). It holds each node's internal domain — `<node>.internal` and everything under it, at
that node's private address ([ADR 0191](0191-the-meshs-resolver-holds-only-the-meshs-own-names.md))
— and listens on the private network only. It is placed on the node every tunnel converges on, so
that it shares the overlay's single point rather than adding one. Which daemon fills the seat stays
the module's business, as the connectivity design says.
**Every node asks it for the mesh's names and nothing else.** `node-resolver-config` routes the
mesh's suffix to `mesh-resolver` and leaves every other name with public resolvers.
**Why a stub on every node.** `resolv.conf` cannot route by domain: the C library asks the servers it
lists in order, for every name, and moves to the next only when one does not answer — an NXDOMAIN
from the first is final. Listing `mesh-resolver` first sends every public name through the tunnel
(option 2); listing a public resolver first means `.internal` is never asked of the mesh. Something on
the node has to look at the name before choosing a server, and that is a stub resolver. Keeping
dnsmasq for it would keep a daemon that reads hosts files and can be told to answer a LAN — the two
ways copies went wrong. systemd-resolved holds no names of its own, routes by domain natively (a
routing domain `~<suffix>` on the server that answers it), and is part of systemd, already installed
on every node and enabled on none.
**So the asking side is a `systemd-resolved` module**, claiming `node-resolver-config` — the same claim
as the `resolv-conf` module it replaces, so the mesh refuses both on one node. It enables the service,
writes its configuration (the mesh resolver for the suffix, public resolvers for everything else), and
writes `/etc/resolv.conf` as a file naming the stub — a file the module owns, not a link to one.
**A container asks the mesh's resolver directly.** The container runtime cannot use a loopback stub
and drops its routing domains, so the runtime's `dns` names `mesh-resolver`, which forwards public
names for the containers that ask it. This is the one place a public name passes through the mesh,
and it is stated rather than hidden.
**`node-dns-resolver` is retired**, and with it every per-node copy: the zones file, the mesh's region
of `/etc/hosts` (the floor connectivity §2 already planned to remove), and the daemon on every node
but the one holding `mesh-resolver`. This narrows ADR 0121's *"the-dns-port → node-dns-resolver"*:
the serving role keeps its distinction from the asking role and moves to mesh scope, as ADR 0121 did
for the private network.
**A LAN's resolver is not the mesh's.** No device that is not a member can reach a private address,
so no member's resolver answers a LAN on the mesh's behalf. A router that hands out a node's address
as a LAN's DNS server is pointed elsewhere before that node stops answering.
**The order is fixed, because every step before the last leaves a working resolver:**
1. `mesh-resolver` is assigned and answers on the private network.
2. Each node's `node-resolver-config` moves from `resolv-conf` to `systemd-resolved`, and the container
runtime's `dns` to `mesh-resolver`.
3. A LAN whose router points at a node's resolver is pointed at its router or a public resolver.
4. `node-dns-resolver` is unassigned from every node, and the hosts region is withdrawn.
## Consequences
- **One answer per name.** A name is wrong in one place or right everywhere; no node can hold a copy
that disagrees, and no operator file on a node is read by the mesh's resolver.
- **The anchor down means no `.internal` names** — which it already meant for `.internal` traffic,
since every tunnel goes through it. Public resolution on every node is unaffected.
- **A container's public resolution depends on the mesh's resolver.** Accepted, and named in the
decision; a container that must resolve public names with the tunnel down is the case it costs.
- **Every node runs systemd-resolved**, through the `systemd-resolved` module. It is installed
everywhere already and enabled nowhere; the mesh still ships no resolver of its own.
- **The runtime's `dns` changes once per node**, which the runtime reads only at start. With
`live-restore` already on, that restart keeps every container running.
- **A LAN loses a resolver it had borrowed.** The router change is an explicit step, done through
the module that manages the router, before the node's resolver goes.
**How each is checked:**
- **One holder:** the seat has capacity one, so a second assignment is refused by the controller.
- **Asking:** on each node, `resolvectl` shows the tunnel's link with `mesh-resolver` and the suffix as
its routing domain; a name under `.internal` is answered by it, and a public name is answered
without it (its query log shows no public name from a node).
- **No copies:** no node but the holder answers DNS on a private or LAN address — every other node's
port 53 is systemd-resolved's loopback stub and nothing else — and no node's `/etc/hosts` carries a
mesh region.
- **A LAN:** the router's DHCP DNS option names no node's address.
## References
- [ADR 0191](0191-the-meshs-resolver-holds-only-the-meshs-own-names.md) — what the mesh's resolver
holds; this record decides where it runs and how nodes reach it.
- [ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md) — the two
resolver seats, and the private network's move to mesh scope this mirrors.
- [Connectivity §2](../03-DESIGN/01-to-be/08-connectivity.md) — serving and asking as two roles;
amended alongside this record.
- [The seats](../03-DESIGN/01-to-be/26-the-seats.md) — the seat table, amended alongside.
@@ -1,102 +0,0 @@
---
topic: what runs on it
status: accepted
date: 2026-10-03
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0152-the-operators-surface-is-a-module-the-console.md
---
# 195. The mesh's tools are found by address, not announced whole
## Context
The console answers MCP on a machine's loopback ([ADR 0152](0152-the-operators-surface-is-a-module-the-console.md),
[to-be 34](../03-DESIGN/01-to-be/34-the-console.md)) and announces, at a session's start, every tool
the mesh can say it has: 228 on 2026-10-03, 110 KB, taken once. Three things are wrong with that,
measured in research [021](../01-RESEARCH/021-finding-a-tool-in-the-mesh/00-overview.md):
- **Size.** Only one client's habit of deferring long lists keeps them out of the model's context.
- **Ambiguity.** A module on two machines is listed once, `node` optional, *whichever answers* — for
postgres and mssql, whose instances hold different data, a call that names no machine asks an
arbitrary one.
- **Staleness.** A tool that arrives after the session started is not listed until it reconnects.
The mesh already has the structure a caller needs: seats held once for the mesh, seats held once per
machine ([ADR 0159](0159-a-tool-call-names-the-machine-and-a-holder-serves-its-seats-verbs.md)), and
modules assigned to machines, each assignment issued its own subjects
([ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)).
The operator's direction: *tools are discoverable, in layers, and asking novox's postgres is not asking
ace's.*
## Considered Options
1. Keep the flat list and rely on the client. Rejected: the ambiguity and the staleness stay, and it
is one client's behaviour.
2. One tool per assignment, the machine in the name. Rejected: the list multiplies, and an address in
a tool's name meets the API's limit — letters, digits, `_` and `-`, at most 64 characters.
3. **A fixed handful of tools that walk the mesh's structure, the address an argument.** Chosen.
4. MCP resources or prompts. Rejected: unevenly supported, and an agent acts through tools.
## Decision
**1. Everything the mesh answers has one address, by the layer it lives in:**
| layer | address | answered by |
|---|---|---|
| a seat held once for the mesh | `<seat>.<verb>` | that seat's holder |
| a seat held once per machine | `<node>/<seat>.<verb>` | that machine's holder |
| a module assigned to a machine | `<node>/<module>.<tool>` | that assignment |
| a module whose instances are interchangeable (ADR 0160) | `<module>.<tool>` as well | any of them |
**A call to a module that is not interchangeable names its machine, or is refused** naming the machines
it runs on. "Whichever answers" is no longer an answer for state a machine holds.
**2. The console announces a fixed set of tools, not the catalogue:**
- **`mesh_overview`** — the mesh's seats with their verbs, and its machines;
- **`mesh_machine`** — one machine: the node seats it holds and the modules assigned to it, each with
its tools by name;
- **`mesh_search`** — words in, matching addresses out with one line each, across every layer;
- **`mesh_describe`** — one address in, its description and argument schema out;
- **`mesh_call`** — an address and its arguments in, the answer out, with the machine that gave it.
Each is answered from the mesh when it is asked, so a tool that arrived a minute ago is found without
the client reconnecting. The names are the API's kind of name; addresses never have to be.
**3. The flat catalogue stays reachable, not announced:** the `mesh` client and a console setting can
still list it whole, for a person reading it or a client that wants it. An agent pointed at the console
sees the five.
> **The mechanism changed — 2026-10-03, by ADR 0197.** Where the console learns what exists: not
> from the catalogue's roster and the controller's printed lists, but from every runtime announcing
> itself on the bus in the NATS services protocol, checked against the controller's records read as
> JSON. The addresses and the five tools stand.
## Consequences
- An agent spends a call or two finding a tool it does not know, and none on one it does; the context
no longer carries 110 KB it mostly never uses.
- The ambiguity is closed by the address, not by a description asking the agent to remember `node`.
- What got harder: an agent that once saw a tool's schema up front now asks for it. `mesh_describe` and
`mesh_search` answering with the schema of a close match keep that to one call.
- The discovery verbs are the console's; the mesh's own records — seats, machines, assignments — are
the controller's, and the console asks it rather than keeping a copy.
## How it is checked
| Rule | Checked by |
|---|---|
| The console announces five tools | the console's test: `tools/list` answers exactly the five |
| An address resolves to one subject per layer | the console's tests: a mesh seat, a node seat, an assignment, an interchangeable module, each called by address over a real bus |
| A non-interchangeable module without a machine is refused, naming its machines | the same tests |
| A tool that arrives after the session started is found | a test registering a module after the console's first answer and finding it by `mesh_search` |
| Live | from a fresh session, *which databases does novox's postgres hold* is answered by novox's postgres, found through the five |
## References
- [ADR 0152](0152-the-operators-surface-is-a-module-the-console.md),
[ADR 0159](0159-a-tool-call-names-the-machine-and-a-holder-serves-its-seats-verbs.md),
[ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)
- Research [021](../01-RESEARCH/021-finding-a-tool-in-the-mesh/00-overview.md)
- [to-be 34](../03-DESIGN/01-to-be/34-the-console.md)
@@ -1,98 +0,0 @@
---
topic: the tiers
status: accepted
date: 2026-10-03
deciders: jochen
reconstructed: false
supersedes-in-part:
- 0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
---
# 196. A node asks the mesh's resolver first, and a public one only when it is silent
## Context
**[ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) gave the
mesh one resolver and had each node ask it for the mesh's names only.** Because `resolv.conf` cannot
route by domain, that needed a stub on every node — a `systemd-resolved` module — and a separate
`dns` for the container runtime, which cannot use a loopback stub. It rejected the simpler shape,
every node sending every query to the mesh's resolver, on the grounds that *"a laptop whose tunnel is
down could resolve nothing at all."*
**That is true only of a `resolv.conf` naming the mesh's resolver alone.** The C library asks the
servers it lists in order and moves to the next when one does not answer within its timeout. A public
resolver listed second is asked exactly when the mesh's is unreachable — the anchor down, the tunnel
down, a laptop behind a captive portal that has not let the tunnel up — and never otherwise. An answer
from the first, including "no such name", is final, so `.internal` is never asked of a public resolver
while the mesh's answers.
**And the container runtime copies a machine's resolvers into its containers when they are not
loopback addresses.** With the mesh's resolver and a public one listed, every container gets both, as
they are, with nothing configured for the runtime.
## Considered Options
**1. Keep ADR 0194's stub.** Public names never touch the mesh, and a node with the anchor down
resolves public names at full speed. It costs a module and a running service on every node, a second
configuration for containers, and the one asymmetry ADR 0194 had to state — containers' public names
through the mesh, nodes' not.
**2. Every node asks the mesh's resolver for everything, with a public resolver as the silent
fallback.** One server answers every node and every container; nothing on a node routes, holds names,
or runs. Chosen.
## Decision
**A node's `/etc/resolv.conf` names the mesh's resolver first and a public resolver second, with a
short timeout and a single attempt.** It is written by the module holding `node-resolver-config` — the
existing `resolv-conf` — which now names `mesh-resolver`'s address instead of the machine's own. The
mesh's resolver answers the mesh's names from what it holds and forwards every other name, giving the
public answer ([ADR 0191](0191-the-meshs-resolver-holds-only-the-meshs-own-names.md) is unchanged: no
public name gets a private answer).
**Containers take the same two resolvers from their machine.** The container runtime's own `dns`
setting is not written; the runtime copies the machine's non-loopback resolvers into every container.
**This replaces, from ADR 0194:** the asking side as a stub (*"So the asking side is a
`systemd-resolved` module"*), the container runtime's `dns` naming `mesh-resolver`, and step 2 of the
migration as written. There is no `systemd-resolved` module. Everything else in ADR 0194 stands — one
`mesh-resolver`, on the node every tunnel converges on, holding each node's internal domain, the
retirement of `node-dns-resolver` and every per-node copy, and a LAN's resolver not being the mesh's.
**The migration, as it now reads:**
1. `mesh-resolver` is assigned and answers on the private network.
2. Each node's `resolv-conf` names `mesh-resolver` first and a public resolver second.
3. A LAN whose router points at a node's resolver is pointed at its router.
4. `node-dns-resolver` is unassigned from every node, and the hosts region is withdrawn.
## Consequences
- **Every name a node or container asks goes through the anchor while it is up.** A public lookup
takes a few milliseconds longer than asking a public resolver directly, and the mesh's resolver sees
every name its nodes look up. It is the operator's own server.
- **With the anchor unreachable, each lookup waits out one timeout, then resolves publicly.**
`.internal` names fail then — as `.internal` traffic does, every tunnel going through the anchor.
- **A LAN is unaffected by this choice.** Devices that are not members never read a node's
`resolv.conf`; they get their resolver from their router, which step 3 points at itself.
- **Nothing new runs on a node.** No stub, no module, no per-node configuration for containers.
- **The runtime's `dns` key goes with `node-dns-resolver`.** The dnsmasq module wrote it into the
runtime's configuration; unassigning that module in step 4 withdraws it, and the runtime reads the
change only when it next starts — with `live-restore` on, that restart keeps every container
running.
**How each is checked:**
- **Order:** each node's `/etc/resolv.conf` lists `mesh-resolver`'s private address first and a public
resolver second, and nothing else.
- **Fallback:** with `mesh-resolver` unreachable from a node, a public name still resolves there, after
the timeout.
- **Containers:** a container started on a node lists the same two resolvers.
- **A LAN:** the router's DHCP DNS option names the router, not a node.
## References
- [ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) — the one
resolver; this record replaces how nodes and containers ask it.
- [ADR 0191](0191-the-meshs-resolver-holds-only-the-meshs-own-names.md) — what the resolver holds.
- [Connectivity §2](../03-DESIGN/01-to-be/08-connectivity.md), amended alongside this record.
@@ -1,85 +0,0 @@
---
topic: what runs on it
status: accepted
date: 2026-10-03
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0195-the-meshs-tools-are-found-by-address-not-announced-whole.md
---
# 197. Every tool announces itself on the bus, in the NATS services protocol
## Context
[ADR 0195](0195-the-meshs-tools-are-found-by-address-not-announced-whole.md) gave every tool an
address and the console five tools to find them. Where the console learns what exists, it inherited
from [to-be 34](../03-DESIGN/01-to-be/34-the-console.md) §3: ask the catalogue for its roster, ask
each module on the roster for its `tools`, and read the machines and assignments from the
controller's printed `node list` and `module list`. Built that way on 2026-10-03, it worked and showed
what is wrong with it:
- **It asks what should exist and infers what does.** The roster holds every module the catalogue
ever registered; 47 of them were reported "not answering" on 2026-10-03, most with no tools at all
and several retired.
- **It parses prose.** Two of the controller's answers are text for a person, and a reworded column
breaks discovery.
The bus already knows what answers. NATS has a services protocol for exactly this: a service answers
`$SRV.PING` and `$SRV.INFO` — every instance, on one request — with its name, its instance, and every
endpoint's subject and metadata, in a published format the NATS tools read. The operator's direction:
*every tool announces itself; the mesh has the full picture, so nothing should be inferred or parsed.*
## Considered Options
1. **Keep asking the roster, and give the controller JSON answers.** Fixes the parsing, keeps the
inference.
2. **Re-serve every tool through a NATS services library.** The announcement for free, but every
runtime's serving path rewritten around a library, in two languages, for no change in behaviour.
3. **Every runtime answers the services protocol's discovery subjects with what it serves; serving
is unchanged.** Chosen.
## Decision
**1. What answers announces itself.** Every runtime that serves tools — each machine's tool runtime,
the per-module runtimes still in containers, and the controller for the seat it holds — answers
`$SRV.PING` and `$SRV.INFO` in the NATS services format: one service per module or seat it serves,
named for it, its instance the machine; one endpoint per tool, its subject and queue exactly as
served, its metadata the tool's description, argument schema, the machine, the seat and scope where
it is a seat's verb, and whether the module's instances are interchangeable.
**2. The console finds what exists by asking the bus,** one `$SRV.INFO` request, every answer
gathered for a short window. What it announces through ADR 0195's five tools is what answered.
**3. What should exist is the mesh's records, read as data.** The controller answers its machines and
modules as JSON, and says which modules declare tools; the console names as *not answering* only an
assignment that declares tools and did not announce them. A module with no tools is never listed.
**4. The grants say so.** Every principal that serves tools may subscribe the services discovery
subjects for what it serves; the console's and every runtime's account may publish the discovery
request. Replies travel to the asker's own inbox as every reply does.
## Consequences
- The standard `nats micro list` and `nats micro info` show the mesh's tools, live, to anybody holding
a credential — the bus's own view, not the mesh's description of it.
- Discovery costs one request and a gathering window, not one request per roster entry.
- What got harder: three runtimes must answer the same format the same way — the Go tool runtime, the
TypeScript runtime the containers still run, and the controller. The format is NATS's, so a test
reads all three with the NATS services client and nothing of the mesh's.
## How it is checked
| Rule | Checked by |
|---|---|
| A runtime announces exactly what it serves | each runtime's test: `$SRV.INFO` answered with one service per served module or seat, its endpoints' subjects the subjects served |
| The format is NATS's | the same tests read the answer with the NATS services client's own types |
| A module with no tools is never listed; an assignment with tools that did not answer is | the console's test, against controller records with both |
| Nothing is parsed from prose | the console reads only JSON answers (code review; the text parsers are deleted) |
| Live | `nats micro list` against the mesh's bus lists every machine's tool runtime and the controller |
## References
- [ADR 0195](0195-the-meshs-tools-are-found-by-address-not-announced-whole.md),
[ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md),
[ADR 0175](0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md)
- [to-be 34](../03-DESIGN/01-to-be/34-the-console.md)
@@ -1,79 +0,0 @@
---
topic: what runs on it
status: accepted
date: 2026-10-03
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md
---
# 198. A module's long-running code is launched by the node's runtime, and reaches the bus through it
## Context
[ADR 0193](0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md) made
every tools bundle a child the node's runtime launches, speaking MCP over stdio, and gave the channel
one bus verb: a tool's emit, published by the runtime as the module. Twenty-three modules still run the
rest of their own code — event handlers, provisioners, a preparation step, three mains — in a container
on the runtime's image, because that code needs what a container gave it: a bus connection that can
*subscribe*, and its module's words. Research [022](../01-RESEARCH/022-where-a-modules-long-running-code-runs/00-overview.md)
measured what it uses. [ADR 0188](0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md)
§1 says a module's own code is never an image.
## Considered Options
1. **The runtime launches it and is its bus.** Chosen.
2. A process per module with its own bus client and credential. Rejected for the reasons ADR 0188
rejected it for tools: a transport in every language's SDK, a credential per module on disk, and a
bus change rebuilding every module.
3. Keep the containers for it. Rejected: ADR 0188's rule stays broken for most of the catalogue.
## Decision
**1. A module's long-running code is a bundle the node's runtime launches and supervises,** exactly as
its tools are: an executable entrypoint, given the runtime's words and its module's
([ADR 0192](0192-a-tools-bundle-declares-what-it-is-given-and-the-runtime-hands-it-to-that-bundle-alone.md)),
started at the runtime's start and again when it exits. A bundle may serve tools, run long, or both.
**2. The runtime is its bus.** The stdio channel carries, beside MCP, the mesh's verbs a module's code
uses: `mesh/publish` (ADR 0193), **`mesh/subscribe`** — the runtime binds that module's durable
consumer, as the module's own runtime did, and delivers each event to the child as a `mesh/event`
request, acknowledging it on the bus only when the child has answered — and **`mesh/ask`**, a tool
call made on the module's behalf. The runtime's account is granted what each module it carries
consumes, and the consumer keeps the module's name, so no event is lost or replayed in the move.
**3. A preparation step is a run-once process the host runs before the runtime starts the module,**
with its module's words and no bus — what it already was.
**4. What a container reached by its network is reached on the machine.** A service by its published
port (`${port:…}`) on loopback; a backend's command-line client as a package of the machine's system,
or, where the system has none, the backend's own driver inside the bundle.
## Consequences
- The per-module containers go, and with them the runtime image as a way module code runs; ADR 0188's
registration rule can then refuse a module's own image without exception.
- One bus connection per machine carries every module's events; a module's handler is a function of
the events it is handed, in any language, with no bus client of its own.
- What got harder: the runtime holds every carried module's consumer and must not acknowledge an event
before the child has handled it — a child that dies mid-event leaves it unacknowledged, and it is
delivered again. The runtime's grants widen to what its modules consume.
- Two modules need code before they can move: their backends' clients exist on no machine's system,
so they talk to the backend through a driver instead.
## How it is checked
| Rule | Checked by |
|---|---|
| A subscribed event reaches the child and is acknowledged only after it answered | the runtime's test over a real bus: a child that answers is acknowledged once; one that dies mid-event is delivered again |
| The module's consumer keeps its name | the composer's test: the durable consumer the runtime binds is the one the module's own runtime bound |
| No module's own code is an image | the catalogue's registration check, without exception, once the last container has moved |
| Live | every moved module's provisioner and handlers act, on their machines, from the node's runtime; `docker ps` shows no runtime-image container on any machine |
## References
- [ADR 0188](0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md),
[ADR 0192](0192-a-tools-bundle-declares-what-it-is-given-and-the-runtime-hands-it-to-that-bundle-alone.md),
[ADR 0193](0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md)
- Research [022](../01-RESEARCH/022-where-a-modules-long-running-code-runs/00-overview.md)
- [to-be 38](../03-DESIGN/01-to-be/38-building-the-operators-machine.md) WP4c
+1 -13
View File
@@ -188,7 +188,7 @@ python3 00-META/checks/index.py fail if stale
- **0185** — [A control plane behind its seat's row serves what it can](0185-a-control-plane-behind-its-seats-row-serves-what-it-can.md)
- **0186** — [A ban list never holds a neighbour, and the mesh's own bans are its own wherever they hang](0186-a-ban-list-never-holds-a-neighbour.md)
- **0187** — [A dead tracker is not the machine's failure](0187-a-dead-tracker-is-not-the-machines-failure.md)
- **0190** — [A seat's work is shared by its holders, and building is the first such role](0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md)
- **0188** — [A provider declares what it derives for each consumer, and the mesh tells both ends](0188-a-provider-declares-what-it-derives-for-each-consumer.md)
### Its tiers, from the bottom up
@@ -222,9 +222,6 @@ python3 00-META/checks/index.py fail if stale
- **0126** — [A module declares its own seats; the mesh reserves its own](0126-a-module-declares-its-own-seats.md)
- **0148** — [The mesh's names are resolved, not copied into every container](0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
- **0151** — [A route's internal name is composed under the node that serves it](0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md)
- **0191** — [The mesh's resolver holds only the mesh's own names; a public name resolves publicly](0191-the-meshs-resolver-holds-only-the-meshs-own-names.md)
- **0194** — [The mesh has one resolver, and every node asks it for the mesh's names](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)
- **0196** — [A node asks the mesh's resolver first, and a public one only when it is silent](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md)
### What runs on them, and how it gets there
@@ -282,9 +279,6 @@ python3 00-META/checks/index.py fail if stale
- **0150** — [A module's own code runs as supervised processes under the module's one account](0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md)
- **0152** — [The operator's surface is a module the mesh assigns: the console](0152-the-operators-surface-is-a-module-the-console.md)
- **0155** — [A definition names no installation: how that is checked, and the three ways a value that did gets out](0155-a-definition-names-no-installation-and-how-that-is-checked.md)
- **0164** — [A setting is declared with its default, its meaning and what changing it costs](0164-a-setting-is-declared-with-its-default-its-meaning-and-what-changing-it-costs.md) *(proposed)*
- **0165** — [`container-runtime` is what a machine can run; that a runtime is running is its holder's health](0165-container-runtime-is-what-a-machine-can-run-and-a-running-runtime-is-its-holders-health.md) *(proposed)*
- **0166** — [The container runtime is a node seat, and the host creates containers through its holder](0166-the-container-runtime-is-a-node-seat-and-the-host-creates-containers-through-its-holder.md) *(proposed)*
- **0173** — [The operator's machine is the mesh's, and a module is whatever it declares](0173-the-operators-machine-is-the-meshs-and-a-module-is-what-it-declares.md)
- **0175** — [One tool runtime per node serves every module's tools, on the host side](0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md)
- **0176** — [The login shell is a node seat held by one shell module, and `execute` is its contract](0176-the-login-shell-is-a-node-seat-and-execute-is-its-contract.md)
@@ -292,12 +286,6 @@ python3 00-META/checks/index.py fail if stale
- **0181** — [The operator account is a node fact, and a home is a placement root](0181-the-operator-account-is-a-node-fact-and-a-home-is-a-placement-root.md)
- **0182** — [Inside a home, the mesh owns the directory and the files it places, writes into the tool's own files, and holds everything else as found](0182-inside-a-home-the-mesh-owns-what-it-places-and-holds-the-rest-as-found.md)
- **0183** — [The Anthropic licence manager is a module holding a seat; it hands each node's agent its token over the bus, sealed; the controller and the host have no part](0183-the-anthropic-licence-manager-is-a-module-and-hands-tokens-to-the-agent-over-the-bus.md)
- **0188** — [A module's own code is bundles in any language, and a tools bundle speaks MCP to the runtime](0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md)
- **0192** — [A tools bundle declares what it is given, and the runtime hands it to that bundle alone](0192-a-tools-bundle-declares-what-it-is-given-and-the-runtime-hands-it-to-that-bundle-alone.md)
- **0193** — [Every bundle the runtime serves is launched, and the runtime knows no language](0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md)
- **0195** — [The mesh's tools are found by address, not announced whole](0195-the-meshs-tools-are-found-by-address-not-announced-whole.md)
- **0197** — [Every tool announces itself on the bus, in the NATS services protocol](0197-every-tool-announces-itself-on-the-bus-in-the-nats-services-protocol.md)
- **0198** — [A module's long-running code is launched by the node's runtime, and reaches the bus through it](0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md)
### How it is built
+15 -71
View File
@@ -5,14 +5,10 @@ code:
- mesh-controller internal/catalogue/filtering.go
- mesh-controller examples/route-proxy
- mesh-controller internal/identity/authority.go
- mesh-controller cmd/mesh-controller/plan.go (the names the roster publishes)
- mesh-host internal/identity/serving.go
- mesh-host internal/apply (the service that reflects a rule set)
updated: 2026-10-03
updated: 2026-10-02
decisions:
- 02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md
- 02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
- 02-DECISIONS/0191-the-meshs-resolver-holds-only-the-meshs-own-names.md
- 02-DECISIONS/0180-the-found-front-end-is-uninstalled-once-a-machine-is-converged.md
- 02-DECISIONS/0170-the-firewall-seat-serves-its-verbs.md
- 02-DECISIONS/0169-a-machine-joins-through-the-tunnel-and-the-bus-is-never-public.md
@@ -292,10 +288,8 @@ expensively enough to be worth restating:
- **A node must not pin its own public name locally.** The duplicate record breaks resolution of
that name for everything else that needs it.
**What the host receives:** what to ask, not what to answer. The mesh has **one resolver**, holding
every node's internal domain; a node asks it first and a public resolver only when it is silent
([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md),
[ADR 0196](../../02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md)).
**What the host receives:** the resolver's configuration, as files, listing every peer's internal
name and overlay address.
**What goes away:** the `/etc/hosts` floor. It exists because a node had to reach the mesh
database before its own DNS existed; with [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)
@@ -353,33 +347,6 @@ not a list of containers.
somebody starts by hand is not the mesh's to configure, and reaching into every container on a
machine — declared or not — is what a nameserver in `resolv.conf` would be for.
### One resolver for the mesh
*2026-10-03.* **The mesh's names live in one place: the module holding `mesh-resolver`**, a mesh-scoped
seat of capacity one, placed on the node every tunnel converges on. It holds one wildcard per node —
`<node>.internal` and everything under it — and listens on the private network only. It answers the
mesh's names from what it holds and forwards every other name, giving the public answer.
**Every node asks it for everything, and a public resolver only when it is silent.** The module
holding `node-resolver-config` writes `/etc/resolv.conf` naming `mesh-resolver` first and a public
resolver second, with a short timeout and one attempt: the C library moves to the second only when the
first does not answer — the anchor or the tunnel down, a captive portal holding the tunnel back — so
public names keep resolving then, and `.internal` is never asked of a public resolver while the mesh's
answers. Containers take the same two from their machine, the runtime copying non-loopback resolvers
into every container, so the runtime is given no `dns` of its own
([ADR 0196](../../02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md),
replacing ADR 0194's per-node `systemd-resolved` stub).
**No node holds a copy.** The per-node resolver, its zones file and the mesh's region of `/etc/hosts`
go: every resolution fault found on 2026-10-03 was a copy disagreeing with the truth — a hosts file
read once at start, an operator's old line beside the mesh's, a node's resolver lent to a LAN. No
member's resolver answers a LAN; a router pointing at one is moved first. *Checked by each node's
`/etc/resolv.conf` naming `mesh-resolver` then a public resolver, by no node but the holder answering
DNS on any address, and by the router's DHCP DNS option naming the router.*
*What follows describes the per-node resolver this replaces — how it was built and why the roles were
split. The split stands; the serving role's scope is what moved.*
### The resolver, built
*2026-08-31.* **A service is reached at `<service>.<node>.internal`** — the first label is the
@@ -454,22 +421,7 @@ that module and nothing else.
the argument for the table in ADR 0009 being a table: the pattern is only obvious once seen, and
the cost of not seeing it is inventing a mechanism that already exists.
### The mesh resolves only its own names; a public name resolves publicly
**The mesh's resolver holds each node's internal domain and nothing else** — `<node>.internal` and
everything under it, so every route's internal name `<label>.<node>.internal` with no line of its own
([ADR 0151](../../02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md)).
**A node's public domains — one or more — are never given a private answer**: it is forwarded and resolves to the
public address, from a member and from anything else the resolver answers — a resolver may serve a
machine's LAN, and a phone on that LAN must get the address it can reach
([ADR 0191](../../02-DECISIONS/0191-the-meshs-resolver-holds-only-the-meshs-own-names.md)). Inside the
mesh, a routed service is reached, and certified by the internal authority, under its internal name.
*Checked by the controller's tests — the roster names the machines and no routed name —
and on a machine by asking its resolver for a public name the mesh serves: the answer is the public
address.*
*What follows is how the mesh got here, kept because the reasoning it rejects is the expensive half to
rediscover.*
### And the public names a proxy serves must resolve in the mesh too
*2026-09-09, found by an internal certificate authority that could not issue.* The mesh writes every
`<node>.internal` name into every declared container and treats the public names a proxy serves as a
@@ -492,13 +444,6 @@ would go stale the day one changes. The mesh propagates the names it was told to
knows nothing about what they mean
([ADR 0066](../../02-DECISIONS/0066-public-routing-is-name-agnostic.md)).
*2026-10-03, withdrawn.* Publishing public names with private answers turned every resolver that also
serves a LAN into an outage for that LAN's non-members — a phone was handed the control-node's tunnel
address for the mail server — while every check, run from a member, passed. Its reason had gone: routes
have internal names since ADR 0151, and the proxy certifies public names from a public authority and
internal names from the internal one. Superseded by the rule at the head of this section
([ADR 0191](../../02-DECISIONS/0191-the-meshs-resolver-holds-only-the-meshs-own-names.md)).
## 3 — Exposure
Settled by [ADR 0007](../../02-DECISIONS/0007-connectivity.md); summarised here because
@@ -926,15 +871,13 @@ is unchanged. **The same code path certifies against an internal authority as ag
only the issuer differs.** That is what makes trusted certificates possible for a mesh whose names
the public internet cannot resolve.
**The internal authority certifies internal names; a public one certifies public names.** The
authority's challenge reaches the name it certifies, so each certifies what it can resolve: the
internal authority a route's `<label>.<node>.internal`, which the mesh resolves, and a public authority
the public name, which public DNS resolves. A proxy holds both, and a public name is never certified
by the internal authority. *Checked by a handshake to a route's internal name that verifies against
the internal root and nothing else, and one to its public name that verifies against the public
roots* ([ADR 0191](../../02-DECISIONS/0191-the-meshs-resolver-holds-only-the-meshs-own-names.md)).
Until 2026-10-03 this paragraph had the internal authority certify public names, which needed them
resolved inside the mesh ([ADR 0066](../../02-DECISIONS/0066-public-routing-is-name-agnostic.md)).
**And it does not work until the routed name resolves inside the mesh** — the §2 finding above,
arriving here because this is what needed it. The authority's challenge reaches the routed name only
once that name is in internal resolution; a public authority is handed that dependency by public
DNS, and an internal one has to be handed it by the mesh. *Checked by a handshake to a routed name
that verifies against the internal root and nothing else — which cannot succeed unless the issuer
first reached the name to certify it*
([ADR 0066](../../02-DECISIONS/0066-public-routing-is-name-agnostic.md)).
## 6 — One statement behind exposure, filtering and certificates
@@ -1019,9 +962,6 @@ The list is worth having in one place, because it is most of the argument:
## Open
- **One resolver ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)).** Not built: every node still runs
`node-dns-resolver`. The migration's four steps are in the record, in order.
- ~~**What happens when the hub is down.**~~ **Resolved** by
[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md), together with `06`'s
matching question — they were one question. Nothing takes over. WireGuard has no failover, the
@@ -1043,6 +983,10 @@ The list is worth having in one place, because it is most of the argument:
operator's to move between meshes, but the manifest layer still stores it as a literal — so today
the composition is a per-node override rather than the design. The interpolation that would let a
module carry a label and a node carry the domain, and the mesh join them, does not yet exist.
- **Publishing route names into internal resolution.** The same ADR requires a granted route to be
resolvable inside the mesh, not only routable from outside it; the mechanism that writes
`<node>.internal` into containers does not yet also write the routed names, which is why an
internal issuer cannot currently validate one without a hand-placed entry.
## The hub adopts the predecessor's tunnel
+2 -26
View File
@@ -4,10 +4,9 @@ status: proposed
code:
- mesh-controller cmd/mesh-builder
- mesh-controller internal/builder
- mesh-catalog modules/build-agent
updated: 2026-10-03
- mesh-catalog modules/builder
updated: 2026-10-01
decisions:
- 02-DECISIONS/0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md
- 02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md
- 02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md
- 02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md
@@ -269,29 +268,6 @@ ships one and wrong for code the mesh built, which has no unit until the mesh wr
**Tools, hooks and consumers are not further modes**, which is the test of whether three is the
right number: they are loaded by a tool host, and a tool host is a process that stays up.
## Where a build runs
**On whichever machine holding the build role is idle** ([ADR 0190](../../02-DECISIONS/0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md)).
The role is `node-build-agent`, a node seat; its holder is the `build-agent` module, assignable to
every machine with a container runtime. The controller asks the role, never a machine: a tier's asks go
onto the seat's one work queue together, and each holder pulls one at a time when it is idle, so a
tier of many images is built by as many machines as hold the seat and are online, and a machine that
is off builds nothing and blocks nothing. What a holding machine needs is what the builder always
needed — a container runtime, the artifact store and the package registry as provisions, a workspace,
the bus credential — said once in the module's manifest. The outcome names the machine that built it.
This is the bus's shared-work pattern, not a build-specific one: any module declaring a node seat with
`accepts` has its work shared by its holders the same way. Building is the first use.
*Built and proven live 2026-10-03.* `build-agent` holds `node-build-agent` on all four machines; the first
build taken by a workstation's agent was a catalogue module at 09:35 UTC; the one-holder `builder` is
retired. The switch found five gaps, each an issue: a worker whose type changed stranded its holder
([206](../../04-ISSUES/206-a-seats-worker-changing-type-strands-the-holder-and-the-build-that-would-fix-it/00-report.md)),
a re-made worker replayed the stream's history ([207](../../04-ISSUES/207-a-re-made-worker-replayed-every-ask-the-stream-kept/00-report.md)),
a seat's worker is made only when the controller starts ([208](../../04-ISSUES/208-a-seats-worker-is-made-only-when-the-controller-starts/00-report.md)),
a module's identifier must fit the tightest backend's key (the slug), and an idle machine's empty fetch
was read as the end (fixed in the controller the same day).
## A build says what it does, as it happens
*2026-10-01 — [ADR 0157](../../02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md).*
+3 -5
View File
@@ -1,6 +1,6 @@
---
layer: to-be
status: in-progress
status: implemented
code:
- mesh-controller internal/catalogue/seats.go
- mesh-controller internal/catalogue/resolve.go
@@ -10,9 +10,8 @@ code:
- mesh-controller cmd/mesh-controller/source.go
- mesh-controller internal/inventory/migrations/0032-a-source-may-live-on-a-seat.sql
- mesh-catalog modules/gitea/module.json
updated: 2026-10-03
updated: 2026-10-01
decisions:
- 02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
- 02-DECISIONS/0161-what-deserves-a-seat.md
- 02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
- 02-DECISIONS/0126-a-module-declares-its-own-seats.md
@@ -127,8 +126,7 @@ convention, which later seats departed from.
| `mesh-npm-package-registry` | `npm-package-registry` | mesh | `npm-package-registry` | the forge |
| `mesh-git` | `git` | mesh | `git` | the forge |
| `mesh-build-machine` | `the-build-machine` | node | — | a builder |
| `mesh-resolver` | — | mesh | — | the mesh's one resolver, holding every node's internal domain ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)) |
| ~~`mesh-dns-port`~~ | `the-dns-port` | node | — | retired by [ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md): the local resolver became the mesh's one |
| `mesh-dns-port` | `the-dns-port` | node | — | the local resolver |
| `mesh-intrusion-prevention` | `the-intrusion-prevention` | node | — | an intrusion-prevention service |
| `mesh-packet-filter` | `the-packet-filter` | node | — | the packet filter |
| `mesh-private-network` | `the-private-network` | node | — | the private network the mesh runs over |
@@ -2,7 +2,7 @@
layer: to-be
status: in-progress
code: [mesh-controller internal/catalogue]
updated: 2026-09-30
updated: 2026-10-02
decisions:
- 02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md
- 02-DECISIONS/0155-a-definition-names-no-installation-and-how-that-is-checked.md
@@ -14,6 +14,7 @@ decisions:
- 02-DECISIONS/0084-which-provider-serves-a-consumer.md
- 02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md
- 02-DECISIONS/0038-the-mesh-assigns-the-port.md
- 02-DECISIONS/0188-a-provider-declares-what-it-derives-for-each-consumer.md
---
# 27 — A module requires, the mesh resolves
@@ -206,6 +207,22 @@ name when nothing sets it. That is the contract half of this design's operator p
the placeholder allows: the definition says which values reach which requirement, and nothing else
does. *How it is checked:* the unit tests named in issue 173, and the plan comparison that closed it.
*A provider says once what it derives for each consumer (2026-10-02,
[ADR 0188](../../02-DECISIONS/0188-a-provider-declares-what-it-derives-for-each-consumer.md),
[issue 124](../../04-ISSUES/124-a-consumer-cannot-be-told-what-its-provider-derived/00-report.md)):*
where a provider **names the resource** it gives each consumer — a bucket, a database, a vhost — the
name is derived per consumer, and a literal `serves` block could not carry it. A served value may
now name the consumer the mesh is serving: `${consumer:as}`, the identity the mesh minted, and
`${consumer:as:dns}`, that same identity written as a DNS label. Nothing else — **the mesh learns no
protocol here; it spells its own name in an alphabet it already knows.** Settings are laid on first,
so an operator may still set a prefix and the mesh derives the rest. The mesh fills it at the one
moment it knows who the consumer is, and the one filled value reaches both ends: the consumer, as
its binding's served facts and as `${bound:<provision>:<key>}` in any file it writes; the provider,
as `derived` on that consumer's entry in its contributions file, so its provisioner is told the name
rather than recomputing it. A consumer that writes the derived value into its own definition instead
of asking for it is refused, naming the placeholder to use. *How it is checked:* the unit tests in
ADR 0188's "how this is checked", each run against the unchanged controller first.
## How a definition reads what was resolved
**One form, naming a requirement and a field of its contract.** A definition that needs the database's
@@ -214,7 +231,9 @@ name in a configuration file writes the same thing: the requirement's name and t
controller fills it at resolution.
This one form replaces the placeholders that exist today, one per mechanism: bound values, secrets,
ports and machine facts.
ports and machine facts. It subsumes the consumer placeholder too — a value a provider derives is
read by the consumer exactly as any other field of the contract is, and `${consumer:…}` is only
how the *provider* states the rule.
**The seat placeholder stays, for the controller alone.** The controller composes its own
declaration and reaches the store and broker it made before any module existed, so it cannot be
+2 -26
View File
@@ -1,11 +1,9 @@
---
layer: to-be
status: designed
status: implemented
code: [mesh-catalog, mesh-tools, mesh-controller]
updated: 2026-10-03
updated: 2026-10-02
decisions:
- 02-DECISIONS/0197-every-tool-announces-itself-on-the-bus-in-the-nats-services-protocol.md
- 02-DECISIONS/0195-the-meshs-tools-are-found-by-address-not-announced-whole.md
- 02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md
- 02-DECISIONS/0159-a-tool-call-names-the-machine-and-a-holder-serves-its-seats-verbs.md
- 02-DECISIONS/0152-the-operators-surface-is-a-module-the-console.md
@@ -95,28 +93,6 @@ prerequisites are listed in that record. When the seat serves them, the console
modules' own, and the person stops opening a shell for the mesh's own questions. Until then the console
says so in its handshake.
## 3a. Found by address, not announced whole (2026-10-03)
*By [ADR 0195](../../02-DECISIONS/0195-the-meshs-tools-are-found-by-address-not-announced-whole.md);
this section governs where it and §2–§3 disagree.* The console announces five tools —
`mesh_overview`, `mesh_machine`, `mesh_search`, `mesh_describe`, `mesh_call` — and every tool the mesh
answers is reached through them by its address: `<seat>.<verb>` for a seat held once for the mesh,
`<node>/<seat>.<verb>` for one held per machine, `<node>/<module>.<tool>` for an assignment, and
`<module>.<tool>` as well for a module whose instances are interchangeable. A module that is not
interchangeable is called with its machine or refused with the machines it runs on. Each discovery
verb asks the mesh when it is called, so nothing is kept for a session's length; the flat catalogue
stays reachable through the `mesh` client and a setting, unannounced.
**What exists is what announced itself** ([ADR 0197](../../02-DECISIONS/0197-every-tool-announces-itself-on-the-bus-in-the-nats-services-protocol.md)):
every runtime answers the NATS services protocol's `$SRV.INFO` with what it serves, and the console
gathers one request's answers; the controller's records, read as JSON, say which assignments with
tools should have answered.
*Found 2026-10-03, measuring for that record:* §3's statement that a stateful module on two machines is
listed once per machine does not hold on the live console — postgres and mssql are listed once, `node`
optional, answered by whichever instance replies. The address replaces that statement rather than
repairing it.
## 4. Where it runs
On whichever machines an operator sits at, by assignment. It is not on the control node by default and
@@ -2,7 +2,7 @@
layer: to-be
status: in-progress
code: [mesh-tools, mesh-controller, mesh-host, mesh-catalog]
updated: 2026-10-03
updated: 2026-10-02
decisions:
- 02-DECISIONS/0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md
- 02-DECISIONS/0173-the-operators-machine-is-the-meshs-and-a-module-is-what-it-declares.md
@@ -11,10 +11,6 @@ decisions:
- 02-DECISIONS/0177-a-unit-may-be-user-scoped-and-the-service-manager-is-a-node-seat.md
- 02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md
- 02-DECISIONS/0149-the-live-mesh-is-the-test-bed.md
- 02-DECISIONS/0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md
- 02-DECISIONS/0192-a-tools-bundle-declares-what-it-is-given-and-the-runtime-hands-it-to-that-bundle-alone.md
- 02-DECISIONS/0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md
- 02-DECISIONS/0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md
---
# 38. Building the operator's machine
@@ -101,19 +97,6 @@ what `tools` answers, and the others serve. The runtime reads `MESH_OPERATOR_ACC
five tools and two seat verbs answer on their subjects; `tools` names the failed bundle; a
membership republished mid-run re-subscribes without a restart.
*Built and proven 2026-10-02* (mesh-tools, branch `feat/the-operators-machine`, commit `6390d1d`).
**WP1b — the launcher beside the loader** ([ADR 0188](../../02-DECISIONS/0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md)).
*mesh-tools, mesh-sdk. A day for the skeleton.* A bundle whose entry is not JavaScript is launched
as a child process with the runtime's environment and spoken to over MCP on stdio: `tools/list`
once, `tools/call` per call; a tool named `<seat>.<verb>` is the seat's implementation. A child
that exits is named as a failed bundle and restarted on the next call. The TypeScript import stays
as the shortcut. Beside it, one skeleton SDK per language of the first set — the stdio loop and the
tool-definition type, nothing else — each proven by one bundle in that language answering one tool
in the runtime's test. **Proof.** The runtime's test: a bundle in a second language, launched, its
tool answering on its subject over a real bus; the TypeScript fixture served through the protocol
with the shortcut off answers the same.
## WP2 — The controller composes one runtime per node
*mesh-controller. Two to three days; the largest package.*
@@ -128,23 +111,12 @@ with the shortcut off answers the same.
declaration gains an `archive` placed under a directory the controller derives, so the host
fetches and unpacks it as it does any artifact. The bundle's digest is what the build recorded.
3. **The runtime's process.** One `process` per node running the runtime from its own bundle
(WP3), `MESH_TOOL_MODULES` composed from the unpacked entrypoints — each as
`<module>=<path>`, and the runtime decides from the file whether it is loaded or launched
(WP1b) — `MESH_OPERATOR_ACCOUNT` and
(WP3), `MESH_TOOL_MODULES` composed from the unpacked entrypoints, `MESH_OPERATOR_ACCOUNT` and
`MESH_OPERATOR_HOME` from the account fact, `restart-on` naming every bundle so a push that
changes one restarts it. A node with no account composes the runtime without the two words.
4. **The gate.** A manifest declaring `tools` and a container built on the runtime's base image is
refused at registration once the runtime module is registered, naming this record. It is the
mechanism that keeps the old pattern from returning by habit. ADR 0188 widens it, after WP4:
a module whose own code is an image artifact is refused, whatever image it is built on.
*Amended 2026-10-02, at WP3.* The gate refuses the pattern **spreading**, not standing: a module
new to the catalogue in that shape, or one that had already moved to a bundle and returns to it,
is refused; a module the catalogue already holds in that shape — judged from the manifest it
holds and what that module's newest build stood on — is rebuilt without complaint. The day the
runtime arrives some thirty such modules stand, each moves in its own change from WP4 on, and a
gate refusing every rebuild in the meantime would stop the catalogue's pipeline to make a point
this record already makes.
mechanism that keeps the old pattern from returning by habit.
**Proof.** Composition tests: a node with three assigned modules, one holding a seat, yields one
process, three archives, one node principal whose grants are the union, and the same three
@@ -166,22 +138,6 @@ runtime's `serve` keeps answering MCP on loopback; the person's end of it keeps
wrote, `tools/list` on loopback answers as before, and the controller's verbs answer through it.
This is the first live step, and it is reversible by re-assigning `mesh-console`.
*Decided 2026-10-02:* `mesh-tools` keeps its name as the module the TypeScript images come from, and
`node-tools` is a second module in the same repository ([ADR 0069](../../02-DECISIONS/0069-a-module-is-a-repository-and-a-path.md))
rather than a rename — the thirty-five manifests that build `on` `mesh-tools` stay true. Three things
WP3 found that the plan did not say: a TypeScript bundle must carry its dependencies and a
`package.json` naming its files as ES modules, which the toolchain now copies in from its own image;
the runtime's credential must be owned by the account the runtime runs as, which the controller
composes; and `MESH_TOOL_MODULES` is empty on a node where the runtime is the only bundle, which the
runtime accepts. *Built 2026-10-02* (mesh-tools `c46f950`, mesh-controller `ca7e81e` `773b561`
`729a5f9`). *Proven live 2026-10-02/03, on all four machines*: the console's container is gone,
`node-tools` runs as a unit the host wrote, as the operator's account, `tools/list` on each loopback
answers with the same 219 tools as before, and the controller's verbs answer through it; `mesh-console`
retired from the catalogue. Three things the step found are issues
[203](../../04-ISSUES/203-a-fresh-assignment-is-pushed-before-its-credential-exists/00-report.md),
[204](../../04-ISSUES/204-a-controller-handover-re-sent-every-node-a-stale-declaration/00-report.md) and
[205](../../04-ISSUES/205-a-package-resource-fails-against-a-stale-package-database/00-report.md).
## WP4 — The first holder moves: the packet filter
*mesh-catalog. Half a day. The live proof of ADR 0175.*
@@ -194,121 +150,6 @@ tool where they need root, which they have, since the runtime runs as the node's
four machines; `docker ps` shows no `mesh-nftables`; `status` is well. Then the fail2ban holder
proposed in an open change follows the same way when it lands.
*Found 2026-10-03, before the step ran:* a bundle imported in-process brings its own copy of the SDK
(WP3's *carry its dependencies*), and the SDK's tool registry is the copy's own — the first module
loaded beside the runtime would have registered its tools where the runtime never looks, and served
nothing, silently. Issue
[209](../../04-ISSUES/209-a-bundles-own-sdk-copy-registers-into-a-registry-the-runtime-never-reads/00-report.md):
the runtime now resolves every bundle's import of the SDK to its own copy, one registry and one
broker per node. Two things the package did not say, settled in the module: the filter's commands
run through `sudo` without a prompt where the runtime is not root, since the operator's account may
escalate as the operator would; and a bundle has no environment of its own, so the tool reads the
filter from the path the manifest's `filtering` names rather than from a variable the container used
to carry, a test holding the two together. Three things a review of the change found: the module's
own bus credential and state directory went with the container, since nothing reads them once the
runtime speaks with the node's (the shell module of WP5 declares neither); the `iptables` package the
image used to carry is now declared on the host; and that the operator's account may escalate without
a prompt is a fact about the machine the mesh neither declares nor checks — true on all four today,
and when it is not, the tool names it by how it failed, which is the only check there is until a
record says where the fact belongs.
*Built 2026-10-03* (mesh-tools `7152148` for issue 209, mesh-catalog `db5e7c8`). *Proven live
2026-10-03, on all four machines*: `node-packet-filter.rules`, `reload` and `remove` answer from the
node's runtime on each — `rules` and the module's own tool list the mesh's table, `reload` loads the
file and answers with the table, `remove` refuses the mesh's own table by name — `docker ps` shows no
`mesh-nftables` on any, the container's credential is gone with it, and `status` is well. One thing
the step found is issue
[210](../../04-ISSUES/210-the-host-re-creates-the-nodes-runtime-on-every-reconcile/00-report.md):
the host re-creates the runtime's process on every reconcile (resolved the same day, mesh-host #80).
*fail2ban followed 2026-10-03* (mesh-catalog `aa5bf7d`), the same shape: container, base images,
credential and state directory gone, the client through `sudo` since the daemon's socket is root's;
proven on all four machines — `status`, `banned` and the module's own `fail2ban_settings` answer from
the runtime, no `mesh-fail2ban` container, the runtime serving both bundles. Two holders moved; of the
thirty-three tool containers the catalogue held, thirty-one remain, and all but these two carried their
module's configuration and secrets in the container's environment, which a bundle does not have — the
question research [020](../../01-RESEARCH/020-what-a-bundled-tool-is-given/00-overview.md) opened
and [ADR 0192](../../02-DECISIONS/0192-a-tools-bundle-declares-what-it-is-given-and-the-runtime-hands-it-to-that-bundle-alone.md)
settled the same day: a tools bundle declares `env` on its artifact, the composer resolves it as a
container's, the runtime hands each bundle its own. That is WP4b below.
## WP4b — Every tool container moves
*mesh-controller, mesh-tools, mesh-catalog. One day. The rest of ADR 0175, under ADR 0192.*
**What changes**, in order: the manifest's tools artifact gains `env` and the catalogue check
refuses a secret's content in it; the composer resolves a bundle's `env` per machine and carries it
beside the bundle's archive, `restart-on` included; the runtime hands each bundle its own
environment — the contributor's argument for an imported bundle, the child's environment for a
launched one — and a test holds two bundles apart. Then the thirty-one remaining tool containers
move in one change: each container's `env` becomes its tools artifact's, mount targets folded into
the host paths they came from, the container, its base images, its Dockerfile and its own bus
credential gone. Last, the registration gate refuses the container shape for every module.
**Proof.** The controller's and the runtime's tests named in ADR 0192; live, every module's tools
answer from the runtime on the machines that run it, `docker ps` shows no tool container on any of
the four, and `status` is well.
*Found 2026-10-03, building it:* of the thirty-one, nine run only tools, three a main of their own,
and twenty import the module's own event handlers and provisioners beside their tools (ADR 0192's
dated note). WP4b moves the tools-only nine; WP4c holds the rest. Two of the nine stay with WP4c as
well — one carries a run-once provisioning step in a second container, one reads an env-file and two
sockets — so seven move here. *Built 2026-10-03:* mesh-sdk #12 (`collectToolsEach`), mesh-tools #33
(each bundle its own environment), mesh-controller #236 and #237 (the words composed, made the
account's to read, and named files restarting the runtime), and the seven modules in one change.
Building it found issue [211](../../04-ISSUES/211-a-bundle-is-built-before-the-toolchain-it-is-compiled-in/00-report.md).
*Proven live 2026-10-03* (mesh-catalog #242, #243): on the one machine that runs them, baserow,
letta, searxng and unifi answer from the runtime with no tool container — each reading its
configuration file as the operator's account — and the runtime serves seventeen tools for six
modules there. confluence, gitlab and jira are assigned nowhere and retire with the predecessor.
Two traps met on the way: a tools bundle whose module declares no `tools` list must say `loads`, or
the composer delivers it nowhere while the build reports success; and a module whose builds are
pinned to an old commit is left out of a merge's plan and must be built from `main` by hand.
## WP4d — Every served bundle is launched; the runtime in Go
*mesh-sdk, mesh-controller, mesh-tools. [ADR 0193](../../02-DECISIONS/0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md).*
**In order.** The SDK's stdio loop serves what a bundle registered under the module it is told it
serves as. The builder writes, beside every TypeScript entrypoint, an executable launcher that
imports it and serves what it registered; the composer names the launcher where it named the
entrypoint. The runtime launches every served entrypoint and imports none; the resolve hook and the
per-registration environment go. Proven live on all four machines. Then the runtime is rewritten in
Go against the same contract — the bus, the memberships and seats, the launcher, the console's MCP
over HTTP — and replaces the TypeScript one, proven the same way.
**Proof.** The tests ADR 0193 names; live, every moved module's tools and both node seats answer from
launched bundles on every machine, and then do again from the Go runtime.
*Built and proven live 2026-10-03.* mesh-sdk #13/#14 (0.1.4, 0.1.5: served as the named module; an
emit travels through the runtime), mesh-controller #239/#240 (a launcher beside every TypeScript
entrypoint; a runtime compiled to a binary runs itself), mesh-host #81 (`./name` is the process's own
binary), mesh-tools #35/#36/#37/#38 (the module named; launch-only; the toolchain requiring 0.1.5; the
runtime in Go). On all four machines node-tools is now the Go binary, launching every served bundle:
both node seats answered from it on every machine and the four moved modules on theirs. Found on the
way: issue [212](../../04-ISSUES/212-a-toolchain-rebuild-keeps-the-sdk-it-cached/00-report.md) (the
toolchain image kept a cached SDK, and the seats' verbs went unanswered on three machines for an hour),
and the controller's plan losing track of its own rebuild when it restarts mid-plan.
## WP4c — The module's own long-running code moves
*Not yet broken down.* Twenty-three containers carry code that is not a tool: event handlers,
provisioners, a step, a main. [ADR 0188](../../02-DECISIONS/0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md)
§3 already says such code is a `process` bundle the host runs. What no record says yet is how that
process is given what its container was: the module's own bus credential and the subscriptions it
consumes with, the words its code reads at import, the packages the image installed (a database's
client), and the service it reaches by a container network name. That begins with a decision record,
after which the twenty-three move and the registration gate refuses the container shape for all.
*Decided 2026-10-03, [ADR 0198](../../02-DECISIONS/0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md):*
the node's runtime launches that code as it launches tools, and is its bus — `mesh/subscribe` and
`mesh/ask` beside `mesh/publish` on the stdio channel, the module's own durable consumer bound by the
runtime and acknowledged only after the child answered. **In order:** the runtime's subscription and
its grants; the SDK's `on` and provisioner bound to the channel; then the modules in three waves — the
provisioners and handlers whose backends are reached on loopback with a system package (postgres,
redis, mosquitto, influxdb, keycloak, umami, cloudflare-dns, grafana, icecast, home-assistant, nodered,
nextcloud, minio), the two whose clients exist on no system (mongodb, mssql: a driver in the bundle),
and last the mesh's own (mesh-catalog, mesh-vault, records, gitea, mailu, audit-logger, lab, and the
three mains).
## WP5 — The shell, on a server first
*mesh-catalog #224, already written. Half a day to assign and prove.*
@@ -1,9 +1,9 @@
---
status: located
status: resolved
opened: 2026-09-26
located-in: [mesh-controller internal/catalogue/declaration.go, mesh-sdk src/provisioner, mesh-catalog modules/minio]
fixed-by:
amended-design:
located-in: [mesh-controller internal/catalogue, mesh-sdk src/provisioner, mesh-catalog modules/minio]
fixed-by: 02-DECISIONS/0188-a-provider-declares-what-it-derives-for-each-consumer.md
amended-design: 03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md
---
# 124 — A consumer cannot be told a value its provider derived for it, so it transcribes one
@@ -63,3 +63,27 @@ compares it to what the provider will actually create. The one wrong instance wa
- What would have caught the wrong instance? A test that resolves a consumer's grant and compares the
bucket in its own configuration against the one the provider would create is a check that could
exist today, for any interface, without the mechanism above.
## Answered, 2026-10-02 — [ADR 0188](../../02-DECISIONS/0188-a-provider-declares-what-it-derives-for-each-consumer.md)
The channel is the provider's own `serves` block, which may now name the consumer the mesh is
serving: `${consumer:as}` and `${consumer:as:dns}`. The mesh fills it once, where it knows who the
consumer is, and delivers the one filled value to both ends — the consumer's binding and its
`${bound:…}` substitutions, and the provider's contributions entry, so a provisioner is told the
name rather than deriving it. Each open question above, answered:
- **Should a provider return values from provisioning?** No. It would make a grant carry data the
provider wrote, make a consumer's declaration wait on its provider's reconcile loop, and put the
rule where nothing can refuse it. The reasoning is in the record.
- **Or should `serves` say a value is derived?** Yes, and the mesh performs the derivation — but it
learns no protocol doing it. The only fact is the identity the mesh itself minted, in one of two
alphabets it already knows.
- **Should a consumer that names the resource be refused?** Yes. A consumer's file that already
contains the value the mesh is about to derive for it is refused at resolution, naming the
placeholder to write instead. That is the check this report asked for, and it is exact rather than
heuristic: a derived value carries the identity minted for this consumer on this machine, which
nothing else would spell out.
minio's `bucketFor` is gone; its manifest serves `"bucket": "${consumer:as:dns}"`. The three
consumers' hand-written bucket names are gone with it — each of them also named the machine the
module happens to run on, which is the second thing wrong with a transcription.
@@ -1,85 +0,0 @@
---
status: located
opened: 2026-10-01
located-in: [mesh-catalog modules/dnsmasq, mesh-controller internal/overlay/generator.go, mesh-controller internal/catalogue/resolve.go (checkResources)]
fixed-by:
amended-design:
---
# 190 — The container runtime's configuration is written by modules that are not the runtime's
## What was observed
The runtime's configuration file and its service are declared by two parties, neither of which is
the runtime:
- **The resolver module** writes the runtime's `dns` key (the machine's private address) and
`live-restore` into the runtime's file, written into rather than over
([ADR 0102](../../02-DECISIONS/0102-the-mesh-writes-into-a-shared-file-never-over-it.md)). It also
declares the runtime's service, reloaded when that file changes. The `dns` key has been written
since the resolver module was converted from its predecessor on 2026-09-23; `live-restore` and the
service were added on 2026-09-30 while fixing
[issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md),
where containers silently resolved through a public resolver.
- **The private network** writes the runtime's `insecure-registries` into the same file, and declares
the same service reloaded on it, as [ADR 0082](../../02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)
and ADR 0102 decided. The controller generates both resources per machine.
On the three machines that run the resolver module, both declare one path and one unit. Nothing refuses
it. The collision check compares the resources of catalogue modules. The private network is computed,
so its resources are produced when a machine's declaration is composed, and the check never sees them.
The machine without the resolver module shows the other half. Its runtime still has the predecessor's
resolver and `live-restore` off, because the only module that sets them is a DNS server. A machine
gets a correct container runtime only as a side effect of being given a resolver.
> **Later the same day, 2026-10-02.** The resolver module and its sibling for the resolver file were
> assigned to the fourth machine ([issue 198](../198-the-lans-dns-server-ran-outside-the-mesh-and-its-filter-closed-it/00-report.md)),
> so all four now have the resolver writing into the runtime's file, and the predecessor's
> `live-restore: false` there is gone. The same work made the runtime's file, as the resolver declares
> it, take no settings: a setting meant for the resolver's own configuration had reached it. The
> collision and the ownership question above are unchanged.
## Why this is here
The operator ruled it a defect, not a design: **a module does not write another software's
configuration.** The need behind each write is real. Containers must resolve the mesh's names
([ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) step 2).
A daemon restart must not stop every container. Every machine on the network must trust the mesh's
registry. But each of these is a fact the runtime must be *given*, and the module that gives it is the
runtime's own. With three writers, nobody can say what the file should contain. Two of the facts are
reloaded when one of them needs a restart (issue 110's first fault). And the moment a module for the
runtime exists, it is refused on every machine with the resolver, or, through the private network's
path, accepted without anyone noticing a collision.
## What resolves it
[ADR 0166](../../02-DECISIONS/0166-the-container-runtime-is-a-node-seat-and-the-host-creates-containers-through-its-holder.md)
gives the runtime a module that holds its seat and owns its file and service.
[ADR 0164](../../02-DECISIONS/0164-a-setting-is-declared-with-its-default-its-meaning-and-what-changing-it-costs.md)
gives that module declared settings with defaults. The fix, once both are accepted:
1. The resolver module drops its runtime file and runtime service. It knows nothing of the runtime.
2. The private network stops generating either resource. ADR 0082's decision stands — being on the
network is what grants the trust, and no module author is involved — and only *who writes it*
moves. The mesh gives the registry to the runtime module as a value. ADR 0082 and ADR 0102 each
get a dated note saying where their mechanism now lives.
3. The runtime module writes `dns`, `live-restore` and `insecure-registries`, each a declared
setting with its cost: `dns` costs a restart, which `live-restore` makes harmless.
4. Steps 1–3 land in one push. A runtime module declaring the file beside a resolver module still
declaring it is refused.
5. The collision check sees a computed module's resources as well, so a second writer cannot come
back through generated code.
## Open questions
- **How the resolver's address reaches the runtime.** Either the resolver seat (`node-dns-resolver`)
delivers an address its holder serves, or the runtime module reads a machine fact and the seat
being held is only a precondition. The first tracks a resolver moving off the private address. The
second needs nothing new.
- **What `dns` defaults to on a machine with no resolver seat held.** Nothing, leaving the runtime's
own behaviour, is the honest default. A public resolver hides exactly the failure issue 110 took a
day to find.
- **The adopted machine's predecessor values.** The runtime module adopting a file with a
hand-written `dns` and `live-restore: false` replaces both. That is intended, and is the one
restart the operator must make on that machine.
@@ -1,78 +0,0 @@
---
status: open
opened: 2026-10-02
located-in: [mesh-catalog modules/mesh-console, mesh-controller cmd/mesh-controller/plan.go (port assignment)]
fixed-by:
amended-design:
---
# 192 — The mesh's tools reach a person only by a registration made by hand
## What was observed
A design session on a workstation had none of the mesh's tools. The console was running on that
machine and answering on its loopback port. It was reached over the bus as the console's account, and
listed every running module's tools and every seat's verbs
([ADR 0152](../../02-DECISIONS/0152-the-operators-surface-is-a-module-the-console.md)). What was missing
was the registration that tells the person's coding agent where the console is. That registration had
been made by hand, once, while migrating the machine, and scoped to the one project directory it was
made in. Every session started anywhere else had no mesh tools. Nothing said so: the agent simply
offered no mesh tools, and the session fell back to a pull-request link for a person to open by hand.
The predecessor did this job itself: it wrote its tool server into the agent's user configuration on
every machine. Migrating removed that entry, as it should have, and no module took the job over.
## Why this is here
Three gaps, each of which would have stopped a module from doing it even if one existed.
**1. The console tells nobody where it is.** Its definition listens on a port and provides nothing.
A module that wanted to point an agent at the console has no requirement it could name, so it
would have to write the address into its own definition as a literal. That is exactly what
[ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md) and
[ADR 0155](../../02-DECISIONS/0155-a-definition-names-no-installation-and-how-that-is-checked.md)
remove.
**2. The console's port is one its definition chose.** The definition names a port, and the mesh
never assigned one: no port assignment exists for the console on any machine. The plan assigns a
machine port only to a port a container publishes through a mapping, "without one the software binds
what it binds". The console runs on the host network with no mapping, but it reads its listening
address from `${port:…}`, so the mesh could move it and does not. That is a module choosing a
machine port, which [ADR 0038](../../02-DECISIONS/0038-the-mesh-assigns-the-port.md) exists to
prevent, through a gap in how the rule is applied rather than a decision against it. A module that
reads its port from the mesh should be assigned one like any other.
**3. Nothing in the mesh owns a person's agent configuration.** No catalogue module writes the agent's
settings, its tool-server registrations, or the rules and skills the predecessor delivered. On the
four machines these are hand-kept, or left over from the predecessor, or missing.
## What a fix looks like (not decided)
- **The console provides its endpoint.** A provision, working name `mesh-tools`, served as the URL on
the machine port the mesh gives it. The console listens only on loopback, so the provider must be on
the consumer's own machine. Co-location already chooses it
([ADR 0084](../../02-DECISIONS/0084-which-provider-serves-a-consumer.md)), and a machine with no
console refuses the consumer, naming the provision.
- **A module for the coding agent requires it** and writes the registration into the agent's
system-wide managed settings. The agent reads tool servers from a `managedMcpServers` key there. That
file is the machine's rather than a user's, so the module owns it whole and no home directory is
named. People keep their own registrations beside it. The agent's separate *exclusive* managed
file is the wrong one: it blocks every registration a person makes and hides the hosted connectors.
The agent's per-user file is rewritten by the agent continuously and sits in a home directory,
which would make its path an operator value. These facts come from the agent's documentation
(managed MCP and managed settings pages), not yet verified on a machine.
- **The same module owns the rest of the agent's configuration** the predecessor delivered: managed
settings and the rules, skills and instructions every session reads. Each declared setting carries
a default (ADR 0164,
proposed on its own branch), so one configuration serves every machine and one machine may differ.
## Open questions
- **Is the agent's configuration one module or several?** Tool registration, managed settings, and
the instruction files have different readers and change at different rates.
- **Whose machine port is the console's?** Should a host-network container that reads its port from
`${port:…}` be assigned one, or should a machine-only listener keep its declared number? The second
needs a decision, because ADR 0038 does not allow it today.
- **Credentials.** The console's authority is the machine's login (ADR 0152). A registration that
reaches it carries no secret today. If the console ever listens beyond loopback, the registration
needs one, from the vault.
@@ -1,112 +0,0 @@
---
status: resolved
opened: 2026-10-02
located-in: [mesh-catalog modules/postgres/client.ts (readOnlyQuery)]
fixed-by: [mesh-catalog PR 209 (postgres), mesh-catalog PR 210 (mssql)]
amended-design:
---
# 193 — The store seat's read-only query is read-only by convention, and its answer is unreadable
## What was observed
Asking the store seat's `query` verb for a count through the console returned this. Rows are each
wrapped in an object under a key named `BEGIN`: the column name, then the value, then the word
`ROLLBACK`. A query returning nothing gave the column name and `ROLLBACK` alone. The answer to
`select count(*) as n from <table>` was:
> `rows: [ {BEGIN: "n"}, {BEGIN: "46"}, {BEGIN: "ROLLBACK"} ]`
A reader can work it out. A program cannot, and a query with two columns loses which value belongs to
which.
## Why this is here
**The cause is the same line that makes the query read-only.** The holder's tool sends
`BEGIN TRANSACTION READ ONLY; <the caller's statement>; ROLLBACK;` to the command-line client as one
string. The client prints a command tag for each of the three statements, and the parser takes the
first line, `BEGIN`, as the header.
**And it is not read-only.** [ADR 0159](../../02-DECISIONS/0159-a-tool-call-names-the-machine-and-a-holder-serves-its-seats-verbs.md)
decided the store seat's `query` verb is "one read-only statement against one database". The only
thing enforcing that is the wrapping transaction, and the caller's statement is pasted inside it as
text. A statement that begins by ending the transaction (a commit, then anything) runs whatever
follows it outside the read-only transaction, with the holder's own role, the administrative one that creates every
consumer's role and database. A rule stated in a decision and enforced by string concatenation is enforced by nothing.
*This is read from the code, not tried against the live store, and it should not be tried there.*
The lab bed is where it gets proven.
Every caller with `invokes` on the store seat's `query` can do this. The console has `invokes: ["*"]`,
so that includes anyone logged in on a machine running the console.
## What a fix looks like
- **One statement, refused otherwise.** Send the caller's statement alone, through the client's
single-statement path (the extended protocol takes one statement per call and refuses more). The
read-only property then comes from the session, not from text around the statement.
- **Read-only by role, not by transaction.** Run the verb as a role that can only read, granted
`pg_read_all_data`, not as the administrative role. A statement that escapes every wrapper still cannot write.
- **Rows as rows.** Parse the client's output with the column names it returns, or use a driver
instead of the command-line client, so a row is an object keyed by its columns.
- **The check 0159 lacks:** a test that sends a commit followed by a write and asserts the write is
refused and nothing changed. Another asserts a two-column row comes back keyed by both columns.
## Proven, 2026-10-02
On a throwaway server — the same engine image, no network, reached over a socket — the module's code
from the catalogue's main branch ran `COMMIT; COPY (select 1) TO PROGRAM '<a command>'` and **the
command ran on the database host** as the server's own user. `COMMIT; DROP TABLE t` executed the drop
outside the read-only transaction; the wrapper's own trailing rollback happened to undo it, which a
caller ending their statement with a commit of their own would get past (not tried). Nothing was tried
against the live store.
The fix (mesh-catalog PR 209) runs the caller's statement as a login granted `pg_read_all_data` and
nothing else, read-only by its role and its session, with a password the mesh mints as one of the
module's own secrets; without that password the call is refused rather than run as the admin. On the
same throwaway server every escape above, and `SET ROLE`, `RESET SESSION AUTHORIZATION`, turning
read-only off, creating a table, altering the role and reading a server file, is refused; a plain
select comes back keyed by its columns. One attempt — turning the transaction's read-only off, then
deleting — got past the first layer and was stopped by the second, which is why both exist.
**Not answered by the statement-count fix proposed above.** The command-line client sends one string
in one message, so several statements still arrive together. They are harmless as the reader, and
refusing them is left to whoever moves the module to a driver.
## The same hole, elsewhere — and two worse ones
The `mssql` module wrapped a caller's statement the same way (`BEGIN TRANSACTION; … ROLLBACK;` as its
administrator) for its `mssql_query` tool. Its command-line client added two holes of its own. Both
were proven on a throwaway server, running the client the way the module ran it:
- **It substitutes `$(NAME)` from its environment into the caller's text**, and the administrator's
password is in that environment. Selecting it as a string returned the password.
- **It reads a line beginning `:!!` as a command that starts a program**, in the container that holds
the administrator's password and the module's bus credentials. Its switch for refusing such commands
makes the shipped version ignore the statement entirely, so the switch cannot be the guard.
None of it was reachable on the live mesh, for a reason that is a defect of its own: the runtime image
never installed the client, so every mssql tool failed (`spawn sqlcmd ENOENT`). The fix (mesh-catalog
PR 210) installs the client at a pinned digest and runs the caller's statement as a login that can
connect and read and do nothing else. Substitution is off. The statement must be one line, placed after
the module's own text, so no line of it can begin a command; a line break is refused before the client
starts. On the throwaway server, writes, `xp_cmdshell`, impersonating the administrator, and joining
the administrators' role were all refused, and the variable came back as the literal text.
**The general lesson**, worth more than either module: *a command-line client is an interpreter with
its own syntax, and a caller's text handed to it is a program in that syntax as well as in SQL.* A
module that passes a caller's text to a client has two languages to defend, and a transaction drawn
around the text defends neither.
## Resolved, 2026-10-02
Both pull requests merged, built and pushed to the two machines that run each module. Checked live, on
every copy, by asking each one who it is:
- the store seat's `query`, and postgres's own tool on each machine, answer as the reader login —
not a superuser, in a read-only transaction — with rows keyed by their columns;
- mssql's tool, on each machine, answers as its reader login, outside the administrators' role, and
returns `$(SQLCMDPASSWORD)` as the literal text it is. Its tools work for the first time.
The escapes themselves were tried only on the throwaway servers above; on the live mesh the check is
the identity a statement runs as, which is what makes every escape a statement that the login cannot do.
@@ -1,45 +0,0 @@
---
status: resolved
opened: 2026-10-02
located-in:
- mesh-controller
fixed-by: mesh-controller PR #233
amended-design:
---
# 203 — A fresh assignment is pushed before its credential exists, and the runtime crash-loops
## What was observed
2026-10-02, the first live assignment of the node tools runtime (to-be 38 WP3). `assign` put the
module on a machine and `push` sent the declaration. The host applied everything: the bundle unpacked,
the unit written and started, the module's `broker` secret file written and owned by the operator
account. The runtime then restarted thirteen times in a minute:
```
mesh-tools: cannot read the broker credential at …/broker: SyntaxError: Unexpected token 'O',
"Oj6j2Ssa-v"... is not valid JSON
```
The file held a 40-byte random secret, not a bus credential. The push's own output had said why,
one line among forty: *the bus's user list leaves out … `<node>.node-tools`. Each is a user that cannot
connect until one is issued.* The credential exists only after `module issue <module> --node <node>`,
a separate act; a second push then carried the real credential and the runtime came up. The same
sequence repeated on the next two machines, by hand, in the right order.
## Why it matters beyond this instance
Every module that speaks on the bus declares an `own-secrets.broker`; the mesh seals *something*
there on assignment and the real credential only on issue. So the first push of any fresh assignment
delivers a process that cannot authenticate and will crash-loop until a person runs a second verb and
a second push. Nothing refuses the first push, and the warning is a line in a long list that is
printed on every push regardless. The design says the mesh issues an assignment's subjects and the
runtime serves what it is issued ([ADR 0160](../../02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md));
an assignment whose credential is not issued is half an assignment, and the mesh lets it through.
## Questions HQ must answer
- Is issuing the credential part of assigning, so `assign` mints it, or must a push refuse a module
whose bus user is unminted, naming the verb?
- Is a placeholder sealed where a credential belongs ever right, or should the resource be absent
until the credential exists, so the host never writes a file the process cannot read?
@@ -1,18 +0,0 @@
# Diagnosis — 203
**2026-10-03.** Two acts, one effect. `assign` records the assignment and, at composition, every
`own-secrets` entry a module declares gets a sealed value from the controller's `Needed` map — a
value minted so the file exists, which for `broker` is a random secret, not a credential. The bus
credential is composed only by `module issue <module> --node <node>` (cmd/mesh-controller/modules.go,
`issueOnTheNewBus` → `issueWith`): it mints the bus user, records its hash, and seals the credential
JSON into the same `broker` need. Nothing joins the two: `push` composes and sends whatever the need
holds, and the only warning is the standing line listing every bus user without a minted credential,
printed on every push regardless of what was just assigned.
Ruled out: the host (it wrote the file it was given, owned as asked); the runtime (it refused a file
that is not JSON, correctly, and said so); the manifest (`own-secrets.broker` is the shape every
module uses).
**Owner:** mesh-controller — the assign path. **Fix direction:** assigning a module that declares
`own-secrets.broker` issues its credential in the same act, idempotently; a push of a module whose bus
user is unminted is refused by name rather than sent with a placeholder.
@@ -1,49 +0,0 @@
---
status: resolved
opened: 2026-10-02
located-in:
- mesh-controller
fixed-by: mesh-controller PR #232
amended-design:
---
# 204 — A controller handover re-sent every node a declaration composed from a stale view
## What was observed
2026-10-02 21:29 UTC. Two machines were assigned the node tools runtime and had the console taken off
them, and were pushed; the host on each applied it (the console removed, the runtime created and
running). Two seconds later, on each, a second declaration arrived that undid it — the host's log:
```
23:29:18 created node-tools.runtime (node-tools): 239 file(s), running as node-tools.service
23:29:18 applied 569 resource(s)
23:29:20 removed node-tools.runtime (node-tools)
23:29:20 forgotten node-tools.interpreter (nodejs)
23:29:24 created mesh-console.needs-broker, created mesh-console.server (mesh-console)
```
The second declaration had the console assigned and the runtime absent: the assignments as they were
a minute earlier. The controller's status knew of one send per machine, the person's. At that minute a
plan from an unrelated merge was rolling a new controller build onto the control node, so an instance
was starting while another was stopping. A third push by hand, two minutes later, restored both
machines and nothing undid it again.
Which instance sent the stale declaration, and from what, is not established: the outgoing one on its
way down, the incoming one at start-up before its view was current, or the rolling plan sending what it
had composed when it was made — the same family as [issue 201](../201-a-push-recreated-the-controller-behind-the-row-its-successor-wrote/00-report.md),
where a plan sent a digest older than the one a successor had written.
## Why it matters beyond this instance
A declaration the mesh sends is the mesh's word on what a machine should be; a machine applies it in
full, including removing what it no longer names. A stale one is therefore not a no-op: it tears down
whatever was assigned since, and the controller's own record does not show the send, so the next
person reads "applied, current" over a machine that is wrong. Two people merging within a minute is
ordinary, and a controller handover happens on every controller merge.
## Questions HQ must answer
- Where does the stale view come from, and is every send recorded so status can show it?
- Should a declaration carry the assignment generation it was composed from, so a host refuses one
older than the last it applied, as it already refuses a declaration it cannot verify?
@@ -1,37 +0,0 @@
# Diagnosis — 204
**2026-10-03, in the controller's code.**
Ruled out: a start-up re-send (a starting controller sends nothing; it asserts the bus, resumes plans,
follows events); a plan sending a recorded declaration (a plan records artifact digests and module
states, and its rollout composes fresh at send time); a cache (every compose reads the store); a path
that does not record its send (push, the cascade and the rollout all record after sending; only the raw
`declare <node> <file>` command did not); the host applying an older sequence (it already refuses a
declaration numbered below the one it kept).
Found, two faults that together produce the evidence:
1. **The sequence number went on at send time, after composing.** Every sending path composed first and
numbered each declaration as it was sent; a multi-machine send composes every machine before sending
any. So a declaration composed *before* an assignment changed and sent *after* a fresher one carried
the higher number — and the host, refusing only lower numbers, applied the older content as the mesh's
newest word. The stale declaration was accepted, so its number was higher, so it was composed earlier
and sent later.
2. **The send record was written on the sender's own context, after the send.** A controller being
replaced in that second has its context cancelled between telling the machine and writing the record;
the machine was told, the record never written, and status showed only the person's earlier send.
The likely sender, consistent with both and with the timing: the outgoing controller's reaction to a
catalogue registration during the build round, which re-sends the machines running the registered
module (one ran on exactly the two machines affected), composed under its hold before the person's
assignments, numbered and sent at 21:29:19–20 as the controller was being replaced. The old container's
log is gone, so the sender is inferred from code and timing; the mechanism is not.
**Owner:** mesh-controller. **Fix direction:** number a declaration before composing it, in every path,
so what was composed earlier is numbered lower whatever order the sends happen in and the host's
existing refusal does its job; record a send on a context that outlives the sender; the raw `declare`
command records too.
Two questions left for HQ: whether status should show the sequence a machine was last sent beside the
digest, and whether a sender's hold should also cover the assignment verbs, which today run between a
hold's compose and its send without waiting.
@@ -1,47 +0,0 @@
---
status: resolved
opened: 2026-10-02
located-in:
- mesh-host
- 00-META/how-we-build.md
fixed-by: mesh-host PR #79 (the host's half; who keeps a machine current is still a decision to take)
amended-design:
---
# 205 — A package resource fails against a stale package database, and nothing keeps it fresh
## What was observed
2026-10-02. The node tools runtime's manifest declares the interpreter as a package. On three machines
the package installed. On the control node the host failed the declaration three times and reported the
machine wrong, stuck:
```
applying "node-tools.interpreter": installing nodejs: pacman exited 1:
error: failed retrieving file 'nodejs-26.5.0-1-x86_64.pkg.tar.zst' from <mirror>: 404
… (every mirror)
error: nodejs: signature from "<packager>" is invalid
```
The machine's package database was from 24 July, ten weeks earlier; the mirrors had long moved on from
the version it asked for, and its keyring was as old. The host asks the package manager to install from
whatever database the machine has and does not refresh it; refreshing on the host's own initiative is
not safe either, because on a rolling distribution a refreshed database plus a single install is a
partial upgrade, which the distribution warns against. The way out was a full system upgrade by the
operator, outside the mesh.
## Why it matters beyond this instance
A `package` resource is one of the host's shapes and every environment module leans on it (design 37).
Its success depends on a machine fact the mesh neither records nor keeps: how old the package database
is. A machine that has not been upgraded in months fails every new package the mesh declares, with an
error that reads as a mirror outage. The mesh says the operator's machine is the mesh's
([ADR 0173](../../02-DECISIONS/0173-the-operators-machine-is-the-meshs-and-a-module-is-what-it-declares.md));
nothing in it says who keeps the machine current enough for its own declarations to apply, or checks.
## Questions HQ must answer
- Is keeping a machine's package database and keyring current a module's job (a `package-manager`
seat holder with a schedule), the host's, or the operator's — and how is it checked?
- Should a `package` resource's failure distinguish "the database is stale" from "the mirror is down",
so the report says what to do?
@@ -1,18 +0,0 @@
# Diagnosis — 205
**2026-10-03.** The host's package step on an Arch machine installs with the package manager against
the database the machine has (mesh-host internal/system/arch.go); it neither refreshes it nor can
safely, since a refresh plus one install is the partial upgrade the distribution warns against. On
the control node the database and keyring were from 24 July; the mirrors no longer served the version
it named, so every mirror answered 404 and the one cached file failed its signature. The host reported
the package manager's output whole, which reads as a mirror outage.
Two owners. The **narrow** half is the host's: classify that failure and say what it is — the database
is stale, the operator must upgrade — rather than relaying forty mirror lines. The **wide** half is a
rule nobody has written: who keeps a machine current enough for its own declarations to apply, and
how that is checked. ADR 0173 makes the machine the mesh's; `00-META/how-we-build.md` says nothing
about its package database. That is a decision (playbook 02), not a code fix: a `package-manager` seat
holder with a schedule, the host, or the operator by rule.
Ruled out: the manifest (`package: nodejs` is correct for the distribution and installed on three
machines the same hour); the network (the mirrors answered, with 404s).
@@ -1,57 +0,0 @@
---
status: resolved
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by: mesh-controller PR #233 (the worker's shape and the plan's order); a credential older than its shape is re-issued by hand, not detected
amended-design:
---
# 206 — A seat's worker changing type strands its holder, and the build that would fix it
## What was observed
2026-10-03, the switch to shared build work ([ADR 0190](../../02-DECISIONS/0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md)).
The controller change made a seat's worker a pull consumer and the build machine pull from it. The
merge's plan put the **build machine** in tier 0 and the controller in tier 1, so the new build machine
rolled first, onto a bus where the running controller had defined the worker as push:
```
mesh-builder: this machine cannot take work from mesh-build-machine: nats: cannot pull subscribe to
push based consumer. The mesh creates that queue and this machine's worker on it, and a build
machine may not create one
```
It restarted every few seconds. The old controller kept asking for tier 1 — the new controller image —
on that worker, and nothing took it. The only thing that would redefine the worker is the controller
that could not be built; there is no verb to run a module's previous build. The way out was a
person running the previous build machine image by hand on the control node until the new controller
had rolled, then removing it.
Two smaller faults surfaced on the way and were each a step of the same handover: the build machine's
credential, issued on 2026-09-28, carried no `claims`, so the new binary's "serve the seat your
credential claims" fell back to the new seat it had no grant for (re-issuing the credential fixed it);
and the build machine's container restarts on its environment file, not on its credential, so the
re-issued credential reached it only because it was already restarting.
## Why it matters beyond this instance
A consumer's type is part of the contract between the controller that defines a worker and the holder
that binds it, and the two are built and rolled by different plans in an order the dependency graph
decides, not the contract. Any future change to a worker's shape — ack wait, filters, type — can strand
every holder the same way, and when the holder is the build machine, the mesh cannot build its way out.
[ADR 0162](../../02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md) orders tiers by
artifacts; it says nothing about what must be *running* before what, and
[issue 201](../201-a-push-recreated-the-controller-behind-the-row-its-successor-wrote/00-report.md)
is the same gap seen from the controller's side.
## Questions HQ must answer
- Does the controller own the worker's shape fully — redefining an existing consumer to the shape it
derives on every start — or does a holder bind whatever shape it finds? Either answer must hold
across a handover where the two are at different versions.
- Should a plan that changes the build machine always roll the controller first, or should the mesh
keep a way to run a module's previous build without a person on the machine?
- A credential issued before claims existed names none: should the mesh re-issue credentials whose
shape is older than what the binary reads, or should every holder treat an unclaimed credential as
the seat its manifest claims?
@@ -1,26 +0,0 @@
# Diagnosis — 206
**2026-10-03.** Three faults in one handover, all the controller's.
1. **The worker's shape is asserted, not reconciled.** The controller creates a seat's worker if
absent (internal/broker, the consumer assertion on start) and leaves an existing one as it is. A
change of shape — here push to pull — therefore never reaches a bus that already has the worker
until somebody deletes it. The new build machine bound a worker whose type its code no longer
speaks.
2. **The plan rolled the holder before the definer.** The merge's plan tiered by artifacts (ADR 0162):
the build machine's image stands on nothing of the controller's, so it came first. For every other
module the order is indifferent; for the holder of the build seat, the controller that defines its
worker must run first, or the build that would bring the controller cannot be taken.
3. **A credential older than its shape.** The build machine's credential was sealed on 2026-09-28,
before credentials carried `claims`; the new binary read none and fell back to the new seat, for
which it had no grant. Re-issuing the credential fixed it; nothing had said it was stale.
Also seen: the build machine's container restarts on its environment file and not on its credential,
so a re-issued credential reaches it only by chance (shared with issue 203's fix direction).
Ruled out: the bus (it refused exactly what the grants and the consumer type said to refuse); the
build machine's new code (it did what its credential told it).
**Fix direction:** the controller reconciles every consumer it owns to the shape it derives, recreating
one whose type changed and saying so; a plan rolls the controller before any holder of the build seat;
a credential whose shape predates what the binary reads is listed and re-issued.
@@ -1,49 +0,0 @@
---
status: resolved
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by: mesh-controller PR #233 (a re-made worker on a history-keeping stream delivers from now on); the two questions on retention and on outcomes for commits already passed stay open for a decision
amended-design:
---
# 207 — A re-made worker replayed every ask the stream kept, and the mesh re-registered its past
## What was observed
2026-10-03, recovering from [issue 206](../206-a-seats-worker-changing-type-strands-the-holder-and-the-build-that-would-fix-it/00-report.md).
The build seat's worker, a push consumer the new controller could not change to pull, was deleted by hand
and the controller restarted. On start it recreated the worker in the shape it derives — pull — and with
the delivery policy a new consumer gets when nothing says otherwise: *every message the stream holds*.
The seat's stream keeps its history. The build machine then took, in order, every build ask since
1 October:
```
a build request arrived for …/mesh-catalog.git (build-1790856308080864567)
[built] baserow from 6afc1160
```
Each outcome was heard and taken in as any build's is. In the minute before the build machine was taken
off the control node to stop it, nine modules were re-registered from the 1 October commit — baserow,
cloudflare-dns, gitlab, grafana, icecast, influxdb, jira, keycloak, letta — the catalogue's recorded
head moved back to that commit, so status listed almost every module as "behind", and the upgrade
policy re-sent the machine running five of them, which replaced their tools containers with the old
images. Recovery: the nine rebuilt from main by hand, the worker re-made by hand as pull delivering only
new asks, the build machine assigned again.
## Why it matters beyond this instance
A work queue that keeps its history is a replay waiting for a consumer that starts from the beginning,
and a new consumer starts there unless told not to. Nothing in the controller's consumer derivation says
where a seat's worker starts, so any re-creation — by hand, or by the reconciliation issue 206's fix
adds — can replay the mesh's whole build history into the catalogue and onto machines. And the taking-in
of an outcome trusts the outcome's commit absolutely: an outcome for a commit older than what the
catalogue holds is registered as if it were news, and rolls out.
## Questions HQ must answer
- Does a seat's work queue keep acknowledged asks at all? If it is a work queue, acknowledged work
should leave it ([design 25](../../03-DESIGN/01-to-be/25-the-bus-on-nats.md) §1 calls it one); if it
keeps history for the record, its worker must be derived to start at new messages, always.
- Should a build outcome for a commit the catalogue has already moved past be recorded and **not**
registered — a build of the past is a fact, not a change — and never sent to machines?
@@ -1,13 +0,0 @@
# Diagnosis — 207
**2026-10-03.** The worker was deleted by hand and the controller, on restart, recreated it with the
server's default delivery policy — every message the stream holds — on a stream that keeps its history.
The build machine then took every ask since 1 October in order, and each outcome was taken in as news:
`takeIn` records the build and registers the manifest it carries, whatever commit it is from. Nothing
in the consumer derivation said where a seat's worker starts; nothing in the taking-in compared the
outcome's commit with what the catalogue already held.
Owner mesh-controller. Fixed in part: a worker re-made by the controller on a history-keeping stream
now delivers from the moment it is made (PR #233). Open for a decision, kept in the report's questions:
whether a seat's work queue should keep acknowledged asks at all, and whether an outcome for a commit
the catalogue has already moved past should be recorded but never registered or rolled out.
@@ -1,42 +0,0 @@
---
status: located
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by:
amended-design:
---
# 208 — A seat's worker is made only when the controller starts, so a holder assigned later finds none
## What was observed
2026-10-03, assigning the first holders of `node-build-agent` ([ADR 0190](../../02-DECISIONS/0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md))
to four machines and pushing them. Every agent came up, authenticated, bound its seat, and restarted:
```
taking build work as a holder of node-build-agent
mesh-builder: this machine cannot take work from node-build-agent: nats: consumer not found. The mesh
creates that queue and this machine's worker on it, and a build machine may not create one
```
The seat's stream existed; its worker did not. The controller makes a seat's worker where it raises
the bus's objects — on start — for the seats that have a holder at that moment, by design: *the queue
before the holder, so work queues until somebody arrives*. Nothing makes the worker when a holder
arrives later: a push asserts the module's own consumers (what it consumes) and not the seat's worker.
The remedy was a controller restart, so the raise ran again with the holder known.
## Why it matters beyond this instance
A seat's first holder is assigned after the controller started in every case but genesis, so every
new role's first holder meets this. The build machine met it on 2026-09-28 the same way ("left the
build machine bound to a consumer nothing had created") and the fix then was to pass the holders at
the raise — which fixed the start, not the arrival. The holder says the right thing and cannot do
anything about it, because a holder may not create its worker (design 25 §3).
## Diagnosis
Owner mesh-controller: the raise runs once (`RaiseSeats` with the holders of the moment); `push` and
`assign` run `EnsureConsumer` only for a module's declared consumption. **Fix direction:** when a
module claiming a seat with `accepts` is assigned, or on every push that composes a holder for such a
seat, ensure the seat's worker as the raise does — the same derivation, the same idempotent assertion.
@@ -1,60 +0,0 @@
---
status: resolved
opened: 2026-10-03
located-in:
- mesh-tools
fixed-by: mesh-tools #32
amended-design:
---
# 209 — A bundle's own copy of the SDK registers into a registry the runtime never reads
## What was observed
2026-10-03, preparing [design 38](../../03-DESIGN/01-to-be/38-building-the-operators-machine.md)
WP4 — the first module whose tools bundle the node's runtime would load beside its own. Before the
manifest changed, a probe on the laptop did what the runtime does: a process holding one copy of
`@novox/mesh-sdk` imported a bundle in another directory that carried its own copy, and the bundle
called `registerModuleTools` as every tools entrypoint does.
```
registrations seen by host after importing bundle: 0
```
The import succeeds, the bundle registers, and the runtime's `collectTools` sees nothing. The
runtime would log the bundle as loaded and answer `tools` for the module with an empty list: silent,
and indistinguishable from a module that serves nothing by choice.
## Why it matters beyond this instance
WP3 found that *a TypeScript bundle must carry its dependencies*, and the toolchain copies them in
so a bundle starts anywhere. Among them is the SDK. The runtime imports a bundle in-process, and
the language resolves a bare import from the importing file's own tree — so every bundle brings a
second SDK into the process: its own registry of tools and its own broker handle. The SDK's registry
is a module-level list, and the runtime reads only the one it imported itself.
Every module from WP4 on is loaded this way. The runtime's tests never met it because their fixture
bundles sit under the runtime's own tree and resolve the same copy. Nothing in design 38 WP1 or WP3
says which SDK a loaded bundle speaks to, and the one-runtime record
([ADR 0175](../../02-DECISIONS/0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md))
assumes without saying that it is the runtime's.
## Diagnosis
Owner mesh-tools (`node-tools`): the runtime imports a bundle's entrypoint and counts the
registrations that appear after it, through the SDK it imported; the bundle's `import "@novox/mesh-sdk/tools"`
resolves to the copy under the bundle's own `node_modules`. **Fix direction:** the runtime resolves
every import of the SDK, from whichever bundle, as if the runtime had written it — one registry,
one broker — and leaves everything else a bundle carries to the bundle's own tree. A test loads a
bundle from a directory holding its own copy of the SDK and asserts its tools are served. The
toolchain keeps copying dependencies in: a bundle launched as a process ([ADR 0188](../../02-DECISIONS/0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md))
needs them, and an imported one is simply not allowed to bring a second SDK.
## Resolution
2026-10-03, mesh-tools #32: the runtime installs a synchronous resolve hook before the first bundle is
imported, sending every import of the SDK, from whichever bundle, to its own copy; a bundle's other
dependencies still resolve from its own tree, and a bundle launched as a process is untouched. The test
loads a bundle from a directory holding its own copy of the SDK and its own dependency, and sees its
tool served with the dependency's answer. Proven live 2026-10-03 by WP4's proof in design 38: the packet filter's bundle, the first loaded beside
the runtime's own, serves its four tools on all four machines.
@@ -1,64 +0,0 @@
---
status: resolved
opened: 2026-10-03
located-in:
- mesh-host
fixed-by: mesh-host #80
amended-design:
---
# 210 — The host re-creates the node's runtime on every reconcile, restarting it every ten minutes
## What was observed
2026-10-03, on the laptop, while proving [design 38](../../03-DESIGN/01-to-be/38-building-the-operators-machine.md)
WP4. The host's log says the same thing at every reconcile, ten minutes apart, since the runtime
module first arrived at 00:04 — 82 times in the day's first thirteen hours:
```
created node-tools.runtime (node-tools): 239 file(s), running as node-tools.service
```
and systemd confirms it: `node-tools.service` is stopped and started at 12:39, 12:49, 12:59, 13:08.
Nothing else in those reconciles changed; every other resource is `kept`. The declaration is the
same one each time — no push happened between the cycles.
## Why it matters beyond this instance
The runtime is every module's tools on the node ([ADR 0175](../../02-DECISIONS/0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md)).
A restart every ten minutes drops every call in flight at that moment, re-subscribes every
membership, and re-imports every bundle; a bundle that is slow to load leaves the node's tools
silent for that long, ten minutes out of every ten. And the log reports it as success, so nothing
in `status` shows a node whose tools blink. The host's rule is that a resource reports `unchanged`
when the machine is as declared; a `process` here never does.
## Diagnosis
Owner mesh-host, `internal/apply/process.go`. The decision is: the digest of the fetched bundle
plus the unit text is `want`; when the host's record of what it last wrote equals `want` and no
`restart-on` resource changed, the process is left alone. Two things were ruled out on the laptop:
the credential the runtime restarts on, which has not been written since the evening before (the
host would also say `updated`, not `created`, for a restart it owed to another resource); and the
unit text, rendered with sorted keys and so stable. What fails is the record. The host's state file
holds an entry for `node-tools.runtime` — applied at the last cycle, with **no `wrote` digest at
all** — while the sibling entry for the runtime's credential carries its digest. The apply sets the
digest on its outcome on every path that installs the daemon, and the loop that records outcomes
copies it into the record for every kind; between the two, a `process` outcome arrives with its
digest empty. **Fix direction:** find where a `process` outcome loses its digest on the way to the
record, and a test that applies the same `process` declaration twice against a recorded store and
asserts the second outcome is `unchanged` with no restart — the test the shape never had. Observed
on the laptop's journal and state; the other three machines' host logs are not readable by the
operator account over SSH, and the behaviour is the host's, not the machine's.
## Resolution
2026-10-03, mesh-host #80. The diagnosis above was right about where and wrong about what: the
record is written every cycle, and the process applier's *unchanged* path returned an outcome that
said nothing about what was written, so the loop recorded it without the digest. The next cycle,
five minutes later, found an empty record and re-created the daemon; the one after was unchanged
and erased the digest again. `created` every other cycle is the ten-minute cadence, and it is why the
state file held the digest on one read and not on the next. The unchanged outcome now carries the
digest forward, as a file's and an archive's do. The test applies one process three times and
asserts the record survives an unchanged apply and no restart is asked; it fails on the code before.
Proven live on the laptop after the host rolled: two reconcile cycles with no `created
node-tools.runtime` line and the runtime's start time unmoved.
@@ -1,47 +0,0 @@
---
status: located
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by:
amended-design:
---
# 211 — A bundle is built in the same tier as the toolchain it is compiled in, against the old toolchain
## What was observed
2026-10-03, merging a runtime change that needed a new SDK release. The SDK was published first;
then the merge of the tools repository planned two tiers, and the first held both the toolchain
images and the runtime's own bundle:
```
> tier 0
mesh-tools asked
node-tools built from 8e30ea9f
```
The bundle was built while its toolchain image was still being built. A TypeScript bundle's
dependencies are copied from the toolchain image it is compiled in
([design 38](../../03-DESIGN/01-to-be/38-building-the-operators-machine.md) WP3), so this one was
compiled and packed against the toolchain as it stood before the merge — carrying the old SDK —
and was recorded as built from the new commit.
## Why it matters beyond this instance
Every bundle compiled by a toolchain depends on the module that publishes that toolchain, and the
planner does not know it: a manifest names its toolchain by `language`, not in `build.on`, so the
dependency is implicit and the tiers are computed without it. Whenever a change touches the
toolchain and a bundle in one merge — or the toolchain's own repository holds a bundle, as this one
does — the bundle is built against the previous toolchain and reported current. Nothing fails; the
bundle simply carries yesterday's dependencies under today's commit.
## Diagnosis
Owner mesh-controller: the planner orders a merge's modules by `build.on`; the builder's toolchain
for a bundle (`ToolchainFor(language)`: the module and artifact it is compiled in) is not part of
that order. **Fix direction:** the planner treats a bundle's toolchain as a `build.on` it did not
have to write — a bundle of language L depends on the module that publishes L's toolchain — so the
bundle lands in the tier after it. A test: a merge touching the toolchain module and a TypeScript
bundle plans the bundle one tier later. Worked around on the day by building the bundle again once
the toolchain was built.
@@ -1,11 +0,0 @@
# 211 — Diagnosis
*2026-10-03.* The planner orders a merge's modules by `inventory.Dependencies`, whose edges come from a
manifest's `build.on`, from what a build recorded it stood on, from the repositories it read, and from
the build machine. A bundle names its toolchain by `language`; the builder takes the toolchain image
(`ToolchainFor(language)`) from what the mesh holds and records nothing of it as stood on. So no edge
ran from a bundle to the module publishing its toolchain, and a merge moving both (mesh-tools: the
images and node-tools) tiered them together. **Fix (mesh-controller, branch
`fix/issue-211-a-bundle-stands-on-its-toolchain`, commit c72f6ca):** `dependenciesOf` adds a `stands-on`
edge from every bundle artifact to its toolchain's module, read from the manifest. Tested: TypeScript
bundle → mesh-tools, Go bundle → mesh-tools-go, image → none; a merge moving both plans two tiers.
@@ -1,44 +0,0 @@
---
status: located
opened: 2026-10-03
located-in:
- mesh-tools
- mesh-controller
fixed-by:
amended-design:
---
# 212 — A toolchain rebuild keeps the SDK it cached, and every bundle built on it carries the old one
## What was observed
2026-10-03, rolling out [ADR 0193](../../02-DECISIONS/0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md).
SDK 0.1.5 was published, then the TypeScript toolchain image was rebuilt, then node-tools and six
bundles were rebuilt on it and pushed. On every machine the bundles carried SDK 0.1.3:
```
bundle SDK: "version": "0.1.3"
```
The launched packet filter and intrusion prevention, which register their seat first, then served
their seat's verbs as their own tools, and `node-packet-filter.*` and `node-intrusion-prevention.*`
stopped answering on three machines until fixed.
## Why it matters beyond this instance
The toolchain image installs its dependencies from a `package.json` that names the SDK by range.
The file did not change, so the image build reused its cached install layer, and "rebuilt after
the release" did not mean "carries the release". Every TypeScript bundle copies its dependencies
from that image ([design 38](../../03-DESIGN/01-to-be/38-building-the-operators-machine.md) WP3),
so an SDK release reaches no bundle until something else happens to change that file. Nothing
reports it: the build succeeds and records the new commit.
## Diagnosis
Owners mesh-tools (the image's recipe) and mesh-controller (the builder that runs it). **Worked
around** on the day by requiring `^0.1.5` in node-tools' `package.json` (mesh-tools #37), which
changes the layer. **Fix direction:** a toolchain build must not trust a cached dependency install —
install from a lockfile that a release updates, or build the install layer without the cache — and
the build should record which SDK version the image carries, so a bundle's record says what it was
compiled against. A check: after an SDK release and a toolchain rebuild, a bundle built on it reports
the released version.
@@ -1,52 +0,0 @@
---
status: open
opened: 2026-10-03
located-in:
- mesh-controller
- mesh-catalog
fixed-by:
amended-design:
---
# 213 — The controller is a Go program, and it still runs in a container
## What was observed
2026-10-03. The controller — the program that holds the `mesh-controller` seat and answers its
verbs (`status`, `nodes`, `push`, `assign` …), composes every machine's declaration and plans the
builds — runs on its machine as a container built from an image:
```
mesh-controller Up … (docker ps on the machine that runs it)
```
It is written in Go and compiles to one static binary, as the node host does. The host is delivered
as a bundle and run as a process; the node's tool runtime now is too
([ADR 0193](../../02-DECISIONS/0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md)).
The controller is the one piece of the mesh's own Go code still shipped as an image.
## Why it matters beyond this instance
[ADR 0188](../../02-DECISIONS/0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md)
§1 says a module's own code is bundles, never an image, and §3 that a bundle that is a service is a
`process` the host runs. The controller breaks the rule it is the mechanism of: the registration
gate that will refuse an image of a module's own code has to exempt the controller, or refuse it.
It also costs what an image costs — a container runtime on its machine as a hard requirement,
eight mounts standing in for files a process would simply read, an image rebuild for a binary
change — and
every restart of it is a container recreation, which is how the controller restarts in the middle
of a plan today.
## What a fix has to settle
- The controller's module declares a Go bundle (`system`, `binary`) and a `process` running it
(`./<binary>`, mesh-host #81), with its credential and store connection as files and words, not
container mounts and a container network name.
- What the container gives it now that a process would not: it runs on the host's network already,
as an unprivileged user (65534), with eight mounts. Each mount named and replaced by a path, and
the user by an account the host declares.
- The handover: the controller restarting itself as a process, on the one machine that runs it,
without a window where nothing answers the mesh's verbs.
Located only by owner; the move is a change of the controller's module and its deployment, not of
its code.
@@ -1,36 +0,0 @@
---
status: located
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by:
amended-design:
---
# 214 — A plan loses track of the controller it rebuilds, and waits on it for ever
## What was observed
2026-10-03, a merge to the controller's repository planned three tiers: the controller itself, then
the build agent, then the route proxy. The controller's build finished and was recorded at 19:24:
```
mesh-controller built 2ebbb799 g14 2026-10-03 19:24
```
Twenty-seven minutes later the plan still read `tier 1 of 3 … mesh-controller asked`, and every later
plan waited behind it. It was stopped by hand; the later tiers were never asked.
## Why it matters beyond this instance
The controller rebuilding itself is the one plan whose first tier replaces the process that runs the
plan. The new controller starts with the plan's state as stored, and the outcome of the build that
produced it arrived to the old one, or between the two. Every merge to the controller's own repository
can end this way, and each one blocks every plan after it until somebody notices.
## Diagnosis
Owner mesh-controller (the planner). **Fix direction:** on start, and whenever a plan waits on a
build, the plan settles an `asked` build against the build records — a build recorded as built from
the plan's commit is that tier's outcome — so a plan resumes after the controller replaced itself.
A test: a plan whose build outcome was recorded while no controller followed it resumes on start.
@@ -1,10 +0,0 @@
# 214 — Diagnosis
*2026-10-03.* A plan learns a tier's outcome only through `planBuilt`, called when a controller takes
in a build result off the bus. A merge to the controller's repository replaces the controller in tier
0; the build that produced the new controller was recorded, but the plan state the new controller
loaded still read `asked`, and no path ever revisited it. **Fix (branch
`fix/issue-214-a-plan-settles-from-the-build-records`, commit d86baeb):** `advanceOnce` settles every
still-asked module from its build records — a build recorded after the ask is that ask's outcome,
built or failed — on every advance and on the 30-second ticker. Tested with a pure helper. Live
proof: the next merge to mesh-controller passes tier 0 on its own.
@@ -1,35 +0,0 @@
---
status: located
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by:
amended-design:
---
# 215 — A module built once at a commit stops following its branch, and merges leave it out
## What was observed
2026-10-03, a catalogue merge changed seven modules. Its plan built six; `unifi` was not in it,
though its change was in the same commit. A second merge the same evening left it out again. Its
build records name a commit where every other module names a branch:
```
unifi built 9c97a8a1 … http://…/mesh-catalog.git at 9c97a8a
```
Built by hand from `main` it was registered again and the change reached its machine.
## Why it matters beyond this instance
A module that was once built at a commit — to pin it during a fix, or by a build asked with a `ref`
— silently stops following its branch: merges plan without it, `status` does not say it is behind its
branch, and nothing says it is pinned. The operator learns it when a change does not arrive.
## Diagnosis
Owner mesh-controller. **Fix direction:** a build asked at a commit does not change the branch a
module follows; or, if pinning is meant, the pin is said — in `module list`, in `status`, and by a
merge's plan naming the module it leaves out and why. A test: building a module at a commit and then
merging a change to it plans it.
@@ -1,10 +0,0 @@
# 215 — Diagnosis
*2026-10-03.* `takeIn` registers a build's `Ref` as the branch the module follows. unifi was once
built with `ref=9c97a8a`, which became its followed ref. `sourceIs` matches a merge only to modules
whose ref is empty or the merged base — so every merge into main left unifi out — and `askTier`
re-asks `Source.Ref`, so every plan that rebuilt unifi built the same old commit again (its build
records all read "at 9c97a8a"). **Fix (branch `fix/issue-215-a-commit-is-never-a-branch-to-follow`,
commit 6784efa):** registration keeps the followed branch when a build names a commit; matching and
re-asking read a recorded commit as the default branch, healing existing records; a merge names the
modules of its repository it leaves out. Store-backed test fails without the fix.
@@ -1,30 +0,0 @@
---
status: located
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by:
amended-design:
---
# 216 — A tools bundle nothing says to load is built, recorded, and never delivered
## What was observed
2026-10-03, seven modules moved their tools from a container to a bundle. Each built, each build was
recorded, and the machines that run them reported current — with none of the bundles on them. The
composer delivers a bundle only when the runtime loads something from it, and that is derived from
the module's `tools` list when the artifact names no `loads`; these modules had neither.
## Why it matters beyond this instance
Every step reported success: the build, the registration, the push, the machine's apply. The tools
were simply absent, and the old container was gone. A rule the composer applies silently is a rule
nobody learns until the tools are missing.
## Diagnosis
Owner mesh-controller (the catalogue's registration check). **Fix direction:** a bundle that nothing
loads, runs or unpacks — no `loads`, no `tools` list on its module, no resource naming it — is refused
at registration, naming the field that would deliver it. A test: such a manifest is refused; adding
`loads` admits it.
@@ -1,9 +0,0 @@
# 216 — Diagnosis
*2026-10-03.* The composer delivers a bundle as an archive only when its `Loads` is non-empty, and
`Loads` derives from the artifact's `loads` or, failing that, from the module's `tools` list. The
seven modules had neither, so their bundles were recorded and never composed into any declaration;
nothing checked it. **Fix (branch `fix/issue-216-a-bundle-nothing-delivers-is-refused`, commit
cf2bb3b):** registration refuses a bundle that nothing loads, runs or unpacks — no `loads`, no `tools`,
no resource naming it, and not the runtime — naming the field that would deliver it. The current
catalogue passes the check.
@@ -1,52 +0,0 @@
---
status: located
opened: 2026-10-03
located-in:
- mesh-tools
fixed-by:
amended-design:
---
# 217 — A refused announcement took down every container's runtime, and the console with it
## What was observed
2026-10-03, rolling out [ADR 0197](../../02-DECISIONS/0197-every-tool-announces-itself-on-the-bus-in-the-nats-services-protocol.md).
The tool runtimes announced themselves by subscribing `$SRV.<verb>.>`; the grants composed for them
allowed only `$SRV.<verb>` and the service's own name and instance. The bus refused the wildcard:
```
Subscription Violation - User "novox.gitea", Subject "$SRV.PING.>"
NatsError: 'Permissions Violation for Subscription to "$SRV.PING.>"'
```
The TypeScript runtime the per-module containers run treats a refused subscription as fatal, so on
the one machine that had received the new images nine containers crash-looped — gitea's runtime,
postgres's, the catalogue, the vault, mongodb, mssql, keycloak, mailu, nextcloud. Gitea's tools went
with them, which closed the usual path for merging the fix.
Then the console stopped answering the controller's verbs, though the controller held every
subscription: the console builds the list that tells a seat's verb from a module's tool by asking the
controller *and* the catalogue, and gives up on both when the catalogue does not answer — so
`mesh-controller.status` was asked of a module subject nobody serves.
## Why it matters beyond this instance
Two properties, each worse than the mistake that exposed it:
- **A runtime dies for an optional subscription.** Announcing is discovery; serving tools and running
provisioners is the work. A refusal of the first should never stop the second.
- **The console's view of the mesh's own verbs depended on a module.** The controller's verbs are how
the operator repairs the mesh; they must not become unreachable because the catalogue is down.
## Diagnosis
Owner mesh-tools. The wildcard is fixed on branch `fix/announce-only-what-the-grants-allow`: both
runtimes subscribe exactly what the grants allow. The console that discovers from what announces itself
(ADR 0195, 0197, on main) asks the bus and the controller, not the catalogue, which removes the second
property once it is deployed. **Still to do:** the TypeScript runtime treats a refused announcement
subscription as a logged warning, not as fatal; a test against a bus with real grants proves the
announcement subscriptions are allowed for every principal kind.
Recovered on the day without the forge's API: the toolchain images built by the controller straight
from the fix branch, every module image rebuilt on them, and the machine pushed.
@@ -1,65 +0,0 @@
---
status: located
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by: mesh-controller#248
amended-design:
---
# 218 — A seat held once for the mesh is answered by a module on a machine that does not hold it
## What was observed
2026-10-03. Asked which databases the mesh's store holds, `mesh-store.databases` answered from the
postgres on one machine with that machine's application databases; the controller's own database
lives on the postgres of the other machine, which the controller's records name as the seat's one
holder:
```
mesh-store scope: mesh delivers: postgres-database holders: [ {node: <the control machine>, module: postgres} ]
```
The discovery console, which reads what answers on the bus ([ADR 0197](../../02-DECISIONS/0197-every-tool-announces-itself-on-the-bus-in-the-nats-services-protocol.md)),
shows the same seat announced from **both** machines running postgres.
## Why it matters beyond this instance
A seat held once for the mesh promises one answerer: the role's holder. A module that implements a
seat's verbs on every machine it runs on, and is let serve them on each, turns "the mesh's store" into
"whichever postgres replied first" — a read against the wrong database that looks like a right one,
and a write would be worse. Every mesh-scoped seat whose implementing module runs on more than one
machine has this shape.
## Where to look
Whether the runtime serves a seat's verbs where its module merely *claims* the seat rather than where
the mesh made it the holder ([ADR 0159](../../02-DECISIONS/0159-a-tool-call-names-the-machine-and-a-holder-serves-its-seats-verbs.md),
[ADR 0160](../../02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)):
the membership the controller issues each assignment, and what the runtime admits from it. A check:
a mesh-scoped seat's verbs are served by exactly the holder the records name, on every machine.
## Root cause
The controller composed each assignment's held seats from what its module *claims*, once per module
and not once per machine. Every machine running postgres was therefore given the store seat's grants
and issued its subjects, and each runtime served the seat's verbs because it serves what it is issued
([ADR 0160](../../02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)).
The runtime behaved as designed. The fault was in what it was issued.
The seat's verbs are not the module's tools. The store's `databases` and `query` are a separate
implementation registered under the seat's name ([ADR 0159](../../02-DECISIONS/0159-a-tool-call-names-the-machine-and-a-holder-serves-its-seats-verbs.md)).
Only that implementation should be withdrawn where the module does not hold the seat. postgres's own
tools stay served on every machine it runs on.
## Fix
The controller now reads the recorded seat holdings when it composes grants and memberships. A seat
held once for the mesh is issued only to the machine and module the records name as its holder. A
seat held once per machine, and a mesh seat with no holder on record, are issued as before. Grants
and memberships come from the same list, so they cannot disagree.
**How it is checked.** A controller test asserts that a claimant on another machine keeps its node
seats and loses the recorded mesh seat. Live, the discovery console's overview must show each
mesh-scoped seat announced from exactly the holder the records name. Status moves to `resolved` once
that holds after the fix is rolled out.