Compare commits

..
41 Commits
Author SHA1 Message Date
jochen baf82b1cdb Plan and designs after ADR 0193, 0195 and 0198: bundles are launched and the runtime is their bus; the manager's daemon is a long-running bundle; the console's five tools
The dated note on ADR 0183 now rests on ADR 0193 and 0198 rather than on a bundle having no way to
call: the manager starts every exchange by the operator's direction, through mesh/ask. To-be 40's
WP4 no longer waits on a record — ADR 0198 is it — and the live proofs count the console's five
tools (ADR 0195).
2026-10-03 23:41:38 +02:00
jochen 88fd7d9739 Design 36 §4: the console is registered in the exclusive managed tool-server file, because the managed-settings key refuses a non-https URL 2026-10-03 23:41:10 +02:00
jochen ce0dc7b55b To-be 40 revised for the tools refactor; the manager starts every exchange (ADR 0183 dated note, designs 36 and 39)
Design 38's WP1-WP4b ran: the node's tool runtime is live on all four machines as the operator
account, tools are bundles given only their declared words, and a bundle has no bus credential.
So the wait on design 38 WP3 is over, the agent module calls nothing and the manager starts every
exchange (key, hand-over, waiting login, reconcile), and the manager's daemon now waits on WP4c's
record instead. Accounts are stated on all four, sudo -n works for each, the agent is installed on
all four; the plan's WP0 shrinks and WP2 gets a configuration-only live proof before any licence.
2026-10-03 23:41:10 +02:00
jochen 7d18b0c1d0 To-be 40: building the operator's agent and its licence manager as work packages
Designs 36 and 39 say what is built; this says in which order and what proves each step, in the
shape to-be 38 gave the operator's machine. Seven packages: the operator states the facts (accounts,
roles, licences); the console provides its endpoint; the licence manager and the agent module are
built and unit-tested in parallel; the manager goes live on the control node; the agent on one
workstation, with the switch and the predecessor's files removed as the proof of the whole; then the
rest of the nodes and the retirement of the two catalogue modules built on the old placement. The
live proofs wait for to-be 38's WP3, because both modules' tools run in the node's tool runtime
(ADR 0175) and a per-module tool container would rebuild what that record retires.
2026-10-03 23:41:10 +02:00
mesh-admin 895c2afad1 Merge pull request 'Issue 218: a mesh seat answered by a non-holder (located, fixed in mesh-controller#248)' (#339) from issues/218-a-mesh-seat-answered-by-a-non-holder into main 2026-10-03 21:32:24 +00:00
jochen 4b8c5e3b11 Issue 218 located: the controller issued a mesh seat to every claimant; fixed in mesh-controller#248 2026-10-03 23:31:13 +02:00
mesh-admin d362155401 Merge pull request 'Issue 218: a mesh seat is answered by a module on a machine that does not hold it' (#338) from issues/218-a-mesh-seat-answered-by-a-non-holder into main 2026-10-03 21:25:27 +00:00
jochen 557760e537 Issue 218: a mesh seat is answered by a module on a machine that does not hold it 2026-10-03 23:25:17 +02:00
mesh-admin 6b1ebd1d4a Merge pull request 'Issue 217: a refused announcement took down every container's runtime, and the console with it' (#336) from issues/217-a-refused-announcement into main 2026-10-03 21:20:06 +00:00
mesh-admin fa73a17ceb Merge pull request 'Issues 211, 214, 215, 216 diagnosed' (#335) from issues/211-214-216-diagnosed into main 2026-10-03 21:19:03 +00:00
jochen 08cea893bd Issue 217: a refused announcement took down every container's runtime, and the console with it 2026-10-03 22:47:28 +02:00
jochen 04625d3e35 Issues 211, 214, 215, 216 diagnosed: root causes and the branches that fix them 2026-10-03 22:28:40 +02:00
mesh-admin d7bb24b181 Merge pull request 'ADR 0198: a module's long-running code is launched by the node's runtime and reaches the bus through it' (#334) from decision/0198-a-modules-long-running-code-is-launched-by-the-runtime into main 2026-10-03 20:21:04 +00:00
jochen 23d6e30b8a ADR 0198: a module's long-running code is launched by the node's runtime and reaches the bus through it; research 022; design 38 WP4c plan 2026-10-03 22:20:53 +02:00
mesh-admin 2043e90f35 Merge pull request 'Issues 214–216: a plan losing its own controller, a module pinned silently, a bundle never delivered' (#333) from issues/214-216-found-building-0192-0193 into main 2026-10-03 20:15:30 +00:00
jochen c9418af42e Issues 214–216: a plan losing its own controller, a module pinned silently, a bundle never delivered 2026-10-03 22:15:19 +02:00
mesh-admin 9c3e77999e Merge pull request 'Design 38 WP4d: every served bundle launched and node-tools in Go, proven live on all four machines' (#332) from design/38-wp4d-proven into main 2026-10-03 20:07:30 +00:00
jochen 6943843fff Design 38 WP4d: every served bundle launched and node-tools in Go, proven live on all four machines 2026-10-03 22:07:20 +02:00
mesh-admin 34325d3566 Merge pull request 'ADR 0197: every tool announces itself on the bus, in the NATS services protocol' (#331) from decision/0196-every-tool-announces-itself into main 2026-10-03 20:04:12 +00:00
jochen fa9e94d863 Renumber to ADR 0197: 0196 landed first on main 2026-10-03 22:03:58 +02:00
jochen 1b34821aa0 ADR 0196: every tool announces itself on the bus in the NATS services protocol 2026-10-03 22:03:34 +02:00
jschoubben 50ddaf8408 Merge pull request 'ADR 0196: a node asks the mesh's resolver first, and a public one only when it is silent' (#330) from decision/0196-nodes-ask-the-mesh-resolver-with-a-public-fallback into main 2026-10-03 20:01:05 +00:00
jschoubben f5d518d256 ADR 0196: a node asks the mesh's resolver first, and a public one only when it is silent
ADR 0194 rejected sending every query to the mesh's resolver because a node with its tunnel down
would resolve nothing; a public resolver listed second answers exactly then. That drops the
systemd-resolved stub and the runtime's dns: containers copy the machine's resolvers. Narrows 0194;
amends connectivity §2.
2026-10-03 21:56:33 +02:00
jschoubben 577ddf0089 Merge pull request 'ADR 0194: the mesh has one resolver, and every node asks it for the mesh's names' (#326) from decision/0194-the-mesh-has-one-resolver into main 2026-10-03 19:51:04 +00:00
jschoubben 13e28e6873 ADR 0194: the no-copies check allows each node's loopback stub 2026-10-03 21:51:03 +02:00
jschoubben 6b6ff76a19 ADR 0194: why every node needs a stub, and the systemd-resolved module that provides it 2026-10-03 21:50:33 +02:00
mesh-admin 8d5e6ef76f Merge pull request 'ADR 0195: the mesh's tools are found by address, not announced whole; research 021; to-be 34 §3a' (#329) from decision/0195-the-mesh-tools-are-found-by-address into main 2026-10-03 19:49:37 +00:00
jochen 77813f4613 ADR 0195: the mesh's tools are found by address, not announced whole; research 021; to-be 34 §3a 2026-10-03 21:49:26 +02:00
mesh-admin 4ee8e3905d Merge pull request 'Issue 213: the controller is a Go program and still runs in a container' (#328) from issues/213-the-controller-runs-in-a-container into main 2026-10-03 19:42:31 +00:00
jochen 216faec69e Issue 213: what the container gives it, measured 2026-10-03 21:42:22 +02:00
jochen e36b1a9e9c Issue 213: the controller is a Go program and still runs in a container 2026-10-03 21:42:11 +02:00
mesh-admin 5886969c75 Merge pull request 'Issue 212: a toolchain rebuild keeps the SDK it cached' (#327) from issues/212-a-toolchain-keeps-a-cached-sdk into main 2026-10-03 19:24:49 +00:00
jochen 5fa43ff755 Issue 212: a toolchain rebuild keeps the SDK it cached 2026-10-03 21:24:39 +02:00
jschoubben 1de4a5f25e ADR 0194: the mesh has one resolver, and every node asks it for the mesh's names
Every resolution fault found on 2026-10-03 was a per-node copy disagreeing with the truth: a hosts
file read once, an operator's old line beside the mesh's, a node's resolver lent to a LAN. Every
tunnel already converges on one node. Retires node-dns-resolver for a mesh-scoped mesh-resolver;
nodes route only the mesh's suffix to it. Narrows 0121; amends connectivity §2 and the seats.
2026-10-03 21:21:49 +02:00
jschoubben 8a1fa37dce Merge pull request 'ADR 0191: domains are a node's — the resolver holds each node's internal domain, nothing else' (#323) from decision/0191-names-by-origin into main 2026-10-03 19:12:26 +00:00
mesh-admin 44eb13acd0 Merge pull request 'ADR 0193: every bundle the runtime serves is launched, and the runtime knows no language; design 38 WP4d' (#325) from decision/0193-every-served-bundle-is-launched into main 2026-10-03 19:05:12 +00:00
jochen 6d53f9168a ADR 0193: every bundle the runtime serves is launched, and the runtime knows no language; design 38 WP4d 2026-10-03 21:05:03 +02:00
mesh-admin 6a1fc71a4c Merge pull request 'Design 38 WP4b: four tools-only modules proven live from the runtime; two delivery traps' (#324) from design/38-wp4b-seven-proven into main 2026-10-03 14:13:40 +00:00
jochen fb76fb7256 Design 38 WP4b: four tools-only modules proven live from the runtime; two delivery traps 2026-10-03 16:13:32 +02:00
jschoubben 2344bfb69b ADR 0191: domains are a node's — one internal, one or more public; the roster is the machines 2026-10-03 16:10:38 +02:00
jschoubben ca8a865e73 ADR 0191: the mesh's names are known by where they were composed, not by their suffix
A progressive insight: the rule and its check were stated as a suffix test; the mesh composes both
names of a route and publishes its internal one. The decision is unchanged.
2026-10-03 15:39:07 +02:00
32 changed files with 1239 additions and 49 deletions
@@ -0,0 +1,48 @@
---
status: graduated
initiated: 2026-10-03
touches: [the console, 03-DESIGN/01-to-be/34-the-console.md, the tool runtime, seats, assignments]
became: [02-DECISIONS/0195-the-meshs-tools-are-found-by-address-not-announced-whole.md, 03-DESIGN/01-to-be/34-the-console.md]
---
# 021 — Finding a tool in the mesh
## What was investigated
How an agent finds the one tool it needs among everything the mesh answers, and how a call names
exactly what it asks — a role the mesh holds once, a role every machine holds, or one assignment of a
module on one machine — rather than receiving the whole catalogue and a name that can mean several
things.
## Why
The operator's observation on 2026-10-03: *Claude should not see all tools at once; they should be
discoverable — and `postgres.list_databases` is wrong, asking one machine's postgres is not asking
another's.* Measured the same day from the console's own answer:
| | |
|---|---|
| tools announced to every session at its start | 228, in 110 KB |
| names (module or seat prefixes) | 43 |
| node seats' verbs, which require `node` | 22 |
| modules with tools on more than one machine | 4 — fail2ban, nftables (every machine), postgres, mssql (two each) |
| modules reported "not answering", most with no tools and several retired | 47 |
The two stateful modules on two machines are listed **once**, with `node` optional and *whichever
answers* when it is left out — though their two instances hold different databases. Design 34 §3 says
such a module is listed once per machine; the live console does not do that. The list is taken once
per session, so a tool that arrives later is invisible until the client reconnects. And only Claude
Code's own deferral of long tool lists keeps the 228 from the model's context; another MCP client
would receive them whole.
## Options
1. **Keep the flat list; rely on the client to defer it.** Rejected: a property of one client, and it
leaves the ambiguity and the stale list.
2. **One flat tool per assignment** (`ace_postgres_list_databases`). Removes the ambiguity, multiplies
the list, and runs into the API's tool-name limit (letters, digits, `_`, `-`, 64 characters).
3. **A small fixed set of tools that walk the mesh's own structure**, with the full address as an
argument: the mesh's seats; a machine's node seats and assignments; a search; a description; a
call. Chosen — see [ADR 0195](../../02-DECISIONS/0195-the-meshs-tools-are-found-by-address-not-announced-whole.md).
4. **MCP resources or prompts for discovery.** Clients support them unevenly, and an agent acts
through tools; a resource it cannot be relied on to read is not a discovery path.
@@ -0,0 +1,41 @@
---
status: graduated
initiated: 2026-10-03
touches: [the tool runtime, the per-module containers, the SDK, the bus grants, 03-DESIGN/01-to-be/38-building-the-operators-machine.md]
became: [02-DECISIONS/0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md, 03-DESIGN/01-to-be/38-building-the-operators-machine.md]
---
# 022 — Where a module's long-running code runs
## What was investigated
Twenty-three modules still run their own code in a container built on the runtime's image. Their tools
can move as bundles ([ADR 0192](../../02-DECISIONS/0192-a-tools-bundle-declares-what-it-is-given-and-the-runtime-hands-it-to-that-bundle-alone.md),
[ADR 0193](../../02-DECISIONS/0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md));
the rest of what those containers run cannot yet. This asks where that code goes and how it reaches
what its container handed it.
## What that code is, measured 2026-10-03
| | modules |
|---|---|
| subscribes to events on the bus | audit-logger (everything), mesh-catalog (two seat events), mesh-vault, records (`gitea.pull.merged`), and postgres, mongodb, mssql, redis, mosquitto logging their own lifecycle |
| provisioners: read the grants the mesh delivered as files, act on the backend, emit | 12 |
| a run-once preparation step | mesh-catalog |
| a command-line client of the backend | psql, mosquitto_ctrl, git (packages on every machine's system); mongosh, sqlcmd (not in its repositories) |
| a service reached by a container name | icecast, mailu-admin, minio, mongodb-server, mssql |
| a main of its own | anthropic-consumer, openai-consumer, route-adapter |
A provisioner needs nothing a launched bundle lacks: files named by its words, its backend, and an emit
that already travels through the runtime. The one thing missing is **a subscription** — events
delivered to the module's code, acknowledged when it has handled them.
## Options
1. **The runtime launches it and is its bus**: the stdio channel gains a subscription; the runtime
binds the module's durable consumer and delivers each event to the child, acknowledging when the
child answers. One bus connection per machine; any language. Chosen.
2. **A process per module with its own bus client and credential.** Every language's SDK would carry
a transport and every module a credential on disk — what ADR 0188 rejected for tools, for the same
reasons.
3. **Keep the containers for this code.** Leaves ADR 0188's rule broken for 23 modules indefinitely.
@@ -9,6 +9,12 @@ extends: 0110-a-seat-is-a-module-assignment-from-a-closed-set.md
# 121. A system seat is named for its scope, and a module may define its own
> **Narrowed, not replaced — 2026-10-03.** *"`the-dns-port` → `node-dns-resolver`"* no longer holds:
> the serving role moves to mesh scope as `mesh-resolver`, one per mesh, and `node-dns-resolver` is
> retired ([ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)). The
> distinction this record kept — serving and asking are two roles, two seats — stands, and
> `node-resolver-config` is unchanged.
> **The mechanism changed — 2026-10-02, by [ADR 0190](0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md).** The naming rule stands. The build role this record made mesh-scoped — *the mesh's single build machine* — is node-scoped now: `node-build-agent`, one holder per machine, every holder taking from one work queue.
## Context
@@ -152,18 +152,19 @@ the node is bound to, and refuses with a notification otherwise.
| An unservable binding refuses rather than lends | a manager test: a worker bound to a dead licence is answered with a refusal, never another licence's token |
| A switch through the console changes the token on the node and nothing in the answer is a token | a live check on one workstation |
> **The mechanism changed — 2026-10-03, by [ADR 0192](0192-a-tools-bundle-declares-what-it-is-given-and-the-runtime-hands-it-to-that-bundle-alone.md).**
> **The mechanism changed — 2026-10-03, by [ADR 0193](0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md)
> and [ADR 0198](0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md).**
> What stands: one manager holding the seat, one rotation source, a token sealed to the receiving
> module's key on request/reply and never an event, the agent module alone writing what the agent
> reads, the identity guard, the host knowing nothing. What moved: the agent module's code is now a
> tools bundle the node's runtime serves ([ADR 0175](0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md),
> [ADR 0188](0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md)),
> and a bundle has no bus credential of its own and answers calls rather than making them (ADR 0192's
> consequences). So **the manager starts every exchange**: it asks each bound node's agent module for
> its public key, hands it a token, asks it for a login waiting to be adopted, and reconciles every node
> on a schedule — which is what "the agent module asks the seat for its current token" and "offers the
> grant to the manager" in the decision above now mean in practice. The manager's own process needs a
> bus credential to make those calls, which is the question design 38's WP4c leaves to a record.
> reads, the identity guard, the host knowing nothing. What moved: both modules' code is bundles the
> node's runtime launches over stdio and is the bus for — `mesh/ask` for a call made on the module's
> behalf, `mesh/publish` and `mesh/subscribe` beside it — so neither holds a bus credential of its own.
> The manager's refresh and visits are a long-running bundle the control node's runtime launches. And,
> by the operator's direction, **the manager starts every exchange**: it asks each bound node's agent
> module for its public key, hands it a token, asks it for a login waiting to be adopted, and reconciles
> every node on a schedule — which is what "the agent module asks the seat for its current token" and
> "offers the grant to the manager" in the decision above now mean in practice. The agent module could
> ask through its runtime; it does not need to.
## References
@@ -87,6 +87,10 @@ that lets one bundle in that language be built, delivered and answer one tool on
Anything beyond that is added when a module needs it. A skeleton that is not proven by one bundle
answering is not a skeleton; it is a promise.
> **The mechanism changed — 2026-10-03, by ADR 0193.** §2's allowance that a TypeScript bundle may
> be imported into the runtime's own process is withdrawn: every served bundle is launched, and the
> build makes each served entrypoint executable. The rest of §2 stands.
## Consequences
- The runtime gains a launcher beside its loader. The loader, the memberships, the seats and the
@@ -11,6 +11,16 @@ supersedes-in-part:
# 191. The mesh's resolver holds only the mesh's own names; a public name resolves publicly
> **Progressive insight — 2026-10-03.** The first implementation told the mesh's names from public
> ones by their spelling — a name ending in the mesh suffix — and this record said so: the Decision
> read *"only names under its own suffix"*, and the roster check *"every name the roster carries ends
> in the mesh suffix"*. The mesh needs no such test, nor any per-route name: domains are a node's. A
> node has **one internal domain**, `<node>.internal`, and every route on it is a name under that domain
> (ADR 0151), answered by one wildcard per node; a node has **one or more public domains**, which public
> DNS answers. So the mesh's resolver holds the nodes' internal domains and nothing else, and the roster
> carries the machines and no routed name. Both sentences now say that; what was decided — a public
> name is never given a private answer — is unchanged.
## Context
**[ADR 0066](0066-public-routing-is-name-agnostic.md) published every routed name into internal
@@ -63,7 +73,8 @@ every public name the mesh serves, is forwarded and resolves publicly. Chosen.
## Decision
**The mesh publishes into internal resolution only names under its own suffix.** A machine's name,
**The mesh's resolver holds each node's internal domain and nothing else** — `<node>.internal` and
everything under it, at that node's private address. A machine's name,
and through it every `<label>.<node>.internal`, resolve to that machine's private address. **A public
name is never given a private answer by the mesh**: it resolves through public DNS to the public
address, from members and non-members alike.
@@ -93,8 +104,8 @@ reachability — the lab — certifies its internal names and has no public name
**How each is checked:**
- **The roster:** the controller's catalogue tests assert that every name the roster carries ends in
the mesh suffix — a routed public name in it fails the build.
- **The roster:** the controller's tests assert that the roster names the machines and nothing
else — a routed name in it, public or internal, fails the build.
- **On a machine:** asking the machine's resolver for a public name the mesh serves returns the
public address, and asking it for that route's internal name returns the private one. Asked from a
non-member on a LAN the resolver answers, the first must hold as well.
@@ -88,6 +88,10 @@ tools answering from the runtime, and the registration gate of
[to-be 38](../03-DESIGN/01-to-be/38-building-the-operators-machine.md) WP2 then refuses the
container shape for every module, as ADR 0188 already provides.
> **The mechanism changed — 2026-10-03, by ADR 0193.** Decision 3's imported path — the environment
> handed to an imported bundle's contributor — has nothing left to do: every served bundle is
> launched, and a launched bundle's environment is its own. The decision stands.
## Consequences
- The manifest gains one field on one artifact kind; the composer gains one more thing to resolve
@@ -0,0 +1,88 @@
---
topic: what runs on it
status: accepted
date: 2026-10-03
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md
---
# 193. Every bundle the runtime serves is launched, and the runtime knows no language
## Context
[ADR 0188](0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md)
§2 made a served bundle a process that speaks MCP over stdio, and kept one exception: a TypeScript
bundle may be imported into the runtime's own process, "a shortcut over the same contract". Every
module's tools today take the shortcut, and it is where the day's defects came from:
- [Issue 209](../04-ISSUES/209-a-bundles-own-sdk-copy-registers-into-a-registry-the-runtime-never-reads/00-report.md):
an imported bundle's own copy of the SDK registered into a registry the runtime never read; fixed
by a resolve hook that redirects every bundle's SDK import to the runtime's copy.
- [ADR 0192](0192-a-tools-bundle-declares-what-it-is-given-and-the-runtime-hands-it-to-that-bundle-alone.md):
thirty modules' environments in one process had to be kept apart by an SDK change and runtime
bookkeeping, where a process of its own has an environment of its own by construction.
- One faulty module can block or crash every other module's tools on its node.
And the shortcut ties the runtime to Node.js: only a JavaScript runtime can import JavaScript. The
operator's direction on 2026-10-03: *a module's tools are written in any language and the builder
builds them; the runtime runs them all and announces them; it should be fully language agnostic,
and node-tools can be rewritten in Go.*
## Considered Options
1. **Keep the shortcut.** Rejected: it is the cause of the three defects above, and it pins the
runtime's language.
2. **Launch every served bundle; the runtime knows how to start each language** (`node` for a
`.js`, exec for a binary). Rejected: the runtime would hold a table of interpreters, and a
runtime in Go would carry Node.js's knowledge for nothing.
3. **Launch every served bundle, and the build makes each served entrypoint executable.** Chosen.
A compiled language's binary is executable already; for an interpreted one the toolchain writes
a launcher beside the entrypoint — for TypeScript, a file that imports the entrypoint and serves
what it registered over stdio, using the bundle's own SDK. The runtime execs what it is given.
## Decision
**1. Every bundle the node's runtime serves is a child process speaking MCP over stdio.** The
in-process shortcut of ADR 0188 §2 is withdrawn. Everything else ADR 0188 §2 says — `tools/list`,
`tools/call`, `<seat>.<verb>` naming a seat's verb, the runtime serving each on the bus — stands.
**2. The runtime knows no language.** It is given an executable per served entrypoint and starts
it, with that bundle's environment ([ADR 0192](0192-a-tools-bundle-declares-what-it-is-given-and-the-runtime-hands-it-to-that-bundle-alone.md))
over its own words, and the module it serves it as. What makes an entrypoint executable is the
toolchain's business: a binary is one; an interpreted language's toolchain writes a launcher.
**3. A launched bundle is told the module it serves as**, so that what it lists unprefixed is that
module's own tools and a seat's verbs are always `<seat>.<verb>`, whichever it registered first.
**4. The runtime may be written in any language.** Nothing it does needs it to share a language
with a bundle; the mesh's runtime moves to Go, against this contract, once the contract is proven
in the runtime that exists.
## Consequences
- A process per served module per node. On the busiest machine that is a score of small children
where there was one process; a Go bundle costs a fraction of a Node.js one.
- The SDK resolve hook (issue 209) and the per-registration environment hand-off (ADR 0192 §3,
imported bundles) have nothing left to do and go; a launched bundle's environment is its own.
- A module's TypeScript tool code does not change: it registers as before, and the generated
launcher serves what it registered.
- A bundle that crashes or hangs takes only its own tools down, and is started again on its next
call, as ADR 0188 already provides for a launched bundle.
## How it is checked
| Rule | Checked by |
|---|---|
| Every served bundle is launched | the runtime's tests: a TypeScript bundle and a bundle in a second language, both launched, both answering over a real bus; a non-executable entrypoint is refused by name |
| The runtime knows no language | the runtime holds no interpreter: it execs the path it is given (code review; the Go runtime has no Node.js dependency at all) |
| A TypeScript served entrypoint is executable | the builder's test: a TypeScript bundle's served entrypoint has a launcher beside it, mode 0755 |
| A seat's verbs are named as the seat's whichever registers first | the SDK's test: a bundle registering its seat first and its own tools second lists `<seat>.<verb>` and its own tools unprefixed |
| Live | the packet filter's and intrusion prevention's seat verbs and the moved tools answer from launched bundles on every machine |
## References
- [ADR 0175](0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md),
[ADR 0188](0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md),
[ADR 0192](0192-a-tools-bundle-declares-what-it-is-given-and-the-runtime-hands-it-to-that-bundle-alone.md)
- [to-be 38](../03-DESIGN/01-to-be/38-building-the-operators-machine.md) WP4d
@@ -0,0 +1,157 @@
---
topic: the tiers
status: accepted
date: 2026-10-03
deciders: jochen
reconstructed: false
supersedes-in-part:
- 0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md
extends: 0191-the-meshs-resolver-holds-only-the-meshs-own-names.md
---
# 194. The mesh has one resolver, and every node asks it for the mesh's names
> **Narrowed, not replaced — 2026-10-03.** How a node asks is decided again by
> [ADR 0196](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md): every
> node and container asks `mesh-resolver` first and a public resolver only when it is silent. There is
> no `systemd-resolved` module and no runtime `dns` naming `mesh-resolver`, and step 2 of the migration
> reads as 0196 states it. Option 2 below was rejected for a laptop with its tunnel down resolving
> nothing; a public resolver listed second answers exactly then. The one resolver, its placement and
> the retirement of every per-node copy stand.
## Context
**Every node runs its own resolver and holds its own copy of the mesh's names.** On the production
mesh on 2026-10-03, each of the four nodes held `node-dns-resolver` with dnsmasq, fed on every push
with a zones file (one wildcard per node) and a region of `/etc/hosts` (the machines), and pointed
its own `/etc/resolv.conf` at itself. The controller computes the names once; four daemons then hold
four copies, each read in its own way.
**Every resolution fault found that day was a copy disagreeing with the truth, not the truth being
wrong:**
- **A copy read once.** dnsmasq reads `/etc/hosts` at start. After the controller stopped publishing
public names ([ADR 0191](0191-the-meshs-resolver-holds-only-the-meshs-own-names.md)), every node's
hosts file was right and every resolver still answered the mail server's public name with a
tunnel address, until each was restarted.
- **A copy beside other copies.** On the workstation, a name resolved to two addresses in rotation:
the mesh's region gave the tunnel address, and two lines the operator had written before the mesh
existed — one in `/etc/hosts`, one in a file the resolver also reads — gave the LAN address. A
comment beside one of them said to delete it once the mesh took over; nothing made that happen.
- **A copy that became somebody else's resolver.** The home server's resolver also answers its LAN
(a listen address added 2026-10-02), and the LAN's router hands that address out as the only DNS
server. Every phone and television on the LAN resolved through a mesh node's private copy, which is
how ADR 0191's outage reached them.
**And the overlay already has one centre.** Every node has exactly one tunnel peer — the anchor —
and routes the whole private range through it. Two nodes on the same LAN reach each other through
the anchor. So a name under `.internal` is only ever useful while the anchor is reachable: a resolver
anywhere else adds a copy without adding an answer anybody can use.
**What a node asks is already a separate role.** The connectivity design split *serving* (answers
the names) from *asking* (decides what the machine asks), because systemd-resolved cannot answer a
wildcard and can only route the mesh's suffix to something that can
([connectivity §2](../03-DESIGN/01-to-be/08-connectivity.md)). ADR 0121 kept them as two seats,
`node-dns-resolver` and `node-resolver-config`, both at node scope. No node runs systemd-resolved
today; each writes `/etc/resolv.conf` as a plain file pointing at its own dnsmasq.
## Considered Options
**1. Keep a resolver on every node, and make the copies more careful.** Restart on every file it
reads, own every file it reads, refuse to listen on a LAN. Each is a fix for one way a copy goes
stale, and the next way is not on the list yet. It keeps four answers to one question.
**2. One resolver for the mesh, and every node sends it every query.** The simplest asking side —
`resolv.conf` names the mesh's resolver and nothing else. Rejected: public resolution then depends on
the tunnel. A laptop whose tunnel is down could resolve nothing at all, and a public name would take
a detour through the anchor for no reason ADR 0191 left standing.
**3. One resolver for the mesh's names; each node asks it for those only.** The mesh's resolver holds
every node's internal domain. Each node's asking role routes the mesh's suffix to it and every other
name to public resolvers. Chosen.
## Decision
**The mesh has one resolver.** It is a module holding a new mesh-scoped seat, **`mesh-resolver`**
(capacity one). It holds each node's internal domain — `<node>.internal` and everything under it, at
that node's private address ([ADR 0191](0191-the-meshs-resolver-holds-only-the-meshs-own-names.md))
— and listens on the private network only. It is placed on the node every tunnel converges on, so
that it shares the overlay's single point rather than adding one. Which daemon fills the seat stays
the module's business, as the connectivity design says.
**Every node asks it for the mesh's names and nothing else.** `node-resolver-config` routes the
mesh's suffix to `mesh-resolver` and leaves every other name with public resolvers.
**Why a stub on every node.** `resolv.conf` cannot route by domain: the C library asks the servers it
lists in order, for every name, and moves to the next only when one does not answer — an NXDOMAIN
from the first is final. Listing `mesh-resolver` first sends every public name through the tunnel
(option 2); listing a public resolver first means `.internal` is never asked of the mesh. Something on
the node has to look at the name before choosing a server, and that is a stub resolver. Keeping
dnsmasq for it would keep a daemon that reads hosts files and can be told to answer a LAN — the two
ways copies went wrong. systemd-resolved holds no names of its own, routes by domain natively (a
routing domain `~<suffix>` on the server that answers it), and is part of systemd, already installed
on every node and enabled on none.
**So the asking side is a `systemd-resolved` module**, claiming `node-resolver-config` — the same claim
as the `resolv-conf` module it replaces, so the mesh refuses both on one node. It enables the service,
writes its configuration (the mesh resolver for the suffix, public resolvers for everything else), and
writes `/etc/resolv.conf` as a file naming the stub — a file the module owns, not a link to one.
**A container asks the mesh's resolver directly.** The container runtime cannot use a loopback stub
and drops its routing domains, so the runtime's `dns` names `mesh-resolver`, which forwards public
names for the containers that ask it. This is the one place a public name passes through the mesh,
and it is stated rather than hidden.
**`node-dns-resolver` is retired**, and with it every per-node copy: the zones file, the mesh's region
of `/etc/hosts` (the floor connectivity §2 already planned to remove), and the daemon on every node
but the one holding `mesh-resolver`. This narrows ADR 0121's *"the-dns-port → node-dns-resolver"*:
the serving role keeps its distinction from the asking role and moves to mesh scope, as ADR 0121 did
for the private network.
**A LAN's resolver is not the mesh's.** No device that is not a member can reach a private address,
so no member's resolver answers a LAN on the mesh's behalf. A router that hands out a node's address
as a LAN's DNS server is pointed elsewhere before that node stops answering.
**The order is fixed, because every step before the last leaves a working resolver:**
1. `mesh-resolver` is assigned and answers on the private network.
2. Each node's `node-resolver-config` moves from `resolv-conf` to `systemd-resolved`, and the container
runtime's `dns` to `mesh-resolver`.
3. A LAN whose router points at a node's resolver is pointed at its router or a public resolver.
4. `node-dns-resolver` is unassigned from every node, and the hosts region is withdrawn.
## Consequences
- **One answer per name.** A name is wrong in one place or right everywhere; no node can hold a copy
that disagrees, and no operator file on a node is read by the mesh's resolver.
- **The anchor down means no `.internal` names** — which it already meant for `.internal` traffic,
since every tunnel goes through it. Public resolution on every node is unaffected.
- **A container's public resolution depends on the mesh's resolver.** Accepted, and named in the
decision; a container that must resolve public names with the tunnel down is the case it costs.
- **Every node runs systemd-resolved**, through the `systemd-resolved` module. It is installed
everywhere already and enabled nowhere; the mesh still ships no resolver of its own.
- **The runtime's `dns` changes once per node**, which the runtime reads only at start. With
`live-restore` already on, that restart keeps every container running.
- **A LAN loses a resolver it had borrowed.** The router change is an explicit step, done through
the module that manages the router, before the node's resolver goes.
**How each is checked:**
- **One holder:** the seat has capacity one, so a second assignment is refused by the controller.
- **Asking:** on each node, `resolvectl` shows the tunnel's link with `mesh-resolver` and the suffix as
its routing domain; a name under `.internal` is answered by it, and a public name is answered
without it (its query log shows no public name from a node).
- **No copies:** no node but the holder answers DNS on a private or LAN address — every other node's
port 53 is systemd-resolved's loopback stub and nothing else — and no node's `/etc/hosts` carries a
mesh region.
- **A LAN:** the router's DHCP DNS option names no node's address.
## References
- [ADR 0191](0191-the-meshs-resolver-holds-only-the-meshs-own-names.md) — what the mesh's resolver
holds; this record decides where it runs and how nodes reach it.
- [ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md) — the two
resolver seats, and the private network's move to mesh scope this mirrors.
- [Connectivity §2](../03-DESIGN/01-to-be/08-connectivity.md) — serving and asking as two roles;
amended alongside this record.
- [The seats](../03-DESIGN/01-to-be/26-the-seats.md) — the seat table, amended alongside.
@@ -0,0 +1,102 @@
---
topic: what runs on it
status: accepted
date: 2026-10-03
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0152-the-operators-surface-is-a-module-the-console.md
---
# 195. The mesh's tools are found by address, not announced whole
## Context
The console answers MCP on a machine's loopback ([ADR 0152](0152-the-operators-surface-is-a-module-the-console.md),
[to-be 34](../03-DESIGN/01-to-be/34-the-console.md)) and announces, at a session's start, every tool
the mesh can say it has: 228 on 2026-10-03, 110 KB, taken once. Three things are wrong with that,
measured in research [021](../01-RESEARCH/021-finding-a-tool-in-the-mesh/00-overview.md):
- **Size.** Only one client's habit of deferring long lists keeps them out of the model's context.
- **Ambiguity.** A module on two machines is listed once, `node` optional, *whichever answers* — for
postgres and mssql, whose instances hold different data, a call that names no machine asks an
arbitrary one.
- **Staleness.** A tool that arrives after the session started is not listed until it reconnects.
The mesh already has the structure a caller needs: seats held once for the mesh, seats held once per
machine ([ADR 0159](0159-a-tool-call-names-the-machine-and-a-holder-serves-its-seats-verbs.md)), and
modules assigned to machines, each assignment issued its own subjects
([ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)).
The operator's direction: *tools are discoverable, in layers, and asking novox's postgres is not asking
ace's.*
## Considered Options
1. Keep the flat list and rely on the client. Rejected: the ambiguity and the staleness stay, and it
is one client's behaviour.
2. One tool per assignment, the machine in the name. Rejected: the list multiplies, and an address in
a tool's name meets the API's limit — letters, digits, `_` and `-`, at most 64 characters.
3. **A fixed handful of tools that walk the mesh's structure, the address an argument.** Chosen.
4. MCP resources or prompts. Rejected: unevenly supported, and an agent acts through tools.
## Decision
**1. Everything the mesh answers has one address, by the layer it lives in:**
| layer | address | answered by |
|---|---|---|
| a seat held once for the mesh | `<seat>.<verb>` | that seat's holder |
| a seat held once per machine | `<node>/<seat>.<verb>` | that machine's holder |
| a module assigned to a machine | `<node>/<module>.<tool>` | that assignment |
| a module whose instances are interchangeable (ADR 0160) | `<module>.<tool>` as well | any of them |
**A call to a module that is not interchangeable names its machine, or is refused** naming the machines
it runs on. "Whichever answers" is no longer an answer for state a machine holds.
**2. The console announces a fixed set of tools, not the catalogue:**
- **`mesh_overview`** — the mesh's seats with their verbs, and its machines;
- **`mesh_machine`** — one machine: the node seats it holds and the modules assigned to it, each with
its tools by name;
- **`mesh_search`** — words in, matching addresses out with one line each, across every layer;
- **`mesh_describe`** — one address in, its description and argument schema out;
- **`mesh_call`** — an address and its arguments in, the answer out, with the machine that gave it.
Each is answered from the mesh when it is asked, so a tool that arrived a minute ago is found without
the client reconnecting. The names are the API's kind of name; addresses never have to be.
**3. The flat catalogue stays reachable, not announced:** the `mesh` client and a console setting can
still list it whole, for a person reading it or a client that wants it. An agent pointed at the console
sees the five.
> **The mechanism changed — 2026-10-03, by ADR 0197.** Where the console learns what exists: not
> from the catalogue's roster and the controller's printed lists, but from every runtime announcing
> itself on the bus in the NATS services protocol, checked against the controller's records read as
> JSON. The addresses and the five tools stand.
## Consequences
- An agent spends a call or two finding a tool it does not know, and none on one it does; the context
no longer carries 110 KB it mostly never uses.
- The ambiguity is closed by the address, not by a description asking the agent to remember `node`.
- What got harder: an agent that once saw a tool's schema up front now asks for it. `mesh_describe` and
`mesh_search` answering with the schema of a close match keep that to one call.
- The discovery verbs are the console's; the mesh's own records — seats, machines, assignments — are
the controller's, and the console asks it rather than keeping a copy.
## How it is checked
| Rule | Checked by |
|---|---|
| The console announces five tools | the console's test: `tools/list` answers exactly the five |
| An address resolves to one subject per layer | the console's tests: a mesh seat, a node seat, an assignment, an interchangeable module, each called by address over a real bus |
| A non-interchangeable module without a machine is refused, naming its machines | the same tests |
| A tool that arrives after the session started is found | a test registering a module after the console's first answer and finding it by `mesh_search` |
| Live | from a fresh session, *which databases does novox's postgres hold* is answered by novox's postgres, found through the five |
## References
- [ADR 0152](0152-the-operators-surface-is-a-module-the-console.md),
[ADR 0159](0159-a-tool-call-names-the-machine-and-a-holder-serves-its-seats-verbs.md),
[ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)
- Research [021](../01-RESEARCH/021-finding-a-tool-in-the-mesh/00-overview.md)
- [to-be 34](../03-DESIGN/01-to-be/34-the-console.md)
@@ -0,0 +1,98 @@
---
topic: the tiers
status: accepted
date: 2026-10-03
deciders: jochen
reconstructed: false
supersedes-in-part:
- 0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
---
# 196. A node asks the mesh's resolver first, and a public one only when it is silent
## Context
**[ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) gave the
mesh one resolver and had each node ask it for the mesh's names only.** Because `resolv.conf` cannot
route by domain, that needed a stub on every node — a `systemd-resolved` module — and a separate
`dns` for the container runtime, which cannot use a loopback stub. It rejected the simpler shape,
every node sending every query to the mesh's resolver, on the grounds that *"a laptop whose tunnel is
down could resolve nothing at all."*
**That is true only of a `resolv.conf` naming the mesh's resolver alone.** The C library asks the
servers it lists in order and moves to the next when one does not answer within its timeout. A public
resolver listed second is asked exactly when the mesh's is unreachable — the anchor down, the tunnel
down, a laptop behind a captive portal that has not let the tunnel up — and never otherwise. An answer
from the first, including "no such name", is final, so `.internal` is never asked of a public resolver
while the mesh's answers.
**And the container runtime copies a machine's resolvers into its containers when they are not
loopback addresses.** With the mesh's resolver and a public one listed, every container gets both, as
they are, with nothing configured for the runtime.
## Considered Options
**1. Keep ADR 0194's stub.** Public names never touch the mesh, and a node with the anchor down
resolves public names at full speed. It costs a module and a running service on every node, a second
configuration for containers, and the one asymmetry ADR 0194 had to state — containers' public names
through the mesh, nodes' not.
**2. Every node asks the mesh's resolver for everything, with a public resolver as the silent
fallback.** One server answers every node and every container; nothing on a node routes, holds names,
or runs. Chosen.
## Decision
**A node's `/etc/resolv.conf` names the mesh's resolver first and a public resolver second, with a
short timeout and a single attempt.** It is written by the module holding `node-resolver-config` — the
existing `resolv-conf` — which now names `mesh-resolver`'s address instead of the machine's own. The
mesh's resolver answers the mesh's names from what it holds and forwards every other name, giving the
public answer ([ADR 0191](0191-the-meshs-resolver-holds-only-the-meshs-own-names.md) is unchanged: no
public name gets a private answer).
**Containers take the same two resolvers from their machine.** The container runtime's own `dns`
setting is not written; the runtime copies the machine's non-loopback resolvers into every container.
**This replaces, from ADR 0194:** the asking side as a stub (*"So the asking side is a
`systemd-resolved` module"*), the container runtime's `dns` naming `mesh-resolver`, and step 2 of the
migration as written. There is no `systemd-resolved` module. Everything else in ADR 0194 stands — one
`mesh-resolver`, on the node every tunnel converges on, holding each node's internal domain, the
retirement of `node-dns-resolver` and every per-node copy, and a LAN's resolver not being the mesh's.
**The migration, as it now reads:**
1. `mesh-resolver` is assigned and answers on the private network.
2. Each node's `resolv-conf` names `mesh-resolver` first and a public resolver second.
3. A LAN whose router points at a node's resolver is pointed at its router.
4. `node-dns-resolver` is unassigned from every node, and the hosts region is withdrawn.
## Consequences
- **Every name a node or container asks goes through the anchor while it is up.** A public lookup
takes a few milliseconds longer than asking a public resolver directly, and the mesh's resolver sees
every name its nodes look up. It is the operator's own server.
- **With the anchor unreachable, each lookup waits out one timeout, then resolves publicly.**
`.internal` names fail then — as `.internal` traffic does, every tunnel going through the anchor.
- **A LAN is unaffected by this choice.** Devices that are not members never read a node's
`resolv.conf`; they get their resolver from their router, which step 3 points at itself.
- **Nothing new runs on a node.** No stub, no module, no per-node configuration for containers.
- **The runtime's `dns` key goes with `node-dns-resolver`.** The dnsmasq module wrote it into the
runtime's configuration; unassigning that module in step 4 withdraws it, and the runtime reads the
change only when it next starts — with `live-restore` on, that restart keeps every container
running.
**How each is checked:**
- **Order:** each node's `/etc/resolv.conf` lists `mesh-resolver`'s private address first and a public
resolver second, and nothing else.
- **Fallback:** with `mesh-resolver` unreachable from a node, a public name still resolves there, after
the timeout.
- **Containers:** a container started on a node lists the same two resolvers.
- **A LAN:** the router's DHCP DNS option names the router, not a node.
## References
- [ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) — the one
resolver; this record replaces how nodes and containers ask it.
- [ADR 0191](0191-the-meshs-resolver-holds-only-the-meshs-own-names.md) — what the resolver holds.
- [Connectivity §2](../03-DESIGN/01-to-be/08-connectivity.md), amended alongside this record.
@@ -0,0 +1,85 @@
---
topic: what runs on it
status: accepted
date: 2026-10-03
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0195-the-meshs-tools-are-found-by-address-not-announced-whole.md
---
# 197. Every tool announces itself on the bus, in the NATS services protocol
## Context
[ADR 0195](0195-the-meshs-tools-are-found-by-address-not-announced-whole.md) gave every tool an
address and the console five tools to find them. Where the console learns what exists, it inherited
from [to-be 34](../03-DESIGN/01-to-be/34-the-console.md) §3: ask the catalogue for its roster, ask
each module on the roster for its `tools`, and read the machines and assignments from the
controller's printed `node list` and `module list`. Built that way on 2026-10-03, it worked and showed
what is wrong with it:
- **It asks what should exist and infers what does.** The roster holds every module the catalogue
ever registered; 47 of them were reported "not answering" on 2026-10-03, most with no tools at all
and several retired.
- **It parses prose.** Two of the controller's answers are text for a person, and a reworded column
breaks discovery.
The bus already knows what answers. NATS has a services protocol for exactly this: a service answers
`$SRV.PING` and `$SRV.INFO` — every instance, on one request — with its name, its instance, and every
endpoint's subject and metadata, in a published format the NATS tools read. The operator's direction:
*every tool announces itself; the mesh has the full picture, so nothing should be inferred or parsed.*
## Considered Options
1. **Keep asking the roster, and give the controller JSON answers.** Fixes the parsing, keeps the
inference.
2. **Re-serve every tool through a NATS services library.** The announcement for free, but every
runtime's serving path rewritten around a library, in two languages, for no change in behaviour.
3. **Every runtime answers the services protocol's discovery subjects with what it serves; serving
is unchanged.** Chosen.
## Decision
**1. What answers announces itself.** Every runtime that serves tools — each machine's tool runtime,
the per-module runtimes still in containers, and the controller for the seat it holds — answers
`$SRV.PING` and `$SRV.INFO` in the NATS services format: one service per module or seat it serves,
named for it, its instance the machine; one endpoint per tool, its subject and queue exactly as
served, its metadata the tool's description, argument schema, the machine, the seat and scope where
it is a seat's verb, and whether the module's instances are interchangeable.
**2. The console finds what exists by asking the bus,** one `$SRV.INFO` request, every answer
gathered for a short window. What it announces through ADR 0195's five tools is what answered.
**3. What should exist is the mesh's records, read as data.** The controller answers its machines and
modules as JSON, and says which modules declare tools; the console names as *not answering* only an
assignment that declares tools and did not announce them. A module with no tools is never listed.
**4. The grants say so.** Every principal that serves tools may subscribe the services discovery
subjects for what it serves; the console's and every runtime's account may publish the discovery
request. Replies travel to the asker's own inbox as every reply does.
## Consequences
- The standard `nats micro list` and `nats micro info` show the mesh's tools, live, to anybody holding
a credential — the bus's own view, not the mesh's description of it.
- Discovery costs one request and a gathering window, not one request per roster entry.
- What got harder: three runtimes must answer the same format the same way — the Go tool runtime, the
TypeScript runtime the containers still run, and the controller. The format is NATS's, so a test
reads all three with the NATS services client and nothing of the mesh's.
## How it is checked
| Rule | Checked by |
|---|---|
| A runtime announces exactly what it serves | each runtime's test: `$SRV.INFO` answered with one service per served module or seat, its endpoints' subjects the subjects served |
| The format is NATS's | the same tests read the answer with the NATS services client's own types |
| A module with no tools is never listed; an assignment with tools that did not answer is | the console's test, against controller records with both |
| Nothing is parsed from prose | the console reads only JSON answers (code review; the text parsers are deleted) |
| Live | `nats micro list` against the mesh's bus lists every machine's tool runtime and the controller |
## References
- [ADR 0195](0195-the-meshs-tools-are-found-by-address-not-announced-whole.md),
[ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md),
[ADR 0175](0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md)
- [to-be 34](../03-DESIGN/01-to-be/34-the-console.md)
@@ -0,0 +1,79 @@
---
topic: what runs on it
status: accepted
date: 2026-10-03
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md
---
# 198. A module's long-running code is launched by the node's runtime, and reaches the bus through it
## Context
[ADR 0193](0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md) made
every tools bundle a child the node's runtime launches, speaking MCP over stdio, and gave the channel
one bus verb: a tool's emit, published by the runtime as the module. Twenty-three modules still run the
rest of their own code — event handlers, provisioners, a preparation step, three mains — in a container
on the runtime's image, because that code needs what a container gave it: a bus connection that can
*subscribe*, and its module's words. Research [022](../01-RESEARCH/022-where-a-modules-long-running-code-runs/00-overview.md)
measured what it uses. [ADR 0188](0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md)
§1 says a module's own code is never an image.
## Considered Options
1. **The runtime launches it and is its bus.** Chosen.
2. A process per module with its own bus client and credential. Rejected for the reasons ADR 0188
rejected it for tools: a transport in every language's SDK, a credential per module on disk, and a
bus change rebuilding every module.
3. Keep the containers for it. Rejected: ADR 0188's rule stays broken for most of the catalogue.
## Decision
**1. A module's long-running code is a bundle the node's runtime launches and supervises,** exactly as
its tools are: an executable entrypoint, given the runtime's words and its module's
([ADR 0192](0192-a-tools-bundle-declares-what-it-is-given-and-the-runtime-hands-it-to-that-bundle-alone.md)),
started at the runtime's start and again when it exits. A bundle may serve tools, run long, or both.
**2. The runtime is its bus.** The stdio channel carries, beside MCP, the mesh's verbs a module's code
uses: `mesh/publish` (ADR 0193), **`mesh/subscribe`** — the runtime binds that module's durable
consumer, as the module's own runtime did, and delivers each event to the child as a `mesh/event`
request, acknowledging it on the bus only when the child has answered — and **`mesh/ask`**, a tool
call made on the module's behalf. The runtime's account is granted what each module it carries
consumes, and the consumer keeps the module's name, so no event is lost or replayed in the move.
**3. A preparation step is a run-once process the host runs before the runtime starts the module,**
with its module's words and no bus — what it already was.
**4. What a container reached by its network is reached on the machine.** A service by its published
port (`${port:…}`) on loopback; a backend's command-line client as a package of the machine's system,
or, where the system has none, the backend's own driver inside the bundle.
## Consequences
- The per-module containers go, and with them the runtime image as a way module code runs; ADR 0188's
registration rule can then refuse a module's own image without exception.
- One bus connection per machine carries every module's events; a module's handler is a function of
the events it is handed, in any language, with no bus client of its own.
- What got harder: the runtime holds every carried module's consumer and must not acknowledge an event
before the child has handled it — a child that dies mid-event leaves it unacknowledged, and it is
delivered again. The runtime's grants widen to what its modules consume.
- Two modules need code before they can move: their backends' clients exist on no machine's system,
so they talk to the backend through a driver instead.
## How it is checked
| Rule | Checked by |
|---|---|
| A subscribed event reaches the child and is acknowledged only after it answered | the runtime's test over a real bus: a child that answers is acknowledged once; one that dies mid-event is delivered again |
| The module's consumer keeps its name | the composer's test: the durable consumer the runtime binds is the one the module's own runtime bound |
| No module's own code is an image | the catalogue's registration check, without exception, once the last container has moved |
| Live | every moved module's provisioner and handlers act, on their machines, from the node's runtime; `docker ps` shows no runtime-image container on any machine |
## References
- [ADR 0188](0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md),
[ADR 0192](0192-a-tools-bundle-declares-what-it-is-given-and-the-runtime-hands-it-to-that-bundle-alone.md),
[ADR 0193](0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md)
- Research [022](../01-RESEARCH/022-where-a-modules-long-running-code-runs/00-overview.md)
- [to-be 38](../03-DESIGN/01-to-be/38-building-the-operators-machine.md) WP4c
+6
View File
@@ -223,6 +223,8 @@ python3 00-META/checks/index.py fail if stale
- **0148** — [The mesh's names are resolved, not copied into every container](0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
- **0151** — [A route's internal name is composed under the node that serves it](0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md)
- **0191** — [The mesh's resolver holds only the mesh's own names; a public name resolves publicly](0191-the-meshs-resolver-holds-only-the-meshs-own-names.md)
- **0194** — [The mesh has one resolver, and every node asks it for the mesh's names](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)
- **0196** — [A node asks the mesh's resolver first, and a public one only when it is silent](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md)
### What runs on them, and how it gets there
@@ -292,6 +294,10 @@ python3 00-META/checks/index.py fail if stale
- **0183** — [The Anthropic licence manager is a module holding a seat; it hands each node's agent its token over the bus, sealed; the controller and the host have no part](0183-the-anthropic-licence-manager-is-a-module-and-hands-tokens-to-the-agent-over-the-bus.md)
- **0188** — [A module's own code is bundles in any language, and a tools bundle speaks MCP to the runtime](0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md)
- **0192** — [A tools bundle declares what it is given, and the runtime hands it to that bundle alone](0192-a-tools-bundle-declares-what-it-is-given-and-the-runtime-hands-it-to-that-bundle-alone.md)
- **0193** — [Every bundle the runtime serves is launched, and the runtime knows no language](0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md)
- **0195** — [The mesh's tools are found by address, not announced whole](0195-the-meshs-tools-are-found-by-address-not-announced-whole.md)
- **0197** — [Every tool announces itself on the bus, in the NATS services protocol](0197-every-tool-announces-itself-on-the-bus-in-the-nats-services-protocol.md)
- **0198** — [A module's long-running code is launched by the node's runtime, and reaches the bus through it](0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md)
### How it is built
+40 -9
View File
@@ -10,6 +10,8 @@ code:
- mesh-host internal/apply (the service that reflects a rule set)
updated: 2026-10-03
decisions:
- 02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md
- 02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
- 02-DECISIONS/0191-the-meshs-resolver-holds-only-the-meshs-own-names.md
- 02-DECISIONS/0180-the-found-front-end-is-uninstalled-once-a-machine-is-converged.md
- 02-DECISIONS/0170-the-firewall-seat-serves-its-verbs.md
@@ -290,8 +292,10 @@ expensively enough to be worth restating:
- **A node must not pin its own public name locally.** The duplicate record breaks resolution of
that name for everything else that needs it.
**What the host receives:** the resolver's configuration, as files, listing every peer's internal
name and overlay address.
**What the host receives:** what to ask, not what to answer. The mesh has **one resolver**, holding
every node's internal domain; a node asks it first and a public resolver only when it is silent
([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md),
[ADR 0196](../../02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md)).
**What goes away:** the `/etc/hosts` floor. It exists because a node had to reach the mesh
database before its own DNS existed; with [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)
@@ -349,6 +353,33 @@ not a list of containers.
somebody starts by hand is not the mesh's to configure, and reaching into every container on a
machine — declared or not — is what a nameserver in `resolv.conf` would be for.
### One resolver for the mesh
*2026-10-03.* **The mesh's names live in one place: the module holding `mesh-resolver`**, a mesh-scoped
seat of capacity one, placed on the node every tunnel converges on. It holds one wildcard per node —
`<node>.internal` and everything under it — and listens on the private network only. It answers the
mesh's names from what it holds and forwards every other name, giving the public answer.
**Every node asks it for everything, and a public resolver only when it is silent.** The module
holding `node-resolver-config` writes `/etc/resolv.conf` naming `mesh-resolver` first and a public
resolver second, with a short timeout and one attempt: the C library moves to the second only when the
first does not answer — the anchor or the tunnel down, a captive portal holding the tunnel back — so
public names keep resolving then, and `.internal` is never asked of a public resolver while the mesh's
answers. Containers take the same two from their machine, the runtime copying non-loopback resolvers
into every container, so the runtime is given no `dns` of its own
([ADR 0196](../../02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md),
replacing ADR 0194's per-node `systemd-resolved` stub).
**No node holds a copy.** The per-node resolver, its zones file and the mesh's region of `/etc/hosts`
go: every resolution fault found on 2026-10-03 was a copy disagreeing with the truth — a hosts file
read once at start, an operator's old line beside the mesh's, a node's resolver lent to a LAN. No
member's resolver answers a LAN; a router pointing at one is moved first. *Checked by each node's
`/etc/resolv.conf` naming `mesh-resolver` then a public resolver, by no node but the holder answering
DNS on any address, and by the router's DHCP DNS option naming the router.*
*What follows describes the per-node resolver this replaces — how it was built and why the roles were
split. The split stands; the serving role's scope is what moved.*
### The resolver, built
*2026-08-31.* **A service is reached at `<service>.<node>.internal`** — the first label is the
@@ -425,15 +456,15 @@ the cost of not seeing it is inventing a mechanism that already exists.
### The mesh resolves only its own names; a public name resolves publicly
**The mesh's resolver holds names under the mesh suffix and nothing else** — every machine, and through
it every route's internal name `<label>.<node>.internal`
**The mesh's resolver holds each node's internal domain and nothing else** — `<node>.internal` and
everything under it, so every route's internal name `<label>.<node>.internal` with no line of its own
([ADR 0151](../../02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md)).
**A public name the mesh serves is never given a private answer**: it is forwarded and resolves to the
**A node's public domains — one or more — are never given a private answer**: it is forwarded and resolves to the
public address, from a member and from anything else the resolver answers — a resolver may serve a
machine's LAN, and a phone on that LAN must get the address it can reach
([ADR 0191](../../02-DECISIONS/0191-the-meshs-resolver-holds-only-the-meshs-own-names.md)). Inside the
mesh, a routed service is reached, and certified by the internal authority, under its internal name.
*Checked by the controller's catalogue tests — every name the roster carries ends in the mesh suffix —
*Checked by the controller's tests — the roster names the machines and no routed name —
and on a machine by asking its resolver for a public name the mesh serves: the answer is the public
address.*
@@ -988,6 +1019,9 @@ The list is worth having in one place, because it is most of the argument:
## Open
- **One resolver ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)).** Not built: every node still runs
`node-dns-resolver`. The migration's four steps are in the record, in order.
- ~~**What happens when the hub is down.**~~ **Resolved** by
[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md), together with `06`'s
matching question — they were one question. Nothing takes over. WireGuard has no failover, the
@@ -1009,9 +1043,6 @@ The list is worth having in one place, because it is most of the argument:
operator's to move between meshes, but the manifest layer still stores it as a literal — so today
the composition is a per-node override rather than the design. The interpolation that would let a
module carry a label and a node carry the domain, and the mesh join them, does not yet exist.
- **Withdrawing public names from internal resolution.** The roster still publishes every routed
public name at its serving node's private address, which ADR 0191 forbids; until the controller
stops, a resolver that answers a LAN hands that LAN's non-members addresses they cannot reach.
## The hub adopts the predecessor's tunnel
+5 -3
View File
@@ -1,6 +1,6 @@
---
layer: to-be
status: implemented
status: in-progress
code:
- mesh-controller internal/catalogue/seats.go
- mesh-controller internal/catalogue/resolve.go
@@ -10,8 +10,9 @@ code:
- mesh-controller cmd/mesh-controller/source.go
- mesh-controller internal/inventory/migrations/0032-a-source-may-live-on-a-seat.sql
- mesh-catalog modules/gitea/module.json
updated: 2026-10-01
updated: 2026-10-03
decisions:
- 02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
- 02-DECISIONS/0161-what-deserves-a-seat.md
- 02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
- 02-DECISIONS/0126-a-module-declares-its-own-seats.md
@@ -126,7 +127,8 @@ convention, which later seats departed from.
| `mesh-npm-package-registry` | `npm-package-registry` | mesh | `npm-package-registry` | the forge |
| `mesh-git` | `git` | mesh | `git` | the forge |
| `mesh-build-machine` | `the-build-machine` | node | — | a builder |
| `mesh-dns-port` | `the-dns-port` | node | — | the local resolver |
| `mesh-resolver` | — | mesh | — | the mesh's one resolver, holding every node's internal domain ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)) |
| ~~`mesh-dns-port`~~ | `the-dns-port` | node | — | retired by [ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md): the local resolver became the mesh's one |
| `mesh-intrusion-prevention` | `the-intrusion-prevention` | node | — | an intrusion-prevention service |
| `mesh-packet-filter` | `the-packet-filter` | node | — | the packet filter |
| `mesh-private-network` | `the-private-network` | node | — | the private network the mesh runs over |
+26 -2
View File
@@ -1,9 +1,11 @@
---
layer: to-be
status: implemented
status: designed
code: [mesh-catalog, mesh-tools, mesh-controller]
updated: 2026-10-02
updated: 2026-10-03
decisions:
- 02-DECISIONS/0197-every-tool-announces-itself-on-the-bus-in-the-nats-services-protocol.md
- 02-DECISIONS/0195-the-meshs-tools-are-found-by-address-not-announced-whole.md
- 02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md
- 02-DECISIONS/0159-a-tool-call-names-the-machine-and-a-holder-serves-its-seats-verbs.md
- 02-DECISIONS/0152-the-operators-surface-is-a-module-the-console.md
@@ -93,6 +95,28 @@ prerequisites are listed in that record. When the seat serves them, the console
modules' own, and the person stops opening a shell for the mesh's own questions. Until then the console
says so in its handshake.
## 3a. Found by address, not announced whole (2026-10-03)
*By [ADR 0195](../../02-DECISIONS/0195-the-meshs-tools-are-found-by-address-not-announced-whole.md);
this section governs where it and §2–§3 disagree.* The console announces five tools —
`mesh_overview`, `mesh_machine`, `mesh_search`, `mesh_describe`, `mesh_call` — and every tool the mesh
answers is reached through them by its address: `<seat>.<verb>` for a seat held once for the mesh,
`<node>/<seat>.<verb>` for one held per machine, `<node>/<module>.<tool>` for an assignment, and
`<module>.<tool>` as well for a module whose instances are interchangeable. A module that is not
interchangeable is called with its machine or refused with the machines it runs on. Each discovery
verb asks the mesh when it is called, so nothing is kept for a session's length; the flat catalogue
stays reachable through the `mesh` client and a setting, unannounced.
**What exists is what announced itself** ([ADR 0197](../../02-DECISIONS/0197-every-tool-announces-itself-on-the-bus-in-the-nats-services-protocol.md)):
every runtime answers the NATS services protocol's `$SRV.INFO` with what it serves, and the console
gathers one request's answers; the controller's records, read as JSON, say which assignments with
tools should have answered.
*Found 2026-10-03, measuring for that record:* §3's statement that a stateful module on two machines is
listed once per machine does not hold on the live console — postgres and mssql are listed once, `node`
optional, answered by whichever instance replies. The address replaces that statement rather than
repairing it.
## 4. Where it runs
On whichever machines an operator sits at, by assignment. It is not on the control node by default and
@@ -124,8 +124,8 @@ the playbooks in the record.
**Other tool servers** a person wants on every machine, or on one, are a declared setting of this module
— mesh layer or node layer — rendered into the same managed file. The person sets them with the
controller's `settings` verb on this module, so the list stays declared state; a tool of this module
cannot set it, because a bundle calls nothing (ADR 0192). The agent's own HTTP-only constraint for managed servers applies; a person's local
controller's `settings` verb on this module, so the list stays declared state; the list is the operator's
choice, set where every setting is set. The agent's own HTTP-only constraint for managed servers applies; a person's local
command-based servers stay their own, in their own file.
**The entry's name is `mesh`.** The hand-made entry both workstations carry today is named after this
@@ -157,9 +157,9 @@ decides it; to-be 39 is the manager's half. This module:
the file matches what was handed over — by fingerprint, never by value.
Switching is the seat's `switch` verb, asked through the console; this module only applies what it is
handed. *2026-10-03:* every exchange is started by the manager, because a tools bundle answers calls and
has no bus credential to make them ([ADR 0192](../../02-DECISIONS/0192-a-tools-bundle-declares-what-it-is-given-and-the-runtime-hands-it-to-that-bundle-alone.md),
ADR 0183's dated note). The tool names follow the catalogue's `<module>_<verb>` form.
handed. *2026-10-03:* every exchange is started by the manager, by the operator's direction (ADR 0183's dated
note); this module is a bundle the node's runtime launches over stdio and answers what it is asked
([ADR 0193](../../02-DECISIONS/0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md)). The tool names follow the catalogue's `<module>_<verb>` form.
## 6. Scope, settings and the order of assignment
@@ -169,7 +169,7 @@ extra tool servers. **Prerequisite:** the manager holds its seat and has adopted
**Order:** the manager assigned and a refresh observed; the console's provision in the catalogue; this
module on one workstation; the six predecessor files and the hand-made console entry removed there; a
new session read to confirm it sees the mesh's instruction file, the console's tools under `mesh`, and
new session read to confirm it sees the mesh's instruction file, the console's five tools under `mesh`, and
its licence; then the rest.
## 7. The package
@@ -193,7 +193,7 @@ installer is rejected: it puts a self-updating binary under the person's home, i
| a switch asked of the seat through the console changes the licence and the token on the node; no tool answer and no log line holds a token | ADR 0183 |
| the API-key binding writes nothing under the home and the agent authenticates through the helper | ADR 0183 |
| the console's provision resolves by co-location; a machine without the console refuses the module by name | ADR 0027, ADR 0152 |
| a new session on the assigned workstation lists the console's tools under `mesh` and answers "which node am I" from the instruction file | the exit of the build |
| a new session on the assigned workstation lists the console's five tools under `mesh` ([ADR 0195](../../02-DECISIONS/0195-the-meshs-tools-are-found-by-address-not-announced-whole.md)) and answers "which node am I" from the instruction file | the exit of the build |
## What this does not settle
@@ -13,6 +13,8 @@ decisions:
- 02-DECISIONS/0149-the-live-mesh-is-the-test-bed.md
- 02-DECISIONS/0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md
- 02-DECISIONS/0192-a-tools-bundle-declares-what-it-is-given-and-the-runtime-hands-it-to-that-bundle-alone.md
- 02-DECISIONS/0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md
- 02-DECISIONS/0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md
---
# 38. Building the operator's machine
@@ -253,6 +255,38 @@ sockets — so seven move here. *Built 2026-10-03:* mesh-sdk #12 (`collectToolsE
(each bundle its own environment), mesh-controller #236 and #237 (the words composed, made the
account's to read, and named files restarting the runtime), and the seven modules in one change.
Building it found issue [211](../../04-ISSUES/211-a-bundle-is-built-before-the-toolchain-it-is-compiled-in/00-report.md).
*Proven live 2026-10-03* (mesh-catalog #242, #243): on the one machine that runs them, baserow,
letta, searxng and unifi answer from the runtime with no tool container — each reading its
configuration file as the operator's account — and the runtime serves seventeen tools for six
modules there. confluence, gitlab and jira are assigned nowhere and retire with the predecessor.
Two traps met on the way: a tools bundle whose module declares no `tools` list must say `loads`, or
the composer delivers it nowhere while the build reports success; and a module whose builds are
pinned to an old commit is left out of a merge's plan and must be built from `main` by hand.
## WP4d — Every served bundle is launched; the runtime in Go
*mesh-sdk, mesh-controller, mesh-tools. [ADR 0193](../../02-DECISIONS/0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md).*
**In order.** The SDK's stdio loop serves what a bundle registered under the module it is told it
serves as. The builder writes, beside every TypeScript entrypoint, an executable launcher that
imports it and serves what it registered; the composer names the launcher where it named the
entrypoint. The runtime launches every served entrypoint and imports none; the resolve hook and the
per-registration environment go. Proven live on all four machines. Then the runtime is rewritten in
Go against the same contract — the bus, the memberships and seats, the launcher, the console's MCP
over HTTP — and replaces the TypeScript one, proven the same way.
**Proof.** The tests ADR 0193 names; live, every moved module's tools and both node seats answer from
launched bundles on every machine, and then do again from the Go runtime.
*Built and proven live 2026-10-03.* mesh-sdk #13/#14 (0.1.4, 0.1.5: served as the named module; an
emit travels through the runtime), mesh-controller #239/#240 (a launcher beside every TypeScript
entrypoint; a runtime compiled to a binary runs itself), mesh-host #81 (`./name` is the process's own
binary), mesh-tools #35/#36/#37/#38 (the module named; launch-only; the toolchain requiring 0.1.5; the
runtime in Go). On all four machines node-tools is now the Go binary, launching every served bundle:
both node seats answered from it on every machine and the four moved modules on theirs. Found on the
way: issue [212](../../04-ISSUES/212-a-toolchain-rebuild-keeps-the-sdk-it-cached/00-report.md) (the
toolchain image kept a cached SDK, and the seats' verbs went unanswered on three machines for an hour),
and the controller's plan losing track of its own rebuild when it restarts mid-plan.
## WP4c — The module's own long-running code moves
@@ -264,6 +298,17 @@ consumes with, the words its code reads at import, the packages the image instal
client), and the service it reaches by a container network name. That begins with a decision record,
after which the twenty-three move and the registration gate refuses the container shape for all.
*Decided 2026-10-03, [ADR 0198](../../02-DECISIONS/0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md):*
the node's runtime launches that code as it launches tools, and is its bus — `mesh/subscribe` and
`mesh/ask` beside `mesh/publish` on the stdio channel, the module's own durable consumer bound by the
runtime and acknowledged only after the child answered. **In order:** the runtime's subscription and
its grants; the SDK's `on` and provisioner bound to the channel; then the modules in three waves — the
provisioners and handlers whose backends are reached on loopback with a system package (postgres,
redis, mosquitto, influxdb, keycloak, umami, cloudflare-dns, grafana, icecast, home-assistant, nodered,
nextcloud, minio), the two whose clients exist on no system (mongodb, mssql: a driver in the bundle),
and last the mesh's own (mesh-catalog, mesh-vault, records, gitea, mailu, audit-logger, lab, and the
three mains).
## WP5 — The shell, on a server first
*mesh-catalog #224, already written. Half a day to assign and prove.*
@@ -74,8 +74,8 @@ Carried from the predecessor, where each rule was earned by an incident:
## 4. Handing a token to a node
**The manager starts every exchange** (ADR 0183's dated note of 2026-10-03): the agent module is a
tools bundle, which answers and calls nothing. The manager asks each bound node's module for its public
**The manager starts every exchange** (ADR 0183's dated note of 2026-10-03), by `mesh/ask` through the
runtime that launched it ([ADR 0198](../../02-DECISIONS/0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md)). The manager asks each bound node's module for its public
key the first time and keeps it. From then on:
- **On rotation**, the manager calls `claude-code.apply@<node>` on every node bound to the rotated
@@ -23,11 +23,15 @@ packages, their order, their sizes and their proofs, and is wrong the moment it
It is the shape [design 38](38-building-the-operators-machine.md) gives the operator's machine.
*Revised 2026-10-03, after design 38's WP1–WP4b ran:* the node's tool runtime is live on all four
machines, tools are bundles it serves and each is given only the words its artifact declares, and a
bundle has no bus credential of its own. Three things in the first version of this plan changed with
that: the wait on design 38's WP3 is over; the agent module calls nothing, so the manager starts every
exchange (ADR 0183's dated note); and the manager's own process now waits on a different question,
named under WP4.
machines, tools are bundles it serves and each is given only the words its artifact declares, every
bundle is a child the runtime launches over stdio and is the bus for
([ADR 0193](../../02-DECISIONS/0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md),
[ADR 0198](../../02-DECISIONS/0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md)),
and the console offers five tools over addresses
([ADR 0195](../../02-DECISIONS/0195-the-meshs-tools-are-found-by-address-not-announced-whole.md)). What changed
in this plan: the wait on design 38's WP3 is over; the manager starts every exchange, by the operator's
direction (ADR 0183's dated note); and the manager's daemon is a long-running bundle the runtime launches,
which ADR 0198 decided the same day — nothing in this plan waits on another record.
## How this is built, and where it is run
@@ -52,8 +56,8 @@ Measured 2026-10-03 on the four machines and in the repositories.
| escalation | passwordless `sudo` for the operator account on all four — a fact about the machines, checked by nobody | how the agent module writes its managed directory under `/etc` |
| the agent itself | installed on all four, at four different versions, all above the one the managed tool-server key needs | declared as the module's package |
| a bundle's words | paths and constants written with `${dir:…}` and `${port:…}` only; a fact the mesh knows reaches a bundle as a file whose path is a word | the agent module's facts file and settings file |
| a bundle calling a tool | **not possible**: a bundle answers calls; it holds no bus credential | the manager starts every exchange |
| a module's own long-running process with a bus credential | **undecided** — design 38's WP4c names it as the question its next record answers | the manager's daemon (WP4) |
| a bundle calling a tool | `mesh/ask` through the runtime that launched it (ADR 0198); no bundle holds a bus credential | how the manager visits every node |
| a module's own long-running code | a bundle the runtime launches and restarts (ADR 0198); the runtime's subscription and grants built, the modules moving in design 38 WP4c's waves | the manager's daemon (WP3, WP4) |
| the vendor's refresh, the sealed box, the grant file | `anthropic-manager` in the catalogue, built on the controller placement ADR 0183 moved away from; assigned to nothing | its client ported into the manager; the module retired (WP6) |
| the credentials write, the strip, the identity read | `anthropic-consumer` in the catalogue; tested; assigned to nothing | ported into the agent module with its tests; the module retired (WP6) |
| the predecessor's manager and consumer | the lease per licence, the expiry floor, the lineage comparison, the identity guard, three touchpoints, cooldowns | ported as logic with its tests |
@@ -68,7 +72,7 @@ WP3 the manager's code, built and tested (mesh-catalog)
│
WP2 live: one workstation, configuration only — no licence yet ── the first live proof
│
WP4 the manager live on the control node ── waits on design 38 WP4c's record (a process's bus credential)
WP4 the manager live on the control node ── its daemon a long-running bundle (ADR 0198)
WP5 the licence end to end on one workstation
WP6 the rest of the nodes, and the predecessor's remains
```
@@ -127,7 +131,7 @@ contains a token. The catalogue's checks pass.
**Proof, live, on one workstation, configuration only.** Assign the module; set the node's role; push.
The agent's managed directory holds the two files; everything under the person's agent directory is
byte-identical to before; a new session lists the mesh's tools under `mesh` and answers *which node am
byte-identical to before; a new session lists the console's five tools under `mesh` and answers *which node am
I* from the managed instruction file. `claude_code_status` answers through the console. No licence is
touched: the module writes the credentials file only when it is handed a token.
@@ -136,7 +140,7 @@ touched: the module writes the credentials file only when it is handed a token.
*mesh-catalog. Two to three days.*
**What is written**, as design 39 says: the manifest (the seat and its verbs, a database, a `secret`
for the key the grants are encrypted with, a tools bundle, a process bundle for the daemon, settings
for the key the grants are encrypted with, a tools bundle and a long-running bundle for the daemon, both launched by the runtime, settings
with defaults); the store's migrations; the refresh with its plan, lease, floor and cadence as pure
functions; the vendor client from `anthropic-manager`; adoption from a file and from a node's waiting
login with the identity guard; usage and its threshold; the visit — key, hand-over, waiting login — per
@@ -149,9 +153,9 @@ node's key.
## WP4 — The manager live on the control node
*The live mesh. Half a day.* **Waits on design 38 WP4c's record** — how a module's own long-running
process is given a bus credential and its subscriptions — because the daemon must call the agent module
on every node. Until that record exists, nothing in this plan works around it: no tool container, no
*The live mesh. Half a day.* The daemon is a long-running bundle the control node's runtime launches
(ADR 0198); it calls each node's agent module by `mesh/ask`. If the runtime's half of ADR 0198 is not
yet live on the control node when this package starts, this package waits for it: no tool container, no
credential copied by hand.
**Order.** Assign the manager on the control node; push. Adopt the API key from a file there. Adopt the
@@ -0,0 +1,11 @@
# 211 — Diagnosis
*2026-10-03.* The planner orders a merge's modules by `inventory.Dependencies`, whose edges come from a
manifest's `build.on`, from what a build recorded it stood on, from the repositories it read, and from
the build machine. A bundle names its toolchain by `language`; the builder takes the toolchain image
(`ToolchainFor(language)`) from what the mesh holds and records nothing of it as stood on. So no edge
ran from a bundle to the module publishing its toolchain, and a merge moving both (mesh-tools: the
images and node-tools) tiered them together. **Fix (mesh-controller, branch
`fix/issue-211-a-bundle-stands-on-its-toolchain`, commit c72f6ca):** `dependenciesOf` adds a `stands-on`
edge from every bundle artifact to its toolchain's module, read from the manifest. Tested: TypeScript
bundle → mesh-tools, Go bundle → mesh-tools-go, image → none; a merge moving both plans two tiers.
@@ -0,0 +1,44 @@
---
status: located
opened: 2026-10-03
located-in:
- mesh-tools
- mesh-controller
fixed-by:
amended-design:
---
# 212 — A toolchain rebuild keeps the SDK it cached, and every bundle built on it carries the old one
## What was observed
2026-10-03, rolling out [ADR 0193](../../02-DECISIONS/0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md).
SDK 0.1.5 was published, then the TypeScript toolchain image was rebuilt, then node-tools and six
bundles were rebuilt on it and pushed. On every machine the bundles carried SDK 0.1.3:
```
bundle SDK: "version": "0.1.3"
```
The launched packet filter and intrusion prevention, which register their seat first, then served
their seat's verbs as their own tools, and `node-packet-filter.*` and `node-intrusion-prevention.*`
stopped answering on three machines until fixed.
## Why it matters beyond this instance
The toolchain image installs its dependencies from a `package.json` that names the SDK by range.
The file did not change, so the image build reused its cached install layer, and "rebuilt after
the release" did not mean "carries the release". Every TypeScript bundle copies its dependencies
from that image ([design 38](../../03-DESIGN/01-to-be/38-building-the-operators-machine.md) WP3),
so an SDK release reaches no bundle until something else happens to change that file. Nothing
reports it: the build succeeds and records the new commit.
## Diagnosis
Owners mesh-tools (the image's recipe) and mesh-controller (the builder that runs it). **Worked
around** on the day by requiring `^0.1.5` in node-tools' `package.json` (mesh-tools #37), which
changes the layer. **Fix direction:** a toolchain build must not trust a cached dependency install —
install from a lockfile that a release updates, or build the install layer without the cache — and
the build should record which SDK version the image carries, so a bundle's record says what it was
compiled against. A check: after an SDK release and a toolchain rebuild, a bundle built on it reports
the released version.
@@ -0,0 +1,52 @@
---
status: open
opened: 2026-10-03
located-in:
- mesh-controller
- mesh-catalog
fixed-by:
amended-design:
---
# 213 — The controller is a Go program, and it still runs in a container
## What was observed
2026-10-03. The controller — the program that holds the `mesh-controller` seat and answers its
verbs (`status`, `nodes`, `push`, `assign` …), composes every machine's declaration and plans the
builds — runs on its machine as a container built from an image:
```
mesh-controller Up … (docker ps on the machine that runs it)
```
It is written in Go and compiles to one static binary, as the node host does. The host is delivered
as a bundle and run as a process; the node's tool runtime now is too
([ADR 0193](../../02-DECISIONS/0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md)).
The controller is the one piece of the mesh's own Go code still shipped as an image.
## Why it matters beyond this instance
[ADR 0188](../../02-DECISIONS/0188-a-modules-own-code-is-bundles-in-any-language-and-a-tools-bundle-speaks-mcp-to-the-runtime.md)
§1 says a module's own code is bundles, never an image, and §3 that a bundle that is a service is a
`process` the host runs. The controller breaks the rule it is the mechanism of: the registration
gate that will refuse an image of a module's own code has to exempt the controller, or refuse it.
It also costs what an image costs — a container runtime on its machine as a hard requirement,
eight mounts standing in for files a process would simply read, an image rebuild for a binary
change — and
every restart of it is a container recreation, which is how the controller restarts in the middle
of a plan today.
## What a fix has to settle
- The controller's module declares a Go bundle (`system`, `binary`) and a `process` running it
(`./<binary>`, mesh-host #81), with its credential and store connection as files and words, not
container mounts and a container network name.
- What the container gives it now that a process would not: it runs on the host's network already,
as an unprivileged user (65534), with eight mounts. Each mount named and replaced by a path, and
the user by an account the host declares.
- The handover: the controller restarting itself as a process, on the one machine that runs it,
without a window where nothing answers the mesh's verbs.
Located only by owner; the move is a change of the controller's module and its deployment, not of
its code.
@@ -0,0 +1,36 @@
---
status: located
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by:
amended-design:
---
# 214 — A plan loses track of the controller it rebuilds, and waits on it for ever
## What was observed
2026-10-03, a merge to the controller's repository planned three tiers: the controller itself, then
the build agent, then the route proxy. The controller's build finished and was recorded at 19:24:
```
mesh-controller built 2ebbb799 g14 2026-10-03 19:24
```
Twenty-seven minutes later the plan still read `tier 1 of 3 … mesh-controller asked`, and every later
plan waited behind it. It was stopped by hand; the later tiers were never asked.
## Why it matters beyond this instance
The controller rebuilding itself is the one plan whose first tier replaces the process that runs the
plan. The new controller starts with the plan's state as stored, and the outcome of the build that
produced it arrived to the old one, or between the two. Every merge to the controller's own repository
can end this way, and each one blocks every plan after it until somebody notices.
## Diagnosis
Owner mesh-controller (the planner). **Fix direction:** on start, and whenever a plan waits on a
build, the plan settles an `asked` build against the build records — a build recorded as built from
the plan's commit is that tier's outcome — so a plan resumes after the controller replaced itself.
A test: a plan whose build outcome was recorded while no controller followed it resumes on start.
@@ -0,0 +1,10 @@
# 214 — Diagnosis
*2026-10-03.* A plan learns a tier's outcome only through `planBuilt`, called when a controller takes
in a build result off the bus. A merge to the controller's repository replaces the controller in tier
0; the build that produced the new controller was recorded, but the plan state the new controller
loaded still read `asked`, and no path ever revisited it. **Fix (branch
`fix/issue-214-a-plan-settles-from-the-build-records`, commit d86baeb):** `advanceOnce` settles every
still-asked module from its build records — a build recorded after the ask is that ask's outcome,
built or failed — on every advance and on the 30-second ticker. Tested with a pure helper. Live
proof: the next merge to mesh-controller passes tier 0 on its own.
@@ -0,0 +1,35 @@
---
status: located
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by:
amended-design:
---
# 215 — A module built once at a commit stops following its branch, and merges leave it out
## What was observed
2026-10-03, a catalogue merge changed seven modules. Its plan built six; `unifi` was not in it,
though its change was in the same commit. A second merge the same evening left it out again. Its
build records name a commit where every other module names a branch:
```
unifi built 9c97a8a1 … http://…/mesh-catalog.git at 9c97a8a
```
Built by hand from `main` it was registered again and the change reached its machine.
## Why it matters beyond this instance
A module that was once built at a commit — to pin it during a fix, or by a build asked with a `ref`
— silently stops following its branch: merges plan without it, `status` does not say it is behind its
branch, and nothing says it is pinned. The operator learns it when a change does not arrive.
## Diagnosis
Owner mesh-controller. **Fix direction:** a build asked at a commit does not change the branch a
module follows; or, if pinning is meant, the pin is said — in `module list`, in `status`, and by a
merge's plan naming the module it leaves out and why. A test: building a module at a commit and then
merging a change to it plans it.
@@ -0,0 +1,10 @@
# 215 — Diagnosis
*2026-10-03.* `takeIn` registers a build's `Ref` as the branch the module follows. unifi was once
built with `ref=9c97a8a`, which became its followed ref. `sourceIs` matches a merge only to modules
whose ref is empty or the merged base — so every merge into main left unifi out — and `askTier`
re-asks `Source.Ref`, so every plan that rebuilt unifi built the same old commit again (its build
records all read "at 9c97a8a"). **Fix (branch `fix/issue-215-a-commit-is-never-a-branch-to-follow`,
commit 6784efa):** registration keeps the followed branch when a build names a commit; matching and
re-asking read a recorded commit as the default branch, healing existing records; a merge names the
modules of its repository it leaves out. Store-backed test fails without the fix.
@@ -0,0 +1,30 @@
---
status: located
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by:
amended-design:
---
# 216 — A tools bundle nothing says to load is built, recorded, and never delivered
## What was observed
2026-10-03, seven modules moved their tools from a container to a bundle. Each built, each build was
recorded, and the machines that run them reported current — with none of the bundles on them. The
composer delivers a bundle only when the runtime loads something from it, and that is derived from
the module's `tools` list when the artifact names no `loads`; these modules had neither.
## Why it matters beyond this instance
Every step reported success: the build, the registration, the push, the machine's apply. The tools
were simply absent, and the old container was gone. A rule the composer applies silently is a rule
nobody learns until the tools are missing.
## Diagnosis
Owner mesh-controller (the catalogue's registration check). **Fix direction:** a bundle that nothing
loads, runs or unpacks — no `loads`, no `tools` list on its module, no resource naming it — is refused
at registration, naming the field that would deliver it. A test: such a manifest is refused; adding
`loads` admits it.
@@ -0,0 +1,9 @@
# 216 — Diagnosis
*2026-10-03.* The composer delivers a bundle as an archive only when its `Loads` is non-empty, and
`Loads` derives from the artifact's `loads` or, failing that, from the module's `tools` list. The
seven modules had neither, so their bundles were recorded and never composed into any declaration;
nothing checked it. **Fix (branch `fix/issue-216-a-bundle-nothing-delivers-is-refused`, commit
cf2bb3b):** registration refuses a bundle that nothing loads, runs or unpacks — no `loads`, no `tools`,
no resource naming it, and not the runtime — naming the field that would deliver it. The current
catalogue passes the check.
@@ -0,0 +1,52 @@
---
status: located
opened: 2026-10-03
located-in:
- mesh-tools
fixed-by:
amended-design:
---
# 217 — A refused announcement took down every container's runtime, and the console with it
## What was observed
2026-10-03, rolling out [ADR 0197](../../02-DECISIONS/0197-every-tool-announces-itself-on-the-bus-in-the-nats-services-protocol.md).
The tool runtimes announced themselves by subscribing `$SRV.<verb>.>`; the grants composed for them
allowed only `$SRV.<verb>` and the service's own name and instance. The bus refused the wildcard:
```
Subscription Violation - User "novox.gitea", Subject "$SRV.PING.>"
NatsError: 'Permissions Violation for Subscription to "$SRV.PING.>"'
```
The TypeScript runtime the per-module containers run treats a refused subscription as fatal, so on
the one machine that had received the new images nine containers crash-looped — gitea's runtime,
postgres's, the catalogue, the vault, mongodb, mssql, keycloak, mailu, nextcloud. Gitea's tools went
with them, which closed the usual path for merging the fix.
Then the console stopped answering the controller's verbs, though the controller held every
subscription: the console builds the list that tells a seat's verb from a module's tool by asking the
controller *and* the catalogue, and gives up on both when the catalogue does not answer — so
`mesh-controller.status` was asked of a module subject nobody serves.
## Why it matters beyond this instance
Two properties, each worse than the mistake that exposed it:
- **A runtime dies for an optional subscription.** Announcing is discovery; serving tools and running
provisioners is the work. A refusal of the first should never stop the second.
- **The console's view of the mesh's own verbs depended on a module.** The controller's verbs are how
the operator repairs the mesh; they must not become unreachable because the catalogue is down.
## Diagnosis
Owner mesh-tools. The wildcard is fixed on branch `fix/announce-only-what-the-grants-allow`: both
runtimes subscribe exactly what the grants allow. The console that discovers from what announces itself
(ADR 0195, 0197, on main) asks the bus and the controller, not the catalogue, which removes the second
property once it is deployed. **Still to do:** the TypeScript runtime treats a refused announcement
subscription as a logged warning, not as fatal; a test against a bus with real grants proves the
announcement subscriptions are allowed for every principal kind.
Recovered on the day without the forge's API: the toolchain images built by the controller straight
from the fix branch, every module image rebuilt on them, and the machine pushed.
@@ -0,0 +1,65 @@
---
status: located
opened: 2026-10-03
located-in:
- mesh-controller
fixed-by: mesh-controller#248
amended-design:
---
# 218 — A seat held once for the mesh is answered by a module on a machine that does not hold it
## What was observed
2026-10-03. Asked which databases the mesh's store holds, `mesh-store.databases` answered from the
postgres on one machine with that machine's application databases; the controller's own database
lives on the postgres of the other machine, which the controller's records name as the seat's one
holder:
```
mesh-store scope: mesh delivers: postgres-database holders: [ {node: <the control machine>, module: postgres} ]
```
The discovery console, which reads what answers on the bus ([ADR 0197](../../02-DECISIONS/0197-every-tool-announces-itself-on-the-bus-in-the-nats-services-protocol.md)),
shows the same seat announced from **both** machines running postgres.
## Why it matters beyond this instance
A seat held once for the mesh promises one answerer: the role's holder. A module that implements a
seat's verbs on every machine it runs on, and is let serve them on each, turns "the mesh's store" into
"whichever postgres replied first" — a read against the wrong database that looks like a right one,
and a write would be worse. Every mesh-scoped seat whose implementing module runs on more than one
machine has this shape.
## Where to look
Whether the runtime serves a seat's verbs where its module merely *claims* the seat rather than where
the mesh made it the holder ([ADR 0159](../../02-DECISIONS/0159-a-tool-call-names-the-machine-and-a-holder-serves-its-seats-verbs.md),
[ADR 0160](../../02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)):
the membership the controller issues each assignment, and what the runtime admits from it. A check:
a mesh-scoped seat's verbs are served by exactly the holder the records name, on every machine.
## Root cause
The controller composed each assignment's held seats from what its module *claims*, once per module
and not once per machine. Every machine running postgres was therefore given the store seat's grants
and issued its subjects, and each runtime served the seat's verbs because it serves what it is issued
([ADR 0160](../../02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)).
The runtime behaved as designed. The fault was in what it was issued.
The seat's verbs are not the module's tools. The store's `databases` and `query` are a separate
implementation registered under the seat's name ([ADR 0159](../../02-DECISIONS/0159-a-tool-call-names-the-machine-and-a-holder-serves-its-seats-verbs.md)).
Only that implementation should be withdrawn where the module does not hold the seat. postgres's own
tools stay served on every machine it runs on.
## Fix
The controller now reads the recorded seat holdings when it composes grants and memberships. A seat
held once for the mesh is issued only to the machine and module the records name as its holder. A
seat held once per machine, and a mesh seat with no holder on record, are issued as before. Grants
and memberships come from the same list, so they cannot disagree.
**How it is checked.** A controller test asserts that a claimant on another machine keeps its node
seats and loses the recorded mesh seat. Live, the discovery console's overview must show each
mesh-scoped seat announced from exactly the holder the records name. Status moves to `resolved` once
that holds after the fix is rolled out.