1 Commits
Author SHA1 Message Date
jschoubben d222091fe9 Issues 245, 246; 240 resolved, 242 located
mesh/merge-gate the mesh is checking this head against every machine
mesh/delivery checking
245: a media server's previews were reached through a link its container never mounted — fixed by
mounting them. 246: settings set replaces the whole layer with nothing to read it first and no
history. 240 is fixed by mesh-controller#47; 242 has backups running and records what the rollout
taught.
2026-10-05 14:54:38 +02:00
139 changed files with 208 additions and 13223 deletions
+1 -4
View File
@@ -56,7 +56,4 @@ a to-be design names a decision, an in-progress/implemented design names its own
located/fixed issue names its owner, a fixed/resolved issue says what fixed it, a graduated located/fixed issue names its owner, a fixed/resolved issue says what fixed it, a graduated
research overview says what it became, and no two issue records share a number (issue 155 — the research overview says what it became, and no two issue records share a number (issue 155 — the
number is how a record is cited, and `main` lags every open pull request, so two people reading it number is how a record is cited, and `main` lags every open pull request, so two people reading it
allocate the same one). And a core issue — one opened from 2026-10-07 whose `located-in` names a core allocate the same one). `python3 00-META/checks/cycle.py`
repository — resolves only with `replay:` (an id in mesh-lab's replays register) or `replay-none:` saying
why none is possible ([ADR 0237](../../02-DECISIONS/0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md));
it failed on a resolved core issue carrying neither before it passed. `python3 00-META/checks/cycle.py`
-25
View File
@@ -16,11 +16,6 @@ What is enforced:
once `resolved`, `fixed-by:` says what fixed it (prose counts -- once `resolved`, `fixed-by:` says what fixed it (prose counts --
"nothing, the capability existed" is an answer). And no two records share a "nothing, the capability existed" is an answer). And no two records share a
number -- the number is how a record is cited. number -- the number is how a record is cited.
replays a core issue -- one opened from 2026-10-07 whose `located-in:` names a core repository --
resolves only with `replay:` (the replay's id in mesh-lab's replays register) or
`replay-none:` (why no replay is possible): ADR 0237, to-be 45 §9. A replay is what
proves a fix fails before and passes after; without one, a fixed class comes back
through a different door (research 031: four such chains in one week).
research a known `status:`; a `graduated` overview says what it `became:`, and every research a known `status:`; a `graduated` overview says what it `became:`, and every
target it names exists. target it names exists.
decisions every accepted record is REACHABLE from the cycle: cited by a design doc's decisions every accepted record is REACHABLE from the cycle: cited by a design doc's
@@ -46,16 +41,6 @@ DESIGN_STATUSES = {"proposed", "designed", "in-progress", "implemented", "abando
ISSUE_STATUSES = {"open", "diagnosing", "located", "resolved", "wontfix"} ISSUE_STATUSES = {"open", "diagnosing", "located", "resolved", "wontfix"}
RESEARCH_STATUSES = {"active", "graduated", "abandoned"} RESEARCH_STATUSES = {"active", "graduated", "abandoned"}
# The core, as to-be 45 names it, by where its code lives: the controller, the node-engine, the node
# tools and the console, the SDK's loops, and the catalogue's bus, forge (the announcer of merges) and
# build agent. A `located-in:` entry naming any of these makes an issue a core issue.
CORE = re.compile(r"\b(mesh-controller|mesh-host|mesh-tools|mesh-sdk|node-tools|"
r"mesh-catalog\s+modules/(nats|gitea|build-agent))\b")
# The day the rule began (ADR 0237): issues opened before it are not held to it.
REPLAYS_FROM = "2026-10-07"
REPLAY_ID = re.compile(r"^R\d+\b")
def rel(path): def rel(path):
return os.path.relpath(path, ROOT) return os.path.relpath(path, ROOT)
@@ -174,16 +159,6 @@ def main():
bad(path, "status %s but located-in is empty" % status) bad(path, "status %s but located-in is empty" % status)
if status == "resolved" and not listy(front, "fixed-by"): if status == "resolved" and not listy(front, "fixed-by"):
bad(path, "status %s but fixed-by says nothing" % status) bad(path, "status %s but fixed-by says nothing" % status)
# A core issue resolves with its replay, or with why none is possible (ADR 0237).
replay = " ".join(listy(front, "replay"))
if replay and not REPLAY_ID.match(replay):
bad(path, "replay: %r names no replay -- an id from mesh-lab's replays register, R<issue>" % replay)
core = any(CORE.search(entry) for entry in listy(front, "located-in"))
opened = str(front.get("opened") or "")
if (status == "resolved" and core and opened >= REPLAYS_FROM and not replay
and not " ".join(listy(front, "replay-none")).strip()):
bad(path, "a core issue resolved with no replay: name it in `replay:` (mesh-lab replays "
"register.go) or say in `replay-none:` why none is possible (ADR 0237)")
# ---- research ---------------------------------------------------------------------- # ---- research ----------------------------------------------------------------------
for path in sorted(glob.glob(os.path.join(ROOT, "01-RESEARCH", "*", "00-overview.md"))): for path in sorted(glob.glob(os.path.join(ROOT, "01-RESEARCH", "*", "00-overview.md"))):
+2 -51
View File
@@ -7,7 +7,7 @@ another — and a mesh you cannot name precisely is a mesh two people describe d
## The mesh and its machines ## The mesh and its machines
- **node** — a machine in the mesh. There are 0..n of them, and each runs the node-engine. A node is - **node** — a machine in the mesh. There are 0..n of them, and each runs the host agent. A node is
just a machine that has joined; being one implies nothing about what it runs. just a machine that has joined; being one implies nothing about what it runs.
- **operator account** — the login name of the person who works on a node, stated on the node - **operator account** — the login name of the person who works on a node, stated on the node
record; empty for a machine nobody logs into. Everything the mesh places under a person's home is record; empty for a machine nobody logs into. Everything the mesh places under a person's home is
@@ -28,13 +28,6 @@ another — and a mesh you cannot name precisely is a mesh two people describe d
control-plane/data-plane, and opaque here). control-plane/data-plane, and opaque here).
- **mesh-controller** — the module that runs the controller. It **claims** the `mesh-controller` - **mesh-controller** — the module that runs the controller. It **claims** the `mesh-controller`
seat at mesh scope, which is what makes it singular. Replaces the module name **`mesh-control`**. seat at mesh scope, which is what makes it singular. Replaces the module name **`mesh-control`**.
- **node-engine** — the program on every node that applies what the controller declares: it receives the
node's declaration, writes the files, runs the services and containers, and reports what it did. It
is the engine, not a module: it owns no file's content, and every file it writes belongs to the module
that declared it. Replaces **"host agent"**, **"the host"** and **`mesh-host`**. *Agent* is avoided
because the word already means two other things here, the build agent and the operator's coding
agent. The code still carries the old name (the `mesh-host` repository, its binary and service
unit) until the rename is made there; a record written before the rename keeps the old name.
(The git repository has been renamed `mesh-control` -> `mesh-controller` on the forge; the module, (The git repository has been renamed `mesh-control` -> `mesh-controller` on the forge; the module,
container and image it produces are `mesh-controller`.) container and image it produces are `mesh-controller`.)
- **foundation** — the store and the broker, raised at genesis before any module system exists. - **foundation** — the store and the broker, raised at genesis before any module system exists.
@@ -80,26 +73,7 @@ another — and a mesh you cannot name precisely is a mesh two people describe d
The set, with who holds each seat, is the overview of what a mesh has The set, with who holds each seat, is the overview of what a mesh has
([26 — The seats](../03-DESIGN/01-to-be/26-the-seats.md)). A seat has a **capacity**: a ([26 — The seats](../03-DESIGN/01-to-be/26-the-seats.md)). A seat has a **capacity**: a
capacity-1 seat is exclusive (one holder); a higher-capacity seat is a **bench** (several holders capacity-1 seat is exclusive (one holder); a higher-capacity seat is a **bench** (several holders
coexist). The first bench is `mesh-dns-resolver`, a *replicated* mesh seat: one holder per machine, coexist).
each on record and each answering the same names ([ADR 0223](../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md)).
The second sort is the **kinded** bench: holders are different modules, each claiming one **kind**,
and a verb's subject carries the kind. `channel` and `intake` are the only ones
([ADR 0234](../02-DECISIONS/0234-the-mesh-holds-a-conversation-with-its-operator.md)).
- **channel / intake** — the two kinded benches the mesh talks to its operator through: `channel`
sends, `intake` turns what arrives into one envelope. A holder of either declares **capabilities**
from the fixed vocabulary `channel-capabilities/1`. Not "notifier" (that is the desktop's node seat,
one holder of kind `desktop`) and not "bot" (that is one service's account).
- **ask** — a request for the operator's input, of a declared kind (yes-no, one-of, text, number, date,
acknowledge). An **authorising ask** is one whose answer performs an action; the controller holds it
and checks its **proofs** (a verified sender, a TOTP code, optionally a security key's touch)
([ADR 0234](../02-DECISIONS/0234-the-mesh-holds-a-conversation-with-its-operator.md)).
- **operator message / input** — what arrives on `intake`, once the router has checked the sender
against the controller's list of the operator's identities: a **trusted** one is an operator message,
addressed to an agent by `@name` or thread or else to the **responder** (`@mesh`, the router's own
participant, which answers from read verbs); anything else is **untrusted input** — data, never
instructions ([ADR 0234](../02-DECISIONS/0234-the-mesh-holds-a-conversation-with-its-operator.md)).
- **reference** — an opaque token in a message's words standing for a detail the content rule keeps out
of them (a path, an address); opened with `detail` only on a `private` channel or at the console.
- **claim** — a module taking a spot on a seat. `claims: [{name, scope}]` in a manifest. A - **claim** — a module taking a spot on a seat. `claims: [{name, scope}]` in a manifest. A
mesh-scoped exclusive claim is how the mesh says "there is one of me". A foundation seat is mesh-scoped exclusive claim is how the mesh says "there is one of me". A foundation seat is
named after the server it guards: the `mesh-controller`, `postgres` and `lavinmq` modules claim named after the server it guards: the `mesh-controller`, `postgres` and `lavinmq` modules claim
@@ -127,29 +101,6 @@ another — and a mesh you cannot name precisely is a mesh two people describe d
- **invokes** — the manifest word for the tools a module calls, `<module>.<tool>` each or `*` for - **invokes** — the manifest word for the tools a module calls, `<module>.<tool>` each or `*` for
every one. A grant on the publish side and nothing else; a module that declares none calls nothing. every one. A grant on the publish side and nothing else; a module that declares none calls nothing.
## How a change reaches the machines
- **delivery** — one commit in one repository on its way to the machines, from its pull request's head being
announced to delivered, failed, superseded or stopped: one pull request, one status, one note
([ADR 0239](../02-DECISIONS/0239-a-delivery-is-owned-by-the-mesh-delivery-module-and-runs-from-commit-to-delivered.md)).
Its states are one table (`proposed`, `checked`, `ready` or `rejected`, `published`, `delivering`, `held`,
and the final four). Replaces **release plan** as a concept: the controller's plan is now the *walk* of one
delivery's trunk commit across machines, its record and nothing more. Not "change" (a diff), not
"pipeline" (to-be 10 retires it), not "deployment" (one stage of it).
- **delivery group** — two or more deliveries sharing a pull request head branch name across the mesh's
repositories, delivered as one unit in an order declared (`after:`) or inferred by the planner. One level:
a group holds deliveries, never groups. Its state is derived from its members, never set. A lone delivery
needs none.
- **delivery plan** — what a delivery does to the mesh: its build plan (the modules it moves and their
dependents, in tiers), its deploy plan (per machine, what it receives and what waits for a person) and its
verdict (the composed machines, the replays). Computed by the controller's planner from a diffset. ADR 0238
called it the **change plan**; that word is retired.
- **mesh-delivery** — the module that owns deliveries and delivery groups, holding the mesh-scoped seat of the
same name. It records and decides; the controller sends, judges and rolls back when it is asked.
- **walk** — the controller's sending of one trunk commit's builds across machines, tier by tier, one machine
first and judged at the gate (ADR 0236). A primitive the delivering stage asks for, not an object anyone
manages.
## How this page is kept ## How this page is kept
A new name for an existing thing lands here first, in the same change that introduces it in code. A A new name for an existing thing lands here first, in the same change that introduces it in code. A
+1 -5
View File
@@ -42,11 +42,7 @@ incident someone must **clear**.
2. Investigate in `01-diagnosis.md` in the same folder — the trail, dated, including what was 2. Investigate in `01-diagnosis.md` in the same folder — the trail, dated, including what was
ruled out. Move `status:` to `diagnosing`, then `located` once the owner is known. ruled out. Move `status:` to `diagnosing`, then `located` once the owner is known.
3. Resolve. Set `status: resolved`, fill `fixed-by:`, and if the root cause was a design gap, 3. Resolve. Set `status: resolved`, fill `fixed-by:`, and if the root cause was a design gap,
run playbook [02](02-graduation.md) and fill `amended-design:`. **A core issue** — its `located-in` run playbook [02](02-graduation.md) and fill `amended-design:`.
names the controller, the node-engine, the node tools, the SDK, or the catalogue's bus, forge or build
agent — resolves with `replay:`, the id of its replay in mesh-lab's replays register, proved to fail on
the commit before the fix and pass on it, or with `replay-none:` saying why none is possible
([ADR 0237](../../02-DECISIONS/0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md); `cycle.py` checks it).
## Rules ## Rules
@@ -1,5 +1,5 @@
--- ---
status: graduated status: active
initiated: 2026-10-04 initiated: 2026-10-04
touches: touches:
- 04-ISSUES/187-the-mesh-tells-nobody-when-it-stops-working/00-report.md - 04-ISSUES/187-the-mesh-tells-nobody-when-it-stops-working/00-report.md
@@ -10,17 +10,7 @@ touches:
- 02-DECISIONS/0208-the-graphical-session-is-one-module-per-piece-on-the-meshs-seats.md - 02-DECISIONS/0208-the-graphical-session-is-one-module-per-piece-on-the-meshs-seats.md
- 02-DECISIONS/0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md - 02-DECISIONS/0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md
- 03-DESIGN/01-to-be/32-what-a-module-declares.md - 03-DESIGN/01-to-be/32-what-a-module-declares.md
- 03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md became: []
- 02-DECISIONS/0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md
- 02-DECISIONS/0212-a-seat-says-what-it-receives-and-the-machines-hotkeys-are-a-seat.md
- 02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md
- 02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md
- 02-DECISIONS/0228-a-value-given-by-hand-lives-only-until-its-modules-first-good-start.md
- 02-DECISIONS/0230-a-consumer-the-mesh-stops-asking-for-is-retired-and-deleted-only-by-a-person.md
- 02-DECISIONS/0232-a-binding-to-a-consumers-data-moves-only-by-a-person.md
became:
- 02-DECISIONS/0234-the-mesh-holds-a-conversation-with-its-operator.md
- 03-DESIGN/01-to-be/46-the-conversation-with-the-operator.md
--- ---
# 028 — The mesh's output channel # 028 — The mesh's output channel
@@ -42,34 +32,7 @@ The effort looks at:
- **the life of a message:** deduplicated while it holds, resolved when it stops, acknowledged or - **the life of a message:** deduplicated while it holds, resolved when it stops, acknowledged or
silenced by the operator; silenced by the operator;
- **the watcher's watcher:** who tells the operator when the parts that would tell them are the - **the watcher's watcher:** who tells the operator when the parts that would tell them are the
ones that failed; ones that failed.
- **the conversation** (widened 2026-10-06): the mesh, its modules and its agents send messages and
**asks** (a question, a choice, a value, an approval), and the operator answers or writes first, over
channels chosen by their declared **capabilities** and by the operator's **work context**. Inputs
from outside (a mail arriving) share the same envelope;
- **asks that authorise:** one layer on top, for the answers that perform an action. These are checked
by the controller, and allowed only on channels whose capabilities prove that the operator answered.
### Why this effort widened rather than a new one opened
On 2026-10-06 the operator asked for these:
- approving and rejecting through Telegram;
- a generic shape in which Telegram simply holds a seat;
- input triggers alongside output channels;
- capabilities that decide which actions may travel on which channel;
- the work context as a factor in choosing the channel;
- asks that are not about permission at all.
That could have opened a new effort. It did not, because:
- every part of it hangs on this effort's open questions: Q1 (where a channel attaches), Q4
(presence), Q7 (answering back) and Q8 (what may leave);
- ADR 0227 kept answering back open **here**, and said this effort's graduation amends to-be 45 §5;
- an answer belongs to the message or ask this seat sends. Splitting them would leave two efforts each
owning half of one conversation.
Input that is not an answer (a mail arriving, a webhook) shares the envelope and the seat shape, and
is designed here only as far as that shape. Its consumers are later work.
## Why ## Why
@@ -106,19 +69,3 @@ for the mesh, sources that call it, and channels that deliver.
2. [The channels](02-the-channels.md): the candidates, Telegram first, weighed on the same questions. 2. [The channels](02-the-channels.md): the candidates, Telegram first, weighed on the same questions.
3. [Open questions](03-open-questions.md): the seat, routing, life of a message, the watcher's 3. [Open questions](03-open-questions.md): the seat, routing, life of a message, the watcher's
watcher, what may leave the mesh. watcher, what may leave the mesh.
4. [Telegram, as the first holder](04-telegram-as-the-first-holder.md): making the bot, the bot
API's limits and semantics, what Telegram sees, and ten defects in the built code.
5. [The other holders, on the same axes](05-the-other-holders-on-the-same-axes.md): ntfy, Matrix,
Pushover, Gotify, mail, Signal, SMS and a dead-man service, as away channel and as the watcher's
path.
6. [A conversation with the operator](06-a-conversation-with-the-operator.md): messages, asks and
operator messages; asks' kinds and life; the kinded benches `channel` and `intake`; the capability
vocabulary; agents as participants; the migration.
7. [The work context, and the desk](07-the-work-context-and-the-desk.md): the signals, the routing
by context and its escalation, presence kept inside the mesh, and the desk as a full participant
(notification actions, the launcher's prompt).
8. [Asks that authorise](08-asks-that-authorise.md): the actions that need a person, the trust
capabilities, the three proofs and the three tiers, the controller's checks, why the desk needs a
factor, and what a compromise can reach.
9. [A proposed decision](09-a-proposed-decision.md): the recommendation, the operator's steps, the
tables, and a record ready for graduation.
@@ -1,7 +1,6 @@
# 03 — Open questions # 03 — Open questions
Each question names the options seen so far. None is decided here. Q1 and Q7 are taken further in Each question names the options seen so far. None is decided here.
[06](06-a-conversation-with-the-operator.md) and [08](08-asks-that-authorise.md).
## Q1. The seat ## Q1. The seat
@@ -1,222 +0,0 @@
# 04 — Telegram, as the first holder of a channel
Telegram was the operator's first required channel ([02](02-the-channels.md)) and it is built:
the output seat's holder carries a Telegram client, and so does the watcher's watcher (to-be 45 §5).
Neither is configured, because no bot exists yet. This document is what the operator needs to make one,
what the mesh's use of the bot API must respect, and what the built code gets wrong against it.
[06](06-a-conversation-with-the-operator.md) makes Telegram one holder of a generic channel seat rather
than the subject of the design. Everything here stays true under that shape: it is the first holder's
analysis.
Facts are as of 2026-10-06, Bot API 10.3 (2026-08-24). Sources are listed at the end.
## What the mesh uses from Telegram
Two programs send, and neither reads anything back yet:
- **The output seat's holder**, on the control node. It sends a message when a condition is raised,
says it again as a reminder, and edits the first message in place when the condition clears.
- **The watcher's watcher**, on a machine that is not the control node. It sends straight to the bot
API over HTTPS when the controller's self-check or the bus has been silent past its bound.
Each holds a **bot token** as its own secret, issued outside the mesh (ADR 0228, `issued-by: outside`),
and a **chat id** as a setting. Each sends plain text: no `parse_mode`, so no markup to escape and none
to inject.
## Making the bot
Telegram has no developer console. A bot is made by talking to Telegram's own bot, BotFather, from an
ordinary Telegram account.
- **An account is required, and an account needs a phone number.** There is no other sign-up. The
number can be a virtual one bought on Telegram's own marketplace, at a price that makes it
irrelevant here.
- **`/newbot`** asks for a display name and a username. The username is 5–32 characters of Latin
letters, digits and underscores, must end in `bot`, and cannot be changed later.
- BotFather answers with the **token**: digits, a colon, then a key. Anyone holding it controls the bot.
- **`/token`** issues a new token for the bot. The old one stops working at once. This is the rotation
path, and it is the only one.
- **`/setjoingroups` → Disable** stops anyone adding the bot to a group. The mesh's bot talks to one
person; a group is only a way for someone else to see what it says.
- **Privacy mode** (`/setprivacy`) governs what a bot sees **in groups**: with it on, only commands
meant for it, replies to it and service messages. In a private chat a bot sees everything the person
writes. With groups disabled, privacy mode does not matter; leave it on.
### A bot cannot speak first
A bot cannot open a conversation. Until the person presses **Start** in the bot's chat, every send to
them fails with a "Forbidden" error. So the operator presses Start once, on each bot.
### Finding the chat id, safely
In a private chat the chat id equals the person's user id. There are two ways to learn it:
- **Read `getUpdates` once by hand.** After pressing Start, a call to `getUpdates` returns the `/start`
message with the chat's id. It works today. Its two weaknesses: the token appears in a command line
(and so in a shell's history) unless read from a file, and it trusts that the `/start` it sees is the
operator's. A bot's username is public, and anyone who finds it can press Start too.
- **A linking verb with a one-time code.** The holder makes a short code and answers with a deep link
(`https://t.me/<bot>?start=<code>`). The operator opens it on the phone; Telegram sends `/start <code>`.
The holder reads it by `getUpdates`, and binds **that** chat and **that** user id only if the code
matches and is fresh. This proves the chat belongs to whoever held the code, and it never shows the
token to anyone. It needs the holder to read updates, which approvals need anyway
([08](08-asks-that-authorise.md)).
The second is the one to build. The first is the stop-gap until it exists.
### One bot or two
The holder and the watcher each have their own secret. They can hold the same token or two.
**Two bots are better:**
- Revoking one does not silence the other. The watcher exists for the day the rest is broken, and
that day must not also be the day its token was rotated away.
- The phone shows which program spoke.
- **Only one program may read a bot's updates.** Two concurrent `getUpdates` callers on one token make
Telegram answer the older with HTTP 409, "terminated by other getUpdates request". The moment the
holder reads answers, the watcher could no longer share its token with anything that reads.
## The bot API, as the mesh uses it
### Limits
- **Rate.** Telegram's FAQ: "In a single chat, avoid sending more than one message per second." In a
group, 20 messages a minute. Broadcast across chats: about 30 a second. The holder's own cap is 20 an
hour, so the limit is never near.
- **Over the limit** the API answers HTTP 429 with `parameters.retry_after`, the seconds to wait
before the request may be repeated. Repeating early prolongs the wait.
- **Length.** A text message is at most 4096 characters after entity parsing. Longer is refused with
HTTP 400, not cut.
- **Callback data** on a button is 1–64 bytes ([08](08-asks-that-authorise.md)).
### Editing
- **`editMessageText`** replaces a sent message's text. For an ordinary bot message there is no time
limit. The 48-hour limit applies only to business messages, and deletion has its own 48-hour limit.
- **An edit notifies nobody.** No sound, no banner, and the message stays where it was in the chat's
history. This is why a clearing is cheap to say by edit. It is also why an edit must never be the
only way something **new** is said.
- **An identical edit is an error:** HTTP 400, "message is not modified". It is harmless and must be
read as success, not as a failure to fall back from.
- **A deleted message** answers "message to edit not found". Falling back to a new message is right
then.
### Loudness
Telegram has **no message priority**. The only lever is **`disable_notification`**: the message
arrives without sound. It cannot break through the phone's do-not-disturb, so an urgent message at
night is as quiet as the phone is set to be. Per-chat notification settings on the phone (a custom
sound, an exception to do-not-disturb) are the operator's, not the mesh's.
### Formatting
With no `parse_mode` the text is shown as written, and nothing in a message can be read as markup.
That is the right default for words that pass a content rule rather than a template. If markup is ever
wanted, `MarkdownV2` needs every reserved character escaped and fails the whole send on one miss, so
plain text or `HTML` with escaping are the safer options.
`disable_web_page_preview` was **deprecated in Bot API 7.0** in favour of
`link_preview_options: {is_disabled: true}`. It still works, but the mesh's messages carry no links
(the content rule refuses URLs), so the parameter can simply be dropped.
### Answering back
Answers reach a bot two ways:
- **Long polling with `getUpdates`**, over outbound HTTPS. Updates wait at Telegram for at most
24 hours.
- **A webhook** (`setWebhook`), which Telegram calls over HTTPS on port 443, 80, 88 or 8443. It may
carry a secret header (`X-Telegram-Bot-Api-Secret-Token`) that proves the call came from the webhook
that was set.
Only one of the two at a time. [08](08-asks-that-authorise.md) chooses between them.
## What Telegram sees
- **Everything in the message.** A bot chat is a "cloud chat": encrypted between the phone and
Telegram, and between Telegram and the bot API caller, and readable by Telegram. Bots cannot take
part in Telegram's end-to-end "secret chats".
- **That is what the content rule is for.** The holder refuses any message carrying an address, a
path or a secret's shape (to-be 45 §5), so what Telegram stores is roles, words and condition keys.
- **Who the operator is.** The account's phone number, and the addresses the phone and the sending
machines connect from. Since September 2024 Telegram's privacy policy says it may disclose a user's
IP address and phone number to judicial authorities on a valid order.
- **Machine names.** A condition key names the machine it is about, and so does the watcher's message
("told by mesh-watcher on …"). The content rule refuses host names with a top-level domain, not bare
machine names. To-be 45 says a message's subject is "a machine's role". Whether a bare machine name
may leave is a decision this effort has not taken; today it does.
## When Telegram is unreachable
- **A send fails at the transport** (no DNS, no connection, timeout). The holder keeps what it held
and tries again every minute. The watcher keeps what it owes and tries again at its next tick.
Neither loses a message while it runs.
- **A send fails because Telegram refuses** (400 or 403). It is permanent for that message. The holder
today treats it like a transport failure and tries again every minute, for ever (see the defects).
- **Telegram being down is invisible to Telegram.** The holder's status says the channel is failing.
The operator sees that only through another channel or by asking. This is what a second holder of a
different kind is for ([05](05-the-other-holders-on-the-same-axes.md)).
- **Telegram is blocked** in some countries and on some networks. An operator travelling should know
the mesh's phone channel may be one of them.
## Cost
- **Free.** No per-message charge. Telegram's paid broadcasts (above 30 messages a second) are far
out of range.
- **One account**, which the operator very likely already has.
- **Two secrets**, one token per bot, both `issued-by: outside`.
## The built code, checked against this
Read from the code repository's main branch on 2026-10-06: the holder's `telegram.go`, `outbox.go`,
`holder.go`, `content.go`, and the watcher's `telegram.go` and `watcher.go`. The two Telegram clients
are copies of each other, kept apart on purpose so the watcher depends on nothing it watches.
What is right:
- **Plain text**, with no `parse_mode`.
- **The token is kept out of every error.** The client rebuilds transport errors from their kind,
because the URL carries the token. It also strips the token from the API's own description.
- **Token and chat id are re-read at each send**, so accepting the secret or changing the setting needs
no restart.
- **A token's shape is checked.** A random value the mesh minted for an un-accepted secret is named as
that, not sent to Telegram to be refused.
- **A missing edit falls back to a new message.**
- **The holder caps itself** at 20 messages an hour and folds bursts into digests, far inside
Telegram's limits.
### Defects
| # | Where | What | Effect | Weight |
|---|---|---|---|---|
| D1 | holder: `telegram.go`, `outbox.go` | Nothing bounds a message to 4096 characters. A long summary, or a digest of long titles, is refused with 400. `failed` keeps the whole batch and retries it every minute. | One oversized message wedges the Telegram channel: everything queued behind it waits for ever. | high |
| D2 | holder: `holder.go` (`h.edit(old, "reopened", false)`) | A condition that clears and is raised again within ten minutes is said by **editing** the first message. On Telegram an edit notifies nobody. | A reopened urgent condition reaches the phone **silently**, high up in the chat's history. | high |
| D3 | holder: `outbox.go` (`r.Sent[name]` keeps only the first id) | Reminders and escalations are new messages, but clearing edits only the first. | The newest thing on the phone still says "STILL OPEN" or "NOW URGENT" after the condition cleared. The "CLEARED" is a silent edit, out of sight. | medium |
| D4 | both: `telegram.go` | `Message.Quiet` and `Message.Urgent` are ignored. `disable_notification` is never set. | A clearing, or a warning, rings as loudly as an urgent message. Telegram's only loudness lever is unused. | medium |
| D5 | both: `telegram.go` (`call`) | HTTP 429's `parameters.retry_after` is not read. The retry is a flat minute. | Harmless at the holder's cap. Under a real flood wait, repeating early prolongs it. | low |
| D6 | holder: `telegram.go`, `outbox.go` | Every refusal (400, 403) is retried like a transport failure. "message is not modified" on an edit is read as a failed edit, and a new message is sent instead. | Permanent errors loop every minute in the log. A no-op edit becomes a duplicate message. | low |
| D7 | both: `telegram.go` | `disable_web_page_preview` is deprecated since Bot API 7.0. | Works today. Moot, since no message carries a link. Drop it. | low |
| D8 | both: `telegram.go` (`call`) | The `json.Marshal` error is discarded. A non-numeric `message_id` (`json.Number`) makes the body empty. | A confusing refusal from Telegram instead of a local error. Ids come from Telegram, so it is unlikely. | low |
| D9 | watcher: `watcher.go` (`Tick`) | A "silent" message that could not be sent is overwritten by the "CLEARED" message when the signal returns. | The operator can receive "heard again" for a silence they were never told of. Better to say both, or one line saying it was silent for N minutes and is back. | low |
| D10 | both | `Ready()` is satisfied by a token and a chat id. It does not know whether the operator pressed Start, or whether the bot was blocked (403). | Status says "ready" until the first send fails. The watcher's own test verb is the only proof. A `getChat` check at status time would say it. | low |
D1 and D2 matter before the channel is configured. D1 can silence the channel. D2 silences exactly the
case (a flapping urgent condition) the operator most needs to hear.
## Sources
- Telegram, *Bots FAQ*: rate limits, paid broadcasts. https://core.telegram.org/bots/faq
- Telegram, *Bot API* (version 10.3, recent changes, `getUpdates` retention, `ResponseParameters`,
`setWebhook`, `link_preview_options`). https://core.telegram.org/bots/api
- Telegram, *Bot features*: BotFather, `/newbot`, `/token`, `/setprivacy`, deep linking.
https://core.telegram.org/bots/features
- `link_preview_options` replacing `disable_web_page_preview` (Bot API 7.0); the removal of the old
argument in python-telegram-bot v22. https://docs.python-telegram-bot.org/en/v22.0/telegram.ext.defaults.html
- "message is not modified" and "message to edit not found" in practice:
https://github.com/tdlib/telegram-bot-api/issues/400
- 409 "terminated by other getUpdates request":
https://community.home-assistant.io/t/help-on-telegram-extension-error-while-getting-updates-conflict-terminated-by-other-getupdates-request-make-sure-that-only-one-bot-instance-is-running-409/177544
- Telegram privacy policy change, September 2024:
https://www.bleepingcomputer.com/news/security/telegram-now-shares-users-ip-and-phone-number-on-legal-requests/
- Bots and secret chats; cloud-chat encryption:
https://www.kaspersky.com/blog/telegram-privacy-security/38444/
- Phone number required; anonymous numbers: https://en.wikipedia.org/wiki/Telegram_(software)
@@ -1,186 +0,0 @@
# 05 — The other holders, on the same axes
Every candidate is judged as a holder of the channel and intake seats in
[06](06-a-conversation-with-the-operator.md): which capabilities it can honestly declare
(the vocabulary is defined there), and what it needs. [02](02-the-channels.md) weighed the same
candidates before any of this was measured. This document replaces its reading of them with facts as
of 2026-10-06, and adds the question 02 could not ask: **can the operator answer through it, and
authorise an action through it?** ([08](08-asks-that-authorise.md)).
Two roles are judged separately, because they want different things:
- **The away channel:** the mesh's urgent messages and its asks, wherever the operator is.
- **The watcher's path:** the message that the mesh itself has gone silent. It must not depend on the
control node, the bus or the controller. The mesh observed has its controller, its bus and its mail
server on the anchor (the control node). Its Matrix server is on the home-server, behind a household
connection.
## The candidates
### ntfy
A small push server. Topics are published to over HTTP; the phone app subscribes.
- **Self-hosted vs the public server.** Self-hosted keeps the words on the operator's machines.
The public `ntfy.sh` takes no sign-up. Its free tier allows 250 messages a day **per IP address**,
shared with whoever else sends from that address. Paid tiers (from about $5–6 a month) give
reserved topics and higher quotas. An unreserved topic on the public server is readable by anyone
who guesses its name.
- **Phone delivery.**
- **Android:** through Google's FCM from the public server, or through the app's own long-lived
connection to a self-hosted server ("instant delivery"), which costs battery.
- **iOS:** cannot be reached by a self-hosted server alone. The server must name an upstream
(`upstream-base-url`, normally `ntfy.sh`), which receives a poll request carrying only a message
id and a hash of the topic, and has Apple wake the phone. The words do not pass through the
upstream. The dependency does.
- **Loudness:** five priorities. The highest gives "really long vibration bursts" and a pop-over on
Android.
- **Answers:** up to three action buttons. An `http` action makes **the phone** send a request,
which needs a route from the phone to the mesh and a credential carried inside the notification.
Nothing tells the server **who** tapped, only that someone holding the notification did. No free-text
reply.
- **As the watcher's path:** self-hosted, it fails with the machine it runs on. The public server
works, at the cost of a guessable topic or a subscription.
### Matrix (a homeserver is already one of the mesh's modules)
The module runs Conduit and Element Web, on the home-server.
- **Reach:** any Matrix client on the phone. Push goes from the homeserver to the client's **push
gateway**: for the stock Element apps, Element's gateway at matrix.org, which hands it to Apple or
Google. A self-hosted gateway needs a self-built app. UnifiedPush (via ntfy) is an option on Android.
- **Push support in Conduit has lagged.** Its own documentation long listed mobile push as missing,
and forks have since reworked pushers. Whether the running version pushes reliably is **unverified**
and must be measured before Matrix is relied on for anything urgent.
- **Answers:** free text, and reactions (`m.reaction` annotations) as one-tap choices. Clients add
emoji variation selectors, which must be normalised before a reaction is read as a choice.
- **Who answered:** the sender's Matrix id is authenticated **by the homeserver**, which the mesh
runs. That is strong for an account on the mesh's own server, and only as strong as that server.
- **Privacy:** end-to-end encrypted if the bot supports it. Otherwise readable by the homeserver,
which is the mesh's own.
- **As the watcher's path:** it survives the control node going down. It does not survive the home's
connection going down, and its phone push still depends on matrix.org's gateway.
- **Cost:** a bot account and its secret. No third party for the words.
### Pushover
A paid push service with a stable API.
- **Cost:** $4.99 one-time per platform after a 30-day trial. 10,000 messages a month per
application.
- **Limits:** 1024 characters, a 250-character title.
- **Loudness:** the strongest of any candidate.
- Priority 1 bypasses the user's quiet hours.
- Priority 2 ("emergency") repeats every `retry` seconds (at least 30) until acknowledged or until
`expire` (at most three hours). It returns a **receipt** that can be polled outbound to learn
whether, and when, it was acknowledged.
- **Answers:** acknowledgement only. No buttons, no reply.
- **As the watcher's path:** yes. It is outbound HTTPS from any machine, to a third party.
- **Where the words go:** to Pushover.
### Gotify
Self-hosted, Android only. Delivers over a WebSocket the app keeps open. There is no official iOS app,
and Apple's restrictions make a self-hosted iOS push impossible without a relay. It has no answer path
beyond opening the app. **Not pursued:** it covers less than ntfy and nothing ntfy does not.
### Mail through an outside provider
- **The mesh's own mail server is on the control node,** so it fails with it.
- **An outside relay** (an SMTP account at a mail provider) reaches the operator from any machine.
- **Urgency:** none. Mail is the digest and the record.
- **Answers:** a reply, slowly. **Who answered is weak:** a From line can be forged, and checking
DKIM only proves the operator's provider sent it.
- **Mail as an intake** (a new mail arriving) is a trigger in its own right
([06](06-a-conversation-with-the-operator.md)), whatever its weakness as a channel for asks that authorise.
### Signal, through `signal-cli`
- **An unofficial client,** and it needs its own phone number, registered with a captcha.
- **It must be kept current.** Signal's own clients expire after three months, and the server then
changes incompatibly. In March 2026 Signal began unregistering accounts whose client lacked a new
protocol feature, and every `signal-cli` account registered before that date was dropped.
- **Privacy:** end-to-end encrypted. Sender identity is strong (Signal's identity keys). Reactions
and replies both work.
- **Weight:** a phone number and a maintenance burden with a hard failure mode. It is the right
choice only for an operator who requires end-to-end encryption on the phone and accepts that burden.
### SMS through a paid gateway
The only candidate that reaches a phone with **no data connection**.
- **Cost:** per message, with an account at a gateway.
- **Privacy:** no encryption. Sender identity on replies is spoofable.
- **Its place:** the last resort of an outside dead-man service (below), which offers SMS and phone
calls on paid plans, rather than a channel of the mesh's own.
### An outside dead-man service
A service the mesh **pings**, which alerts by its own means when the pings stop. It is not a channel
of the mesh: it is the one thing that still speaks when **every** machine, or the home's connection
and the anchor together, are gone. [03](03-open-questions.md) Q6 named it the cheapest answer that
also covers "the whole house is offline".
- **Healthchecks.io** (also self-hostable, which defeats the point here) monitors 20 checks free,
without a card.
- It notifies through Telegram, Signal, Matrix, ntfy, Pushover, mail and others.
- Paid plans add SMS, WhatsApp and phone-call credits.
## The table
Capabilities are those of [06](06-a-conversation-with-the-operator.md). ✓ declared honestly,
— not, ~ conditional (the note says on what).
| Holder | reaches-away | loud | silent | edit | choice | reply | verified-sender | exact-render | code-factor | private | off the control node | cost |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Telegram | ✓ | — (do-not-disturb wins) | ✓ | ✓ | ✓ | ✓ | ✓ (user id) | ✓ | ✓ (code as a reply) | — | ~ (a holder on another machine) | free |
| desktop notifier | — | ~ (critical urgency) | ✓ | ✓ | ~ (actions, read by nobody) | — | — (any program of the account) | ✓ | — | ✓ | — | none |
| ntfy, self-hosted | ✓ | ✓ (priority 5) | ✓ | — | ~ (http action, phone → mesh) | — | — | ✓ | — | ✓ (iOS: relay sees ids) | — | a module |
| ntfy.sh | ✓ | ✓ | ✓ | — | ~ | — | — | ✓ | — | — | ✓ | free (250/day/IP) or ~$5/mo |
| Matrix (own server) | ~ (push unverified) | — | ✓ | ✓ | ✓ (reactions) | ✓ | ✓ (own homeserver) | ✓ | ✓ | ~ (E2E if the bot does it) | ~ (home-server, not anchor) | a bot account |
| Pushover | ✓ | ✓✓ (emergency, repeats) | ✓ | — | ~ (acknowledge only) | — | ✓ (for acknowledge) | ✓ | — | — | ✓ | $4.99 once |
| mail, outside relay | ✓ (slow) | — | ✓ | — | — | ✓ | — (forgeable) | ✓ | ~ (code in a reply) | — | ✓ | an account |
| Signal (`signal-cli`) | ✓ | — | ✓ | ✓ | ✓ (reactions) | ✓ | ✓ | ✓ | ✓ | ✓ | ~ | a number + upkeep |
| SMS gateway | ✓ (no data needed) | ✓ | — | — | — | ~ | — | ✓ | — | — | ✓ | per message |
## Reading it
**For the away channel and for asks,** Telegram is the only candidate that is all of these at
once: free, on both phone platforms, without a server of the mesh's own, and able to carry every tier
of authorising ask ([08](08-asks-that-authorise.md)), including a code as a second proof. Its price is that
Telegram reads the words, which is what the content rule is for.
- **Matrix** is the self-hosted equivalent for asks. It waits on a measurement of Conduit's push.
- **Signal** is the end-to-end-encrypted equivalent, at a maintenance cost that has already broken
every installation once this year.
**For waking the operator,** Pushover's emergency priority is the only thing that repeats until
acknowledged and gets through quiet hours. Telegram cannot. It is a reasonable **second** away holder
for urgent conditions only, if the operator wants to be woken. It cannot carry an answer beyond
"acknowledged".
**For the watcher's path,** the requirement is independence from what it watches:
- **Telegram, sent directly** from a machine that is not the control node, with its own bot, meets it
(to-be 45 §5).
- **ntfy self-hosted does not.**
- **Matrix only half meets it** (it is on the home-server, and its push goes through matrix.org).
- **What none of them covers** is the watcher's own machine, or the home's connection, going down
together with the anchor. Only an outside dead-man service covers that.
## Sources
- ntfy, *Configuration* (iOS upstream relay, FCM, access control): https://docs.ntfy.sh/config/
- ntfy, *Publishing* (priorities, actions, `http` action, message size): https://docs.ntfy.sh/publish/
- ntfy.sh pricing: https://ntfy.sh/#pricing. Its free-tier rate limit is per IP:
https://github.com/binwiederhier/ntfy/issues/1963
- Pushover API (length, quota, priorities, emergency retry/expire, receipts): https://pushover.net/api
- Pushover pricing: https://pushover.net/pricing
- Gotify, platform support: https://play.google.com/store/apps/details?id=com.github.gotify and
https://guancyxx.cn/en/blog/ntfy-vs-gotify-vs-nostr
- Element push gateway (Sygnal at matrix.org):
https://github.com/vector-im/element-android/blob/develop/docs/notifications.md
- Matrix reactions (`m.annotation`): https://github.com/uhoreg/matrix-doc/blob/aggregations-reactions/proposals/2677-reactions.md
- Conduit changelog: https://conduit.rs/changelog/
- signal-cli, registration with a captcha: https://github.com/AsamK/signal-cli/wiki/Registration-with-captcha
- signal-cli, unregistration of outdated clients in 2026: https://github.com/AsamK/signal-cli/issues/1993
- Healthchecks.io pricing: https://healthchecks.io/pricing/. Its Telegram integration:
https://healthchecks.io/integrations/telegram/
@@ -1,332 +0,0 @@
# 06 — A conversation with the operator
The operator's directions, 2026-10-06, in substance:
- **Telegram is one output channel among many to come.** The setup must be generic, and Telegram
simply fulfils a seat.
- **The same holds for input:** a new mail, a new message from the operator.
- **Each channel has capabilities.**
- **Approval is only an example.** An agent may just as well want to ask a simple question, and "it
doesn't have to be permission related".
So the core of this effort is not a notifier and not an approval path. It is **a conversation with
the operator, held over channels**:
- the mesh, its modules and its agents **say** things and **ask** things;
- the operator **answers**, or **writes first**;
- each exchange goes over whichever channel is right for its needs and for where the operator is.
This document is that general model. [07](07-the-work-context-and-the-desk.md) is how the work context
chooses the channel. [08](08-asks-that-authorise.md) is one layer on top: asks whose answer performs an
action, and the checks that makes necessary. [04](04-telegram-as-the-first-holder.md) and
[05](05-the-other-holders-on-the-same-axes.md) are the first holders.
## What exists, and how it is shaped
Read from the catalogue's main branch on 2026-10-06.
- **The output seat's holder is one module doing three jobs.** It claims `operator-channel` (mesh
scope, serving `open`, `history` and `notify`). It consumes the controller's three condition events.
It holds the Telegram bot token as its own secret. It reaches the desktop through the `node-notifier`
seat's `send`.
- The router, the Telegram client and the desktop adapter are one process.
- `notify` is a served verb, not the work queue to-be 32 §3 sketched.
- **The desktop notifier** is the node seat `node-notifier`, held by the dunst module on each
graphical machine.
- **The watcher's watcher** has its own Telegram client and token, and is not assigned yet.
- **Nothing reads anything back.** The operator speaks to the mesh only through an agent session or a
shell.
## The three things said
- **A message:** the mesh tells the operator something. A condition raised, a reminder, a clearing,
a notice from a module. It expects no answer. It may be edited later (a clearing).
- **An ask:** someone wants the operator's input, of a declared kind. The answer goes back to whoever
asked.
- **An operator message:** the operator writes first. It is a message to an agent, a note to the mesh,
or an answer to an ask written in the thread instead of tapped.
Inputs from outside that are not the operator (a mail arriving, a webhook) are the same kind of
envelope as an operator message, with a sender that is not the operator. This effort designs their
shape only. Their consumers are later work.
## Asks
### Kinds
| Kind | The operator gives | Required capabilities of the channel |
|---|---|---|
| `yes-no` | yes or no | `choice`, or `reply` (read as yes or no) |
| `one-of` | one of up to eight labelled options | `choice`, or `reply` with the option's number |
| `text` | free text | `reply` |
| `number`, `date` | a value of that type, within bounds | `reply`. Parsed and checked by the seat's holder; asked again once if it does not parse. |
| `acknowledge` | "seen" | `choice` |
An ask whose answer **performs an action** (approve a retirement, delete data) is the same ask with
an **authorising** flag. It adds requirements of trust, which [08](08-asks-that-authorise.md)
defines. Everything else about it is as below.
### What an ask carries
- an id;
- the asker (its bus principal and its machine);
- the kind, with its options or bounds;
- the words;
- a priority (urgent or normal);
- optionally a timeout and a default;
- optionally a conversation handle (where the asker is talking with the operator);
- optionally a group, for batching.
The words pass the content rule like every message.
### Its life
**open → answered | defaulted | expired | cancelled**
- **Answered:** the first answer wins. Copies of the ask shown on other channels are edited to say
where it was answered.
- **Defaulted:** the timeout passed and the ask declared a default. The asker receives the default,
marked as a default, never as the operator's answer.
- **Expired:** the timeout passed with no default. The asker is told.
- **An authorising ask never defaults to performing.** It expires. ADR 0230's rule, "a timer is the
mesh acting alone again, only later", applies to every ask that authorises.
- **Cancelled:** the asker no longer needs it (`ask cancel`), or its owner sees it is moot (the
condition behind it cleared). Its copies are edited to "no longer needed".
### How the asker gets the answer
- The answer is emitted as **`ask-answered`**, with the ask's id and, where there is one, the
conversation handle.
- An asker may **wait**: `ask` with a wait of up to a few minutes returns the answer if it comes in
time, and otherwise returns the id.
- An agent working through a long task can **poll** `asks <id>`.
### Batching
- **Asks to the same channel within the burst window go out together.** A heading says how many are
open ("3 questions waiting"), and each ask follows as its own message, so each can be answered and
edited alone.
- **An asker may hold at most three open asks.** A fourth is refused to it, in words, so an agent in
a loop cannot flood the operator.
- **Asks count against the router's hourly cap** like messages. Answers do not.
### History
**`asks`** lists open asks, and closed ones for 30 days: the asker, the kind, the outcome, the channel,
and the answer. A free-text answer is the operator's own words. It stays in the router's state and is
never forwarded to a channel other than the one it came from.
## Who holds what
### The output seat's holder becomes the conversation's router
It holds `operator-channel` and gains asks:
- `ask` (create),
- `ask cancel`,
- `asks`,
- `answer` (called by intake holders, below),
- events `ask-opened`, `ask-answered`, `ask-closed`.
It **owns** every ask that authorises nothing.
An authorising ask is **owned by the controller** ([08](08-asks-that-authorise.md)). The router carries
it like any other, but its answer goes to the controller, which alone can perform.
### Where a channel sits: [03](03-open-questions.md) Q1, asked again
Q1 settled that **one seat speaks for the mesh** and that sources never learn channels. It assumed
each channel attaches by contribution. With many channels to come, and channels that answer, that is
the question to settle.
The mesh's precedents:
- A **seat** has one holder at its scope (ADR 0121, ADR 0126).
- A **node seat** has one holder per machine, and a verb's subject carries the machine (design 33 §4).
- A **bench** is a seat with several holders. The only one is `mesh-dns-resolver`: the same module,
one per machine. "Making another one is a decision, recorded" (ADR 0223).
- A **work queue** is shared by a seat's holders (ADR 0190).
- A **contribution** is content another module hands to a seat's holder (ADR 0210, ADR 0212).
The options:
- **a. Each channel contributes itself to the output seat** (Q1 a, to-be 45 §5).
- A contribution is content a holder places.
- A channel is running code: it holds a secret, keeps a connection, reads answers, fails on its own.
- Making it fit puts every channel's client back in the router. **Rejected.**
- **b. One seat per channel kind.** The router learns every seat; a new channel is a change to the
router. **Rejected.**
- **c. Channel modules found by a manifest field and called by module address.** Callers use seats,
never modules (ADR 0126). **Rejected.**
- **d. One monolithic notifier,** every channel built into the router. Shared secrets and failures, a
release per channel. **Rejected.**
- **e. A channel module per service, with its own approval or question path** (a "Telegram module"
that decides things). It locks the conversation to one service, and the next channel repeats it.
**Rejected.**
- **f. Kinded benches.** Two mesh seats, **`channel`** (out) and **`intake`** (in). Their holders are
different modules, each claiming a **kind** (`telegram`, `desktop`, `ntfy`, `matrix`, `mail`, …).
- One holder per kind; two claiming one kind is refused at registration.
- A verb's subject carries the kind, as a node seat's carries the machine:
`mesh.seat.channel.tool.send.<kind>`.
- A new channel is a new module claiming a new kind, with no change to the router.
- **Chosen.**
Option f needs a second bench, of a new sort (different modules, keyed by kind), which ADR 0223 says
must be decided. It also needs a claim that carries a kind and capabilities.
## The channel seat (out)
**`channel`**, mesh scope, a kinded bench.
### Served
- **`send`:** words, a priority, whether silent, an optional **ask block** (the kind, the options each
with an opaque token, the ask id) and an optional thread (the conversation handle, or the message
this replies to). Answers the channel's id for what it sent, or a refusal in words.
- **`edit`:** replace a sent message by its id, where `edit` is declared.
- **`standing`:** ready, not configured (naming what is missing, never a value), or failing (since
when, why); the last delivery; the declared capabilities.
### Emitted
- **`delivered`** and **`failed`**, the latter marked **permanent** or **transient** (defect D6 in
[04](04-telegram-as-the-first-holder.md)).
### Honoured
- The declared maximum length, by cutting and saying so (D1).
- Silence where declared (D4).
- An identical edit is a success (D6).
- No own secret in any error or event.
### The content rule
The content rule is the router's, applied before anything reaches a holder that does not declare
`private`.
## The intake seat (in)
**`intake`**, mesh scope, a kinded bench. A service that is read and written by one program (a
Telegram bot, [04](04-telegram-as-the-first-holder.md)) is held by one module claiming both seats under
one kind.
A holder turns what arrives into **one envelope**:
| Field | What it is |
|---|---|
| `id` | Unique, for deduplication. |
| `kind` | The holder's kind. |
| `what` | `message` (written first), `choice` (a button or a reaction), `reply` (written in an ask's thread), `mail`, `call` (a webhook), `seen` (activity, for the work context). |
| `sender` | The identity on that service, and whether the service authenticated it. |
| `operator` | Whether that identity is on the **controller's** list of the operator's identities. Filled from that list, never from the holder's own. |
| `conversation` | An opaque handle. Sending on it reaches the same chat, room or mail thread. |
| `in-reply-to` | The ask or message it answers, if any. |
| `payload` | The text, or the option's token. **No secrets:** a code for an authorising ask never travels in an envelope ([08](08-asks-that-authorise.md)). |
| `at` | When. |
**Answers go to the ask's owner by request and reply:**
- the router's `answer` for an ordinary ask;
- the controller's for an authorising one.
**Everything else is an event on the seat,** which consumers take by `what`:
- the router takes `seen` for the work context;
- an agent bridge takes `message` from the operator;
- a future mail rule takes `mail`.
**One gap to close:** the shared library cannot yet publish on a seat's event subjects (design 32 §1).
## The capability vocabulary
**Fixed and versioned:** `channel-capabilities/1`. A holder declares capabilities in its claim, and
the catalogue refuses a word outside the vocabulary. **Each capability has a contract test** the
holder's build runs and a **drill** its `standing` can run. One that fails its drill is reported and
withdrawn from routing until it passes.
The words are in three groups. The first two serve every conversation. The third exists only for
asks that authorise, and is defined in [08](08-asks-that-authorise.md).
### Delivering
| Capability | Promise | Test / drill |
|---|---|---|
| `deliver` | It arrives, or `failed` says why. | Against a test double; a live test message. |
| `reaches-away` | It reaches a phone away from the operator's machines. | Declared by kind; the operator acknowledges a drill. |
| `loud` | It can break through the phone's quiet hours. | The service's override is set for urgent. |
| `silent` | It can arrive without sound. | The service's silent flag is set. |
| `edit` | A sent message can be replaced in place. | Edit and read back, against a double. |
| `max-length:N` | Up to N characters arrive whole; longer is cut, and the cut is said. | N+1 characters give a cut message, not a failure. |
| `reaches-when-mesh-down` | Delivering needs neither the bus nor the control node. | Checked by the holder's placement and its send path. |
| `private` | The words stay on the operator's machines, or are end-to-end encrypted. | Declared by kind; reviewed. |
### Conversing
| Capability | Promise | Test / drill |
|---|---|---|
| `choice` | The operator can pick one offered option in one act, and the pick comes back. | A simulated tap gives a `choice` envelope with the option's token. |
| `reply` | The operator can answer in free text, and it comes back. | A simulated reply gives a `reply` envelope. |
| `threads` | An answer is tied to the message it answers. | A reply to message A carries A in `in-reply-to`. |
| `operator-first` | The operator can write to the mesh unprompted. | A simulated message gives a `message` envelope. |
### Trusting
`verified-sender`, `exact-render`, `code-factor` and `key-factor` are defined in
[08](08-asks-that-authorise.md). A channel without them still converses fully. It just cannot carry
an answer that performs an action.
## Agents in the conversation
An agent is a participant. It asks, and it receives.
- **An agent whose operator is at its own terminal** asks there, in the terminal. The seat is not
needed.
- **An agent working unattended** (in the background, on a schedule, or with the operator stepped
away) asks **through the seat**. The router puts the ask where the operator is now
([07](07-the-work-context-and-the-desk.md)): the desk if they are at it, Telegram if they are away
or talking through Telegram. The agent waits for, or polls, the answer.
- **An agent the operator talks to through Telegram** receives the operator's messages as intake
envelopes. It answers on the conversation handle. Its asks carry that handle, so they appear in the
same chat.
- **An agent's words reach the operator only as an asker's words,** and the operator's answer reaches
the agent only as the operator's. An agent relaying "the operator said yes" is not an answer to
anything. That is why an answer comes from a channel holder, never from the asker
([08](08-asks-that-authorise.md)).
## The watcher, in this shape
The watcher's watcher stays **outside** the seats, deliberately:
- it must speak when the bus and the control node are what failed;
- the seats live on the bus.
It is a minimal `deliver` + `reaches-when-mesh-down` sender with its own bot. It never reads, and never
asks. Its sibling outside the mesh is the dead-man service ([05](05-the-other-holders-on-the-same-axes.md)).
## From today to this shape
1. **Fix the first holder in place:** D1–D4 of [04](04-telegram-as-the-first-holder.md). No change of
shape.
2. **The operator configures Telegram** ([09](09-a-proposed-decision.md)). The mesh starts telling.
3. **The vocabulary and the kinded bench in the catalogue,** and seat events in the shared library.
4. **Split Telegram out** into its own module holding `channel` and `intake` under `telegram`. The
router keeps the desktop adapter as `channel/desktop` and `intake/desktop`.
5. **Asks in the router,** with buttons and replies on Telegram and actions on the desktop
([07](07-the-work-context-and-the-desk.md)). Agents can ask from here on.
6. **The work context** as the router's ordering ([07](07-the-work-context-and-the-desk.md)).
7. **Authorising asks in the controller** ([08](08-asks-that-authorise.md)).
8. **Operator-first messages,** and an agent bridge consuming them.
9. **Further holders as wanted:**
- Pushover for waking;
- Matrix once its push is measured;
- mail out for the digest, mail in as an intake.
The watcher changes only at step 1.
## What this revisits in [03](03-open-questions.md)
- **Q1:** channels attach as holders of kinded benches (f), not as contributions.
- **Q3:** the life of a message gains the life of an ask. `edit` and `silent` are declared, and decide
how a clearing is said (D2, D3).
- **Q4:** presence becomes the work context ([07](07-the-work-context-and-the-desk.md)).
- **Q7:** answering back is the conversation itself. Its authorising layer is
[08](08-asks-that-authorise.md).
- **Q8:** the content rule stays, applied by the router. Whether a `private` holder may be exempted
is left to graduation.
@@ -1,117 +0,0 @@
# 07 — The work context, and the desk
The operator, 2026-10-06:
- "If dunst can also show buttons, we could prefer to use desktop notifications instead of Telegram
for approval actions when working in a session."
- "The work context is an important factor when deciding the correct output channel."
[06](06-a-conversation-with-the-operator.md) decides **which channels may carry** a message or an ask:
those whose declared capabilities satisfy it. This document decides **which of those comes first**,
from where the operator is working. It also makes the desk a full participant in the conversation.
## The rule
> **A message or ask goes to the most direct channel in the operator's current context, among those
> whose capabilities already satisfy it. Unanswered in time, it escalates along a fixed chain.
> Context orders the candidates; it never adds one. Context never lowers the bar.**
The last sentence matters most for asks that authorise ([08](08-asks-that-authorise.md)). Being at the
desk never makes a click count as more than it proves.
## The signals
| Signal | Source | Read as |
|---|---|---|
| A graphical session unlocked, with input in the last 5 minutes, on machine M | The `node-lock-screen` seat's holder on M (the screen-lock module), whose tools already read the idle time and the lock state. Underneath: logind's `LockedHint` and `IdleSinceHint`. Proposed: an event on each change, not a poll. | **at the desk on M** |
| That session locked, or idle longer | the same | **not at the desk** |
| An agent asking from machine M | The ask's asker names its machine. An agent module's own "session active" event, when one exists. | **working with an agent on M**. It strengthens "at the desk on M"; alone it proves nothing. |
| A verified intake `message`, `reply` or `choice` in the last 15 minutes | The intake seat ([06](06-a-conversation-with-the-operator.md)) | **in a conversation** on that kind |
| An ask carrying a conversation handle | The ask | **that conversation**, whatever else is true |
| The hour, against quiet hours | The router's setting | **night**: only urgent wakes |
## The contexts, and where things go
"The away channel" is the operator's setting (Telegram, to begin with). "The loud holder" is an
optional second away holder for waking (Pushover, [05](05-the-other-holders-on-the-same-axes.md)).
| Context | An ask goes to | Urgent message | Warning | Unanswered or unacknowledged → |
|---|---|---|---|---|
| **In a conversation through Telegram** (the ask carries its handle, or Telegram activity is newer than any desk input) | that chat, in the thread | that chat | that chat, silent | after 10 min (urgent) or 1 h: also the desk, if active |
| **At the desk on M** (with or without an agent there) | the desk on M, if it can carry the ask; otherwise the away channel, and the desk says where it went | the desk on M, and the away channel silently | the desk on M | after 5 min (urgent) or 30 min: the away channel, with sound |
| **Away** (no unlocked active session, no recent conversation) | the away channel | the away channel | the away channel, silent | after 15 min (urgent): the loud holder, if configured |
| **Night, away** | non-urgent asks wait for the morning; urgent as away | the away channel and the loud holder | the morning digest | as away |
| **The desk locks while an ask is shown there** | moves at once to the away channel | — | — | — |
- **Every copy of an ask stays valid until one answer wins.** The others are edited to say where it
was answered.
- **The context's channel cannot carry the ask** (a free-text ask at a desk without a prompt, or an
authorising ask the desk cannot prove): the next in the chain carries it, and the context's channel
says where it went.
- **Nothing can carry it:** the router says so, as a condition of its own.
### Presence stays in the mesh
- **Presence facts are events on the bus,** consumed by the router.
- **They are kept as current state only:** a key-value entry per machine and per intake kind,
overwritten, never a history.
- **They never appear in a message's words,** so they never reach a channel that is not `private`.
- **No module keeps them as a timeline of the operator's day.** A consumer that wants one is a
decision of its own.
## The desk as a participant
Read from the catalogue's main branch and the tools' current documentation, 2026-10-06.
### What exists
- **The `node-notifier` seat** is held on each graphical machine by the dunst module. Its `send` runs
`notify-send --print-id` with an urgency, an application name and an optional replace id. It
carries **no actions** today.
- **libnotify's `notify-send`** (0.8 and later) takes `--action=NAME=Label`, repeatable, and `--wait`.
It prints the chosen action's name when one is chosen, and nothing when the notification is closed.
Underneath, the notification server emits `ActionInvoked` with the notification's id and the
action's key.
- **dunst shows actions:**
- `do_action`, which the module binds to the **middle** click, invokes the default or only action;
- otherwise it opens the **context menu**, which in this mesh is the `node-launcher` seat's
dmenu-compatible menu;
- `dunstctl action` and `dunstctl context` do the same from a command line.
- **The `node-launcher` seat's `menu` verb** shows a list in the operator's session and answers the
chosen line. A dmenu-compatible menu also accepts typed text that is not a listed line, which makes
it a free-text prompt.
- **The graphical session is X11.**
### What it takes
- **`send` gains actions:** a list of (token, label).
- **It still answers at once.** A notification may be answered minutes later.
- **The holder listens for `ActionInvoked`** and emits the chosen token as an event on its node seat.
- **The router's desktop adapter,** which holds `channel/desktop` and `intake/desktop`, turns that
into a `choice` envelope.
- **For `text`, `number` and `date` asks,** the notification's single action opens the launcher's
prompt, and what is typed comes back as a `reply`.
### What the desk can declare
| Group | Capabilities |
|---|---|
| Delivering | `deliver`, `silent` (low urgency), `loud` (critical urgency stays until dismissed), `edit` (replace id), `private` |
| Conversing | `choice` (actions), `reply` (through the launcher's prompt), `threads` (the ask's id is carried) |
| Trusting | **not** `verified-sender`. `exact-render` yes. `code-factor` (a prompt) yes. `key-factor` yes where a security key is plugged in. See [08](08-asks-that-authorise.md). |
So the desk carries **every ordinary ask**: yes or no, one of, text, number, date and acknowledge.
It needs no account anywhere. An agent working unattended on the workstation asks a clarifying question,
and the operator, at the desk, answers it in the notification.
It carries an **authorising** ask only with a factor the controller verifies itself, because a click
on an X11 desk proves that someone was there, not that the operator clicked
([08](08-asks-that-authorise.md)).
## Sources
- `notify-send(1)`, `--action` and `--wait`: https://man.archlinux.org/man/notify-send.1.en
- Desktop notifications and actions: https://wiki.archlinux.org/title/Desktop_notifications
- dunst documentation (mouse actions, `do_action`, the context menu): https://dunst-project.org/documentation/
- logind's `LockedHint`, `IdleHint`, `IdleSinceHint`:
https://freedesktop.org/software/systemd/man/org.freedesktop.login1.html
@@ -1,247 +0,0 @@
# 08 — Asks that authorise
Most asks inform their asker and change nothing ([06](06-a-conversation-with-the-operator.md)). Some
answers **perform an action**: approving a retirement, confirming that a binding moves, deleting data.
This document is the layer those asks need on top of the conversation. It is the controller's checks,
the trust a channel must prove, and what a compromise can reach.
The operator, 2026-10-06:
- "Make sure I can approve and reject stuff via the Telegram channel."
- In a terminal session with an agent, having to open Telegram is acceptable.
- But when talking to an agent **through** Telegram, the operator cannot switch to a desktop session.
- At the desk, desktop buttons would be preferred ([07](07-the-work-context-and-the-desk.md)).
## Where this stands against what was decided
- To-be 45 §5 says: "No answering back in this form".
- ADR 0227 kept it open on purpose: "routing by presence, quiet hours, **answering back** and the
external dead-man service stay open in 028, whose graduation amends to-be 45."
- So this answers 028's Q7. It does not reverse ADR 0227.
## What there is to authorise
Read from the controller's main branch on 2026-10-06. The condition store marks conditions only a
person resolves (`resolver: operator`): from the start for retirement, clean-up and binding conditions,
and once a healer's budget is spent.
| Action | The verb today | Asked for by | Reversible | Tier |
|---|---|---|---|---|
| Approve the retirement set waiting | `retire approve <node> <provider> --why` | `retire-waiting` (urgent) | yes: asking for a consumer again re-enables it | approve |
| Reject it | `retire reject … --why` | `retire-waiting` | yes | approve |
| Approve a set rejected before | `retire approve …` | `retire-rejected` (warning) | yes | approve |
| Confirm that a binding moves, once its data is moved | `pin <node> <provision> <from> <module>` | `binding-kept` (urgent) | the pin, yes. The data, not by the mesh. | approve |
| End a stuck plan | `plans stop` / `close <id> --why` | `stalled`, `sent-not-reported` once escalated | no, but it destroys nothing | approve |
| Send a machine its declaration by hand | `push <node> --why` | `sent-not-reported` once escalated | n/a | approve |
| Reset a bus consumer's position | `broker consumer-reset --why` | `consumer-behind` once escalated | no: messages skipped or redelivered | approve (graduation to confirm) |
| Try a healer's repair once more | the healer's ordinary path | any escalation (H1–H5), `healers-braked` | as the repair is | approve |
| Silence a condition | `conditions silence <key> --for --why` | any | yes, it ends by itself (at most 7 days) | acknowledge |
| Delete one retired consumer's data | `cleanup delete <node> <provider> <consumer> --why` | `cleanup-waiting` (after 30 days) | **no** | destroy |
| Delete everything retired longer than N days | `cleanup delete --older-than N --confirm --why` | `cleanup-waiting` | **no** | destroy |
| Change the operator's identities, the away channel, or a factor's enrolment | (new) | — | yes, but it changes who may authorise | destroy |
The build queue verbs and `replay --register` are not offered as asks: no condition asks for them.
## Trust, as capabilities and proofs
The conversation's vocabulary ([06](06-a-conversation-with-the-operator.md)) gains four words used
only here:
| Capability | Promise | Test / drill |
|---|---|---|
| `verified-sender` | The holder proves the answer came from the operator's own account on that service, by the service's authentication, through a holder **no agent shares**: not on the operator's account, not on a machine where agents run as the operator. | A choice from an identity not on the list is dropped and reported. The holder's placement is checked. |
| `exact-render` | The ask is shown as the controller rendered it, by the holder itself. No asker or agent composes the words the operator authorises. | Rendered text equals the controller's, byte for byte, against a double. |
| `code-factor` | The holder can carry a code the operator types to the controller, by request and reply, and never judges it. | A code is never in an event, and is deleted from the conversation where the service allows. |
| `key-factor` | The holder can run a security key's assertion over the controller's challenge, and hand the controller the result. | The challenge is the controller's, and the signature is checked by the controller. |
From these, three **proofs** that the operator is the one answering:
- **P1, a verified sender:** a Telegram tap, through a holder on a machine no agent runs on.
- **P2, a code:** from the operator's authenticator, verified by the controller.
- **P3, a key touch bound to the ask:** the controller's challenge is a hash of the ask's id and the
state digest. A FIDO2 assertion with user presence (the key waits for a touch) is checked against the
operator's enrolled credential. It proves a physical touch for **this** ask and no other.
## Three tiers
| Tier | Required |
|---|---|
| **acknowledge** | `choice`, `exact-render`. Silencing is open to agents already, and announced; a proof adds nothing. |
| **approve** | `choice`, `exact-render`, and **one** proof (P1, P2 or P3). Single use, bound to the exact state shown, expiring when that state changes or after 24 h. |
| **destroy** | `choice`, `exact-render`, and **two** proofs, at least one of them P2 or P3. Valid 10 minutes after it is shown; at most one destroy answered per 10 minutes. |
| Where the operator answers | Proofs it offers | acknowledge | approve | destroy |
|---|---|---|---|---|
| Telegram | P1 (tap), P2 (code as a reply) | tap | tap | tap and code |
| the desk, with a security key | P3 (touch), P2 (code in a prompt) | click | click and touch | click, touch and code |
| the desk, without a key | P2 (code in a prompt) | click | click and code | not possible: carried by the away channel |
| an agent's terminal | none | none | none | none: an agent asks, it never answers |
| the console (a shell) | P2 (code) | — | break-glass: a code | none |
### The rule
1. **Every authorising verb declares its tier** in the controller's verb table.
2. **An authorising ask is offered only on a channel whose capabilities satisfy its tier.** An answer
arriving from any other channel is refused, and the refusal is said there.
3. **The operator's away channel must satisfy every tier.** The self-check verifies it. A setting
that would make a tier possible only at a desk is refused, unless the operator chose that for the
tier explicitly. While working through Telegram, everything can be completed in Telegram.
4. **The work context chooses among the channels that qualify; it never makes one qualify**
([07](07-the-work-context-and-the-desk.md)).
5. **No agent authorises.** An agent asks. It holds no verb that performs an authorising action. The
record names the agent that asked.
## Why the desk needs a factor
- **X11 does not isolate the clients of one display.** Any of them can inject input (the XTEST
extension, as `xdotool` does) and read keystrokes.
- **`dunstctl action` invokes a notification's action** for any program of the account.
- **The desktop holder runs as the operator's account,** whose files, the holder's bus credential
included, every agent on that account can read.
So a click at the desk, its `ActionInvoked` and the desktop holder's envelope can all be produced by
an agent. An unlocked session with recent input proves a person was there, not that the person
clicked. The desk declares no `verified-sender`. A factor the **controller** verifies gets around
that.
| Factor at the desk | Can an agent on the account fake it? | Judgement |
|---|---|---|
| **A security key's touch, bound to the ask** (P3) | No: the touch is physical, and the signature covers this ask's id and state. | **Preferred.** One click, one touch, no phone. Needs a key and an enrolment. |
| **A code typed into the launcher's prompt** (P2) | It cannot know the code. On X11 it can read keystrokes and race to use the code first, and each step's code is accepted once. | **Acceptable** without a key. Costs picking up the phone. The race is a residual risk until the session leaves X11. |
| **The screen's unlock or a fingerprint** | Yes: only the local holder sees the result. | **Rejected.** |
The desk would earn `verified-sender` only if agents ran under an account of their own, without the
operator's display, session bus or holders' credentials, on a compositor that isolates clients. That
is a question for the agent modules' placement. It is noted, not proposed.
## The controller holds authorising asks
Today a grant covers a whole verb (`seat:mesh-controller.retire` allows approving and listing alike).
A hand-act records `by` from the calling bus principal, which for a channel would be the module, not
the person. Both call for the controller to hold these asks itself.
- **The verb table gains a field.** The controller's verb definition (a name, a description, input and
output schemas) gains **`authorises`**: the tier, and the arguments that make up the exact state a
person must see (for `retire approve`, the set of consumers).
### Three verbs
- **`authorise request`:** anyone may call it, an agent or the router on a condition's behalf. It
carries the action (verb and exact arguments), why, and optionally a conversation handle.
- The controller renders the ask: what is asked, the exact state, who asks, what each option does.
- It stores it with an opaque id (10 random base32 characters), its expiry and a digest of the
state shown, and emits `ask-opened` with the controller as owner.
- The router carries it like any ask. **Nothing is performed.**
- **`authorise answer`:** only intake holders are granted it. It carries the id, the option, the
sender's identity, and the code or key assertion where the tier needs them. The controller checks,
and refuses at the first failure:
1. the ask is open and not expired;
2. the caller is the holder of an intake kind, by the controller's own seat records, never by the
request's claim;
3. that kind's declared capabilities, from the controller's records, satisfy the tier, and the
proofs present are enough;
4. a P1 answer: the sender is on the controller's list of the operator's identities for that kind;
5. a code: valid for the current or previous 30-second step, and unused;
6. a key assertion: it verifies against the enrolled credential, over this ask's challenge, with the
user-presence bit set;
7. the state **now** has the digest it had when shown. Otherwise the ask is void and a new one is
requested.
Then it performs the action as itself, records the hand-act, closes the ask with a compare-and-set
(so a second answer on another channel loses), and emits `ask-answered`.
- **`authorisations`:** open and recent authorising asks.
### The authorising verbs refuse to be called directly
`retire approve|reject`, `cleanup delete`, and `pin` while a `binding-kept` names it refuse a caller
unless the call comes through `authorise answer`, or carries a valid code as **break-glass** at the
console. Break-glass is recorded as such, and announced on every channel.
- `retire approve` must accept the set it approves (`expect`). Today it re-reads the set when it
runs, so "approve what you were shown" (ADR 0230) does not hold end to end.
- `conditions silence` stays callable. A silence an agent sets is said on the away channel, with its
why.
### The record
The hand-act gains:
- **`via`:** the kind and the holder;
- **`requested-by`:** the agent principal or the condition key;
- **`ask`:** the id;
- **`proofs`:** which of P1, P2 and P3 were present.
`by` reads "the operator, as <kind> identity <id>". Every copy of the ask is edited to the outcome
("approved by the operator on telegram at 14:02 UTC"), and its buttons are removed.
## Telegram, carrying them
- **Buttons.** `callback_data` is at most 64 bytes, so a button carries only `a1:<id>:<option>`.
Everything else is in the controller.
- **A tap** arrives as a `callback_query` with the tapping user's id and the chat. The holder:
1. drops it, and reports, unless both are the operator's;
2. calls `answerCallbackQuery` at once, because the phone shows a spinner until it does;
3. calls `authorise answer`;
4. edits the message to the outcome or the refusal.
- **A destroy ask** answers the tap with a `ForceReply` prompt for the code. The holder reads the
operator's reply, deletes it from the chat (bots may delete incoming messages in private chats), and
hands it to the controller by request and reply. It is never put in an event.
- **Long polling, not a webhook.**
- `getUpdates` needs no route into the mesh.
- A stolen token can steal updates, and a second reader shows as HTTP 409, but it cannot inject an
update.
- A webhook needs a public route, and a stolen token can redirect it.
- The offset is kept in the holder's state, and ids are single use anyway.
- **Placement.** The Telegram holder runs where no agent runs as the operator. Its own declaration of
`verified-sender` depends on it.
## Agents
- **In a terminal:**
1. The agent calls `authorise request`.
2. The router carries the ask to where the operator is ([07](07-the-work-context-and-the-desk.md)):
the desk, with a factor, or Telegram.
3. The agent sees `ask-answered`.
Nothing the agent says counts.
- **Through Telegram:** the request carries the conversation handle, so the ask appears in the same
chat, rendered by the holder, and the operator taps in place, destroy included.
## If something is compromised
| What is lost | What the attacker can do | What limits it |
|---|---|---|
| The bot token | Read what the bot is sent from then on. Steal taps (visible as 409). Send the operator fake messages. | It cannot call `authorise answer`. Revoke with BotFather's `/token`. |
| The operator's Telegram account, on a new device | Approve or acknowledge. | Telegram's two-step password. Every authorisation is announced on the other channels. Approve is reversible. Destroy needs a code. |
| The phone, unlocked | Everything, including destroy, if the authenticator is open on it. | An authenticator behind biometrics. One destroy per 10 minutes, announced. Backups (research 030). A setting turning destroy off for the away channel. |
| An agent on the operator's account | Click at the desk, read the desk's keystrokes, read the desktop holder's credential. | No tier accepts the desk without a code or a key touch the controller verifies. A code read off X11 is good for one step, and the race is said above. |
| A channel module, or its bus account | Forge P1 for approve or acknowledge. | It cannot forge P2 or P3. Destroy needs one of them. |
| Telegram itself | Read the words. In principle, forge a tap. | The content rule. Destroy needs P2 or P3, which Telegram never sees. |
**The factors' secrets are the controller's own.**
- The TOTP seed is made by the mesh. It is enrolled by showing its URI once, only to a terminal, and
never through a channel or an event.
- A security key is enrolled by registering its credential's public key.
- Re-enrolling either is a destroy ask.
## The other holders, for authorising
- **Matrix** can declare everything Telegram does: reactions, replies, a sender authenticated by the
mesh's own homeserver, codes. It is the self-hosted carrier once its push is measured.
- **ntfy:** an `http` action makes the phone call the mesh, with a credential inside the notification.
Nobody knows who tapped, and there is no reply. No tier.
- **Pushover:** acknowledgement, read back by polling a receipt. At most `acknowledge`.
- **Mail:** a reply can carry a code, but the sender is forgeable. No tier, except as break-glass.
## Sources
As in [04](04-telegram-as-the-first-holder.md) and [07](07-the-work-context-and-the-desk.md), and:
- Bot API `callback_data` (1–64 bytes), `answerCallbackQuery`, `ForceReply`, `deleteMessage`:
https://core.telegram.org/bots/api
- `answerCallbackQuery` is required even with no text: https://gramio.dev/telegram/methods/answercallbackquery
- X11 and its clients (input injection, keystroke reading):
https://hackindex.io/services/x11/exploitation/x11-session-hijacking and
https://www.semicomplete.com/projects/xdotool/
- `fido2-assert` (user presence, verifying an assertion):
https://developers.yubico.com/libfido2/Manuals/fido2-assert.html
- ntfy `http` actions: https://docs.ntfy.sh/publish/
- Pushover receipts: https://pushover.net/api
@@ -1,238 +0,0 @@
# 09 — A proposed decision, and what the operator does now
This is the effort's reading as of 2026-10-06, written so that playbook 02 can turn it into a record
and amend to-be 45 §5. It is a proposal: nothing here is decided until it graduates.
## The recommendation, short
1. **The mesh holds a conversation with the operator over channels.** It sends messages and asks;
the operator answers or writes first. A message, an ask and an operator message are the three
things said ([06](06-a-conversation-with-the-operator.md)).
2. **Channels and intake are seats.** One kinded bench each, `channel` and `intake`. Each holder is a
module of its own, claiming a kind and declaring capabilities from a fixed, versioned vocabulary,
each with a contract test and a drill.
3. **Asks are general.** Yes or no, one of, text, number, date, acknowledge, each requiring its own
capabilities. They have timeouts, defaults, cancellation, batching and history. The answer returns
to the asker as an event, and an asker may wait.
4. **The output seat's holder is the router.** It chooses among the channels that satisfy a message or
ask, by **work context**: in a Telegram conversation, Telegram; at the desk, the desk; away, the
away channel. Unanswered, it escalates. **Context never lowers the bar** ([07](07-the-work-context-and-the-desk.md)).
5. **The desk is a full participant.** Notification actions and the launcher's prompt carry every
ordinary ask, with no account anywhere.
6. **Asks that authorise are a layer on top,** held by the controller. Three tiers (acknowledge,
approve, destroy), and three proofs the operator is answering (a verified Telegram sender, a code,
a security key's touch bound to the ask). The controller checks the channel's declared
capabilities from its own records and verifies codes and key assertions itself
([08](08-asks-that-authorise.md)).
7. **No agent authorises.** An agent asks; the operator answers where they are. On an X11 desk, where
an agent could click for them, only a code or a key touch counts. In a Telegram conversation, the
ask appears in that chat and is completed there, destroy included.
8. **Telegram is the first holder and the away channel.** It is free, on both phone platforms, needs
no server of the mesh's own, and is the only candidate that carries every tier
([04](04-telegram-as-the-first-holder.md), [05](05-the-other-holders-on-the-same-axes.md)).
9. **The watcher's watcher stays outside the seats,** with its own bot, on a machine that is not the
control node. A free outside dead-man service covers the rest. Pushover is the optional holder for
waking. Matrix is the self-hosted carrier once its push is measured.
10. **The built Telegram code needs D1–D4 fixed before it is configured**
([04](04-telegram-as-the-first-holder.md)).
## What the operator does
Minimal, in order. Steps 1–6 are possible today. The desk needs nothing from the operator: no account,
no bot.
1. **Make two bots.** In Telegram, open BotFather and run `/newbot` twice: one for the mesh's
conversation, one for the watcher. Keep each token out of agent sessions.
2. **Close them to groups:** `/setjoingroups` → Disable, for each bot.
3. **Press Start** in each bot's chat. A bot cannot write first.
4. **Turn on Telegram's two-step verification,** if it is not on.
5. **Find your chat id.** Until the linking verb exists, call `getUpdates` once, reading the token from
a file, and take `message.chat.id` from your `/start`. It is the same for both bots.
6. **Give the mesh the values,** through the controller, never on disk:
- accept the mesh bot's token as the output seat's holder's own secret `telegram-token`, and set
its `telegram-chat-id`;
- assign the watcher to a machine that is not the control node, accept the watcher bot's token,
and set its chat id;
- push both machines, and run each module's test verb.
7. **Later, once built:**
- make a free dead-man check, and give its ping address to the mesh;
- enrol an authenticator, at a plain terminal;
- optionally, enrol a security key, to approve at the desk with one touch.
## What would change in the code (proposal, not built)
- **Now, in the output seat's holder and the watcher:** D1 (cut at 4096 characters), D2 (a reopening
is a new message), D3 (a clearing reaches the newest message), D4 (`disable_notification` for
warnings and clearings), then D5–D10.
- **Catalogue:**
- a claim carries `kind` and `capabilities`;
- a kinded bench refuses two holders of one kind;
- the vocabulary `channel-capabilities/1` and its contract tests;
- the shared library publishes on a seat's event subjects.
- **Router:**
- asks (`ask`, `ask cancel`, `asks`, `answer`, and the events `ask-opened`, `ask-answered`,
`ask-closed`);
- the work context, from presence events;
- the escalation chain.
- **Desktop:** `node-notifier.send` gains actions, and the holder emits the chosen one. The
screen-lock holder emits lock and idle changes.
- **Controller:**
- `authorises` on the verb definition;
- `authorise request`, `authorise answer`, `authorisations`;
- `retire approve` takes `expect`;
- the authorising verbs refuse direct calls without a code;
- the hand-act gains `via`, `requested-by`, `ask` and `proofs`;
- the operator's identities, the TOTP seed and enrolled keys as its own;
- a self-check probe: the away channel satisfies every tier.
- **Modules:**
- a `telegram` module holding `channel` and `intake` (long polling, buttons, replies, `ForceReply`
codes, a linking verb with a one-time deep-link code), placed where no agent runs as the operator;
- an agent bridge for operator messages;
- the dead-man ping.
## The tables
### Ask kinds × what a channel needs
| Ask | deliver | choice | reply | threads | Trust (only if it authorises) |
|---|---|---|---|---|---|
| a message (no answer) | ✓ | | | | |
| acknowledge | ✓ | ✓ | | | tier acknowledge: `exact-render` |
| yes-no | ✓ | ✓ or | ✓ | ✓ | tier approve or destroy, if it authorises |
| one-of | ✓ | ✓ or | ✓ (a number) | ✓ | as above |
| text, number, date | ✓ | | ✓ | ✓ | never authorises |
### Authorising tiers × proofs
| Tier | Needs | Telegram | Desk with a key | Desk without a key |
|---|---|---|---|---|
| acknowledge | `choice`, `exact-render` | tap | click | click |
| approve | + one proof | tap (P1) | click + touch (P3) | click + code (P2) |
| destroy | + two proofs, one of them P2 or P3 | tap + code (P1 + P2) | click + touch + code (P3 + P2) | carried by Telegram |
### Surfaces × declared capabilities
✓ declared, ~ conditional, blank not.
| Surface | deliver | reaches-away | loud | silent | edit | choice | reply | threads | operator-first | verified-sender | exact-render | code-factor | key-factor | private | reaches-when-mesh-down |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| telegram | ✓ | ✓ | | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ (placed apart from agents) | ✓ | ✓ | | | |
| desktop (dunst + launcher) | ✓ | | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | | | ✓ | ✓ | ~ (a key plugged in) | ✓ | |
| matrix (own server) | ✓ | ~ (push unmeasured) | | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | | ~ | |
| ntfy | ✓ | ✓ | ✓ | ✓ | | ~ (http action) | | | | | ✓ | | | ~ | ~ (ntfy.sh) |
| pushover | ✓ | ✓ | ✓ | ✓ | | ~ (acknowledge) | | | | ✓ | ✓ | | | | ✓ |
| mail (outside relay) | ✓ | ✓ | | ✓ | | | ✓ | ✓ | ✓ | | ✓ | ~ | | | ✓ |
| an agent's terminal | — | | | | | | ~ (relayed) | | | | | | | | |
| the console | — | | | | | | | | | | ✓ (own output) | ✓ | | | |
| the watcher's sender | ✓ | ✓ | | | | | | | | | | | | | ✓ |
### Where things go, by context
The full table is in [07](07-the-work-context-and-the-desk.md). In one line each:
- **in a Telegram conversation:** that chat;
- **at the desk:** the desk, or the away channel when the desk cannot carry it;
- **away:** the away channel;
- **night:** urgent only;
- **unanswered:** the next in the chain;
- **the desk locks:** it moves away.
## The proposed record
> **Title.** The mesh holds a conversation with its operator over channels that are seats, chosen by
> capability and work context, and an answer that performs an action is authorised by the controller.
>
> **Context.** The output channel was built in its minimal form (ADR 0227, to-be 45 §5): one holder
> carrying the router, a Telegram client and a desktop adapter, and no answering back.
> - The operator expects many channels and many inputs.
> - The operator wants agents and modules to ask questions, not only for permission.
> - The operator wants the work context to choose the channel, and every authorisation completable on
> the away channel.
> - Today, actions that need a person are verbs any granted principal can call, agents included, and
> a hand-act records the calling principal, not the person.
>
> **Considered options.**
> 1. Channels as contributions to the output seat. A channel is running code with a secret and
> answers, not content a holder places.
> 2. One seat per channel kind. The router learns every seat.
> 3. Channel modules found by a manifest field. Callers use seats, never modules (ADR 0126).
> 4. One notifier with every channel built in. Shared failure, shared secrets, a release per channel.
> 5. A per-service module with its own approval or question path. Locks the conversation to one
> service.
> 6. **Kinded benches for out and in, a capability vocabulary, a router that holds the conversation
> and orders channels by work context, and an authorising layer held by the controller. Chosen.**
>
> For answers:
> - a webhook, or **long polling (chosen)**;
> - authorising actions called directly by the channel module, or **requested and answered through
> the controller (chosen)**;
> - trusting a desktop click, or **requiring a code or a key touch the controller verifies (chosen)**.
>
> **Decision.**
> - **The conversation.**
> - Two mesh seats are kinded benches: `channel` (send, edit, standing) and `intake` (one envelope
> per input).
> - Each holder is its own module, claims one kind, and declares capabilities from
> `channel-capabilities/1`, each with a contract test and a drill.
> - The output seat's holder routes messages and asks (yes-no, one-of, text, number, date,
> acknowledge) by required capability, then by work context, then by severity, escalating when
> unanswered. It says when nothing can carry something.
> - Asks have timeouts, defaults (never for an authorising ask), cancellation, a per-asker limit,
> batching and 30 days of history. Answers return to the asker as events.
> - Presence is current state on the bus, never a history, never in a message's words.
> - **Asks that authorise.**
> - The controller holds them: `authorise request` (anyone; never performs), `authorise answer`
> (intake holders only), `authorisations`.
> - It checks the caller's declared capabilities from its own records, the sender against its own
> list of the operator's identities, codes and key assertions itself, and the exact state's digest,
> single use and expiry.
> - Tiers acknowledge, approve and destroy require none, one and two proofs. The away channel must
> satisfy every tier, checked by the self-check, unless the operator chose otherwise for a tier.
> - No agent authorises. The authorising verbs refuse direct calls except as break-glass with a
> code.
> - The hand-act records `via`, `requested-by`, `ask` and `proofs`.
> - **First holders.**
> - Telegram is the first holder of both seats and the away channel, placed where no agent runs as
> the operator.
> - The desktop holds both for the desk.
> - The watcher's watcher stays outside the seats with its own bot, and an outside dead-man service
> is pinged by the self-check and the watcher.
>
> **Consequences.**
> - The output seat's holder loses its Telegram client to a module of its own and gains asks and the
> work context.
> - The catalogue gains `kind` and `capabilities` on a claim, and a second kind of bench.
> - `node-notifier.send` gains actions.
> - The controller's verb definition gains `authorises`, and `retire approve` takes the set it
> approves.
> - Agents' grants lose authorising verbs.
> - To-be 45 §5 is amended: the operator answers.
> - Telegram sees the words, held to the content rule, and never a factor's secret.
>
> **How it is checked.**
> - **Catalogue tests:**
> - an unknown capability is refused;
> - two holders of one kind are refused;
> - each declared capability's contract test runs in its holder's build.
> - **Router tests:**
> - an ask goes only to channels whose capabilities satisfy it;
> - context reorders but never adds a channel;
> - an authorising ask never defaults;
> - a cancelled ask's copies are edited;
> - a fourth open ask from one asker is refused.
> - **Controller tests,** one per refusal of `authorise answer`:
> - from a non-intake principal;
> - from a kind that does not meet the tier;
> - from an identity not on the list;
> - with too few proofs;
> - with a used code;
> - with a key assertion over another ask's challenge;
> - with a stale digest;
> - a second answer.
>
> Also: a direct `retire approve` without a code.
> - **Self-check probe:** the away channel meets every tier.
> - **Live drills:**
> - an agent's question answered at the desk;
> - the same with the desk locked, answered on the phone;
> - an approve and a destroy on a test condition, answered on the phone, with the hand-acts read.
@@ -1,66 +0,0 @@
---
status: graduated
initiated: 2026-10-06
touches:
- 00-META/how-we-build.md
- 01-RESEARCH/017-a-mesh-that-heals-itself/00-overview.md
- 01-RESEARCH/028-the-meshs-output-channel/00-overview.md
- 01-RESEARCH/019-a-warm-twin-of-the-running-mesh/00-overview.md
- 02-DECISIONS/0141-the-host-delivers-its-own-successor.md
- 02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md
- 02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md
- 02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md
- 03-DESIGN/01-to-be/06-the-controller.md
- 03-DESIGN/01-to-be/09-the-node-lifecycle.md
- 03-DESIGN/01-to-be/25-the-bus-on-nats.md
- 03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md
- 04-ISSUES/187-the-mesh-tells-nobody-when-it-stops-working/00-report.md
became:
- 02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md
- 03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md
---
# 031 — A core that cannot fail silently
**What.** The principles the mesh's core must hold — the controller, the machine host, the bus, the
console and the path a change takes through them — so that it is fully diagnosable, monitors itself,
heals what it knows how to heal, and upgrades itself without a person standing by. And the mechanisms
and the order in which to build them.
**Why.** The operator, 2026-10-06: *"I still notice a lot of race issues, and commands being ignored, or
no feedback, no logs, no monitoring. Our mesh core setup must be fully diagnosable, with active
monitoring, self-healing, self-upgradeable, self-monitoring. The core principles must be very sturdy, no
ambiguities, clear plan of execution, fail-proof setup."*
The record bears it out. In the six days to 2026-10-06, 92 issue reports were opened. Of the 48 read here
as core failures, **every one was noticed because a person or an agent looked**, and **none was raised by
the mesh unasked**. Four faults came back through a different door after their first fix, because each
fix closed an instance and left its class open. One merge was skipped by the bus, and twenty-three over three
days have no matching action; a provider failed for twenty-three hours with only its own
journal saying so; seven databases were dropped on one unreadable file.
**What it touches.** The controller (its verbs, `status`, plans, a lease), the host (its apply and
report), the bus (its advisories and its upgrade), the console, the build path, the output channel of
research 028, and the self-healing intent of research 017, which this effort extends from the loops that
converge modules to the core that runs those loops. 017 deferred heartbeats, conditions and advisories
until the bus was NATS; it is now.
**Documents.**
- [01 — The evidence](01-evidence.md): 48 issues classified by class of failure (races, dropped
commands, no feedback, logs only, two writers, manual repair, self-upgrade, CI-versus-live, third
party), with time to detect and how each was noticed.
- [02 — Principles](02-principles.md): nine, each with what exists, what is missing and **how it is
checked**; and the candidates weighed and not kept.
- [03 — Mechanisms and roadmap](03-mechanisms-and-roadmap.md): conditions, watchdogs from a signals table,
a self-check (`doctor`), the output channel, healers with a hand-act log, a controller lease and report
sequences, staged core upgrades with rollback, a facts snapshot for merge checks, a lab replay of every
incident; five phases, ordered by risk removed; the five largest risks today.
**Graduated 2026-10-06.** The operator approved the conclusion the same day (*"do the research and
implement it"*). The nine principles and the plan became
[ADR 0227](../../02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md);
the mechanisms, the tables and the six phases became
[to-be 45](../../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md). The two measurements owed
— the signals table's bounds and a week of the hand-act log — are Phase 0's work there, not
preconditions of the decision.
@@ -1,202 +0,0 @@
# 01 — The evidence, classified by class of failure
Every issue report opened between 2026-09-30 and 2026-10-06 that bears on the mesh's core — the
controller, the machine host, the bus, the console, the build path — read in full and classified by
**the class of failure**, not by the component it was found in. A component view says "fix the host";
a class view says "the same thing is wrong in four places", which is what a principle is for.
## The count
| | |
|---|---|
| issue reports opened 2026-10-01 to 2026-10-06 | **92** (about fifteen a day) |
| of those (and a few from the days before), classified below as core failures | **48** distinct issues |
| classified in more than one class | 17 of 48 |
| a fault that **came back** after a fix of the same symptom | 4 chains: 200 → 265, 230 → 264, 257 → 261 → 267, 175 → 184 → 248 |
| noticed because a person or an agent looked — at a stalled plan, a wrong outcome, a journal, a test run by hand, a review | **48 of 48** |
| of those, the mesh's own answer carried the fact for whoever asked (a refusal, a `status` line, a push's output) | 4 (233, 244, 259, 263) |
| raised by the mesh to anyone, unasked | **0 of 48** |
The recurrences matter most. Each fix was correct for its instance and left the class standing, so the
same symptom came back through a different door days later. That is the measurement behind the
operator's mandate: point fixes are converging on the instances, not on the class.
## The classes
Nine classes, as the mandate frames them. The table under each is the evidence; *detected* is the time
from the fault's start to the moment anyone knew; *noticed by* is how.
### (a) Races: concurrent actors without an ordering
Two actors act on the same thing, and the order of arrival — not an explicit order — decides the outcome.
| Issue | The two actors | Detected | Noticed by |
|---|---|---|---|
| 204 | an outgoing and an incoming controller both sent declarations | 2 min | a person saw a module undone |
| 201 | a plan's push carried a controller digest older than its successor had written | 10 min crash loop | a person, nothing answered |
| 214 | the controller rebuilding itself; the outcome reached the old one or neither | 27 min | a person asked the plan twice |
| 219 | an older build finishing later replaced a newer one | hours | a person reading builds |
| 234 | seven declarations arrived during an eleven-minute apply; the newest was composed from a stale view and undeclared four modules | 8 min of removals | a person, the modules were gone |
| 254 | three plans for three merges ran at once, each asking the same builds | hours | a person, one plan stuck "building" |
| 256 | a first machine's report landed between a module's two sends and read as stale | 7 min | a person |
| 257, 261 | the machine's five-minute reconcile and a delivery took the apply lock in the wrong order | 5 min (257), 29 s visible undo (261) | a person |
| 267 | a reconcile's report queued behind a delivery's apply overtook it at the controller | until a hand push | a person |
| 265 | a push reloaded the bus's permissions while its own answer was still owed | 54 of 103 pushes over two days | a person, "did not answer in time" |
**What they share.** Every one is a receiver that kept *the last thing written* rather than *the
newest thing by an explicit order*. Declarations carry a sequence (issue 107); **reports do not**
(267 says so: "the report does not carry the declaration's sequence number, so the digest decides").
Plans are ordered by when they were made only since ADR 0218. Builds are ordered since issue 219.
Controllers have no epoch, so two instances can both act (204). Ordering was added one message kind at
a time, each after a race in it was seen.
### (b) Commands silently ignored, arguments dropped
| Issue | What was dropped | Effect |
|---|---|---|
| 244 | the console removed `node` from every mesh-seat verb's schema and call | `plan` could only refuse; **`push <one machine>` arrived empty and pushed every machine** |
| 259 | a named push's flush sent every machine a build a policy held back | a fault met on every machine at once, not one |
| 202 | a module whose setting was unset was *left out* of the machine | the resolver vanished from a declaration, no error |
| 188 | a refusal inside "who is on the network" dropped a machine | 40 min, every symptom pointed elsewhere |
| 231 | a misspelled placeholder written to a file as literal text | passes every check |
| 241 | an unreadable contributions file read as "nobody asks" | **seven databases dropped and recreated empty** |
| 255 | the journal verb read nothing and said "-- No entries --" | a refusal that reads as a quiet service |
| 246 | a runtime that answered late was treated as absent | the console said modules "run nowhere" |
**What they share.** A receiver that could not tell *nothing was asked* from *something was lost on the
way*, and chose a default. In 241 and 244 the default was the most destructive reading available.
### (c) Outcomes not fed back to the caller
| Issue | What the caller was told | What happened |
|---|---|---|
| 200, 265 | "did not answer in time" | the push ran; the answer was refused by the bus |
| 176 | the console's build tool neither waits nor registers | — |
| 229 | `plans` answers once in prose; nothing waits for a plan | an agent went round the mesh with `curl` |
| 230, 264 | a host stood aside for its successor and its report was cancelled | the plan waited for ever, reading `late: false` |
| 186 | the build machine dropped 26 of 43 asks; nothing counts asks against outcomes | inferred two hours later |
| 237 | `assign` answered "held" and "does not resolve for lack of it" in one breath | a person or agent would loop |
Since 2026-10-06 the controller answers within ten seconds and keeps every call's outcome under an id
(`calls`, issue 265). Read live the same night: **that log holds the last hundred calls in the
controller's memory**, so a controller restart — which every merge to the controller's own repository
causes — forgets every outcome it held. And `status`, a read-only verb, took **18 seconds** to answer,
twice in a row, so even the health question is answered only through the "still running, ask `calls`"
path.
### (d) Failures visible only as log lines
| Issue | Where it was said | For how long |
|---|---|---|
| 179 (recurred) | the identity provider's journal, every five seconds | **23 hours**, about 31 000 refused logins |
| 184 | the controller's log: slow consumer, heartbeats dropped | 24 min deaf |
| 187 | five faults in one day, each found by reading a container's log hours later | hours each |
| 183, 217, 265 | a `Permissions Violation` line from the bus client library | days |
| 233 | `status` said `refused`, correctly; nothing said it had lasted | 1.5 h |
| 243 | nothing: machines silently ignored lower licence generations | until a login waited three minutes |
| 248 | the controller's event loop stopped logging at 15:17 | hours; "status showed every plan done" |
| 266 | nothing: a merge was skipped by the bus | **23 unmatched merges over three days** |
| 238 | a ban list of 400 entries | the operator's own address banned for four weeks |
`status` printed its all-well sentence through 179, 248 and 266. ADR 0224 made the first of those break
it. The other two have no signal that `status` reads.
### (e) State that two writers own
| Issue | The two writers |
|---|---|
| 190, 222 | the runtime's configuration written by modules that are not the runtime, and by the controller |
| 201 | the controller seat's row written by a successor, read by a predecessor pushed back in |
| 239 | two definitions, in two repositories, held one module name |
| 245 | "behind" answered by a commit comparison beside the plan that already knows |
| 257, 261, 267 | the machine's state written by both the delivery and the five-minute reconcile |
| 250 | a merge announced by the forge's tool and by its poll |
| 179 | the identity provider's admin password: the mesh minted one, the database kept another |
The operator's direction on 245 is the principle in their own words: *"a second answer to the same
question is how the two came to disagree."*
### (f) Manual repair needed
Counted from the reports' own "what unblocked it" sections:
| Repair by hand | Issues | Times |
|---|---|---|
| a push by hand to unstick a plan waiting on a report | 230, 257, 264, 267 | at least 4 |
| a controller restart to recreate a missing object or let go of a held message | 208, 248 | 2 |
| a one-off program run as the controller, outside the service | 201, 248 | 2 |
| the identity provider's admin reset through its bootstrap command | 179 | 2 |
| a kept file restored on a machine by hand | 233 | 1 |
| a stuck plan closed by hand | 214, 254 | 2 |
| a ban lifted by hand | 238 | 1 |
| a consumer remade from now | 248 | 1 |
ADR 0224 (*detected automatically, repaired where safe, loud where not*) is the first rule that turns
one of these into a mechanism. 248's `broker consumer-reset` and 254's `plans close` turned two into
verbs a person runs. Every other row is still a hand on a machine.
### (g) Self-upgrade fragility
The core updates itself: the controller rebuilds and replaces itself, the host delivers its own
successor (ADR 0141), the runtime and the console are modules, and the bus is a module on the control
node.
| Issue | What the self-upgrade broke |
|---|---|
| 201 | the controller pushed back to an older build than its own successor's row |
| 204 | two controllers both sending during a handover |
| 213, 223 | the controller ran as a container, and a new mesh installed it so |
| 214 | the plan that rebuilds the controller lost track of it |
| 230, 264 | the host that stands aside loses the report of the apply that delivered it (fixed twice) |
| 245 | a rebuild of everything replaced the bus's container: **every runtime lost the bus for a minute** |
| 248, 266 | a controller restart is where merges go missing: most of 266's 23 lie in such windows |
| 217 | a refused announcement crash-looped nine runtimes, closing the path that would merge the fix |
A machine's first-in-line rollout (ADR 0218) protects modules. It does not protect the core from itself:
the health a plan waits for is "reported applied", which a controller that cannot plan, a host that
cannot report or a bus that drops messages can each satisfy. Nothing rolls back. The recovery in 201
was the mesh's own binary run by hand from the newer image.
### (h) Checks that pass in CI and fail live
| Issue | The environmental fact the check did not have |
|---|---|
| 177 | the store-backed tests are skipped by the quick check, and the mesh runs none of a module's tests |
| 236 | the host's declaration validation is not run by the catalogue check |
| 262 | musl takes an NXDOMAIN for IPv6 as final; glibc does not |
| 263 | the real machine names make a consumer's identity 23–26 characters against a bound of 20 |
| 202 | the controller's test against the real catalogue, run by nobody until that day |
| 228 | the host's removal has no case for a `user` — found by a review, not a test |
Each check was right about the world it was given. None of them was given the mesh's world: its machine
names, its catalogue, its host's validation, its C libraries.
### (i) Third-party bugs
| Issue | |
|---|---|
| 266 | the bus server's 2.10 release skips messages on a consumer with several filter subjects |
| 265 | a reload of the bus's authorization forgets every reply permission already granted (documented server behaviour, read from its source) |
| 262 | a resolver answering NXDOMAIN where NODATA is correct, met by musl's stricter reading |
The lesson of 266 is not "upgrade the bus" — it is that nothing compared *what was announced* with
*what was acted on*, so a dependency's bug was silent for three days. A defence in depth (watch the
outcome, not the transport) would have caught it whichever layer was wrong.
## What would have prevented or caught each class
| Class | Would have been prevented by | Would have been caught by |
|---|---|---|
| (a) races | every message ordered by its writer, stale refused by every receiver | a lab replay of the interleaving |
| (b) dropped | refusing an unknown or unreadable input by name | a schema walk over every verb |
| (c) no feedback | answer at once with an id; outcome kept durably | a watchdog on "asked and never finished" |
| (d) logs only | — | a condition in `status` and a notification |
| (e) two writers | one writer per piece of state | a registry of writers checked at composition |
| (f) manual repair | a healer for every repair done twice | a counter of hand acts |
| (g) self-upgrade | one machine first, health-gated, rolled back | a lab upgrade with a deliberately broken build |
| (h) CI vs live | checks fed the real mesh's facts | the same, before merge |
| (i) third party | pinning and testing the version that runs | an end-to-end count of announced vs acted |
The two columns are the principles of [02](02-principles.md). Read by count, **the "caught by" column
is the cheapest and widest**: a watchdog and a condition would have shortened most of the 48 from "a
person noticed" to minutes, whatever the cause.
@@ -1,247 +0,0 @@
# 02 — Principles for the core, each with how it is checked
Nine principles. Each is stated as a rule a reviewer can refuse a change against, carries the classes of
[01](01-evidence.md) it answers, says what exists already, and says **how it is checked** — the
repository's own rule ([how-we-build §5](../../00-META/how-we-build.md)): a rule that states no check is
indistinguishable from a wrong one.
They extend, not replace, the six of [research 017](../017-a-mesh-that-heals-itself/01-the-intended-behaviour.md)
(a loop compares with what is; healing is the ordinary path again; a repair never destroys; nothing fails
silently; what cannot be fixed goes to an agent; correctness, not only liveness). Those are about the
loops that converge modules. These are about **the core that runs those loops**: the controller, the
host, the bus, the console and the path a change takes through them. 017 deferred heartbeats, conditions
and advisories until the bus was NATS. It is now, so that deferral has expired.
**The core**, for this document: the controller, the machine host and its launcher, the bus server, the
tool runtime and the console, the build seat, and the forge's announcer of merges. Everything a change
passes through before a module's own code runs.
---
## P1 — One writer per piece of state
Every piece of state the mesh keeps has exactly one writer, named. Anyone else who wants it changed asks
that writer; nobody writes beside it, and nobody computes a second answer to a question it already
answers.
- **Answers:** (e), most of (a). Issues 190, 201, 204, 222, 239, 245, 250, 257/261/267.
- **Exists:** the controller is the only writer of stream definitions (to-be 25); ADR 0222 §3 (the
controller writes no file a seat's holder owns); the collision check at composition (no two modules
declare one path, unit, name or package); the operator's direction on 245.
- **Missing:** a single writer for a *machine's applied state* (the delivery and the five-minute
reconcile both apply and both report — three issues in two days); a single *controller* (two instances
can both act during a handover, 204: nothing holds a lease); a single announcer per event kind (250
was found by counting duplicates by hand).
- **How it is checked:**
1. A **writers table** — state kind, its writer, where it is kept — is part of the to-be design, and a
test in each core repository asserts that the code paths that write each kind are the one named
(by a lint over the store's write calls and the bus subjects each component publishes on, the
latter read from the grants the controller composes: a subject two components may publish on is
refused at composition unless the table says it is shared).
2. **Live:** the controller holds a **lease** (a key-value entry with a revision) and every message it
sends carries the lease's epoch; a host refuses a declaration from an older epoch and says so. A
probe ([03](03-mechanisms-and-roadmap.md), the self-check) asserts one lease holder and no message
from a stale epoch in the last interval.
## P2 — Everything that changes state carries its writer's order, and every receiver refuses what is older
Declarations, reports, plans, builds, calls and announcements each carry `(writer, epoch, sequence)`.
Every receiver keeps the highest it has accepted per writer, and **refuses** — with a line in the mesh's
own words and a counter — anything older. Arrival order never decides.
- **Answers:** (a). Issues 201, 204, 214, 219, 234, 256, 257, 261, 264, 267.
- **Exists:** a declaration carries a sequence and a host keeps the newest (issue 107, to-be 25 §3); a
newer build wins over an older one finishing later (219); a newer merge's plan supersedes an older
one (ADR 0218 §3); a report about a declaration the mesh has moved past no longer replaces the stored
account (267).
- **Missing:** a **report carries no sequence** — the digest last recorded as sent decides, which is why
each of 256, 257 and 267 needed its own rule. No epoch on the controller. A plan's state carries no
revision, so two instances can both advance it (214).
- **How it is checked:**
1. A contract test per message kind, in the receiver's repository: deliver `n`, then `n−1`; the state
names `n` and a refusal is counted. Deliver from epoch `e−1` after `e`: refused. A new message kind
without such a test fails a check that lists every subject the component consumes against the
tests that name it.
2. **Live:** the refusals counter is a signal (P5): zero is normal, a burst is a condition naming the
writer that sent stale.
## P3 — Every command is acknowledged at once, and its outcome is kept where it can be read later
A call is answered within a bound the caller can rely on — with its result, or with an id. Its outcome is
kept **durably**, outlives the process that ran it, and can be read by id or waited on. Nothing is fired
and forgotten, and no answer depends on what the command does to the transport carrying it.
- **Answers:** (c). Issues 176, 186, 200, 229, 230, 237, 264, 265.
- **Exists:** since 265, a seat's call answers within ten seconds or says "running" with an id, `push`
answers before it acts, and `calls` keeps the last hundred calls and their answers; the build path
says "asked, not waited for" and ADR 0219 §3 makes every cancel/kill leave an outcome.
- **Missing:** `calls` lives in the controller's memory, so **a controller restart forgets every
outcome** — and the controller restarts on every merge to its own repository. Nothing waits on a plan
(229). Observed the night this effort began: the read-only `status` took 18 s, so the health question
itself is answered through the "still running" path.
- **How it is checked:**
1. A test that walks every verb the controller announces: each answers within `AnswerWithin`, with a
result or an id (already partly built for 265).
2. A test that restarts the controller between a call and the read of its outcome: the outcome is
still there.
3. **Live:** a probe calls `status` and asserts it answers *in full* within the bound; a call
`running` for longer than its verb's declared bound is a condition (P5).
## P4 — Nothing is dropped silently: an input that is unknown, unreadable or unmet is refused by name
A receiver that cannot read, place or honour an input refuses it and says which, where and why. It never
substitutes a default — above all never "empty" — for "I could not tell". An unknown argument, key,
placeholder, seat or subject is refused, naming it.
- **Answers:** (b). Issues 188, 202, 231, 241, 244, 246, 255, 259.
- **Exists:** the host's parser refuses unknown keys (ADR 0007, 0045); an undeclared setting or endpoint
is refused (ADR 0164, 0138); a verb takes only its declared arguments (ADR 0154) and, since 244, the
controller and the console refuse an undeclared one by name; a seat the mesh does not answer is
refused by name (ADR 0222 §1); ADR 0219 §3, "nothing dropped is silent".
- **Missing:** the rule is in a dozen records and in no principle, so each new reader re-meets it. The
destructive form — **an unreadable input read as "nothing asked", then acted on** (241) — has no
general guard.
- **How it is checked:**
1. Per component, a test that feeds each input reader an unreadable, malformed and foreign input and
asserts a refusal, never an empty result. A reader whose error path returns an empty value fails a
lint that the core repositories run (the shape is mechanical: an error branch that returns the
zero value of a collection).
2. Schema walks for every verb (244's tests) and every placeholder namespace (231).
3. **Destructive deltas are braked**: a reconcile that would withdraw more than a bound of what it
holds (one consumer, one module, a fraction set per provider) stops and raises a condition instead
(P7). Checked by a test that empties the input and asserts nothing is withdrawn.
## P5 — Every expected signal has a watchdog: absence is itself a condition
Every signal the core expects on a cadence or after an act — a heartbeat, a report after a send, a
plan's progress, a build's outcome after its ask, the controller's event loop taking something in, a
provider's standing, an announcement turning into an action — has a declared bound. Silence past the
bound is raised as a condition, naming what was expected, from whom, since when.
- **Answers:** (c), (d), (i). Issues 179, 184, 186, 187, 230, 243, 248, 257, 264, 266, 267.
- **Exists:** a machine's "last heard from — out of touch N m" in `node show`; ADR 0090's stuck machine
(three identical reports); ADR 0224's provider standing, shown with its silence after thirty minutes;
266's catch-up of merges not acted on after ten minutes (the first true *announced-versus-acted*
watchdog); ADR 0162's bound on a plan, which 230 found never fires (`late: false` for ever).
- **Missing:** a list of the signals, their bounds and their owners; the plan bound working; the event
loop's last-taken age (187); asks counted against outcomes (186); the bus's own advisories (slow
consumer, maximum deliveries, permission violations) read as observations instead of log lines.
- **How it is checked:**
1. A **signals table** — signal, emitter, cadence or trigger, bound, condition raised — kept in the
to-be design, and compiled into the controller. A test generated from it suppresses each signal in
turn and asserts the named condition is raised within its bound and cleared when the signal
returns.
2. **Live:** the self-check (P6) reports, for every row, the age of the newest signal, so a row that
never fires is itself visible.
## P6 — The mesh checks itself continuously, against live facts, and says what it found outward
The invariants the design states are probed **against the running mesh** on a schedule, not only in unit
tests. A violation is a **condition** — durable, with since-when, evidence and who can resolve it —
shown in `status` and **sent to the operator** through the output channel. The checker's own heartbeat is
watched from somewhere it does not run.
- **Answers:** (d), (h). Issues 177, 187, 238, 245, 253, 262.
- **Exists:** `status` itself (assembled from reports, as how-we-build §5 requires); 017's *condition*
shape; 028's output seat, researched only; individual live checks done by hand (262: "the live check
stays by hand, on each resolver"); 253's measurement before the collector's first run — the one time
in the window a check ran before the damage.
- **Missing:** a scheduled runner, a condition store, a channel out, and a watcher's watcher.
- **How it is checked:**
1. Every invariant in the to-be design that names a live check is a **probe** in the self-check's
registry; a check over the design documents counts invariants with a stated live probe against
those without, and the number may only go down.
2. The self-check publishes a heartbeat; a second machine's watcher raises "the self-check is silent"
through a channel that does not depend on the control node.
3. **Live, once:** a deliberately broken invariant on a lab mesh appears in `status` and as a
notification within one probe interval.
## P7 — A known failure heals itself, under a brake, and every repair is said
A failure that has been repaired by hand twice is a failure the mesh must repair itself: by running the
ordinary path again (017 P2), never by destroying (017 P3), with a budget and a back-off, and with one
line and one event saying what it did and why. When the budget is spent, or the only repair destroys,
it is a condition and a notification, not a retry.
- **Answers:** (f). Issues 179, 208, 214, 230, 233, 248, 254, 257, 264, 267.
- **Exists:** ADR 0224 §5, *detected automatically, repaired where safe, loud where not*, applied to the
identity provider's admin; the host's launcher rolls back once to known-good and then halts (ADR
0141); the provisioner's `holds` re-provisions what a backend lost (017/02); `broker consumer-reset`
and `plans close` as verbs.
- **Missing:** the general mechanism. The most frequent hand act in the window — **a push by hand to
unstick a plan waiting on a report** — has no healer: the controller could ask the machine to report
again (the machine knows what it applied) before waiting longer.
- **How it is checked:**
1. Every hand act on the core is done through a verb that records it (who, what, why) — the **hand-act
log**. Its count per week is a reported number; an act recorded twice for the same cause is a
condition asking for a healer.
2. Each healer ships with a test that induces its failure, asserts the repair and the event, and
asserts the brake after the budget.
## P8 — The core upgrades itself one machine at a time, health-gated, and rolls back on its own
A new controller, host, runtime or bus reaches one machine first; it is **healthy** only when the
self-check's probes for that component pass there (not merely when it "reported applied"); the rest
follow only then. A component that does not become healthy within its bound is rolled back to the last
known good **by something other than itself**, and the rollback is said. The component being replaced is
never the only witness of its successor's success.
- **Answers:** (g). Issues 201, 204, 213, 214, 217, 230, 245, 248, 264, 266.
- **Exists:** ADR 0218's one machine first, for modules, with "applied and current" as the gate; ADR
0141/0142's side-by-side host versions, known-good and the launcher's single rollback; ADR 0185's
controller serving what it can when it is behind its seat's row; 264's report kept on disk across the
hand-over.
- **Missing:** a health definition per core component; a gate stronger than "reported"; rollback for
the controller, the runtime and the bus; a lease hand-over between controllers (P1); a planned,
rehearsed path for the bus, which is still one process on one machine and whose next upgrade (266:
2.10 → 2.11) is one-way.
- **How it is checked:**
1. **Lab:** a deliberately broken build of each core component (one that starts and does nothing; one
that crashes; one that cannot reach the bus) is merged on a lab mesh. Each is rolled back without a
hand, the mesh ends on the previous build, and a condition and a notification say so.
2. **Live:** every core rollout leaves a record — first machine, health verdict, time to verdict,
rolled back or not — readable through `plans`.
## P9 — A check is fed the real mesh's facts before a change is merged
A check whose verdict depends on the environment — names and their lengths, the machines that exist,
the catalogue as it is, the host's validation, the C library, the server versions — runs against **the
mesh's real facts**, exported and anonymised, before merge. A dependency's version that the mesh runs is
the version its tests run.
- **Answers:** (h), (i). Issues 177, 202, 228, 236, 262, 263, 266.
- **Exists:** 266's `TestTheImageIsTheServerTestedHere` (the bus image's release equals the tested
server's); ADR 0223's composition test that renders the resolver's machine list; 202's test against the real
catalogue (run by hand).
- **Missing:** the export of facts; a merge gate that composes every real machine's declaration with the
change and runs the host's validation over it (which would have refused 236, 263 and 202 in their own
pull requests); a resolver test under musl as well as glibc.
- **How it is checked:**
1. The controller exports a **facts snapshot** (machines, names, assignments, seats, catalogue
commit — no secrets, no addresses) daily; the core repositories' merge check composes every machine
from it with the change applied and runs the host's validation; a pull request that makes any
machine fail to compose or validate fails its check, naming the machine's role and the module.
2. The snapshot's age is a signal (P5).
---
## Candidates weighed and not kept as principles
- **"Every invariant has a live probe, not only a unit test"** — merged into P6; it is how P6 is built.
- **"Environment-dependent checks run against the real mesh's facts"** — kept as P9; the third-party
case (i) folded into it, because pinning and testing the version that runs is the same act.
- **"Self-healing is the default"** — kept as P7 but narrowed to *known* failures, those repaired by
hand twice. A default of healing everything heals what is not understood, which is how a repair
destroys (241's reconcile was, in its own terms, healing).
- **"No loop blocks on long work"** (184, 248, 175) — not a separate principle: a blocked loop is a
signal gone silent (P5, the loop's last-taken age) and a design defect each owner fixes; stating it
as a principle adds a rule with no mesh-wide check.
- **"Clear plan of execution"** from the mandate — not a principle about the mesh; it is the roadmap in
[03](03-mechanisms-and-roadmap.md), and each phase there states its own verification.
## How the principles relate
P2 and P1 **prevent** the races. P4 **prevents** the silent drops. P3, P5 and P6 **catch** whatever the
first three miss, at the cost of minutes, not hours. P7 and P8 **repair**. P9 **moves** the catching
before merge. The order of the roadmap follows from that: catching first, because it is cheapest and
covers every class, including the ones nobody has met yet.
@@ -1,221 +0,0 @@
# 03 — Mechanisms and a phased roadmap
The principles of [02](02-principles.md) need few new things. Most of the parts exist in some form;
what is missing is the connective tissue that makes a fact the mesh already has reach someone without
being asked. This document names the mechanisms, then orders them into phases **by risk removed per unit
of effort**, each phase with deliverables and a verification that says it is done.
Effort is given in **focused working days** of one agent-and-operator pair, at the pace the record shows
(a located issue to a merged fix in under a day is common). The figures are for ordering, not promises.
## The mechanisms
### M1 — Conditions (P5, P6)
017's *condition*, built: a durable fact about something the mesh owns — what is wrong, since when, the
evidence, what was tried, who can resolve it, and whether it may clear itself. Kept in a key-value
bucket the controller writes (one writer, P1), keyed by subject (`plan/<id>`, `machine/<role>`,
`provider/<module>/<consumer>`, `core/<component>`). Raised and cleared by observation only; a person
can **silence** one for a stated time, never resolve it. `status` becomes, first, the list of open
conditions; the all-well sentence is "no open conditions". ADR 0224's provider standing is the first
condition kind and moves into it unchanged.
### M2 — Watchdogs from a signals table (P5)
One table, compiled into the controller, of every signal the core expects. Its first rows, each from an
issue in [01](01-evidence.md):
| Signal | Bound (to be measured, then set) | Condition raised | Issue |
|---|---|---|---|
| machine heartbeat | 3 × interval | machine silent (asleep is a declared state, ADR 0211) | 187 |
| report after a send | the machine's last apply duration × 3, at least 2 min | sent, not reported — then **ask the machine to report again** (M5) | 230, 257, 264, 267 |
| plan tier progress | per tier, from build and apply durations | plan stalled at tier N, waiting on X | 214, 230, 254 |
| controller event loop took something | 2 min while the stream has pending | controller deaf | 184, 248 |
| merge announced → plan made or "nothing reads it" | 10 min (exists, 266) | merge never acted on | 248, 266 |
| build asked → outcome | build's own declared timeout | ask lost | 186 |
| call `running` → finished | the verb's declared bound | call hung | 265 |
| provider standing repeated | 30 min (exists, ADR 0224) | provider silent | 179 |
| bus advisories: slow consumer, maximum deliveries, permission violation | any | bus refused or dropped X for Y | 183, 187, 217, 265 |
| self-check heartbeat | 2 × its interval, watched from a second machine | the watcher is silent | — |
| facts snapshot age | 2 days | merge checks run on stale facts | 263 |
The bus advisories are the cheapest row: the server already publishes them on its system subjects, and
the controller only has to subscribe (read-only) and translate each into the mesh's words, naming the
call or the consumer, as 265 now does for a refused reply.
### M3 — The self-check: `doctor` (P6)
A controller verb, `doctor`, and the same code run every few minutes by the controller itself. Each run
executes the **probe registry** — the live form of the design's invariants — and raises or clears
conditions. First probes, each an invariant that a person checked by hand in the window:
- every machine's declaration composes, and every host would accept it (236, 263);
- every holder of the mesh's resolver answers a machine name for IPv4 and NODATA for IPv6 (262);
- every seat on record has a live holder that answers (208, 218);
- every kept archive is held by a manifest (253 — the controller's collection command already reports it);
- exactly one controller holds the lease (P1);
- every durable consumer's position is near its stream's head (248's replay, 266's skip);
- no address the mesh owns is in a ban list (238);
- `status` answers in full within its bound (P3).
`doctor` with no argument answers the last run's verdict at once (P3) and, with `run`, runs now under an
id. Its own heartbeat is a signal (M2), watched from a second machine.
### M4 — The output channel (P6)
[Research 028](../028-the-meshs-output-channel/00-overview.md)'s seat, built minimally first: **one**
channel the operator chose (028 records it), plus the desktop notifier where the operator is. A condition
is sent when raised, once more if it lasts past a bound, and when it clears. Deduplicated by the
condition's key. The watcher's watcher (028's open question) is the second-machine watchdog of M3,
sending through a channel that does not pass through the control node.
### M5 — Healers (P7)
A healer is a registered response to one condition kind: its repair (the ordinary path again), its
budget, its brake, and the event it emits. First healers, all from hand acts in [01](01-evidence.md) §(f):
| Condition | Repair | Brake |
|---|---|---|
| sent, not reported | ask the machine to report what it last applied (it keeps it since 264); if that names another declaration, send again | twice, then condition |
| plan stalled on a superseded or finished wait | close the plan with its note (`plans close`, done by the mesh) | once |
| seat holder without its worker | raise the seat's objects again (208) | once per holder |
| consumer far behind on a history stream | `broker consumer-reset` (248) — **only** for the consumers the table marks resettable | once, then condition |
| provider admin refuses the mesh's secret | ADR 0224 §5 (exists) | exists |
And the **hand-act log**: every repair a person makes on the core goes through a verb that records who,
what and why. Its weekly count is the measure of P7.
### M6 — Order and epochs (P1, P2)
- The controller takes a **lease** in a key-value bucket before it acts and renews it; its revision is
the **epoch** every declaration and plan write carries. A starting controller waits for the lease;
the outgoing one stops sending when it loses it. That closes 204 and makes 201/214 detectable.
- A **report carries the sequence** of the declaration it is about; the controller keeps the highest
per machine and refuses older accounts by sequence, not by a digest lookup (256, 257, 267 become one
rule).
- **The host has one apply queue.** A delivery and the reconcile are two reasons to enqueue the same
act; the queue applies the newest declaration once, and makes one report (257/261/267 become
impossible rather than handled).
- Every receiver's refusal of something stale is a counted line (P2's live check).
### M7 — Staged, reversible core upgrades (P8)
- **A health definition per core component**, written as probes in M3's registry: the controller
answers `status` in bound and holds the lease; the host has reported its current declaration; the
runtime has announced and answers a PING; the bus has every stream and every durable consumer and
passes a request/reply round trip.
- **The gate:** ADR 0218's first machine is judged by those probes, not by "applied".
- **Rollback by a witness that is not the new build:** the host's launcher for the host (exists); for
the controller, the previous controller's process kept installed beside it, re-started by the host
when the new one does not take the lease in bound; for the runtime, the host's known-good the same
way. Each rollback is a condition, so it is said.
- **The bus is planned, not rolled.** A bus upgrade is a declared maintenance step: streams snapshotted,
the step announced as a condition while it runs, every consumer's position checked after (M3). The
one-way 2.10 → 2.11 upgrade of 266 is the first. Whether the bus should become a cluster of three so
that it can be upgraded live is a question for its own effort.
### M8 — The facts snapshot and the merge gate (P9)
The controller exports a facts snapshot (machines by role and name length, assignments, seats, catalogue
commit) to a place the build seat reads. The core repositories' and the catalogue's merge checks compose
every machine with the change and run the host's validation. The resolver module's tests run under musl
and glibc. A bus, store or library version the mesh runs is the one its tests run (266's pattern,
generalised).
### M9 — The lab replay (all)
[Research 019](../019-a-warm-twin-of-the-running-mesh/00-overview.md)'s warm twin, or a throwaway lab of
three containers, running **scripted replays of each incident** in the window: a reconcile due during a
push; a host self-update during its report; a bus reload during a call; a consumer with several filters
under mixed traffic; a missing consumer made with the server's default; an unreadable contributions
file; a controller rebuilding itself mid-plan; two controllers at once. Each replay asserts the
principle's outcome (refused stale, condition raised, healed, rolled back). They run on every merge to a
core repository.
---
## The roadmap
Ordered by **risk removed per day**. Detection comes first because it covers every class at once,
including the ones not met yet; prevention second; repair and staging third; the pre-merge and lab work
last because they are larger and pay off over months.
### Phase 0 — Finish what is in flight (2–3 days)
- Land and roll out the located fixes: 244, 264, 265, 266 (including the bus's planned 2.11 upgrade, done
as M7's first planned bus step), 267, 257/261.
- Make `calls` durable (a key-value bucket, bounded by count and age), and bring `status` inside its own
answer bound.
- Start the **hand-act log** now, before anything else, so every later phase is measured against a
baseline.
- **Done when:** the four located core issues resolve with their live checks; a controller restart
keeps `calls`; `status` answers in full within ten seconds.
### Phase 1 — The mesh says when it is wrong (5–8 days)
- M1 conditions (ADR 0224's standing moved into them); M2 watchdogs for the first eight rows; the bus
advisories subscribed and translated; M3 `doctor` with the first probes; M4 with one channel and the
second-machine watcher.
- **Done when:** on a lab mesh, suppressing each signal in the table raises its condition within its
bound and sends a notification; clearing it clears both. Live: a week of conditions read back, every
one either real or a bound corrected.
- **Risk removed:** every class in [01](01-evidence.md) moves from "noticed by a person, hours later" to
"said by the mesh, minutes later".
### Phase 2 — Order and one writer (5–7 days)
- M6: the controller's lease and epoch; a report's sequence; the host's single apply queue; stale
refusals counted.
- The writers table and the signals table written into a to-be design, with their compile-time checks.
- P4's lint for empty-on-error readers in the core repositories, and the destructive-delta brake in the
provisioner harness.
- **Done when:** the lab replays of 204, 257/261/267 and 241 end with "refused stale", "one report" and
"withdrawal braked" respectively; the contract test for every consumed subject exists.
- **Risk removed:** class (a), the largest by count, and the destructive half of (b).
### Phase 3 — Healers (3–5 days)
- M5's first healers and the rule that a repair done by hand twice asks for one (from the hand-act log).
- **Done when:** a lab mesh recovers from each induced failure in the healers table with no hand, says so,
and brakes after its budget. Live: a week in which the hand-act log has no repeat.
### Phase 4 — Core upgrades that roll back (8–12 days)
- M7: health definitions; the gate; rollback for the controller and the runtime; the bus as a planned
step.
- **Done when:** on a lab mesh, a deliberately broken build of the controller, the host and the runtime
is each rolled back without a hand, the mesh ends on the previous build, and the rollback is a
condition and a notification. Live: the next three core rollouts each record a health verdict.
- **Risk removed:** class (g) — the failures that take the control path itself down.
### Phase 5 — Checks before merge, and the replay suite (8–12 days, then ongoing)
- M8: the facts snapshot and the compose-and-validate merge gate; the libc matrix; versions tested as
run.
- M9: every incident in the window as a scripted replay, run on every core merge; each new core issue
adds its replay as its "how it is checked".
- **Done when:** the replays of 236, 262, 263 and 266 fail on the commit before their fix and pass after;
a new core issue cannot resolve without a replay or a stated reason why none is possible.
**Total, roughly six to eight weeks of focused work**, with the first visible change — the mesh saying
when it is wrong — inside the first two.
## The five largest risks today
Ranked by likelihood × damage, from the window's evidence and the mesh as read the night this effort
began:
1. **A stall nobody is told about.** A plan, a merge or a report that stops is noticed only by someone
looking (230, 248, 257, 264, 266, 267; issue 187 still open). Every release goes through a plan.
2. **A destructive act on a misread input.** 241 dropped seven databases on one unreadable file; 234
undeclared four modules from a stale composition; 245's advice rebuilt everything and cut the bus;
253's collector would have deleted every archive. There is no general brake on a large withdrawal.
3. **The core replacing itself with nothing to roll it back.** A controller build that starts but cannot
plan (201, 214) halts every later change, because the controller is what plans the fix. Only the
host has a launcher rollback.
4. **The bus as a single, un-upgradeable process.** It runs a release that skips messages (266), its
reload drops owed replies (265), replacing it cuts every machine off (245), and its next upgrade is
one-way and needs a restart.
5. **Concurrent actors with no lease or report order.** Several agent sessions, plans and reconciles act
at once; on the night this began, four named pushes, one per machine, started within two seconds. The digests
catch most of it now, one rule per message kind; the next message kind will not have one.
@@ -1,85 +0,0 @@
---
status: graduated
initiated: 2026-10-06
became: [02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md, 03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md]
touches:
- 03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md
- 03-DESIGN/01-to-be/18-building-a-module.md
- the module manifest
- the node-engine
- the release gate
---
# 032 — A module says how it is healthy
## What is investigated
Whether, and how, every module should declare **how its own health is checked**, as a field of its
manifest, so the mesh can judge a module the way it already judges its core parts.
Today the release gate ([ADR 0236](../../02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)) judges a
module on its first machine by what the mesh can see from outside: the declaration applied, no new
condition about the module or its machine since the send, and its tools answering. It does not see
whether the module's own service works. A container that applied cleanly and then restarts in a loop,
a web application whose port is open but whose pages fail, a database that accepts connections and
refuses queries — each passes the gate and is caught only through what it breaks later, if anything
notices at all. To-be 45 names this gap ("no container state in the gate").
The questions:
1. **What a health declaration says.** The kinds of check a module may declare (a container's own
health state, an HTTP request and the answer expected, a TCP connect, a command run inside the
service, a query, a tool of the module's own that answers "healthy"), its interval, its timeout,
how many failures in a row count, and a start period during which failure does not count.
2. **Who runs the checks, and where the result goes.** The node-engine on the machine, the node tools,
or the module itself; how the result reaches the controller (the report, an event, a seat verb); and
what the release gate, the self-check and the healers do with it.
3. **What the mesh already has to build on.** Container runtimes' own healthchecks, which many images
ship and which the mesh neither reads nor sets today; systemd's own state for services; the tools a
module already serves.
4. **What "healthy" covers.** Liveness (it runs), readiness (it serves), and whether a module's health
may depend on what it requires (a provider down makes its consumers unhealthy — said once, at the
provider, not once per consumer).
5. **The cost and the noise.** How often checks may run across all modules on a small machine, and how
a check avoids the single-sample flaw the self-check already met (issue 277).
6. **Migration.** How every catalogue module gets a declaration, what a module without one is judged
by, and whether `module check` should require one.
## Why
The mesh now rolls a module out on its own, one machine first, and puts the previous build back when
the first machine is not healthy. That promise is only as good as "healthy" is, and for a module it is
today judged from the outside. Making each module say how its health is checked turns "applied" into
"working", for the gate, for the self-check and for whoever asks.
## What it touches
The module manifest and `module check`; the node-engine, which would run or read the checks; the
release gate and the doctor probes of to-be 45; the healers, which may restart what stays unhealthy;
the operator's conversation (ADR 0234), which carries what stays unhealthy; and every catalogue module,
each of which would gain a declaration.
## Where it stands
Evidence gathered on the live mesh, options weighed and a decision drafted, 2026-10-07:
- [01 — The evidence](01-evidence.md): 125 catalogue modules, 68 of them running something long-lived
(49 a container, 19 a service); no manifest can declare a check and nothing reads one; 19 of 73
long-running catalogue containers carry an image check, two of which were wrong in the mesh's hands;
every restart count is 0 because a recreate loses it, and the runtime's event history is about a minute
long. Of seven recent incidents, liveness alone would have caught the crash loop, readiness the
eleven-hour silent web app, and only a module's own tool the refused identity-provider admin.
- [02 — Options](02-options.md): who runs the checks, where the result goes, liveness and readiness,
dependency-aware health, the field's shape and the migration.
- [03 — Recommendation](03-recommendation.md): liveness judged for every long-running resource at once;
readiness declared per resource in a field named `health` and run by the node-engine; the state in the
report, the condition raised by the controller on the second look, so the gate needs no new rule; a
provider down said once at the provider; a proposed decision text with how each rule is checked, and
the migration of the 68 modules.
**Graduated 2026-10-07** on the operator's word (*"yes, turn it into a decision"*), as
[ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md) and
[to-be 48](../../03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md). The record takes the proposed
decision, adding: a fifth state, `held`, for a resource under a maintenance step; a ceiling of five minutes on
grace plus failing looks, so a broken start is said inside the gate's bound; and the gate's *own health* reading
the stated health, so a resource still starting is not yet a pass.
@@ -1,196 +0,0 @@
# 01 — The evidence
Measured on 2026-10-07 on the live mesh of four machines (a home server, a control node, a workstation
and a laptop), read only: the three catalogue repositories at their trunk, the controller's and the
node-engine's source at their trunk, each machine's container runtime and service manager, and the
controller's own answers. Counts, not anecdotes. Machines are named by role.
## 1. What the catalogue runs
The three catalogue repositories hold **125 module manifests**: 113 in the main catalogue, 12 in the
media catalogue. The photos repository holds application code and no manifest; its two instances are
modules in the main catalogue.
Each manifest sorted by the **longest-lived thing it runs** (a container that stays up first, then a
service the manifest says is running, then the mesh's own process that stays up, then a bundle only):
| What a module runs | Modules |
|---|---|
| at least one container that stays up | **49** |
| no such container, but a service unit stated `running` | **19** |
| only its bundle — tools and handlers hosted by the node tools | **50** |
| only files, directories and packages | **7** |
Underneath:
| Resource | Count | Of which |
|---|---|---|
| container | **85** in 50 modules | 77 stay up, 6 run once (a step), 2 on a schedule; **71 distinct images** |
| service (an existing unit put in a state) | **31** in 21 modules | 26 stated `running`, 5 stateless (the machine's lifecycle) |
| process (the mesh's own code in a unit it writes) | **20** in 18 modules | 17 run once, 2 on a schedule, **1** stays up |
| bundle artifact | **101** in 99 modules | |
And what a check could be built from:
- **50** modules declare `listens` (an endpoint, a port, a protocol) — **48 of the 49** container
modules. A TCP or HTTP check has its target named already.
- **53** modules declare `tools`. **One** tool in the whole catalogue is a health tool by name
(`dbus_health`); **18** modules have a tool named `…_status`.
- **67** modules `require` something; the most-required provisions are `route` (36 consumers),
`x11-display` (15) and `postgres-database` (12). A database down is, today, potentially twelve
consumers failing at once.
- **No manifest declares a health check.** The container resource has no field for one — its fields are
image, environment, files, ports, volumes, arguments, names, networks, capabilities, logging, the
restart triggers and the run-once and schedule modes — and the node-engine passes the runtime no
health option. Whatever check runs is the image's own, run by the runtime by default.
## 2. What the images already ship, and what the runtime says today
Every container on every machine inspected — its state, its health, its restart count, and its
**image's** healthcheck from the image configuration:
| | |
|---|---|
| containers on the four machines | **115** (94 running) |
| of those, declared by the catalogue and running | **79 instances** of **73** of the 77 long-running containers (4 are assigned nowhere) |
| long-running catalogue containers whose **image ships a HEALTHCHECK** | **19 of 73 (26 %)** |
| …in how many modules | **7 of 45** modules with a running container (mail, the hosted database suite, a spreadsheet app, the chat client, a flow editor, the media server, the certificate authority) |
| …modules whose every container has one | **4** |
| running container modules with **no health state at all** | **38 of 45** |
| catalogue containers reporting `healthy` / `unhealthy` / `starting` | **19 / 0 / 0** |
| containers reporting `unhealthy` | 5, all exited, all on the workstation, all predating the mesh (none declared) |
**Two of the nineteen image checks were wrong under the mesh's own configuration** until a catalogue fix:
- the hosted database suite's studio — the framework binds the address the runtime puts in `HOSTNAME`, so
it answered only on its network address while its image's check asked `localhost`: *unhealthy for ever
while working* (fixed 2026-10-05);
- the flow editor — the image's check reads its settings from a path the module mounted elsewhere: *ran
fine but reported unhealthy forever* (fixed 2026-09-30).
So 2 of 19 shipped checks (≈ 10 %) gave a false *unhealthy* in the mesh's hands. Read blindly, they would
have failed two good builds at the gate.
The intervals the images chose range from **2 s to 60 s**: a database every 2 s, three web apps every
5 s, an API gateway every 10 s, the rest 30 s or 60 s. Start periods range from none to 360 s (the
antivirus). Retries from 3 to 20.
### The restart count the runtime keeps is lost
**All 94 running containers report a restart count of 0** — including the agent server that, per
[issue 268](../../04-ISSUES/268-letta-printed-its-passwords-into-its-log/00-report.md), had restarted about
a hundred times in a crash loop days before. The counter belongs to a container, and the node-engine
recreates a container whenever its declaration or a file it reads changes; the fix recreated it and the
history went with it. A restart count is only evidence if something outside the container keeps it.
### The runtime's event history is about a minute long
The runtime keeps its most recent ~250 events in memory. On the home server, a query for the last 30
minutes and for the last minute both answered ~250 events, **every one an `exec_*` event** — the image
health checks running (84 creates, 85 dies, 84 starts). A 24-hour query for container lifecycle events,
through the mesh's own `docker_events` tool, answered **zero**. The images' checks crowd every lifecycle
event out of the history within a minute. A container that died an hour ago leaves no trace there.
## 3. What the service manager says today
- **Failed system units:** 0 on three machines; 2 on the workstation, both mounts that the mesh does not
declare. One failed user unit on each of two machines, neither the mesh's.
- The mesh's own units (`mesh-ca-trust`, `mesh-filter`, the power units, the controller, the node tools,
the node-engine): all `active`, **`NRestarts` 0**.
- systemd already gives, per unit and for free: `ActiveState`/`SubState`, `is-failed`, `NRestarts`,
the time it entered its state, and — through a unit's own `ExecStartPost`/`WatchdogSec` where software
supports it — readiness. None of it is read by the mesh for a module's unit.
## 4. What the node-engine reports today
The node-engine's report (its `Report` message) carries: what applied and what failed per resource,
refusals, the declaration's digest and order, held and stray containers, reachable sockets on an adopted
machine, packet filters and the firewall found, windows (a scheduled step holding containers still), the
machine's profile, the engine's own build, outward links, rollbacks its witnesses decided, and whether it
reads `witness`. Separately, an `Alive` heartbeat with its interval.
**It carries no container state, no unit state and no health.** The controller's `node` answer for the
home server shows the consequence: its assigned modules, its strays, its filters and its capabilities —
and nothing about whether any of its 45 catalogue containers is running.
The node-engine has one judging mechanism already: the **witness** (to-be 45 §8) — a process's `witness`
field, `lease` or `ping` or `none`, judges a new build of the controller or the node tools and restores
the previous one when it is not healthy in bound. Only the two core processes use it.
A **run-once step** is already a health gate in practice: a step exiting non-zero fails the apply and
everything placed after it. **23** steps exist (6 containers, 17 processes); one of them, in the hosted
database suite, is a loop that waits up to five minutes for its analytics service's `/health` before the
services that need it start.
## 5. What the release gate judges a module by
The gate on a plan's first machine ([ADR 0236](../../02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md);
`judgeHealth` in the controller) passes a module there when, on three judgings at least 40 s apart over at
least two minutes, within ten minutes of the send:
1. no witness on the machine put a core build back since the send;
2. the machine's last report is current and **applied** (failed or refused is broken);
3. **no condition** raised since the send names the module on that machine, or the machine as a whole
(issue 281 sorted the latter out);
4. for the controller, the node-engine and the node tools: their own health definitions (lease held and
ready; the engine's build reported; the node tools answering the bus);
5. for any other module **that declares tools**: the machine's node tools serve them.
That is the whole of it for a catalogue module. For the 49 container modules nothing in 1–5 looks at a
container: point 5 asks the node tools, which host the module's bundle, not its container. **A container
that applied and then crash-loops passes all five.** To-be 45's Phase 4 note says so: *"container state in
the node-engine's report, without which a container that crash-loops after its compose applied is seen
only through what it breaks"*.
None of the self-check's probes (D1–D13, the data and delivery probes, H-controller, H-engine, H-tools,
H-bus, DG) reads a module's container or unit state either.
## 6. The incidents, and what a declaration would have done
Every incident of the last ten days where a module applied and did not work, or looked as if it did not:
| Incident | What the gate and the self-check saw | Caught by, after | Would a declaration have caught it? |
|---|---|---|---|
| The agent server crash-looped: its database lacked the vector extension the provider never created (catalogue fix 2026-10-05) | applied; tools served (its bundle runs in the node tools, not the container); no condition | a person reading its log for another reason ([issue 268](../../04-ISSUES/268-letta-printed-its-passwords-into-its-log/00-report.md)), **about a hundred restarts**; the change readying it for the home server merged 2026-09-30, the provider's fix 2026-10-05 | **Yes, by liveness alone** — *running and not restarting* — inside the gate's ten minutes; an HTTP check on its declared `web` endpoint the same |
| A web app accepted TCP and answered no HTTP request, its database unreachable after a firewall change ([issue 145](../../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)) | *"all doing what they were told"* | a person, **after eleven hours** | **Yes, by an HTTP readiness check** within two looks (a minute at 30 s). **Not** by a TCP check — the port was open — and not by liveness |
| The identity provider's admin refused the secret the mesh minted; every consumer's client failed ([issue 179](../../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)) | the server up and serving; its image ships no check | a person, the module having been broken since it moved to the mesh; **twice** (2026-10-01, again 2026-10-05) | **Only by a module's own check** — *the admin logs in* — which the module now runs itself and announces (ADR 0224). No container check sees it |
| The studio and the flow editor read unhealthy while working (§2) | nothing: the mesh does not read health | a person reading `docker ps` | The opposite case: **a check read without proving it would have rolled back two good builds.** A declaration must be owned by the module and proved before it is trusted |
| Two media managers refused a new recycle-bin folder; their settings step exited non-zero ([issue 279](../../04-ISSUES/279-a-folder-cannot-be-owned-by-the-account-a-module-runs-as/00-report.md)) | the apply failed → the gate failed, at once | the gate | Already caught: a step is a gate. A health declaration adds nothing here |
| The media server's event handler threw on an empty answer and was offered each event five times ([issue 276](../../04-ISSUES/276-a-handler-that-did-its-work-was-offered-it-five-times/00-report.md)) | — | the bus watchdog (S9), `max-deliveries` | **No, and it should not**: the server was healthy; handling an event is the event contract's, not health's |
| A provisioner runtime restarted several times at start until the overlay was up, then worked ([issue 058](../../04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md)) | — | the lab | A warning: **a restart-counting check without a start period reads churn that stops as a crash loop** |
Of the seven: **two** a declaration catches that nothing catches today (the crash loop, by liveness; the
silent web app, by readiness), **one** only a module's own tool can catch (the admin), **one** shows what
an unproved check does wrong, and **three** are not health at all or are already caught.
## 7. Cost and noise
- **Reading every container's state on a machine**, health and restart count included: one `ps` of
38–49 containers took **17–21 ms**; one inspect of all of them **30–37 ms** (control node and home
server, five and three runs).
- **What the images' own checks cost today**, from the runtime's record of each check's start and end
(95 checks over 19 containers): **median 33 ms, slowest 138 ms**. On the home server they run
**77 checks a minute** — 8 containers, three of them every 2–5 s — for about **5 s of exec time a
minute**. On the control node, 11 containers at 30 s: 22 checks a minute, under 1 s.
- **The smallest machines in this mesh** have 12 cores (the control node) and 31 GB (the laptop); the busiest runs 45 catalogue containers.
One check per long-running resource at 30 s is **90 checks a minute** on the busiest machine — about
3 s of exec time a minute for in-container checks, a few hundred milliseconds for HTTP or TCP checks
the engine makes itself. The cost is not CPU; it is that every in-container check is an `exec` in the
runtime's event history, which on the home server is already one minute long.
- **Single-sample noise.** [Issue 277](../../04-ISSUES/277-one-unanswered-question-was-an-urgent-alert-nobody-could-read/00-report.md):
one DNS question unanswered on a machine starting twenty containers was raised urgent; thirty asked by
hand a moment later were all answered. Its rule — *a finding one look can be wrong about is raised on
the second look in a row* — is the runtime's own `retries` under another name; the images use 3 to 20.
The gate's own three passes 40 s apart are a third form of the same idea.
## 8. What this says
1. **Most of the catalogue has no health anywhere.** 38 of 45 running container modules, all 19
service-only modules and every bundle module have nothing that says *working* beyond *applied*.
2. **What exists is not read and not owned.** 19 image checks run, 2 were wrong in the mesh's hands,
and neither the report nor the gate nor the self-check reads any of them.
3. **The runtime's memory is too short to rely on.** Restart counts vanish with each recreate; the
event history is a minute long. Liveness must be observed and kept by the node-engine.
4. **Liveness alone would have caught the worst incident; readiness the longest; only a module's own
tool the identity provider's.** All three kinds are needed, and none is enough alone.
5. **The cost is small and the noise is known.** Tens of milliseconds a look; the two-look rule exists.
@@ -1,141 +0,0 @@
# 02 — Options
Each question of the [overview](00-overview.md), the options for it, and what the
[evidence](01-evidence.md) says about each. The recommendation is [03](03-recommendation.md).
## A. Who runs the checks
### A1 — The container runtime's own HEALTHCHECK, read by the node-engine
The image's check, or one the module sets, run by the runtime inside the container; the node-engine reads
the `health` state it keeps.
- **For:** 19 images already ship one; it runs inside the container's namespace, so a check of
`localhost` sees what the program sees; the runtime handles interval, timeout, retries and start period.
- **Against:** containers only — nothing for the 19 service-only modules, the one long-running process,
or anything a module's own tool knows. Every check is an `exec`, and on the home server the 77 a minute
already push every lifecycle event out of the runtime's history (01 §2). Two of 19 shipped checks were
wrong in the mesh's configuration; an image check the module never stated is a check nobody owns. The
runtime does nothing when a container turns unhealthy (it restarts only on exit), so the state is only
useful to whoever reads it.
### A2 — The node-engine runs every check itself
HTTP and TCP from the machine, `exec` through the runtime, a unit's state through the service manager, a
tool through the node tools.
- **For:** one runner and one reader per machine, for every hosting form; it keeps what the runtime
forgets (restarts across recreates, 01 §2); a check from outside the container tests the path a caller
takes, which issue 145 says matters; HTTP and TCP cost no `exec`.
- **Against:** the engine grows a scheduler; a check that only makes sense inside the container (a CLI,
a pid file) still needs an `exec`, which A1 does better.
### A3 — The module's own tool answers "healthy"
The bundle exposes a health tool; something calls it.
- **For:** the only kind that catches a failure of function — the identity provider's refused admin
(issue 179) — and the module knows what "working" means for it.
- **Against:** a bundle is hosted by the node tools, not by the service; a tool that answers proves the
bundle is up, not the server. One tool in the catalogue is a health tool today. And a module that judges
itself is the thing ADR 0227 rule 8 refuses for the core: *the component being replaced is never the
judge*. For readiness of function it is the only option; for liveness it must not be the only one.
### A4 — A mix, by kind (the shape the evidence points at)
The engine owns every check and every verdict. It runs HTTP, TCP and unit checks itself; it delegates an
in-container command to the runtime by setting the container's HEALTHCHECK from the declaration (so the
runtime's retries and start period do the timing) and reads the state; it asks a module's health tool
through the node tools. Liveness — running and not restarting — it observes for every long-running
resource with no declaration at all.
## B. Where the result goes
| Option | For | Against |
|---|---|---|
| **B1 — A field of the report**: per module, per long-running resource, a *state* (healthy, unhealthy, starting, unknown), since when, the failing streak and the restarts the engine counted | the report is how a machine states facts about itself; the gate already reads the last report; one place to read | a report is sent after an apply, not on a change — a container that goes bad at 03:00 waits for the next report |
| **B2 — An event on each transition** (`module.<m>.<machine>.health` changed) | immediate; the controller can keep the state | events are lost or replayed; an event alone is a sample, not a state |
| **B3 — A condition, raised by the controller** | the gate already fails a module on a new condition naming it on that machine (01 §5, point 3); the operator's conversation (ADR 0234) and the healers (ADR 0231) already act on conditions; the two-look rule and the content rule already apply | a condition is the controller's word, raised from something — it needs B1 or B2 under it |
| **B4 — A seat verb the controller calls** (`health <module>`) | always current | a pull per module per judging; a machine that is slow to answer reads as unhealthy |
B1 + B2 + B3 together is how the core already works: the engine states its state in the report and
emits the change; the controller keeps the last state per machine, raises the condition on the second
look, and clears it on the look that no longer sees it. That makes **the gate need no new rule**: an
unhealthy module raises a condition naming it, and point 3 already holds the gate on it.
## C. Liveness and readiness
- **Liveness** — *it runs*: a container running and not restarted within a window; a unit `active` and
not `failed`, `NRestarts` not climbing; a process up. Observable for every long-running resource with
no declaration. It alone catches the worst incident of the window (the crash loop). It needs a **start
period** and a **settle rule**, or it reads issue 058's churn-that-stops as a crash loop.
- **Readiness** — *it serves*: the declared check passes. Only the module can say what serving means.
- **Function** — *it does its job*: the identity provider's admin logs in. A module's own tool, and only
for what no endpoint shows.
Options: judge liveness only (cheap, no declarations, misses issue 145 and 179); judge readiness only
(misses nothing that readiness sees, but every module must declare before anything is judged); **judge
liveness everywhere at once and readiness where declared**, making the declaration required over a
migration. The last keeps the gate meaningful from the first day.
What the mesh **does** with each is a separate choice. Restarting a container that is unhealthy is what
an orchestrator's liveness probe does; the runtime here does not, and a restart hides the failure the
gate is meant to see. Options: the healers restart what stays unhealthy (ADR 0231's shape: act on what
observation raised, say whether it worked), or nothing restarts on health and the condition reaches a
person. The evidence has no case where a restart would have fixed anything — the crash loop was
restarting already.
## D. Dependency-aware health
A module requiring a provision fails when its provider fails. With 12 consumers of the database
provision and 36 of a route, a provider down would be a dozen conditions said once each, and a dozen
gates failed for something none of them did.
| Option | What it costs |
|---|---|
| D1 — Ignore it | twelve conditions for one fault; the operator learns which one matters by reading all of them; the gate blames the consumers |
| D2 — A consumer's check names the provision it exercises; when that provider is unhealthy on the record, the consumer's finding is **held under the provider's** — said as *waiting on* the provider, not raised on its own — and its gate waits rather than fails | one condition, at the provider; needs the controller to know which provider answers which consumer — it does: it composes the grants ([issue 181](../../04-ISSUES/181-an-assignment-does-not-record-which-provider-answers-it/00-report.md) is about recording it) |
| D3 — Order the judging: providers first, consumers only when their providers are healthy | simple, but a consumer broken on its own is not said while a provider is down, which is the moment it is most needed |
D2 is how issue 281 already treats a machine-level condition: what is about the machine is the
machine's, and is never pinned on the module the gate is kept on.
## E. The manifest field
Research may sketch a shape; a design doc may only name the field. Where the declaration lives:
- **E1 — On the resource** (`health` on a container, a process, a service): the check belongs to the
thing it checks; a module with three containers states three. Matches how `restart-on` and `witness`
already sit on the resource.
- **E2 — One per module** (a top-level `health`): one verdict, but a module of eleven containers (the
mail module) cannot say which one is wrong.
The sketch for E1 — a field named `health`, holding:
| Part | Meaning | Default |
|---|---|---|
| kind | `runtime` (the image's own check, adopted as is), `http` (an endpoint from `listens`, a path, the status expected), `tcp` (an endpoint from `listens`), `exec` (a command in the container), `unit` (the unit's own readiness), `tool` (one of the module's tools answering) | — |
| every | interval | 30 s; not under 10 s |
| timeout | how long one look may take | 5 s; under `every` |
| after | looks failing in a row before it is unhealthy | 3; **not under 2** (issue 277) |
| grace | after a start, how long failure does not count | 60 s |
| needs | the provision whose provider it exercises, for D2 | none |
An HTTP or TCP check names an endpoint by its `listens` name, never a port or an address, so a check
follows the machine's ports the way the endpoint does. A `runtime` kind is an explicit adoption of the
image's check, so a module that ships one *says* it does, and `module check` can prove it on a lab.
## F. Migration
- **F1 — Required at once** — every long-running resource declares, or `module check` refuses. 68
modules to touch before the next merge; nothing is judged until all are done.
- **F2 — Liveness at once, declaration required by a date** — every long-running resource is judged by
liveness from the first build; `module check` warns, then refuses after a stated date; a catalogue-wide
test counts the modules still without one, and the count only goes down.
- **F3 — Optional for ever** — modules without one are judged by liveness and tools served. The 38 of 45
container modules with nothing today would stay at liveness.
What a module without a declaration is judged by in F2 and F3: liveness of every long-running resource
plus today's five points (01 §5). A bundle-only module (50) and a files-only module (7) need no
declaration: they run nothing long-lived of their own, and the node tools serving their tools is their
liveness.
@@ -1,120 +0,0 @@
# 03 — Recommendation
From the [evidence](01-evidence.md) and the [options](02-options.md): the node-engine judges every
long-running thing a module runs — **liveness for all of them, at once, with no declaration**, and
**readiness where the module declares how** — states it in its report, and the controller raises it as a
condition on the second look, so the gate, the self-check and the operator's conversation act on it with
no new rule of their own. Every catalogue module that runs something long-lived declares its check
within a stated migration, and `module check` then requires it.
## Proposed decision text
Ready for graduation through playbook 02. The number is assigned then.
> **A module says how it is healthy, and the node-engine judges it**
>
> **Context.** The release gate (ADR 0236) judges a catalogue module on its first machine by what the
> mesh sees from outside: the declaration applied, no new condition about it or its machine, its tools
> served. None of that looks at what the module runs. Of 125 catalogue modules, 49 run a container that
> stays up and 19 a service unit; no manifest can declare a health check, the node-engine reports no
> container or unit state, and no probe reads one. 19 of 73 long-running catalogue containers have an
> image check the mesh never reads, two of which were wrong in the mesh's configuration. In ten days a
> container crash-looped about a hundred times and a web app answered nothing for eleven hours, both
> while every check passed.
>
> **Decision.**
>
> 1. **Liveness is judged for every long-running resource, with no declaration.** A container that stays
> up, a process that stays up and a service stated `running` are *alive* when running and not
> restarted more than once within the settle window after their grace period. The node-engine
> observes this itself on each tick and keeps the restarts it counted across recreates; it never
> relies on the runtime's restart count or event history. A resource held still by an open window (ADR 0189) is
> neither alive nor dead: it is said as held, and judged again when the window closes.
> 2. **A module declares how each long-running resource is ready**, in a field named `health` on that
> resource: one kind — the image's own check adopted by name, an HTTP request to a declared endpoint
> and the status expected, a TCP connect to a declared endpoint, a command in the container, the
> unit's own readiness, or one of the module's tools — with an interval (default 30 s, not under
> 10 s), a timeout under the interval, a number of failing looks in a row (default 3, **not under 2**),
> and a grace period after a start. An endpoint is named by its `listens` name, never by a port or an
> address. A check of function that no endpoint shows is a module's own tool, and only in addition to
> a check the module does not run itself.
> 3. **The node-engine runs every check and owns every verdict.** HTTP, TCP and unit checks it makes
> itself; a command in the container it hands to the runtime as that container's healthcheck and reads
> the state; a tool it asks through the node tools. Nothing else on the machine judges a module.
> 4. **The state goes in the report, its change on the bus, and a condition is the controller's.** Each
> report carries, per module and long-running resource, a state — healthy, unhealthy, starting,
> unknown — since when, the failing streak and the counted restarts; each transition is emitted as an
> event. The controller keeps the last state per machine and raises `module.<module>.<machine>`
> unhealthy as a condition when two consecutive states say so, and clears it on the first that does
> not. The release gate, unchanged, holds a module on a condition naming it; at its bound the build is
> put back.
> 5. **A provider down is said once, at the provider.** A check names the provision it exercises. While
> that provision's provider is unhealthy on the record, the consumer's finding is held under the
> provider's condition — said as waiting on it — and the consumer's gate waits rather than fails.
> What a consumer finds while its provider is healthy is its own.
> 6. **Nothing is restarted for being unhealthy.** The runtime restarts what exits, as now. A condition
> reaches a person or a healer (ADR 0231); a healer that restarts on health is its own decision.
> 7. **A declaration is proved before it is trusted.** `module check` refuses a `health` field that names
> an endpoint the module does not declare, an interval or count below the floor, or a tool the module
> does not serve. A lab bed applies every changed declaration and requires it healthy within its grace;
> an image check adopted by name is proved the same way, because two of nineteen were wrong.
> 8. **Every catalogue module that runs something long-lived declares one.** `module check` warns from
> the decision and refuses a long-running resource without `health` after the migration's date. A
> module running nothing long-lived — its bundle only, or files and packages — declares none: the node
> tools serving its tools is its liveness, as the gate judges today.
>
> **Consequences.** Liveness alone, from the first build, would have caught the crash loop inside the
> gate's ten minutes. Readiness would have caught the silent web app in a minute instead of eleven hours.
> The identity provider's refused admin needs the module's own tool, which it now has. The node-engine
> grows a small scheduler and a report field; the controller a condition kind; the gate nothing. A
> machine of 45 containers spends tens of milliseconds a look reading state, and about 3 s a minute on
> in-container commands at the default interval.
## How each rule is checked
| Rule | Checked by |
|---|---|
| 1 Liveness without declaration | node-engine unit tests over a fake runtime and service manager: a container recreated keeps its counted restarts; a restart inside grace is not counted; two restarts within the settle window after grace make it unhealthy. A lab bed that replays the crash loop (a container whose program exits at start) fails the gate within its bound |
| 2 The field and its floors | `module check` refuses each out-of-range part, with a test per refusal; a catalogue-wide test parses every `health` field |
| 3 The engine owns the verdict | a test that a declared command becomes the container's healthcheck and nothing else sets one; a test that an HTTP check dials the endpoint's current port after a port change |
| 4 Report, event, condition | controller tests: one unhealthy state raises nothing and is listed unconfirmed; two raise; a healthy state clears; the gate holds a module on that condition (an existing gate test, extended with this condition's kind). A lab replay of issue 145 — a database made unreachable — raises the consumer's condition within two looks |
| 5 Said once at the provider | a controller test with one unhealthy database provider and three consumers failing: one condition, at the provider, the consumers listed as waiting; the consumers' gates wait, not fail |
| 6 No restart on health | a node-engine test that an unhealthy container is not restarted, recreated or stopped |
| 7 Proved before trusted | the lab bed's run of every changed declaration on every catalogue merge; the replay of the studio's false *unhealthy* (a check that asks `localhost` while the program binds elsewhere) fails the bed, not a machine |
| 8 Every long-running module declares | a catalogue-wide test that counts the modules with a long-running resource and no `health`: it may only go down, and is zero by the date; after it, `module check` refuses |
## The migration
**Today:** 68 modules run something long-lived (49 with a container, 19 with a running service only);
57 run nothing long-lived and need no declaration.
1. **Build first, without declarations.** The node-engine's liveness, the report field, the events, the
controller's condition and its two-look rule, and the dependency hold. From this step every
long-running resource is judged by liveness. The catalogue-wide counter starts at 68.
2. **The seven modules whose images ship checks** adopt them by name — and, for the two that were
wrong, the fixed configuration is what the lab proves. 19 of their containers are covered at once.
3. **The other 42 container modules, and the containers without an image check in three of the seven,** declare an HTTP check on their declared endpoint where they serve
HTTP, TCP where they serve something else, a command where neither shows readiness. 48 of 49
already declare the endpoint the check needs.
4. **The 19 service-only modules** declare the unit's own readiness, or a TCP check where the unit
listens; most are machine software (a resolver, a time daemon, a session manager), where `active` and
not `failed` is already the honest answer.
5. **Modules whose function no endpoint shows** add a tool check: the identity provider (its admin logs
in), the database provider (it can create in a consumer's database), the broker. Identified as they
are met, not in advance.
6. **The date:** when the counter reaches zero, or six weeks after step 1, whichever is first. From then
`module check` refuses a long-running resource without `health`.
**What a module without a declaration is judged by, until then and for ever if it runs nothing
long-lived:** liveness of everything it runs that stays up, and the gate's five points as they stand —
applied, no witness put it back, no new condition naming it or its machine, the core's own definitions,
its tools served.
## What this does not decide
- Whether a healer restarts what stays unhealthy (rule 6 leaves it to its own record, under ADR 0231).
- Health for scheduled work — whether the last scheduled run succeeded is data the windows of ADR 0189
already carry, and belongs to that record.
- Health of what a module's events do (issue 276) — the event contract's, and the bus watchdog's.
- The core's own definitions (to-be 45 §8), which stand; a core component may later declare its own
through the same field.
@@ -38,11 +38,6 @@ hand.
### What a domain module turns out to be, and why it is not the one refused above ### What a domain module turns out to be, and why it is not the one refused above
> **Narrowed, not replaced — 2026-10-06, by [ADR 0226](0226-the-private-network-is-assigned-by-its-own-name-and-the-proxy-names-its-public-issuer.md).** The catalogue no longer
> holds a domain module: `networking` was retired, and every machine is assigned the private network's
> own module. A module with requirements and no files stays a thing a module may be; "why it is worth
> having" below describes what this record decided for `networking`, not what runs.
*Written 2026-08-29, from building it. The heading above reads as a contradiction of what now *Written 2026-08-29, from building it. The heading above reads as a contradiction of what now
exists and is not one — but only if the difference is stated, so it is stated here.* exists and is not one — but only if the difference is stated, so it is stated here.*
@@ -9,10 +9,6 @@ extends: 0027-a-provision-names-what-the-consumer-is-coupled-to.md
# 44. A public name is provisioned, not registered by hand # 44. A public name is provisioned, not registered by hand
> **The mechanism changed — 2026-10-06, by [ADR 0226](0226-the-private-network-is-assigned-by-its-own-name-and-the-proxy-names-its-public-issuer.md).** The interface stands; its one provider,
> `cloudflare-dns`, left the catalogue, assigned nowhere and required by nothing. A mesh that needs a
> public name made for it adds a provider of `public-dns` again.
## Context ## Context
The mesh names and resolves its own machines internally: the overlay generates The mesh names and resolves its own machines internally: the overlay generates
@@ -108,14 +108,6 @@ declared slug is a strictly better escape hatch than an opaque hash. **D stays r
## Consequences (of E) ## Consequences (of E)
> **The mechanism changed — 2026-10-06, by [ADR 0225](0225-a-consumers-identity-is-bounded-by-the-provision-it-requires.md).**
> Option C, named above as the later refinement, is taken: each offer states the longest identity its
> backend keeps, and a consumer is bounded by the provision it requires rather than by 20 everywhere.
> 20 stays the bound of the object store and of a provider that is told its consumers and does not
> say. The overflow is refused before merge by the catalogue check, and at composition the consumer
> is left out of its provider's grants and reported — never the provider's machine refused. What
> stands: the identity is said once, the slug is the remedy, nothing is hashed or truncated.
- A module manifest gains an optional `slug`; a node may carry one too. `ConsumerIdentity` prefers - A module manifest gains an optional `slug`; a node may carry one too. `ConsumerIdentity` prefers
the slug over the cleaned name for each half. `identityLimit` becomes 20 (the true minimum), and the slug over the cleaned name for each half. `identityLimit` becomes 20 (the true minimum), and
`CheckIdentity` refuses at `module add` / assignment — now with a message naming the slug to set. `CheckIdentity` refuses at `module add` / assignment — now with a message naming the slug to set.
@@ -51,14 +51,6 @@ the builder) — cost with no new property.
restarts the runtime when that file changes — the `/etc/hosts` pattern for the content, the restarts the runtime when that file changes — the `/etc/hosts` pattern for the content, the
nftables pattern for the reload. No module author is involved; being on the network is what nftables pattern for the reload. No module author is involved; being on the network is what
grants the trust, because being on the network is what the trust *is*. grants the trust, because being on the network is what the trust *is*.
> **The mechanism changed — 2026-10-05, by [ADR 0222](0222-a-module-is-told-where-a-mesh-seats-holder-is-reached-and-the-controller-writes-no-file-a-seats-holder-owns.md).**
> What stands: being on the network grants the trust, the overlay is the transport security, no
> module author chooses it, and the trust is written into the runtime's file rather than over it
> and reloaded rather than restarted (ADR 0102). What moved: the controller no longer injects the
> file or the service. The container runtime's own module writes `insecure-registries`, told where
> this machine reaches the store by `${seat:mesh-artifact-store:reach}`, because the controller
> writes no file a seat's holder owns ([issue 190](../04-ISSUES/190-the-runtimes-configuration-is-written-by-modules-that-are-not-the-runtime/00-report.md)).
3. **No accounts (issue 042), recorded as the position it always was.** Reading and pushing 3. **No accounts (issue 042), recorded as the position it always was.** Reading and pushing
require presence on the overlay and nothing else. The boundary is enforced, not assumed: the require presence on the overlay and nothing else. The boundary is enforced, not assumed: the
registry's `listens` is `from: mesh`, the firewall derives from it, and the overlay admits only registry's `listens` is `from: mesh`, the firewall derives from it, and the overlay admits only
@@ -45,13 +45,6 @@ an operator obligation, and an obligation enforced by nothing is issue 057 resta
declaration is computed from the whole mesh; delivering a mesh that is knowingly inconsistent and declaration is computed from the whole mesh; delivering a mesh that is knowingly inconsistent and
merely saying so would make "push succeeded" mean less than it says. merely saying so would make "push succeeded" mean less than it says.
> **The mechanism changed — 2026-10-05, by [ADR 0221](0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md).**
> A named push still flushes every other machine that is behind, compared against what each was last
> sent. A machine whose modules would move to a build their upgrade policy records, or that an open plan
> has not sent it yet, is no longer flushed: the push names it and leaves it for `push <node>`. "Behind
> for an unrelated reason" no longer covers a held upgrade
> ([issue 259](../04-ISSUES/259-a-named-push-sent-every-machine/00-report.md)).
## Consequences ## Consequences
- One push is sufficient for a cross-node consumer: the provider's grants arrive from the same - One push is sufficient for a cross-node consumer: the provider's grants arrive from the same
@@ -64,15 +64,6 @@ node therefore trusts the mesh's registry as soon as it is on the private networ
networking no longer touches the runtime. The hosts file networking writes is still written whole networking no longer touches the runtime. The hosts file networking writes is still written whole
and stays held until networking is taken; a converge preview names it among the files it replaces. and stays held until networking is taken; a converge preview names it among the files it replaces.
> **The mechanism changed — 2026-10-05, by [ADR 0222](0222-a-module-is-told-where-a-mesh-seats-holder-is-reached-and-the-controller-writes-no-file-a-seats-holder-owns.md).**
> Writing into a shared file, adding to a list and reloading rather than restarting all stand. What
> moved is the writer: the networking module no longer writes the runtime's trust or declares its
> service. The container runtime's own module writes it into its own file and reloads its own
> service ([issue 190](../04-ISSUES/190-the-runtimes-configuration-is-written-by-modules-that-are-not-the-runtime/00-report.md)).
> The controller test named under *How it is checked* below, which held the networking module to
> declaring the runtime's file, is replaced by one holding the networking module to declaring neither,
> and by one holding the runtime's module to the trust.
## Consequences ## Consequences
- The runtime's file on a machine in use keeps its data directory, its logging settings and - The runtime's file on a machine in use keeps its data directory, its logging settings and
@@ -293,12 +293,3 @@ travel, which is the other half and was never in question.
> request/reply over the bus, never through the vault and never as a file the host writes. ADR 0183 > request/reply over the bus, never through the vault and never as a file the host writes. ADR 0183
> states that as a bounded exception — one vendor, tokens that live hours, one recipient per message — > states that as a bounded exception — one vendor, tokens that live hours, one recipient per message —
> and a second such channel is a decision of its own. > and a second such channel is a decision of its own.
> **The mechanism changed — 2026-10-06, by [ADR 0228](0228-a-value-given-by-hand-lives-only-until-its-modules-first-good-start.md).**
> What stands: a delivered value the vault cannot replace, such as an external API key, is not rotated
> by the vault, and rotating it means an operator delivering a new one. What moved: the controller had
> read that as covering **every** value given to it, and refused to rotate any of them. 0228 says what
> cannot be replaced is a value an outside party issues, which a module's definition now marks
> (`"issued-by": "outside"`); a given value for a secret the module reads at start is rotated like a
> made one, and one given through `secret accept` is replaced on its own after the module's first good
> start under the mesh.
@@ -9,10 +9,6 @@ extends: 0110-a-seat-is-a-module-assignment-from-a-closed-set.md
# 117. A machine's uplink is a seat: the mesh configures the manager, never the link # 117. A machine's uplink is a seat: the mesh configures the manager, never the link
> **The mechanism changed — 2026-10-06, by [ADR 0226](0226-the-private-network-is-assigned-by-its-own-name-and-the-proxy-names-its-public-issuer.md).** The seat stands; the catalogue holds two
> of its three modules. `dhcpcd`, held by no machine, left it. What this record says of dhcpcd is what
> a module for it must do, if one is written again.
## Context ## Context
The mesh installs on top of a machine's own networking. The private network's generator says The mesh installs on top of a machine's own networking. The private network's generator says
@@ -60,14 +60,6 @@ with the server held still for its duration. Plain collection, not `--delete-unt
mesh keeps is still a manifest in the store, so it is still referenced, so its blobs stay — the mesh keeps is still a manifest in the store, so it is still referenced, so its blobs stay — the
dangerous flag is not needed at all once the mesh is the one deciding. dangerous flag is not needed at all once the mesh is the one deciding.
> **Progressive insight — 2026-10-05.** "What the mesh keeps is still a manifest in the store" was true
> of images and false of archives: the builder published every archive as a bare blob no manifest names,
> and the store's collector keeps only what a manifest names. Its first night would have deleted every
> archive the mesh keeps ([issue 253](../04-ISSUES/253-the-stores-collector-would-delete-every-archive-the-mesh-keeps/00-report.md)).
> The decision stands — the mesh decides, the store reclaims with plain collection. What changes is how an
> archive is published: with a manifest that holds it, so the sentence becomes true of archives too. Until
> every kept archive is held, the collector runs as a dry run.
**3. What the mesh keeps, stated as three reasons rather than a number.** A digest is kept because: **3. What the mesh keeps, stated as three reasons rather than a number.** A digest is kept because:
- **a definition names it** — every artifact reference in any module's current recorded manifest, - **a definition names it** — every artifact reference in any module's current recorded manifest,
@@ -11,11 +11,6 @@ extends: 0191-the-meshs-resolver-holds-only-the-meshs-own-names.md
# 194. The mesh has one resolver, and every node asks it for the mesh's names # 194. The mesh has one resolver, and every node asks it for the mesh's names
> **The mechanism changed — 2026-10-05, by [ADR 0223](0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md).** `mesh-resolver` (named `mesh-dns-resolver` in the
> set) is no longer of capacity one: it is a replicated seat, held on the anchor and on the home
> server, each holding every node's internal domain from the same roster. One resolving module, the
> mesh's names held only by its holders, and the retirement of every per-node copy stand.
> **Narrowed, not replaced — 2026-10-03.** How a node asks is decided again by > **Narrowed, not replaced — 2026-10-03.** How a node asks is decided again by
> [ADR 0196](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md): every > [ADR 0196](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md): every
> node and container asks `mesh-resolver` first and a public resolver only when it is silent. There is > node and container asks `mesh-resolver` first and a public resolver only when it is silent. There is
@@ -10,13 +10,6 @@ supersedes-in-part:
# 196. A node asks the mesh's resolver first, and a public one only when it is silent # 196. A node asks the mesh's resolver first, and a public one only when it is silent
> **Narrowed, not replaced — 2026-10-05, by [ADR 0223](0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md).** The public resolver listed second goes: musl asks
> every listed server at once and takes the first reply, so a public "no such name" for a mesh name
> won it, and every Alpine build on the home server failed. Every machine now lists the mesh's
> resolvers — two holders of `mesh-dns-resolver`, its own first on a holder — and nothing else. Every
> node and container asking the mesh's resolver for every name, with no stub and no runtime `dns`,
> stands. The "fallback" consequence and check below describe what this record decided, not what runs.
## Context ## Context
**[ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) gave the **[ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) gave the
@@ -9,10 +9,6 @@ extends: 0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-nam
# 199. A module that answers names declares its zone, and a node's hosts file is one module's # 199. A module that answers names declares its zone, and a node's hosts file is one module's
> **Decided to change — 2026-10-05, by [ADR 0223](0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md), not yet built.** The `hosts` module is renamed `hostname`
> and its seat `node-hostname`, and it owns `/etc/hostname` as well as `/etc/hosts`; the operator's
> lines are kept as this record decides. Until that is built, everything below stands as decided.
## Context ## Context
**[ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) and **[ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) and
@@ -64,13 +64,6 @@ instead of `_`, which is the whole of the difference between the mesh's identifi
the one buckets, vhosts and hostnames use. A provider that needs a prefix or a suffix writes it the one buckets, vhosts and hostnames use. A provider that needs a prefix or a suffix writes it
around the placeholder, because a served value is a string. around the placeholder, because a served value is a string.
> **The mechanism changed — 2026-10-06, by [ADR 0225](0225-a-consumers-identity-is-bounded-by-the-provision-it-requires.md).**
> The identity is no longer capped at twenty characters for every provision: each offer states the
> bound its backend keeps, and twenty is the bound of the object store and of a provider that does not
> say. What stands: the identity is still the mesh's, and `dns` is still that name with its separator
> written `-`. An offer serving `${consumer:as:dns}` now bounds its consumers at 63 or less, so the
> label still fits.
The rejected alternative is **the provider returning values from provisioning** — the natural The rejected alternative is **the provider returning values from provisioning** — the natural
channel, since the provider is what derived them. It is rejected for three reasons, in order of channel, since the provider is what derived them. It is rejected for three reasons, in order of
weight. It inverts the delivery the mesh is built on: a grant would carry data the provider wrote weight. It inverts the delivery the mesh is built on: a grant would carry data the provider wrote
@@ -46,13 +46,6 @@ filesystem where there is one. Its key is a mesh secret.
**A night that fails reaches the operator,** and a machine with data and no good backup in 48 hours **A night that fails reaches the operator,** and a machine with data and no good backup in 48 hours
shows in the mesh's status. shows in the mesh's status.
> **The mechanism changed — 2026-10-06, by [ADR 0233](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md).** The decision stands: backups guard
> against mistakes and stay on the machine. What moved is the declaration: a module declares its data in
> its `data` section, with a class and how each item is protected, and the holder's lines are derived
> from it; a line written by hand is refused. The media library, left out here, is now declared
> irreplaceable and protected by the redundancy of its array, which the holder watches, because there
> is no room to copy it.
## Consequences ## Consequences
- Adding a store provider means declaring its dump; the catalogue check can refuse a store provider - Adding a store provider means declaring its dump; the catalogue check can refuse a store provider
@@ -1,107 +0,0 @@
---
topic: the mesh
status: accepted
date: 2026-10-05
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md
---
# 218. A plan sends grants before code, rolls a module out one machine first, and a newer merge takes over an older plan
## Context
On 2026-10-05 the delivery path was watched through a day of merges, by several sessions at once. Three
things went wrong, each recorded as an issue with its evidence.
- **Code arrived before the right to use it** ([issue 249](../04-ISSUES/249-a-modules-new-state-is-refused-until-a-push-the-merge-did-not-make/00-report.md)).
A merge gave a module a new key-value state. The plan sent the new bundle to every machine, and only
then issued the memberships that grant the state. On three machines the module's new state was refused
for two minutes, until a push made by hand. The order is written into the code on purpose: memberships
"after the declaration, because the runtime it is for arrives with it". That reason holds only for a
first assignment, and even then a membership is kept on the bus for the runtime that connects later
([ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)).
- **No machine went first.** The module's upgrade policy sends one machine at a time, but a plan's rollout
ignores it and sends every machine running the module at once. One at a time also never waited for the
first machine to come up healthy: it stopped only if the publish itself failed. A change was therefore
everywhere before anything had seen it run.
- **Plans for successive merges ran over each other** ([issue 254](../04-ISSUES/254-plans-for-successive-merges-run-over-each-other-and-one-was-left-open/00-report.md)).
Three merges to the catalogue within four minutes made three plans. Each sent the build agent to every
machine and asked for the same builds. One was still "building" hours later, with nothing left for it to
wait on. [ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md) decides one plan per merge
and says nothing about the next merge arriving while one is open. [Issue 219](../04-ISSUES/219-an-older-build-that-finishes-later-replaces-a-newer-one/00-report.md)
settled only which build's output wins.
## Considered Options
1. **Debounce merges:** wait a window before planning, so close merges make one plan. Rejected: it only
delays the overlap, does nothing for merges further apart than the window, and makes every merge slower.
2. **Queue plans:** a new plan waits until the older one is done. Rejected: the older plan builds what the
newer merge is about to replace, then the newer one builds it again.
3. **A newer merge's plan takes over the older plan's unfinished work, a plan rolls a module out one
machine first, and grants travel before code.** Chosen.
## Decision
**1. Grants before code.** Every send — a plan's rollout and a push alike — issues the memberships for
the machines it is about to send to before it sends their declarations, after raising the buckets they
name. When the composed list of bus users changes, the machine that holds the bus is sent first, because
that list travels in its declaration. A membership that could not be issued fails the send, and the send
is tried again. It is never reported as done "until the next push".
**2. One machine first.** A plan rolls a module out according to the module's upgrade policy
([ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md) §3). Unless the policy says
*together*:
- the module is sent to one machine first, the first by name of the machines running it;
- the rest are sent only once that machine has reported the new declaration applied and current;
- a first machine that reports a failure, or does not report in time, stops the module's rollout there.
The plan names the machine and the reason, and the other machines keep what they ran.
The plan records which machine went first, so a controller replaced mid-rollout resumes from there. A
policy of *together* keeps today's behaviour.
> **The mechanism changed — 2026-10-06, by [ADR 0236](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md).** What still stands: one machine
> first, the first by name, the rest only after it, a failed first machine stopping the module and the
> plan, *together* as it was. What moved: "reported applied and current" is no longer enough — the first
> machine is judged at a gate by the component's health, three times over at least two minutes within ten,
> before the rest are sent; and a first machine that fails, by its report or at the gate, is not left on
> the build that failed it: the previous build is registered again and sent to it, once per build.
**3. A newer merge takes over an older plan.** When a merge into a repository's branch makes a plan,
every open plan for the same repository and branch made before it is superseded, ordered by when each
plan was made, never by commit:
- the modules the older plan had not yet built join the newer plan's set, before its tiers are computed;
- the older plan ends in a state of its own, *superseded*, naming the plan that took it over.
Builds the older plan already asked for still finish and register; issue 219's ordering keeps the newer
one current. A person can also close a plan that waits on nothing, by its id. The plan is marked closed
by hand and never resumed.
## Consequences
- A module that gains a state, an event or a tool can use it from its first start on every machine.
- A change reaches one machine before the rest. A change that breaks its first machine stops there, with
the reason in the plan.
- Successive merges build each module once, for the newest commit. The build agent is sent to the
machines once per run of merges, not once per merge.
- **What got harder:** a rollout takes one machine's report longer than before. A module that must change
everywhere at once says *together* in its policy. A plan's record now has a superseded state that
readers of the plans must know.
## How it is checked
| Rule | Checked by |
|---|---|
| grants before code | the controller's test: a send records memberships issued before any declaration; the machine holding the bus is sent first when the user list changes; a failed membership fails the send |
| one machine first | the controller's test: with a one-at-a-time policy, one machine is sent, the rest only after its applied and current report; a failed first machine stops the module; *together* sends all at once |
| a newer merge takes over | the controller's test: an older open plan for the same repository and branch is superseded, its unbuilt modules folded in; a plan for another repository is left alone; a superseded plan is not open |
| live | the next merge to the catalogue that gives a module a new state: no refusal of that state on any machine, the first machine named in the plan, one plan open per repository |
## References
- [Issue 249](../04-ISSUES/249-a-modules-new-state-is-refused-until-a-push-the-merge-did-not-make/00-report.md), [issue 254](../04-ISSUES/254-plans-for-successive-merges-run-over-each-other-and-one-was-left-open/00-report.md)
- [ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md) — plans and tiers, extended here
- [ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md) — memberships, kept on the bus
- [to-be 30](../03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md) — the design this amends
@@ -1,120 +0,0 @@
---
topic: the mesh
status: accepted
date: 2026-10-05
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md
---
# 219. The build queue is controlled through the controller and the build seat
## Context
The operator asked for tools to control the mesh's builds: see what is queued and running, cancel,
clear, stop a build immediately, pause and continue, restart and replay. On the day of the request none
existed. The controller could ask for a build and list finished ones. Nothing could see an ask waiting on
the build seat's work queue, or one being built. Nothing could take an ask back, and a running build
could only be stopped by restarting the machine's build agent, which hands the ask to another holder.
How builds run today, read from the code and the bus
([ADR 0190](0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md),
[ADR 0157](0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md)):
- an ask is a message on the seat's work-queue stream;
- every holder pulls one at a time from one shared worker, and acknowledges after it has announced the
outcome;
- a build says it started, logs its steps, and announces what it built, all on the bus;
- an ask delivered five times without an answer stays in the stream for a week, with nothing saying so.
Three facts constrain any control:
- **Only the controller may act on the bus's streams and consumers** (design 25 §3, enforced in the bus's
user list). Holders may only take from their worker and acknowledge.
- **A plan waits for a build's outcome and has no timeout.** An ask that disappears without one leaves
its plan waiting, said only as late after half an hour.
- **The bus server in use cannot pause a consumer**; that arrived in a later version.
## Considered Options
1. **Pause and cancel by remaking the worker consumer.** Rejected: a remade worker delivers from now on,
so every queued ask would be skipped silently. Issues 206 and 207 are this mistake both ways round.
2. **Upgrade the bus first and use its consumer pause.** Not now: it gives pause and continue, and
nothing else on the list, and upgrading the bus is a change of its own.
3. **Queue actions are the controller's verbs, process actions are the build seat's verbs, and every
action that drops work leaves a failed outcome.** Chosen.
## Decision
**1. The queue is the controller's.** Its verbs act on the work-queue stream, the one thing only it may
touch:
| verb | what it does |
|---|---|
| `queue` | lists every ask: waiting, in flight (with the machine building it, from its started event), and dead (delivered as often as allowed, still in the stream) |
| `cancel <id>` | takes back a waiting or dead ask; an ask in flight is refused, and `kill` is named instead |
| `clear` | cancels every waiting ask, and with `--dead` every dead one |
| `rebuild <module or build>` | asks again for a module's current source, or a past build's source and ref, under a new id |
| `replay <build>` | asks again for a past build at its commit, as a **dry run** unless told to register |
| `kill <id>`, `pause [node]`, `resume [node]` | pass the request on to the build seat on the right machine, or on every machine |
**2. The process is the build seat's.** Each holder serves verbs on its own machine:
- `current`: the build running here, its step and how long, and whether this holder is paused;
- `kill`: stops a running build at once. The build's whole process group is ended, along with every
container it started. Its outcome is announced as failed, and the ask is acknowledged, so it is not
delivered again;
- `pause` and `resume`: a paused holder takes nothing new, and a running build finishes. The flag
survives the holder's restart.
A holder restarted mid-build keeps today's behaviour: the ask is not acknowledged, and another holder
takes it.
**3. Nothing dropped is silent.** Cancel, clear and kill each leave a failed outcome for the ask's id,
taken in like any other. A plan waiting on that build fails, and says why, instead of waiting. A plan keeps
the id it asked for, so it can match its outcome exactly. A holder checks a cancelled ask before it builds
it, so an ask taken in the instant it was cancelled is not built.
**4. A plan says what its builds are waiting on, and a failed plan can go on.**
- A plan whose build waits on a paused seat says the seat is paused, and on which machines, and is not
counted late while it waits.
- `plans retry <id>` asks again for the modules a failed plan could not build, under new ids. The plan
resumes at that tier and goes on through its later ones. A plan another has superseded, or one
already done, is refused.
- `rebuild` of a module that an open or failed plan has not yet built joins that plan, so the plan and
the build are one thing.
**5. Replay does not move the mesh backwards unasked.** A replayed build is a dry run: built, its log
kept, nothing registered. With `--register` it is registered. If a newer build of the module is already
registered, that is refused unless `--older` is said as well: registering an older commit makes it the
current one, and the rollout policy sends it to the machines ([issue 207](../04-ISSUES/207-a-re-made-worker-replayed-every-ask-the-stream-kept/00-report.md)).
## Consequences
- Every build in the mesh can be seen, taken back, stopped or asked again from the console. None of it
needs a shell on a machine.
- A cancelled or killed build shows in the build records as failed, with who stopped it. Its plan fails
saying the same.
- **What got harder:** a holder now serves verbs as well as taking work, and keeps one small flag on
disk. Pause is per holder, so "pause the mesh" is the controller asking every holder in turn. A holder
away at the time misses it, and the answer names that holder.
## How it is checked
| Rule | Checked by |
|---|---|
| the queue is read and classified right | the controller's test: waiting, in flight and dead asks told apart from the stream and the worker; a live test on a throwaway bus |
| nothing dropped is silent | the controller's test: cancel and clear delete the ask and record a failed outcome; a plan asked for that id fails |
| a cancelled ask is not built | the holder's test: an ask taken after its cancel is answered failed without building |
| kill stops everything it started | the holder's test: the build's process group and its labelled containers are ended; the outcome is failed and the ask acknowledged |
| pause survives a restart | the holder's test: the flag is read back at start, and nothing is taken while it is set |
| plans follow the queue | the controller's test: a plan waiting on a paused seat says so and is not late; `plans retry` resumes a failed plan and its later tiers are asked; `rebuild` joins the plan that holds the module |
| replay is safe | the controller's test: a dry run by default, and `--register` over a newer build refused without `--older` |
## References
- [ADR 0190](0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md) — holders share the seat's work, one at a time
- [ADR 0157](0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md) — a build's own events are its record
- [issue 207](../04-ISSUES/207-a-re-made-worker-replayed-every-ask-the-stream-kept/00-report.md) — a remade worker replayed every ask
- [to-be 18](../03-DESIGN/01-to-be/18-building-a-module.md) — the design this amends
@@ -1,161 +0,0 @@
---
topic: what runs on it
status: accepted
date: 2026-10-05
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md
---
# 220. What a machine asks needs its uplink held, and the retired resolver pieces go
> **Decided to go — 2026-10-05, by [ADR 0223](0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md), not yet built.** `/etc/resolv.conf` becomes the `node-uplink`
> holder's file, and `resolv-conf`, `node-resolver-config` and this record's dependency of it on
> `node-uplink` retire with it. Until that is built, everything below stands as decided.
## Context
**Three things about a machine's resolver were left half done when the mesh moved to one resolver.**
On the production mesh on 2026-10-05, read from the controller's `seats` verb:
- **`node-dns-resolver` has no holder on any node, and no module in the catalogue claims it.**
[ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) retired it
and the controller kept its row deliberately, *"deleted once nothing claims it"*, because removing a
seat a machine still holds makes that machine unresolvable. That condition now holds. The row still
stands in the controller's compiled set and in the store's seat table, and the overview still lists
it, unheld, beside the seats a mesh actually has.
- **The rule that keeps `/etc/resolv.conf` the mesh's is checked by nothing.**
[ADR 0117](0117-a-machines-uplink-is-a-seat.md) found that a network manager rewrites the resolver
file on every connectivity change unless it is told not to, and gave that telling to the module
holding `node-uplink`. It said the condition *"only if NetworkManager runs"* is expressed by
assigning the manager's module. Nothing makes anybody do so: `resolv-conf` can be assigned to a
machine with no uplink holder, and the file is then replaced the first time a laptop changes
network while every surface of the mesh reads green. Today every node holding
`node-resolver-config` also holds `node-uplink` — NetworkManager on the home server, the
workstation and the laptop, systemd-networkd on the anchor — by care, not by check.
- **The catalogue still carries a systemd-resolved split-DNS module**, `resolved-split-dns`, claiming
`node-resolver-config`. [ADR 0196](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md)
chose against a stub on every node and says *"There is no `systemd-resolved` module."* It is
assigned nowhere. `resolv-conf`'s own resolver file still tells its reader that systemd-resolved or
NetworkManager may be assigned *instead* — the opposite of how the roles now divide.
**A dependency mechanism already exists.** [ADR 0207](0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md)
made a module depend on the node seats that apply its resources, derived rather than stated, judged
over the node's whole set of assignments, refused at `assign` naming the seat and its possible holders,
and refused at composition. [ADR 0210](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md)
derived a second kind from contributions. What is missing is a dependency that belongs to a *role*
rather than to what a module declares.
## Considered Options
**For the retired seat:**
1. **Keep the row until the build seat's retired row goes too, and delete both together.** Rejected:
the two have nothing in common but having been retired; one is unclaimed now and the other is not
yet known to be.
2. **Remove it from the compiled set only.** Rejected: seeding adds a seat a release ships and never
removes one ([ADR 0122](0122-a-seat-is-data-a-rename-is-a-database-update.md)), so the store's
row, which is the live set, would stay.
3. **Remove it from the compiled set and delete the store's row in a numbered migration**, as the
artifact store's rename did for its old row. Chosen.
**For the resolver file and the uplink:**
1. **Leave it to the operator.** Rejected: it is the failure ADR 0117 describes — the file silently
replaced — with the one difference that the operator was told.
2. **`resolv-conf` declares the manager's settings itself.** Rejected by ADR 0117 already: which
setting depends on which manager runs, and a resolver module that knew about network managers would
be the wrong module knowing the wrong thing.
3. **A manifest field in `resolv-conf` naming `node-uplink`.** Rejected for the reason ADR 0207
rejected its own option 2: a second module claiming the same seat would have to restate it, and
one that forgot would pass.
4. **The seat carries what its holder needs beside it.** `node-resolver-config` names `node-uplink`;
any module claiming the former depends on the latter, derived from the claim and judged exactly as
ADR 0207 judges a resource's dependency. Chosen.
**For the split-DNS module:** keep it for a machine that wants systemd-resolved in charge, or remove it.
Kept, it is a second answer to a question ADR 0196 settled, and a claimant the catalogue offers
without a record allowing it. Removed.
## Decision
**1. `node-dns-resolver` is deleted from the mesh's set.** The controller's compiled set no longer
carries it, and a numbered migration of the controller's store deletes its row and any alias naming
it. No alias is kept: nothing was renamed, and a manifest still claiming it should be refused at
registration, naming the seat. This completes ADR 0194's retirement; nothing it decided changes.
**2. A seat may name the node seats its holder needs held on the same node.** A module claiming such a
seat depends on each of them. The dependency is a third source beside ADR 0207's resources and ADR
0210's contributions, and everything ADR 0207 §3 and §4 say of those applies unchanged: met by any
module assigned to the node, the claimant included; judged over the node's whole set; refused at
`assign` naming the seat and the catalogue's possible holders; refused at composition; and only said,
never refused, when no module in the catalogue could hold the needed seat. Unassigning the needed
seat's last holder beneath a dependent is refused, naming the dependent. What a seat needs is part of
the mesh's definition of the role: compiled with the set, never stored, as ADR 0212 keeps what a seat
receives. Adding a need to a seat is a decision, recorded.
**3. `node-resolver-config` needs `node-uplink`.** The holder that writes the resolver file is right
only while the network manager is told to leave it alone, and that telling is the uplink holder's
(ADR 0117). Every manager the catalogue knows — NetworkManager, systemd-networkd, dhcpcd — holds
`node-uplink`, so the refusal always has a remedy to name.
**4. `resolved-split-dns` leaves the catalogue.** `resolv-conf` is the only module claiming
`node-resolver-config`. Its resolver file's comment says the uplink's holder is required beside it,
rather than naming alternatives to assign instead.
## Consequences
- **The set reads as the mesh is.** Thirty-seven seats in the compiled set; the overview no longer
lists a role nothing can fill.
- **A machine cannot be given the mesh's resolver file without its network manager being told to keep
off it.** A machine with no manager at all — a static configuration — needs the smallest holder,
`dhcpcd`, or a new module holding `node-uplink` for its way of configuring the link. That is the
point: such a machine has to say what manages its link before the mesh writes a file the manager
could overwrite.
- **Order of assignment on a new machine**: the uplink holder before or with `resolv-conf`, in one
`assign` when together. On the production mesh nothing changes: every node already holds both.
- **The uplink becomes harder to take away.** Unassigning a machine's manager module while
`resolv-conf` stays is refused; replacing one manager with another is one act assigning the new and
unassigning the old, or the dependent goes first.
- **A machine wanting systemd-resolved has no module for it.** A future need for one is a new record,
not a revival of the removed module.
- **Changing `resolv-conf`'s comment rewrites `/etc/resolv.conf` on every node once**, with the same two
nameserver lines and options; only the comment differs.
- **The merge order matters.** The controller's tests read the catalogue beside them, and the two
changes are judged together: the controller's change and the catalogue's removal merge together,
the catalogue's first or in the same window, and the controller rolls out only once its test suite
passes against the merged catalogue.
## How it is checked
| Rule | Checked by |
|---|---|
| `node-dns-resolver` is not in the set, and the set has thirty-seven seats | mesh-controller's closed-set unit test on the compiled seats |
| The store's row goes with it | the migration, and after rollout the controller's `seats` verb listing no `node-dns-resolver` |
| `node-resolver-config` needs `node-uplink`, and a module claiming it depends on the uplink with nothing in its manifest | mesh-controller's seat-dependency tests on the seat definition and on a synthetic claimant |
| What a seat needs survives loading the set from the store | a unit test loading the store's rows, which carry no such column |
| `resolv-conf` without an uplink holder is refused at `assign`, naming `node-uplink` and its possible holders; beside one, or with one in the same act, it passes; a composition without one is refused | the same tests, and `assign` live |
| Unassigning the uplink's last holder beneath `resolv-conf` is refused | an unassign test |
| In the catalogue, `resolv-conf` depends on the uplink, dhcpcd, NetworkManager and systemd-networkd each hold it, and `resolv-conf` is the only claimant of `node-resolver-config` | a mesh-controller test reading the catalogue beside it |
| Two modules deciding what a machine asks are still refused on one node | the resolver test, now with a synthetic second claimant |
| Every node of the live mesh holding `node-resolver-config` also holds `node-uplink` | the controller's `seats` verb, read before this was decided and after it rolls out; `status` reports no unheld dependency |
## References
- [ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) — retired
`node-dns-resolver`; this record deletes it.
- [ADR 0196](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md) — no
stub, and so no systemd-resolved module.
- [ADR 0117](0117-a-machines-uplink-is-a-seat.md) — the uplink's holder keeps the manager off the
resolver file; this record makes that a checked dependency.
- [ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md) — serving and
asking as two seats.
- [ADR 0207](0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md),
[ADR 0210](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md) — the dependency mechanism
this extends.
- [ADR 0122](0122-a-seat-is-data-a-rename-is-a-database-update.md) — the set as data, which is why a
deletion is a migration.
- [The seats](../03-DESIGN/01-to-be/26-the-seats.md) and
[connectivity §2](../03-DESIGN/01-to-be/08-connectivity.md), amended alongside.
- mesh-controller `internal/catalogue/seats.go`, `internal/catalogue/seat_dependencies.go`, and the
store migration deleting the row; mesh-catalog `modules/resolv-conf`.
@@ -1,129 +0,0 @@
---
topic: the mesh
status: accepted
date: 2026-10-05
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
---
# 221. A push sends no build a policy or a plan holds back, except to the machine it names
## Context
[ADR 0083](0083-one-push-leaves-the-mesh-consistent.md) makes a named push finish what it starts. After
the named machine is sent, every other machine whose declaration differs from what it was last sent is
sent too. The case it was written for is a grant: assigning a consumer changes the provider's
declaration on another machine. ADR 0083 accepted that a machine behind *for an unrelated reason* is
flushed as well, and called that correct rather than a cost.
On 2026-10-05 that reasoning met an upgrade policy
([issue 259](../04-ISSUES/259-a-named-push-sent-every-machine/00-report.md)). A change to the resolver
modules was merged with the policy `record`, so that each machine would take it only when pushed, one
at a time: the anchor first, each checked before the next. `push <anchor>` sent all four machines the
new build, and so did a later `push <laptop>`. A fault in the change
([issue 260](../04-ISSUES/260-the-resolver-started-before-its-zones-file-existed/00-report.md)) was met
on every machine at once.
The cascade compares one digest per machine, and a digest cannot say why a machine differs. Under
`record`, every machine running the module differs from the merge on, so every machine is flushed. The
same holds for a plan's rollout under [ADR 0218](0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md):
while the plan waits on its first machine, the rest differ, and any named push elsewhere sends them the
build the plan is holding. [Issue 249](../04-ISSUES/249-a-modules-new-state-is-refused-until-a-push-the-merge-did-not-make/00-report.md)
met the same confusion for the machine holding the bus, which a send adds when its user list changed.
It narrowed that check to the user list, but the machine, once added, is still sent its whole
declaration.
So `record` held a change back from nothing but the merge, and "one machine first" held it back only
from the plan's own sends.
## Considered Options
1. **Report instead of cascade:** list every machine that is behind and send none. Rejected: it is the
option ADR 0083 rejected, and for the same reason. [Issue 057](../04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md)'s provider would wait for a second push
nothing tells anyone to make.
2. **Send a held machine its consequences and keep its held modules at their last build.** The cascade
would compose the machine with each held module pinned at the build it was last sent, so a grant
still reaches it and the upgrade does not. Rejected: the catalogue resolves every machine against
one manifest per module, the current one. Pinning means resolving a machine's set against a mix of
current and older manifests, read back from the build records, beside the current settings, seats
and grants. The result matches neither what the machine runs nor what the mesh would send it, so
neither `plan` nor `status` could show it. A grant composed for the new build may not fit the old
one. It is a second composition path to get one edge case right, and the edge case has a one-word
remedy: name the machine.
3. **Keep, with every send, which build of each module the machine was sent, and leave a machine any
of whose modules a policy or a plan holds back.** Chosen.
## Decision
**1. A send records the builds it carried.** With the digest of every declaration it sends, the mesh
keeps which build of each module the declaration carried: the commit the module's current build was
made from. A module left out of the declaration keeps the build it was last sent. A declaration sent by
hand records that what it carried is not known.
**2. A push does not send a machine it did not name a held build.** A machine reached by a named push's
cascade, or added because it holds the bus, is not sent when any module it runs would move to a build
that:
- its upgrade policy records rather than rolls out, or
- an open plan has not yet sent it: the plan is still building the module, or has sent it to its first
machine and this is not that machine.
A module the machine was never sent counts as a move. A machine whose last send's builds are not known
counts as held. It was sent before this was kept, or by hand, so a held upgrade cannot be told apart
from anything else.
**3. It is named, not hidden.** The push says which machine it left, which module and which builds, why,
and that `push <node>` sends it. For the machine holding the bus it also says that the bus may refuse
what this push's machines were newly granted until that machine is sent.
**4. Everything else stays as it is.** A machine with nothing held is flushed exactly as ADR 0083
decides: a grant, a peer, a setting. The machine a push names is sent everything, held builds included.
A push that names no machine sends every machine. `push --behind` still sends every machine that is
behind, held or not. It is the remedy `record` names when an upgrade is announced ("`push --behind`
when you want them"), and the one command for taking a recorded upgrade everywhere. Narrowing it would
leave `record` with no way to say "now".
> **The mechanism changed — 2026-10-06, by [ADR 0236](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md).** What still stands: a send records
> the builds it carried, a held machine is left and named, the named machine is sent everything, and
> `push --behind` takes a recorded upgrade everywhere. What moved: `record` is no longer the default — a
> build rolls out one machine first, gated, unless a person, the module, its irreplaceable data or the bus
> says otherwise — and one case of "the machine a push names is sent everything" is refused: a machine
> whose bus would move, which is replaced only as the planned step `bus upgrade`.
## Consequences
- `record` and "one machine first" hold a change back from every push that does not name the machine.
Walking a change through the mesh is `push <anchor>`, check, `push <next>`.
- **The cost:** a held machine that is also owed a grant from this push waits for its own push. The
provider in issue 057's case, if a held upgrade is pending on it, is not sent its new grant, and its
consumer is refused until the provider is pushed. The push names that machine and the remedy, so the
wait is announced, not silent. When the holder of the bus is held, a new grant may be refused by the
bus until it is sent, and the push says so.
- On the first push after this ships, no machine's last send has its builds recorded yet. Each is held
from cascades until it is pushed once: by name, in a whole-mesh push, or by `push --behind`.
- A manifest handed over by hand does not change the commit a module records, so a cascade still sends
its change. Only builds have a commit to compare.
- ADR 0083's consequence that a machine behind for an unrelated reason is flushed is narrowed. It still
holds for every reason except a build held back.
## How it is checked
| Rule | Checked by |
|---|---|
| a send records the builds it carried, and "not known" | the controller's test: the builds kept with a send round-trip, an empty send is known and empty, a send by hand reads as not known |
| a left-out module keeps its last build | the controller's test: a module left out of a declaration records the build it was last sent |
| a recorded upgrade holds a machine a push did not name | the controller's test: with a module under `record` moved, the cascade of a push naming the anchor leaves the laptop's last send unchanged and names it, the module, both builds and `push laptop` |
| held and owed something else: not sent, both said | the same test: a newly placed machine changes the laptop's peers while the upgrade is held; the laptop is not sent and the output says why and that what else it is owed waits |
| a roll-out policy is not held | the same test: with the policy set to roll out, the laptop is sent and its new build recorded |
| a consequence nothing holds is still sent (issue 057) | the controller's test: a placed machine's peers reach the others by cascade; a machine whose last send is not known is held |
| a plan waiting on its first machine holds the rest | the controller's test: a machine the plan has not reached is held, the first machine is not; a module sent everywhere, failed, or in a closed plan holds nothing |
| live | the next change merged under `record`: `push <anchor>` sends the anchor alone and names every other machine running the module |
## References
- [Issue 259](../04-ISSUES/259-a-named-push-sent-every-machine/00-report.md), [issue 260](../04-ISSUES/260-the-resolver-started-before-its-zones-file-existed/00-report.md), [issue 249](../04-ISSUES/249-a-modules-new-state-is-refused-until-a-push-the-merge-did-not-make/00-report.md)
- [ADR 0083](0083-one-push-leaves-the-mesh-consistent.md): the cascade, narrowed here
- [ADR 0218](0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md): one machine first, now held from a push as well
- [ADR 0010](0010-delivery.md): delivery, and what `push --behind` answers
- [to-be 30](../03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md): the design this amends
@@ -1,150 +0,0 @@
---
topic: the mesh
status: accepted
date: 2026-10-05
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md
---
# 222. A module is told where a mesh seat's holder is reached, and the controller writes no file a seat's holder owns
## Context
[Issue 190](../04-ISSUES/190-the-runtimes-configuration-is-written-by-modules-that-are-not-the-runtime/00-report.md)
found the container runtime's configuration file written by modules that are not the runtime's. The
first half is closed: the resolver module no longer writes into that file, and the catalogue's runtime
module (the holder of `node-container-runtime`) writes `live-restore` and reloads its own service. The
second half stands. The controller's private network still generates two resources on every machine
on the network: one writes `insecure-registries`, naming the mesh's artifact store, into the runtime's
file ([ADR 0082](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md) §2,
[ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md)), and the other declares the
runtime's service, reloaded on that file. So two parties declare one path and one unit on every
machine, and nothing refuses it, because the collision check runs over catalogue manifests and the
private network's resources only exist once its generator has answered for a machine.
On 2026-10-05 the operator made the rule general: **the controller never writes a file a seat's holder
owns; it tells the owner.** The hosts file had already moved the same way
([ADR 0199](0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)).
What the runtime's module needs is one fact: the address this machine reaches the mesh's artifact
store at. The controller already composes that address for every machine at every push. It puts it
into every image and archive reference the mesh built, and never stores it. Today a module has no way
to ask for it. A binding gives a consumer the provider's address and a credential. The `${seat:…:<port>}`
placeholder gives only a port, and only for the store and the broker the controller itself dials.
Two facts about the move itself, read from the host's code:
- The private network writes one member into a list. The host adds a list member and records it per
resource, so between the two writers declaring it and one of them going, the member is owned by
whichever record added it first.
- The host removes every resource no longer declared before it applies anything, so in the apply
where the private network's record goes and the runtime module's member arrives, the member leaves
and comes back with no reload in between.
## Considered Options
1. **Keep the private network writing the trust.** Rejected: it is the defect issue 190 names. Three
writers became two, and the second is computed code no check can see.
2. **An operator setting on the runtime module naming the registry**, as issue 190 first proposed
under [ADR 0164](0164-a-setting-is-declared-with-its-default-its-meaning-and-what-changing-it-costs.md).
Rejected: where the store is, is a fact the mesh holds, not a choice an operator makes. A setting
would be a copy that goes wrong the day the store moves.
3. **The runtime module requires `artifact-store` and reads the address from its binding
(`${bound:…}`).** Rejected. A binding makes the module the store's consumer, and that mints a
credential for every machine's runtime, which needs none: reading from the store needs presence on
the private network and nothing else (ADR 0082 §3). It also makes a cycle: the store runs as a
container in the runtime, and the runtime would require the store before it could be assigned.
4. **A placeholder that answers where a mesh seat's holder is reached, with no binding.** Chosen. It
is the reasoning of `${seat:…:<port>}` taken one step further: nothing is required, nothing is
granted, and the answer is an address the mesh already holds in the clear.
## Decision
**1. `${seat:<mesh seat>:reach}` is where this machine reaches the holder of a seat the mesh holds once,
as host:port.** It is filled where every placeholder is, in a file's content and in an environment
value, from the address the controller composes for that machine at that push. It requires nothing
and grants nothing, and no credential comes with it.
- **Only `mesh-artifact-store` is answered.** Another seat is refused by name, never answered with
nothing, so a module asking a question the mesh does not answer learns that at composition. A seat
is added to what is answered by a record saying why its address is needed.
- **The answer may be empty**: no machine on the private network holds the seat yet, which is how
genesis begins. In a file written into as JSON, an empty member is dropped from its list, and a list
left with no members is dropped, so the software is never told an empty name.
**2. The runtime's module states the registry's trust.** Its `daemon` resource writes
`insecure-registries`, naming `${seat:mesh-artifact-store:reach}`, beside `live-restore`, and it
reloads its own service. ADR 0082's decision stands: being on the private network is what grants the
trust, the private network is the transport security, and no module author chooses it. Only who writes
it moves, from the private network to the runtime's module.
**3. The controller writes no file a seat's holder owns.** The private network generates neither the
runtime's file nor its service.
**4. What the mesh computes is held to the collision check.** At composition, each computed module's
resources, as its generator answers for that machine, are checked beside the other modules' resources.
The check is the same one: no two modules on a machine declare one path, unit, name or package. A
collision is refused by name. The generator is asked once, and what is checked is exactly what is
declared.
**5. A unit still held is not given back when another record of it goes.** When the host removes a
service resource whose unit another declared resource still gives a state, it forgets that record and
leaves the unit as it is. What the record going had found is not the machine's to restore while
another resource holds the unit, and restoring it could stop the runtime only for the remaining
resource to start it again in the same apply.
## Consequences
- **Rollout has a fixed order, in three steps.**
1. The controller learns the placeholder. Nothing uses it yet.
2. The catalogue's runtime module uses it. A controller that does not know the placeholder would
send it through unfilled, so this waits until step 1 is deployed.
3. The controller stops generating the private network's two resources and checks generated
resources for collisions. With step 3 before step 2, machines would lose the trust. The host
change in Decision 5 is deployed before step 3.
This replaces issue 190's single push. That was needed when two writers of one scalar key would
have been refused. A list member written by two records is tolerated by the host.
- **In the apply of step 3**, the private network's records go first: the member leaves the list and
the service record is forgotten. Then the runtime module's file is applied, and the member is added
back, recorded as the runtime module's. The runtime is reloaded once. The address is unchanged, so
the trust never lapses for a pull.
- **A machine holding the store, before any network exists**, is answered with the loopback address
it reaches the store at. A runtime already trusts loopback, so this adds a redundant member, not a
wrong one.
- **A second writer cannot return through generated code.** It is refused at composition, naming
both modules and what they share.
- **A module can now learn where the artifact store is without being its consumer.** That is a
capability, and it is bounded: one seat today, and widened only by a record.
## How it is checked
| Rule | Checked by |
|---|---|
| `${seat:mesh-artifact-store:reach}` is filled with this machine's address for the store, in a file and in an environment value | mesh-controller `TestASeatsReachIsWhereThisMachineReachesItsHolder`, and through the whole composition `TestTheRuntimesTrustIsComposedFromTheSeatsReach` |
| An empty answer adds nothing to a list, and other members stay | mesh-controller `TestAnUnansweredReachAddsNothingToAList` |
| Any other seat is refused by name | mesh-controller `TestAReachTheMeshDoesNotAnswerIsRefused` |
| The catalogue's runtime module writes `live-restore` and the trust, and writes no trust when no store is reachable | mesh-controller `TestTheRuntimesModuleTrustsTheMeshsRegistry`, which composes the catalogue's manifest beside it |
| The private network declares neither the runtime's file nor its service | mesh-controller `TestTheNetworkWritesNothingOfTheRuntimes` |
| A generated resource colliding with a module's is refused by name; disjoint ones compose; the generator is asked once | mesh-controller `TestAGeneratedResourceCollidingWithAModulesIsRefused`, `TestAGeneratedResourceBesideAModulesOwnIsComposed`, `TestAComputedModuleIsAskedAboutTheNodeItIsFor` |
| The member moves between records in one apply without leaving the list; one reload; a unit still held is forgotten, never stopped or disabled; the plan says so | mesh-host `TestTheRegistryMovesToTheRuntimesModuleWithoutLeavingTheList`, `TestThePlanForgetsARecordOfAUnitStillHeld` |
| After rollout, every machine's runtime trusts the store's address | the runtime module's `docker_daemon_config` tool on each machine, read before step 3 and after it |
## References
- [Issue 190](../04-ISSUES/190-the-runtimes-configuration-is-written-by-modules-that-are-not-the-runtime/00-report.md),
steps 2 and 5.
- [ADR 0082](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md): the trust and why it
is plain HTTP; its mechanism moves here.
- [ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md): writing into a shared file;
its writer of the runtime's trust moves here.
- [ADR 0079](0079-the-foundation-seats-are-named-after-their-servers.md): a mesh seat names the
server, which is what lets a placeholder ask about it.
- [ADR 0166](0166-the-container-runtime-is-a-node-seat-and-the-host-creates-containers-through-its-holder.md):
the runtime module given the registry as a value.
- [ADR 0199](0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md):
the hosts file left the controller the same way.
- mesh-controller `internal/catalogue/seat_into.go`, `internal/catalogue/declaration.go`
(`generatedHere`), `internal/overlay/generator.go`; mesh-catalog `modules/docker`; mesh-host
`internal/apply/apply.go` (orphan removal).
@@ -1,204 +0,0 @@
---
topic: the tiers
status: accepted
date: 2026-10-05
deciders: jochen
reconstructed: false
supersedes-in-part:
- 0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md
extends: 02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
---
# 223. The mesh has two resolvers, and a machine lists only them
## Context
**[ADR 0196](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md) wrote
every machine's `/etc/resolv.conf` as the mesh's resolver first and a public resolver second.** Its
reasoning is the C library's: servers are asked in order, the next only when one does not answer, and
an answer from the first — "no such name" included — is final. That holds for glibc. It does not hold
for musl, the C library of every Alpine image: musl sends the query to every listed server at once
and takes the first reply.
**On the production mesh on 2026-10-05 that made the anchor's own name unresolvable from the home
server's builds.** The build agent on the home server runs its steps in Alpine containers on the host
network, so they read the machine's `resolv.conf` as it is. Asked for `<anchor>.internal`, the public
resolver — which has no such name and is the nearer of the two — answered NXDOMAIN first, and musl
took it. Reproduced six times out of six inside the build agent's own container; every build on that
machine failed fetching from the anchor. The same lookup from glibc on the same machine answered every
time. [Issue 262](../04-ISSUES/262-an-alpine-container-could-not-find-a-machine-by-its-mesh-name/00-report.md)
was the same library failing on a different answer (NXDOMAIN for a missing IPv6 record); fixing that
did not touch this one, because here the wrong answer comes from a server that should never have been
asked.
**The public line was there for one case: the mesh's resolver unreachable.** ADR 0196 kept public
names resolving with the anchor down, the tunnel down, or a laptop behind a captive portal. That case
is real and rare; the musl case is every lookup of a mesh name from every Alpine container on any
machine that is not the anchor.
**A seat's work can already be shared by several holders.** [ADR 0190](0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md)
made the build role node-scoped with every holder pulling from one queue. A mesh-scoped seat has
always had exactly one holder: by derivation, or on record since
[ADR 0131](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md), where a handover replaces the
holder in one write and every other eligible assignment stands beside it, silent.
## Considered Options
1. **The mesh's resolver alone.** Every mesh name and every public name answered consistently, by one
server. Rejected alone: with the anchor down, no machine resolves anything — the reason ADR 0194
gave against it, and still true.
2. **A local forwarder on every machine** — a small resolver on loopback that asks the mesh's resolver
for the mesh's suffix and public resolvers for the rest, with `resolv.conf` naming only it. Rejected:
it is ADR 0194's per-node stub again, which ADR 0196 removed — a daemon and a module on every node,
and a container cannot use a loopback resolver, so containers would need a second configuration.
Every resolution fault found on 2026-10-03 was a per-node copy disagreeing with the truth.
3. **Split by kind of machine** — servers list the mesh's resolver alone, laptops keep a public
fallback. Rejected: a laptop runs Alpine containers too, and a rule that differs by machine is a
rule nobody can state about the mesh.
4. **Two mesh resolvers, and nothing else listed.** The same module, the same roster and the same zones
on two machines, and every machine lists both. Whichever answers first gives the same answer, so
musl's race is harmless; a public name still resolves through either. Chosen.
## Decision
**1. Now: the mesh has two resolvers.** The seat `mesh-dns-resolver` may be held on more than one
machine — the anchor and the home server — each holder answering the same mesh names: the same
module, the same machine list, the same zones, rendered by the controller into each.
- **A seat may be replicated.** It is an attribute of the seat in the mesh's definition, compiled with
the set and never stored, as what a seat needs ([ADR 0220](0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md))
and what it receives are. `mesh-dns-resolver` is the only replicated seat. Making another one is a
decision, recorded. In the glossary's words it is the mesh's first **bench** — a seat whose
holders coexist — and *replicated* says what kind: every holder answers the same thing.
- **Every holder is on record, and each is added by an act**: `seat mesh-dns-resolver --add
<node>/<module>` records one more holder beside those on record. Assigning the module is not enough:
an assignment not on record stands beside the holders, eligible and silent, exactly as ADR 0131 says,
and two claimants with nothing on record are refused as for any mesh seat. `--to` still hands the
seat over, leaving exactly one holder. `--add` on a seat held once is refused, naming `--to`.
- **One per machine still**: two modules on one machine claiming it are refused. And a seat held once
stays held once: a second claimant on another machine is refused while nothing is on record, and a
store recording two holders of such a seat is refused, naming the seat.
- **Every machine's `/etc/resolv.conf` lists every holder's private address, the holder on the machine
itself first if it is one, then the rest in name order, and no public resolver.** The controller
gives a module's roster template the holders of each replicated seat, ordered so; the module holding
`node-resolver-config` writes the file from it. On a holder, its own private address is first: the
resolver listens there and on loopback, and the private address is the one a container on that
machine can reach.
- **The requirement stays.** What writes the file still requires `wildcard-resolution`, so a machine is
refused when nothing in the mesh resolves, rather than given a file listing nothing. A holder answers
its own requirement; any other machine is bound to the first holder by name. Nothing reads that
binding's address any more — the file lists every holder — and the binding is kept for the refusal
and the order of delivery.
**2. Next, decided and not yet built: `/etc/resolv.conf` belongs to the uplink's holder.** The program
that manages the machine's network already has to be told to keep off the file
([ADR 0117](0117-a-machines-uplink-is-a-seat.md)); instead, the `node-uplink` holder writes it, given
the resolvers by the mesh: NetworkManager through its global DNS configuration, dhcpcd through static
nameservers, and systemd-networkd's module declaring the file itself. `resolv-conf` and the seat
`node-resolver-config` then retire, and ADR 0220's dependency of `node-resolver-config` on `node-uplink`
goes with them. One owner for the file, and it is the program that would otherwise rewrite it.
> **Progressive insight — 2026-10-05.** Built, the managers' own mechanisms turned out not to write
> the mesh's file: NetworkManager's global DNS configuration and dhcpcd's resolv.conf hook each write
> `/etc/resolv.conf` in their own form — their own header, their own options line — and dhcpcd reads
> its configuration only at its next start, so a change of resolvers would wait for one. The record
> said NetworkManager and dhcpcd would be given the resolvers "through its global DNS configuration"
> and "through static nameservers"; instead every one of the three modules declares the file itself,
> from one template, and keeps its manager off it as before (`dns=none`, `nohook resolv.conf`, and
> nothing for systemd-networkd). The decision — the uplink's holder owns the file, and `resolv-conf`,
> `node-resolver-config` and ADR 0220's dependency retire — stands. How parts 2 and 3 are checked is
> in [connectivity §2](../03-DESIGN/01-to-be/08-connectivity.md).
**3. Next, decided and not yet built: a machine's names are one seat's.** `/etc/hosts` and
`/etc/hostname` belong to one seat for the machine's identity. The `hosts` module, holding
`node-hosts-file` ([ADR 0199](0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)),
is renamed **`hostname`** and holds the seat renamed **`node-hostname`**; it writes `/etc/hostname`
and the machine's own `127.0.1.1` line, and keeps every operator's line as ADR 0199 does today.
- **Why one seat for both files**: the mesh's only content in `/etc/hosts` is the machine's own name,
and nothing owns `/etc/hostname` today. Two files saying one fact belong to one owner.
- **Why `node-hostname`**: a system seat is named `node-` for its scope and then for its role
([ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md)), and the role
is the machine's name. `node-identity` was considered and rejected: a machine's identity in the mesh
already means its key and certificate. `node-host` was rejected for the reason the module name
`host` was: "the host" is what the [node-engine](../00-META/glossary.md) was called until now, and
its repository and binary still carry `mesh-host`.
- **Why a module and not part of the node-engine**: the node-engine applies every module's resources
and owns no file's content; every file it writes belongs to the module that declared it. A file's
content belongs to a seat's holder, and a seat's holder stays replaceable — another module can hold
`node-hostname` on a machine that names itself another way, and the node-engine does not change.
- The rename goes through the seat set's rename (an alias keeps `node-hosts-file` resolving,
[ADR 0122](0122-a-seat-is-data-a-rename-is-a-database-update.md)), so nothing claiming the old name
breaks in between.
## Consequences
- **musl and glibc agree.** Every listed server answers a mesh name the same way, so the race musl
runs has one possible outcome; a public name resolves through either holder's upstreams.
- **Accepted cost: with no mesh resolver reachable, a machine has no DNS at all** until it can reach
one. The anchor is the hub of the private network, so with it down the private network is down too,
and the second resolver is reachable only on its own machine and from machines on its own LAN. The
mesh assumes the anchor is up about 99.9% of the time. The same holds for a laptop behind a captive
portal that keeps the tunnel down: it resolves nothing, the portal's own name included, until the
tunnel is up.
- **The resolver's options change from one attempt to two**, one second each. With no public resolver
to fall back to, a single dropped datagram would otherwise fail a lookup on a machine that reaches
only one resolver.
- **Holdings are keyed by seat and assignment.** The store's holding table takes a numbered migration;
every existing row is one per seat and satisfies the new key. A handover replaces every holder in one
transaction.
- **Unassigning a holder takes its own row only**: the other resolver keeps holding. Removing the
second resolver is unassigning it, then `push --behind`.
- **The rollout order matters.** The controller that knows replicated seats and renders the holders
rolls out first; the catalogue's `resolv-conf`, which reads the holders, second — a controller without
them cannot render it; and only then is the second holder added, because an older `resolv-conf`
reading its one binding would name the first holder by name order, which may be the new one, and
still list the public resolver.
- **A container keeps the resolvers it started with.** A container on the default bridge copies its
machine's `resolv.conf` when it starts; one on a user-defined network is answered by the runtime's
embedded resolver, which forwards to the servers it read at start. Either keeps the public resolver
until restarted.
- **Parts 2 and 3 change who writes two files on every machine**, each a handover between modules on
the same path; they wait for their own build handoff.
## How it is checked
| Rule | Checked by |
|---|---|
| `mesh-dns-resolver` is replicated, and no other seat is; it survives loading the set from the store | mesh-controller unit test over the compiled set and over a store row without the attribute |
| Two holders on record compose on both, with no refusal, and each lists itself first | mesh-controller resolution and composition test with the catalogue's `dnsmasq` and `resolv-conf` |
| A machine holding no resolver lists both holders, by name, and no public resolver | the same test on a third machine, asserting no public address appears |
| A holder answers its own requirement although another holder sorts first | a resolution test on the second holder (issue 258's case kept) |
| A seat held once still refuses a second claimant, and refuses two holders on record | resolution tests on `mesh-store` |
| Two claimants of the replicated seat with nothing on record are refused, naming the handover | a resolution test |
| Every consumer is bound to the same holder whatever order the mesh was resolved in | a unit test on the holder among providers |
| The store keeps several holders of one seat, once each; a handover leaves one; unassigning takes only its own row | a store test against a live database, through the migration |
| `seat --add` is refused for a seat held once | the controller's `seat` command |
| The resolver's machine list has one host record per machine (issue 262's missing check) | a mesh-controller composition test counting host records |
| `resolv-conf` lists only the seat's holders, with two short attempts | a mesh-controller test reading the catalogue's `resolv-conf` |
| Live: every machine's `/etc/resolv.conf` lists both holders' private addresses, its own first on a holder, and nothing else; an Alpine container on the host network of the home server resolves `<anchor>.internal` every time | after rollout, read through each machine's tools, and `getent hosts` in an Alpine container repeated ten times |
| Parts 2 and 3 | not yet built; their checks are written with their handoff |
## References
- [ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) — the
mesh's resolver; now held on two machines, each holding every node's internal domain.
- [ADR 0196](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md) —
replaced in part: its public resolver second goes. Every node and container asking the mesh's
resolvers for every name, with no stub and no runtime `dns`, stands.
- [ADR 0190](0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md) —
several holders sharing a role, for a node seat.
- [ADR 0131](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md) — holders on record, handover,
eligible and silent.
- [ADR 0220](0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md) — the
uplink's dependency, which part 2 retires.
- [ADR 0117](0117-a-machines-uplink-is-a-seat.md), [ADR 0199](0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md),
[ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md),
[ADR 0122](0122-a-seat-is-data-a-rename-is-a-database-update.md).
- [Issue 258](../04-ISSUES/258-every-machine-bound-the-resolver-to-itself/00-report.md),
[issue 262](../04-ISSUES/262-an-alpine-container-could-not-find-a-machine-by-its-mesh-name/00-report.md).
- [Connectivity §2](../03-DESIGN/01-to-be/08-connectivity.md) and [the seats](../03-DESIGN/01-to-be/26-the-seats.md),
amended alongside.
- mesh-controller `internal/catalogue/seats.go`, `resolve.go`, `roster.go`,
`cmd/mesh-controller/seats.go`, `holdings.go`, and the store migration keying a holding by seat and
assignment; mesh-catalog `modules/resolv-conf`, `modules/dnsmasq`.
@@ -1,141 +0,0 @@
---
topic: the mesh
status: accepted
date: 2026-10-06
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0040-what-a-module-is.md
---
# 224. A provider that keeps failing a consumer is a problem the controller reports
## Context
**A provider's provisioner is the only thing in the mesh that knows whether a provision was made.**
The controller composes a grant, delivers the contributions file and the minted secret, and the node
reports that it applied every resource. Whether the provider then made the consumer's database, client
or bucket is known to the provisioner loop alone ([ADR 0040](0040-what-a-module-is.md),
[ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)), and until now it
said so only in its own journal.
**On 2026-10-05 that cost a day.** The identity provider's database was moved that morning. Its
realm's administrator kept the password it had before, the module's minted one was inert, and its
provisioner failed every consumer every five seconds — about 31,000 refused logins from shortly after
midnight until it was repaired by hand that night. Every surface the mesh has said the mesh was well:
the machines applied what they were sent, the modules were current, `status` printed its all-well
sentence. The same fault had been found and repaired by hand four days earlier
([issue 179](../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)).
[Issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
already taught that "the mesh and the machines agree" is not "it works", and `status` says so under its
all-well line. This is the narrower, checkable half of that gap: not whether a consumer can reach its
provider, which only dialling answers, but whether the provider has told the mesh it cannot do its job.
## Considered Options
1. **Leave it to the journal, and to a person reading it.** Rejected: that is what happened, twice.
2. **The controller dials every provision.** Rejected for now: it is issue 145's open question, it
needs the controller to hold or borrow every consumer's credential, and it would still not say
*why* a provider fails.
3. **The provider reports its standing through the node's report.** Rejected: the node-engine applies
resources and knows nothing of what a module's code does after it starts; the provisioner runs in the
node's tool runtime, which speaks to the bus, not to the host.
4. **The provider announces, as an event, a consumer it keeps failing; the controller keeps the newest
word and `status` names it.** Chosen.
## Decision
**1. A provider announces a consumer it keeps failing.** When a provisioner has failed one consumer —
its create, its periodic check, or reading the consumer's minted secret — for five minutes without a
single success in between, it emits `provisioner.failing`, naming the consumer, the consumer's
machine, the provision, the class of error and the error's first words, since when, and how many
attempts. It says it again every fifteen minutes while it lasts. The first success after that is
`provisioner.recovered`; so is a consumer the mesh stopped asking for, and so is the first success
for each consumer after the provider starts, because a provider restarted after announcing a failure
has forgotten it.
- **The classes** are `credentials-rejected`, `unreachable`, `secret-unreadable` and `refused` (any
other refusal), read from the error's words unless the provider's own code classes it better. They are
what a person reading `status` needs before opening the journal.
- **No secret travels.** The error is the provider's own text with the consumer's password removed, as
its log line already is.
> **The mechanism changed — 2026-10-06, by [ADR 0230](0230-a-consumer-the-mesh-stops-asking-for-is-retired-and-deleted-only-by-a-person.md).**
> What stands: a consumer the mesh stopped asking for is announced recovered. What moved: a consumer the
> provider still holds is no longer withdrawn on the first pass that misses it; it is said recovered when
> it is *retired*, after five passes and ten minutes, or a person's approval. A provider says what it
> retires, approves, re-enables and deletes on an event of its own, `provisioner.retirement`, permitted
> the same way as the two events here.
**2. Every provider may say it, whatever its manifest lists.** The permission to publish the two events
is derived for every module that receives contributions; no manifest declares them. A provider whose
manifest forgot them would otherwise have its announcement refused by the bus and fail as silently as
before.
**3. The controller follows both events from every module, keeps the newest failing word per provider
module, its machine and the consumer, and removes it on recovery.** Who said it is read from the subject
the bus let the provider publish on, never from the body. A recovery that arrives while the store is
away is held and delivered again, because it is said once. The controller's subscription names the two
events with a wildcard for the emitter — the one pattern on its list — rather than a list of providers
somebody would have to extend.
**4. `status` names every consumer a provider still assigned where it ran says it keeps failing**, and
such a consumer breaks the all-well sentence. Its JSON carries them as `failing`; `node show` lists those
whose provider or consumer is on that machine. A provider no longer assigned there has nothing running
to fail anybody, and is not asked about. A standing not said again for thirty minutes is shown with how
long it has been silent: the provider stopped saying anything, and its last word is all the mesh has.
**5. Where a provider's failure has one known cause it can repair safely, it repairs it.** The identity
provider's administrator refusing the mesh's secret is the first: the module now checks the
administrator's login on start, every five minutes and whenever its provisioner is refused, and repairs
a refusal through the server's own bootstrap command — a temporary administrator, the real one's
password set to the mesh's, the temporary one removed, nothing printed — then checks again and says
what it did. A repair that fails is braked, from ten minutes doubling to six hours, and announced; while
the administrator is refused the provisioner stops asking the server, so a lockout policy is not
provoked, and its consumers are still announced failing. The mechanism is the module's
([issue 179](../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)),
the rule is this record's: **detected automatically, repaired where safe, loud where not.**
## Consequences
- **The loop is carried by every Go provider until the Go SDK has one.** The provisioner loop is the
SDK's ([ADR 0039](0039-what-the-sdk-holds-and-refuses.md)); the Go SDK has none yet, so the two Go
providers carry identical copies and a test in each fails when they differ. A TypeScript provider
announces nothing until the TypeScript SDK's loop does the same; until then its consumers fail as
silently as before, and the next provider to be ported to Go closes that gap for itself.
- **The controller's event consumer takes one message at a time**, so a standing can wait behind a
build being acted on for some minutes. Fifteen-minute repetition makes that harmless for a failure;
a recovery is held, never dropped.
- **The store keeps one row per provider module, its machine and consumer**, through a numbered
migration.
- **The rollout order matters.** The controller that derives the permission and follows the events
first; the catalogue's providers second — an older controller refuses nothing that matters, but
every announcement is then refused by the bus and logged by the provider.
## How it is checked
| Rule | Checked by |
|---|---|
| A consumer failed for five minutes without a success is announced, again every fifteen, and recovered on the first success | provider loop tests in mesh-catalog, postgres and keycloak (`standing_test.go`) |
| A failing periodic check and an unreadable secret count, not only a failing create | the same tests |
| A withdrawn consumer and the first success after a start are announced recovered | the same tests |
| The two Go providers carry the same loop | `harness_same_test.go` in each, comparing the files |
| Every module that receives contributions is granted the two events, and no other module is | mesh-controller inventory test over the derived declaration and the bus permissions |
| The controller may hear the two events from any module, and not every event | the same test over the controller's permissions |
| The emitter is read from the subject; a failing word is kept, a recovery cleared, and a recovery held while the store is away | mesh-controller link tests with a fake store |
| One row per provider, machine and consumer; recovered removes only its own | an inventory test against a live database, through the migration |
| `status` names it and is not all-well; the JSON carries it; `node show` names it on both machines; an unassigned provider's word is not a problem | mesh-controller command tests against a live database |
| The identity provider repairs a refused administrator, verifies, brakes a failed repair, never puts a secret in a command line, and stops asking the server while refused | keycloak module tests with a fake server and a fake container runtime; a live test against a throwaway server when asked for |
| Live: after rollout, `status` shows nothing failing on a healthy mesh; a provider made to fail a consumer for five minutes appears, and disappears on its next success | read through the console after rollout |
## References
- [Issue 179](../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)
— the failure, twice, and the repair this record makes automatic.
- [Issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
— the scope of the all-well sentence, which this narrows and does not close.
- [ADR 0040](0040-what-a-module-is.md), [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md),
[ADR 0039](0039-what-the-sdk-holds-and-refuses.md), [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md).
- [The module protocol](../03-DESIGN/01-to-be/19-the-module-protocol.md), provisioning, amended alongside.
- mesh-controller `internal/link/standing.go`, `internal/inventory/standing.go` and its migration,
`cmd/mesh-controller/standing.go`; mesh-catalog `modules/postgres` and `modules/keycloak`.
@@ -1,148 +0,0 @@
---
topic: what runs on it
status: accepted
date: 2026-10-06
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md
---
# 225. A consumer's identity is bounded by the provision it requires, judged before merge, and never refuses its provider
## Context
[ADR 0049](0049-a-consumers-identity-fits-the-tightest-backend.md) bounds every consumer's identity,
`mesh_<machine>_<slug-or-module>`, by the tightest backend anywhere in the mesh: an S3 access key's 20
characters. It chose that over per-interface bounds (its option C) because one constant unblocked the
object store, and named C as the refinement "if non-S3 consumers are paying for S3's limit often
enough to mind". [Issue 263](../04-ISSUES/263-every-consumer-pays-for-the-tightest-backends-name-limit/00-report.md)
is that point. The operator: "the 20 character limit has bitten us multiple times".
What happened on 2026-10-06, and what it shows:
- **The bound applied where it meant nothing.** A change made the network-manager modules require the
mesh's resolver provision. The resolver mints no credential and keeps no name: its provider is not
even told who its consumers are (it receives nothing). `networkmanager`'s identity, 23 to 26
characters on real machine names, was refused all the same.
- **It was found late and far from its cause.** The catalogue's module check and the controller's
tests passed. ADR 0049 says the refusal comes at assignment; this was a new requirement on modules
already assigned, so no assignment saw it. It surfaced when the provider composed its grants.
- **It refused the wrong machine, wholly.** The controller judged the identity while composing the
*provider's* declaration, and one refusal there fails the whole composition. The anchor holds the
resolver, so the anchor — every module on it — could not be pushed until a slug was changed
elsewhere.
The catalogue's providers keep very different names. Read from each provider's code: the object store
keeps the identity as an access key (20) and a bucket name; PostgreSQL as a role and a database (63);
MongoDB as a database (63); SQL Server as a login and a database (128); the identity provider as a
client id (255); the forge as a user name (40); the mail server as a mailbox's local part (64); the
public DNS provider as a label (63); the cache, the vault, the message broker and the time-series
store as names with no limit worth stating. The resolver and both route providers keep no name of
their consumers at all.
## Considered Options
1. **Keep one bound, raise or lower it.** Any single number is wrong for most provisions: 20 refuses
a database consumer for an object store's key, 63 lets the object store fail at provision time
again ([issue 034](../04-ISSUES/034-mesh-login-exceeds-s3-access-key-limit/00-report.md)). Rejected —
it is the cause.
2. **Per-interface bounds in a mesh-wide table of provisions.** Puts the fact in the controller, which
would then know what an S3 key is. The mesh is name-agnostic about what a provider does with an
identity (`ConsumerIdentity`); a table would make it learn every backend. Rejected.
3. **Each offer states its own bound; the bound a consumer meets is that of the provision it
requires** (ADR 0049's option C, placed where [ADR 0202](0202-a-provider-declares-what-it-derives-for-each-consumer.md)
already places what a provider derives: in its own definition). Chosen.
4. **Derive a different identity per provision** (C's own "against": one consumer, several names).
Not needed: the identity stays one name, said once, and must fit every provision the module
requires. Only the *bound* is per provision. Rejected as unnecessary.
And for where the refusal lands:
5. **Keep refusing the provider's composition.** Rejected — it makes one consumer's name a reason
no push reaches a machine that did nothing wrong.
6. **Refuse the consumer's whole machine.** Right machine, still too wide, and still late. Not chosen;
the consumer is named instead, and refused in its own pull request (below).
## Decision
**1. An offer states the longest consumer identity its backend keeps.** In the provider's
definition, beside the provision: `identity` with a `max` and the words for what keeps it (`in`), so a
refusal can say "an S3 access key keeps 20"; `in` alone for a backend with no limit worth stating; or
`false` for a provision that keeps no name derived from its consumer. The bound applied to a consumer
is that of the provision it requires, from the offer of the module answering it — not the tightest
backend in the mesh.
**Unsaid, the bound follows from what the provider is told.** A provider that receives the provision,
or serves its consumers a value built from their identity, is told who each consumer is and may make
a name of it in a backend nobody measured: it keeps ADR 0049's 20. A provider told neither keeps
nothing of its consumers: no bound. A provider whose definition is not at hand is held to 20.
**2. An overflow is refused before merge.** The catalogue check (`module check`, and the controller's
test over the real catalogue) judges every module's identity, built on the longest machine name of
the mesh, against the bound of every provision it wants that a module in the catalogue offers — the
tightest where several offer it. The pull request that introduces an overflow — a new requirement, a
lowered bound, a longer module name — is the one that fails, naming the module and the longest slug
that would fit. The longest machine name is a parameter of the check; its default is the mesh's own
longest, raised in the same change that names a longer machine.
**3. One consumer's identity never refuses its provider's machine.** When the provider's declaration
is composed, a consumer whose identity overflows the bound is left out of the grants and returned
beside them; every other consumer is granted and the declaration composes. The consumer is named
where an operator looks: on the push and plan of the provider, on the plan of the consumer's own
machine, and in `status` (and its document), which does not call the mesh well while one stands.
**4. An identity served as a DNS label stays inside one.** The mesh writes an identity into a label
without truncating it ([ADR 0202](0202-a-provider-declares-what-it-derives-for-each-consumer.md)),
so an offer that serves `${consumer:as:dns}` must bound its consumers at 63 or less.
What does not change: the identity is still derived once and said to both ends (ADR 0049, issue 023);
the slug is still the remedy; truncation and hashing are still refused.
## Consequences
- A consumer of a database or the identity provider keeps a legible name on a long machine name; only
a consumer of the object store still needs a slug of a few characters. Requiring a keyless
provision costs nothing in name length.
- The catalogue's providers each state a bound in their offer, and the check reads it there. A new
provider that is told its consumers and says nothing keeps the old 20: safe, and visible in review.
- A consumer left out of the grants holds a binding and credential its provider never created. It
fails to authenticate; that is reported by name in `status` until a slug fixes it, rather than
hidden behind a machine nobody could push.
- The check's default machine-name length is a fact about one mesh carried in code. A mesh that
names a longer machine must raise it, or pass its own to `module check`; nothing reminds it to.
- Old controllers refuse the new field (manifests are parsed strictly), so the controller ships before
the catalogue that states bounds.
- ADR 0049's consequence "`identityLimit` becomes 20 … `CheckIdentity` refuses at `module add` /
assignment" now holds only for a provision that keeps 20, and is checked before merge and reported
at composition rather than refused at assignment. ADR 0049 carries a note saying so.
**How each rule is checked.**
- *Rule 1:* catalogue unit tests — an offer's stated bound is the one applied; an unstated one is 20
for a provider told its consumers and none for one told nothing; `false` is none; the field parses
strictly and a bound without `in`, or shorter than any identity, is refused when the definition is
parsed. Over the real catalogue: the resolver provision bounds nothing and the object store 20.
- *Rule 2:* `module check` refuses a module whose identity overflows what it requires, naming the
slug length, and passes a long name requiring a keyless provision; a test runs the same judgement
over every manifest in the real catalogue on the default machine-name length, so the catalogue's
own pull request fails on an overflow.
- *Rule 3:* a controller test against a real store reproduces the night it was found — the network
manager, no slug, on a six-character machine, requiring the resolver provision, beside a consumer
that overflows an object store: the provider's declaration composes, the network manager is granted,
the overflowing consumer is left out, named by the push, carried by `status` and its document, and
the mesh is not called well.
- *Rule 4:* a test that a DNS-label identity under a bound over 63 is refused when the definition is
parsed.
## References
- [Issue 263](../04-ISSUES/263-every-consumer-pays-for-the-tightest-backends-name-limit/00-report.md) —
the observation.
- [ADR 0049](0049-a-consumers-identity-fits-the-tightest-backend.md) — the bound this refines; its
option C.
- [ADR 0202](0202-a-provider-declares-what-it-derives-for-each-consumer.md) —
a provider declares what it derives; the bound is declared beside it.
- `mesh-controller` `internal/catalogue/identity.go` (`IdentityBound`, `CheckIdentityWithin`,
`IdentityProblems`, `Resolution.Overflowing`), `manifest.go` (`Offer.Identity`,
`IdentityBoundOf`), `cmd/mesh-controller/plan.go` (`grantsFor`), `check.go`, `status.go`.
- `mesh-catalog` — each provider's offer states its bound.
@@ -1,158 +0,0 @@
---
topic: the tiers
status: accepted
date: 2026-10-06
deciders: jochen
reconstructed: false
supersedes-in-part:
- 0009-modules-and-the-graph.md
extends: 02-DECISIONS/0007-connectivity.md
---
# 226. The private network is assigned by its own name, and the proxy names its public issuer
## Context
**The operator, 2026-10-06: "too many network-related modules … we simply have a private network,
some docker stuff, a firewall and a public and internal certificate resolver."** Counted on the
production mesh that day, from the controller's module listing and each machine's plan:
| module | on | what it is |
|---|---|---|
| `networking` | all four machines | requirements only: `private-network`, nothing else. Ships nothing |
| `mesh-wireguard` | on none directly; on all four through `networking` | the private network; its resources are computed by the controller |
| `public-acme` | the anchor and the home server | runs nothing; offers `acme-ca`, pointing at Let's Encrypt's production directory |
| `route-proxy` | the anchor and the home server | the public front door; the only consumer of `acme-ca` |
| `dhcpcd` | none | a third holder of `node-uplink`, for a machine whose link is dhcpcd's |
| `cloudflare-dns` | none | the only provider of `public-dns`; nothing in the catalogue requires `public-dns` |
**`networking` is the domain module [ADR 0009](0009-modules-and-the-graph.md) argued for in its
"what a domain module turns out to be" section**: a module holding nothing, so `assign networking`
finds one VPN and takes it, and choosing another is assigning that instead. It has never been more
than a second name for one module: one VPN exists, every machine runs it, and the name resolution it
once also required became a fact (ADR 0199). Its cost is real: every machine carries two modules for
one network, genesis assigns a module that is not the network, and the resolver's hint for a machine
off the network names the bundle rather than what it installs.
**`public-acme` is a provision with one consumer and one answer.** It was introduced so a lab without
a public issuer could answer `acme-ca` from `step-ca` instead (ADR 0066); `step-ca` now offers only
`internal-acme-ca`, and the mesh itself is the test bed (ADR 0149). What remained was a module
assigned beside every proxy to say "Let's Encrypt", and a pin when two answers stood on one machine
(issue 258).
**The proxy's account directory is named after the issuer as rendered.** `route-proxy` keeps each
authority's ACME account and certificates under `/var/lib/route-proxy/acme`, in a directory named by
a digest of the directory URL and of the root bundle its `trust` step copies in (the proxy's
`forThisAuthority`). The binding rendered the URL as `https://acme-v02.api.letsencrypt.org:443/directory`
— with the port — and the root as the system bundle of the pinned `trust` image. Spelled any other
way, the proxy sees a new authority, registers a new account and orders every routed name again. On
2026-10-06 that is 25 names under one registered domain on the home server, all issued on 2026-10-01,
and 26 on the anchor: a reissue of the home server's would reach Let's Encrypt's 50-per-week limit for that
domain inside the week.
## Considered Options
1. **The private network a default of every machine**, with no assignment. Rejected: "a machine is
on the private network because it was assigned the module" is what made the network a module at
all (connectivity §1, 2026-08-29); a default is a second path to the same state, and a machine
that should stay off would need an exception mechanism that does not exist.
2. **Keep `networking`, document it better.** Rejected: it answers a question — which VPN — that has
had one answer since it was written, and costs a module on every machine to do so. Choosing
another VPN is still assigning it; the bundle never added anything to that.
3. **Assign `mesh-wireguard` directly on every machine and retire the bundle.** Chosen.
4. **`route-proxy` provides `acme-ca` itself**, keeping the provision. Rejected: nothing else
consumes it, and a provision whose only consumer is its provider is a field, not an edge.
5. **The public issuer as a setting of `route-proxy`.** Rejected: a setting has no default (ADR
0112), so every mesh would need it set before the proxy resolves, and the one value that must not
change — the URL as rendered, port included — would become a value anybody can change with
`settings set`. Changing the public issuer of a live proxy is a reissue of every certificate it
holds; it should take an edit of the module and a decision, not a command.
6. **`route-proxy` states Let's Encrypt in its own `acme.env`, byte for byte what the binding
rendered, and `public-acme` retires.** Chosen.
## Decision
**1. The private network is assigned by its own name.** Every machine is assigned `mesh-wireguard`
directly. The `networking` module is retired: the controller no longer ships it, genesis assigns
`mesh-wireguard`, and the resolver's refusal for two machines that share no private network names
`mesh-wireguard`. Another VPN is still chosen by assigning it instead, and the node-scoped claim
`the-private-network` still refuses two. This supersedes ADR 0009's section *what a domain module
turns out to be*: the mechanism — a module with requirements and no files — stands, and the
catalogue no longer holds one.
**2. A module the controller stops shipping is retired at its next start, never from under a
machine.** `module forget` refuses a module the controller ships, because the next start would put
it back; so a retired one could be removed by nothing. At start the controller removes every module
it recorded as its own and no longer ships — unless a machine is still assigned it, or the mesh still
holds settings, secrets or ports for it, in which case it is kept and the start says why. This removes
a module's *definition* from the catalogue, never a consumer's data: what a consumer leaves behind is
[ADR 0230](0230-a-consumer-the-mesh-stops-asking-for-is-retired-and-deleted-only-by-a-person.md)'s, and
a module holding anything is never removed here.
**3. The proxy names its public issuer.** `route-proxy` no longer requires `acme-ca`; its `acme.env`
states Let's Encrypt's production directory exactly as the binding rendered it. `public-acme` leaves
the catalogue and the mesh. The internal issuer stays a provision (`internal-acme-ca`, from
`step-ca`): it is a module that runs, held on one machine and consumed by every proxy and by
`ca-trust`. The proxy binary's own default stays Let's Encrypt *staging*, for anything that runs it
without the module.
**4. `dhcpcd` and `cloudflare-dns` leave the catalogue.** Neither is assigned anywhere; nothing
requires `public-dns`. `node-uplink` keeps two holders (NetworkManager, systemd-networkd), and the
`public-dns` interface of [ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md) has no
provider until a mesh needs one.
`avahi` and `netcheck` are not decided here.
## Consequences
- **The rollout changes no machine's network or certificates.** Assigning `mesh-wireguard` beside
`networking` and then unassigning `networking` leaves every machine's declaration of the private
network as it was; the proxy's `acme.env` is unchanged, so its `trust` step does not run again and
its account directory is the one it has. The one file that goes is the proxy's binding to
`acme-ca`, which nothing reads.
- **The order is assign, then unassign.** Unassigning `networking` first takes `mesh-wireguard` off
every machine that does not also run `dnsmasq` (which requires the mesh's addressing) at the next
push — the private network down. The controller's retirement refuses nothing here; it only keeps
the bundle while it is assigned.
- **The public issuer moves only with a plan for the account.** The proxy's account directory depends
on the URL as spelled and on the pinned `trust` image's root bundle. Moving that image's digest
reorders every certificate on both proxies; a test holds both values and says so.
- **The lab's beds that assigned `networking`, or pointed the proxy at `step-ca` for public names,
need changing before they run again.** A lab proxy now orders public names from Let's Encrypt's
production directory, so a bed that routes public names must not assign the module as it stands.
- **Genesis from an older host binary assigns `networking`, which a newer controller refuses.** The
host's change is merged before the controller's.
- **A mesh that wants another public issuer edits `route-proxy`**, rather than assigning a different
provider. That is deliberate (option 5).
## How it is checked
| Rule | Checked by |
|---|---|
| The controller ships one network module and no bundle; it resolves on its own | mesh-controller `internal/catalogue/provided_test.go` |
| A refusal for want of the private network names `mesh-wireguard` | the same file, `TestARefusalForWantOfThePrivateNetworkNamesItsModule` |
| Another VPN is chosen by assigning it, and WireGuard is not dragged in; two VPNs collide | the same file |
| A module no longer shipped is retired at start; kept while assigned or holding anything; a registered module is never touched | mesh-controller `internal/inventory/retire_test.go`, against a live database |
| Genesis assigns `mesh-wireguard`, never `networking` | mesh-host `internal/bootstrap/network_module_test.go` |
| `route-proxy` requires no `acme-ca`; its `acme.env` is byte for byte what the binding rendered; its `trust` image is the pinned one | mesh-controller `internal/catalogue/public_issuer_test.go` against the sibling catalogue |
| `public-acme`, `dhcpcd`, `cloudflare-dns` are not in the catalogue; nothing in it requires `acme-ca` or `public-dns` | the same file |
| Live: every machine still on the private network; every public name's certificate serial and expiry unchanged; no new file in a proxy's account directory | the rollout's checks, before and after each step: the controller's network view, `wg show` through each machine's tools, `openssl s_client` against every routed public name, and a listing of `/var/lib/route-proxy/acme` |
## References
- [ADR 0009](0009-modules-and-the-graph.md) — its section on a domain module is superseded; the rest
stands.
- [ADR 0007](0007-connectivity.md), [connectivity](../03-DESIGN/01-to-be/08-connectivity.md) §1 and §5,
amended alongside.
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — the proxy requiring an issuer; the public one
is now its own fact.
- [ADR 0044](0044-a-public-name-is-provisioned-like-any-capability.md) — `public-dns`, without a
provider in the catalogue.
- [ADR 0117](0117-a-machines-uplink-is-a-seat.md) — `node-uplink`, now with two holders.
- [ADR 0112](0112-a-module-definition-names-no-node-mesh-or-path.md), [ADR 0149](0149-the-live-mesh-is-the-test-bed.md).
- [Issue 258](../04-ISSUES/258-every-machine-bound-the-resolver-to-itself/00-report.md) — the pin two
issuers on one machine needed.
- mesh-controller `internal/overlay/generator.go`, `cmd/mesh-controller/modules.go`, `stores.go`,
`internal/inventory/catalogue.go`, `internal/catalogue/resolve.go`; mesh-host
`internal/bootstrap/phase2.go`; mesh-catalog `modules/route-proxy`, and the removal of
`modules/public-acme`, `modules/dhcpcd`, `modules/cloudflare-dns`.
@@ -1,226 +0,0 @@
---
topic: the mesh
status: accepted
date: 2026-10-06
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md
---
# 227. The core holds nine rules, each checked, and is built to them in six phases
## Context
**The core is the controller, the node-engine and its launcher, the bus server, the node tools and the
console, the build seat, and the forge's announcer of merges** — everything a change passes through
before a module's own code runs. On 2026-10-06 the operator asked for it to be *"fully diagnosable, with
active monitoring, self-healing, self-upgradeable, self-monitoring … very sturdy, no ambiguities, clear
plan of execution, fail-proof setup"*, and approved [research 031](../01-RESEARCH/031-a-core-that-cannot-fail-silently/00-overview.md)'s
conclusion the same day: *"do the research and implement it"*.
The evidence is [031/01](../01-RESEARCH/031-a-core-that-cannot-fail-silently/01-evidence.md):
- **92 issue reports in six days; 48 of them core failures.** Every one of the 48 was noticed because a
person or an agent looked. **None was raised by the mesh unasked.** In four the mesh's own answer
carried the fact for whoever asked; in none did it tell anybody.
- **Four faults came back through a different door after their first fix** (200 → 265, 230 → 264,
257 → 261 → 267, 175 → 184 → 248): each fix closed an instance and left the class open.
- **The classes, by count:** races between actors with no explicit order (10 issues); commands or
arguments dropped with a default chosen in their place (8, two of them destructive —
[241](../04-ISSUES/241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md)
dropped seven databases on one unreadable file,
[244](../04-ISSUES/244-a-verb-whose-schema-is-empty-cannot-be-called-through-the-console/00-report.md)
pushed every machine when one was named); outcomes never fed back (8); failures visible only as log
lines (9, the longest twenty-three hours); state with two writers (7); repairs by hand (at least
fifteen acts, the commonest a push by hand to unstick a plan waiting on a report); self-upgrade
breaking the core (10); checks that pass in CI and fail on the mesh's real facts (6); third-party
faults nobody compared against outcomes (3, one of them
[266](../04-ISSUES/266-a-merge-on-the-bus-was-never-handed-to-the-controller/00-report.md):
23 merges unacted on over three days).
- **The parts mostly exist, one rule per message kind.** Declarations carry a sequence; reports do not.
Builds and plans are ordered; controllers have no epoch. `calls` keeps outcomes, in memory, and a
controller restart — every merge to its own repository — forgets them. ADR 0224 made one failure
kind a problem `status` reports and repairs where safe; nothing generalises it.
**Checked against GENESIS.** The mission's core value *failure must be loud — prefer failing to lying*
is the rule this record enforces on the core. The context says *human agents are few, often one, and
usually asleep; anything requiring a human to notice it will be noticed late* — the 48-of-48 count is
that sentence measured. The effect promises that *the mesh notices when something is wrong before you
do* and asks *with the context, not a log line*. Nothing in 031 conflicts with GENESIS; the effort is
the gap between the effect and the as-is, counted.
## Considered Options
1. **Keep fixing issues one at a time.** Rejected: that is what the window did, and four classes
recurred through a different door. Point fixes converge on instances, not classes, and the next
message kind or the next reader of an unreadable file starts from nothing.
2. **Monitoring from outside: a metrics stack with alert rules** (an exporter per component, a time
series store, an alert manager). Rejected: it observes symptoms the core would still keep to itself
(a report never sent is not a metric), it is a second answer to questions the controller already
answers (how-we-build §5: a report is read from the system), it adds three services to a mesh with one
operator, and it repairs nothing. The mesh's facts are in the controller; the watcher belongs there,
with one watcher of the watcher outside it.
3. **Prevention first: order and one writer before anything else.** Rejected as the *first* step, kept
as the second. Prevention covers the classes already met; detection covers every class including
those not met yet, and is cheaper per day. Phase 1 makes the mesh say when it is wrong; Phase 2
removes the largest class.
4. **Heal everything by default.** Rejected: a default of healing heals what is not understood, which is
how a repair destroys — 241's reconcile was, in its own terms, healing. Healing is narrowed to
*known* failures, ones repaired by hand twice, under a brake.
5. **A three-server bus cluster, so the bus can be upgraded live.** Not decided here: it is a change of
the foundation's shape and needs its own effort. Until then a bus upgrade is a planned, announced
step (rule 8).
6. **The lab as the place every core change is proven.** Rejected by [ADR 0149](0149-the-live-mesh-is-the-test-bed.md),
which stands: the live mesh is the test bed. Lab replays are kept for what must not be done to the
live mesh on purpose — two controllers at once, a deliberately broken controller build, a suppressed
signal — and each rule still has a live check.
7. **Nine principles as stated, the mechanisms M1–M9 and the phased roadmap of 031/03.** Chosen.
## Decision
**The nine rules below hold for the core. A change to a core repository is refused in review if it
breaks one, and each rule is checked as its row in *How it is checked* says.** They extend, and do not
replace, [research 017](../01-RESEARCH/017-a-mesh-that-heals-itself/01-the-intended-behaviour.md)'s
six for the loops that converge modules.
1. **One writer per piece of state.** Every kind of state the core keeps has exactly one writer, named
in the writers table of [to-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md).
Anyone else asks that writer. Nobody computes a second answer to a question it already answers. The
controller holds a **lease**; only the holder acts.
2. **Everything that changes state carries its writer's order, and every receiver refuses what is
older.** Declarations, reports, plans, builds, calls and announcements carry writer, epoch and
sequence. A receiver keeps the highest it accepted per writer and refuses anything older, with a line
in the mesh's words and a counter. Arrival order never decides.
3. **Every command is answered at once, and its outcome is kept where it can be read later.** Within
the verb's declared bound, with a result or a call id. The outcome is durable: it outlives the process
that ran it, and is read or waited on by id. No answer depends on what the command does to the
transport carrying it.
4. **Nothing is dropped silently.** An input that is unknown, unreadable or unmet is refused by name —
which input, where, why. A default is never substituted for *I could not tell*, above all never
"empty". A reconcile that would withdraw more than a bound of what it holds stops and raises a
condition instead.
5. **Every expected signal has a watchdog; absence is a condition.** Every signal the core expects on
a cadence or after an act has a row in the signals table: emitter, trigger, bound, condition raised.
Silence past the bound is a condition naming what was expected, from whom, since when.
6. **The mesh checks itself continuously against live facts and says what it found outward.** The
design's invariants are probes run on a schedule against the running mesh (`doctor`). A violation is
a **condition** — durable, with since-when, evidence and who can resolve it — shown in `status` and
sent to the operator through the output channel. The checker's heartbeat is watched from a machine
that is not the control node, through a channel that does not pass through it.
7. **A known failure heals itself, under a brake, and every repair is said.** A failure repaired by
hand twice gets a healer: the ordinary path again, never a destructive act, with a budget and a
back-off, and one event saying what it did. A spent budget, or a repair that could only destroy, is
a condition, not a retry. Every repair a person makes on the core goes through a verb that records
who, what and why (the **hand-act log**).
8. **The core upgrades itself one machine at a time, health-gated, and rolls back on its own.** A new
controller, node-engine or node tools build is judged on its first machine by that component's
health probes, not by "reported applied". One not healthy within its bound is rolled back to the
last known good **by something other than itself**, and the rollback is a condition. The component
being replaced is never the only witness of its successor. A bus upgrade is a planned, announced
maintenance step, never a plain rollout.
9. **A check is fed the real mesh's facts before a change is merged.** A check whose verdict depends
on the environment runs against a facts snapshot exported by the controller — every machine
composed with the change and validated by the node-engine's validator — and a dependency's version
the mesh runs is the version its tests run.
**The plan of execution is the six phases of to-be 45**, in that order, each ending at its own *done
when*:
| Phase | What it delivers | Rules |
|---|---|---|
| 0 — finish what is in flight | the located core fixes rolled out; `calls` durable; `status` inside its bound; the hand-act log; the bus's planned upgrade as the first maintenance step; the durations the bounds are measured from | 3, 7, 8 |
| 1 — the mesh says when it is wrong | conditions; watchdogs from the signals table; the bus's advisories; `doctor`; the output channel in its minimal form and the second-machine watcher | 5, 6 |
| 2 — order and one writer | the controller's lease and epoch; a report's sequence; one apply queue on every machine; stale refusals counted; the writers table enforced; the empty-on-error lint; the withdrawal brake | 1, 2, 4 |
| 3 — healers | the first healers, each braked; a repeated hand act asks for one | 7 |
| 4 — core upgrades that roll back | a health definition per core component; the gate; rollback by a witness; the bus as a planned step | 8 |
| 5 — checks before merge, and replays | the facts snapshot and the compose-and-validate merge gate; versions tested as run; every core incident a replay | 9, and all |
**The output channel is built in the smallest form research 028 allows.** One mesh seat,
`operator-channel`, accepting `notify` as a work queue and holding its open messages in its own state
([028 Q1](../01-RESEARCH/028-the-meshs-output-channel/03-open-questions.md), option a); the controller
emits condition events and the holder decides what is sent (028 Q5, option a); the Telegram channel the
operator required and the desktop notifier as the two channels; deduplicated by the condition's key; no
answering back except through the mesh's own verbs; a message carries roles and words, never an
address, a path or a secret, refused by the holder otherwise (028 Q8). Routing by presence, quiet hours,
answering back and the external dead-man service stay open in 028, whose graduation amends to-be 45.
> **The mechanism changed — 2026-10-06, by [ADR 0234](0234-the-mesh-holds-a-conversation-with-its-operator.md).**
> What still stands: the output channel is built first in this minimal form, as Phase 1 of to-be 45, and
> everything above about conditions, deduplication, the content rule and the watcher. What moved, as
> this paragraph said it would on 028's graduation: "no answering back" no longer holds. The operator
> answers and is asked over channels that are holders of two kinded benches, `channel` and `intake`,
> not contributions to `operator-channel`, whose holder becomes the router; an answer that performs an
> action is checked and performed by the controller. The design is
> [to-be 46](../03-DESIGN/01-to-be/46-the-conversation-with-the-operator.md).
**ADR 0224's provider standing becomes the first condition kind**, unchanged in what it says and
when; its storage moves into the condition store.
## Consequences
- **`status` changes meaning.** It becomes, first, the list of open conditions; the all-well sentence
is "no open conditions", silenced ones included. A silenced condition is still open; silence stops
only its messages, for a stated time, with a reason, recorded as a hand act. Nobody resolves a
condition by hand: it clears when observation says so.
- **The bounds are measured, not guessed.** The signals table's first bounds are provisional; Phase 0
records the durations they are set from, and Phase 1's first live week corrects every bound that
raised a condition that was not real. A corrected bound is a change to the table, reviewed like code.
- **The controller's restart stops being a forgetting.** Calls, conditions, the hand-act log and the
lease live in key-value buckets ([ADR 0201](0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md));
a new holder of the lease marks a call the old one left running as abandoned, which is said.
- **The node-engine gains an order it did not have.** One apply queue replaces the delivery path and
the five-minute reconcile as two appliers; a report carries the declaration's sequence; a declaration
from an older controller epoch is refused. Issues 257, 261 and 267 become impossible rather than
handled. The controller–node-engine wire changes, so the rollout order is the node-engine first
(it accepts both shapes), the controller second.
- **More is checked before merge, and merges get slower.** The compose-and-validate gate runs every
machine through the validator on every core and catalogue merge. That is minutes, against the hours
each of 202, 236 and 263 cost.
- **A third party now carries operational words.** Telegram is not end-to-end encrypted for bots, so the
holder's content rule is the only thing between a condition's evidence and someone else's server. It
is enforced by the holder and tested there, not trusted to each source.
- **The watcher outside the control node holds the channel's secret.** That is one more machine with a
bot token, accepted for the one case nothing else covers: the control node or the bus is what failed.
- **The roadmap is six to eight weeks of focused work.** The first visible change — the mesh saying
when it is wrong — is inside the first two. Until Phase 2 lands, races are still caught by watchdogs,
not prevented.
- **The rules are not yet in how-we-build.** They are this record's until they are carried into
how-we-build §2 through playbook 05 (constitution sync), which is a separate change.
## How it is checked
| Rule | Checked by | From |
|---|---|---|
| 1. one writer | the writers table in to-be 45; a test per core repository that the code paths writing each kind are the named writer's; the controller refusing, at composition, a bus subject two components may publish on unless the table says it is shared; live, the `doctor` probe *exactly one lease holder, no stale-epoch message in the last interval* | Phase 2 |
| 2. order | a contract test per consumed message kind in the receiver's repository (deliver *n*, then *n−1*: refused and counted; epoch *e−1* after *e*: refused); a check listing every consumed subject against the tests that name it; live, the stale-refusals signal | Phase 2 |
| 3. answered, kept | a test walking every verb the controller announces (answers within its bound, with a result or an id); a test restarting the controller between a call and the read of its outcome; live, the `doctor` probe *status answers in full within ten seconds* and the *call hung* signal | Phase 0 |
| 4. nothing dropped | per component, a test feeding each input reader an unreadable, malformed and foreign input (a refusal, never an empty result); a lint in each core repository refusing an error branch that returns an empty collection; schema walks over every verb and placeholder namespace; a test emptying a provider's input and asserting nothing is withdrawn | Phase 2 |
| 5. watchdogs | a test generated from the signals table that suppresses each signal in turn and asserts its condition is raised within its bound and cleared when it returns; live, `doctor` reporting the age of the newest signal of every row | Phase 1 |
| 6. self-check, outward | a check over the to-be designs counting invariants with a live probe against those without, which may only go down; the second-machine watcher raising *self-check silent* when the controller is stopped; once, on a lab mesh, a broken invariant appearing in `status` and as a message within one probe interval | Phase 1 |
| 7. healers | a test per healer inducing its failure, asserting the repair, the event and the brake after the budget; the hand-act log's weekly count in `status`; a cause recorded twice raising *healer wanted* | Phase 3 |
| 8. staged upgrades | on a lab mesh, a broken build of the controller, the node-engine and the node tools (one that starts and does nothing, one that crashes, one that cannot reach the bus) each rolled back with no hand, ending on the previous build, said as a condition and a message; live, every core rollout's record (first machine, verdict, time to verdict, rolled back or not) readable through `plans` | Phase 4 |
| 9. real facts | the merge gate in mesh-controller, mesh-host and mesh-catalog failing a change that makes any machine of the snapshot fail to compose or validate, naming the machine's role and the module; a test that a pinned dependency's version equals the version in the snapshot; the replays of 236, 262, 263 and 266 failing on the commit before their fix | Phase 5 |
| the plan | each phase's *done when* in to-be 45, recorded in that design when met, with the design's status following | each phase |
## References
- [Research 031](../01-RESEARCH/031-a-core-that-cannot-fail-silently/00-overview.md) — the evidence,
the principles with the candidates not kept, the mechanisms and the roadmap.
- [To-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) — the design this record
authorises.
- [Research 017](../01-RESEARCH/017-a-mesh-that-heals-itself/00-overview.md) — the condition, and
the six principles for the loops; [research 028](../01-RESEARCH/028-the-meshs-output-channel/00-overview.md)
— the output channel, of which the minimal form is taken here.
- [ADR 0224](0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md) —
*detected automatically, repaired where safe, loud where not*, generalised here.
- [ADR 0141](0141-the-host-delivers-its-own-successor.md), [ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md),
[ADR 0218](0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md),
[ADR 0149](0149-the-live-mesh-is-the-test-bed.md), [ADR 0201](0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md).
- Issues [187](../04-ISSUES/187-the-mesh-tells-nobody-when-it-stops-working/00-report.md),
[204](../04-ISSUES/204-a-controller-handover-re-sent-every-node-a-stale-declaration/00-report.md),
[241](../04-ISSUES/241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md),
[248](../04-ISSUES/248-the-controllers-event-consumer-replayed-a-week-and-held-every-merge-behind-it/00-report.md),
[264](../04-ISSUES/264-a-self-updating-engine-lost-the-report-of-the-apply-that-delivered-it/00-report.md),
[265](../04-ISSUES/265-a-push-outlived-its-caller-and-its-answer-was-refused/00-report.md),
[266](../04-ISSUES/266-a-merge-on-the-bus-was-never-handed-to-the-controller/00-report.md),
[267](../04-ISSUES/267-a-reconciles-report-overtook-the-apply-that-followed-it/00-report.md).
@@ -1,143 +0,0 @@
---
topic: what runs on it
status: accepted
date: 2026-10-06
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0113-the-vault-makes-every-secret.md
---
# 228. A value given by hand lives only until its module's first good start
## Context
**Two of a module's own secrets leaked into logs on 2026-10-06**, and the operator approved replacing
both. Each module's definition says it reads the secret when it starts (`taken: at-start`), which is
the form the mesh rotates by making a new value and starting the module again
([ADR 0114](0114-a-shared-credential-rotates-over-two-credentials.md),
[to-be 13](../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md)). The controller refused both:
> not rotated: … holds "…" as a value given to the mesh, not made by it, and the mesh will not replace
> what it cannot read (ADR 0113). Change it where it lives, then `secret accept …` with the new value
Both values had been given by hand long before, when the modules were moved from an older setup, and
nothing outside the mesh uses either. The way out the refusal names has a person or an agent make a
value and feed it to `secret accept` — a secret passing through hands, which is the thing the mesh
exists to avoid.
**The refusal reads [ADR 0113](0113-the-vault-makes-every-secret.md) wider than it decides.** 0113 says
*"A delivered value the vault cannot replace, such as an external API key, is not rotated by the vault:
rotating it means an operator delivering a new one"*. What the vault cannot replace is a value only an
outside party can issue — a vendor's key, a bot's token, a licence — because no value of the mesh's
would work in its place. A secret the module reads at start and nobody else holds is not that: the old
value is not needed to replace it, because the module is the only reader and it reads the new one when
it starts. The controller applied the rule to **every** value it had been given, because the store
records only *made* or *accepted*, and nothing in a module's definition says who issued a value.
The operator, the same day: *"Our mesh should most definitely be able to rotate 'custom provided'
passwords — in fact, ideally we immediately rotate them after assigning and running the module for the
first time so the 'custom pwd' is gone"*, and *"a custom pwd is only useful in case we're adopting an
existing running container into our mesh."*
Counted in the catalogue on 2026-10-06: 48 own secrets across 32 modules; 6 say `taken: at-start`, none
says `applied` and 42 say neither. Of the 6 at-start ones, one is an outside party's key (an OpenAI
key). Of the rest, at least ten are keys or tokens an outside party issues.
## Considered Options
1. **Keep the refusal; add a separate verb that turns a given value into a made one** (`secret mint`,
with a reason). Rejected: it keeps a step whose only effect is to say "yes, really" to a rotation
the definition already permits, and it leaves every given value in force until a person remembers
to take it. The operator asked for the opposite default.
2. **Rotate a given at-start value like a made one, and stop there.** Better, and still leaves the
given value in force indefinitely after an adoption — the value a person handled stays the live one
until somebody asks.
3. **Rotate it like a made one, and replace a newly given value on its own once the module has started
on it**, with the exception said where it belongs: in the module's definition, for a value an
outside party issues. Chosen.
4. **Refuse `secret accept` for a secret the mesh may make, except on an adopted machine.** Rejected:
the mesh cannot tell a fresh install from one carrying data in from elsewhere on a converged machine
— a restored volume holds the password it was made with — and the replacement after the first good
start already bounds what a needless given value costs. It is accepted and said.
## Decision
**A module's own secret says who may make its value.** By default the mesh may. An entry says
`"issued-by": "outside"` when only a party outside the mesh can issue the value — a vendor's API key, a
bot's token, a licence. The parser refuses any other word.
**A secret the mesh may make** — read at start, and not issued outside — **is the mesh's to replace,
whoever gave the value it holds:**
- **`secret rotate` replaces a given value as it replaces a made one**: made anew, sealed to the
machine and the operator, recorded as made, and the machine sent so the module starts on it. A
rotation may say why, and the why is recorded in the hand-act log.
- **A value given by hand lives only until the module's first good start under the mesh.** Given
through `secret accept`, it is marked; when the machine's clean report arrives — every resource
applied, nothing failed or refused — **for the declaration it was last sent, sent after the value
was given**, the controller replaces the value with one it makes, sends the machine, says so in its
log and states `secret-replaced` as the controller seat's fact, never the value. On an adopted
machine the module must also be taken: until then the mesh runs nothing of it. The mark is cleared
as the value is replaced, so it happens once.
- **A given value exists to adopt something already running that holds it.** A module the mesh
installs fresh needs none; `secret accept` for such a secret is accepted, says that the value lives
until the first good start, and says on a converged machine that a fresh install needs no value.
**What stays as given, refused with the reason:**
- a value issued outside the mesh — never replaced; `secret rotate` names the issuer as the one to ask
and says how old the given value is, and `secret accept` delivers the new one;
- a value the module **applies** to a backend that takes it once, until the staged rotation of
[ADR 0114](0114-a-shared-credential-rotates-over-two-credentials.md) exists;
- a value the mesh's own code accepted — a bus account it issued is its word to a broker, and is never
marked;
- a value given **before** this record, which is not marked: the mesh does not decide for a person
that something given long ago is used nowhere else. `secret rotate` replaces it when asked.
**What this does not change.** A pair credential an operator delivers is still never replaced by a
made one ([ADR 0092](0092-an-operator-delivers-a-pair-credential.md)): it is held at both ends of a
provision, and this record is about a secret one module holds. A secret whose definition says neither
`at-start` nor `applied` is still not rotated. That the vault makes every shared secret, and its custody,
stand as 0113 decided.
## Consequences
- Rotating a leaked secret a module reads at start is one call, whoever gave it, with no value in
anybody's hands.
- An adoption leaves no hand-given value in force once the module runs under the mesh. A person who
needs the new value — a password typed at a login page — recovers it with the operator key, as for
any value the mesh makes.
- **The catalogue must mark every secret an outside party issues.** An at-start secret left unmarked
is one the mesh will replace with a random value. The two OpenAI keys in the catalogue are marked
with this change; the other outside keys say neither `at-start` nor `applied`, so they are not
rotated either way, and marking them is tidying, not a prerequisite.
- **The controller ships before the catalogue marks anything.** The parser refuses a field it does not
know, so a definition saying `issued-by` is refused by a controller older than this record.
- What got harder: the controller now acts on a machine's report by itself, sending it once more after
the first good start of a module given a value. A replacement that cannot be sent is kept sealed and
carried by the next push, and said.
## How it is checked
| Rule | Checked by |
|---|---|
| A given value read at start rotates like a made one | Controller inventory test: a value accepted for an at-start secret rotates, and is recorded as made. |
| A given value is replaced once, after the first good start on it | Controller inventory test: a report of a declaration sent before the value was given replaces nothing; a report of an older declaration replaces nothing; the clean report of the declaration sent after it replaces it; the same report again, and the next declaration's report, replace nothing. |
| An outside party's value is never replaced | Controller inventory tests: a value accepted for a secret marked `issued-by: outside` is not marked and not replaced after a start; `rotate` refuses it naming the issuer and the value's age; a definition changed to say `outside` after the value was given keeps it. |
| An applied value, and one the mesh's own code accepted, stay as given | Controller inventory test: neither is marked or replaced after a start; `rotate` of an applied value is refused as not stageable. |
| An adopted machine waits for the take | Controller inventory test: a module held as found keeps its given value through a clean report; after the take, the next clean report replaces it. A module no longer assigned is never replaced. |
| A good start is a clean account of a declaration | Controller test: a refusal, a failure, or a bare word that the machine is there is not one. |
| The definition says who issues a value | Catalogue test: `issued-by` is read, written back as read, and any word but `outside` is refused. |
| The fact is stated, and permitted | Broker test: the controller's grant and its seat's emits both name `secret-replaced`, and nothing else is added. |
| A rotation through the console carries why | Controller test: the `rotate` verb passes why and cause to `secret rotate`; why beside a provision is refused as passed over. |
## References
- [ADR 0113](0113-the-vault-makes-every-secret.md): the decision this extends — what the vault cannot
replace is what an outside party issued
- [ADR 0114](0114-a-shared-credential-rotates-over-two-credentials.md): read at start and applied, and
the staged rotation an applied secret waits for
- [ADR 0092](0092-an-operator-delivers-a-pair-credential.md): a delivered pair credential, unchanged
- [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md): an adopted machine, and the take
- [To-be 13](../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md): the design this amends
- [To-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) §7: the hand-act log a rotation's why is recorded in
@@ -1,177 +0,0 @@
---
topic: the mesh
status: accepted
date: 2026-10-06
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md
---
# 229. The core's order is a lease the store remembers, and an epoch a machine is sent once it reads one
## Context
Phase 2 of [to-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) — *order and one
writer* — was built on 2026-10-06 in the controller, the node-engine, the providers' loop and the
SDK's. The design states the lease, the epoch, the report's order and the brake in a paragraph each;
building them met questions the paragraphs do not answer, and each was answered in code. This record
is those answers, so the design can say them and the next build does not answer them again.
Four of them could not be left to taste:
- **The node-engine decodes a declaration strictly** and refuses a key it does not know, whole. An
epoch sent to every machine from the first controller that has one is refused by every machine
not yet updated — and a refused declaration is also the one that would have updated its
node-engine, so a machine asleep through the rollout could never catch up.
- **The epoch is a bucket revision, and a bucket can be raised again from nothing.** A bus whose data
directory was replaced starts its revisions at one, and every machine that heard epoch 57 would
refuse the next controller for ever.
- **A command at a shell sends declarations too** — the installer's first pushes, a lab step, a
person repairing a mesh whose controller is down ([issue 201](../04-ISSUES/201-a-push-recreated-the-controller-behind-the-row-its-successor-wrote/00-report.md)).
"Only the holder acts" read strictly would stop the installer at its first push.
- **A controller's grant on the bus is composed by the controller, and the live bus holds the list the
previous build composed.** A controller that needs a grant for its lease's bucket before it may act
cannot send the list that grants it.
## Considered Options
1. **Send the epoch to every machine and roll the node-engine out first**, as ADR 0227's consequences
suggest. Rejected: it relies on every machine being up for the node-engine's rollout, and strands
one that is not.
2. **The epoch outside the signed bytes**, where a strict decoder does not look. Rejected: a broker
could then give an old declaration a new epoch, and the order of declarations is the one property
their signature does not already protect.
3. **The epoch inside the signed bytes, sent to a machine only once its node-engine has said it reads
one.** Chosen.
4. **A one-off command takes no part in the lease**, and is refused while a controller serves.
Rejected for the installer above.
5. **A controller that cannot take the lease refuses to act.** Rejected for the grant above: it could
never send the user list that grants it.
## Decision
**The lease.** The controller's lease is the key `holder` in the bucket `mesh-controller_lease`, whose
age is fifteen seconds; the holder renews it every five by compare-and-set at the revision it last
wrote. Its value names the instance (machine, process, start), its build, its epoch, and when it was
taken and renewed. **A holder stops acting three seconds before its key could expire unrenewed**, by
its own clock, whatever its renewing goroutine is doing; a renewal refused or failed is the lease lost,
said, the instance's epoch recorded as lost, and the process exits to be started again as a candidate.
A holder that stops gives the key back, so the next takes it at once. The serving controller takes the
lease before it asserts the bus's objects, and a candidate waits, said once per holder.
**The epoch never goes backwards.** Every epoch issued is kept in the controller's store with its
instance, when it was taken and how it ended: given back (`released`), lost by its own renewal
(`lost`), or found gone by the next holder (`expired`). The highest is a floor: a lease bucket whose
revisions are at or under it was raised again from nothing, and its stream is compacted past the floor
before the key is taken — said as a condition, S12.
**A controller the bus refuses the lease, with nobody holding it, serves without one**, as every
controller did before: its declarations carry no epoch, which no node-engine refuses; it says so as
the urgent condition S12 names, and tries again every five seconds. The first push of the machine
holding the bus sends the user list that grants it. One that finds another took the lease meanwhile
stops.
**A command run at a shell** acts under the holder's epoch, read at the moment it acts, while a
controller holds the lease — it is the same mesh's word, composed under the same hold of each machine —
and under a lease of its own, given back as it ends, while none does. A process with no bus configured
claims no epoch.
**What carries the order on the wire** — the contract, written once on each side (the controller's
`internal/link/order.go`, the node-engine's `internal/link/messages.go`):
| Where | Key | Holds |
|---|---|---|
| a declaration, inside the signed envelope | `epoch` | the lease epoch it was composed under; absent claims none |
| | `sequence` | its number for that machine (issue 107); absent claims none |
| a report | `epoch`, `sequence` | the order of the declaration the report is about |
| | `report_sequence` | the node-engine's own number for the report, kept on disk, growing across restarts and self-updates; **its presence says the node-engine reads `epoch`** |
| | `older_than` | on a refusal of a declaration older than one applied: the order of the one held |
| | `refused_older` | how many declarations the node-engine has refused as older, ever, on every report |
- **A machine is sent `epoch` only while its latest account carried a `report_sequence`.** One that
stops carrying it — a node-engine rolled back — is sent none again.
- **A node-engine refuses by epoch only when both declarations claim one**: an older epoch, or the
same epoch and a lower sequence. A declaration with no epoch is taken by its sequence, so a
controller rolled back to a build without the lease is never stranded; that gives up the epoch's
protection for as long as such a build runs.
- **The controller keeps the account of each machine by its order**: by epoch where both claim one,
then by sequence, then by report sequence, and refuses an older account — said, counted, and the
machine's facts in it kept as before. An account without a report sequence, from an older
node-engine, is judged by the digest it names (issue 267), and clears the order kept.
- **What the mesh would send a machine is composed with the epoch it was last sent**, as with its
sequence: a new holder of the lease is not a change of the machine.
**Stale refusals name their writer** (S13): a declaration refused as older is counted against the
controller epoch it claimed, named from the store's record of epochs with how that epoch ended; an
account the controller refused is counted against the machine's node-engine; and what a machine's
`refused_older` rose by beyond the refusals heard is counted as refusals whose reports were lost.
**One writer is enforced where grants are composed.** The writers table is compiled into the
controller with, for each state the bus carries, the subjects a write of it publishes to and who its
writer is; a principal whose grant overlaps another writer's subject is refused at composition, naming
the state and its writer. The controller's own grant loses `mesh.control.>`, which it never published
and which made it a second writer of every machine's report. The stream-definition row names stream
creation, update and deletion and durable consumer creation — not every consumer creation, because a
module watching its own bucket makes an ordered consumer that defines nothing the mesh keeps.
**A plan is written by compare-and-set** on a revision the store keeps beside it, and carries the epoch
that wrote it; a write against a plan moved since is refused, and nothing is written by a process that
may not act.
**The empty-on-error lint** is a test over the repository's own source: a `return` inside an
`if err != nil` branch that answers an empty collection and no error fails it, unless a comment
`empty-on-error: <why>` on the line or above says why empty is the truth there.
**Every consumed message kind has a contract**: how an older one is refused, with the tests that deliver
the newer and then the older, or why none is needed. A check fails a kind the controller can be handed
without one, and one naming a test that does not exist.
**The withdrawal brake** (the providers' loop, Go and the SDK's alike): a pass that would withdraw more
than one consumer, or more than half of those held where it holds more than one, withdraws nothing;
each consumer it keeps is announced `provisioner.failing` with the class `withdrawal-braked`, so the
controller raises it as a condition (ADR 0224); and while the mesh goes on not asking for them, one is
let go every hour, said. A consumer asked for again is kept. **The brake stops the bulk, not the
intention**: an unassignment of many completes without a hand, a mistaken one costs at most one consumer
an hour while the operator is told — and withdrawal destroys no data since issue 241.
> **Replaced in part — 2026-10-06, by [ADR 0230](0230-a-consumer-the-mesh-stops-asking-for-is-retired-and-deleted-only-by-a-person.md).**
> The paragraph above no longer stands; the operator decided against releasing one consumer an hour.
> A consumer the mesh stops asking for is now *retired* — disabled, reversibly, and marked to delete —
> once the same result holds for five passes and ten minutes; a set of more than three, or more than
> half of those held, waits for a person's `retire approve`; and only `cleanup delete` deletes. Every
> other decision in this record stands. The *how it is checked* row for the brake is replaced by 0230's.
**What Phase 1 decided while building** — [issue 270](../04-ISSUES/270-phase-1-watches-signals-that-come-later-and-hears-only-the-controllers-faults/00-report.md)'s
decisions 1 to 6: the condition history as a bucket of its own, the events' shape, a key's last token,
the probe DW, which deleted consumers are said, and the bounds the design left to the build — is taken
into to-be 45 as built, and S12 and D5 move to Phase 2, S14 to Phase 5, as that issue asked.
## Consequences
- **The rollout is the node-engine first, then the controller**, as ADR 0227 says — but the order no
longer has to be exact. A machine on an older node-engine is sent no epoch; one that updates later
is sent it from its first ordered report.
- **The first controller with a lease on the live mesh serves without it** until the user list that
grants its bucket reaches the bus, and S12 says so, urgent. A push of the machine holding the bus
ends it.
- **Two controllers at once are now the lease's to settle**, not the consumers' binding (issue 213's
standing by), which stays as a second guard.
- **The TypeScript providers' brake is said in their journal only**: the SDK's loop does not announce a
provider's standing at all (ADR 0224 is in the Go loop), so a braked withdrawal there is not a
condition until it does.
- **The lint runs in the controller's repository.** The node-engine's and the node tools' carry their
own copy when they take it up; until then they are not covered.
## How it is checked
| What | Checked by |
|---|---|
| the lease: one holder, a waiting second, a handover at a higher epoch, a loss acting on nothing | `internal/lease` tests and the controller's `TestTwoControllersOneActs`, against a real bus and store |
| an epoch never issued twice | `TestAnEpochIsNeverIssuedTwiceWhenTheBucketStartsOver`; live, D5 and S12 |
| the contract | `internal/link/order_test.go`, beside the node-engine's own tests of the same cases |
| an account kept by its order | the report's contract tests (`heard_order_test.go`); live, S13 |
| one writer at composition | `internal/broker/writers_test.go`: the table is the design's, a whole mesh composes, a second writer is refused |
| a plan by compare-and-set | `TestAPlanIsWrittenByCompareAndSetUnderTheLease` |
| empty on error | `internal/lint`, over the whole repository on every test run |
| every consumed kind | `TestEveryConsumedKindHasAContract` |
| the brake | the providers' `brake_test.go` (issue 241 replayed with seven consumers) and the SDK's test |
@@ -1,219 +0,0 @@
---
topic: the mesh
status: accepted
date: 2026-10-06
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0229-the-cores-order-is-a-lease-the-store-remembers-and-an-epoch-a-machine-is-sent-once-it-reads-one.md
---
# 230. A consumer the mesh stops asking for is retired, and deleted only by a person
## Context
**This is the operator's decision, taken on 2026-10-06**, the day the withdrawal brake of
[ADR 0229](0229-the-cores-order-is-a-lease-the-store-remembers-and-an-epoch-a-machine-is-sent-once-it-reads-one.md)
was built. The model below — the three states, the stable result, the threshold that hands the act to a
person, the cleanup verbs, the thirty-day reminder — is the operator's direction, recorded here as given.
What the record adds is the evidence it rests on and the facts the build had to settle.
**What the brake did.** A provider's loop ([ADR 0224](0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md),
in the Go providers and the SDK's) that would withdraw more than one consumer in a pass, or more than
half of those it held, withdrew nothing and announced each kept consumer as failing; then, while the mesh
went on not asking, it let one go every hour, with no person involved. So a mistaken unassignment of
seven consumers still completed by itself in seven hours, and a person who did not read `status` in that
time lost all seven.
**What withdrawal does today**, since [issue 241](../04-ISSUES/241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md):
nothing a provider withdraws is destroyed in the providers that issue names — the relational database
locks the role and keeps the database; the SQL Server provider disables the login; the document store
strips the user's roles; the object store revokes the key and keeps the bucket; the mail server disables
the mailbox; the forge prohibits the login; the analytics site is kept. But three things were still
wrong:
- **The identity provider deleted the client.** Its withdrawal was a `DELETE` of the consumer's client,
outside issue 241's list; the secret, redirects and mappers went with it.
- **Nothing recorded that a withdrawal happened, or when.** A locked role is indistinguishable from one
an operator locked by hand. On the home server one consumer's login has been locked by an earlier
withdrawal, its database kept, with no record of when or why; on the control node six databases set
aside during issue 241's recovery sit under a dated name, which is the only record of them.
- **A restart forgot.** The loop withdrew only what *this process* had made. A consumer unassigned while
its provider was down was never withdrawn at all — its login stayed open, indefinitely.
**And nothing ever deleted anything**, so withdrawn data accumulates with no surface that says it is
there.
## Considered Options
1. **Keep the hourly release.** Rejected by the operator: it lets the mesh finish, unattended, an act the
bound exists to question. Many changes at once means a person is at work, and that person can confirm
once.
2. **Delete after a grace period.** Rejected: a timer is the mesh acting alone again, only later; the
loss in issue 241 was data nobody had decided to lose.
3. **Withdraw at once, below the bound, as before.** Rejected: one pass is one read of one file, and
issue 241 was one bad read. A result has to hold before it is acted on.
4. **Three states — active, retired, deleted — where retiring is reversible and automatic only when
stable and small, and deleting is always a person's act through the controller, executed by the
provider.** Chosen.
## Decision
**1. A consumer is ACTIVE, RETIRED or DELETED.**
- **Retired**: the mesh stopped asking for it. The provider **disables its access** — reversibly — and
**marks its login and data "to delete", with when and why**. Nothing is deleted. Asked for again, the
provider's ordinary create re-enables it at once, as it was, and clears the mark. A provider that
has no non-destructive way to disable does **not** fall back on its removal: it is **mark-only** —
the consumer keeps its access, is marked retired and needing deletion, and is said and announced (§2).
- **Deleted**: only through `cleanup delete`, after a person approved it.
**2. What "disable" means, per provider kind** — each the smallest reversible switch that stops the
consumer reaching its data, keeping everything else; or, where a provider has none, nothing at all:
| Provider | Retired is | Kept | Mark |
|---|---|---|---|
| postgres | the role set `NOLOGIN`, its open sessions ended | the database, its owner and grants, every byte; `CONNECT` is not revoked and the database still accepts connections, so the nightly dump still backs it up | the role's comment, a JSON object naming the consumer's machine, when and why |
| keycloak, the identity provider | the client `enabled: false` — the server refuses its authorization and token requests | its secret, redirects, mappers and every other setting | the client's attributes `mesh.retired` and `mesh.retired-why` |
| a TypeScript provider whose adapter has `retire` (and, if its create does not undo it, `reenable`) | what its `retire` disables | what its `retire` keeps | what its `retire` writes |
| a TypeScript provider without `retire` — **mark-only** | **nothing changes in the backend: the consumer keeps its access** | everything | the loop's own record, "retired, access kept, needs deletion", said loudly, announced, shown by `cleanup list` |
**`remove` is never called to retire.** A TypeScript adapter's `remove` is what it did on withdrawal, and
several destroy something a person has not decided to lose — the secrets vault unlinks the ledger file
holding the secret, the public DNS provider deletes the record, the message brokers delete the user. So
an adapter without a non-destructive `retire` is mark-only, and `remove` is reached only as the delete of
a provider without one of its own, through a person's `cleanup delete`. Today **every TypeScript provider
is mark-only**, because none has a `retire` yet: redis, mosquitto, minio, umami, mssql, mongodb,
influxdb, the secrets vault (mesh-vault), mailu, gitea and cloudflare-dns. Each leaves mark-only by
adding a `retire` that disables — the SQL Server, document-store, object-store, mail and forge providers
already have such a switch in their `remove` (disable the login, strip the roles, revoke the key,
disable the mailbox, prohibit the login), moved into `retire`.
A mark-only retirement is kept in the provider's process, not its backend: a restart forgets it, and the
consumer, still reachable, is no longer listed. That is the cost of a provider that cannot disable, and
the reason to give each a `retire`.
A database renamed aside — by postgres's `postgres_retire_database` tool or by hand, the
`<name>_deleted_<date>` form — is listed beside the retired consumers, retired since that date, so a
person sees it and can delete it.
**3. Stable removals.** A provider retires a consumer only after it has seen **the same set** of
consumers no longer asked for **both** in **five consecutive passes** that read the contributions file
**and** for at least **ten minutes** since the first of them (the constant `StableFor` in the Go loop,
`STABLE_FOR_MS` in the SDK's). At the loop's pass interval of five seconds, five passes alone are
twenty-five seconds — shorter than a controller restart, a store reconnecting or a file half written —
so the ten minutes are what a real hiccup has to outlast, and the five passes keep one slow pass from
counting as agreement. A pass that could not read the file is not a result and starts both again; so does
a different set. **Additions and changes to an existing consumer — create, rotate — act on the first pass
and are never delayed.**
**4. Too many is a person.** A stable set of **more than three consumers, or of more than half of those
the provider holds where it holds more than one**, retires nothing. The provider **waits**: it says so in
its journal and announces it, the controller raises an **urgent** condition listing what would be
retired, and it waits for `retire approve <node> <provider-module>` or `retire reject <node>
<provider-module>`. **An unassignment the controller itself made goes through the same threshold** — the
operator's words: *too many changes means a human is actively working on it*, so one confirmation. The
bound is the old brake's half, with three in place of one: one, two or three consumers of a larger
provider go without a hand; two of two do not.
- **Approve** retires exactly the set waiting — the provider refuses any other set, so a person approves
what they were shown.
- **Reject** keeps the set active; the provider does not ask about it again while the set stays the same.
The rejection is held by the provider's process: a restart asks again, loudly. A rejected set can still
be approved later.
- **Settled**: a waiting or rejected set ends without retiring when the mesh asks for any of it again or
the set changes; a new stable set starts the count over.
**5. The backend remembers, not the process.** Each provider lists from its own backend what the mesh
made, active and retired, with the mark. So a restart forgets nothing; a consumer the backend holds
active and the mesh no longer asks for is retired by the same rules after a restart as before one; and a
consumer **found disabled without the mark** — withdrawn before this record — is **adopted** as retired:
marked, said, announced, its clock starting then. What the mesh made is known by the mark, or, before
the mark existed, by its shape (postgres: a role valid until *infinity*, which only the
mesh's role statement sets, owning a database of its own name; keycloak: the
`mesh.provisioned` attribute it already had). Anything else is somebody else's and is never listed,
retired or deleted.
**6. Cleanup is the controller's verbs, executed by the provider.**
- `cleanup list` — every retired consumer per provider: its age, its size where the backend can say,
and why.
- `cleanup delete <node> <provider-module> <consumer> --why` — one.
- `cleanup delete --older-than <days> --why` — lists what it would delete; deletes only with the
operator's explicit `--confirm`.
- **The provider deletes**, because it owns its backend: the controller asks the provider's tool on that
machine through the mesh, as any module's tool is asked, and never touches a backend itself. A provider
refuses to delete anything active or asked for, and anything not retired.
- **Every deletion is written to the hand-act log**, as is every approval and rejection — including one
made by asking a provider's tool directly rather than through the controller. They are a person's decision
by design, not a repair, so they never count toward the hand-act log's *healer wanted* (to-be 45 S15):
a healer may not withdraw or delete data.
**7. Every provider serves the same four tools**, the protocol between the verbs and the providers:
`provisioner_retirement` (what is held, waiting, rejected and retired), `provisioner_retire_approve`,
`provisioner_retire_reject` and `provisioner_delete`, each act requiring a why. And says one event,
`provisioner.retirement`, whose `change` is `waiting`, `settled`, `approved`, `rejected`, `retired`,
`reenabled`, `deleted` or `adopted`; the controller derives the permission to publish it for every module
that receives contributions, as ADR 0224 does the standing events.
**8. Nothing is silent.** Retire, approve, reject, re-enable and delete are each announced and said in the
provider's journal; waiting is an urgent condition and a rejection a warning one, both naming the
provider and its machine; and a self-check probe raises a **low-severity `cleanup-waiting` condition**
for any provider holding something retired **more than thirty days**.
**9. The withdrawal brake paragraph of ADR 0229 is replaced by this record.** The rest of 0229 — the
lease, the epoch, the report's order, one writer at composition, the plan by compare-and-set, the lint,
every consumed kind's contract — stands unchanged.
## Consequences
- **The first rollout surfaces what is already there.** On the home server the consumer locked by an
earlier withdrawal is adopted as retired; the control node's six set-aside databases are listed, and
raise `cleanup-waiting` thirty days after their date; any consumer a backend holds open that the mesh no
longer asks for becomes a retirement — waiting for a person if there are more than three. Nothing is
deleted by the rollout.
- **Retired data is still backed up and still takes space** until a person deletes it. That is the
point; the thirty-day condition is what keeps it from being forgotten.
- **A person who rejects must come back.** A rejected set is kept active and said as a warning until the
mesh asks for it again or someone approves it.
- **The TypeScript providers get the stable count, the ten minutes, the threshold and mark-only** from
the SDK's loop on their next build (every catalogue provider's range accepts it). Mark-only is safe —
nothing is destroyed — and weak: the consumer keeps its access, and a restart forgets the mark. A set
over the bound there waits with no way to approve it until each provider passes its module name and an
announcer to the loop. Until each adds a `retire` and that wiring, they are covered for safety and not
for disabling or cleanup, and the record says so rather than claiming otherwise.
- **The Go loop is still two identical copies**, in postgres and keycloak,
held together by a test — now over three files. Its home is the Go SDK ([ADR 0039](0039-what-the-sdk-holds-and-refuses.md));
moving it there means a tagged SDK release before the catalogue can use it, a separate change.
- **The rollout order**: the controller first (it derives the permission to publish the new event and
hears it; an older controller makes the provider's announcement a refusal the provider logs), then the
catalogue's two Go providers, then the SDK's release and each TypeScript provider as it is rebuilt.
## How it is checked
| What | Checked by |
|---|---|
| a transient empty list for four passes retires nothing; five passes in twenty-five seconds retire nothing, the same set held ten minutes does; ten minutes in fewer than five passes do not; an unreadable pass and a changed set restart the count; additions and changes are not delayed | the catalogue's shared `retirement_test.go` (identical in both Go providers) and the SDK's `retire.test.ts` |
| over the bound waits, is announced and said, again every fifteen minutes; approve takes only the exact set; reject keeps it and is not asked again; a set asked for again settles | the same tests |
| retired is disabled with data intact; asked for again it is enabled as it was | postgres's `live_test.go` and keycloak's `live_retire_test.go`, against throwaway servers |
| delete removes only that consumer, never an active or asked one | the same live tests, and `retirement_test.go` |
| a restart retires what the backend holds unasked and adopts what it finds disabled | `retirement_test.go`, the provider's inventory tests |
| the conditions, verbs, hand acts and the thirty-day probe | the controller's tests of the event, the verbs against a fake provider on a real bus, and the probe |
| an adapter without `retire` is mark-only: `remove` is not called on retirement, the consumer is marked with its access kept and said loudly, and `remove` runs only on a person's delete; `reenable` runs before create for a retired consumer asked for again | the SDK's `retire.test.ts` and `sdk.test.ts` |
| the two Go copies agree | `harness_same_test.go` in each provider, over the three files |
| live, after rollout | `cleanup list` names the adopted and set-aside entries; `conditions` shows no `retire-waiting` on a mesh nobody is changing |
## References
- [ADR 0229](0229-the-cores-order-is-a-lease-the-store-remembers-and-an-epoch-a-machine-is-sent-once-it-reads-one.md)
— replaced in part: its withdrawal brake paragraph. Everything else in it stands.
- [ADR 0227](0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md) rule 4,
*nothing dropped silently* — which this keeps: a stable result over the bound stops and raises a
condition.
- [ADR 0224](0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md)
— the provider's events and their derived permission, which the retirement event follows.
- [Issue 241](../04-ISSUES/241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md)
— what withdrawal destroyed, and what it stopped destroying.
- [To-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) §2, §4, §7 and Phase 2, and
[the module protocol](../03-DESIGN/01-to-be/19-the-module-protocol.md), amended alongside.
- mesh-catalog `modules/postgres` and `modules/keycloak` (`retirement.go`, `retire_pg.go`, `oidc.go`);
mesh-sdk `src/provisioner` 0.1.12; mesh-controller's `retire` and `cleanup` verbs and its probe.
@@ -1,158 +0,0 @@
---
topic: the mesh
status: accepted
date: 2026-10-06
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md
---
# 231. A healer acts on what observation raised, and only observation says it worked
## Context
Phase 3 of [to-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) — *healers* — was
built on 2026-10-06 in the controller and the node-engine. §7 of the design gives each healer a row:
the condition, the repair, the budget and what happens then. Building them met questions the rows do
not answer, and each was answered in code. This record is those answers.
Five could not be left to taste:
- **Where a budget is counted.** A budget kept in the controller's memory is reset by every restart,
and a controller restarting in a loop is exactly when a healer must not start counting again. A
budget kept in the condition is lost when the condition clears and reopens — which a repair that
half-works makes happen every few minutes.
- **When an act has failed.** A repair is made in a moment; whether it worked is known only when the
watchdog or the probe that raised the condition looks again — thirty seconds for a watchdog, five
minutes for a probe.
- **What a healer may send.** H1's repair is "send the current declaration again", which is what a
named push does — and a named push sends a machine the builds a policy or a plan is holding back
([ADR 0221](0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md)).
A person naming a machine means it; a healer does not.
- **What triggers H4.** The design names S9's `slow-consumer`. The controller hears a slow consumer for
its own connection only until the bus has a system account
([issue 270](../04-ISSUES/270-phase-1-watches-signals-that-come-later-and-hears-only-the-controllers-faults/00-report.md)),
and a slow *connection* is not a consumer a reset repairs. What 248 met — a durable consumer a week
behind its stream — is what D6 finds.
- **How the machine is asked to report.** The design gives the node-engine a `report` verb. A machine's
bus user may not answer anybody's inbox, and granting it that is a second channel out of every
machine.
## Considered Options
1. **Budgets in memory, success when the repair returns.** Rejected: a restart resets every budget, and
"the act returned" is the healer's opinion of its own work — the second answer to a question
rule 1 says has one.
2. **Budgets in a bucket on the bus.** Rejected: the bus's user list must grant a new bucket before the
first heal could be counted, so the first controller with healers would heal uncounted on the live
mesh until the bus's machine was pushed.
3. **Budgets in the controller's store, each act begun before it is made; success only as the
condition's clearing; H4 on D6's own kind; the machine answering through its ordinary report.**
Chosen.
## Decision
**A healer is a row of the healer registry**, compiled into the controller: the condition kinds it
answers, its repair, its budget (acts against one *budget key* within a window), its **settle** (how long
after an act the observation is given before the next act or the escalation), what happens when the
budget is spent, where the repair runs, and the event each act is said as. A test generated from the
registry holds every row to a kind the mesh raises — a row of the signals table, a probe of the
self-check, a provider's event — a budget, a settle inside its window, and its event, and every healer
the controller runs to an induced failure that sees it act and brake.
**The healers of Phase 3:**
| Healer | Answers | Repair | Budget, settle |
|---|---|---|---|
| H1 | `sent-not-reported` (S2) | ask the machine's node-engine to report again; if it does not then report the declaration it was sent, or says nothing within 45 s, send it again — never moving a build a policy or a plan holds back | 2 per condition in 6 h; 3 min |
| H2 | `stalled` (S3) on a wait that is **superseded** (a newer plan of the same repository and branch exists) or **finished** (every module of every tier built or failed, every one that rolls out sent) | close the plan with its note — `superseded`, naming the newer plan, or `done`; what it asked still builds | 1 per plan in 24 h; 2 min |
| H3 | `holder-silent` (D3), `consumer-lost` (D6, S9) | assert the bus's streams, consumers and seat workers — the assertion every send makes (issue 208) | 1 per object in 1 h; 6 min |
| H4 | `consumer-behind` (D6) on a consumer the stream table marks resettable | `broker consumer-reset` | 1 per consumer in 24 h; 6 min |
| H5 | `provider-failing`, the identity provider's administrator refusing the mesh's secret | the provider's own repair (ADR 0224 §5), registered and not run by the controller | the provider's |
- **A condition a healer does not apply to is left alone.** H2 passes over a plan waiting on something
still to come; H4 over any consumer not marked resettable. Passing over is not an attempt: no budget,
no escalation, said once in the controller's journal. **Only the controller's own events consumer is
marked resettable**, with why in the table: what a reset drops is caught up — a merge by the catch-up
pass that reads the forge (issue 266), a build's outcome from the build records (issue 214), a
provider's failing word said again within a quarter of an hour (ADR 0224). A module's consumer is
never: nothing would catch up for it.
- **A send by a healer is a push that did not name the machine** (ADR 0221): a machine whose
composition would move a held build is not sent again, and the attempt says so and why.
- **D6's far-behind finding has its own kind, `consumer-behind`**, so H4 answers it and nothing a reset
cannot repair; a probe's registry row names the kinds its findings carry besides its own.
**Every act is begun in the controller's store before it is made**, under the lease and carrying its
epoch, and finished after it: `acted`, `failed`, or `escalated`. The budgets and the brake are counted
from there, so a controller dying mid-act has still spent it, and one restarting in a loop resets
nothing. Only the controller holding the lease heals; one serving without it (S12) heals nothing — a
repair is the one act that can always wait.
**Success is never a healer's to say.** An act is kept in its condition's `tried` as `healer Hn`, the
resolver becomes `healer:Hn`, and the condition stays open until the watchdog or probe that raised it
no longer observes it. A condition cleared and reopened within ten minutes carries what was tried. **A
spent budget** — the budget's acts made, the last one's settle past, the condition still open — hands
the condition to the operator: resolver `operator`, severity `urgent`, an attempt saying what was tried.
An observation does not lower an escalated condition's severity again, and no healer touches it until it
clears.
**Every act is said** as the controller seat's event `healer-acted`: the healer, the condition and its
kind, the act, the outcome, where the budget stands, the epoch, and the verb that shows more. A heal is
never written to the hand-act log, which is how S15 tells a repair the mesh made from one a person had
to. `healers` lists the registry, the acts lately and the brake; `status` counts the week's heals.
**The mesh-wide brake.** Twelve acts in an hour, all healers together, and every healer stops: the
urgent condition `mesh.healers.braked` names which healer acted on what, and the brake holds until an
hour after the last act. A healer looping is then at most a dozen acts, said, never the incident.
Escalations are sayings, not acts, and are not counted.
**The `report` verb, as built.** The controller asks on `mesh.node.<machine>.ask.report`, on core NATS,
which each machine's bus user may subscribe to for its own name only. The node-engine enqueues a
reconcile — the one apply queue's ordinary act — whose account of the declaration it keeps is said
whether or not it is news; a delivery waiting meanwhile is applied and reported instead. **The answer is
the machine's ordinary report** on its own report subject: nothing new is published and nobody's inbox
is answered. A node-engine older than the verb hears nothing, and H1 then sends again, as a person did.
**S15, as built.** The watchdog reads the hand-act log's fortnight every half minute; a cause recorded
twice within fourteen days raises `mesh.hand-acts.<cause>.healer-wanted`, naming the acts, who and why —
and, where a healer answers that cause, that it was not enough. It clears when fewer than two acts of
that cause remain within the fortnight.
## Consequences
- **The controller, the bus's user list and the node-engines roll out in any order**: a machine whose node-engine cannot hear the question is sent again instead, and a
bus whose user list does not yet grant `healer-acted` refuses only the event — the act and `tried` still
say it. A push of the machine holding the bus ends that.
- **H1 sends what a push of the machine would, minus what is held back.** A send that went unreported
because the machine holds a held build is not repaired by a healer: its two attempts say why, and the
operator's `push <machine>` is still the way, recorded as a hand act.
- **Twelve an hour is a guess**, like Phase 1's first bounds: corrected from the live week, in the
registry, reviewed like code.
- **A heal that half-works costs a few acts, then a person.** The settle and the budget make the
slowest probe the pace: H3 and H4 act at most once per object an hour and a day.
- **Not built:** the induced failures on a lab mesh (mesh-lab); a healer for the commonest remaining hand
acts that have none — a ban lifted, a kept file restored — which S15 will now ask for by name.
## References
- [To-be 45 §7](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) — the design this record
details.
- [Research 031, evidence §(f)](../01-RESEARCH/031-a-core-that-cannot-fail-silently/01-evidence.md) — the
hand acts the healers replace.
- [Issue 208](../04-ISSUES/208-a-seats-worker-is-made-only-when-the-controller-starts/00-report.md) — H3 as
the send's own assertion, for D3 and D6 alike.
- mesh-controller and mesh-host, branch `feat/a-core-that-cannot-fail-silently-phase-3`.
## How it is checked
| What | Checked by |
|---|---|
| every healer answers a raised kind, with a budget, a brake, an event and an induced failure | `TestEveryHealerAnswersAKindTheMeshRaisesWithABudgetABrakeAndItsEvent` |
| H1 asks, sends again, keeps `tried`, says `healer-acted`, escalates and stops | `TestH1AsksAMachineToReportAndSendsItAgain` |
| H2 closes a superseded plan and leaves one still waiting | `TestH2ClosesAPlanAnotherHasTakenOver` |
| H3 against a real bus: repaired, cleared by the probe, escalated after its budget | `TestNatsH3AssertsAMissingConsumerAgainAndBrakesAfterItsBudget` |
| H4 against a real bus: reset, cleared by the probe, and the controller still hears | `TestNatsH4ResetsTheControllersEventsConsumerAndItStillDelivers`, `TestH4ResetsOnlyWhatTheTableMarksResettable` |
| the brake, and no heal without the lease | `TestTheBrakeStopsEveryHealerAndSaysSo`, `TestNoHealerActsWithoutTheLease` |
| the `report` verb | mesh-host `internal/link/asked_test.go`, against a real bus; the grant in `TestNatsAskToReportReachesTheMachine` |
| S15 | the signals table's generated test, and `TestARepeatedHandActNamesItsCauseAndItsHealer` |
| live | `healers`, each condition's `tried`, and the hand-act log's weekly count in `status` |
@@ -1,143 +0,0 @@
---
topic: what runs on it
status: accepted
date: 2026-10-06
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md
---
# 232. A binding to a consumer's data moves only by a person
## Context
**A rule written for the resolver moved a machine's databases.** [Issue 258](../04-ISSUES/258-every-machine-bound-the-resolver-to-itself/00-report.md)
found that a machine running its own provider of a mesh-wide provision bound its consumers to that
provider, even when the mesh's seat for the provision was held elsewhere. Its fix made the seat's
holder, or a pin naming another machine, win over the local provider. That is right for the mesh's
resolver: every resolver gives the same answer, so a consumer moved between two loses nothing.
The fix applied to every seat that delivers a provision. The store's seat delivers the relational
database. On the home server, which runs its own store with five applications' databases in it, the
fix's first push re-bound all five to the store on the control node, which holds the seat. That store
did what a provider does for a new consumer: it made each one an empty database. The applications
started, ran their first migrations, and served empty data for about twenty hours. Nothing warned
([issue 273](../04-ISSUES/273-a-rule-for-the-resolver-moved-a-machines-databases/00-report.md)).
Nothing was lost, only because the old store kept everything.
The incident has two parts:
1. **Issue 258's rule applies to two kinds of provision as if they were one.** It does not say which
kind it is for.
2. **Nothing in the mesh knows where a consumer's data is.** A resolution answers one question:
*which provider would I choose now?* Every input to that answer can change under a consumer without
anybody meaning to move it: the seat's holder (a handover, ADR 0131), a pin added or removed, a
provider assigned or unassigned. Only the pair secrets in the store held a trace of the old
binding, and only because a new pair was minted beside it.
## Considered Options
1. **Revert issue 258's fix.** Rejected. It was correct for the resolver, and the resolver would
break again.
2. **Make the store's seat not deliver its provision.** Rejected. That fixes one seat, and the next
seat that delivers a provision which keeps data repeats the incident.
3. **Refuse any resolution that changes a binding.** Rejected. It would refuse the resolver's moves,
which are the point of 258, and a person moving a database on purpose would have no way to do it.
4. **Know which provisions keep their consumers' data. For those, keep each consumer where it was
last sent, say every move the mesh would have made, and let only a pin move it.** Chosen.
## Decision
**1. An offer says whether it keeps its consumers' data.** The field is `keeps-consumer-data` on
`provides`.
- If unsaid, a provider that **grants** each consumer a credential of its own keeps that consumer's
data. The grant makes an account (a role and its database, a key and its bucket, a client), and
what the consumer writes under it stays with that provider.
- A provider that grants nothing keeps nothing of anybody's. Examples: the resolver, a CA, the
artifact store, a route.
- The property belongs to the provision's name, as brokering does (ADR 0009): if any provider says
the provision keeps data, it keeps data.
**2. Issue 258's rule is for provisions that keep nothing.** For a provision that keeps data, the
seat's holder on another machine does **not** overrule a provider beside the consumer. A pin naming
another machine still does, because a pin is a person.
**3. Where each such consumer was sent is recorded, and a resolution keeps it there.** One row per
consumer and provision is written when a declaration carrying the binding is sent. The provider is
recorded by name, so a provider's machine leaving the mesh does not erase where the data is.
- **Recorded binding, different provider chosen, no pin naming the new one:** the recorded provider
keeps answering. The move is said.
- **Recorded provider no longer provides the provision:** the machine's set is **refused**, naming the
pin that would confirm the move. It is never answered by the provider chosen instead. This is ADR
0009's stance, already applied to a pin naming a provider that is gone.
- **New consumer with no record:** it binds as resolved, and is recorded when first sent.
**4. Only a pin moves it.** The pin names the provider, for that machine and provision. The send
that carries the move records the new provider and keeps where it was. The data is moved by a person
before the push; the mesh does not move data.
**5. Nothing is silent.** A kept move is printed by the push that composed it and raised at once as
an **urgent** condition:
> would move X's P from A to B — its data is on A; kept there. `pin …` to confirm a move (and move
> the data first)
A self-check probe ([to-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) §4)
checks every machine every run. It raises:
- a kept move, as urgent;
- a pinned move not yet sent, as a **warning**, so the data goes first;
- any consumer about to be sent another provider than the one on record with no pin naming it, as
**urgent**. The resolver makes this impossible; the probe exists for the day it is not.
> **The mechanism changed — 2026-10-06, by [ADR 0233](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md).** The decision stands: a binding to
> a consumer's data moves only by a person. What moved is how a provision is known to keep data: a
> provider now says it in its `data` section, as a class for its consumers — irreplaceable, valuable or
> rebuildable keep data, cache and none do not — and `module check` refuses a provider that grants
> without saying it. `keeps-consumer-data` on the offer and the inference from `grants` are read only
> from a definition already on the shelf.
## Consequences
- **The first push after rollout records what each machine is bound to then.** A consumer moved
silently before the rollout and never moved back would be recorded where it was moved to. The
rollout therefore starts by checking that every binding to data points where the data is. On
2026-10-06 every one did: the five moved consumers had already been pinned back.
- **A pin is per machine and provision, not per consumer.** Pinning one consumer of a machine
elsewhere pins all its consumers of that provision. That is how pins have always worked, and the
warning names every consumer that would move.
- **Unassigning a store beside its consumers now refuses their machine** until a person pins them
elsewhere. Before this, they moved silently to whichever store answered.
- **A retired consumer's binding stays on record** ([ADR 0230](0230-a-consumer-the-mesh-stops-asking-for-is-retired-and-deleted-only-by-a-person.md)).
Assigned again, the consumer returns to the provider that kept its retired data.
- **No module definition changes.** Every provider in the catalogue that keeps data already grants
credentials, and none that keeps nothing does. A definition states the field only when the default
is wrong for it.
## How it is checked
| What | Checked by |
|---|---|
| the incident: a machine running its own store and its consumers, the store's seat held elsewhere — the consumers stay on their own store, and the resolver beside them still follows its seat | mesh-controller `internal/catalogue/bound_test.go`, and `cmd/mesh-controller/bindings_test.go` through the real stores; both fail without the fix |
| a recorded consumer is kept when the seat's holder changes, and the move is said with its pin | `bound_test.go` |
| a pin moves it; a gone provider, or a store beside it unassigned, is refused and never answered elsewhere | `bound_test.go`, `bindings_test.go` |
| an offer's `keeps-consumer-data` is read and written, and unsaid follows the grant | `bound_test.go` |
| a binding is recorded on send, a move keeps where it was, and a provider's machine leaving keeps the record | `internal/inventory/bindings_test.go` |
| a push raises a kept move at once; the probe says kept (urgent), pinned and not yet sent (warning), and unasked (urgent) | `bindings_test.go` |
| live, after rollout | the self-check passes its binding probe on every run, and `conditions` holds no `binding-kept` on a mesh nobody is changing |
## References
- [Issue 273](../04-ISSUES/273-a-rule-for-the-resolver-moved-a-machines-databases/00-report.md):
the incident.
- [Issue 258](../04-ISSUES/258-every-machine-bound-the-resolver-to-itself/00-report.md): the rule
this record limits.
- [ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md): a seat's holder answers. This
record extends it, because the holder now answers only where nothing is kept.
- [ADR 0009](0009-modules-and-the-graph.md): refuse rather than guess.
- [ADR 0230](0230-a-consumer-the-mesh-stops-asking-for-is-retired-and-deleted-only-by-a-person.md):
the provider keeps a consumer's data until a person deletes it.
- mesh-controller PR #86: `catalogue/bound.go`, the `binding` table (migration 0071), and the
self-check's binding probe.
@@ -1,281 +0,0 @@
---
topic: what runs on it
status: accepted
date: 2026-10-06
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0232-a-binding-to-a-consumers-data-moves-only-by-a-person.md
---
# 233. A module declares the data it holds, and the mesh protects and watches it from that declaration
## Context
**This is the operator's decision, taken on 2026-10-06**, the day after five applications ran for
twenty hours on empty databases ([issue 273](../04-ISSUES/273-a-rule-for-the-resolver-moved-a-machines-databases/00-report.md)).
In the operator's words: modules must indicate the "data storage which should be monitored/protected",
"in the module's manifest", "in a generic, configurable way". The ranking of what is precious is also the
operator's, given the same day and recorded as given:
> "data loss would be painful but not the end of the world. Only my plex storage is sacred and the photo
> sites as well (I'm not sure everyone has a good backup themselves but it's their own responsibility)."
> "we can't backup the plex storage, we don't have any storage for it, it's on a RAID5 which should be
> safe as far as we can afford it. The plex metadata storage is not important" — and the metadata
> "should be part of the standard backup plan of course".
What the mesh knew about data before this record, measured on the code and the catalogue that day:
- **Where a consumer's data is** was inferred, not said. [ADR 0232](0232-a-binding-to-a-consumers-data-moves-only-by-a-person.md)
made a binding sticky when the provider *grants* a credential; no manifest in the catalogue stated
`keeps-consumer-data`. A provider whose grants hold nothing — a cache — was sticky all the same.
- **Backups were a second list.** [ADR 0214](0214-backups-guard-against-mistakes-and-stay-on-the-machine.md)
had a module write `run` and `path` lines to the `node-backup` seat by hand; thirteen modules did,
across two catalogues. The rule that "the catalogue check refuses a store provider that declares none"
was never in `module check`: one controller test applied it to the catalogue checked out beside it,
against a list of six provisions kept by hand.
- **Unassigning a module kept its data only by accident of the host's rule.** The node-engine removes an
orphaned directory only when it is empty; a directory with anything in it is **kept and forgotten** —
its record dropped from the machine's store, the report saying "kept … remove it by hand once you know
what it is". Nothing anywhere then knew the data was there, and the one instruction given was the one
that destroys it.
- **Nothing was measured.** No size, no last write, no check that a backup covered what mattered, and
nothing that could tell an empty copy of a database from the full one beside it — exactly the shape
of issue 273.
- **The media library was out of every design** (ADR 0214 named it as not backed up), and the photo
sites' storage lives in two providers, the object store and the document store, as consumer data
whose preciousness only the photo application knows.
## Considered Options
1. **Keep ADR 0232's inference and add a list of protected paths in the backup module.** Rejected: a
second list is the failure ADR 0214 rejected ("whatever is not listed is unprotected, silently"), and
an inference cannot say that a grant holds a cache.
2. **Keep `keeps-consumer-data` on the offer, required, and add a separate `backups` field.** Rejected:
two places for one fact, and a yes/no cannot carry the operator's ranking — a cache's consumers and a
photo library's both "keep data".
3. **Every declared item backed up, irreplaceable or not; the media library copied too.** Rejected by
the operator: there is no storage for a copy of the media library. A class that forces a backup would
make the most precious data undeclarable.
4. **A class per item decides everything, including the protection.** Rejected for the same reason:
how precious data is and how it is protected are two facts. The media library is the most precious
thing on the mesh and the one thing that cannot be copied.
5. **A per-module check in code ("postgres is a store").** Rejected: never per-module code; a new
module would be unprotected until someone remembered it.
6. **Array health watched by a new `node-storage` seat.** Deferred, not adopted: the only thing needed
today is to *read* an array's health under declared data, and the backup holder already reads that
data on every machine that keeps any. A seat for storage is right the day the mesh *acts* on disks
(scrubs, replacing a member); that is not decided here.
7. **One `data` section, a class for how precious, a protection for how it is kept, and everything
derived from them.** Chosen.
## Decision
**1. One section, `data`, says every kind of data a module keeps.**
- `own`: a list of items, each `id`, `path`, `class`, and how it is protected. `path` names one of the
module's directories, `${dir:<id>}`, or an operator's path it was given, `${access:<id>}`, or a path
beneath either — never a machine path (ADR 0112).
- `consumers`: per provision the module grants, the `class` of what it keeps for each consumer and
`in` — the own item that holds it, or a provision the module itself requires (an identity provider
keeps its clients in its database). **Any provider that `grants` must say this; `module check`
refuses it otherwise.** The offer's `keeps-consumer-data` (ADR 0232) is still read from a definition
already on the shelf, and refused by `module check` in a new one: it is said once, here.
- `kept-by`: per provision the module requires, the `class` of what of its own lives with that
provider. The photo sites' application says its objects and its database are irreplaceable; the
object store and the document store are then held to that class for the machine and consumer
concerned.
- Each entry may carry `why`, a line for the reviewer.
**2. Four classes, ranked by the operator.** A fixed vocabulary; a fifth is a decision.
| Class | Is | Backed up | On unassign | Alerts |
|---|---|---|---|---|
| `irreplaceable` | what must never be lost: the media library, the photo sites' storage | a protection is **required**: a backup, or a declared redundancy | retired, never removed | **urgent** |
| `valuable` | anybody's work, painful to lose: every store, mailboxes, repositories, the vault, the agent's home | **yes, by default** — the standard nightly plan, where the machine has a backup holder | retired, never removed | warning |
| `rebuildable` | made again from elsewhere: a dump, a clone, thumbnails, Plex's metadata | **yes, by default**; may say `none` (the artifact store, downloaded models) | forgotten | none |
| `cache` | disposable | never | forgotten | none |
For consumers, `none` also exists: the provision keeps nothing of anybody's (a resolver). A binding to
a consumer's data is sticky (ADR 0232) when the class is irreplaceable, valuable or rebuildable — not
for a cache or none. **Redis is `cache`**: no module in this catalogue requires its provision, its name
says what it is for, and a consumer keeping data in a cache would be the consumer's mistake, not the
mesh's to make sticky.
**3. Protection is its own field.** `backup` is `copy` (the holder reads the path as it stands),
`none`, or `{dump, into}` — a command that writes a consistent copy into another item, which is copied
(a running store's files are not a consistent copy). Unsaid, anything but a cache is copied. Or
`redundancy: "<why that is enough>"`: the item is protected by the redundancy of the storage it is on,
and the mesh watches that storage instead. **An irreplaceable item has a backup or a redundancy;
`module check` refuses it with neither.** `within` bounds the age of its last good backup (48 hours
unsaid, ADR 0214); `active` says it is written all the time and how long quiet is a fault.
**4. The backup holder's lines are derived from the section.** The controller composes, from every
module on a machine, the `backup` lines it always did and a second kind, `data`: every item with its
class, its path, the path that covers it and its protection. **A backup line written by hand is
refused by `module check`** and, beside a data section, not placed: one list. A module keeping
anything irreplaceable depends on the machine's `node-backup` holder (ADR 0207) — the holder backs it
up or watches its array — so it cannot be assigned where nothing would. Valuable and rebuildable data
is backed up where a holder is and refuses no machine for want of one.
**5. The holder measures; the controller judges — and nothing large is ever walked.** The holder
measures every item that is not a cache by the item's own `measure`:
- `walk` (the default, for small items): every file summed and the newest change found, in one walk as
root — **at most once a day**, after the night, and **stopped after ten minutes or two million
files**, said as "measured partially": a partial size is a lower bound, shown and never compared;
- `dataset`: the ZFS dataset holding the path — its size from the filesystem's own counters — and the
newest change among the path's top-level entries; hourly, nothing walked. Items on one dataset share
its size, so a shrink of it is one condition naming them all;
- `shallow`: the newest change among the top-level entries only, and no size; hourly.
A large item never says `walk`: the media library is `dataset`, Plex's metadata and previews, the
artifact store and the models are `shallow`. A file walk over a pool of that size every hour is load on
the array that protects the library, and competes with the server reading it. The holder also reads the
redundant storage each item is on: a ZFS pool's health and its own verdict, an md array's members, a
btrfs filesystem's error counters. It answers this, with each item's newest good backup, in
`node-backup.backed-up`; it decides nothing. **The holder is the measurer**, not the node-engine: it
already reads every declared path as root, it is a module that can be replaced, and the core gains no
work. The controller's self-check asks every holder on every run.
**6. Unassigning keeps the data, and the mesh remembers it.** An irreplaceable or valuable item in a
module's own directory that its machine no longer declares is **retired** (ADR 0230's meaning): kept,
recorded with when and why, listed by `cleanup list`, said by `unassign` when it is done, and a warning
after thirty days. Assigned again, it is no longer retired. It is deleted only by `cleanup delete
<machine> <module> <item> --why`, a hand act, which the machine's backup holder executes **only after
taking a last restore point of it**, tagged as retired, so the deletion can be undone until a person
forgets that restore point; the holder refuses any path declared now. An item on an operator's path —
the media library — is **never** retired and never deleted: the mesh does not own it. The node-engine's
rule stays (a directory with anything in it is never removed); its report no longer says "remove it by
hand".
**7. What is watched, and how loud.** Each condition is **urgent for irreplaceable data and a warning
for valuable data**; rebuildable data and caches raise none. A consumer's data is as precious as the
stricter of its provider's class and its own `kept-by`; a provider's item is held to the strictest
class kept in it.
| Condition | Raised when |
|---|---|
| `data-shrank` | an item holds less than half of its largest size in seven days, and at least 16 MiB less; one condition per dataset for items measured from a dataset's counters; never from a partial size |
| `data-missing` | an item's path is gone |
| `empty-replacement` | a copy in use is less than half the size (and 16 MiB less) of a copy of the same thing kept elsewhere: an item retired on another machine, or a consumer's data at another provider — issue 273's shape |
| `data-held-twice` (warning) | a consumer has active data at two providers whose sizes cannot be compared |
| `data-quiet` | an item that says `active` has not been written within it |
| `backup-stale` | an item's newest good backup is older than its bound, or there is none — for valuable data, only where a holder measured it |
| `array-degraded` | the redundant storage under watched data is not healthy, or cannot be read; one per array |
| `protection-missing` | an item says `redundancy` and is on storage the holder cannot read as redundant |
| `data-unmeasured` (warning) | a machine's holder did not answer |
| `cleanup-waiting` (warning) | an item retired more than thirty days |
**8. A verb reads it.** `data` lists every item on every machine with its class, protection, path,
size, newest write, newest good backup and its bound, the array under it, and what is retired.
**9. Every catalogue module that holds data declares it** (the table below), and the catalogue check
refuses a container that writes one of its directories, mounted whole, without a data item covering
it — a directory the mesh does not know about is data the mesh cannot protect.
### The classes, per module
The operator's ranking applied to every module that holds data, **for the operator to correct**: only
the media library and the photo sites' storage are irreplaceable; everything people wrote is valuable.
| Module | Own data | Its consumers' | Kept with a provider |
|---|---|---|---|
| plex | the media library, eight operator paths: **irreplaceable**, by **redundancy**, measured from its dataset; metadata: rebuildable, backed up, measured shallow; **previews: rebuildable, not backed up** (Plex makes them again; too large to copy nightly); working data: rebuildable; transcode: cache | | |
| photos | | | object store and document store: **irreplaceable** |
| postgres, mssql, mongodb | the store: valuable, by dump; the dumps: rebuildable | valuable | |
| minio, influxdb | data (and influx's config): valuable | valuable | |
| mesh-vault | state, ledger, root: valuable | valuable | |
| mailu | mail, signing keys, data, calendars, webmail, the queue: valuable; filter, certificates, autoconfiguration, fetch state: rebuildable; virus signatures, its redis: cache | valuable | |
| gitea | data: valuable | the package registry: rebuildable | |
| keycloak | | clients: rebuildable, in its database | |
| umami | | page views: valuable, in its database | |
| mosquitto | retained messages: rebuildable | rebuildable | |
| redis | its snapshot: cache | **cache** | |
| nextcloud, baserow, grafana, matrix, nodered, unifi, step-ca, audit-logger | their data: valuable | | |
| home-assistant | its configuration and history: valuable, written all the time (a day) | | |
| nats | the bus's streams and buckets: valuable, written all the time (a day) | | |
| n8n | data: valuable; scratch: cache | | |
| supabase | database: valuable, **by `pg_dumpall`** inside its container; storage, functions: valuable; the dump and its seeded configuration: rebuildable | | |
| claude-code | the operator's agent's home: valuable | | |
| ssh-client | the operator's keys: valuable | | |
| radarr, sonarr, lidarr, bazarr | the database: valuable, by sqlite's own backup; settings: valuable; covers and the night's copy: rebuildable; lidarr's start-up scripts: cache | | |
| bookshelf, jackett, kometa, nzbget, ombi, qbittorrent, tautulli | configuration: valuable | | |
| distribution, ollama | the artifact store, the models: rebuildable, **not backed up** (too large; made again), measured shallow | | |
| only-office | its working data, database, fonts: rebuildable; logs, queue, cache: cache | | |
| records, route-proxy | the record's clone, issued certificates: rebuildable; the authority's roots: cache | | |
| restic | the repository: rebuildable, not backed up — it is the copy, never the only copy of anything; its state: cache | | |
| build-agent, icecast, lab, model-usage, searxng | cache | | |
**Uncertain, for the operator:** the bus's streams are copied as live files — the bus's image carries
no tool to snapshot a stream safely, so a consistent copy needs the bus's own backup command run from
somewhere that has it, which is not built; mailu's queue is valuable though transient; the agent's home
and the operator's keys are valuable, and are backed up only on machines that hold `node-backup`
(today, neither workstation does).
> **Progressive insight — 2026-10-06.** The first of these items is resolved, and was a fact about what
> was built rather than part of the decision. The record said the bus's streams "are copied as live
> files" and that a consistent copy "is not built". It now is: the nats module's streams item is
> protected by a dump — `backup: {dump, into}`, this record's own mechanism — that runs a snapshot
> program built into the bus's image, through the server's snapshot API, under the bus module's own
> account granted that API and nothing else; the live store is no longer copied, and a restore builds
> a new store beside the live one ([ADR 0235](0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md)).
> The bus's row in the table above stands: its streams and buckets are valuable, written all the time.
## Consequences
- **The first run of the self-check records what every machine holds**, and the first night backs up
what was not backed up before: on the home server, every application's own data and Plex's metadata;
on the control node, the bus's streams and the certificate authority. Measured read-only on
2026-10-06, the home server's new nightly data is about 11 GB of application data (the chat server's
database 5.6 GB, the database platform's 2.1 GB before its dump, the network controller 1.1 GB, the
viewing history 0.9 GB, dashboards 0.5 GB, the rest under 0.2 GB each) and Plex's metadata, at most
the 136 GB its disk holds — under 150 GB against 33.6 TB free where the repository is: under half a
percent. The previews, on the media pool, are left out. Later nights cost what changed.
- **The media library is watched, not copied.** Its array is read every hour; a degraded pool is
urgent. A pool that loses a vdev still loses the library — the operator's accepted risk, now said
rather than silent. SMART pre-failure warnings are not read: the pool's own device error counters are,
and reading SMART is the work of a storage seat when one is decided.
- **A provider that grants and a container that writes a directory must declare their data** before
`module check` passes; a module outside the catalogue meets this in its own repository, and a module
already on the shelf is read as it is.
- **Retired data takes space until a person deletes it**, as ADR 0230's retired consumers do.
- **An empty replacement of a store's consumers is a warning**, not urgent: the stores' data is
valuable by the operator's ranking. A consumer that says `kept-by` irreplaceable — the photo sites —
makes it urgent.
- **Rollout order**: the controller first (it reads the new section and composes the holder's new file;
an older holder ignores what it is not given), then the catalogues and the photo application, then a
push of every machine. The controller's grant gains `node-backup.backed-up`, and the installer's first
user list with it.
## How it is checked
| What | Checked by |
|---|---|
| the section parses and refuses what it must: unknown class, a machine path, an undeclared directory or access, irreplaceable with no protection, a cache backed up, a dump into nothing | mesh-controller `internal/catalogue/data_test.go` |
| a granting provider says its consumers' class; nobody says it on the offer; no backup line by hand; a written directory is declared; data kept with a provider as irreplaceable is protected there | `data_test.go`, and `module check` over the whole catalogue (`TestModuleCheckPassesTheCatalogue`, `TestEveryCatalogueModuleDeclaresItsData`) |
| stickiness follows the class; an older definition still follows its grant | `data_test.go` (`TestKeepingConsumerDataFollowsTheClass`) |
| the holder's lines are derived — dumps, copies, redundancy, a home directory, an operator's path — and only irreplaceable data needs a holder | `data_test.go` (`TestTheBackupHoldersLinesAreDerivedFromTheData`) |
| unassigning keeps irreplaceable and valuable data retired, says so, and lists it; an empty replacement of it elsewhere is said — through the real stores and a real bus | `cmd/mesh-controller/data_test.go` (`TestNatsUnassigningIrreplaceableDataRetiresItAndAnEmptyReplacementIsUrgent`), `internal/inventory/data_test.go` |
| **issue 273 replayed**: five consumers' real databases at one provider, empty ones at another — each an empty replacement, urgent where the consumer keeps irreplaceable data there; a deliberate move is silent | `data_test.go` (`TestAnEmptyReplacementOfAConsumersDataIsSaid`, `TestADeliberateMoveIsNoEmptyReplacement`) |
| shrink, missing, quiet, backup age, array health and missing protection, by class | `data_test.go` |
| the holder reads the data file, measures, reads ZFS and md, and deletes only after a last restore point, never a declared path | mesh-catalog `modules/restic/cmd/restic-backups/data_test.go` |
| the node-engine keeps a directory with anything in it and says so without "by hand" | mesh-host `internal/apply/apply_test.go` |
| the controller's grant names `node-backup.backed-up`, and the installer's carries it | mesh-controller `internal/broker` (`TestTheInstallersFirstUserListIsWhatTheControllerWouldCompose`, the golden) |
| live, after rollout | `data` lists every declared item with a measurement and a recent backup; `doctor` passes D13; `conditions` holds no `array-degraded` while the pool is healthy |
## References
- [Issue 273](../04-ISSUES/273-a-rule-for-the-resolver-moved-a-machines-databases/00-report.md) — the incident.
- [ADR 0232](0232-a-binding-to-a-consumers-data-moves-only-by-a-person.md) — extended: stickiness now
follows the class, and `keeps-consumer-data` moved into the data section.
- [ADR 0230](0230-a-consumer-the-mesh-stops-asking-for-is-retired-and-deleted-only-by-a-person.md) — the
meaning of retired, applied to a module's own data.
- [ADR 0214](0214-backups-guard-against-mistakes-and-stay-on-the-machine.md) — its backup lines are now
derived; its exclusion of the media library is replaced by redundancy watched.
- [ADR 0051](0051-shared-data-is-the-operators.md) — an operator's path is the operator's: never retired
or deleted.
- [To-be 32](../03-DESIGN/01-to-be/32-what-a-module-declares.md), [to-be 43](../03-DESIGN/01-to-be/43-backups-against-mistakes.md),
[to-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) — amended alongside.
- mesh-controller `internal/catalogue/data.go`, `internal/inventory/data.go` (migration 0072),
`cmd/mesh-controller/data.go`; mesh-catalog `modules/restic`; mesh-host `internal/apply/apply.go`.
@@ -1,462 +0,0 @@
---
topic: the mesh
status: accepted
date: 2026-10-06
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md
---
# 234. The mesh holds a conversation with its operator, over channels that are seats, and an answer that performs an action is authorised by the controller
## Context
**The output channel was built in its smallest form** ([ADR 0227](0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md),
[to-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) §5): one mesh seat,
`operator-channel`, whose single holder carries the router, a Telegram client and a desktop adapter in
one process, and **no answering back**. ADR 0227 left routing by presence, quiet hours, answering back
and the outside dead-man service open in [research 028](../01-RESEARCH/028-the-meshs-output-channel/00-overview.md),
"whose graduation amends to-be 45". This record is that graduation.
The evidence, from 028:
- **The mesh tells nobody.** 15 of the 236 issue reports say the fault was found because a person
happened to look; 74 describe something failing silently. On the day 028 opened the mesh knew of a
machine refusing every declaration for ninety minutes, a machine out of touch for ten minutes, and a
failure repeated thirteen times, and told nobody ([028/01](../01-RESEARCH/028-the-meshs-output-channel/01-what-the-mesh-already-knows.md)).
- **The built Telegram code has ten defects**, two of which silence exactly what the operator most
needs to hear: nothing bounds a message to Telegram's 4096 characters, so one long message wedges the
channel (D1); and a reopened urgent condition is said by an edit, which on Telegram rings nothing (D2)
([028/04](../01-RESEARCH/028-the-meshs-output-channel/04-telegram-as-the-first-holder.md)).
- **Actions that need a person are ordinary verbs.** `retire approve`, `cleanup delete` and `pin` while
a binding is kept are callable by any principal granted the verb, agents included, and a hand-act
records the calling bus principal, not the person ([028/08](../01-RESEARCH/028-the-meshs-output-channel/08-asks-that-authorise.md)).
`retire approve` re-reads the set when it runs, so "approve what you were shown"
([ADR 0230](0230-a-consumer-the-mesh-stops-asking-for-is-retired-and-deleted-only-by-a-person.md))
does not hold end to end.
- **The desk cannot prove who clicked.** The graphical session is X11: any client of the display can
inject input and read keystrokes, `dunstctl action` invokes a notification's action for any program
of the account, and the desktop holder's bus credential is readable by every agent running as the
operator ([028/08](../01-RESEARCH/028-the-meshs-output-channel/08-asks-that-authorise.md)).
The operator's directions, 2026-10-06:
- Telegram is one channel among many to come; it simply fills a seat. Input (a mail, a message from the
operator) has the same shape.
- Each channel has capabilities, and they decide what may travel on it.
- Agents and modules must be able to **ask** the operator anything, not only for permission.
- The **work context** chooses the channel: in a Telegram conversation, Telegram; at the desk, the desk.
- Approving and rejecting must be possible through Telegram, and while talking to an agent through
Telegram, everything must be completable there.
- **There is no hardware security key.** The proof the operator gives is a code from the authenticator
app on the phone. A key may be added later and must never be required.
Reviewing the proposed record, the operator found four questions it left open, and on 2026-10-06
asked for each to be decided here rather than left as a gap:
- **A lost phone locks the operator out for good.** Re-enrolling a factor is a destroy ask, and a destroy
ask needs a code from the factor that was lost.
- **Nothing says who answers an operator message.** "An agent bridge takes `message`" names no
addressee, and a message nobody takes is silence.
- **The content rule refuses what an ask sometimes needs.** "Which of these two paths should be kept?"
cannot be asked in roles and words.
- **"The operator" was the holder's flag.** The research's envelope carried an `operator` field filled
"from the controller's list", but nothing said who fills it or what an envelope from anyone else may
do.
### Against GENESIS
- **Mission — intake, process, deliver; agents, some of whom are human.** A human agent "acts through a
shell, a desktop, a message from a phone". The conversation is that modality made a first-class part
of the mesh, for asking as well as telling. The distinction this record draws is **not** human versus
non-human as a category: it is which **identity** can present which **proof**. A spawned agent holds
no enrolled factor, so it cannot authorise; nor can a person without one.
- **Core value — failure must be loud.** A conversation that cannot carry something says so, and the
away channel is checked to carry every tier.
- **Core value — sovereignty.** Telegram is an outside dependency, taken deliberately: free, on both
phone platforms, no server of the mesh's own, and the only candidate that carries every tier today.
It holds a seat, so it is replaceable by a holder of the mesh's own (Matrix, once its push is
measured) without changing anything that speaks to the seat. It never sees a factor's secret.
- **Context — one human, usually asleep; nodes are personal and mobile.** Hence routing by where the
operator is, escalation when unanswered, quiet hours, and a watcher's watcher that does not depend on
the control node.
- **Effect — "when something genuinely needs a decision, you are asked, with the context, not a log
line".** That sentence is this record's purpose.
Nothing in 028 conflicts with GENESIS.
## Considered Options
### Where a channel attaches
1. **Each channel contributes itself to the output seat** (to-be 45 §5 as written). Rejected: a
contribution is content a holder places ([ADR 0210](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md));
a channel is running code with a secret, a connection and answers of its own. Making it fit puts
every channel's client back in the router.
2. **One seat per channel kind** (`telegram-channel`, `desktop-channel`, …). Rejected: the router must
learn every seat, and every new channel is a change to the router.
3. **Channel modules found by a manifest field and called by module address.** Rejected: callers use
seats, never modules ([ADR 0126](0126-a-module-declares-its-own-seats.md)).
4. **One notifier with every channel built in.** Rejected: shared secrets, shared failure, a release
of the router per channel.
5. **A per-service module with its own question or approval path** (a "Telegram module" that decides).
Rejected: it locks the conversation to one service, and the next channel repeats it.
6. **Two kinded benches, `channel` (out) and `intake` (in); a fixed capability vocabulary; the output
seat's holder as the router that orders by work context; an authorising layer held by the
controller.** Chosen.
### How answers come in
- **A webhook.** Rejected: it needs a public route into the mesh, and a stolen token can redirect it.
- **Long polling.** Chosen: outbound only; a stolen token can steal updates (visible as a conflict) but
cannot inject one.
### Who performs an authorised action
- **The channel module calls the authorising verb itself.** Rejected: the hand-act would name the
module, a compromised channel could act, and every channel would repeat the checks.
- **The asker performs it on hearing "yes".** Rejected: an agent relaying "the operator said yes" is
not an answer to anything.
- **The controller holds the ask, checks the answer and performs.** Chosen.
### What proves the operator answered
- **A click at the desk.** Rejected: on X11 an agent on the operator's account can produce it.
- **The screen's unlock or a fingerprint reader.** Rejected: only the local holder sees the result,
and an agent can forge what it reports.
- **A security key's touch (FIDO2), required.** Rejected: the operator has no key. Kept as an
optional proof for whoever enrols one; never required for any tier.
- **A TOTP code from the operator's authenticator, verified by the controller.** Chosen as the proof
every tier above acknowledge can rest on.
- **A verified Telegram sender**, through a holder no agent shares. Chosen as a proof for approve, never
enough alone for destroy.
### Whether context may lower the bar
- **Being at the desk, or in a recent conversation, counts as presence and so as proof.** Rejected:
presence proves someone was there, not who answered.
- **Context orders the channels that already qualify, and never makes one qualify.** Chosen.
### The seat shape, against ADR 0223
[ADR 0223](0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md) made `mesh-dns-resolver`
the first bench, of the **replicated** sort (the same module, one holder per machine, every holder
answering the same), and said making another bench is a decision, recorded.
- **Model channels as a replicated bench.** Rejected: the holders are different modules that answer
differently, and a send must reach one chosen holder, not any.
- **Model each holder as a node seat keyed by machine.** Rejected: a channel's identity is its service,
not where it runs.
- **A second sort of bench, kinded.** Chosen, decided here (below).
### Recovering the operator's factor
- **Re-enrolling only through a destroy ask.** Rejected: losing the phone loses the code the ask needs;
a permanent lockout of the mesh's only person.
- **Recovery through the away channel** (a Telegram tap re-enrols). Rejected: the phone that was lost
is the away channel, and a stolen phone would recover itself.
- **A verb on the bus that re-enrols for the operator's principal, without a code.** Rejected: any
holder of that principal, an agent on the operator's account included, could take the factor over.
- **Recovery codes kept in clear in the controller's state.** Rejected: whoever reads the state holds
ten proofs.
- **Ten one-time recovery codes made at enrolment, kept only as hashes, each one P2 proof; a last resort
that needs root on the control node, run there and refused over the bus.** Chosen.
### Who answers an operator message
- **Every operator message to every agent.** Rejected: each agent would act on words meant for
another.
- **To whichever agent spoke last.** Rejected: a guess, wrong exactly when two agents are working.
- **Drop what nobody is addressed by.** Rejected: silence, the failure this effort exists to end.
- **Addressed by `@name` or by the thread it is written in; otherwise to the mesh's own responder,
which answers from the controller's read verbs or says it did not understand.** Chosen.
### Asks that need specifics the content rule refuses
- **Lift the content rule for `private` holders, or for Telegram.** Rejected: Telegram reads every
word, and a holder's declaration is not a reason to let a path leave the machines.
- **Refuse silently, as before.** Rejected: the asker cannot tell what to change.
- **Machine names allowed in words; anything else concrete carried as a reference only a private
surface opens; a refusal names the offending part to the sender.** Chosen.
### Who counts as the operator
- **The holder's own allow-list.** Rejected: a channel module, or its bus account, would decide who
the operator is.
- **The service's own verification alone** (a Telegram account is authenticated). Rejected: it proves
an account, not that the account is the operator's.
- **Drop everything not from the operator.** Rejected: mail and webhooks are inputs the mesh wants, as
data.
- **The controller's list decides; everything else is marked untrusted and may be data for consumers
that accept it, never the operator's words.** Chosen.
## Decision
### 1. The conversation
**The mesh holds a conversation with its operator.** Three things are said: a **message** (the mesh
tells; no answer expected), an **ask** (someone wants the operator's input, of a declared kind), and an
**operator message** (the operator writes first). An input from outside that is not the operator (a
mail, a webhook) is the same envelope with another sender; this record decides its shape only, and its
consumers are later work.
### 2. Channels and intake are kinded benches
- **A kinded bench is the mesh's second sort of bench**: a mesh seat whose holders are **different
modules, each claiming one kind**. Two claims of one kind are refused at registration. A verb's
subject carries the kind, as a node seat's carries the machine.
- **`channel`** (out) and **`intake`** (in) are kinded benches, and the only ones. Making another is a
decision, recorded, as ADR 0223 requires of every bench.
- `channel` serves `send`, `edit` and `standing`, and emits `delivered` and `failed` (permanent or
transient). `intake` emits one envelope per input. A service read and written by one program is held
by one module claiming both seats under one kind.
- **A claim on a kinded bench carries `kind` and `capabilities`.** Capabilities come from the fixed,
versioned vocabulary **`channel-capabilities/1`**; a word outside it is refused. Each capability has
a contract test the holder's build runs and a drill its `standing` can run; one that fails its drill
is withdrawn from routing and reported until it passes.
- The vocabulary has three groups: delivering (`deliver`, `reaches-away`, `loud`, `silent`, `edit`,
`max-length:N`, `reaches-when-mesh-down`, `private`), conversing (`choice`, `reply`, `threads`,
`operator-first`) and trusting (`verified-sender`, `exact-render`, `code-factor`, `key-factor`).
### 3. Who counts as the operator
- **Only a sender on the controller's list of the operator's identities is the operator.** The list
is kept per intake kind, in the controller's own state. While no factor exists, an identity is
enrolled only at a terminal; once one does, adding or removing an identity is a destroy ask.
- **Every envelope is marked `trusted: true` or `trusted: false` by a check against that list, made on
the controller's side** — the router asks the controller, never reads a holder's word for it. A raw
intake event is untrusted by construction; only the router's re-emitted operator message is trusted.
- **An untrusted envelope** may start work for consumers that declare they accept untrusted input (a
mail rule). It can **never** answer an ask, authorise, or reach an agent as the operator's words.
- **The rule for agents and consumers: untrusted input is data, not instructions.** An agent that is
handed one may read it, quote it and report on it, and never follows what it says.
### 4. Who answers an operator message
- **An operator message is addressed:** to an agent by `@name` at its start, or by being written in an
agent's thread (its conversation handle); otherwise to **the responder**, the mesh's own participant,
addressed as `@mesh`.
- **The responder is part of the router.** It answers questions about the mesh from the controller's
read verbs only — status, conditions, asks, and the bindings and data on record — and lists what it
can answer when asked or when it does not understand. It performs nothing.
- **Nothing is answered with silence.** A message that is unaddressed and not understood gets a reply
saying so, and how to address someone.
- **An agent registers as addressable** with the router: its name (unique, bound to its bus
principal), its owner, and what it handles. `@` names come only from that register.
- **A message to a registered agent that is not running is kept**, bounded per agent, and the operator
is told it will be read when the agent next runs. **A message to a name not registered is refused**,
with the names that are.
### 5. The router
- **The holder of `operator-channel` is the conversation's router.** It loses its Telegram client to a
module of its own and keeps no channel's client in its process.
- **It routes by required capability, then by work context, then by severity**, and escalates along a
fixed chain when unanswered. When nothing can carry something, it says so as a condition of its own.
- **Context orders, never qualifies.** The work context chooses among channels whose capabilities
already satisfy the message or ask. It never adds one, and never lowers what an ask requires.
- **The content rule stays the router's**, applied before anything reaches any holder (§6). A holder
declaring `private` is not exempted.
- **Presence is current state on the bus**: a value per machine and per intake kind, overwritten, never
a history, and never in a message's words.
### 6. What a message may name, and references
- **The mesh's own machine names may appear in a message's words.** Domains, addresses, paths and
anything shaped like a secret may not.
- **A message or ask may carry references.** A sender attaches a detail (a path, an address, a log
excerpt) with a label; the router keeps it and puts only an opaque reference and its label in the
words. A secret is refused even as a reference.
- **A dereferenced detail is shown only on a channel declaring `private`, and at the console.** Today
that is the desk and the console. On Telegram, and on any channel not declaring `private`, only the
reference's label is shown. The router's verb `detail <ref>` answers only the console and the intake
holders of `private` kinds, and its answer travels only back to them.
- **A refusal by the content rule is said to the sender, naming the offending part**, so the asker can
rephrase or move the specific into a reference. The offending part is never sent to a channel.
### 7. Asks
- **Kinds:** `yes-no`, `one-of` (at most eight options), `text`, `number`, `date`, `acknowledge`, each
requiring its capabilities (`choice` or `reply`).
- **Life:** open → answered | defaulted | expired | cancelled. The first answer wins; every other copy is
edited to say where it was answered. A default is delivered marked as a default, never as the
operator's answer.
- **An authorising ask never defaults.** It expires.
- An asker holds **at most three open asks**; a fourth is refused in words. Asks count against the
router's hourly cap. Asks burst-batched to one channel go out together under one heading, each its
own message. **History is kept 30 days.**
- **The answer returns to the asker as an event**; an asker may also wait a bounded time, or poll.
### 8. Asks that authorise
**An ask whose answer performs an action is held by the controller.** The router carries it like any
other ask; only the controller performs.
- **Every authorising verb declares its tier** in the controller's verb table (`authorises`: the tier,
and the arguments that make up the exact state the operator must see).
- **The controller's verbs:** `authorise request` (anyone may call it; it renders the ask, stores it
with a digest of the state shown, and performs nothing); `authorise answer` (granted to intake holders
only); `authorisations` (open and recent).
- **The proofs** that the operator, and not someone else, answered:
- **P1, a verified sender:** a tap from the operator's own Telegram account, through a holder placed
where no agent runs as the operator;
- **P2, a TOTP code** from the operator's authenticator app, verified by the controller — valid for
the current or previous 30-second step, and accepted once;
- **P3, a security key's touch bound to the ask** — **optional**. It counts where a key is enrolled,
and no tier ever requires it.
- **The tiers:**
| Tier | Required | On the away channel (Telegram) | At the desk |
|---|---|---|---|
| acknowledge | `choice`, `exact-render` | tap | click |
| approve | the above and **one** proof | tap (P1) | click **and a code** (P2) |
| destroy | the above and **two** proofs, at least one of them P2 | tap and a code (P1 + P2) | click, code and key touch (P2 + P3) where a key is enrolled; otherwise carried by the away channel |
Acknowledge performs only what any granted principal may already do (silencing, announced); it is
not an authorisation, and needs no proof.
- **A desk click alone never authorises.** The desk declares no `verified-sender`; at the desk, approve
and destroy rest on a code the controller verifies (or a key touch, where enrolled).
- **Every tier is completable on the away channel.** The self-check verifies it every run; a setting
that would make a tier possible only at a desk is refused unless the operator chose it for that tier
explicitly.
- **The controller checks from its own records, never the request's claims:** the ask is open and
unexpired; the caller holds an intake kind; that kind's declared capabilities meet the tier; a P1
sender is on the controller's list of the operator's identities; a code is valid and unused; a key
assertion verifies over this ask's challenge; and the state now has the digest it had when shown.
It refuses at the first failure, then performs as itself, closes the ask with a compare-and-set (a
second answer loses), and emits `ask-answered`.
- **No agent authorises.** An agent asks; the operator answers where they are. Agents' grants lose the
authorising verbs. Those verbs refuse direct calls except through `authorise answer`, or as
**break-glass** at the console with a code, recorded and announced as such.
- **`retire approve` takes `expect`**, the set it approves, and refuses if the set differs.
- **The hand-act records** `via` (kind and holder), `requested-by` (agent principal or condition key),
`ask` (its id) and `proofs` (which were present). `by` names the operator as that kind's identity.
- **The factors' secrets are the controller's own.** The TOTP seed is made by the mesh and shown once,
to a terminal, never through a channel or an event. Re-enrolling a factor, or changing the operator's
identities or the away channel, is a destroy ask — except recovery, below, which exists because a lost
factor cannot answer one.
- **A code travels by request and reply only**, from the holder to the controller; never in an
envelope or an event, and deleted from the conversation where the service allows.
**Recovering the factor.**
- **At TOTP enrolment the controller makes ten one-time recovery codes**, shown once to the terminal
together with the seed, never through a channel or an event, and stored only as hashes.
- **Each recovery code counts as one P2 proof, once.**
- **`factor recover`**, given a recovery code at a terminal, re-enrols the TOTP factor (a new seed,
and ten new recovery codes replacing the rest). It is recorded as a hand-act and **announced loudly on
every channel**.
- **The count of recovery codes left is visible**, and the self-check warns when three or fewer remain.
- **The last resort:** root on the control node (at it, or by its SSH key) runs `factor enrol
--break-glass` there, locally. The controller refuses it over the bus. It is recorded and announced as
break-glass.
- **The controller refuses to enable any tier that requires P2 until recovery codes exist.**
### 9. First holders
- **Telegram** is the first holder of both seats and the away channel: its own module, kind
`telegram`, long polling, buttons, replies, codes by a forced reply, a linking verb; placed where no
agent runs as the operator, on which its `verified-sender` depends.
- **The desk** holds both seats as kind `desktop`, through the machine's `node-notifier` (gaining
actions) and `node-launcher` (a prompt for text, numbers, dates and codes). It carries every ordinary
ask with no account anywhere.
- **The watcher's watcher stays outside the seats**, with its own bot, on a machine that is not the
control node. An outside dead-man service is pinged by the self-check and by the watcher.
### 10. The order of work
**Fix D1–D4 of the built Telegram code first**, before the channel is configured; then the desk's
actions; then the seats and the router, with the operator's identities, untrusted marking and
references; then asks, the responder and addressing; then authorising with TOTP and its recovery; then
Telegram live. The
phases and their owners are in [to-be 46](../03-DESIGN/01-to-be/46-the-conversation-with-the-operator.md).
## Consequences
- **To-be 45 §5 is amended**: the operator answers. ADR 0227's minimal form is the first step of this
design, not reversed; ADR 0227 left this to 028's graduation and carries a dated note pointing here.
- **The catalogue gains a second sort of bench**, `kind` and `capabilities` on a claim, two seats in
the closed set ([ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md)), and the
vocabulary with its contract tests.
- **The shared library must publish on a seat's event subjects** (to-be 32 §1's open gap), and the
permission model must grant it; until then the seats' events cannot be emitted as designed.
- **`node-notifier.send` gains actions**, and the screen-lock holder emits lock and idle changes.
- **The controller grows**: the `authorises` field, three verbs, refusals on direct calls, the
hand-act fields, the operator's identities, the TOTP seed and enrolled keys as its own state, and a
self-check probe for the away channel.
- **The operator's identities, the factor and its recovery codes are the controller's state**, and the
controller gains `factor enrol`, `factor recover` and a local-only break-glass enrolment.
- **The router gains a register of addressable agents, the responder, kept messages for agents not
running, and references with `detail`.** A message may now name a machine; it still may not name a
domain, an address or a path.
- **Every agent and consumer of input carries a rule:** untrusted input is data, not instructions.
- **Harder:** approving at the desk now costs picking up the phone for a code. That is the price of an
X11 desk shared with agents; it falls if agents run under an account of their own on a display that
isolates clients, which is noted, not decided.
- **A residual risk is accepted:** on X11 an agent can read a code as it is typed at the desk and race
to use it. Each code is accepted once and only within its step; the session leaving X11 closes it.
- **Telegram sees the words**, held to the content rule, and never a factor's secret. If the operator's
Telegram account is taken, approve is reachable (announced, reversible), destroy is not without the
code.
- **The predecessor wording in to-be 45** ("each a module contributing itself to the seat") is replaced
there, not here; this record's option 1 says why.
## How it is checked
| Rule | Checked by |
|---|---|
| A capability outside `channel-capabilities/1` is refused | catalogue registration test |
| Two holders claiming one kind on a kinded bench are refused; only `channel` and `intake` are kinded | catalogue registration test over the compiled seat set |
| Each declared capability has a contract test that runs in its holder's build | catalogue lint: a declared capability without its test fails the build |
| A capability failing its drill is withdrawn from routing and reported | router test with a holder whose drill fails |
| An ask goes only to channels whose capabilities satisfy it | router test |
| Context reorders but never adds a channel, and never lowers what an ask requires | router test: the same ask in every context yields subsets of the same qualified set |
| An authorising ask never defaults | router test |
| A cancelled or answered ask's other copies are edited | router test against channel doubles |
| A fourth open ask from one asker is refused | router test |
| Presence is never written as history and never in a message's words | router test over its state buckets; the content rule's test |
| A desk click alone never authorises | controller test: `authorise answer` from kind `desktop` with no code is refused for approve and destroy |
| A TOTP code is verified by the controller, once, within its step | controller tests: wrong code, used code, code from two steps back, all refused |
| Destroy needs two proofs, one of them a code; a key is never required | controller tests: P1 alone refused; P1 + P2 accepted; P2 + P3 accepted; no tier's requirement names P3 |
| `authorise answer` refuses a non-intake caller, a kind below the tier, an identity not on the list, a key assertion over another ask's challenge, a stale digest, a second answer | one controller test per refusal |
| No agent authorises; authorising verbs refuse direct calls | controller test: direct `retire approve`, `cleanup delete`, `pin` (while kept) refused without a code; catalogue check that no agent module's grant names an authorising verb |
| `retire approve` approves only the set shown | controller test with `expect` differing from the current set |
| The hand-act records `via`, `requested-by`, `ask`, `proofs` | controller test reading the hand-act after an authorised action |
| Every tier is completable on the away channel | self-check probe, every run; raises a condition when not |
| A code never appears in an event | intake holder contract test (`code-factor`) |
| D1–D4 are fixed before Telegram is configured | holder tests: 4097 characters arrive cut and said; a reopening is a new message; clearing edits the newest; warnings and clearings are silent |
| Only a sender on the controller's list is the operator; `trusted` is set from the controller, never the holder | router test: an envelope a holder marks as the operator's, from an identity not on the list, is re-emitted `trusted: false` |
| An untrusted envelope never answers an ask, authorises, or reaches an agent as the operator's words | router test (`answer` refused); controller test (`authorise answer` refused); router test (no addressed operator message emitted for it) |
| Adding or removing an operator identity is a destroy ask | controller test: a change without two proofs is refused |
| Untrusted input is data, not instructions | the rule stated in every agent module's instructions and every intake consumer's definition; review, and a live drill: a mail saying "approve the retirement" changes nothing |
| An operator message reaches the addressed agent, by `@name` or thread; otherwise the responder | router test per route |
| Never silence: an unaddressed message not understood gets a reply saying how to address | router test |
| The responder only reads | catalogue check: its grant names only read verbs |
| A message to a registered agent not running is kept (bounded) and said; to an unknown name, refused with the known names | router tests |
| Machine names pass the content rule; domains, addresses, paths and secrets do not | content-rule test table |
| A dereferenced detail reaches only a `private` channel or the console; elsewhere only the label | router test: `detail` from a non-`private` holder refused; a message to Telegram carries the label only |
| A content-rule refusal names the offending part to the sender, and to no channel | router test |
| Ten recovery codes made at enrolment, shown once, stored only as hashes; each one P2, once | controller tests over enrolment and the store (no clear code in state); a used recovery code refused |
| `factor recover` re-enrols, is recorded and announced on every channel | controller test reading the hand-act and the announcement |
| The self-check warns at three or fewer recovery codes | self-check probe test |
| Break-glass enrolment runs only locally on the control node | controller test: `factor enrol --break-glass` over the bus is refused |
| No tier requiring P2 is enabled before recovery codes exist | controller test |
| End to end | live drills: an agent's question answered at the desk; the same with the desk locked, answered on the phone; an approve and a destroy on a test condition, answered on the phone with a code, hand-acts read |
## References
- [Research 028](../01-RESEARCH/028-the-meshs-output-channel/00-overview.md), above all
[06](../01-RESEARCH/028-the-meshs-output-channel/06-a-conversation-with-the-operator.md),
[07](../01-RESEARCH/028-the-meshs-output-channel/07-the-work-context-and-the-desk.md),
[08](../01-RESEARCH/028-the-meshs-output-channel/08-asks-that-authorise.md) and the proposed record in
[09](../01-RESEARCH/028-the-meshs-output-channel/09-a-proposed-decision.md).
- [ADR 0227](0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md): the
minimal output channel this record grows, and the place it left answering back.
- [ADR 0223](0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md): the first bench, and
the rule that another is a recorded decision.
- [ADR 0230](0230-a-consumer-the-mesh-stops-asking-for-is-retired-and-deleted-only-by-a-person.md):
approve what you were shown; a timer is the mesh acting alone again.
- [ADR 0208](0208-the-graphical-session-is-one-module-per-piece-on-the-meshs-seats.md): the desktop
notifier as a node seat.
- [Issue 187](../04-ISSUES/187-the-mesh-tells-nobody-when-it-stops-working/00-report.md): the class.
- [To-be 46](../03-DESIGN/01-to-be/46-the-conversation-with-the-operator.md): the design.
@@ -1,163 +0,0 @@
---
topic: what runs on it
status: accepted
date: 2026-10-06
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md
---
# 235. The bus is backed up by its own snapshot of each stream, taken under the bus module's account
## Context
[ADR 0233](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md)
made every module declare its data and derived the night's backup from it. It left one item flagged
for the operator: **the bus's streams were copied as live files.** The bus keeps everything it holds in
JetStream — every stream, and every key-value bucket, which is a stream too: conditions and their
history, calls, the hand-act log, the controller's lease, assignments, events, every module's state.
The backup holder read the store's directory as it stood while the server wrote it. A copy taken that
way can hold a block half-written, or an index from before the block it indexes, and may not restore;
nothing would say so until the day it was needed. The operator asked for a safe snapshot the same day.
What was true when this was decided, read from the bus on the control node without changing it:
- 19 streams, 141,703 messages, 117 MB: the events stream 97 MB (139,322 messages), the calls
bucket 17.6 MB, the machines' declarations 1.6 MB, the rest under 200 KB each. All on disk; none in
memory.
- The server's own way to hand a stream out whole is the **snapshot API**: a request names a subject
to deliver to; the server answers with the stream's configuration and its state, then sends an
archive in chunks, each carrying a reply subject the server waits on past an 8 MiB window, and ends
with an empty message. The server goes on taking writes throughout; nothing is paused or
reconfigured; only a second snapshot of the same stream at once is refused. The client library the
mesh uses offers no helper for it; the operator's command-line tool and the server implement it.
- **Only the controller could reach the JetStream API** ([design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md)
§3–§4): it is the one writer of stream definitions, and its grant is the whole API.
- The bus's image is the upstream server and an entrypoint; nothing in it could take a snapshot. Its
module declared no account on the bus.
- A dump is a shell command the backup holder runs as root before it copies the item the dump writes
into ([to-be 43](../03-DESIGN/01-to-be/43-backups-against-mistakes.md)); the stores' dumps run
`docker exec` into their own container. A dump can name the module's directories and nothing else —
no machine path (ADR 0112), and no reference to where a bundle is unpacked.
## Considered Options
1. **Keep copying the live files.** Rejected: it is the thing that may not restore, and a backup that
may not restore is worse than none, because it is believed.
2. **Stop the bus for a cold copy each night.** Rejected: every machine, every tool call and every
hand act depends on the one server; an outage a night to make a copy the server can give while
running is a cost with nothing bought.
3. **The controller takes the snapshots** — it already holds the whole JetStream API — on a schedule
before the night, into its own state directory. Rejected: it puts backup work in the core, whose
rule is that it does as little as it can and fails loudly (ADR 0227); the bus's protection would
depend on the controller's machine and schedule, and a dump of one module would be a file of
another; and the reading would be done under the widest authority on the bus.
4. **A dedicated snapshot principal with an issuance path of its own.** Rejected: a second way to mint
and seal a bus credential, when `module issue` already mints a module's account and seals it to the
machine.
5. **The program in the nats module's tools bundle**, run by the node's runtime. Rejected: a dump
cannot name where a bundle is unpacked, and a restore needs the server itself, which the bundle does
not carry.
6. **The program in the bus's own image, run by the backup holder's dump under the bus module's own
account, granted the snapshot API and nothing else.** Chosen.
## Decision
**1. The bus's streams are protected by a dump, and the live store is no longer copied.** The nats
module's `jetstream` item says `backup: {dump, into: snapshots}`. The dump runs, in the bus's own
container, a program built into the bus's image, `mesh-nats-snapshot`, with the module's credential on
its standard input; it writes one archive to standard output into the module's `snapshots` directory
(rebuildable), renamed into place only when whole. The holder backs up that directory. **No transition
period copies both**: the restore points of the live files taken before stay in the repository for
the rotation's two weeks, two months and half a year, so the old copies are at hand until they age
out, while the snapshot protects the bus from its first night.
**2. The module holding `mesh-broker` is the bus, and its account may snapshot the bus and do nothing
else.** When it declares an own secret named `broker`, the controller composes its user
`<machine>.<module>` with exactly: the stream names, one stream's information, the snapshot request for
any stream, the acknowledgement subjects the server puts on the chunks, and its own inbox. Nothing that
defines, changes, purges or writes a stream — the writers table (to-be 45 §1) holds that at
composition, the same check that refuses any second writer. It answers nothing. A bus module that
declares anything else to say or hear on the bus is refused by `module check`, because it would be
granted nothing. Its tools stay the node runtime's to serve. **The controller's own grant, and so the
installer's first user list, are unchanged.**
**3. A snapshot is the server's, taken gently.** One stream at a time, in name order, with a pause
between two; each chunk acknowledged as it arrives so the server's window keeps moving; a bound per
stream (ten minutes) and for the whole (thirty). A memory stream holds nothing across a restart and
cannot be snapshotted; it is listed as skipped with why. Any stream failing fails the whole night for
the bus: a partial copy where a whole one is expected is the silent failure this exists to prevent.
The archive is a tar: first a **manifest** — when it was taken and how long it took, the server's
version, and per stream its subjects, messages, bytes, first and last sequence, consumers, and the size
and SHA-256 of each file beside it — then, per stream, the server's configuration and state and the
server's own archive of it, consumers included, in the layout the operator's command-line tool also
reads.
**What a snapshot promises, as measured, not assumed:** every message up to the last sequence the
manifest gives for a stream, exactly. The server states a stream when the snapshot starts and reads its
blocks a moment later, so a stream written during its snapshot carries some of what arrived in that
moment as well — the first test against a server being written every two milliseconds restored to
sequence 6025 where the manifest said 6024, and at a hundred megabytes to 90075 where it said 90024,
with some messages of that tail present and some not. A restore ends at or after the manifest's
sequence, and one that ends before it is refused.
**4. Restoring never touches the live bus.** A stream is restored only where it does not exist, and on
the live bus every stream exists. So the same program builds a **new store beside the live one**: it
starts the bus's own server binary, from the bus's own image, on loopback, with the bus's one account,
restores every stream, holds each to the manifest, and stops it. Swapping the new store in for the
live one, with the bus stopped, is a person's act and a planned bus step (to-be 45). The steps are in
to-be 43 and the module's README.
**5. Nothing new watches it.** A failed snapshot is a failed night for the bus, and its item's last good
backup ages into the existing `backup-stale` condition, read by D13; the holder measures the snapshots
directory like any item. Each night's runtime and size are in the manifest, and the archive's size is
the holder's measurement of the item.
## Consequences
- **The bus's image changes, so the bus restarts once**: the image now carries the program (built
from a Go toolchain declared beside the server's base). That restart is a planned bus step, done by a
person, not a side effect of a push.
- **Rollout order**: the controller first (it composes the new user once the module declares an
account; until then nothing changes); then the catalogue; then `module issue nats --node <the bus's
machine>` **before** that machine's next push — a push refuses a module whose declared account was
never issued, and says so; then the push, as the planned bus step. The first night after it is the
first snapshot.
- **The size of a night**: at most the streams' bytes — 117 MB today, less once compressed. A local run
of the program at that size (105 MiB, 90,000 messages, written to throughout) took half a second for
the snapshot and a third of a second to restore the largest stream. The archive's compression is
per block, so a night's restore point shares what did not move with the night before; the events
stream ages out a week at a time, so most of it moves within the week.
- A restore is a person's work of four steps, not a verb. A verb that swaps a store would stop the bus
from a tool call; that stays a hand act until the planned bus step (to-be 45) is a verb itself.
- A stream's messages written during its snapshot may be partly there; nothing the mesh keeps depends
on the last few milliseconds of a night.
## How it is checked
| What | Checked by |
|---|---|
| a snapshot of every stream and bucket of a server being written to, restored into a fresh server and into a new store served by a third: every message to the manifest's sequence identical, deletes still deleted, buckets and an object store working, consumers where they were; writes during the snapshot all accepted | mesh-catalog `modules/nats/snapshot/live_test.go` (`TestASnapshotOfALiveBusRestoresToTheSameContent`, throwaway `nats:2.11`) |
| nothing restored over a live stream; a damaged archive refused | the same test |
| the manifest's dump, exactly, in the module's image with the module's configuration; a refused night fails and keeps the last snapshot | `modules/nats/snapshot/image_test.go` |
| the composed grant is the snapshot API and an inbox; it writes nothing (the writers table); only the broker seat's holder gets it | mesh-controller `internal/broker/snapshot_test.go`, `internal/inventory/busrecords_test.go`, the composition golden |
| that grant, composed, against a real server: a whole snapshot taken; every write and every wider subscription refused | `internal/broker/snapshot_live_test.go` (`TestTheComposedSnapshotUserCanSnapshotAndCannotWrite`) |
| a bus module declaring more on the bus is refused; the catalogue's bus is protected by the snapshot | `internal/catalogue/bus_snapshot_test.go` |
| the installer's first user list still matches the controller's | `TestTheInstallersFirstUserListIsWhatTheControllerWouldCompose` |
| live, after rollout | `node-backup.backed-up` shows the bus's item with a recent good backup; `doctor` passes D13; `nats_users` shows the bus module's user with four grants |
## References
- [ADR 0233](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md)
— its uncertain item, resolved here.
- [ADR 0214](0214-backups-guard-against-mistakes-and-stay-on-the-machine.md),
[to-be 43](../03-DESIGN/01-to-be/43-backups-against-mistakes.md) — the night, the dump, restoring
beside.
- [Design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md) §3–§4 — the controller remains the one
writer of stream definitions.
- [To-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) — the writers table; the bus
as a planned step.
- mesh-catalog `modules/nats` (`snapshot/`, `Dockerfile`, `module.json`, `README.md`); mesh-controller
`internal/broker` (`BusSnapshotGrants`), `internal/inventory/busrecords.go`,
`internal/catalogue/manifest.go`.
@@ -1,314 +0,0 @@
---
topic: the mesh
status: accepted
date: 2026-10-06
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md
---
# 236. A build is judged on its first machine and put back by something other than itself, and so it rolls out unattended
> **The mechanism changed — 2026-10-06, by [ADR 0239](0239-a-delivery-is-owned-by-the-mesh-delivery-module-and-runs-from-commit-to-delivered.md).** The gate, the rollback, the
> witness, the policy and the bus step stand as decided here. What moved: a **release plan** is no longer an
> object of its own. A plan is the *walk* of one delivery's trunk commit, and the backlog walk of §4a credits
> each build it carries to the delivery that published it. While `mesh-delivery` is held, a walk that moves
> no core module starts on that module's word, and *held* is a delivery's state, released by a person through
> `mesh-delivery`'s `release`. These are the gate rules ADR 0239 applies when a member of a delivery group fails.
## Context
**The operator said "start phase 4"** of [to-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md)
on 2026-10-06: a health definition per core component, the gate on a plan's first machine, rollback by
a witness that is not the new build, the bus as a planned step ([ADR 0227](0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md)
rule 8). The same day's hand-act log says what the missing phase costs:
- **47 pushes by hand in one day**, the cause the log records most and the one that raised
`healer-wanted` (S15). By what they did: **21 walked a new node-engine build through the four machines
one at a time**, each pushed only after the one before was looked at; **8 carried a new controller's
grant into the bus's user list**, which the controller's own rollout does not send; **2 were `push
--behind`** to deliver catalogue builds that their policy held back; 16 were steps of other work a person
was walking through.
- **Every module but five had the upgrade policy `record`** — built on a merge, sent to no machine
until a person pushed. `record` was the column's default since [ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md)
§3, and nothing told a `record` somebody chose from the default it always was. The five that rolled out
were set by hand: the controller, the catalogue, the build agent, the intrusion filter and the records
module.
- **The one-machine-first rollout ([ADR 0218](0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md)
§2) judged the first machine by "reported applied".** A build that applies and then serves no tools,
crashes after the report, or makes the machine's own conditions fire passes that test. Nothing put a
failed build back: a first machine that failed stopped the plan, and the machine stayed on the build
that failed it.
- **The witness for the core is being built beside this.** mesh-host's node-engine keeps the previous
controller and node tools bundles beside the new ones, judges the new controller by the lease and the
node tools by their answer to the services protocol's ping, puts the previous build back when either is
not healthy in its bound, and says so in every report while it stands (mesh-host PR #40). The controller
must grant it what it reads, raise what it says, and never send what it put back again.
- **A merge that deleted a module's directory made the controller ask the build seat to build it.**
The merge of 2026-10-06 that folded `public-acme` into the proxy failed its plan: "has no module.json at
modules/public-acme"; the four other modules of its tier were built and never sent.
- **The bus's planned step had a snapshot to take since [ADR 0235](0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md)**:
the bus machine's backup holder snapshots every stream through the bus module's dump.
**Checked against GENESIS.** *Anything requiring a human to notice it will be noticed late*: a person
walking a build through four machines is that sentence, every day. *Failure must be loud*: a build put
back without saying it, or never put back, is the silent failure the core exists to end. Nothing here
conflicts with GENESIS.
## Considered Options
1. **Keep `record` as the default and heal the pushes** — a healer that pushes what is behind. Rejected:
it is a push with no judgement, the cascade of [ADR 0221](0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md)
with the hand taken out; a broken build would reach every machine as fast as before, and nothing would
put it back.
2. **Roll everything out, gate nothing, and rely on the conditions to say what broke.** Rejected: rule 8
says the core is judged by health, not by "applied"; a condition after a build is everywhere says what
broke after it broke everywhere.
3. **Gate only the core, and leave modules to `record`.** Rejected: the pushes are mostly the core's, but
the catalogue's are the ones that stay behind the longest, and a module has health the controller can
read as well as the core has.
4. **A health definition per core component, a gate on every plan's first machine that judges a module by
its own health, a rollback there by the ordinary path once per build, a witness on the machine for what
cannot put itself back, and then rolling out as the default.** Chosen.
## Decision
**1. A core component's health is a definition, written as probes.** In the self-check's registry
(to-be 45 §4), run every five minutes on every machine and by the gate on a first machine:
| Probe | Component | Healthy when |
|---|---|---|
| H-controller | controller | the lease is held, renewed within its age, by a controller that says it is ready: its self-check ran, and `status` answered in full within ten seconds in that run (D9) |
| H-engine | node-engine | every machine heard from has reported its current declaration, under a node-engine build it names; on its first machine, under the new build |
| H-tools | node tools | every machine heard from that runs them has them answering the bus's discovery |
| H-bus | bus | every stream and durable consumer the mesh defines is on the bus, and a request crosses it to the machines' node tools and back |
**2. Every release plan's first machine passes a gate before the rest are sent.** ADR 0218's first
machine is judged from the moment it was sent the build: it reported the build applied; no witness on it
put the build back; no condition was raised since about the machine, or about the module on it; and the
component's health holds — a core component's definition above, or a module's own: its tools are served
on that machine where it has tools and the machine runs the node tools that serve them. **Healthy three
times, at least forty seconds apart and two minutes after the send, within ten minutes of it.** A module
on one machine is judged the same way; only then does the plan go on. A policy of *together* is not
gated: it is the module saying it must change everywhere at once.
> **The mechanism changed — 2026-10-07, by [ADR 0240](0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md).**
> What stands: the three judgings, their spacing and bound, and the points above. What moved: *a module's own
> health* is no longer only its tools served — every long-running resource of the module on that machine must
> also be stated healthy by the node-engine, a resource still starting is not yet a pass, and an unhealthy one
> is the condition `module.<module>.<machine>.unhealthy` that *no condition raised since the send* already
> holds the module on. Designed in [to-be 48](../03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md).
**3. A build that fails its gate is put back there, once, and never sent again on its own.** The plan
stops; the build is marked failed at its gate; the module's registered build goes back to the build the
first machine ran before — the newest successful build of that commit still kept ([ADR 0189](0189-the-store-keeps-what-the-records-name.md)
keeps five) — and that machine is sent it, by the ordinary send. The verdict is written before the send
and is never written over, so a controller replaced between the two does not send it twice. A build
marked failed is refused registration if its outcome is heard again, and `plans retry` refuses the plan;
a newer merge or a `rebuild` makes a new build, judged again. It is said as the condition
`build.<module>.<machine>.rolled-back` (a warning: the machine runs what it ran before) or
`core.<component>.<machine>.rolled-back` (urgent), `rollback-failed` (urgent, the operator's) when
nothing could be put back — no earlier build kept, the machine had never had the module, the send refused
— and as the event `rolled-back`. The probe DG keeps each until a newer build of the module passes.
**4. A build rolls out unattended by default.** With the gate and the rollback in place, a module's
build is sent one machine first, judged, then the rest, unless something says otherwise:
- **a person**, through `upgrade <module> roll-out|record|default`, `record` with why; kept, said with
who and why, and taken back by `default`;
- **the module**, in its manifest: `upgrade` with policy `roll`, `together` or `record`, the last two
with why;
- **its data**: a module that declares irreplaceable data ([ADR 0233](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md)),
its own or kept with a provider, records — a rollback cannot undo what a new build does to data that
cannot be had again;
- **the bus** records whatever anyone says (rule 6 below).
The store's `record` rows were the old default and become no choice; a person's `roll-out` is kept.
The catalogue keeps `record` where a module says why — the network path a rollback could not cross, and
the providers every consumer on a machine drops with, among them the two holding the photos. The
resulting policy of every module the mesh held on 2026-10-06:
| Policy | Modules | Why |
|---|---|---|
| record — the bus | nats | a planned step a person starts (rule 6) |
| record — irreplaceable data | plex, photos | the media library; the photos and their albums (the operator's ranking, ADR 0233) |
| record — the network path | dnsmasq, nftables, networkmanager, systemd-networkd, sshd | a build that cuts a machine off from the bus cannot be put back from outside it; sshd is the way in when the mesh cannot reach it, which no gate sees |
| record — a provider whose restart costs | postgres, mssql, mongodb, minio, keycloak | every consumer on the machine drops with it, a new version may change its data's format in place, and minio and mongodb hold the photos |
| roll — a person's choice, kept | mesh-controller, mesh-catalog, build-agent, fail2ban, records | — |
| roll — the default | every other module: 111 of the 129, ten of them on no machine, the node-engine and the node tools among them | gated on the first machine; the node-engine and node tools also witnessed on the machine |
The private network itself is provided by the controller and moves only with it; the modules that run on
no machine move nothing.
**4a. No build reaches a machine without a gate — the backlog included.** A send carries a machine's
whole declaration, so a plan sending one module, a cascade, a healer's resend or a whole-mesh push would
carry every other build waiting there. Under the old default builds were registered and sent nowhere;
on the day the default becomes `roll`, the next send of anything would restart them all at once, on every
machine. So:
- **a gated send carries everything waiting on its machine, and its gate judges all of it** — a plan's
first machine, a release plan's machine; a pass is each build's verdict, a failure puts back what was
found wanting, on the machines that were sent it;
- a module a send exists for may move where it goes — a policy of *together*, a rollback; a build that
passed a gate on one machine may go to the others;
- a person's send naming a machine (`push <node>`, the bus step) carries what it carries;
- **every other send is refused, or leaves the machine, while a build no gate has seen waits there** —
a plan's "rest", a cascade, a healer's, a whole-mesh push, the bus's user list carried — said with what
waits and the remedy;
- **a rebuild that made the same artifacts from the same manifest is no move.**
> **Progressive insight — 2026-10-06.** This rule assumed that a rebuild changing nothing makes the same
> artifacts. That holds for an archive or a bundle, which the builder packs deterministically. It does
> not hold for an image: every build of an unchanged source makes a new image digest. So a catalogue
> merge that never touched the bus rebuilt it, the new digest read as a new bus build, and every send
> to the control node was refused until a planned bus upgrade
> ([issue 280](../04-ISSUES/280-a-rebuild-of-an-unchanged-source-was-read-as-a-new-bus/00-report.md)).
> The rule said "a rebuild that made the same artifacts from the same manifest". It now also covers
> **a rebuild made from the same source**: the module's tree at the commit, the trees of the contexts
> it read, and its bases and toolchains by digest, hashed by the builder as the build's source
> fingerprint. Such a rebuild is registered with the artifacts of the build it repeats, so no machine
> is sent a new digest. Identical artifacts remain the second way to be no move, and the only way for
> a build the registry decides. The decision stands: a rebuild that changes nothing is no move. Only
> the test for "changes nothing" was wrong.
**The release plan walks what waits.** Whenever builds no gate has seen wait on machines and no plan that
has started walks them, the mesh opens a release plan: every such machine heard from, **one at a time,
the control node last**, each sent everything waiting there and judged by the gate before the next is
sent. A machine not heard from when its turn comes is left. A build asked outside a plan (a `rebuild`)
waits for it too, instead of being sent one machine after another unjudged. **A release plan that fails
puts back what failed and stops; the next one opens only when a person says `upgrade release-backlog
--why`**, and until then `release-held` says what waits. `upgrade backlog` lists it, read-only.
Chosen over holding the whole backlog for a person's release (safer by one human glance, but every
merge after the switch would then stall behind it) and over waves of a few modules (a send cannot carry
part of a declaration — ADR 0221's second option — so the bound that can be kept is one machine at a
time, which is the one kept). Measured on 2026-10-06 at the switch: the catalogue merge that rebuilt
103 modules for a change to the build agent made **88 of them byte-identical** to the builds before
(no move) and 15 different; every machine had already been sent all of them by hand that evening, so
**no build waited on any of the four machines** when this was decided.
**5. The controller's rollback is the node-engine's on its machine; the contract is written once on
each side.** mesh-controller `internal/lease/witness.go` and mesh-host `internal/witness/contract.go`,
held field for field:
- **What the host reads.** A direct get of the lease bucket's one key, `holder`: the lease's holder as
the controller writes it (instance, host, epoch, taken, renewed; anything else ignored). The node
principal of a machine assigned the controller is granted that one subject; every machine's node
principal is granted the ping of its own node tools. Reads only: the writers table still refuses any
write.
- **When it is healthy.** The holder's host is the machine, it took the key at or after the moment the
host started the new build (two seconds of skew), and renewed it within the key's fifteen seconds:
asked every five seconds, within sixty of the start. The node tools: they answer the ping within five
seconds, within sixty of the start.
- **What it starts.** The previous bundle it kept, as the previous declaration ran it.
- **What it says.** `rollbacks` on every report while the verdict stands — component, from, to, outcome,
why, when — and `witness` with the contract's version. The controller raises
`core.<component>.<machine>.<outcome>`: urgent for rolled-back, not-reversible, restore-failed and
halted; a warning for nothing-to-restore and unwitnessed; cleared by the first report without it. On a
first machine, a verdict made since the send fails the gate, so the build is marked and the registered
build put back.
- **What the controller adds.** It writes in the lease's value whether it is ready (rule 1), which the
host does not read and the gate does: the host rolls back a controller that never holds the lease, the
gate one that holds it and never becomes ready, by sending the previous build, which the host applies
as any declaration. A process's `witness` and `not-reversible` are the host's to read; the controller
sends neither yet — the host's defaults by name are the two witnesses above — and sends them only to a
machine whose report carries `witness`.
- **And the grant that follows the controller.** Once a new controller passes its gate, it sends the
machine holding the bus the user list it composes, when that differs from the one last sent and nothing
held back would go with it: the eight pushes of the day that carried a controller's grant by hand.
**6. The bus is never rolled out; its upgrade is a step a person starts.** Its policy is `record`
whatever is said, `upgrade` refuses it a roll-out, a plan builds it and sends nothing, a cascade holds its
machine (ADR 0221), and **no send reaches its machine while a new bus build waits for it** — a push naming
it, a plan's send for another module there, a healer's, a rollback's — each refused with the remedy. On
2026-10-06 a plan's send of the catalogue to the control node carried the bus's rebuilt image with it,
and the bus restarted under every machine with nobody having asked (`record` held a cascade, not the
machine a send was for). A rebuild that made the same artifacts from the same manifest is no move. A
change of the bus's **user list** is not a restart: the bus's image reloads its server in place when the
list it is written changes, and that stays the ordinary path. The step is `bus upgrade --why …`: refused unless the person says whether the new version can
be reverted (`--reversible`, or `--irreversible` as their explicit word that it runs anyway); the bus
machine's backup holder snapshots the streams first (ADR 0235) — a person who took one by hand says where
— and only then is the machine sent; recorded as a hand act; `bus-maintenance` (the probe DB) is open
while it runs, and the step ends done when the machine reported the new bus applied and H-bus passes, or
failed after fifteen minutes, said urgent with its snapshot as the way back while the bus is not healthy.
**7. A module deleted at its source is not built.** The forge's announcer says which of a merge's files
it deleted; a module whose manifest is among them is forgotten where nothing holds it and said otherwise
— never asked to build. A build that finds no manifest at a module's path, from an announcer that does
not say which files went, marks the module deleted in its plan, which goes on.
## Consequences
- **A merge reaches every machine with no hand**, one machine first, judged for at least two minutes —
and a build that breaks its first machine is put back there and goes nowhere else. The 21 pushes that
walked a node-engine build, and the 2 that delivered held catalogue builds, are the plan's.
- **A rollout is slower by the gate**: two minutes at the least per tier that rolls out, ten at the most
before a verdict. A plan with a module on one machine waits on that machine's gate before its next
tier.
- **The controller's grant gains one verb that acts**, `node-backup now`, used only by the bus step, and
an event, `rolled-back`; the installer's first user list carries both (mesh-host).
- **A new bus build stalls every send to the bus's machine until a person runs the step.** A plan whose
first machine is that machine waits, said in the plan and, past its bound, as `stalled`. That is the
point: the bus is replaced when a person is there, not as a side effect.
- **A change to the build agent rebuilds most of the catalogue** (its modules are built on what it
builds), which is how the bus came to be rebuilt by a merge that did not touch it. Whether that tiering
is right is left open here; a rebuild that changes nothing is at least no longer a move of the bus.
> **Progressive insight — 2026-10-06.** The open question rested on a wrong cause. This consequence
> said "a change to the build agent rebuilds most of the catalogue", and the measurement in decision 4
> said "the catalogue merge that rebuilt 103 modules for a change to the build agent". That merge did
> not change the build agent. It changed a file of the catalogue's reference module, which no machine
> runs. Its definition was not in the merge, so the controller read the file as shared code and
> rebuilt everything built from the repository
> ([issue 278](../04-ISSUES/278-a-module-held-by-no-machine-was-read-as-shared-code/00-report.md)).
> The build agent stood in tier 0 only because everything else is built by it. A built-by edge
> orders a plan and never widens it, so a change to the build agent rebuilds the build agent alone.
> The controller's test checks this over the real dependency relation. The open question is answered:
> the tiering was right. What was wrong is now fixed: the forge's announcer says which directories
> hold a module at the merge commit, and only a file in none of them is shared. The measurement
> stands as a count; only its cause was wrong. Everything this record decided stands.
- **Thirteen modules still wait for a person**, each saying why in the catalogue or by its data; `status`
lists the machines behind them and `push <machine>` walks them, as before. A person who wants another
held says `upgrade <module> record --why …`.
- **What the gate cannot see.** A container that crash-loops after its compose applied is seen only
through what it breaks — its tools, a provider's failing word, the machine's own conditions — until the
node-engine reports container state, which is mesh-host's to add. A build whose migration cannot be
undone is put back all the same; marking such a build `not-reversible` for the witness is not composed
yet.
- **A merge after a release plan failed may stall** on a machine where a build waits for the person's
release: its first send carries and judges what waits there, but its "rest" waits. Said in the plan
and by `release-held`.
- **Rollout order**: mesh-host first (its genesis user list, and the witness, which reads nothing the
controller does not yet grant until then); the controller second, whose migration turns the store's
default `record` rows into no choice; the catalogue third — its manifests carry `upgrade`, which a
builder older than the controller refuses.
## How it is checked
| Rule | Checked by |
|---|---|
| the gate holds back the rest | the controller's test against a real store: a build that fails its gate on the first machine is put back there, registered at the previous build, marked, said as a condition and an event, never sent to the second machine, never registered or rolled back again, and `plans retry` refuses it |
| a passing build rolls everywhere | the controller's test: three healthy judgings, then the rest sent and the plan done, its verdict kept, no condition |
| the bound | the controller's test: a first machine never healthy fails at the bound and is put back |
| the health definitions | the controller's test per component: tools not served, a condition since the send, a witness's verdict, node tools not answering, a lease held by an older controller or one not ready |
| the policy | the catalogue's test: default roll, a module's word, irreplaceable data, the bus whatever it says; a record without why refused; the store's test: the current builds read the derived policy |
| the bus is never rolled out | the controller's test: the bus records whatever it says, a person's roll-out is refused, a plan sends nothing, no send may reach its machine while a new bus build waits — a rebuild with the same artifacts and manifest excepted — the step refuses without its word on reversibility and without its snapshot, and starts with both |
| no build without a gate | the controller's test: the backlog released one machine at a time, the first judged before the second is sent, each pass kept, a rebuild with the same bytes no move, a send that judges nothing refused; a release that fails puts back what it carried on its first machine, goes no further, holds the next until a person releases it with why; a cascade does not carry a build no gate has seen, and does once one passed |
| a deleted module is not built | the controller's test: a merge deleting a module's manifest asks no build and forgets it; a build finding no manifest leaves the plan going |
| the witness's contract | the lease package's test of the host's rule; the broker's test of the grants (the ping on every machine, the lease's key where the controller runs, no write); the controller's test that a witness's verdict is its condition while reports carry it and cleared after |
| live | the next merge to the catalogue sends one machine first and the rest after its gate, with no push; `plans <id>` shows the gate's record; the next week's hand-act log has no push for a build that rolled out |
## References
- [To-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) §8 and Phase 4 — the design
this builds and amends.
- [ADR 0218](0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md),
[ADR 0221](0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md),
[ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md) — the rollout, the hold, the policy,
whose mechanisms move here.
- [ADR 0233](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md),
[ADR 0232](0232-a-binding-to-a-consumers-data-moves-only-by-a-person.md) — the data a rollout never
moves; [ADR 0235](0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md) — the snapshot the
bus step takes.
- mesh-controller, mesh-host, mesh-catalog: the branches `feat/core-upgrades-that-roll-back`;
mesh-host's witness `feat/core-upgrades-roll-back`.
@@ -1,170 +0,0 @@
---
topic: the mesh
status: accepted
date: 2026-10-06
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md
---
# 237. A change is judged against the mesh that runs, before it merges, on the build seat
> **Narrowed, not replaced — 2026-10-06, by [ADR 0238](0238-a-commit-is-the-build-at-hand-one-commit-one-change-plan-checked-off-the-trunk-and-published-only-on-it.md).**
> Decision 4's *which* and *what* moved: a pull request is checked when the mesh's module graph says it
> reaches a module — the planner's own answer, not *a repository it builds a module from into that
> branch* — and every repository of the mesh runs its own `merge-check.sh` as a second status,
> `mesh/repo-check`, a warning when it has none, instead of *the gate alone*. The gate itself is the
> build seat's, no longer each repository's script. The facts, the gate's rules, the replays, where a
> check runs and what it is given stand as decided here.
## Context
**The operator approved starting Phase 5** of [to-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md)
on 2026-10-06: the facts snapshot, the merge gate, versions tested as run, and the replays of
[ADR 0227](0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md) rule 9. The
design says what they are; building them left decisions it does not make — where a check runs on a mesh
that has no CI, what a snapshot must hold for a check to compose a machine faithfully, where a replay lives
when its incident is in one repository's logic and when it is in what the mesh runs, and how a test suite
that failed in parallel can be one a gate is allowed to read.
The evidence is research 031's class (h), *checks that pass in CI and fail live*, and the incidents of the
window since:
- **236**: a manifest passed `module check` and was refused whole by the node-engine on the first machine
it was assigned to. **263**: a real machine's name made a consumer's identity 26 characters against a
bound of 20, and the provider's whole machine could not be pushed. **262**: an answer glibc forgave was
final to musl. Each check was right about the world it was given; none was given the mesh's.
- **278**: a file of a module no machine held read as shared code, and a merge rebuilt 103 modules. Its
width could have been read before the merge; nobody was shown it.
- **The controller's suite failed when its packages ran in parallel** against one shared bus, and one test
hung under the race detector: the live tests assert, read and remove the mesh's own objects by their
fixed names, so two packages at once were one deleting what the other read. A red suite read as noise.
- **The bus tests ran a release the mesh did not run** (2.10's instructions in the tests' own comments
after the bus had moved to 2.11), and the 2.10 release that skipped messages (266) had been tested by
nobody's tests.
## Options weighed
**Where a check runs.**
1. *An outside CI.* Rejected: a second system to run, watch and keep in step, with no access to the mesh's
facts, its artifact store or its toolchains — the very things a check must be fed.
2. *The live controller composes the change.* Rejected as the gate: a change to the controller is judged by
the controller it changes, so the running one cannot judge it; and a composition run inside the serving
controller is load and risk on the control node for every pull request.
3. **The build seat, as one more kind of work on its queue** — chosen. The build machine already holds the
repositories, a container runtime, the artifact store and the mesh's Go toolchain; the controller already
asks it for work and watches every ask (S6).
**Where the facts are kept.** A key-value bucket on the bus would need a grant for every reader and a bus
that carries a document of a megabyte. **The artifact store, under `facts:latest`** — the design's choice —
is read by any machine of the mesh over plain HTTP, keeps one version, and is where the build seat already
reads and writes.
**Where a replay lives.** All in mesh-lab would put a test of the controller's planning in a repository that
cannot import it. All in their own repositories would leave the bus's release and the resolver's answers —
what the mesh *runs*, not what it wrote — with no home. **Both, by kind** — chosen.
**What the tests' bus is.** A shared bus per run needs serialised packages and still leaks between tests.
**A server per test, linked in at the release the mesh runs** — chosen.
## Decision
1. **The facts snapshot.** The controller (lease holder) composes it every ten minutes and keeps it in the
artifact store as `facts:latest` when its content moved or the one kept is a day old; the manifest it
replaced is deleted by digest, so the nightly collector takes it and the store keeps one. It holds every
machine — under a stable pseudonym of the same length as its name — with its roles in words, system,
C library, architecture, builds, reported capabilities, assignments, pins, settings and the names (never
the values) of secrets a person gave; every seat and holder; every module the mesh holds, as its manifest,
with its source and build edges; the bus's, store's and node-engine's versions as they run; and how each
machine's declaration composes today. No secret, no address: a key or value reading as a secret is
withheld, an address becomes one from a documentation range, and machine, site, account and domain
names are replaced wherever they appear. `facts`, `facts compose`, `facts export` read and write it by
hand. **S14** raises `facts-stale` past two days.
2. **The merge gate is the controller's `merge-gate`.** It raises two throwaway stores from the snapshot
through the controller's own records — the mesh as it is, and with the change's manifests in place of the
mesh's — composes every machine twice in each (the first time making the credentials a push makes), and
runs the node-engine's validator over each body. **What composes in the first and not in the second is
the change's**, named by the machine's roles and the module; what was already broken is said and fails
nothing. It also fails: a manifest the judging controller cannot read; a consumer the change leaves out
of a grant or a credential it leaves bound elsewhere; a module a machine runs removed from its source;
two definitions of one module name; a stored setting the changed definition cannot keep; and a module no
machine runs yet that the node-engine would refuse on the first machine that could run it (tried there
and taken back). It **warns** when a merge would rebuild more than twelve modules, and says when shared
code is why — the width read with the same rule the merge handler uses, the changed directories that
hold a definition read from the tree as the forge's announcer reads them.
3. **A catalogue change is judged by the controller the mesh runs** (version skew caught: a field only a
newer controller reads fails the pull request, not the registration after it); while the running one
predates the gate, by the controller's main, said. A controller change is judged by itself. A node-engine
change is judged by the running controller built with the change's validator in place of the one it
vendors.
4. **The check runs on the build seat.** The forge's announcer — the one that announces merges — announces
each new head of an open pull request (`pull.updated`) and marks it pending; the controller, for a
repository it builds a module from into that branch, asks the build seat a check: the head, and beside it
the controller the mesh runs and its main, the catalogue, the node-engine it runs, and mesh-lab. The
builder reads the snapshot, raises a throwaway store and bus **of the versions the snapshot says run**,
builds the judge, and runs the repository's own `merge-check.sh` — or the gate alone for a repository
with none — **in the mesh's Go toolchain, in a container of its own with no container runtime socket**:
a pull request is code nobody has approved yet. Then mesh-lab's replays from its main — reviewed code —
with the socket, against that bus and the change's own catalogue. The verdict is pass, warning, fail, or
**error, never read as a pass**, said by the controller as `checked`; the forge's holder sets it as the
head commit's status `mesh/merge-gate` and, when it is not a pass, comments with the check's own account.
Nothing a check does is recorded or registered. Whether the status is required to merge is the
operator's setting on the forge.
5. **A replay lives where its incident is.** One of a component's logic is a test in that repository,
named `TestReplay<issue>`, written only with what the component had before the fix, so it can be laid
over the older commit. One of what the mesh runs — the bus server's release, the resolver the catalogue
configures under musl and glibc — is in mesh-lab's `replays/`. mesh-lab's `register.go` names every
replay with its issue and fix, and `replays/cmd/prove` runs each on the commit before its fix (it must
fail, or, for a check that did not exist, not build) and on its fix (it must pass).
6. **A core issue resolves with its replay or a stated reason.** An issue opened from 2026-10-07 whose
`located-in` names a core repository — mesh-controller, mesh-host, mesh-tools, mesh-sdk, or the
catalogue's bus, forge or build-agent module — cannot be `resolved` without `replay:` (a register id) or
`replay-none:` (why none is possible). The issues the replays were written from carry theirs.
7. **The tests' bus is a server of their own, of the release the mesh runs.** The controller's suite links
the bus server in at the version go.mod pins and starts one per test; a test holds that pin to the
catalogue's bus image and, given a snapshot, to the release the mesh runs. The suite runs its packages in
parallel and under the race detector; a test reading timing, not state, is rewritten to read state. A
person may still point a run at a bus of their own with `MESH_TEST_NATS_EXTERNAL=1`.
## Consequences
- **A pull request to the core or the catalogue waits minutes for its verdict**, and the controller's
merge check runs the whole suite. That is the cost ADR 0227 accepted against the hours each of 202, 236
and 263 cost.
- **The gate is as faithful as the snapshot.** What the snapshot does not carry — a machine's adopted
state's detail, its tunnels, ports chosen by hand — a machine may compose differently in the gate than on
the mesh; then the gate says the machine composes on the mesh and not as raised, and judges it by what the
change adds. A gap that hides a real failure is found by the live probe D1, which stays.
- **The build agent pulls the toolchain, store and bus images for every check**, and mesh-lab's replays
pull the resolver and two C libraries' images from the public registry.
- **A bus upgrade is a pin moved in two places**: the catalogue's image and the controller's go.mod, which
a test holds equal. That is the rule, not a cost of it.
- **The forge's holder stays TypeScript** for this change; porting it to Go is its own piece of work.
## How it is checked
| Rule | Checked by |
|---|---|
| the snapshot carries no secret, address or name, and every machine's length | the controller's test raising a mesh with secrets, settings holding a password, an address and a machine's name, a credential made, and asserting none is in the snapshot; the scrubber's own tests |
| the snapshot is kept, one version, and said when stale | the artifact store's live test against the store's registry (put, read back by tag, the replaced one collected); S14's suppression in the generated signals test |
| the gate fails 236, 263, version skew, a removal, and says 278's width | the controller's gate tests, one per incident, against a raised mesh |
| a check runs the repository's script beside the mesh's versions, leaves nothing, and is never a pass when it cannot run | the builder's live tests against a container runtime and a registry |
| the verdict reaches the pull request, an error never as a success | the forge module's tests |
| 236, 262, 263, 266 (and 273) fail before their fix and pass after | `replays/cmd/prove` in mesh-lab |
| a core issue resolves only with a replay or a reason | `00-META/checks/cycle.py` |
| the tests' bus is the mesh's release, and the suite is deterministic | `internal/testbus`'s tests; the suite run three times in parallel under the race detector |
## References
- [ADR 0227](0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md) rule 9,
[to-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) §9 and Phase 5.
- [ADR 0225](0225-a-consumers-identity-is-bounded-by-the-provision-it-requires.md),
[ADR 0232](0232-a-binding-to-a-consumers-data-moves-only-by-a-person.md),
[ADR 0236](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md).
- Issues [236](../04-ISSUES/236-the-catalogue-check-passes-a-manifest-the-host-refuses/00-report.md),
[262](../04-ISSUES/262-an-alpine-container-could-not-find-a-machine-by-its-mesh-name/00-report.md),
[263](../04-ISSUES/263-every-consumer-pays-for-the-tightest-backends-name-limit/00-report.md),
[266](../04-ISSUES/266-a-merge-on-the-bus-was-never-handed-to-the-controller/00-report.md),
[273](../04-ISSUES/273-a-rule-for-the-resolver-moved-a-machines-databases/00-report.md),
[278](../04-ISSUES/278-a-module-held-by-no-machine-was-read-as-shared-code/00-report.md).
@@ -1,193 +0,0 @@
---
topic: the mesh
status: accepted
date: 2026-10-06
deciders: jochen
reconstructed: false
supersedes-in-part:
- 0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md
extends: 02-DECISIONS/0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md
---
# 238. A commit is the build at hand: one commit, one change plan, checked off the trunk and published only on it
> **Superseded in part — 2026-10-06, by [ADR 0239](0239-a-delivery-is-owned-by-the-mesh-delivery-module-and-runs-from-commit-to-delivered.md).** Decision 6's state machine is
> built as a **delivery**'s, owned by the module `mesh-delivery` rather than by the controller. Its states are
> renamed (`releasing` → `delivering`, `done` → `delivered`) and a delivery group sits above it. Decision 7's
> note and status link are written by the forge's holder when it hears a delivery's transition, and the link
> leads to the delivery's view on the pull request. The *change plan* is called the **delivery plan**. Decisions
> 1 to 5 stand as decided here.
## Context
[ADR 0237](0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md) put a
check before every merge and ran it on the build seat. Turning it on for every repository of the mesh, the
same day, showed four places where it asked the wrong question.
- **Which pull requests are checked was the repository's choice, not the mesh's.** The controller checked a
pull request when the mesh built a module from that repository into that branch, and ran the
repository's `merge-check.sh`, or the gate alone. A repository nothing is built from (the decision
records, the lab) was not checked at all, and its pull request kept the forge's *pending* forever. A
repository with no script ran only the gate, and nobody was told its own tests never ran.
- **The gate had its own idea of what a change touches.** The merge handler, the release planner's
what-if, the gate's *rebuild width* and the gate's composition each read a change's files in their own
way. Two of them agreed with each other; none of them was guaranteed to agree with the plan a merge
then made.
- **The rule they shared was wrong.** A file in no module's directory was read as *shared code*, and
everything built from the repository was rebuilt for it. The builder never reads such a file: it clones
the repository and builds within the module's own directory (the whole repository for a module built
from its root), plus any other repository an artifact's recipe names. Adding `merge-check.sh` at the
catalogue's root planned 103 rebuilds ([issue 280](../04-ISSUES/280-a-rebuild-of-an-unchanged-source-was-read-as-a-new-bus/00-report.md),
*Left open*). [Issue 278](../04-ISSUES/278-a-module-held-by-no-machine-was-read-as-shared-code/00-report.md)
and [issue 252](../04-ISSUES/252-a-merges-changed-modules-were-read-wrong/00-report.md) had each patched
an exception into the same rule.
- **Nothing stopped a commit off the trunk becoming a module's version.** `build --ref <branch>` of a
feature branch registered what it built, and the next push could send it. A pull request's head was
only kept out of the catalogue by the check never asking to register it — a convention.
## Considered Options
**What decides whether a pull request is checked.**
1. *The repository opts in*, with a script or a registration. Rejected: a repository that has not opted in
is exactly the one nobody is watching, and the forge said *pending* for it forever.
2. **The mesh's module graph** — chosen. The controller holds every module, the repository and directory
it is built from, and the build records' other repositories. Every pull request the forge holds is
announced, and the graph says what it reaches.
**Where what a change touches is computed.**
1. *In the gate, beside the planner.* Rejected: two computations of one fact drift, and the drift is found
when a merge does something its check did not say.
2. **In the planner, once** — chosen. One function maps a changed file onto modules, one function answers
what a merge would move and build; the merge handler, the what-if, the gate and the check all ask them.
**What a file outside every module's directory touches.**
1. *Everything built from the repository* (the shared-code rule). Rejected: no build reads it — the cost
was 103 rebuilds for a script at the root.
2. **Nothing** — chosen. A build reads its module's directory and the repositories its recipes name;
a repository a recipe names is the build record's, already followed.
**How a commit off the trunk is kept from being published.**
1. *By convention*: checks do not ask to register. Rejected: a person building a branch by hand registered
it, and a definition nobody had reviewed reached a machine ([issue 240](../04-ISSUES/240-a-dry-run-build-is-recorded-and-rolled-out/00-report.md)'s class).
2. **By refusal in code**: the build seat says which branches hold the commit it built; the controller
records and never registers a build off its module's trunk — chosen.
## Decision
1. **The commit is the build at hand, and where it sits decides what may happen.** A commit **not on the
trunk** — a pull request's head, any branch — is only *checked*: its change plan computed, the gate run
(composed machines validated, replays), its own tests run, the verdict posted. Nothing of it is
registered or sent to any machine; anything built to check it is never registrable. A commit **on the
trunk** is *published* — its builds registered as the module's version — and *pushed*, by the gated
release. **The trunk** is the branch the module follows; for a module new to the catalogue, its
repository's default branch as the forge names it. Enforced in code: the build seat reads, from its own
fresh clone, the default branch and every branch holding the commit, and says them with the outcome;
the controller records a build off its module's trunk and refuses to register it, naming the trunk —
which covers `build --ref`, `rebuild` and `replay --register` alike. A check's or a dry run's outcome
is never registered. A build seat older than this rule says nothing, and its build is registered as
before and said, so the rule can reach the mesh through the build seat it is built into.
2. **One commit, one change plan.** A change plan is computed from a diffset — a repository, the branch it
merges into, the commit at hand — by the planner: the **build plan** (the modules the merge moves
itself, their dependents in tiers; whether each is a move is its build's source fingerprint, issue
280), the **deploy plan** (per machine, what it is sent in the build plan's order; what waits there
for a person; steps that are not an ordinary send flagged: a planned bus step, a provider whose
consumers are sent again, a module keeping data per ADR 0232/0233), and the **verdict** (the composed
machines, the replays). It is one object: a pull request's check posts it as its result, the release
after the merge follows the same planner and says what differs when another merge landed meanwhile,
and `plans` answers it for any base and head. It is kept — its id, its inputs, its result — so the
pull request, the release and the operator refer to the same thing.
3. **The planner maps a change onto the modules, once.** A changed file touches exactly the modules whose
build reads it: a module's own directory, the whole repository for a module built from its root, and
a repository a recipe names (the build record's `read`). A file read by no build — at the root, in a
directory no module is in — touches nothing and is said as such. A directory the change adds a
`module.json` in, which the graph does not hold, is a new module: checked, never built by the merge.
One function does the mapping and one answers the whole reach; the merge handler, the release
planner's what-if, the merge gate and a pull request's check call them, and nothing else maps a path.
4. **The graph decides what is checked, in two layers** (replacing ADR 0237 decision 4's *for a
repository it builds a module from into that branch* and *or the gate alone*). The forge's holder
announces every pull request it holds, with the directories holding a module at its head, the files
it deletes and whether its head has a `merge-check.sh`. The controller answers every one:
- a change that reaches a module, or adds one: the build seat runs **the gate**, status
`mesh/merge-gate` — the touched manifests through `module check` (a problem the base branch already
had is said and fails nothing), every machine composed with the definitions of the modules the plan
moves or adds (not every definition in the tree), the replays — judged by the controller the mesh
runs, a controller change by itself, a node-engine change by the running controller with its
validator;
- every repository of the mesh — one a module is built from on some branch, one of the core's owner,
one the change adds a module to — gets **its own check**, status `mesh/repo-check`: its
`merge-check.sh`, in the toolchain it declares (`# mesh-check-toolchain: go|typescript`); a
repository with none is a **warning**, never a pass and never silent;
- a change that reaches nothing is a pass that says so — a fact, not a missing check — and a
repository outside the mesh is told nothing more.
5. **Required statuses are the operator's setting, applied by the forge's holder.** `mesh/merge-gate` is
required on the trunk of every repository a module is built from, the list derived from the graph;
`mesh/repo-check` as well on the core repositories. The forge module's
`gitea_branch_protection_get` / `gitea_branch_protection_set` read and set it; a new rule refuses direct
pushes.
6. **A change plan is a state machine, one table.** States and the transitions allowed between them,
each with its guard, compiled once; any other transition is refused. `proposed` (a head) → `checked`
→ `rejected` | `ready`; `ready` → `published` (guard: the commit on the trunk, its builds registered)
→ `releasing` (per machine: sent → judging → passed | failed → rolled-back) → `done` | `failed` |
`superseded` (a newer plan for the same repository and trunk) | `stopped` (a person, with why) |
`held` (waits for a person: the release backlog, a bus step, a module whose policy records) →
`releasing` (guard: a person's decision, with why). It replaces the release plans' implicit states
(open, building, judging, done, failed, superseded, stopped, held), so there is one machine, not two.
Every transition is kept with the plan and the commit, said on the bus, and appended to the commit's
note; the plan resumes from its recorded state under the lease after a restart; a state held past its
bound raises a condition, and healer H2 works from the table.
7. **The plan is attached to its commit.** The `mesh/merge-gate` status links to the kept plan and
describes it in a line ("builds gitea → the control node; no bus step; 4/4 compose"). A git note under
`refs/notes/mesh-plan` on the pull request's head carries the plan's id and the predicted summary; on
the merge commit on the trunk, the executed plan — what was built, sent where, the gates' verdicts,
rollbacks, timings — written by the mesh's own forge account after the release, never by a person's
or an agent's hand. The note writer appends: running twice adds nothing, and an outcome amends the note,
never history. `git log --notes=mesh-plan` shows each commit's plan.
## Consequences
- **A module built from a branch other than its repository's default keeps that branch as its trunk.**
Three applications are built from such a branch; a module new to the catalogue from one is refused until
the branch is the default or merged into it.
- **The build seat must be updated before the trunk rule bites.** Until the build seat built from this
change is held, it says nothing and its builds register as before, said in the controller's log.
- **The repository's own check runs code nobody has approved, and is given nothing to leak**: no container
runtime socket, no package-registry credential. A suite that needs the mesh's own packages from the
registry says those tests did not run; the Go toolchain gains a C compiler (the race detector) and Python
(the decision records' checks), and the TypeScript toolchain git.
- **A rebuild is narrower**: a file at a repository's root rebuilds nothing; the release planner and the
gate say the files no build reads.
- **What is decided here and not yet built** is listed in to-be 45's Phase 5 and stays there until it is:
the kept change plan and its comparison at release, the state machine replacing the release plans'
states, the notes, the status's link.
## How it is checked
| Rule | Checked by |
|---|---|
| a commit off its module's trunk is recorded and never registered; a check or a dry run never is | the controller's test of the publish rule (on the trunk, off it, a module following another branch, a feature branch named by hand, a check, a dry run, a build seat that says nothing); the builder's test reading the trunk and the branches holding a commit from a fresh clone |
| one function maps a changed file onto modules, and every reader asks the planner | the controller's tests that a pull request's reach is the merge's (moved, dependents, width, unread) for a root file, a module's directory and both; the merge handler's and the gate's width tests on the same rule; a review of the callers, listed in the pull request |
| a file no build reads touches nothing; a repository a recipe names reaches its packager | the merge-handler tests (a root file, a directory with no manifest, a file among the modules), the gate's width test (a root file and a README rebuild nothing), the check's test of a packaged repository |
| every pull request is answered: a gate when it reaches a module, a repository check for the mesh's repositories (a warning without a script), a pass that says so otherwise | the controller's scope tests; the builder's test of a check with nothing to run; the builder's live tests (a script run beside the mesh's versions, failing, in a toolchain the mesh does not hold, past its bound, the gate with a stand-in judge) |
| a manifest problem the base already had fails nothing | the builder's live gate test with a manifest broken on the base branch |
| the statuses and the plan reach the pull request; an error is never a success | the forge module's tests of both statuses, the plan's line and comment |
| the change plan says what each machine receives and what is not an ordinary send | the controller's change-plan test (a module's own change, the bus with its dependents and a module that waits, a root file) |
| a branch protection rule is set as asked and a new one refuses pushes | the forge module's protection tests |
| the state machine refuses a transition not in its table; notes are append-only | not yet built: a test walking the table, as the signals and healers tables are walked; the note writer's idempotence test |
## References
- [ADR 0237](0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md),
[ADR 0236](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md),
[ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md),
[ADR 0132](0132-a-seat-carries-the-tools-its-holder-must-serve.md).
- [to-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) §9 and Phase 5;
[to-be 30](../03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md).
- Issues [252](../04-ISSUES/252-a-merges-changed-modules-were-read-wrong/00-report.md),
[278](../04-ISSUES/278-a-module-held-by-no-machine-was-read-as-shared-code/00-report.md),
[280](../04-ISSUES/280-a-rebuild-of-an-unchanged-source-was-read-as-a-new-bus/00-report.md).
- mesh-controller #101, mesh-catalog #98, mesh-host #43, mesh-tools #18, mesh-tools-go #1, mesh-sdk #11,
mesh-lab #53, mesh-media-catalog #12, and the applications' own checks.
@@ -1,376 +0,0 @@
---
topic: the mesh
status: accepted
date: 2026-10-06
deciders: jochen
reconstructed: false
supersedes-in-part:
- 0238-a-commit-is-the-build-at-hand-one-commit-one-change-plan-checked-off-the-trunk-and-published-only-on-it.md
extends: 02-DECISIONS/0238-a-commit-is-the-build-at-hand-one-commit-one-change-plan-checked-off-the-trunk-and-published-only-on-it.md
---
# 239. A delivery is owned by the mesh-delivery module and runs from commit to delivered
## Context
[ADR 0238](0238-a-commit-is-the-build-at-hand-one-commit-one-change-plan-checked-off-the-trunk-and-published-only-on-it.md)
made the commit the build at hand and gave it one change plan, and decided (its decision 6) that the plan is a
state machine in one table. It did not say who owns that machine, and nothing built it. On 2026-10-06 the
operator decided that a new module owns it, named it `mesh-delivery`, and anchored it in
[to-be 10](../03-DESIGN/01-to-be/10-delivery.md), whose title is *Modules and delivery*.
What the day showed:
- **"Did my change go out?" had no owner.** The forge's holder knew the pull request, the controller knew the
merge's plan, the build seat knew the builds, the gate knew the first machine, and the hand-act log knew
the 47 pushes made by hand. A person joined those five records by hand to answer for one commit.
- **The release plan's lifecycle was implicit.** `building`, `rolling`, `done`, `failed`, `superseded`, and
the release backlog's `held` were strings in `inventory.Plan.State` and in the note beside it. No table said
which state may follow which. A plan could be closed by hand, by healer H2, by a newer merge, or by a gate.
Each did it in its own code, and none said so on the forge.
- **Nearly every change crossed repositories, in an order that was not written down.** A catalogue manifest
used a field only a newer controller parses, so the controller had to go first. The node-engine's genesis
lock had to mirror the controller's grants. A toolchain image had to be built before the builds that use it.
Each order was known to the person merging and to nothing in the mesh. A wrong order failed at
registration (version skew) or on a machine.
- **The controller is the one component that cannot be judged by itself.** Every orchestration put in it
adds to what has to be right before anything can be repaired. The witness ([ADR 0236](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)
§5) exists because of that.
To-be 10 says *"There is no pipeline as a state machine. No stage list something can be omitted from, and
no run to lose."* That was written against a pipeline that decides what to build. It still holds: what is
built is decided by comparing source with artifacts, and nothing here changes that. What it lacked is a
record of one commit's journey through the mesh that has an owner. That record must stay durable when the
controller is replaced, and it must refuse an impossible step instead of drifting into one.
**Checked against GENESIS.** *Anything requiring a human to notice it will be noticed late*: a person joining
five records is that sentence. *Failure must be loud*: a delivery stopped half-way, with nothing saying where,
is the silent failure the core exists to end. Nothing here conflicts with GENESIS.
## Considered Options
**Who owns a delivery.**
1. *The controller, as ADR 0238 decision 6 left it implicitly.* Rejected. The controller would own the
delivery of the controller. Its states would sit in the store the controller migrates. Every new rule about
order, holds and groups would be one more thing the component that cannot be judged by itself must get
right first.
2. **A module, `mesh-delivery`, holding a mesh-scoped seat of the same name.** Chosen. One owner per thing.
The controller stays the owner of what it already owns: machine declarations, bus objects, grants,
registration and the gate's primitives. The delivery is the new module's.
**What it is called.** *Change* was rejected: it is the forge's and git's word for a diff, and it names the
input, not the journey. *Upgrade-planner* was rejected: `upgrade` already names a module's policy (ADR 0236
§4), and the thing is more than a planner. *Release* was rejected: a release plan is one stage of the journey,
the stage this decision folds in, and keeping the word would keep two concepts. *Pipeline* was rejected
because to-be 10 retires it. **Delivery** is to-be 10's word for the whole of it.
**What a delivery is.**
1. *A merge.* Rejected: a pull request is checked long before it merges, and a merge's commit is not the one
the forge shows the verdict on.
2. *A group of commits across repositories, as the unit.* Considered on the day and replaced. A group as the
only object has no one commit status, no one note and no one pull request. Everything the forge shows is
per commit.
3. **One commit in one repository, and a delivery group of deliveries one level above it (the composite).**
Chosen. A delivery is 1:1 with git and the forge: one pull request, one head commit, one status, one note.
A group holds deliveries and never groups. A lone delivery needs no group.
**How a delivery joins a group.**
1. *By timing* (merged close together). Rejected: guessing is the fault being removed.
2. *By a trailer, a field in the pull request, or a verb.* Rejected as the mechanism. Each is a second thing a
person must keep in step with the branch they already named, and a verb is a hand act on every
cross-repository change.
3. **By the pull request's head branch name, across the mesh's repositories.** Chosen.
[Playbook 07](../00-META/process/07-feature-branches.md) already requires one feature to be one branch
name in every repository it touches. The mesh reads the name the person already gave. Two or more open
pull requests with the same head branch form one group. A pull request whose branch name is used in no
other repository is a lone delivery. To keep a change out of a group, rename its branch.
**Dependencies between separately planned deliveries (`needs:`).** Dropped. A delivery that depends on another
is in the same feature and belongs in its group. A dependency on a delivery already delivered needs nothing.
A second linking mechanism would be a second graph to keep acyclic and to show.
**Where the walk across machines runs.**
1. *Ported into mesh-delivery*: the module asks for each tier's build, picks the first machine, asks for one
gated send, polls the judgement and asks for the rest. Rejected for now. The walk holds the fixes of issues
214, 219, 249, 254, 256, 280 and 281. The bootstrap rule below requires the controller to keep a walk for
its own updates and for mesh-delivery's. Porting it means two walks, and one of them would be less tested.
2. **The walk of one trunk commit across machines stays the controller's primitive, and mesh-delivery decides
when it may start, records each step, and stops it.** Chosen. Sending, the gate and the rollback are the
primitives the controller keeps. The walk is their sequence for one commit, as `push <machine>` is a
person's sequence of one. Everything above that sequence moves to mesh-delivery: whether and when, in what
order, waiting for whom, superseded by what, stopped by whom, and the record.
**Where a commit's check is triggered.**
1. *mesh-delivery hears every pull request and asks the controller to check it.* Rejected. A mesh-delivery
that is down would check nothing, including the pull request that fixes mesh-delivery, and that is a mesh
that cannot be fixed.
2. **The controller keeps answering every pull request's head itself (ADR 0238 decision 4), and mesh-delivery
records the verdict as the delivery's `checked` and asks only for what the controller cannot know: the
group's composed check.** Chosen.
**When a published delivery's builds are registered.**
1. *At the merge, as before, and the walk held back until its turn.* Rejected. A send carries a machine's
whole declaration, and a gated send carries every registered build that waits on its machine
([ADR 0236](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)
§4a). A build registered before its turn would reach the machines with the next send of anything else
there. The group's order would be broken by a send nobody made for it, and a catalogue member registered
ahead of its controller is the version skew the order exists to prevent.
2. **When its walk starts.** Chosen. The merge opens the walk and asks nothing. Its first tier is asked when
the delivery's word comes, and a later tier is asked once the earlier one runs, as before. `published` is
the trunk commit accepted with its walk opened and waiting. Its builds are registered in `delivering`.
**Where a delivery is kept.**
1. *A database from the store's provider.* Rejected. It would make mesh-delivery depend on a provider whose
policy is `record` (ADR 0236 §4) and whose restart drops every consumer, adding a second dependency to the
bus mesh-delivery cannot work without anyway. Its queries are a few dozen keys a day.
2. **Key-value state the module declares, on the bus ([ADR 0201](0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md)).**
Chosen. Current state per delivery and per group. History is the events, the git notes and a bounded list
of transitions in each delivery. The seat has one holder, so the state has one writer.
## Decision
**1. Three things, one owner.**
- **A delivery** is one commit in one repository: a pull request's head, or a commit that reached the trunk
without one. Its id is `<owner>/<repository>@<commit, twelve characters>`. It carries its **delivery plan**,
the ADR 0238 change plan under its new name: the **build plan** (the modules it moves, their dependents in
tiers, a move or not by source fingerprint), the **deploy plan** (per machine, what it receives in that
order, what waits there for a person, the steps that are not an ordinary send), and the **verdict** (the
composed machines and the replays). It also carries its transitions, its per-machine steps and, once
merged, the commit it landed on the trunk as.
- **A delivery group** is two or more deliveries with the same head branch name in the mesh's repositories.
It has one level, an order among its members, and a state derived from theirs, never set on its own.
- **mesh-delivery** is the module that owns both. It holds the mesh-scoped seat `mesh-delivery`, one holder,
in the controller's compiled seat set, and serves the seat's verbs: `deliveries`, `show`, `groups`,
`what-if`, `stop`, `release`, `recheck`, `stalled`, `close`, `table`.
**2. A delivery is a state machine, one compiled table.** Each row of the table is a transition: a from-state,
an event, a to-state and a guard. Any transition not in the table is refused and the refusal is said. A test
walks the table. Each state carries its bound and what healer H2 may do once the bound has passed.
| From | Event | To | Guard |
|---|---|---|---|
| (none) | announced | proposed | a pull request's head the forge announced |
| (none) | appeared | held | a trunk commit whose walk waits, with no pull request known: merged before the seat was held, or pushed to the trunk |
| (none) | adopted | delivering | a walk already running: at the switch, or on the controller's own path |
| proposed | checked | checked | the verdict names this commit |
| checked | accepted | ready | the gate passed or warned; the repository's own check did not fail |
| checked | refused | rejected | the gate failed or could not run |
| rejected | recheck | proposed | a person asked, or the head was announced again |
| ready | recheck | proposed | a person asked, or its group changed |
| proposed, checked, ready, rejected | new head | superseded | a newer head of the same pull request |
| proposed, checked, ready, rejected | closed | stopped | the pull request closed unmerged |
| ready | merged | published | merged on the trunk its modules follow, and its walk opened; or nothing for a walk to move |
| proposed, checked, rejected | merged unchecked | held | merged without a passing check; a verdict known with it is taken first |
| published | done | delivered | nothing for a walk to move |
| published, ready | hold | held | merged, and its group's composed check did not pass for the heads that merged, or no walk was opened for it within ten minutes |
| published | go | delivering | its walk started: let go by this owner in its group's order, by a person, or on the controller's own path |
| held | go | delivering | its walk started on a word that was not this owner's: the controller's own path, or a person's `plans go` |
| held | release | delivering | a person's decision, with why |
| delivering | done | delivered | every machine of its deploy plan passed or was left as its policy says |
| delivering | failed | failed | its walk failed: a first machine's gate (what it carried put back, ADR 0236 §3), a build, a machine |
| delivering | stop | stopped | its walk was stopped through this owner |
| published, delivering | superseded | superseded | a newer delivery to the same trunk took over its walk (ADR 0218) |
| any state not final | stop | stopped | a person, with why, or its group stopped by a member before it |
`delivered`, `failed`, `superseded` and `stopped` are final. A delivery never goes back to a state before
`published` once it has left one. A row is either *observed*, taken whenever its guard holds, or an *act*,
taken only when a person or healer H2 asks (recheck, release, stop). The observed rows are taken in the
table's order until none holds. A delivery therefore stands where its facts put it, whatever order the facts
arrived in. Inside `delivering`, each machine has its own row in a second table: `sent` → `judging` →
`passed` | `failed`, `failed` → `rolled-back`. A step may pass through states between two readings of the
walk. A move the table does not reach, such as `passed` → `failed`, is refused. Steps are read from the
walk's own record (rule 7), never guessed.
**3. The trunk rule ([ADR 0238](0238-a-commit-is-the-build-at-hand-one-commit-one-change-plan-checked-off-the-trunk-and-published-only-on-it.md)
decision 1) is the guard on `published`.** A commit off the trunk can reach `ready` and no further, and the
walk that publishes it exists only for a merge into a branch a module follows. Only a commit on the trunk is
published, and only a published delivery is delivered. A check builds nothing today. When one does, what it
builds goes under a scratch namespace of the artifact store that registration refuses.
**4. A delivery group.**
- **Membership** is the head branch name (above). A group that has started delivering is closed: a new pull
request with the same branch name is a lone delivery, and the reason is given.
- **Order** is declared and inferred. It is declared by a line `after: <repository>` in a member's pull
request description. It is inferred by the controller's planner, which holds the graph, through the verb
`delivery-order`, from four rules:
- a member that moves a module goes before a member whose modules are built by it or stand on it
(toolchains and the build agent before their dependents);
- a member that moves the controller goes before a member that changes any module's manifest (version skew:
the newer controller parses what the newer manifest says);
- a member that moves the node-engine goes before the controller's member, because the witness reads
nothing the controller does not yet grant and a controller sends nothing an older engine refuses (ADR 0236
*Rollout order*);
- otherwise, the order of the repositories' names, so the order is the same on every reading.
A declared order that contradicts an inferred one is a cycle. The group is rejected, naming the cycle.
- **Checked** means all members' heads are composed together as one future state of the mesh: one gate over
every machine with every member's definitions, judged by the group's own controller when a member moves it.
The verdict is posted on every member's head as `mesh/delivery-group`. The group is `ready` only when that
verdict passed and every member is `ready`.
- **Delivering** follows the order, member by member. The next member starts only when the one before is
`delivered`.
- **A failed member stops the group cleanly.** The members after it are `stopped`, and the reason names the
member and why. **What was already delivered stays.** Each earlier member passed its own gate on its first
machine and its own judgement on every machine. A later member's failure says nothing about it, and putting
back a build that passed would be a rollback with no verdict behind it. **The failed member's own builds are
put back at its gate, once (ADR 0236 §3)**, by the controller as today. A person releases what was stopped
once the fix is merged, which makes a new delivery.
- The group's state is derived: `checking` while any member is `proposed` or the composed check is
outstanding; `rejected` when any member or the composed check is; `ready`; `delivering`; `delivered` when
all are; `failed` or `stopped` when a member is. Nothing writes a group's state except this derivation.
**5. Every transition is said four ways, and the owner of each says it.** mesh-delivery keeps the transition
in its state before anything else, and does nothing else for it until it is kept. What it owes outside is
kept with the delivery and retried until done:
- the event `mesh-delivery.transition` (a group's change of state: `mesh-delivery.group`), with no secret and
no address;
- a line on the commit's note under `refs/notes/mesh-plan`, asked of the forge's holder (`gitea_note_append`).
The line goes on the head while the delivery is off the trunk, on the commit it landed as once it is on it,
and at its end with what was executed. The forge's holder appends it in its own repository as the forge's
own account. A line the note already holds is not appended again;
- the delivery's view, one comment on the pull request kept current (`gitea_delivery_view`);
- the status `mesh/delivery` on the head, and `mesh/delivery-group` on each member's head
(`gitea_commit_status`).
Every status links to the pull request where the view is, `mesh/merge-gate` among them. The forge's holder
also says when a pull request closes unmerged (`pull.closed`), and says the head and its statuses with a
merge.
**6. The boundary.** mesh-delivery never sends to a machine, never writes a declaration, a bus object or a
grant, and never registers a build. It asks:
- **the controller**, through the seat's verbs: `delivery-plan` (the planner's one answer for a diffset),
`delivery-order`, `delivery-check` (compose a group's heads, or check one head again), `deliver` (let one
published trunk commit's walk start), `delivery-stop` (end a walk, with why), `delivery-walks` (the walks
and their steps);
- **the build seat**, through the controller's `delivery-check`. A group's check is a check ask like any
other;
- **the forge's holder**, through its tools: the notes, the view and the statuses.
The controller keeps composition, registration, the trunk rule, sending, the gate, the rollback and the walk
of one commit's tiers across machines, and says each step of a walk as the event `plan-moved`.
**7. The walk's record is the controller's, and the delivery's state is mesh-delivery's.** A plan the
controller keeps is from now on the walk of one delivery, read by the repository and commit it walks. Its
`building` and `rolling` are the walk's progress, and its end is reported to the delivery, which decides the
delivery's state from it. **A release plan stops being a concept of its own.** The backlog walk of ADR 0236
§4a is the walk of every build that waits on its machines, and each build it carries is credited to the
delivery that published it. `plans` stays as the walk's record. `deliveries` is what a person reads.
**8. The bootstrap rule.** mesh-delivery cannot gate or deliver itself, and the mesh must never depend on it to
be repaired:
- A walk that moves **the controller, the node-engine, the node tools, the bus or mesh-delivery** runs on the
controller's built-in path: started by the merge, judged at the gate, witnessed on the machine (ADR 0236
§5), the previous build kept and switched back to. mesh-delivery records it and does not start it.
- Any other walk, **while the seat `mesh-delivery` has a holder on record**, is built and published by the
merge and then waits for mesh-delivery's `deliver` before its first send. With no holder on record, the
controller starts it itself, exactly as before this record.
- **If mesh-delivery is down**, a walk waiting for it waits. Past its bound, the stall is said as the condition
`stalled` with the remedy.
> **Progressive insight — 2026-10-07.** The first build did not do what this said. It took waiting walks out
> of the stall watchdog (S3), on the reasoning that a silent mesh-delivery is D3's `holder-silent`. That
> covers a holder that is down. It does not cover one that is up and never says go: a bug, or a delivery
> stuck in its own table. Such a walk would have waited for ever with no condition open, the silent failure
> this record exists to end. The wait is now said by the controller itself, whatever mesh-delivery says of
> itself. A new row of the signals table, **S16**, raises `plan.<walk>.waiting` once a walk has waited 30
> minutes, urgent after 4 hours, naming `plans go` as the way on. The condition's name is `waiting`, not
> `stalled`: a waiting walk is no tier late, and H2's `stalled` repairs would not fit it. The decision
> stands; only its first build was wrong. A person delivers by hand: `push <machine>`, which carries what waits there
(ADR 0236 §4a), or `plans go <plan> --why`, which gives the walk a person's word in place of
mesh-delivery's. Nothing in the controller waits for mesh-delivery to repair the controller or
mesh-delivery.
- A pull request's own check is triggered by the controller, not by mesh-delivery (above). The fix for a
broken mesh-delivery is checked, merged and delivered without it.
**9. Durable across restarts.** The state is mesh-delivery's key-value state. A restarted holder reads it back
before it takes an event. It then reads the controller's `delivery-walks`, and again every half minute, so a
walk's move it missed while it was down, or one a command made that says nothing, is found by comparison and
never lost. A pull request merged while it was down is made a delivery from the forge's word: its head and
the statuses the forge holds on it. A state held past its
bound is listed by `stalled`. The controller's self-check reads that list and raises
`delivery.<delivery>.stalled`. **Healer H2 works from the table**: it closes a stalled delivery only by a
transition the table allows for that state (`superseded` when a newer delivery took over, `delivered` when
the walk's record says done), through mesh-delivery's `close`.
**10. The switch, without a gap.** Before mesh-delivery is assigned, the controller walks every merge as it
does today. Once the seat has a holder on record, a walk that starts after that waits for its word. Walks
already open finish where they are: the holder adopts each open plan as a delivery in `delivering`, and an open
release plan's carried builds are credited to their deliveries, or listed as carried with no delivery. Only a
walk's start reads whether a holder is on record, so no walk is ever both started by the controller and
waiting for mesh-delivery. Unassigning mesh-delivery returns the mesh to the controller's own path.
## Consequences
- **"Did my change go out?" is one verb.** `deliveries` and `show <delivery>` answer it from the pull request
to the last machine. The pull request's statuses link to the same answer, and the commit's note keeps it
after the state is pruned.
- **A cross-repository change merges in an order the mesh enforces.** A member that would cause version skew
waits for the member it depends on. A group that cannot be ordered is rejected before anything merges.
- **A merge no longer sends anything until mesh-delivery says so** while it is held, except the core and
mesh-delivery itself. A down mesh-delivery means walks wait, said as a condition. A person can still deliver
by hand.
- **The controller gains six verbs, one event (`plan-moved`, in the installer's first user list too) and a
`go` for `plans`, and loses no primitive.** Its release planner's lifecycle,
supersession and hold decisions remain its code until mesh-delivery is held. After that they run only on the
built-in path. Removing them is a later decision, once mesh-delivery has delivered for a while.
- **The forge's holder gains the notes, the view and two statuses.** A note is written inside the forge's own
container, in the bare repository, as the forge's user. The forge's API reads notes and writes none, and a
push from anywhere else would need a credential to the forge that nothing else should hold.
- **A module named as its seat says its events as itself.** A consumed `<seat>.<event>` names the seat's event
only when the seat emits it. Otherwise it names the event of the module of that name (`mesh-delivery.transition`).
The controller derives subscriptions by that rule.
- **What is not built by this record**: porting the walk itself into mesh-delivery (option 1 of *where the walk
runs*); a web view
of deliveries beyond the pull request's comment; requiring `mesh/delivery-group` on the trunks, which is the
operator's setting through the forge module's protection tool; the scratch namespace, until a check builds
something.
> **Progressive insight — 2026-10-07.** *What is not built* named the self-check reading `stalled` and H2
> calling `close` (to-be 47 Phase B). Both are now built: probe D14 raises `delivery.<delivery>.stalled`, and
> H2 asks mesh-delivery's `close` for a state the table gives it, never for one that is the operator's.
> Decision 9 is unchanged.
## How it is checked
| Rule | Checked by |
|---|---|
| a waiting walk is said whatever mesh-delivery says | the signals table's generated test for S16 (inside 30 minutes nothing, past it `plan.<walk>.waiting`, cleared when the walk starts) and the test that it is urgent past four hours and names `plans go` |
| a stalled delivery is the controller's condition, healed by the table | the controller's D14 test (nothing with no holder on record, one condition per stalled delivery, the operator's where the table gives H2 nothing, nothing when the holder is down, which D3 says); its H2 test (`close` asked only for the delivery H2 may move, a refusal is no repair); the test asking a stand-in owner over a real bus, and that the controller's grant names `stalled` and `close` and nothing else of the seat |
| a transition not in the table is refused; the table is walked | mesh-delivery's test that walks every row and every pair not in the table; the machine-step table walked the same way |
| a pull request's head goes proposed → checked → ready, or rejected | mesh-delivery's test from a `pull.updated` and a `checked` |
| a merge publishes, delivers and is delivered; a failed gate fails it; a newer commit supersedes it | mesh-delivery's tests driven by the controller's `plan-moved` snapshots: done, gate failed with rollback, superseded |
| the trunk rule | mesh-delivery's test that an off-trunk commit never reaches published; the controller's publish-rule test (ADR 0238) |
| a group's membership, order, composed check and clean stop | mesh-delivery's group tests: by branch name, declared and inferred order, a cycle rejected, ready only when all are, a failed member stopping those after it and leaving those before it delivered |
| a restart resumes | mesh-delivery's test that a holder restarted over the same state resumes each delivery and reconciles a walk it missed |
| the bootstrap rule | the controller's tests: a walk moving a core module or mesh-delivery never waits; another waits only while the seat has a holder on record and asks no build while it waits, so a push carries nothing of it; `plans go` starts it with mesh-delivery down, with why, recorded as a hand act; unassigned, the mesh is back on the controller's own path |
| the inferred order | the controller's `delivery-order` test: built by, version skew, engine before controller, declared, by name the same on every reading, and a contradiction named as a cycle |
| a group is checked as one future state | the controller's gate test: a catalogue change needing a provision only another repository's change adds fails alone and passes composed with it; the builder's check test, on a container runtime, clones a group's other head and hands it to the gate, and runs no member's own script |
| a walk moved by a verb is said | the controller's test that `delivery go`, run as a command of its own, says `plan-moved` on the bus |
| every transition is persisted, said, noted and reflected | mesh-delivery's test that each transition is put before it is emitted; the forge module's note test (append-only, idempotent) and view test (statuses carry its link) |
| the seat is in the set and its holder is one | the controller's seat test; the catalogue's manifest test of mesh-delivery |
| live | the next merge with mesh-delivery held waits for `deliver`, and `deliveries` shows it from head to delivered; `git log --notes=mesh-plan` shows the note on its merge commit |
## References
- [ADR 0238](0238-a-commit-is-the-build-at-hand-one-commit-one-change-plan-checked-off-the-trunk-and-published-only-on-it.md)
decision 6, which this record gives an owner and builds;
[ADR 0236](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)
§3, §4a and §5, whose release plan becomes the delivering stage and whose gate rules decide what stays;
[ADR 0218](0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md),
[ADR 0201](0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md),
[ADR 0231](0231-a-healer-acts-on-what-observation-raised-and-only-observation-says-it-worked.md).
- [To-be 47](../03-DESIGN/01-to-be/47-delivery-from-commit-to-delivered.md), the design;
[to-be 10](../03-DESIGN/01-to-be/10-delivery.md), the delivery it carries out;
[playbook 07](../00-META/process/07-feature-branches.md), whose branch name is a group's membership.
- mesh-controller, mesh-catalog and mesh-host: the branches `feat/mesh-delivery`.
@@ -1,263 +0,0 @@
---
topic: what runs on it
status: accepted
date: 2026-10-07
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md
---
# 240. A module says how it is healthy, and the node-engine judges it
## Context
The release gate ([ADR 0236](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)
§2) judges a catalogue module on its first machine by what the mesh sees from outside: the build reported
applied, no witness put it back, no condition raised since the send about the machine or the module there,
and its tools served. None of that looks at what the module *runs*. To-be 45 names the gap in its Phase 4:
*container state in the node-engine's report, without which a container that crash-loops after its compose
applied is seen only through what it breaks.*
Measured on the live mesh for [research 032](../01-RESEARCH/032-a-module-says-how-it-is-healthy/00-overview.md)
([evidence](../01-RESEARCH/032-a-module-says-how-it-is-healthy/01-evidence.md)), 2026-10-07:
- **125 catalogue modules; 68 run something long-lived** — 49 a container that stays up, 19 a service
unit stated `running` and no such container. 50 run only their bundle in the node tools, 7 only files,
directories and packages.
- **No manifest can declare a health check, and nothing reads one.** The node-engine's report carries no
container or unit state; no probe of the self-check reads one.
- **19 of 73 long-running catalogue containers have an image that ships a check** (7 of 45 container
modules). The mesh never reads it. **Two of those were wrong in the mesh's hands**: the studio and the
flow editor read *unhealthy* while working, because their checks ask an address the program does not
bind in the mesh's configuration. Read without proof, they would have rolled back two good builds.
- **Every restart count is 0**, because a recreate loses it; the runtime's event history on the home
server is about a minute long, pushed out by its own check executions.
- **Of seven recent incidents**, liveness alone would have caught the agent server's crash loop (about a
hundred restarts, found by a person reading its log for something else); an HTTP readiness check the
web application that accepted TCP and answered nothing for eleven hours
([issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md));
only the module's own check the identity provider's refused administrator
([issue 179](../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)).
A provisioner runtime that restarted until the overlay was up and then worked
([issue 058](../04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md))
is the warning: a restart count without a start period reads churn that stops as a crash loop.
- **One sample is not a finding.** A single unanswered question raised an urgent alert nobody could read
([issue 277](../04-ISSUES/277-one-unanswered-question-was-an-urgent-alert-nobody-could-read/00-report.md));
the self-check now raises such a finding on the second look (to-be 45 §4).
The mesh now rolls a module out on its own and puts the previous build back when the first machine is not
healthy. That promise is only as good as "healthy" is, and for a module it is judged today from the outside.
**Checked against GENESIS.** *Failure must be loud* — a crash loop behind every passing check is work that
reported success and did nothing. *Anything requiring a human to notice it will be noticed late* — eleven
hours, and a hundred restarts, each found by a person. *Evidence over assertion* — "applied" is an
assertion about a declaration; "it answers" is a measurement. *The mesh notices when something is wrong
before you do* ([effect](../00-META/effect.md)). *Long-lived user services rather than an orchestrator*
([context](../00-META/context.md)) is why this record judges and reports, and restarts nothing on health:
an orchestrator's liveness restart is the part it does not take. *The mesh is a guest on a personal node*
bounds the cost: the checks' floors below. Nothing here conflicts with GENESIS.
## Considered Options
**Who runs the checks.**
1. *The container runtime's own check, read by the node-engine.* Rejected as the whole answer: containers
only — nothing for the 19 service-only modules or anything only a module's tool knows; every look is an
execution inside the container, which on the home server already pushes every lifecycle event out of
the runtime's history; an image check the module never stated is a check nobody owns, and two of 19 were
wrong. Kept as one *kind*, adopted by name and proved.
2. *The module's own tool answers "healthy", and that is the judgement.* Rejected as the judge: a bundle is
hosted by the node tools, so a tool that answers proves the bundle up, not the server; and a component
that judges itself is what [ADR 0227](0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md)
rule 8 refuses for the core. Kept as one kind, for function no endpoint shows, and only beside a check
the module does not run itself.
3. *The node-engine runs everything itself, including commands, on its own schedule.* Rejected in part: a
command that only makes sense inside the container is better timed by the runtime's own retries and
start period than by an execution per look from outside.
4. **The node-engine owns every check and every verdict, and runs each kind where it is cheapest.**
Chosen: HTTP, TCP and unit checks it makes itself; a command it hands to the runtime as that
container's check and reads; a tool it asks through the node tools. One runner and one reader per
machine, for every hosting form, and it keeps what the runtime forgets.
**Where the result goes.**
1. *Only a field of the report.* Rejected alone: a report follows an apply, so a container that goes bad at
03:00 waits for the next one.
2. *Only an event on each change.* Rejected alone: events are lost or replayed; an event is a sample, not
a state.
3. *A verb the controller calls per module when it judges.* Rejected: a pull per module per judging, and a
machine slow to answer reads as unhealthy.
4. **State in the report, its change on the bus, and a condition the controller raises on the second
look.** Chosen. It is how the core already says its own state; the gate, the self-check, the healers
([ADR 0231](0231-a-healer-acts-on-what-observation-raised-and-only-observation-says-it-worked.md)) and
the operator's conversation ([ADR 0234](0234-the-mesh-holds-a-conversation-with-its-operator.md)) all
already act on conditions.
**What is judged.**
1. *Liveness only.* Rejected: no declarations needed, but it misses issue 145 (the port was open and the
program running) and issue 179.
2. *Readiness only, where declared.* Rejected: nothing is judged until every module has declared, and
liveness alone would have caught the worst incident of the window.
3. **Liveness for every long-running resource at once, readiness where declared, the declaration required
over a migration.** Chosen: the gate means something from the first build.
**What is done with an unhealthy module.**
1. *Restart it, as an orchestrator's liveness probe does.* Rejected for this record: the evidence holds no
case where a restart would have fixed anything — the crash loop was restarting already — and a restart
hides the failure the gate is meant to see. Whether a healer restarts what stays unhealthy is left to a
record of its own under ADR 0231.
2. **Nothing restarts on health; the condition reaches a person or a healer.** Chosen.
**A provider down.** With 12 consumers of the database provision and 36 of a route:
1. *Ignore it.* Rejected: twelve conditions for one fault, and twelve gates failed for something none of
them did.
2. *Judge providers first, consumers only once their providers are healthy.* Rejected: a consumer broken on
its own is not said while its provider is down, which is when it is most needed.
3. **A consumer's check names the provision it exercises; while that provision's provider is unhealthy on
the record, the consumer's finding is held under the provider's condition and its gate waits.** Chosen.
It is how [issue 281](../04-ISSUES/281-a-tier-sent-one-module-at-a-time-blamed-a-module-for-its-machine/00-report.md)
already treats a machine-level fault: what is the machine's is never pinned on a module.
**Where the declaration sits.**
1. *One per module.* Rejected: a module of eleven containers (the mail module) could not say which is wrong.
2. **On each long-running resource**, beside the resource's other fields. Chosen.
**The migration.**
1. *Required at once.* Rejected: 68 modules to change before the next merge, and nothing judged until all
are.
2. *Optional for ever.* Rejected: 38 of 45 container modules would stay at liveness, and a field that is
optional for ever is a field half the catalogue never gets.
3. **Liveness at once; the declaration required by a date, counted down.** Chosen.
## Decision
**1. Liveness is judged for every long-running resource, with no declaration.** A container that stays up,
a process that stays up and a service stated `running` are *alive* when running and not restarted more than
once within the settle window after their grace period. The node-engine observes this itself on each tick
and keeps the restarts it counted across recreates and across its own restarts; it never relies on the
runtime's restart count or event history. A resource held still by an open maintenance step
([ADR 0189](0189-the-store-keeps-what-the-records-name.md)) is neither alive nor dead: it is said as held,
and judged again when the step ends.
**2. A module declares how each long-running resource is ready**, in a field named `health` on that
resource: one kind — the image's own check adopted by name, an HTTP request to a declared endpoint and the
status expected, a TCP connect to a declared endpoint, a command in the container, the unit's own
readiness, or one of the module's tools — with an interval (default 30 s, **not under 10 s**), a timeout
under the interval, a number of failing looks in a row before it is unhealthy (default 3, **not under 2**),
and a grace period after a start (default 60 s), in which failure does not count. The grace and the failing
looks together are at most five minutes, so a resource broken from its start is said within the gate's
bound. An endpoint is named by its `listens` name, never by a port or an address, so a check follows the
machine's ports as the endpoint does. A check of function that no endpoint shows is a module's own tool, and
only in addition to a check the module does not run itself.
**3. The node-engine runs every check and owns every verdict.** HTTP, TCP and unit checks it makes itself;
a command in the container it hands to the runtime as that container's check and reads the state; an
adopted image check it reads the same way; a tool it asks through the node tools. Nothing else on the machine
judges a module, and nothing else sets a container's check.
**4. The state goes in the report, its change on the bus, and a condition is the controller's.** Each
report carries, per module and long-running resource, a state — healthy, unhealthy, starting, held,
unknown — since when, the failing streak and the restarts counted; each change is emitted as an event, and a
state that is not healthy is said again while it lasts. The controller keeps the last state per machine,
raises `module.<module>.<machine>.unhealthy` when two consecutive statements say so, and clears it on the
first that does not. **The gate needs no new rule:** ADR 0236 §2's *a module's own health holds* now reads
the module's stated health — a judging is healthy only when every long-running resource of the module on that
machine is healthy, so a resource still starting is not yet a pass — and its *no condition raised since the
send* holds the module on this condition. Every start begins in `starting`, so a condition from before the
send clears at the new build's start and anything after it is the new build's. At the gate's bound the
build is put back, as ADR 0236 §3 says.
**5. A provider down is said once, at the provider.** A check names the provision it exercises. While
that provision's provider for this consumer is unhealthy on the record, the consumer's finding is held under
the provider's condition — listed there as waiting on it, raised as nothing of its own — and the consumer's
gate waits rather than fails. What a consumer finds while its provider is healthy is its own.
**6. Nothing is restarted for being unhealthy.** The runtime restarts what exits, as now. A condition
reaches a person through the operator's conversation, or a healer; a healer that restarts on health is its
own decision.
**7. A declaration is proved before it is trusted.** `module check` refuses a `health` field that names an
endpoint the module does not declare, an interval, timeout or count outside its bounds, or a tool the module
does not serve. A bed from mesh-lab, run by the catalogue's check on the build seat, starts every resource
whose declaration or image changed and requires its check to say healthy within its grace; an image check
adopted by name is proved the same way, because two of nineteen were wrong. This proves the check's
wiring — that it can see the program working — not the change, which the live mesh and the first machine's
gate still judge ([ADR 0149](0149-the-live-mesh-is-the-test-bed.md) stands).
**8. Every catalogue module that runs something long-lived declares one.** A catalogue-wide count of the
long-running resources without `health` may only go down. `module check` warns from this decision, and
refuses a long-running resource without `health` once the count reaches zero, or six weeks after liveness is
first judged live, whichever is first. A module running nothing long-lived — its bundle only, or files and
packages — declares none: the node tools serving its tools is its liveness, as the gate judges today.
## Consequences
- **Liveness alone, from the first build, would have caught the crash loop inside the gate's ten minutes.**
Readiness would have caught the silent web application in about a minute instead of eleven hours. The
identity provider's refused administrator needs the module's own tool, which it already has (ADR 0224 §5).
- **The node-engine grows a small scheduler, a field of its report and an event; the controller a condition
kind and its two-look rule; the gate a reading of state it already had a place for.** A machine of 45
containers spends tens of milliseconds a look reading state, and about 3 s a minute on in-container
commands at the default interval.
- **A module whose first machine was unhealthy before the send no longer passes the gate for being
unchanged in its unhealthiness** — the start resets it to `starting`, and the new build is judged on its
own.
- **What a check names becomes load-bearing.** A check naming the wrong provision hides a consumer's own
fault under its provider; the bed and the first machine's gate are what catch that.
- **Harder:** a slow starter — a module that needs more than five minutes to be ready — cannot say so
within these bounds and fails its gate. That is accepted until one exists; it would be a change to the
bound, recorded.
- **A provider's health and its consumers' failures meet in two places**: this record's condition (the
provider's resource is unhealthy) and ADR 0224's standing (the provider keeps failing a consumer). They
are different facts, both kept; the hold of rule 5 reads only this record's.
- **Data is not health.** Whether a module's data is there, measured and backed up stays
[ADR 0233](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md)'s,
watched by the self-check's D13.
## What this does not decide
- Whether a healer restarts what stays unhealthy (rule 6 leaves it to its own record, under ADR 0231).
- Health for scheduled work — whether the last scheduled run succeeded belongs to the record of the
scheduled step.
- Health of what a module's events do ([issue 276](../04-ISSUES/276-a-handler-that-did-its-work-was-offered-it-five-times/00-report.md))
— the event contract's, and the bus watchdog's.
- The core's own definitions (ADR 0236 §1, to-be 45 §8), which stand; a core component may later declare its
own through the same field.
## How it is checked
| Rule | Checked by |
|---|---|
| 1 Liveness without declaration | node-engine tests over a fake runtime and service manager: a container recreated keeps its counted restarts, and so does the engine restarted; a restart inside grace is not counted; two restarts within the settle window after grace make it unhealthy; a resource under a maintenance step is held. A mesh-lab replay of the crash loop (a container whose program exits at start) fails its gate within the bound |
| 2 The field and its bounds | `module check` refuses each out-of-range part and each endpoint named by port or address, a test per refusal; a catalogue-wide test parses every `health` field |
| 3 The engine owns the verdict | a node-engine test that a declared command becomes the container's check and nothing else sets one; a test that an HTTP check dials the endpoint's current port after a port change |
| 4 Report, event, condition, gate | controller tests: one unhealthy statement raises nothing and is listed unconfirmed; two raise; a healthy one clears; the gate fails a judging while a resource is starting or unhealthy and holds a module on this condition (the existing gate test, extended); a condition from before the send clears at the new build's start. A replay of issue 145 — a database made unreachable — raises the consumer's condition within two looks |
| 5 Said once at the provider | a controller test with one unhealthy database provider and three consumers failing: one condition, at the provider, the consumers listed as waiting; their gates wait, not fail; a consumer failing while its provider is healthy is raised on its own |
| 6 No restart on health | a node-engine test that an unhealthy container is not restarted, recreated or stopped |
| 7 Proved before trusted | the bed's run of every changed declaration in the catalogue's check; the replay of the studio's false *unhealthy* (a check that asks an address the program does not bind) fails the bed, not a machine |
| 8 Every long-running module declares | the catalogue-wide count of long-running resources without `health`, compared with the number kept in the catalogue: a merge may lower it and never raise it; after the date, `module check` refuses |
| live | every machine's report carries a state for every long-running resource; `conditions` raises `module.<module>.<machine>.unhealthy` for a module stopped on purpose and clears it when it runs |
## References
- [Research 032](../01-RESEARCH/032-a-module-says-how-it-is-healthy/00-overview.md) — the evidence, the
options and the recommendation this record takes.
- [To-be 48](../03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md), the design.
- [ADR 0236](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)
§2, whose module health this gives a content;
[ADR 0227](0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md) rules 5, 6
and 8 and [to-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) §2, §4 and §8 — the
condition, the second look, the witness that is never the component;
[ADR 0224](0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md),
[ADR 0231](0231-a-healer-acts-on-what-observation-raised-and-only-observation-says-it-worked.md),
[ADR 0233](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md),
[ADR 0234](0234-the-mesh-holds-a-conversation-with-its-operator.md),
[ADR 0189](0189-the-store-keeps-what-the-records-name.md),
[ADR 0149](0149-the-live-mesh-is-the-test-bed.md),
[ADR 0237](0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md).
- Issues 058, 145, 179, 181, 276, 277, 281.
-23
View File
@@ -194,20 +194,6 @@ python3 00-META/checks/index.py fail if stale
- **0207** — [A module depends on the node seats that apply its resources](0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md) - **0207** — [A module depends on the node seats that apply its resources](0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md)
- **0210** — [A tool's configuration is its seat holder's, and every other module extends it through the seat](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md) - **0210** — [A tool's configuration is its seat holder's, and every other module extends it through the seat](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md)
- **0212** — [A seat says what it receives, and the machine's hotkeys are a seat](0212-a-seat-says-what-it-receives-and-the-machines-hotkeys-are-a-seat.md) - **0212** — [A seat says what it receives, and the machine's hotkeys are a seat](0212-a-seat-says-what-it-receives-and-the-machines-hotkeys-are-a-seat.md)
- **0218** — [A plan sends grants before code, rolls a module out one machine first, and a newer merge takes over an older plan](0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md)
- **0219** — [The build queue is controlled through the controller and the build seat](0219-the-build-queue-is-controlled-through-the-controller-and-the-build-seat.md)
- **0221** — [A push sends no build a policy or a plan holds back, except to the machine it names](0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md)
- **0222** — [A module is told where a mesh seat's holder is reached, and the controller writes no file a seat's holder owns](0222-a-module-is-told-where-a-mesh-seats-holder-is-reached-and-the-controller-writes-no-file-a-seats-holder-owns.md)
- **0224** — [A provider that keeps failing a consumer is a problem the controller reports](0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md)
- **0227** — [The core holds nine rules, each checked, and is built to them in six phases](0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md)
- **0229** — [The core's order is a lease the store remembers, and an epoch a machine is sent once it reads one](0229-the-cores-order-is-a-lease-the-store-remembers-and-an-epoch-a-machine-is-sent-once-it-reads-one.md)
- **0230** — [A consumer the mesh stops asking for is retired, and deleted only by a person](0230-a-consumer-the-mesh-stops-asking-for-is-retired-and-deleted-only-by-a-person.md)
- **0231** — [A healer acts on what observation raised, and only observation says it worked](0231-a-healer-acts-on-what-observation-raised-and-only-observation-says-it-worked.md)
- **0234** — [The mesh holds a conversation with its operator, over channels that are seats, and an answer that performs an action is authorised by the controller](0234-the-mesh-holds-a-conversation-with-its-operator.md)
- **0236** — [A build is judged on its first machine and put back by something other than itself, and so it rolls out unattended](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)
- **0237** — [A change is judged against the mesh that runs, before it merges, on the build seat](0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md)
- **0238** — [A commit is the build at hand: one commit, one change plan, checked off the trunk and published only on it](0238-a-commit-is-the-build-at-hand-one-commit-one-change-plan-checked-off-the-trunk-and-published-only-on-it.md)
- **0239** — [A delivery is owned by the mesh-delivery module and runs from commit to delivered](0239-a-delivery-is-owned-by-the-mesh-delivery-module-and-runs-from-commit-to-delivered.md)
### Its tiers, from the bottom up ### Its tiers, from the bottom up
@@ -245,8 +231,6 @@ python3 00-META/checks/index.py fail if stale
- **0194** — [The mesh has one resolver, and every node asks it for the mesh's names](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) - **0194** — [The mesh has one resolver, and every node asks it for the mesh's names](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)
- **0196** — [A node asks the mesh's resolver first, and a public one only when it is silent](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md) - **0196** — [A node asks the mesh's resolver first, and a public one only when it is silent](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md)
- **0199** — [A module that answers names declares its zone, and a node's hosts file is one module's](0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md) - **0199** — [A module that answers names declares its zone, and a node's hosts file is one module's](0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)
- **0223** — [The mesh has two resolvers, and a machine lists only them](0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md)
- **0226** — [The private network is assigned by its own name, and the proxy names its public issuer](0226-the-private-network-is-assigned-by-its-own-name-and-the-proxy-names-its-public-issuer.md)
### What runs on them, and how it gets there ### What runs on them, and how it gets there
@@ -332,13 +316,6 @@ python3 00-META/checks/index.py fail if stale
- **0214** — [Backups guard against mistakes, stay on the machine, and are declared by the module that owns the data](0214-backups-guard-against-mistakes-and-stay-on-the-machine.md) - **0214** — [Backups guard against mistakes, stay on the machine, and are declared by the module that owns the data](0214-backups-guard-against-mistakes-and-stay-on-the-machine.md)
- **0215** — [The machine's message bus is a node seat, and it is never restarted live](0215-the-machines-message-bus-is-a-node-seat-and-is-never-restarted-live.md) - **0215** — [The machine's message bus is a node seat, and it is never restarted live](0215-the-machines-message-bus-is-a-node-seat-and-is-never-restarted-live.md)
- **0216** — [The agent's configuration is registered through its module, at three scopes, and served as one plugin](0216-the-agents-configuration-is-registered-through-its-module-at-three-scopes-and-served-as-one-plugin.md) - **0216** — [The agent's configuration is registered through its module, at three scopes, and served as one plugin](0216-the-agents-configuration-is-registered-through-its-module-at-three-scopes-and-served-as-one-plugin.md)
- **0220** — [What a machine asks needs its uplink held, and the retired resolver pieces go](0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md)
- **0225** — [A consumer's identity is bounded by the provision it requires, judged before merge, and never refuses its provider](0225-a-consumers-identity-is-bounded-by-the-provision-it-requires.md)
- **0228** — [A value given by hand lives only until its module's first good start](0228-a-value-given-by-hand-lives-only-until-its-modules-first-good-start.md)
- **0232** — [A binding to a consumer's data moves only by a person](0232-a-binding-to-a-consumers-data-moves-only-by-a-person.md)
- **0233** — [A module declares the data it holds, and the mesh protects and watches it from that declaration](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md)
- **0235** — [The bus is backed up by its own snapshot of each stream, taken under the bus module's account](0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md)
- **0240** — [A module says how it is healthy, and the node-engine judges it](0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md)
### How it is built ### How it is built
+32 -111
View File
@@ -7,22 +7,14 @@ code:
- mesh-controller internal/identity/authority.go - mesh-controller internal/identity/authority.go
- mesh-controller cmd/mesh-controller/plan.go (the names the roster publishes) - mesh-controller cmd/mesh-controller/plan.go (the names the roster publishes)
- mesh-controller internal/catalogue/zones.go (the zones a module answers, ADR 0199) - mesh-controller internal/catalogue/zones.go (the zones a module answers, ADR 0199)
- mesh-controller internal/catalogue/seats.go (mesh-dns-resolver, node-hostname, node-uplink) - mesh-controller internal/catalogue/seats.go (mesh-dns-resolver, node-hosts-file)
- mesh-controller internal/catalogue/resolve.go (one owner per path, a rendered fact's included, ADR 0223) - mesh-catalog modules/dnsmasq (the mesh's one resolver)
- mesh-controller cmd/mesh-controller/holdings.go (the holders of a replicated seat, ADR 0223) - mesh-catalog modules/resolv-conf (what a node asks)
- mesh-controller internal/catalogue/roster.go (each replicated seat's holders, for a template) - mesh-catalog modules/hosts (a node's /etc/hosts)
- mesh-catalog modules/dnsmasq (the mesh's resolvers)
- mesh-catalog modules/networkmanager, modules/systemd-networkd (what a node asks, written by its uplink's holder)
- mesh-catalog modules/route-proxy (the public issuer named by the proxy, ADR 0226)
- mesh-controller internal/overlay/generator.go (the private network's module, assigned by its own name, ADR 0226)
- mesh-catalog modules/hostname (a node's /etc/hostname and /etc/hosts)
- mesh-host internal/identity/serving.go - mesh-host internal/identity/serving.go
- mesh-host internal/apply (the service that reflects a rule set; a whole file handed to its new owner) - mesh-host internal/apply (the service that reflects a rule set)
updated: 2026-10-06 updated: 2026-10-03
decisions: decisions:
- 02-DECISIONS/0226-the-private-network-is-assigned-by-its-own-name-and-the-proxy-names-its-public-issuer.md
- 02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md
- 02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md
- 02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md - 02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md
- 02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md - 02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md
- 02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md - 02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
@@ -124,15 +116,7 @@ settings, and absent from a machine nobody gave it to.
|---|---|---|---| |---|---|---|---|
| the WireGuard one | a private network, **and the mesh's own addressing** | | *the* private network, one per node | | the WireGuard one | a private network, **and the mesh's own addressing** | | *the* private network, one per node |
| the names one | name resolution | the mesh's own addressing | | | the names one | name resolution | the mesh's own addressing | |
| ~~`networking`~~ | | both of the above | | | `networking` | | both of the above | |
> **Amended 2026-10-06, by [ADR 0226](../../02-DECISIONS/0226-the-private-network-is-assigned-by-its-own-name-and-the-proxy-names-its-public-issuer.md).** The names
> became a fact (ADR 0199), leaving `networking` a module that required one other and nothing else,
> assigned on every machine. It is retired: every machine is assigned the WireGuard module —
> `mesh-wireguard` — by its own name, and genesis does the same. Choosing another VPN is still
> assigning it instead; the claim still refuses two. A module the controller stops shipping is
> retired at its next start, and kept while any machine is assigned it. The paragraphs below record
> why the bundle was built.
**Three rather than one, because WireGuard is one VPN of several.** Naming the module after the **Three rather than one, because WireGuard is one VPN of several.** Naming the module after the
job — `networking` — and putting WireGuard inside it is the retired *flavor* idea wearing a job — `networking` — and putting WireGuard inside it is the retired *flavor* idea wearing a
@@ -314,10 +298,10 @@ expensively enough to be worth restating:
- **A node must not pin its own public name locally.** The duplicate record breaks resolution of - **A node must not pin its own public name locally.** The duplicate record breaks resolution of
that name for everything else that needs it. that name for everything else that needs it.
**What the host receives:** what to ask, not what to answer. The mesh has **two resolvers**, each **What the host receives:** what to ask, not what to answer. The mesh has **one resolver**, holding
holding every node's internal domain; a node lists both and nothing else every node's internal domain; a node asks it first and a public resolver only when it is silent
([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md), ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md),
[ADR 0223](../../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md)). [ADR 0196](../../02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md)).
**What goes away:** the `/etc/hosts` floor. It exists because a node had to reach the mesh **What goes away:** the `/etc/hosts` floor. It exists because a node had to reach the mesh
database before its own DNS existed; with [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) database before its own DNS existed; with [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)
@@ -375,84 +359,36 @@ not a list of containers.
somebody starts by hand is not the mesh's to configure, and reaching into every container on a somebody starts by hand is not the mesh's to configure, and reaching into every container on a
machine — declared or not — is what a nameserver in `resolv.conf` would be for. machine — declared or not — is what a nameserver in `resolv.conf` would be for.
### The mesh's resolvers ### One resolver for the mesh
*2026-10-03, revised 2026-10-05.* **The mesh's names live in one module, held on two machines: the *2026-10-03.* **The mesh's names live in one place: the module holding `mesh-resolver`**, a mesh-scoped
holders of `mesh-dns-resolver`**, a mesh-scoped seat that is *replicated* — held on the anchor and on seat of capacity one, placed on the node every tunnel converges on. It holds one wildcard per node —
the home server, each running the same module with the same machine list and the same zones, rendered `<node>.internal` and everything under it — and listens on the private network only. It answers the
by the controller into each. Each holds one wildcard per node — `<node>.internal` and everything mesh's names from what it holds and forwards every other name, giving the public answer.
under it — and one host record per node, listens on its private address and loopback only, answers
the mesh's names from what it holds and forwards every other name, giving the public answer
([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md),
[ADR 0223](../../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md)).
Each holder is on record, added by `seat mesh-dns-resolver --add <node>/<module>`; an assignment of
the module not on record stands beside them, eligible and silent, and a seat held once stays held
once. *Checked by the controller's resolution tests with two holders on record, with two claimants
and nothing on record, and on a seat held once.*
**Every node asks them for everything, and nothing else.** `/etc/resolv.conf` lists every holder's **Every node asks it for everything, and a public resolver only when it is silent.** The module
private address — the holder on the machine itself first if it is one, then the rest by name — with a holding `node-resolver-config` writes `/etc/resolv.conf` naming `mesh-resolver` first and a public
short timeout and two attempts, and **no public resolver**. ADR 0196 listed a public resolver second, resolver second, with a short timeout and one attempt: the C library moves to the second only when the
for the anchor being unreachable; a C library that asks every listed server at once and takes the first does not answer — the anchor or the tunnel down, a captive portal holding the tunnel back — so
first reply — musl, so every Alpine container — took the public resolver's "no such name" for a mesh public names keep resolving then, and `.internal` is never asked of a public resolver while the mesh's
name, and every build on the home server failed. With only the mesh's resolvers listed, whichever answers. Containers take the same two from their machine, the runtime copying non-loopback resolvers
answers first gives the one answer. The cost is stated: a machine that reaches no mesh resolver has no into every container, so the runtime is given no `dns` of its own
names until it does, and with the anchor down the second resolver is reachable only on its own machine
and its own LAN. Containers take the same lines from their machine, the runtime copying non-loopback
resolvers into every container, so the runtime is given no `dns` of its own
([ADR 0196](../../02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md), ([ADR 0196](../../02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md),
replacing ADR 0194's per-node `systemd-resolved` stub). *Checked by the controller's composition tests replacing ADR 0194's per-node `systemd-resolved` stub).
on both holders and on a third machine, for every uplink module, and live by each machine's
`/etc/resolv.conf` and an Alpine container on the home server resolving the anchor's name every
time.*
**The file is the uplink's holder's** ([ADR 0223](../../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md)). **No node holds a copy.** The per-node resolver, its zones file and the mesh's region of `/etc/hosts`
A network manager rewrites `/etc/resolv.conf` on every connectivity change unless it is told not to go: every resolution fault found on 2026-10-03 was a copy disagreeing with the truth — a hosts file
([ADR 0117](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md)), so the module holding
`node-uplink` — the one telling it — writes the file itself, and there is one owner for it, the one
whose program would otherwise overwrite it. Each of the catalogue's managers' modules
(NetworkManager, systemd-networkd; dhcpcd's left the catalogue with
[ADR 0226](../../02-DECISIONS/0226-the-private-network-is-assigned-by-its-own-name-and-the-proxy-names-its-public-issuer.md), held by no machine) renders the same template from the resolver's holders and
requires the mesh's resolver, so a machine is refused when nothing in the mesh resolves rather than
given a file listing nothing. The managers' own mechanisms were weighed and not used: NetworkManager's
global DNS and dhcpcd's static nameservers each write the file in their own form — their own header,
their own options line — so neither can write the mesh's file byte for byte, and dhcpcd reads its
configuration only at its next start; each manager is told to keep off the file and the module
declares it. No other module may write that path, as a file or as a rendered fact: two modules on one
node declaring one path are refused. The module that wrote the file before, its seat
`node-resolver-config` and that seat's need of the uplink beside it
([ADR 0220](../../02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md))
retire. The file changes owner in one apply on each machine: the node-engine hands a whole file to the
resource declaring its path now rather than removing it first, so a machine is never without it.
*Checked by the controller's tests that the uplink modules carry one identical template and that
nothing else in the catalogue writes the path, a resolution test refusing a second writer, and the
node-engine's handover test, in which the file is present at every step of the apply.*
**A machine's names are one seat's** ([ADR 0223](../../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md)).
The `hosts` module is renamed `hostname` and holds `node-hostname`, the seat once named
`node-hosts-file`, whose former name resolves to it. It writes `/etc/hostname` from its `hostname`
setting — with no default: what a machine calls itself is the operator's, and the mesh's name for a
machine and its own need not agree — and the machine's `127.0.1.1` line, keeping the operator's lines.
A new name takes effect at the next boot; nothing sets it live, because a graphical session's X
authority is keyed by the name the session started under. *Checked by a resolution test that a
module claiming the old name and one claiming the new are one seat on one machine, and a composition
test that `/etc/hostname` is the setting, left out naming the key when nothing sets it.*
**No node holds a copy of its own.** The two holders hold the same rendering of one roster, never a
list anyone edits. The per-node resolver, its zones file and the mesh's region of `/etc/hosts`
go, and the per-node resolver's seat with them, deleted from the set once nothing claimed it
([ADR 0220](../../02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md)): every resolution fault found on 2026-10-03 was a copy disagreeing with the truth — a hosts file
read once at start, an operator's old line beside the mesh's, a node's resolver lent to a LAN. No read once at start, an operator's old line beside the mesh's, a node's resolver lent to a LAN. No
member's resolver answers a LAN; a router pointing at one is moved first. *Checked by each node's member's resolver answers a LAN; a router pointing at one is moved first. *Checked by each node's
`/etc/resolv.conf` listing the resolver's holders and nothing else, by no node but a holder answering `/etc/resolv.conf` naming `mesh-resolver` then a public resolver, by no node but the holder answering
DNS on any address, and by the router's DHCP DNS option naming the router.* DNS on any address, and by the router's DHCP DNS option naming the router.*
**Names that are neither a node nor a route.** A module that answers names declares a zone (a **Names that are neither a node nor a route.** A module that answers names declares a zone (a
setting) and the listen that answers it; the controller hands the `mesh-dns-resolver` holder every setting) and the listen that answers it; the controller hands the `mesh-dns-resolver` holder every
zone with its module's node address and published port, and each holder forwards that zone there and zone with its module's node address and published port, and the holder forwards that zone there and
answers nothing in it itself — the lab answers `<machine>.incus` for its running scenarios this way. answers nothing in it itself — the lab answers `<machine>.incus` for its running scenarios this way.
An operator's own names, unrelated to the mesh, live in `/etc/hosts`'s kept region, held per node by An operator's own names, unrelated to the mesh, live in `/etc/hosts`'s kept region, held per node by
the `node-hostname` seat's holder and changed through its tools; the controller holds none of them the `node-hosts-file` seat's holder and changed through its tools; the controller holds none of them
([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)). ([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)).
*Checked by the holder's configuration carrying one forwarding rule per declared zone, and by a push *Checked by the holder's configuration carrying one forwarding rule per declared zone, and by a push
leaving the hosts file's operator region byte for byte.* leaving the hosts file's operator region byte for byte.*
@@ -935,17 +871,6 @@ defaults to the public authority's *production* endpoint. Two consequences, and
worse than the lab problem that found it — every certificate experiment on a real node consumes worse than the lab problem that found it — every certificate experiment on a real node consumes
production issuance quota, and a retry loop can exhaust it for a week. production issuance quota, and a retry loop can exhaust it for a week.
> **Amended 2026-10-06, by [ADR 0226](../../02-DECISIONS/0226-the-private-network-is-assigned-by-its-own-name-and-the-proxy-names-its-public-issuer.md).** Configurable in the
> proxy binary, which defaults to the authority's *staging* endpoint; fixed in the mesh. The public
> issuer was a provision, `acme-ca`, answered by a module that ran nothing and was assigned beside
> every proxy; it is the proxy module's own `acme.env` now, Let's Encrypt's production directory
> spelled exactly as the binding rendered it. The proxy names each authority's account directory after
> that spelling and the root it trusts, so changing either is a new account and every certificate
> ordered again: moving the public issuer is an edit of the module with a plan for the account, never
> an assignment. The internal issuer stays a provision, `internal-acme-ca`.
> *How it is checked:* mesh-controller `internal/catalogue/public_issuer_test.go` holds the rendered
> file byte for byte and the `trust` image's digest, against the catalogue.
**The mesh CA is not a bootstrap concern.** A joining node verifies the controller against the **The mesh CA is not a bootstrap concern.** A joining node verifies the controller against the
fingerprint in its token ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)), fingerprint in its token ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)),
so nothing needs the CA before membership. It certifies internal names afterwards, and that is so nothing needs the CA before membership. It certifies internal names afterwards, and that is
@@ -1110,14 +1035,10 @@ The list is worth having in one place, because it is most of the argument:
## Open ## Open
- **One resolver ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)).** *Correction of fact, - **One resolver ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)).** Not built: every node still runs
2026-10-05:* no node holds `node-dns-resolver` any more, and `node-dns-resolver`. The migration's four steps are in the record, in order.
[ADR 0220](../../02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md) deletes the seat. What stood here before: *"Not built: every node still runs Nor are zones or the hosts file's holder ([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)): the
`node-dns-resolver`. The migration's four steps are in the record, in order."* Nor did it workstation moves to the one resolver only once both exist, its lab and operator names depending on them.
stay true that *"the workstation moves to the one resolver only once"* zones and the hosts file's
holder ([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md))
exist: every node, the workstation included, asks the one resolver, and every node holds
`node-hosts-file`.
- ~~**What happens when the hub is down.**~~ **Resolved** by - ~~**What happens when the hub is down.**~~ **Resolved** by
[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md), together with `06`'s [ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md), together with `06`'s
+2 -28
View File
@@ -5,9 +5,8 @@ code:
- mesh-controller internal/builder - mesh-controller internal/builder
- mesh-controller cmd/mesh-controller (build, build --behind, push, status) - mesh-controller cmd/mesh-controller (build, build --behind, push, status)
- mesh-controller internal/inventory/builds.go - mesh-controller internal/inventory/builds.go
updated: 2026-10-06 updated: 2026-09-29
decisions: decisions:
- 02-DECISIONS/0239-a-delivery-is-owned-by-the-mesh-delivery-module-and-runs-from-commit-to-delivered.md
- 02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md - 02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md
- 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md - 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md
- 02-DECISIONS/0010-delivery.md - 02-DECISIONS/0010-delivery.md
@@ -101,9 +100,7 @@ That is the same shape the host uses on a machine, one layer up:
| the host | machine state | declarations | | the host | machine state | declarations |
**There is no pipeline as a state machine.** No stage list something can be omitted from, and no **There is no pipeline as a state machine.** No stage list something can be omitted from, and no
run to lose. *What* is built stays decided by this comparison. *One commit's journey* through the mesh is a run to lose.
delivery with a state table of its own: it records and orders that journey and never decides what to build
(below, *A delivery has an owner*).
### An artifact is current, or it is not ### An artifact is current, or it is not
@@ -260,26 +257,3 @@ A verification mechanism was drafted for this and withdrawn. It would have repor
would not have prevented it, and the part of it that was hard — deciding which network position to check would not have prevented it, and the part of it that was hard — deciding which network position to check
from — existed only because the rule was wrong. Whether the mesh should check that a grant works is still from — existed only because the rule was wrong. Whether the mesh should check that a grant works is still
open, in issue 145; it is not the remedy for a configuration error. open, in issue 145; it is not the remedy for a configuration error.
## A delivery has an owner
*2026-10-06, by [ADR 0239](../../02-DECISIONS/0239-a-delivery-is-owned-by-the-mesh-delivery-module-and-runs-from-commit-to-delivered.md);
designed in full in [to-be 47](47-delivery-from-commit-to-delivered.md).*
The comparison above decides what is built, and that does not change. What it never had is an owner for
**did my change go out?** for one commit: from its pull request's head through its check, the merge, its
builds and every machine's gate. Five records answered it in parts, and a person joined them.
**A delivery is one commit in one repository, and the module `mesh-delivery` owns it.** Its states are one
compiled table: proposed, checked, ready or rejected, published, delivering, held, and four final states.
A transition the table does not hold is refused. Two or more deliveries that share a head branch name are a
**delivery group**, delivered in an order the planner infers and the pull requests may declare, and checked
together as one future state of the mesh.
**This answers the first open question above.** A fit artifact does not declare itself. Off the trunk it is
only checked. On the trunk it is published, and its walk across the machines starts when its delivery says
so: by itself for the core, which cannot wait for anything to be repaired, and by `mesh-delivery` for
everything else while that module is held. A person can always deliver by hand.
The controller keeps the comparison, the planner, the gate, sending and rollback. `mesh-delivery` keeps the
record and the order, and asks.
+3 -5
View File
@@ -7,9 +7,8 @@ code:
- mesh-controller internal/catalogue/build.go - mesh-controller internal/catalogue/build.go
- mesh-controller internal/inventory/secrets.go - mesh-controller internal/inventory/secrets.go
- mesh-controller cmd/mesh-builder - mesh-controller cmd/mesh-builder
updated: 2026-10-06 updated: 2026-09-30
decisions: decisions:
- 02-DECISIONS/0226-the-private-network-is-assigned-by-its-own-name-and-the-proxy-names-its-public-issuer.md
- 02-DECISIONS/0069-a-module-is-a-repository-and-a-path.md - 02-DECISIONS/0069-a-module-is-a-repository-and-a-path.md
- 02-DECISIONS/0037-where-a-module-lives.md - 02-DECISIONS/0037-where-a-module-lives.md
- 02-DECISIONS/0009-modules-and-the-graph.md - 02-DECISIONS/0009-modules-and-the-graph.md
@@ -38,9 +37,8 @@ three things the mesh now does separately:
| turning one on for one node | **assignment**, which is per node already | | turning one on for one node | **assignment**, which is per node already |
| keeping related things together | **`requires`**, and a module with requirements and no files of its own | | keeping related things together | **`requires`**, and a module with requirements and no files of its own |
`networking` was exactly that last row: it shipped nothing, required a private network, and `networking` is exactly that last row: it ships nothing, requires a private network and name
assigning it brought one — retired by [ADR 0226](../../02-DECISIONS/0226-the-private-network-is-assigned-by-its-own-name-and-the-proxy-names-its-public-issuer.md) once one VPN was the resolution, and assigning it brings both. So the module count does not return, because the thing
only answer it ever gave, and every machine is assigned that module by its own name. So the module count does not return, because the thing
that made it return — *a module is expensive, so put several things in one* — is gone. A module that made it return — *a module is expensive, so put several things in one* — is gone. A module
here is cheap: a manifest and, usually, nothing else. here is cheap: a manifest and, usually, nothing else.
@@ -5,15 +5,12 @@ code:
- mesh-controller internal/inventory/secrets.go - mesh-controller internal/inventory/secrets.go
- mesh-controller cmd/mesh-controller/rotate.go - mesh-controller cmd/mesh-controller/rotate.go
- mesh-controller examples/postgres-provisioner - mesh-controller examples/postgres-provisioner
- mesh-controller internal/inventory/given.go updated: 2026-10-01
- mesh-controller cmd/mesh-controller/given.go
updated: 2026-10-06
decisions: decisions:
- 02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md - 02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md
- 02-DECISIONS/0009-modules-and-the-graph.md - 02-DECISIONS/0009-modules-and-the-graph.md
- 02-DECISIONS/0085-a-secret-is-a-provision.md - 02-DECISIONS/0085-a-secret-is-a-provision.md
- 02-DECISIONS/0086-a-secret-reaches-a-process-as-a-file.md - 02-DECISIONS/0086-a-secret-reaches-a-process-as-a-file.md
- 02-DECISIONS/0228-a-value-given-by-hand-lives-only-until-its-modules-first-good-start.md
--- ---
# 13 — Credentials, and moving them # 13 — Credentials, and moving them
@@ -147,35 +144,8 @@ rotates only the first: `secret rotate <node> <module> <name>` makes it anew, se
and the operator, and sends the machine, so the module starts again on it. A secret that says neither and the operator, and sends the machine, so the module starts again on it. A secret that says neither
is refused with the word to write, because a credential rotated under software that never reads it is refused with the word to write, because a credential rotated under software that never reads it
again is the fault of issue 179 made deliberately; an applied one is refused until the staged form is again is the fault of issue 179 made deliberately; an applied one is refused until the staged form is
built. `rotate` is a verb on the controller's seat with built; an accepted one is refused as ADR 0113 says. `rotate` is a verb on the controller's seat with
both shapes, so the console asks for either. A provider that shares its one credential with every both shapes, so the console asks for either. A provider that shares its one credential with every
consumer ([ADR 0158](../../02-DECISIONS/0158-a-provider-with-one-credential-shares-it-with-every-consumer.md)) consumer ([ADR 0158](../../02-DECISIONS/0158-a-provider-with-one-credential-shares-it-with-every-consumer.md))
rotates the same way, with every holder's copy remade and every holding machine sent together. *How it is checked:* the tests named in issue 180, and a rotates the same way, with every holder's copy remade and every holding machine sent together. *How it is checked:* the tests named in issue 180, and a
live rotation through the console of a secret a module reads at start. live rotation through the console of a secret a module reads at start.
### A value given by hand
*Amended 2026-10-06 by [ADR 0228](../../02-DECISIONS/0228-a-value-given-by-hand-lives-only-until-its-modules-first-good-start.md).*
Until then a value given to the mesh for an own secret was refused rotation outright, read as ADR 0113's
*"the mesh will not replace what it cannot read"*. What the mesh cannot replace is a value only an
outside party can issue; a secret the module reads at start and nobody else holds is replaced without
reading the old value. So an own secret also says who may make its value: by default the mesh, and
`"issued-by": "outside"` for a vendor's key, a bot's token or a licence.
For a secret the mesh may make — read at start, not issued outside:
- `secret rotate` replaces a given value as it replaces a made one, and records it as made. It may say
why, through the console's `rotate` as well, and the why goes to the hand-act log.
- **A value given through `secret accept` lives only until the module's first good start under the
mesh.** The signal is the machine's clean report — everything applied, nothing failed or refused — of
the declaration it was last sent, when that declaration was sent after the value was given; on an
adopted machine, after the module is taken. The controller then makes a value, seals it, sends the
machine, logs it and states the seat fact `secret-replaced`, once. A given value exists to adopt
something already running that holds it; for a fresh install `secret accept` says none is needed.
An outside party's value, a value the module applies, a value the mesh's own code accepted (a bus
account it issued), and a value given before 0228 stay as given. The first is refused with the issuer
named and its age; the last is replaced when a person asks `secret rotate`. A pair credential an operator
delivers is unchanged ([ADR 0092](../../02-DECISIONS/0092-an-operator-delivers-a-pair-credential.md)).
*How it is checked:* the tests ADR 0228 names, and a live `rotate` of a given at-start secret through
the console.
+1 -42
View File
@@ -5,10 +5,8 @@ code:
- mesh-controller cmd/mesh-builder - mesh-controller cmd/mesh-builder
- mesh-controller internal/builder - mesh-controller internal/builder
- mesh-catalog modules/build-agent - mesh-catalog modules/build-agent
updated: 2026-10-06 updated: 2026-10-04
decisions: decisions:
- 02-DECISIONS/0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md
- 02-DECISIONS/0219-the-build-queue-is-controlled-through-the-controller-and-the-build-seat.md
- 02-DECISIONS/0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md - 02-DECISIONS/0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md
- 02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md - 02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md
- 02-DECISIONS/0189-the-store-keeps-what-the-records-name.md - 02-DECISIONS/0189-the-store-keeps-what-the-records-name.md
@@ -189,7 +187,6 @@ disagrees with it.
| `accesses` | operator-owned paths it may use and must not own ([ADR 0051](../../02-DECISIONS/0051-shared-data-is-the-operators.md)) | | `accesses` | operator-owned paths it may use and must not own ([ADR 0051](../../02-DECISIONS/0051-shared-data-is-the-operators.md)) |
| `certificate` | a certificate for a name it serves | | `certificate` | a certificate for a name it serves |
| `grants` | credentials it must create for its consumers | | `grants` | credentials it must create for its consumers |
| `data` | the data it holds — its own, what it keeps for its consumers, what it keeps with a provider — each with a class and how it is protected; backups, retirement and the self-check's watch are derived from it ([ADR 0233](../../02-DECISIONS/0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md)) |
| `filtering` | rules beyond its own ports | | `filtering` | rules beyond its own ports |
| `computed` | marks a module the controller generates rather than an author writing | | `computed` | marks a module the controller generates rather than an author writing |
| `build.artifacts` | what it produces; an artifact's `context` may be a URL or a path on the git seat (`seat: git`), composed by the mesh that builds it | | `build.artifacts` | what it produces; an artifact's `context` may be a URL or a path on the git seat (`seat: git`), composed by the mesh that builds it |
@@ -349,21 +346,6 @@ scheduled step with the server held still — which is what `while-stopped` exis
keeps is still a manifest and so still referenced, and the dangerous flag is not needed once the keeps is still a manifest and so still referenced, and the dangerous flag is not needed once the
mesh is the one deciding. mesh is the one deciding.
**An archive is held by a manifest of its own** (2026-10-05,
[issue 253](../../04-ISSUES/253-the-stores-collector-would-delete-every-archive-the-mesh-keeps/00-report.md)).
The sentence above held for images and not for archives: an archive was published as a bare blob that no
manifest names, and the store's collector keeps only what a manifest names, so a nightly collection
would have removed every archive the mesh keeps, the current ones included. So:
- an archive is published with a manifest that names it and nothing else, built from the archive's digest
and size alone so it can be computed again from the record;
- the sweep makes sure every archive it keeps is held that way before it lets anything go, and lets go of
an archive by removing its manifest first;
- the reference a machine fetches is unchanged.
The collector runs as a dry run until the controller reports no kept archive unheld; only then does it
collect for real.
A machine behind by more than five builds of a module, recreating a container, cannot pull what it A machine behind by more than five builds of a module, recreating a container, cannot pull what it
was running. It is already a machine the mesh reports as behind, and the answer is the current was running. It is already a machine the mesh reports as behind, and the answer is the current
declaration. declaration.
@@ -397,26 +379,3 @@ not of the recipe: one artifact declared per target, one build each.
A component's version stops being stamped in at link time. It is unpacked into a directory named for A component's version stops being stamped in at link time. It is unpacked into a directory named for
its version, so it reads its version from its own path, and a build no longer has to know what it will its version, so it reads its version from its own path, and a build no longer has to know what it will
be called. be called.
## The build queue is controlled
*2026-10-05 — [ADR 0219](../../02-DECISIONS/0219-the-build-queue-is-controlled-through-the-controller-and-the-build-seat.md).*
What is asked of the build seat can be seen and controlled from the console. The controller holds the
queue, and each machine's build agent holds its own process.
- **The controller's verbs:**
- `queue` lists every ask, waiting, in flight or dead;
- `cancel` and `clear` take asks back;
- `rebuild` asks again under a new id;
- `replay` asks again for a past build at its commit, as a dry run unless told to register, and never
over a newer build without saying so;
- `kill`, `pause` and `resume` are passed on to the seat on the right machine.
- **The build seat's verbs, on each machine:**
- `current` says what is building here and whether this holder is paused;
- `kill` ends the build's whole process group and every container it started;
- `pause` and `resume` set a flag the holder reads before it takes the next ask, which survives its
restart.
Every action that drops work leaves a failed outcome, so a plan waiting on that build fails and says why
rather than waiting. A plan keeps the id of every build it asked for. A plan waiting on a paused seat says so and is not counted late; a failed plan can be resumed with `plans retry`, and a `rebuild` of a module a plan holds joins that plan.
+1 -32
View File
@@ -5,12 +5,8 @@ code:
- mesh-sdk src - mesh-sdk src
- mesh-tools src/broker-nats.ts (and broker-amqp.ts until the rollout) - mesh-tools src/broker-nats.ts (and broker-amqp.ts until the rollout)
- mesh-controller internal/link - mesh-controller internal/link
- mesh-controller cmd/mesh-controller (status) updated: 2026-09-26
- mesh-catalog modules/postgres, modules/keycloak (the Go provisioner loop)
updated: 2026-10-06
decisions: decisions:
- 02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md
- 02-DECISIONS/0230-a-consumer-the-mesh-stops-asking-for-is-retired-and-deleted-only-by-a-person.md
- 02-DECISIONS/0095-the-control-plane-is-the-way-to-ask-a-module.md - 02-DECISIONS/0095-the-control-plane-is-the-way-to-ask-a-module.md
- 02-DECISIONS/0106-the-bus-is-nats.md - 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0128-the-mesh-bus-is-required-not-ambient.md - 02-DECISIONS/0128-the-mesh-bus-is-required-not-ambient.md
@@ -236,32 +232,6 @@ A provider ships the provisioner that creates instances of what it offers
same provision, so a grant addressed to a node alone does not name a consumer, and withdrawing one same provision, so a grant addressed to a node alone does not name a consumer, and withdrawing one
would take another's away. would take another's away.
### A provider says which consumer it keeps failing
Whether a provision was made is known to the provider's loop alone. A consumer the loop has failed for
five minutes without one success — its create, its periodic check, or reading the secret the mesh
minted for it — is announced as `provisioner.failing`, naming the consumer, its machine, the class of
error and since when, and again every fifteen minutes while it lasts; the first success, a withdrawal,
and the first success after the provider restarts are `provisioner.recovered`. Every module that
receives contributions may publish both, derived and never declared. The controller keeps the newest
failing word per provider, machine and consumer, and `status` names each one until it recovers
([ADR 0224](../../02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md)).
### A consumer the mesh stops asking for is retired
A consumer is active, retired or deleted
([ADR 0230](../../02-DECISIONS/0230-a-consumer-the-mesh-stops-asking-for-is-retired-and-deleted-only-by-a-person.md)).
The provider's loop retires one only after the same set has gone unasked in five consecutive passes
that read the contributions file and for ten minutes; retiring disables its access reversibly and marks
it, in the provider's own backend, with when and why — or, for a provider with no way to disable, only
marks it, its access kept, and never calls its removal; asked for again, the ordinary create re-enables
it. A set of
more than three, or of more than half of those held, waits for a person. Every provider serves four
tools for this — `provisioner_retirement`, `provisioner_retire_approve`, `provisioner_retire_reject`,
`provisioner_delete` — and says `provisioner.retirement`, permitted like the two events above. The
controller's `retire` and `cleanup` verbs ask those tools on the provider's machine; nothing else deletes
a consumer.
### Checked, and it agrees (2026-09-16) ### Checked, and it agrees (2026-09-16)
This looked like the sharpest disagreement and was not one. The live wire is the contributions file This looked like the sharpest disagreement and was not one. The live wire is the contributions file
@@ -285,5 +255,4 @@ two implementations drift from while both believe they conform.
| The envelope is the envelope | An emitted event is compared header by header against a fixture; a missing required header fails, an unknown `x-` header is accepted. | | The envelope is the envelope | An emitted event is compared header by header against a fixture; a missing required header fails, an unknown `x-` header is accepted. |
| Delivery is at-least-once | A fixture delivered twice is handled once. | | Delivery is at-least-once | A fixture delivered twice is handled once. |
| A grant names a consumer | A grant fixture is read by every implementation and yields the same module and the same node. | | A grant names a consumer | A grant fixture is read by every implementation and yields the same module and the same node. |
| A provider that keeps failing a consumer says so | The loop's tests drive five minutes of failure to one `provisioner.failing` and a success to `provisioner.recovered`; the controller's tests carry it from the bus to `status`. |
| A partial SDK is legitimate | An implementation claiming the floor and events passes those suites and is listed for them; a module using tools in that language is refused at build time with the reason. | | A partial SDK is legitimate | An implementation claiming the floor and events passes those suites and is listed for them; a module using tools in that language is refused at build time with the reason. |
@@ -5,9 +5,8 @@ code:
- mesh-host internal/bootstrap - mesh-host internal/bootstrap
- mesh-host cmd/mesh-bootstrap - mesh-host cmd/mesh-bootstrap
- mesh-lab test/integration/one-node-mesh.test.ts - mesh-lab test/integration/one-node-mesh.test.ts
updated: 2026-10-06 updated: 2026-09-21
decisions: decisions:
- 02-DECISIONS/0226-the-private-network-is-assigned-by-its-own-name-and-the-proxy-names-its-public-issuer.md
- 02-DECISIONS/0067-genesis-is-a-pivot.md - 02-DECISIONS/0067-genesis-is-a-pivot.md
- 02-DECISIONS/0073-the-installer-carries-a-builder.md - 02-DECISIONS/0073-the-installer-carries-a-builder.md
- 02-DECISIONS/0071-where-genesis-gets-its-source.md - 02-DECISIONS/0071-where-genesis-gets-its-source.md
@@ -73,7 +72,7 @@ catalogue be missing from a test for weeks without anything complaining.
| 15 | the catalogue is built and run | the module graph | without it the mesh cannot say what it holds, what a change reaches, or what must be rebuilt | | 15 | the catalogue is built and run | the module graph | without it the mesh cannot say what it holds, what a change reaches, or what must be rebuilt |
| 16 | the catalogue asks for what it missed | the builds made before it existed are replayed | on a fresh mesh those are always the base, the store and the catalogue itself ([issue 050](../../04-ISSUES/050-the-catalogue-knows-nothing-built-before-it/00-report.md)) | | 16 | the catalogue asks for what it missed | the builds made before it existed are replayed | on a fresh mesh those are always the base, the store and the catalogue itself ([issue 050](../../04-ISSUES/050-the-catalogue-knows-nothing-built-before-it/00-report.md)) |
| 17 | the controller is rebuilt from its own repository | and rolled out through the module path | the moment the mesh stops depending on the installer for anything | | 17 | the controller is rebuilt from its own repository | and rolled out through the module path | the moment the mesh stops depending on the installer for anything |
| 18 | `mesh-wireguard` is assigned **and the node placed** | a private network | assigning installs the module; placing says where this machine is on it. Both, or the machine has no peers. Assigned by its own name since [ADR 0226](../../02-DECISIONS/0226-the-private-network-is-assigned-by-its-own-name-and-the-proxy-names-its-public-issuer.md) retired the `networking` bundle | | 18 | `networking` is assigned **and the node placed** | a private network, and names | assigning installs the module; placing says where this machine is on it. Both, or the names file is written empty |
| 19 | the packet filter is assigned | rules generated from what modules declared | until this, every rule the mesh computes has never been applied to anything | | 19 | the packet filter is assigned | rules generated from what modules declared | until this, every rule the mesh computes has never been applied to anything |
## Phase three — machines arrive ## Phase three — machines arrive
@@ -148,7 +147,7 @@ two.
| 5 | `builder` | — | turns source into artifacts | | 5 | `builder` | — | turns source into artifacts |
| 6 | `mesh-tools` | build inputs | the base everything with code compiles against. **Runs nowhere** | | 6 | `mesh-tools` | build inputs | the base everything with code compiles against. **Runs nowhere** |
| 7 | `mesh-catalog` | the module graph | what is held, what a change reaches, what must be rebuilt | | 7 | `mesh-catalog` | the module graph | what is held, what a change reaches, what must be rebuilt |
| 8 | `mesh-wireguard` | `private-network` | the private network, computed per machine by the controller; assigned directly ([ADR 0226](../../02-DECISIONS/0226-the-private-network-is-assigned-by-its-own-name-and-the-proxy-names-its-public-issuer.md)) | | 8 | `networking` | `private-network`, naming | requirements only — assigning it brings `mesh-wireguard` and `mesh-names` |
| 9 | `dnsmasq` + one of `resolved-split-dns` / `resolv-conf` | `wildcard-resolution` | names that actually resolve, on top of `mesh-resolver`'s data | | 9 | `dnsmasq` + one of `resolved-split-dns` / `resolv-conf` | `wildcard-resolution` | names that actually resolve, on top of `mesh-resolver`'s data |
| 10 | `step-ca` | `acme-ca` | certificates for `.internal` | | 10 | `step-ca` | `acme-ca` | certificates for `.internal` |
| 11 | `firewall` | *claims* `the-packet-filter` | rules generated from what modules declared | | 11 | `firewall` | *claims* `the-packet-filter` | rules generated from what modules declared |
+1 -9
View File
@@ -8,9 +8,8 @@ code:
- mesh-catalog modules/nats (to be written) - mesh-catalog modules/nats (to be written)
- mesh-sdk src (the protocol's NATS binding, step 3) - mesh-sdk src (the protocol's NATS binding, step 3)
- mesh-tools node-tools/internal/bus (a module's state, ADR 0201) - mesh-tools node-tools/internal/bus (a module's state, ADR 0201)
updated: 2026-10-06 updated: 2026-10-04
decisions: decisions:
- 02-DECISIONS/0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md
- 02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md - 02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md
- 02-DECISIONS/0167-a-membership-carries-what-its-module-receives-and-who-the-mesh-is.md - 02-DECISIONS/0167-a-membership-carries-what-its-module-receives-and-who-the-mesh-is.md
- 02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md - 02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md
@@ -176,13 +175,6 @@ call is a timeout the caller already handles.
Streams and consumers are objects the controller creates at genesis and asserts on start; a module Streams and consumers are objects the controller creates at genesis and asserts on start; a module
declares nothing about them. The controller is the only writer of stream definitions. declares nothing about them. The controller is the only writer of stream definitions.
*Added 2026-10-06, [ADR 0235](../../02-DECISIONS/0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md):*
one other user reads the streams whole, and writes none of them. The module that is the bus — the one
holding `mesh-broker` — has an account of its own whose only grant is the snapshot API: the stream
names, a stream's information, the snapshot request and its acknowledgements, and its own inbox. The
night's backup takes every stream through it ([to-be 43](43-backups-against-mistakes.md)). The
controller stays the only writer of stream definitions; the writers table holds that at composition.
### The store window, and what moving it into the server changes ### The store window, and what moving it into the server changes
The guarantee ([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)) is The guarantee ([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)) is
+8 -47
View File
@@ -4,21 +4,14 @@ status: in-progress
code: code:
- mesh-controller internal/catalogue/seats.go - mesh-controller internal/catalogue/seats.go
- mesh-controller internal/catalogue/resolve.go - mesh-controller internal/catalogue/resolve.go
- mesh-controller internal/catalogue/seat_dependencies.go
- mesh-controller internal/inventory/seats.go - mesh-controller internal/inventory/seats.go
- mesh-controller internal/inventory/migrations/0039-a-seat-is-held-by-one-assignment-on-record.sql - mesh-controller internal/inventory/migrations/0039-a-seat-is-held-by-one-assignment-on-record.sql
- mesh-controller internal/inventory/migrations/0062-a-replicated-seat-has-several-holders-on-record.sql
- mesh-controller cmd/mesh-controller/holdings.go
- mesh-controller cmd/mesh-controller/seats.go - mesh-controller cmd/mesh-controller/seats.go
- mesh-controller cmd/mesh-controller/source.go - mesh-controller cmd/mesh-controller/source.go
- mesh-controller internal/inventory/migrations/0032-a-source-may-live-on-a-seat.sql - mesh-controller internal/inventory/migrations/0032-a-source-may-live-on-a-seat.sql
- mesh-catalog modules/gitea/module.json - mesh-catalog modules/gitea/module.json
- mesh-controller internal/inventory/migrations/0063-a-machines-names-are-one-seats.sql updated: 2026-10-03
- mesh-controller internal/inventory/migrations/0064-the-resolver-file-is-the-uplinks.sql
updated: 2026-10-05
decisions: decisions:
- 02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md
- 02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md
- 02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md - 02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md
- 02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md - 02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
- 02-DECISIONS/0161-what-deserves-a-seat.md - 02-DECISIONS/0161-what-deserves-a-seat.md
@@ -43,7 +36,7 @@ A seat has four properties, fixed by the mesh rather than by any module:
| property | is | | property | is |
|---|---| |---|---|
| name | what a definition names and an assignment holds, and what a person reads in the list | | name | what a definition names and an assignment holds, and what a person reads in the list |
| scope | node, site or mesh: where its capacity applies. Every seat in the set has a capacity of one, so one holder per scope — except a *replicated* mesh seat, which may be held on several machines at once, one holder per machine, each on record ([ADR 0223](../../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md)). `mesh-dns-resolver` is the only one. Replicated is part of the seat's definition, compiled with the set and never stored | | scope | node, site or mesh: where its capacity applies. Every seat in the set has a capacity of one, so one holder per scope. A bench, a seat with several holders, is a word the glossary keeps and no seat uses yet |
| delivers | the provision its holder answers for, or nothing | | delivers | the provision its holder answers for, or nothing |
| decision | the record that made it a seat | | decision | the record that made it a seat |
@@ -74,17 +67,6 @@ needs ([28](28-building-the-bus.md), task 5.3). **And a holding is the assignmen
holder takes the row with it, so a seat never points at something that is not running anywhere, and holder takes the row with it, so a seat never points at something that is not running anywhere, and
the seat falls back to derivation rather than to nothing. the seat falls back to derivation rather than to nothing.
**A replicated seat has several holders on record** (revision, 2026-10-05,
[ADR 0223](../../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md)).
`seat <name> --add <node>/<module>` records one more holder beside those on record, judged exactly
as a handover is; `--to` still leaves exactly one. Each holder is added by that act and never by being
assigned: an assignment not on record is eligible and silent, and two claimants with nothing on record
are refused, as for any mesh seat. `--add` on a seat held once is refused, naming `--to`, and a store
recording two holders of a seat held once is refused at resolution, naming the seat. A requirement the
seat delivers is answered on a holder by itself, and elsewhere by the first holder in name order. What
the holders are is given to a module's roster template, its own machine first — how every machine's
resolver file lists both of the mesh's resolvers.
The handover refuses what would make the new holder wrong before anything is written: the seat must The handover refuses what would make the new holder wrong before anything is written: the seat must
exist, the assignment must exist, and the module must be able to hold the seat — claim it at its scope exist, the assignment must exist, and the module must be able to hold the seat — claim it at its scope
and provide what it delivers, judged against the store's row and not against anything compiled into a and provide what it delivers, judged against the store's row and not against anything compiled into a
@@ -146,13 +128,13 @@ convention, which later seats departed from.
| `mesh-npm-package-registry` | `npm-package-registry` | mesh | `npm-package-registry` | the forge | | `mesh-npm-package-registry` | `npm-package-registry` | mesh | `npm-package-registry` | the forge |
| `mesh-git` | `git` | mesh | `git` | the forge | | `mesh-git` | `git` | mesh | `git` | the forge |
| `mesh-build-machine` | `the-build-machine` | node | — | a builder | | `mesh-build-machine` | `the-build-machine` | node | — | a builder |
| `mesh-resolver` | — | mesh | — | the mesh's resolvers, each holding every node's internal domain ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)); replicated, held on the anchor and the home server ([ADR 0223](../../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md)) | | `mesh-resolver` | — | mesh | — | the mesh's one resolver, holding every node's internal domain ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)) |
| ~~`mesh-dns-port`~~ | `the-dns-port` | node | — | retired by [ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md): the local resolver became the mesh's one; deleted from the set, and from the store's table, once nothing claimed it ([ADR 0220](../../02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md)) | | ~~`mesh-dns-port`~~ | `the-dns-port` | node | — | retired by [ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md): the local resolver became the mesh's one |
| `node-hostname` | `node-hosts-file` (renamed by [ADR 0223](../../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md); the former name an alias) | node | — | owns a machine's names: `/etc/hostname`, written from its holder's `hostname` setting with no default and taking effect at the next boot, and in `/etc/hosts` the machine's own lines and the operator's kept region, changed through its verbs `entries`, `add`, `remove` ([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)). Held by the module `hostname`, formerly `hosts` | | `node-hosts-file` | — | node | — | owns `/etc/hosts`: the machine's own lines and the operator's kept region, changed through its verbs `entries`, `add`, `remove` ([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)) |
| `mesh-intrusion-prevention` | `the-intrusion-prevention` | node | — | an intrusion-prevention service | | `mesh-intrusion-prevention` | `the-intrusion-prevention` | node | — | an intrusion-prevention service |
| `mesh-packet-filter` | `the-packet-filter` | node | — | the packet filter | | `mesh-packet-filter` | `the-packet-filter` | node | — | the packet filter |
| `mesh-private-network` | `the-private-network` | node | — | the private network the mesh runs over | | `mesh-private-network` | `the-private-network` | node | — | the private network the mesh runs over |
| ~~`mesh-resolver-configuration`~~ | `the-resolver-configuration` | node | — | retired by [ADR 0223](../../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md): `/etc/resolv.conf` is written by the holder of `node-uplink`, the program that would otherwise rewrite it, and the need of the uplink beside it ([ADR 0220](../../02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md)) goes with it; deleted from the set, and from the store's table, once nothing claims it | | `mesh-resolver-configuration` | `the-resolver-configuration` | node | — | whichever of the alternative resolver configurations is chosen |
| `mesh-showcase` | `the-showcase` | node | — | the showcase module | | `mesh-showcase` | `the-showcase` | node | — | the showcase module |
The controller holds **the mesh's own** entries in code, and a test asserts their size and that The controller holds **the mesh's own** entries in code, and a test asserts their size and that
@@ -220,28 +202,10 @@ Nothing reaches that state by accident: an unknown manifest field is refused out
protocol was written as one. Checked by a registration test accepting a node seat with no protocol protocol was written as one. Checked by a registration test accepting a node seat with no protocol
and by the showcase manifest, which declares one. and by the showcase manifest, which declares one.
Most node seats deliver nothing. They say which module is this machine's packet filter, or which Most node seats deliver nothing. They say which module is this machine's packet filter, or which of
module writes its resolver file, and a second holder is refused. That is the whole of two alternative resolver configurations it runs, and a second holder is refused. That is the whole of
their job, and it is a real one: it is the mesh saying what a machine is, in words a person can read. their job, and it is a real one: it is the mesh saying what a machine is, in words a person can read.
## A seat that needs another beside it
*2026-10-05* ([ADR 0220](../../02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md)). **A seat may name the node seats its holder needs held on the
same node**, because some roles are right only while another is filled beside them. The holder of the
resolver configuration writes `/etc/resolv.conf`, and that file stays the mesh's only while the
machine's network manager is told to leave it alone — which is what the uplink's holder does
([ADR 0117](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md)). So the resolver configuration
needs the uplink.
**A module claiming such a seat depends on each seat it needs**, exactly as a module declaring a
service depends on the service manager
([ADR 0207](../../02-DECISIONS/0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md)):
derived from the claim and never written in a manifest, met by any module assigned to the node, judged
over the node's whole set, refused at assignment naming the seat and the modules that could hold it,
and refused at composition. Taking the needed seat's last holder from beneath a dependent is refused
too. What a seat needs is the mesh's definition of the role, compiled with the set and never stored,
and adding a need is a decision.
## The overview ## The overview
The controller lists every seat in the set with its scope, what it delivers, and its holder as a node The controller lists every seat in the set with its scope, what it delivers, and its holder as a node
@@ -295,8 +259,5 @@ checked as their tables say:
| `secret` has one provider, the holder of `mesh-vault` | 0161: a second claimant of the seat is refused by name (`CanHold`); *correction of fact, 2026-10-01: no parser rule ever reserved the word, the seat does the work*. | | `secret` has one provider, the holder of `mesh-vault` | 0161: a second claimant of the seat is refused by name (`CanHold`); *correction of fact, 2026-10-01: no parser rule ever reserved the word, the seat does the work*. |
| A singular fact about machines is a placement of capacity one, refused by name | 0161: the overlay command's test for a second hub; the store's unique index. | | A singular fact about machines is a placement of capacity one, refused by name | 0161: the overlay command's test for a second hub; the store's unique index. |
| A holder of `node-uplink` is the dialect the machine runs | 0161: the host reports `uplink-<manager>` in its profile with every report; a resolution test refuses the other holder naming the capability. | | A holder of `node-uplink` is the dialect the machine runs | 0161: the host reports `uplink-<manager>` in its profile with every report; a resolution test refuses the other holder naming the capability. |
| A seat's holder has the seats it needs beside it: the resolver configuration is refused at `assign` without the uplink held on its node, and the uplink's last holder cannot be taken from beneath it | 0220: seat-dependency tests on the definition, on a synthetic claimant, at assign, at unassign and at composition, and one reading the catalogue for the uplink's possible holders. |
| A replicated seat has every holder on record and composes on each; a seat held once still refuses a second holder, on record or not; `--add` is refused for it | 0223: resolution and composition tests with two holders on record and a third machine, with two claimants and nothing on record, and on `mesh-store`; a unit test that only `mesh-dns-resolver` is replicated, surviving the store's rows; store tests keeping several holders once each, a handover leaving one, and unassigning one taking only its row. |
| A retired seat leaves the set once nothing claims it, from the compiled set and the store's table both | 0220: the closed-set test's count, and the store migration that deletes the row. |
| Holdings are derived, and the overview lists every seat | 0118: the `seats` command test, including an unheld seat. | | Holdings are derived, and the overview lists every seat | 0118: the `seats` command test, including an unheld seat. |
| A build source on the seat records no address; an unheld seat refuses only self-hosted builds | 0111's tests. | | A build source on the seat records no address; an unheld seat refuses only self-hosted builds | 0111's tests. |
@@ -2,9 +2,8 @@
layer: to-be layer: to-be
status: in-progress status: in-progress
code: [mesh-controller internal/catalogue] code: [mesh-controller internal/catalogue]
updated: 2026-10-06 updated: 2026-10-02
decisions: decisions:
- 02-DECISIONS/0225-a-consumers-identity-is-bounded-by-the-provision-it-requires.md
- 02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md - 02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md
- 02-DECISIONS/0155-a-definition-names-no-installation-and-how-that-is-checked.md - 02-DECISIONS/0155-a-definition-names-no-installation-and-how-that-is-checked.md
- 02-DECISIONS/0115-one-assignment-of-a-module-per-node.md - 02-DECISIONS/0115-one-assignment-of-a-module-per-node.md
@@ -251,11 +250,7 @@ value in a container's environment is refused when the definition is parsed, wit
**A module is assigned at most once to a node**, and that pair is the assignment's identity **A module is assigned at most once to a node**, and that pair is the assignment's identity
([ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md)). Its directories, ([ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md)). Its directories,
containers, login, broker account and settings are keyed by it, as today, and a login still fits the containers, login, broker account and settings are keyed by it, as today, and a login still fits the
backend that keeps it ([ADR 0049](../../02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md)): tightest backend ([ADR 0049](../../02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md)).
the bound is the one the offer of each provision it requires states, a keyless provision states none,
and an overflow is refused by the catalogue check before merge and left out of the provider's grants,
reported, at composition — never a refusal of the provider's machine
([ADR 0225](../../02-DECISIONS/0225-a-consumers-identity-is-bounded-by-the-provision-it-requires.md)).
**A module may run on many nodes, and one assignment may hold a seat** **A module may run on many nodes, and one assignment may hold a seat**
([ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md) (superseding [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md))). The definition ([ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md) (superseding [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md))). The definition
@@ -2,11 +2,8 @@
layer: to-be layer: to-be
status: proposed status: proposed
code: [] code: []
updated: 2026-10-06 updated: 2026-10-01
decisions: decisions:
- 02-DECISIONS/0238-a-commit-is-the-build-at-hand-one-commit-one-change-plan-checked-off-the-trunk-and-published-only-on-it.md
- 02-DECISIONS/0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md
- 02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md
- 02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md - 02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md
- 02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md - 02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md
- 02-DECISIONS/0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md - 02-DECISIONS/0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md
@@ -132,58 +129,6 @@ controller replaced mid-plan resumes from the store. `status` lists open plans a
has waited too long. The transition discipline for breaking changes in the list above is still has waited too long. The transition discipline for breaking changes in the list above is still
unwritten, and still the next thing. unwritten, and still the next thing.
## How a plan sends (2026-10-05)
Revision, [ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md). Three rules on how a plan delivers what it built.
- **Grants travel before code.** A send issues the memberships for the machines it is about to send to
before their declarations. The machine holding the bus is sent first when the list of bus users changes.
A membership that could not be issued fails the send.
- **One machine first.** Unless a module's upgrade policy says *together*, a plan sends it to one machine,
the first by name, and to the rest only once that machine reports the new declaration applied and
current. A first machine that fails stops the module's rollout there, with the reason in the plan.
- **A newer merge takes over.** A merge's plan supersedes every older open plan for the same repository
and branch, and takes in the modules they had not yet built. A plan that waits on nothing can be closed
by hand, by its id.
What a merge changed is read from the forge whole, page by page
([issue 252](../../04-ISSUES/252-a-merges-changed-modules-were-read-wrong/00-report.md)). A changed path in a
module directory the mesh does not hold yet is that module's own, not shared code, when its definition is
among the changed paths.
Revision, [issue 278](../../04-ISSUES/278-a-module-held-by-no-machine-was-read-as-shared-code/00-report.md)
(2026-10-06). Whether a directory is a module is a fact of the repository at the merge commit, and the
forge's announcer says it. With each merge it lists the directories holding the changed files that hold a
definition at that commit. A changed path in one of them is that module's own, whether or not the mesh
holds the module and whether or not its definition changed. Only a path in no such directory is shared
code. From an announcer that does not list them, the rule above stands. A change to the build agent
rebuilds the build agent alone: what it builds is ordered after it in a plan, never added to one for its
sake.
Revision, [ADR 0238](../../02-DECISIONS/0238-a-commit-is-the-build-at-hand-one-commit-one-change-plan-checked-off-the-trunk-and-published-only-on-it.md)
(2026-10-06), closing [issue 280](../../04-ISSUES/280-a-rebuild-of-an-unchanged-source-was-read-as-a-new-bus/00-report.md)'s
*left open*. **There is no shared code.** A changed file touches exactly the modules whose build reads
it — the module's own directory, the whole repository for a module built from its root, a repository a
recipe names — and a file no build reads, at the root or in a directory holding no module, rebuilds
nothing. One function maps a changed file onto modules and one answers what a merge moves and builds;
the merge handler, the what-if, the merge gate and a pull request's check ask them. Only a commit on the
module's trunk is registered.
## What a push does not send (2026-10-05)
Revision, [ADR 0221](../../02-DECISIONS/0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md).
A named push still sends every other machine its work left behind, but no longer a build that an upgrade
policy or a plan is holding back
([issue 259](../../04-ISSUES/259-a-named-push-sent-every-machine/00-report.md)).
- **A send records the builds it carried**: which build of each module the machine was sent.
- **A machine a push did not name is left** when a module it runs would move to a build its policy
records rather than rolls out, or that an open plan has not sent it yet. A machine whose last send's
builds are not known is left too. The push names the machine, the module, the builds and the reason,
and says `push <node>` sends it.
- The machine a push names, a push of every machine, and `push --behind` send held builds as before.
Walking a change through the mesh under `record` is one `push <node>` per machine.
## Why now, and why not yet ## Why now, and why not yet
**Why it matters:** self-update is the difference between a mesh a person maintains by typing **Why it matters:** self-update is the difference between a mesh a person maintains by typing
@@ -12,10 +12,8 @@ code:
- mesh-tools src/main.ts - mesh-tools src/main.ts
- mesh-catalog modules/mesh-catalog - mesh-catalog modules/mesh-catalog
- mesh-tools node-tools/internal/runtime (a module's state, ADR 0201) - mesh-tools node-tools/internal/runtime (a module's state, ADR 0201)
updated: 2026-10-07 updated: 2026-10-04
decisions: decisions:
- 02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md
- 02-DECISIONS/0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md
- 02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md - 02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md
- 02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md - 02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md
- 02-DECISIONS/0126-a-module-declares-its-own-seats.md - 02-DECISIONS/0126-a-module-declares-its-own-seats.md
@@ -267,32 +265,6 @@ module's assignment** — what a module stored is data
([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)) — and one whose ([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)) — and one whose
declaration is gone is reported, never removed by the mesh. declaration is gone is reported, never removed by the mesh.
**A module declares the data it holds.** *Added 2026-10-06,
[ADR 0233](../../02-DECISIONS/0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md).*
State on the bus is one kind of data; the rest lives on a machine's disks or with a provider, and the
mesh protects only what it is told about. One section, `data`, says all of it:
| `data.` | says |
|---|---|
| `own` | each item the module keeps: `id`, `path` (`${dir:<id>}` or `${access:<id>}`, or beneath one), `class`, and its protection — `backup` (`copy`, `none`, or `{dump, into}`) or `redundancy` with why that is enough — with optional `measure` (`walk`, `dataset`, `shallow`), `within`, `active`, `why` |
| `consumers` | per provision it grants: the `class` of what it keeps for each consumer, and `in` — the own item, or the required provision, holding it |
| `kept-by` | per provision it requires: the `class` of what of its own the provider keeps |
The classes are `irreplaceable`, `valuable`, `rebuildable` and `cache` (and `none`, for consumers),
ranked by the operator; the protections, the bindings that do not move, the backup holder's lines, what
an unassignment retires and what the self-check raises are all derived from them, never written per
module. The backup holder's composed `backup` and `data` lines are the bus-free half: a file the
holder reads, like every contribution (ADR 0212). The rules are in the record; `module check` holds
them.
**A module says how each long-running resource is healthy.** *Added 2026-10-07,
[ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md).*
A container or process that stays up, and a service stated `running`, carries a field named `health`: one kind
of readiness check — the image's own adopted by name, HTTP or TCP to an endpoint named under `listens`, a
command in the container, the unit's own readiness, or one of the module's tools — with its interval,
timeout, failing looks, grace and the provision it exercises. Liveness needs no declaration. The node-engine
runs every check; `module check` holds the bounds. The design is [to-be 48](48-a-module-says-how-it-is-healthy.md).
## 5. Seats ## 5. Seats
A module declares a seat with its protocol, and the mesh enforces one holder at its scope A module declares a seat with its protocol, and the mesh enforces one holder at its scope
@@ -3,14 +3,9 @@ layer: to-be
status: in-progress status: in-progress
code: code:
- mesh-controller: internal/catalogue/seats.go (node-backup), internal/catalogue/seat_contributions.go (BackupSeat, CheckBackup, a contribution's directories) - mesh-controller: internal/catalogue/seats.go (node-backup), internal/catalogue/seat_contributions.go (BackupSeat, CheckBackup, a contribution's directories)
- mesh-catalog: modules/restic (the holder; it measures the declared data and deletes a retired item, ADR 0233), every module's `data` section - mesh-catalog: modules/restic (the holder), modules/postgres, modules/mssql, modules/mongodb, modules/minio, modules/influxdb, modules/mesh-vault, modules/mailu, modules/gitea, modules/nextcloud (backup contributions)
- mesh-controller: internal/catalogue/data.go (the lines derived from a module's data, ADR 0233) updated: 2026-10-05
- mesh-catalog: modules/nats/snapshot (the bus's snapshot and its restore, ADR 0235)
- mesh-controller: internal/broker (the bus module's snapshot-only account, ADR 0235)
updated: 2026-10-06
decisions: decisions:
- 02-DECISIONS/0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md
- 02-DECISIONS/0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md
- 02-DECISIONS/0214-backups-guard-against-mistakes-and-stay-on-the-machine.md - 02-DECISIONS/0214-backups-guard-against-mistakes-and-stay-on-the-machine.md
- 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md - 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md
- 02-DECISIONS/0085-a-secret-is-a-provision.md - 02-DECISIONS/0085-a-secret-is-a-provision.md
@@ -30,13 +25,8 @@ bucket, runs a bad migration — and not for the day a disk dies (ADR 0214).
It keeps one encrypted, deduplicating repository on the machine and runs a nightly scheduled step It keeps one encrypted, deduplicating repository on the machine and runs a nightly scheduled step
(ADR 0053). The repository's key is a secret provisioned to the holder (ADR 0085), held in the (ADR 0053). The repository's key is a secret provisioned to the holder (ADR 0085), held in the
vault, so a person can open the repository without the holder running. vault, so a person can open the repository without the holder running.
- **A module declares its data in its manifest,** naming no node and no absolute path (ADR 0112). - **A module declares its data in its manifest,** naming no node and no absolute path (ADR 0112),
*Amended 2026-10-06 ([ADR 0233](../../02-DECISIONS/0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md)):* in one of two forms:
a module no longer writes backup lines; it declares its data in its `data` section — each item's
class and how it is protected — and the lines below are **derived** from that, one list. Anything
but a cache is in the nightly backup by default; an item may say `none`, or, if it is irreplaceable
and cannot be copied, `redundancy` with why, and the holder then watches the array it is on. A
backup line written by hand is refused by `module check`. The two forms the derivation produces:
- **a dump** — for a store provider: the command that writes a consistent, logical copy of each - **a dump** — for a store provider: the command that writes a consistent, logical copy of each
database it serves, run by the provider in its own container, its output handed to the holder. database it serves, run by the provider in its own container, its output handed to the holder.
The provider covers every consumer it provisions, so a module that only *uses* a database The provider covers every consumer it provisions, so a module that only *uses* a database
@@ -48,25 +38,6 @@ bucket, runs a bad migration — and not for the day a disk dies (ADR 0214).
module assigned is covered the next night; a module unassigned stops being backed up, and its module assigned is covered the next night; a module unassigned stops being backed up, and its
restore points age out by the rotation, never at once. restore points age out by the rotation, never at once.
## The holder measures what is declared
*Added 2026-10-06, [ADR 0233](../../02-DECISIONS/0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md).*
The mesh composes a second file for the holder: every data item every module on the machine declares,
with its class, its path, the path that covers it and its protection. Once an hour and after every
night the holder measures each item that is not a cache, by the item's `measure`: a bounded walk at
most daily for a small item (stopped after ten minutes or two million files and said as partial), a ZFS
dataset's own counters hourly, or the top-level entries only — never an hourly walk of anything large —
and reads the redundant storage it is on (a ZFS pool's health, an md array's
members, a btrfs filesystem's error counters). `backed-up` says each item with that measurement and its
newest good backup. The holder judges nothing: the controller's self-check (to-be 45, D13) keeps the
readings and raises what they show.
**It deletes a retired item, when a person has decided.** An item the mesh retired — its module gone
from the machine — is removed by the holder only when the controller's `cleanup delete` asks, and only
after it has taken a last restore point of it, tagged as retired: a deletion can be undone until that
restore point is forgotten. It refuses any path declared on the machine now, its own repository, and
anything too near the root.
## A night ## A night
Each declared dump runs and writes a full logical copy; each declared path is read as it stands. All Each declared dump runs and writes a full logical copy; each declared path is read as it stands. All
@@ -75,37 +46,6 @@ has not seen, so the object store's first night costs its full size and later ni
changed. Then the rotation prunes to 14 daily, 8 weekly and 6 monthly snapshots. A dump that fails changed. Then the rotation prunes to 14 daily, 8 weekly and 6 monthly snapshots. A dump that fails
fails the night for that module only; the others are still taken. fails the night for that module only; the others are still taken.
## The bus: the server's own snapshot, never its files
*Added 2026-10-06, [ADR 0235](../../02-DECISIONS/0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md).*
The bus keeps everything it holds in streams — its key-value buckets are streams too — and its store
is written all the time, so a copy of the store's files may not restore. Its streams are protected by a
dump like a database's: in the bus's own container, a program built into the bus's image asks the
server for each stream through the server's snapshot API, one at a time, acknowledging each chunk so
the server's flow control keeps moving; the server goes on taking writes, and no stream is paused or
reconfigured. It runs under the bus module's own account, which may take snapshots and do nothing
else. It writes one archive into the module's snapshots directory, which the night copies: a manifest
— when, how long, and per stream its messages, bytes, sequences, consumers and the checksum of each
file — then each stream's configuration and the server's archive of it. The live store is not copied.
Every message up to the last sequence the manifest gives for a stream is in the archive, exactly; a
stream written during its snapshot may carry part of what arrived while its blocks were read, after
that sequence. A failure fails the bus's night, keeps the previous archive, and is said as
`backup-stale` like any other.
**Restoring the bus** is done beside it, never over it — a stream is restored only where it does not
exist, and on the live bus every one exists:
1. `restore` the bus module's snapshots from a named night, beside the live ones;
2. check the archive against its manifest (`mesh-nats-snapshot verify`);
3. build a new store from it with the bus's own image: the program starts the bus's server on loopback
with the bus's one account, restores every stream, holds each to the manifest and stops it;
4. swap the new store in for the live one with the bus stopped — a person's act and a planned bus step
(to-be 45), in one command line, because the machine's node-engine starts a stopped container again
at its next reconcile. The store kept aside is removed by a person once the bus is seen whole.
The module's README has the commands.
## The verbs ## The verbs
On the seat, for a person or an agent: On the seat, for a person or an agent:
@@ -126,9 +66,6 @@ a manifest. Not off-site: a lost machine loses its backups with its data, by the
## Proving it ## Proving it
- A night that fails, or does not run, reaches the operator's output channel, naming the module. - A night that fails, or does not run, reaches the operator's output channel, naming the module.
- The bus's snapshot is proven by restoring it, in tests against the bus's own release: a server
written to throughout its snapshot, restored into a fresh server and into a new store served by a
third, compared message by message (ADR 0235).
- Weekly, the holder checks the repository's integrity and restores the newest dump of one database, - Weekly, the holder checks the repository's integrity and restores the newest dump of one database,
in rotation, into a throwaway instance with no network, comparing table row counts with the live in rotation, into a throwaway instance with no network, comparing table row counts with the live
database. database.
@@ -137,17 +74,10 @@ a manifest. Not off-site: a lost machine loses its backups with its data, by the
## Not in this design ## Not in this design
- Copies off the machine, and encryption to anyone but the mesh's own vault. - Copies off the machine, and encryption to anyone but the mesh's own vault.
- Copying the media library: there is no room for it. It is declared irreplaceable and protected by the - The media library and anything else a module declares no data for.
redundancy of the array it is on, which the holder watches (ADR 0233); a lost array loses it, by the - Recreatable things: container images, the artifact registry, caches.
operator's accepted risk.
- Anything a module declares no data for.
- Recreatable things a module says `backup: none` of — container images, the artifact registry,
downloaded models — and caches.
## How it is checked ## How it is checked
The catalogue check (`module check`, and the catalogue-wide test beside it) refuses a provider that The catalogue check refuses a module that provides a store seat and declares no dump. The weekly
grants without saying what it keeps for its consumers, an irreplaceable item with neither a backup nor a
redundancy, and any backup line written by hand (ADR 0233). The store rule this design stated before
was held only by a controller test over the catalogue checked out beside it, never by `module check`. The weekly
restore test above, and the 48-hour status line, are the running checks. restore test above, and the 48-hour status line, are the running checks.
@@ -1,760 +0,0 @@
---
layer: to-be
status: in-progress
code: [mesh-controller, mesh-host, mesh-tools, mesh-catalog, mesh-sdk, mesh-lab]
updated: 2026-10-07
decisions:
- 02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md
- 02-DECISIONS/0239-a-delivery-is-owned-by-the-mesh-delivery-module-and-runs-from-commit-to-delivered.md
- 02-DECISIONS/0238-a-commit-is-the-build-at-hand-one-commit-one-change-plan-checked-off-the-trunk-and-published-only-on-it.md
- 02-DECISIONS/0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md
- 02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md
- 02-DECISIONS/0234-the-mesh-holds-a-conversation-with-its-operator.md
- 02-DECISIONS/0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md
- 02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md
- 02-DECISIONS/0229-the-cores-order-is-a-lease-the-store-remembers-and-an-epoch-a-machine-is-sent-once-it-reads-one.md
- 02-DECISIONS/0231-a-healer-acts-on-what-observation-raised-and-only-observation-says-it-worked.md
- 02-DECISIONS/0230-a-consumer-the-mesh-stops-asking-for-is-retired-and-deleted-only-by-a-person.md
- 02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md
- 02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md
- 02-DECISIONS/0141-the-host-delivers-its-own-successor.md
- 02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md
---
# 45 — A core that cannot fail silently
**The core says when it is wrong, refuses what is stale or unreadable, repairs what it knows how to
repair, replaces itself one machine at a time with something other than itself watching, and is checked
against the mesh's real facts before a change merges** ([ADR 0227](../../02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md),
from [research 031](../../01-RESEARCH/031-a-core-that-cannot-fail-silently/00-overview.md)).
**The core**, here: the controller, the node-engine and its launcher, the bus server, the node tools and
the console, the build seat, and the forge's announcer of merges.
This document is the build's specification. §1–§9 are the parts; §10 is the order they are built in,
each phase with what it delivers, in which repository, and when it is done. Every bound marked
*provisional* is set from the durations Phase 0 records and corrected in Phase 1's first live week.
## The parts, and how a fact reaches the operator
```
signals ──────────► watchdogs ─┐ ┌─► status (open conditions first)
(heartbeats, reports, │ │
plan progress, advisories) ├─► CONDITION STORE ───┼─► condition events ─► operator-channel ─► Telegram
doctor probes ─────────────────┤ (controller is │ holder desktop notifier
(live invariants, every 5 min) │ its only writer) └─► healers ─► repair, braked ─► event
events (provider failing, …) ──┘ │
└─► hand-act log (people)
second machine: watcher ── hears doctor's heartbeat? ── no, past bound ──► Telegram directly (not the bus)
```
Owning repositories: **mesh-controller** (the condition store, watchdogs, `doctor`, healers, the lease,
calls, the hand-act log, the facts snapshot), **mesh-host** (the node-engine: the apply queue, report
order, epoch refusal, the `report` verb, rollback witnessing), **mesh-tools** (the node tools and the
console: their heartbeat, passing every argument, health answers), **mesh-catalog** (the
`operator-channel` seat's holder and channels, the watcher, the providers' retirement, the catalogue's
merge gate), **mesh-sdk** (the TypeScript providers' loop and its retirement), **mesh-lab** (the
replays and the induced-failure scenarios).
---
## 1. The writers table (rule 1)
Every kind of state the core keeps, and the one component that writes it. Anyone else asks.
| State | Writer | Kept in | Others |
|---|---|---|---|
| a machine's declaration | controller (lease holder) | the bus, last per subject | read |
| a machine's applied state and its report | the node-engine's apply queue (§6) | the machine; the report on the bus | the reconcile and a delivery *enqueue*, never apply |
| the controller lease | the controller instance holding it | key-value `mesh-controller_lease` | a candidate waits |
| plans and their tiers | controller (lease holder), compare-and-set on the plan's revision | the controller's store | read through `plans` |
| conditions | controller | key-value `mesh-controller_conditions`, and their history `mesh-controller_condition-history` | raise or clear only through observations the controller reads |
| calls and their outcomes | controller | key-value `mesh-controller_calls` | read by id |
| the hand-act log | controller, through the verbs that act | key-value `mesh-controller_hand-acts` | — |
| the healers' acts and their brake | controller (lease holder), each act begun before it is made | the controller's store | read through `healers`; each act said as `healer-acted` and in its condition's `tried` |
| stream definitions and bus permissions | controller | the bus | — |
| builds and their outcomes | the build seat's holder | its own state | the controller asks |
| a merge announced | **one** announcer per forge (the hook, or the poll when the hook is absent — never both) | the bus | — |
| a pull request's head announced | the forge's announcer, the one that announces its merges | the bus | the controller asks the build seat to check it |
| a pull request's merge check | the build seat's holder that ran it, said as the controller's `checked` | the bus | the forge's holder sets it as the pull request's status |
| a provider's standing | the provider | the provider's events | the controller keeps the newest word as a condition |
| a consumer's retirement: active, retired (when, why), deleted | the provider | its own backend's mark; said on `provisioner.retirement` | the controller asks the provider's tools; a person approves, rejects and deletes through the controller's verbs |
| the operator-channel's open messages | the seat's holder | its own key-value state | — |
| a machine's declared data: what it measured, what is retired, what was deleted (ADR 0233) | controller, from what each `node-backup` holder answers | the controller's store, `data_item` and `data_reading` | the holder measures and says; it records nothing the controller keeps |
| the facts snapshot | controller | the artifact store, `facts/latest` | the build seat reads |
A bus subject two components may publish on is refused when the controller composes grants, unless
this table marks it shared. Changing a writer is a change to this table, through a decision.
**How it is enforced** ([ADR 0229](../../02-DECISIONS/0229-the-cores-order-is-a-lease-the-store-remembers-and-an-epoch-a-machine-is-sent-once-it-reads-one.md)):
the table is compiled into the controller, each row the bus carries with the subjects a write of it
publishes to — a declaration `mesh.node.*.declare`, a report `mesh.control.*.report` (its own machine's
only), the controller's buckets, stream creation, update and deletion and durable consumer creation, a
build's outcome, a merge announced, a provider's standing — and its writer among the principals. A
grant that overlaps another writer's subject is refused at composition, naming the state and its
writer. A test holds the compiled table to this one, row for row. The controller does not publish
`mesh.control.>`: it would be a second writer of every machine's report.
## 2. The condition store (rules 5, 6)
**A condition is a durable fact about something the mesh owns that is wrong.** It is raised and cleared
by observation only.
**Where.** A key-value bucket, `mesh-controller_conditions`, written by the controller alone, one entry
per open condition, with no age. Every transition — raised, changed, silenced, cleared — is also
appended to a history kept ninety days in a bucket of its own, `mesh-controller_condition-history`,
read through `conditions history`: a bucket has one age for every key, and an open condition must
outlive the history.
**The key** names the thing and the kind, so the same fault said again is the same condition:
`<scope>.<id>.<kind>`, where scope is one of `machine`, `plan`, `call`, `build`, `merge`, `provider`,
`seat`, `bus`, `core`, `probe`, `mesh`. Examples of the shape: `machine.<node>.silent`,
`plan.<id>.stalled`, `provider.<module>.<node>.<consumer>.failing`, `core.controller.<node>.rolled-back`,
`bus.<consumer>.slow-consumer`. The last token is the kind's short word where it has one (`failing` for
`provider-failing`), the kind elsewhere; a key is opaque to every reader, and the kind is the field.
*Added 2026-10-07, [ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md):*
the scope `module`, for a module's own health on a machine — `module.<module>.<machine>.unhealthy`, raised by
the controller on the second statement from the node-engine that a long-running resource is unhealthy, and
cleared on the first that does not. Designed in [to-be 48](48-a-module-says-how-it-is-healthy.md).
**The fields:**
| Field | Holds |
|---|---|
| kind | the condition kind, from the signals table, the probe registry or an event kind |
| subject | scope, id, and the machine it concerns when there is one |
| severity | `urgent` (needs the operator now) or `warning` (when they can) — two levels, no more |
| summary | one line in the mesh's words |
| evidence | the newest observations, at most ten, each with its time |
| source | the signals-table row, probe or event that raised it |
| raised, last observed | times; and how many observations since raised |
| tried | each healer attempt: when, what, outcome |
| resolver | `self` (clears on observation), `healer:<name>`, `operator`, or `agent` |
| silenced | until when, by whom, why — empty when not silenced |
| epoch | the controller lease epoch that last wrote it |
**The life of one:**
```
(absent) ──observation past bound──► OPEN ──healer tries──► OPEN (tried += …)
│ ▲ │ budget spent
│ └─ seen again ◄───────┘──► OPEN, resolver: operator, severity: urgent
│
silence (verb, ≤ 7 days, reason, hand act) ──► OPEN, silenced (no messages)
│
observation says resolved ──► CLEARED: entry removed, transition kept in history
```
- **Nobody resolves a condition by hand.** It clears when the signal returns or the probe passes.
- **A condition cleared and raised again within ten minutes** reopens with its count increased; it is
not a new message.
- **The verbs:** `conditions` (open ones, filtered by scope, severity or machine), `conditions show
<key>`, `conditions silence <key> --for <duration> --why <text>`, `conditions history`.
- **The events**, emitted by the controller on its own subjects: `condition-raised`,
`condition-changed` (severity or resolver changed; not every observation), `condition-cleared`. The
operator-channel's holder and any other surface consume them; the controller learns nothing about
telling. Each carries the condition at the top level, with `event`, `at`, `change` (raised,
reopened, severity, resolver, silenced, silence-ended, cleared), `why`, `was`, `cleared` and `show`
beside it; a silence and its ending are `condition-changed`, a reopening `condition-raised` with
change `reopened`.
- **`status`** lists open conditions first, urgent before warning, oldest first, silenced ones marked
with their expiry. The all-well sentence requires none open, silenced included.
- **ADR 0224's provider standing** is the first kind: `provider-failing`, raised by the provider's
event, cleared by its recovery, its silence after thirty minutes the row S8 below.
- **Retirement** ([ADR 0230](../../02-DECISIONS/0230-a-consumer-the-mesh-stops-asking-for-is-retired-and-deleted-only-by-a-person.md))
raises three, keyed by the provider module and its machine: `retire-waiting` (urgent) while a stable
set too large to retire alone waits for `retire approve` or `retire reject`, listing it; `retire-rejected`
(warning) while a rejected set is kept active though the mesh no longer asks for it; and
`cleanup-waiting` (warning), from the probe D11, while anything retired is older than thirty days. The
first two clear on the provider's `retired`, `settled` or `approved` word, and all three when the
provider is no longer assigned there.
- **Data** ([ADR 0233](../../02-DECISIONS/0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md))
raises, from D13, each **urgent for irreplaceable data and a warning for valuable data** — the
operator's ranking — and none for rebuildable data or caches: `data-shrank` (less than half its
largest size in seven days, and 16 MiB less), `data-missing`, `empty-replacement` (a copy in use
less than half a copy of the same thing kept elsewhere — issue 273's shape — for an item or for a
consumer's data at a provider), `data-quiet` (an item that says `active`, unwritten past it),
`backup-stale`, `array-degraded` (one per array under watched data) and `protection-missing` (an item
said to be on redundancy, found on plain storage); and as warnings `data-held-twice`,
`data-unmeasured` and `cleanup-waiting` for an item retired over thirty days. Keyed by machine,
module and item, or by the array. A shrink a person meant is silenced with why; everything else
clears on observation.
## 3. The signals table (rule 5)
Compiled into the controller. A watchdog per row; the test generated from the table suppresses each
signal and asserts its condition. `doctor signals` shows, for every row, the age of its newest signal.
| Row | Signal | Emitter | Trigger | Bound | Condition kind | Severity | Healer |
|---|---|---|---|---|---|---|---|
| S1 | machine heartbeat | node-engine | its interval | 3 × interval, *provisional* | `silent` — not raised while the machine has declared itself asleep or shut down (ADR 0211) | warning; urgent after 30 min for the control node | — |
| S2 | report after a send | node-engine | each declaration sent | max(2 min, 3 × that machine's last apply duration), *provisional* | `sent-not-reported` | warning | H1 |
| S3 | plan tier progress | controller's plan | each tier entered | from the tier's build and apply durations, *provisional* | `stalled` (names the tier and what it waits on) | warning | H2 |
| S4 | the controller's event loop takes a message | controller | while its consumer has pending messages | 2 min | `controller-deaf` | urgent | — |
| S5 | a merge announced becomes a plan or *nothing reads it* | announcer → controller | each merge | 10 min (exists, issue 266) | `merge-not-acted` | urgent | — |
| S6 | a build asked → its outcome | build seat | each ask | max(20 min, 3 × the p90 of measured builds), 1 h while nothing is measured, *provisional* (the build seat declares no timeout) | `ask-lost` | warning | — |
| S7 | a call running → finished | controller | each call | the verb's bound: push, rotate and command 30 min; assign and unassign 15 min; doctor 3 min; others 10 min | `call-hung` | warning | — |
| S8 | a provider's failing word repeated | provider | every 15 min while failing (ADR 0224) | 30 min | `provider-silent` | warning | — |
| S9 | bus advisories: slow consumer, maximum deliveries, permission violation, consumer deleted | bus server's system subjects | any | any occurrence | `slow-consumer`, `max-deliveries`, `refused`, `consumer-lost` — each naming the call, consumer or module in the mesh's words | warning | H3 for `consumer-lost` |
| S10 | the self-check's heartbeat | controller's `doctor` | every run | 2 × its interval, watched **from the second machine** (§5) | `self-check-silent` | urgent | — |
| S11 | node tools heartbeat | node tools | its interval | 3 × interval, *provisional* | `tools-silent` | warning | — |
| S12 | the controller lease renewed | controller | every 5 s | 15 s; a holder that lost the lease or was found expired, and a lease bucket found raised again from nothing, said for an hour; a controller serving without the lease, while it does | `lease-lost` | urgent | — |
| S13 | stale refusals | every receiver (rule 2) | each refusal | more than 5 from one writer in 5 min | `stale-writer` (names the writer: the controller epoch a refused declaration claimed, with its instance and how its lease ended; or the machine whose older accounts the controller refused) | warning | — |
| S14 | facts snapshot exported | controller | when it moved, and daily | 2 days, or none kept by a controller up that long (Phase 5) | `facts-stale` | warning | — |
| S15 | a hand act with a cause already recorded | hand-act log | each act | the second within 14 days; clears when fewer than two remain within 14 days | `healer-wanted` (names the cause, and the healer that was not enough where one answers it) | warning | — |
| S16 | a walk waiting for its delivery's word is let go | mesh-delivery, through `deliver` (ADR 0239) | each merge whose walk waits | 30 min; urgent after 4 h | `waiting` (names the walk and `plans go`) | warning | — |
The bus advisories (S9) cost one read-only subscription: the server already publishes them. The
controller translates each into a condition naming the thing in the mesh's words, as the refused reply
of issue 265 is translated today. An advisory clears after an hour without another; a deleted consumer
is said only for one the mesh names and expects, and clears when it exists again. Until the bus has a
system account the controller hears a slow consumer and a refused subject for its own connection only
([issue 270](../../04-ISSUES/270-phase-1-watches-signals-that-come-later-and-hears-only-the-controllers-faults/00-report.md),
open question 2).
## 4. The self-check: `doctor` (rule 6)
A **probe registry** in the controller: each probe is a live invariant of a design, with an id, an
interval, a timeout and the condition kind it raises. The controller runs the registry every five
minutes; each probe has thirty seconds. A probe that errors or times out is said in that run's
verdict at once — an unanswered probe is never a pass — and raises `probe-failed` for itself when the
next run cannot run it either.
**A finding one look can be wrong about is raised on the second look in a row**
([issue 277](../../04-ISSUES/277-one-unanswered-question-was-an-urgent-alert-nobody-could-read/00-report.md)):
a question over the network that went unanswered, a time measured once, a reading a burst can move.
Within a run such a question is asked again before it counts; across runs the finding is raised when
the previous run saw it too, kept while it is open however it is seen, and cleared by the run that no
longer sees it. Held back, it is listed in the verdict as unconfirmed. A definite answer — an address
that is wrong, a stream that is not there — is raised at once. The watchdogs hold a row that cannot
read its facts to the same rule over two ticks. Checked by the controller's tests of the self-check over
runs and of a blind row over ticks.
**A condition's summary names machines and says the rest in words**; an address, a domain, a path or a
raw error is its evidence, which stays inside the mesh. The condition store holds every summary to the
operator channel's content rule (ADR 0234 §6) and says one that would be withheld in words, keeping it
whole in the evidence; a lint over every condition raised in the controller's test suite fails the
producer that wrote it.
| Probe | Asserts | From |
|---|---|---|
| D1 | every machine's declaration composes, and passes the node-engine's validation (the validator is a package of mesh-host the controller and the merge gate import — one validator) | 236, 263 |
| D2 | every holder of the mesh's resolver answers a machine name for IPv4, and NODATA for IPv6 | 262 |
| D3 | every seat on record has a live holder that answers (`holder-silent`, H3) | 208, 218 |
| D4 | every kept archive is held by a manifest | 253 |
| D5 | exactly one lease holder — the key names this controller at the epoch it acts under, and the controller's record holds no other epoch open; no message from a stale epoch refused in the last interval | 204 |
| D6 | every durable consumer exists with its definition and is near its stream's head (within 1000 messages, *provisional*); a missing one is `consumer-lost` (H3), one far behind `consumer-behind` (H4) | 248, 266 |
| D7 | every stream the controller defines exists with its definition | 208 |
| D8 | no address the mesh owns is in a ban list | 238 |
| D9 | `status` answers in full within ten seconds | 265 |
| D10 | every machine runs the node-engine and node tools builds its plan says, or is inside a plan's window | version split |
| D11 | no provider holds a consumer retired more than thirty days: each provider assigned is asked `provisioner_retirement` on its machine (ADR 0230); one that does not serve it yet is named, not failed | ADR 0230 |
| D12 | every consumer of a provision that keeps its data is bound where it was last sent, or moves by a pin (`binding-kept`, `binding-moving`, `binding-moved`) | ADR 0232 |
| D13 | every item of data a machine declares is measured, is there, holds what it held, is written where it says it is, is backed up within its bound or sits on healthy redundant storage, and is no empty replacement of a copy kept elsewhere; an irreplaceable or valuable item a machine no longer declares is retired, not forgotten. Each machine's `node-backup` holder is asked `backed-up`; every provider of kept consumer data `provisioner_retirement` (with each held consumer's size) — see §2 for the kinds | ADR 0233, issue 273 |
| DW | the watchdogs of §3 ran within three of their intervals: the watchers are watched, and raise S10 when the self-check stops | rule 6 |
| H-* | the health probes of §8, run for every core component on every machine: H-controller, H-engine, H-tools, H-bus (as built, [ADR 0236](../../02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)) | rule 8 |
| DG | every build that failed its gate, and every core build a witness put back, is said — `build.<module>.<machine>.rolled-back`, `core.<component>.<machine>.<outcome>`, `rollback-failed` — until a newer build passes or the reports stop saying it (ADR 0236) | rule 8 |
| DB | a bus upgrade a person started is said as `bus-maintenance` while it runs, and ends healthy within its bound or is said `bus-upgrade-failed`, with its snapshot (ADR 0236) | rule 8 |
- **`doctor`** answers the last run's verdict at once: per probe, pass, fail or failed-to-run, and age.
**`doctor run`** runs now under a call id. **`doctor probes`** lists the registry; **`doctor
signals`** the table's ages.
- **Every run ends with a heartbeat** event carrying the run's id and counts. That is S10.
- **The registry is the design's live form.** A check over the to-be designs counts invariants that
name a probe against those that do not; the number without may only go down.
## 5. The output channel, minimal form (rule 6)
From [research 028](../../01-RESEARCH/028-the-meshs-output-channel/00-overview.md), the smallest form
that works. **Amended 2026-10-06 by [ADR 0234](../../02-DECISIONS/0234-the-mesh-holds-a-conversation-with-its-operator.md)**,
028's graduation: this form is the first step of [to-be 46](46-the-conversation-with-the-operator.md),
the conversation with the operator, which replaces the points marked below. Phase 1 builds this
section; to-be 46 §13 builds what follows from it.
- **One mesh seat, `operator-channel`**, held once. It **accepts** `notify` as a work queue, so a
message waits for a holder; it keeps its **open messages** in its own key-value state, so a restart
forgets nothing; it **serves** `open` and `history`.
- **The holder consumes the controller's condition events** and decides what is sent. The controller
calls nobody.
- **Two channels**: **Telegram** (a bot to the operator's chat) and the **desktop notifier** of the
machine the operator is at. *Amended:* a channel is not a contribution to this seat. Each is the
holder of a kind on the kinded benches `channel` and `intake`, a module of its own holding its own
secrets, and this seat's holder becomes the router (to-be 46 §2, §3, §8).
- **A message** is: the condition's key, its subject (a machine's role, a module, a plan), kind,
severity, the one-line summary, since when, and the verb that shows more. **Deduplicated by the key.**
- **When:** on `condition-raised`; once more if still open after 1 hour (urgent) or 12 hours
(warning); on `condition-cleared`, by editing the first message where the channel can. A silenced
condition sends nothing. Urgent goes to both channels; warning to the desktop notifier when the
operator's session is there, otherwise to Telegram.
- **Rate:** at most twenty messages an hour; the excess is folded into one message naming them all.
- **What may leave the mesh:** roles and words. A message carrying an address, a path or anything
shaped like a secret is refused by the holder and raises `channel-refused` instead.
*Amended:* the mesh's own machine names may appear too; a concrete detail travels as a reference only
a private surface opens, and a refusal names the offending part to the sender (to-be 46 §9).
- **Answering back.** *Amended:* the minimal form had none, and acknowledging was `conditions
silence`, through the mesh. The operator now answers and is asked over the intake seat; an answer
that performs an action is checked and performed by the controller (to-be 46 §5, §6, §10). `conditions
silence` stays callable.
- **The watcher's watcher.** A module, `mesh-watcher`, assigned to one machine that is not the control
node, holds the Telegram channel's secret too. It listens for the self-check heartbeat (S10) and for
the bus itself; when either has been silent past its bound it sends to Telegram **directly over
HTTPS, not through the bus**, and says so again when they return. It is the only sender that does not
pass through the control node.
## 6. Order: lease, epoch, report sequence, one apply queue (rules 1, 2)
**The lease.** A controller instance acts — sends a declaration, writes a plan, a condition or a call
— only while it holds the key `holder` in `mesh-controller_lease`: written with compare-and-set, a
fifteen-second time to live, renewed every five seconds. The **epoch** is the bucket revision at which
it was taken. A starting controller waits for the key to be absent or expired. One that fails a renewal
stops acting at once and exits, so its service manager restarts it as a waiting candidate.
As built ([ADR 0229](../../02-DECISIONS/0229-the-cores-order-is-a-lease-the-store-remembers-and-an-epoch-a-machine-is-sent-once-it-reads-one.md)):
- **The gate is the clock.** A holder stops acting three seconds before its key could expire
unrenewed, whatever its renewing goroutine is doing; every act passes that gate at the moment it is
made — a declaration at its send as well as where it was composed.
- **The epoch never goes backwards.** The controller's store keeps every epoch issued, its instance,
and how it ended (`released`, `lost`, `expired`); a lease bucket at or under the highest is moved past
it before the key is taken, and S12 says so.
- **A holder that stops gives the key back**, so the next takes it at once.
- **Without the bucket's grant, nobody holding it,** a controller serves unleased — no epoch on what it
sends, S12 urgent — and takes the lease once the user list that grants it reaches the bus.
- **A command at a shell** acts under the holder's epoch, read as it acts, or under a lease of its own
while nobody holds one.
```
controller A (epoch 41) ──renew──renew──╳ (renewal refused)──► stops sending, exits
controller B ──wait──────────────take (epoch 57)──► acts; marks A's running calls abandoned
node-engine accepts 41 … then 57; refuses anything from 41 after 57, counted, reported
```
**What carries the order:**
| Message | Carries | The receiver keeps, per writer | Refuses |
|---|---|---|---|
| declaration | epoch, sequence | highest epoch, then highest sequence | older epoch; same epoch, lower sequence — only where both claim an epoch |
| report | machine, the declaration's epoch and sequence it is about, the node-engine's own report sequence (kept on disk, increasing across restarts and self-updates) | highest epoch where both claim one, then declaration sequence, then report sequence | an account of an older declaration; an older report |
| plan write | the plan's revision, the epoch | — (compare-and-set) | a write against a revision already moved |
| build outcome | the build's ask id and order | newest per module | an older build finishing later (exists, 219) |
| call | call id, epoch | — | a finish from an epoch that is not the holder's, recorded as abandoned |
| merge announcement | forge, repository, commit, the announcer's sequence | newest per repository | a duplicate (one announcer) |
**Every refusal** is one line in the receiver's log in the mesh's words, a counter, and — from the
node-engine — a report naming the refused declaration, so the controller sees it. The counter feeds
S13.
**The wire** — the contract, written once on each side: a declaration's `epoch` and `sequence` are
top-level keys of its signed envelope; a report carries back the `epoch` and `sequence` of the
declaration it is about, its own `report_sequence`, `older_than` (the order held) on a refusal of an
older declaration, and `refused_older` (how many ever) on every report. Every key is optional, zero is
"no order claimed". **A machine is sent `epoch` only while its latest account carried a
`report_sequence`**, because an older node-engine refuses an unknown key, whole; a declaration without an
epoch is taken by its sequence, so a controller rolled back to a build without the lease is never
stranded. A report without a `report_sequence` is judged by the digest it names (issue 267). What the
mesh would send a machine is composed with the epoch it was last sent: a new holder is not a change of
the machine.
**One apply queue on every machine.** The node-engine has one worker that applies; a delivery, the
five-minute reconcile, a self-update hand-over and a manual `apply` are four reasons to **enqueue**.
The queue holds at most one pending request, coalesced. When the worker starts it takes the newest
declaration held at that moment, applies it once, and sends one report naming the sequence it applied.
Nothing else in the node-engine applies.
**The `report` verb.** The node-engine answers, on request, the report of the last declaration it
applied, from what it keeps on disk. It is what healer H1 asks. As built ([ADR 0231](../../02-DECISIONS/0231-a-healer-acts-on-what-observation-raised-and-only-observation-says-it-worked.md)):
asked on `mesh.node.<machine>.ask.report`, its own only, the node-engine enqueues a reconcile whose
account is said whether or not it is news — a delivery waiting is applied and reported instead — and the
answer is its ordinary report; nobody's inbox is answered.
**Durable calls.** Every call's record — verb, arguments with secrets removed, caller, started, epoch,
state (`running`, `finished`, `failed`, `abandoned`), the answer bounded in size, finished — lives in
`mesh-controller_calls`, the last thousand or fourteen days. A new lease holder marks the previous
holder's running calls `abandoned`, which is said.
## 7. Healers and the hand-act log (rule 7)
**A healer** is a registered response to one condition kind: its repair (the ordinary path again), its
budget, its back-off, its brake, and the `healer-acted` event naming the condition, the act and the
outcome. A healer may not withdraw, delete or recreate data; such a repair is a condition for the
operator.
| Healer | Condition | Repair | Budget, then |
|---|---|---|---|
| H1 | `sent-not-reported` | ask the machine's node-engine for `report`; if it does not then report the declaration it was sent, send it again — never moving a build a policy or a plan holds back | twice per condition in 6 h, then resolver `operator`, urgent |
| H2 | `stalled` on a wait that is superseded (a newer plan of its repository and branch) or already finished (every module built or failed, every one that rolls out sent) | close the plan with its note, as `plans close` does: `superseded` or `done` | once |
| H3 | a seat holder that does not answer (D3, `holder-silent`), or a consumer the mesh expects missing (D6, S9, `consumer-lost`) | assert the bus's objects again — the assertion every send makes (208) | once per object per hour |
| H4 | `consumer-behind` (D6) on a consumer the stream table marks *resettable* — the controller's own events consumer alone | `broker consumer-reset` (248) | once a day per consumer, then condition |
| H5 | `provider-failing` with the administrator refusing the mesh's secret | ADR 0224 §5 (exists, in the module) | as ADR 0224 |
As built ([ADR 0231](../../02-DECISIONS/0231-a-healer-acts-on-what-observation-raised-and-only-observation-says-it-worked.md)): the registry is compiled into the controller with a test generated
from it; each act is begun in the controller's store before it is made, under the lease, and the
budgets are counted from there; a healer that does not apply to a condition leaves it alone; **only
observation says a repair worked** — the act is kept in the condition's `tried` as `healer Hn` and said
as `healer-acted`, and the condition clears when what raised it no longer sees it; a spent budget hands it
to the operator, urgent, and no healer touches it again. **The mesh-wide brake:** twelve acts in an hour,
all healers together, stop every healer — `mesh.healers.braked`, urgent — until an hour after the last.
A heal is never a hand act. `healers` lists the registry, the acts and the brake.
**The hand-act log.** Every verb that repairs by hand — a named `push` outside a plan, `plans close`,
`broker consumer-reset`, `conditions silence`, `retire approve`, `retire reject` and `cleanup delete`
(ADR 0230; and an approval, rejection or deletion a provider reports that did not come through the
controller), and `hand-act record` for an act done outside the mesh —
takes a required `--why` and writes an entry: who, which verb and arguments, why, when, the condition
key it addresses if any, and a **cause** (the condition kind, or a word the person gives). `status`
shows the week's count. A cause recorded twice within fourteen days raises `healer-wanted` (S15) —
except an act recorded by a verb that is a person's decision by design, which is no repair and never
counts, whatever cause it gives. The controller keeps one table of the verbs that write the log, and
each says whether it records a repair or a decision, and why: approving or rejecting a retirement and
deleting what was retired (ADR 0230); `bus upgrade`, because the bus is never rolled by the mesh, and
`upgrade release-backlog`, because after a failed release plan the next opens only on a person's word
(ADR 0236); and `secret rotate` when its cause is a leak (`leaked-in-logs`) — a person judges what was
disclosed, one leak rotates several values, and a leak that recurs is a defect of the module that
prints them, an issue against it, not a healer that rotates. A rotation for any other cause counts. A
push by hand, `plans stop` and `close`, `broker consumer-reset`, `conditions silence` and `hand-act
record` (an act outside the mesh, whose kind the mesh cannot tell) count; so does a verb the table does
not list. The table's test finds every verb the controller records and fails on one it does not list.
A `healer-wanted` already open for a decision clears on the next watchdog tick, as any condition its
row no longer sees.
> **Progressive insight — 2026-10-06.** This paragraph said before: "except `retire-waiting` and
> `cleanup-waiting`: approving a retirement and deleting what was retired are a person's decision by
> design (ADR 0230), and no healer may take them over." The exception was keyed on two causes, and the
> first planned bus upgrades (two within a day, one recorded after the fact and one through `bus
> upgrade`) raised `healer-wanted` for `bus-upgrade`, as did two rotations after one leak for
> `leaked-in-logs`. What makes an act a decision is the verb that recorded it, not the word it gives as
> its cause, so the exception is read from the verbs. Nothing decided changes: ADR 0230's acts stay
> exempt, and pushes by hand — what roll-out by default (ADR 0236) exists to end — still count.
## 8. Staged core upgrades and rollback (rule 8)
**Health, per core component**, as probes in the registry (§4):
| Component | Healthy when | Witness that rolls it back |
|---|---|---|
| controller | holds the lease within 60 s of starting; `status` answers in full within 10 s; `doctor` ran once | the node-engine on the control node, which keeps the previous controller build installed beside the new one and reads the lease bucket |
| node-engine | has reported its current declaration under its own build | its launcher (ADR 0141), which keeps the known-good |
| node tools | announced, and answer a ping within 5 s | the node-engine, which keeps the known-good |
| bus | every stream and durable consumer present (D6, D7); a request/reply round trip from every machine | none — a planned step, below |
**The gate.** ADR 0218's first machine for a core component is judged by that component's health
probes passing three consecutive times over two minutes, within ten minutes of the apply. Only then do
the other machines follow. "Reported applied" is not enough.
```
merge ─► plan ─► first machine applies new build ─► health probes ×3 within 10 min?
├─ yes ─► the rest follow, one tier at a time
└─ no ──► witness restores previous build
─► condition core.<component>.<node>.rolled-back (urgent)
─► plan halts that component, records why
```
- **Every core rollout leaves a record** in its plan: component, first machine, from and to build,
verdict, time to verdict, rolled back or not. Read through `plans`.
- **A rolled-back build is not retried** by the same plan. A newer merge makes a new plan.
**As built** ([ADR 0236](../../02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)):
- **Every release plan is gated, not only the core's.** A module is judged by its own health: it reported
applied, no witness put it back, no condition was raised since its send about the machine or about the
module there, and its tools are served on that machine where it has tools. Three healthy judgings at
least forty seconds apart and two minutes after the send, within ten minutes of it. A policy of
*together* is not gated.
- **The rollback is the ordinary path once per build.** The build is marked failed at its gate (never
registered again, `plans retry` refuses it); the module's registered build goes back to the build the
first machine ran before, still kept by ADR 0189; that machine is sent it. The verdict is written before
the send. Said as `build.<module>.<machine>.rolled-back` (warning) or `core.<component>.<machine>.
rolled-back` (urgent), `rollback-failed` (urgent, the operator's) when nothing could be put back, and as
the event `rolled-back`.
- **The witness, as built on the host** (mesh-host `internal/witness/contract.go`, the controller's half
`internal/lease/witness.go`): the controller's is the lease alone — the holder on this machine, taken
since the start, renewed within fifteen seconds, within sixty seconds of the start — and the node tools'
is their answer to the services protocol's ping within five seconds, within sixty. The controller's
"status in bound, doctor ran once" is the controller's own word in the lease's value, which the gate
reads and the host does not: a controller that holds the lease and never becomes ready is put back by
the gate sending the previous build. A witness says its verdict in `rollbacks` on every report while it
stands; the controller raises `core.<component>.<machine>.<outcome>` from it — urgent for rolled-back,
not-reversible, restore-failed and halted — and clears it with the first report without it.
- **No build reaches a machine without a gate**: a gated send carries and judges everything waiting on
its machine; every other send — a plan's rest, a cascade, a healer's, a whole-mesh push — is refused
or leaves the machine while a build no gate has seen waits there; a rebuild with the same artifacts and
manifest is no move. What waits is walked by a **release plan**, one machine at a time, the control
node last, each judged; one that fails holds the next until `upgrade release-backlog --why`.
- **A new controller that passes its gate sends the bus's machine the user list it composes**, when that
changed and nothing held back would go with it.
- **The default policy is to roll** (ADR 0236 §4); `record` stays where a module says why, keeps
irreplaceable data, or is the bus.
**The bus is a planned step.** A bus upgrade is a maintenance step a person starts through the
controller: the streams are snapshotted; a `bus-maintenance` condition is open for the step's duration;
the bus is replaced; afterwards D6, D7 and the round trip must pass, or the step is reported failed and
the snapshot is the way back. A step whose new version cannot be reverted (the bus's 2.10 → 2.11 is one)
says so before it starts and runs only on a person's explicit word, recorded as a hand act. Whether the
bus becomes a cluster that can be upgraded live is left to its own effort.
As built (ADR 0236): the verb is `bus upgrade --why … --reversible|--irreversible`; the snapshot is the
bus machine's backup holder backing the bus module up now ([ADR 0235](../../02-DECISIONS/0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md)),
or one a person took and names with `--snapshot-taken`; the step's bound is fifteen minutes; **no send
reaches the bus's machine while a new bus build waits for it** — a push, a plan's send for another
module, a healer's — after a plan's send restarted the bus unasked on 2026-10-06; a change of the bus's
user list is reloaded in place, not a restart. *2026-10-06:* the snapshot
and the way back from it exist — the bus image's own snapshot program, the same one the night's backup
runs, and a restore that builds a new store beside the live one for a person to swap in
([ADR 0235](../../02-DECISIONS/0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md),
[to-be 43](43-backups-against-mistakes.md)).
## 9. Before merge: facts and replays (rule 9)
**The facts snapshot.** The controller exports daily, and after any change of machines, assignments or
seats: every machine (a stable pseudonym of the same length as its name, its role, operating system,
C library, architecture, node-engine and node tools builds), assignments, settings keys and their
non-secret values, seats and their holders, the catalogue commit, and the versions the mesh runs of
the bus server, the store and the node-engine. No secret and no address: an address is replaced by one
from a documentation range. It is kept in the artifact store as `facts/latest`, where the build seat
reads it.
**The merge gate.** In mesh-controller, mesh-host and mesh-catalog, a check composes every machine of
the snapshot with the change applied and runs the node-engine's validator over each. A change that
makes any machine fail to compose or validate fails its check, naming the machine's role and the
module. The resolver module's tests run under both C libraries the snapshot lists. A test in each core
repository asserts that a pinned dependency's version equals the version the snapshot says runs.
Revision, [ADR 0238](../../02-DECISIONS/0238-a-commit-is-the-build-at-hand-one-commit-one-change-plan-checked-off-the-trunk-and-published-only-on-it.md)
(2026-10-06). **The commit is the build at hand.** A commit off its module's trunk is only checked; a
commit on it is published and pushed, and the controller refuses to register a build of one off it. **The
module graph, not the repository, decides what a pull request's check runs**: every pull request the
forge holds is announced, and the controller asks the planner — the one function that maps a changed file
onto modules, and the one answer of what a merge moves and builds — what a merge of the head would reach.
A changed file touches exactly the modules whose build reads it; a file no build reads touches nothing. A
change reaching a module, or adding one, runs **the gate** on the build seat (`mesh/merge-gate`): the
touched manifests through `module check`, failing only what the change brings; every machine composed with
the definitions of the modules the plan moves or adds; the replays. Every repository of the mesh runs **its
own** `merge-check.sh` (`mesh/repo-check`) in the toolchain it declares, a warning when it has none. The
check's result is the commit's **change plan** — the build plan, the deploy plan machine by machine with
what is not an ordinary send, and the verdict — posted on the pull request. The plan is one object with
a state machine, kept, followed by the release and written to the commit as a note; which parts are built
is Phase 5's.
**The replays.** mesh-lab carries a scripted scenario for each core incident, asserting the rule's
outcome, run on every merge to mesh-controller, mesh-host and mesh-tools:
| Replay | Incident | Asserts |
|---|---|---|
| R1 | two controllers at once (204) | the second waits; nothing from the stale epoch is applied |
| R2 | a reconcile due during a push (257, 261, 267) | one apply, one report, the newest sequence |
| R3 | the node-engine self-updates during its report (230, 264) | the report arrives under the new build |
| R4 | the bus's authorization reloads during a call (265) | the call's outcome is readable by id |
| R5 | a consumer with several filter subjects under mixed traffic (266) | no announcement is skipped; S5 fires if one is |
| R6 | an unreadable contributions file (241) | refused by name; nothing retired; a condition |
| R7 | the controller rebuilds itself mid-plan (214) | the plan continues under the new epoch |
| R8 | a broken controller, node-engine and node tools build | each rolled back with no hand; condition and message |
| R9 | each signal of §3 suppressed | its condition within its bound, cleared on return |
A new core issue resolves with its replay added, or with a stated reason none is possible. The live
mesh stays the test bed ([ADR 0149](../../02-DECISIONS/0149-the-live-mesh-is-the-test-bed.md)): a
replay covers what must not be done to it on purpose, and every rule keeps a live check.
**As built** ([ADR 0237](../../02-DECISIONS/0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md)):
- **The snapshot** is composed every ten minutes and kept as `facts:latest` when it moved or is a day
old; the one it replaced is let go of, so the store keeps one. It carries, beyond the list above, each
machine's roles in words, its reported capabilities, pins, settings, the names of secrets a person gave,
and how its declaration composes today; every module the mesh holds as its manifest, with its source
and build edges. `facts`, `facts compose`, `facts export`.
- **The gate is the controller's `merge-gate`**: the mesh as the snapshot says it is and the mesh with the
change, each raised in a throwaway store through the controller's own records and every machine composed
twice and validated; what composes in the first and not the second is the change's. It also fails a
manifest the judging controller cannot read (version skew), a consumer left out of a grant, a module
removed while a machine runs it, two definitions of one name, a stored setting the change cannot keep,
and a new module the node-engine would refuse on the first machine that could run it; it warns when a
merge would rebuild more than twelve modules, saying when shared code is why (issue 278).
- **It runs on the build seat**, asked by the controller when the forge's announcer says a pull request's
head moved: the repository's own `merge-check.sh` in the mesh's Go toolchain with no container runtime
socket, beside a throwaway store and bus of the versions the snapshot says run, then mesh-lab's replays.
The verdict is the head commit's status `mesh/merge-gate`, an error never a pass. A catalogue change is
judged by the controller the mesh runs; a node-engine change by it built with the change's validator.
- **Replays live where their incident is**: a component's logic in its repository as `TestReplay<issue>`,
what the mesh runs (the bus's release, the resolver under musl and glibc) in mesh-lab's `replays/`, all
named in its register and proved by `replays/cmd/prove`. A core issue opened from 2026-10-07 resolves only
with `replay:` or `replay-none:` (`cycle.py`).
- **The controller's tests run a bus of their own per test**, linked in at the release the mesh runs and held
to it by a test, so the suite runs in parallel and under the race detector.
---
## 10. The phases
Ordered by risk removed per day: detection first, because it covers every class including those not
met yet. A phase is done when its *done when* holds; this section records the date when it does.
### Phase 0 — Finish what is in flight
| Repository | Delivers |
|---|---|
| mesh-controller | the located fixes of 244, 265, 266, 267 rolled out; `calls` moved into `mesh-controller_calls`; `status` answered from a summary the event loop keeps current, inside ten seconds; the hand-act log with `--why` on the repairing verbs and `hand-act record`; recording the durations the bounds come from (apply duration per machine, heartbeat gaps, plan tier durations, build durations) |
| mesh-host | the fixes of 264 and 257/261 rolled out to every machine |
| mesh-tools | the console passing a mesh seat's `node` (244) rolled out |
| mesh-catalog | the bus's 2.10 → 2.11 upgrade, done as the first planned bus step by hand (snapshot, announced, checked after), recorded as a hand act |
**Done when:** the four located core issues resolve with their live checks; a controller restart keeps
every call's outcome; `status` answers in full within ten seconds five times in a row; the hand-act log
has a week of entries; the durations are recorded for every machine.
### Phase 1 — The mesh says when it is wrong
| Repository | Delivers |
|---|---|
| mesh-controller | the condition store, its verbs, history and events (§2); ADR 0224's standing moved into it; watchdogs for S1–S11 and S13 with bounds set from Phase 0's durations (S12 is Phase 2's, S14 Phase 5's); the bus advisories subscribed and translated (S9); `doctor` with D1–D4, D6–D10, DW and its heartbeat (§4; D5 is Phase 2's); `status` led by open conditions; the test generated from the signals table |
| mesh-host | the heartbeat carries its interval; D1's validator published as a package the controller imports |
| mesh-tools | the node tools' heartbeat (S11) |
| mesh-catalog | the `operator-channel` seat and its holder; the Telegram channel and the desktop notifier contributing to it; `mesh-watcher` on a machine other than the control node (§5) |
| mesh-lab | R9: each signal suppressed in turn |
**Done when:** on a lab mesh, suppressing each signal raises its condition within its bound and sends a
message; restoring it clears both. Stopping the controller makes the watcher send *self-check silent*
within twice the self-check's interval. Live: a week of conditions read back, every one real or its
bound corrected in the table.
### Phase 2 — Order and one writer
| Repository | Delivers |
|---|---|
| mesh-controller | the lease and epoch (§6); plans written by compare-and-set; a report kept by sequence; abandoned calls marked; S12, S13, D5; the writers table enforced at grant composition; contract tests for every consumed subject and the check listing them; the empty-on-error lint |
| mesh-host | one apply queue; the report sequence kept on disk; epoch refusal reported; the `report` verb; the contract tests and lint |
| mesh-tools | refusing an unreadable or unknown input by name; the lint |
| mesh-catalog, mesh-sdk | retirement in the providers' loop (ADR 0230, replacing ADR 0229's brake): a consumer no longer asked for is retired — disabled, marked, its data kept; or, by a provider that cannot disable, only marked, its access kept — only once the same result holds for five passes and ten minutes; `remove` is never called to retire; a set of more than three, or more than half of those held, waits for a person and raises a condition; the four retirement tools on every provider |
| mesh-lab | R1, R2, R6 |
**Done when:** R1 ends with nothing from the stale epoch applied, R2 with one report of the newest
sequence, R6 with nothing retired; every consumed subject has its contract test.
**Built 2026-10-06** ([ADR 0229](../../02-DECISIONS/0229-the-cores-order-is-a-lease-the-store-remembers-and-an-epoch-a-machine-is-sent-once-it-reads-one.md)):
the controller's lease and epoch, plans by compare-and-set, accounts kept by order, S12, S13 by writer,
D5, the writers table enforced at composition, a contract for every consumed kind, the empty-on-error
lint; the node-engine's one apply queue, report order and epoch refusal; the withdrawal brake in the Go
providers' loop and the SDK's. **Not yet:** the replays R1, R2 and R6 in mesh-lab; mesh-tools' refusals
and lint; the lint in mesh-host; the SDK's loop announcing a provider's standing, without which its
brake is said in its journal only.
**Amended 2026-10-06** ([ADR 0230](../../02-DECISIONS/0230-a-consumer-the-mesh-stops-asking-for-is-retired-and-deleted-only-by-a-person.md),
the operator's decision): the brake's hourly release is replaced by retirement — three states (active,
retired, deleted), five stable passes and ten minutes before a consumer is retired, a provider that
cannot disable marking only (every TypeScript provider today), a person for more than three or more
than half, `retire approve|reject`, `cleanup list|delete`, `retirement` events and D11. Built on branches
in mesh-catalog (both Go providers), mesh-sdk (0.1.12) and mesh-controller, not yet merged. **Not yet:**
the TypeScript providers giving their adapters a `retire` that disables (until then each is mark-only:
its retired consumers keep their access), passing their module name and an announcer to the loop,
without which a set over the bound there waits with no way to approve it, and adapters that list and
delete what they retired.
### Phase 3 — Healers
| Repository | Delivers |
|---|---|
| mesh-controller | the healer registry and H1–H4, each braked and said; S15 |
| mesh-lab | an induced failure per healer |
**Done when:** a lab mesh recovers from each induced failure with no hand, says so, and brakes after
its budget. Live: a week with no cause repeated in the hand-act log.
**Built 2026-10-06** ([ADR 0231](../../02-DECISIONS/0231-a-healer-acts-on-what-observation-raised-and-only-observation-says-it-worked.md)): the healer registry
and H1–H4 in the controller, H5 registered as the provider's own; the mesh-wide brake; `healer-acted`;
`healers`; S15 live; the node-engine's `report` verb. Each healer's induced failure runs in the
controller's tests against a real store, and H3 and H4 against a real bus: induced, repaired, cleared by
the probe's own next run, and handed to the operator after the budget. **Not yet:** the induced failures
on a lab mesh (mesh-lab), and so the phase's *done when*; the live week.
**Amended 2026-10-06** ([ADR 0233](../../02-DECISIONS/0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md),
the operator's decision, before Phase 4): a module declares the data it holds; D13 and its conditions;
the `data` verb; `cleanup` extended to a module's own retired data, deleted by the machine's backup
holder after a last restore point; the controller's grant names `node-backup.backed-up`. Built on
branches in mesh-controller (migration 0072), mesh-host (the installer's user list; the kept-directory
report), mesh-catalog (the holder's measuring and deletion, every module's data, the providers' held
sizes), the media catalogue and the photo application, not yet merged. **Not yet:** the TypeScript
providers saying their consumers' sizes (until then a consumer at two of them is `data-held-twice`, a
warning, not compared), and SMART read under an array.
### Phase 4 — Core upgrades that roll back
| Repository | Delivers |
|---|---|
| mesh-controller | the health probes of §8; the gate on a core component's first machine — as built, on every plan's first machine (ADR 0236); the rollout record in the plan; the bus maintenance step as a verb; the default upgrade policy `roll` (ADR 0236) |
| mesh-host | keeping the previous controller and node tools builds; restoring one when its health is not met in bound, watching the lease bucket for the controller |
| mesh-tools | answering the health ping |
| mesh-lab | R3, R7, R8 |
**Done when:** on a lab mesh, a broken build of the controller, the node-engine and the node tools —
one that starts and does nothing, one that crashes, one that cannot reach the bus — is each rolled
back with no hand, the mesh ends on the previous build, and a condition and a message say so. Live: the
next three core rollouts each record a health verdict.
**Built 2026-10-06** ([ADR 0236](../../02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md),
the operator's "start phase 4"): in mesh-controller (migration 0073), the probes H-controller, H-engine,
H-tools, H-bus, DG and DB; the gate on every plan's first machine, its record in the plan (`plans <id>`)
and in the store; the rollback by the ordinary path, once per build; the witnesses' verdicts as
conditions; the grants the witness reads; `bus` and `bus upgrade`; the default policy `roll`, `upgrade`
listing every module's policy and where it comes from; the user list carried after a controller passes;
a module deleted at its source forgotten, not built. In mesh-host, the witness (keeping the previous
controller and node tools, restoring one not healthy in bound, the launcher's for the node-engine) and the
installer's grants. In mesh-catalog, the modules that keep `record` say why, and the forge's announcer
says which files a merge deleted. On branches, not yet merged. **Not yet:** R3, R7, R8 on a lab mesh, and
so the *done when*; container state in the node-engine's report, without which a container that
crash-loops after its compose applied is seen only through what it breaks; *(designed 2026-10-07 by
[ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md) as
[to-be 48](48-a-module-says-how-it-is-healthy.md), whose phases track it)* composing `not-reversible` for
a build whose migration cannot be undone; the next three core rollouts' verdicts.
### Phase 5 — Checks before merge, and the replays
| Repository | Delivers |
|---|---|
| mesh-controller | the facts snapshot and S14 |
| mesh-controller, mesh-host, mesh-catalog | the compose-and-validate merge gate; versions tested as run; the resolver's tests under both C libraries |
| mesh-lab | R4, R5 and the rest of the window's incidents; running the replays on every core merge |
**Done when:** the replays of 236, 262, 263 and 266 fail on the commit before their fix and pass after;
a new core issue cannot resolve without a replay or a stated reason.
**Built 2026-10-06** ([ADR 0237](../../02-DECISIONS/0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md),
the operator's "start Phase 5"), on branches, not yet merged: in mesh-controller the facts snapshot and
S14, `merge-gate`, the check asked of and run by the build seat, `merge-check.sh`, the replays of 263 and 273,
and a bus per test at the mesh's release; in mesh-catalog the forge's announcer of pull requests' heads and
the verdict as their status, and `merge-check.sh`; in mesh-host `merge-check.sh` and the grants in the
installer's user list; in mesh-lab the replays of 262 and 266, the register and the prover; in hq the
replay rule in `cycle.py`. The prover ran the replays of **236, 262, 263, 266 and 273: each fails on the
commit before its fix** (236's check did not exist there) **and passes on it**. **Not yet:** R4, R5 and the
rest of the window's incidents as replays; the merge gate's first runs against the live mesh's snapshot;
the status required to merge, which is the operator's setting on the forge; mesh-tools' and mesh-sdk's
`merge-check.sh`.
**Revised 2026-10-06** ([ADR 0238](../../02-DECISIONS/0238-a-commit-is-the-build-at-hand-one-commit-one-change-plan-checked-off-the-trunk-and-published-only-on-it.md),
the operator's "enable the checks on every repository of the mesh", "the logic acts on our module graph"
and "the commit is the build at hand"), on branches, not yet merged:
| Repository | Delivers |
|---|---|
| mesh-controller | the planner's one mapping of a changed file onto modules and its one answer of a merge's reach, asked by the merge handler, the what-if, the gate and the check; a file no build reads touches nothing (issue 280's *left open*); every pull request answered — the gate when it reaches a module, the repository's own check for the mesh's repositories, a pass that says so otherwise; the gate run by the build seat with its judge chosen from the graph, composing the plan's definitions only, failing only a manifest problem the change brings; the change plan computed and carried with the verdict; a build of a commit off its module's trunk recorded and never registered; its own manifest naming every verb of its seat |
| mesh-catalog | the forge's announcer saying, with every pull request, the directories holding a module at its head, the files it deletes and whether it has a `merge-check.sh`; both statuses and the change plan on the pull request; `gitea_branch_protection_get` and `gitea_branch_protection_set`; the catalogue's own check |
| mesh-host, mesh-tools, mesh-sdk, mesh-lab, mesh-media-catalog, and the applications built from their own repositories | each its own `merge-check.sh`; mesh-tools' tests each on a bus of their own |
| mesh-tools-go, mesh-tools | a C compiler and Python in the Go toolchain, git in the TypeScript one |
| hq | its own `merge-check.sh`, running `records.py`, `index.py` and `cycle.py` |
**Not yet:** the change plan kept (id, inputs, result) and the release comparing its plan with it; the
state machine replacing the release plans' states, with its table, its test, its conditions and healer H2;
the status's link to the kept plan; the notes under `refs/notes/mesh-plan`; check builds in a scratch
namespace (the check builds no module today, so there is none to keep apart yet); the statuses made
required — prepared as calls of the forge module's tool, applied by the operator.
**Given an owner 2026-10-06** ([ADR 0239](../../02-DECISIONS/0239-a-delivery-is-owned-by-the-mesh-delivery-module-and-runs-from-commit-to-delivered.md)): the state machine, the
notes and the status's link are built as a delivery's, owned by the module `mesh-delivery`. They are designed
in [to-be 47](47-delivery-from-commit-to-delivered.md), and their progress is tracked there.
## What is not decided here
- The bus as a cluster of three, to upgrade it live.
- Routing by presence, quiet hours, answering back through a channel, and an external dead-man
service — decided by [ADR 0234](../../02-DECISIONS/0234-the-mesh-holds-a-conversation-with-its-operator.md),
designed in [to-be 46](46-the-conversation-with-the-operator.md).
- A condition that needs judgement handed to an agent as work — research 017.
@@ -1,508 +0,0 @@
---
layer: to-be
status: designed
code: []
updated: 2026-10-06
decisions:
- 02-DECISIONS/0234-the-mesh-holds-a-conversation-with-its-operator.md
- 02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md
---
# 46 — The conversation with the operator
**The mesh, its modules and its agents tell the operator things and ask them things; the operator
answers, or writes first. Each exchange travels on a channel whose declared capabilities can carry it,
chosen by where the operator is working. An answer that performs an action is checked and performed by
the controller, never by the channel or the asker** ([ADR 0234](../../02-DECISIONS/0234-the-mesh-holds-a-conversation-with-its-operator.md),
from [research 028](../../01-RESEARCH/028-the-meshs-output-channel/00-overview.md)).
This design grows the minimal output channel of [to-be 45](45-a-core-that-cannot-fail-silently.md) §5.
That form stays the first step; §13 below is the order in which it becomes this one.
## The parts
```
sources ROUTER (holder of operator-channel) channel bench (out)
─────── ─────────────────────────────────── ───────────────────
controller's conditions ──┐ messages and asks ┌──► channel/telegram
modules (notify) ────────┼──► required capability → work context ──┼──► channel/desktop
agents (ask) ────────┤ → severity → escalation └──► channel/<next kind>
controller (authorise │ presence (current state only)
request) ────────────────┘ ▲ ▲
│ │ answer (ordinary asks) intake bench (in)
seen ───┘ └────────────────────────────── intake/telegram
intake/desktop
CONTROLLER ◄──── authorise answer (authorising asks, intake holders only) ─┘
│ checks capabilities, identity, code, digest; performs; records the hand-act
└──► ask-answered
watcher's watcher (outside the seats, own bot, not on the control node) ──► Telegram directly
self-check and watcher ──► outside dead-man service
```
## 1. The three things said
- **A message:** the mesh tells. A condition raised, a reminder, a clearing, a notice from a module.
No answer is expected. It may be edited later.
- **An ask:** someone wants the operator's input, of a declared kind (§6). The answer goes back to
whoever asked.
- **An operator message:** the operator writes first — to an agent, to the mesh, or answering an ask in
its thread instead of tapping.
Input from outside that is not the operator (a mail, a webhook) is the same envelope (§3) with
another sender. Only its shape is designed here; its consumers are later work.
## 2. The `channel` seat (out)
**A mesh seat, kinded bench.** A kinded bench is the second sort of bench, beside the replicated one
([ADR 0223](../../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md)): its
holders are different modules, each claiming one **kind** (`telegram`, `desktop`, later `matrix`,
`ntfy`, `mail`, …). Two claims of one kind are refused at registration. A verb's subject carries the
kind, as a node seat's carries the machine. A new channel is a new module claiming a new kind, with no
change to the router. `channel` and `intake` are the only kinded benches.
**A claim on it carries** `kind` and `capabilities` (§7).
**Served:**
- **`send`** — the words, a priority, whether silent, an optional ask block (the ask's id, its kind,
the options each with an opaque token) and an optional thread (a conversation handle, or the message
this replies to). Answers the channel's own id for what it sent, or a refusal in words.
- **`edit`** — replace a sent message by its id, where `edit` is declared.
- **`standing`** — ready; not configured (naming what is missing, never a value); or failing (since
when, why); the last delivery; the declared capabilities and the result of each drill.
**Emitted:** `delivered`; `failed`, marked **permanent** or **transient**.
**Honoured by every holder:** its declared maximum length, by cutting and saying so; silence where
declared; an identical edit is a success; no secret of its own in any error or event.
## 3. The `intake` seat (in)
**A mesh seat, kinded bench**, as `channel`. A service read and written by one program (a Telegram
bot) is held by one module claiming both seats under one kind.
**A holder turns what arrives into one envelope:**
| Field | What it is |
|---|---|
| `id` | unique, for deduplication |
| `kind` | the holder's kind |
| `what` | `message` (written first), `choice` (a button or reaction), `reply` (written in an ask's thread), `mail`, `call` (a webhook), `seen` (activity, for the work context) |
| `sender` | the identity on that service, and whether the service authenticated it |
| `trusted` | whether the sender is on the **controller's** list of the operator's identities. A holder never sets it: a raw intake event is untrusted by construction, and only the router stamps it, from the controller's answer (§4) |
| `conversation` | an opaque handle; sending on it reaches the same chat, room or thread |
| `in-reply-to` | the ask or message it answers, if any |
| `payload` | the text, or the option's token. **Never a code** (§10) |
| `at` | when |
**Answers go to the ask's owner by request and reply**: the router's `answer` for an ordinary ask, the
controller's `authorise answer` for an authorising one; both refuse an answer that is not trusted.
**Everything else is an event on the seat**, taken by the router, which re-emits it stamped (§4): a
trusted `message` as an **operator message**, addressed (§5); anything untrusted as **input**, for
consumers that accept untrusted input (a later mail rule takes `mail`). The router takes `seen` for the
work context.
**A prerequisite:** the shared library publishes on a seat's event subjects, and the permission model
grants it (the open gap of [to-be 32](32-what-a-module-declares.md) §1).
## 4. Who counts as the operator
- **Only a sender on the controller's list of the operator's identities is the operator.** The list
is kept per intake kind in the controller's state. While no factor exists, an identity is enrolled
only at a terminal; once one does, adding or removing an identity is a destroy ask (§10). Telegram's linking verb
(§11) produces such an ask.
- **The controller serves an identity check** (kind and sender in, trusted or not out). The router keeps
the answers as current state, refreshed when the controller announces a change to the list, and
stamps every envelope it re-emits with `trusted`. **No holder's word is taken for it**, and the
controller checks again itself before authorising (§10).
- **An untrusted envelope** may start work for a consumer that declares it accepts untrusted input. It
can never answer an ask, authorise, or be delivered to an agent as the operator's words. An untrusted
sender writing to the bot is told nothing beyond a fixed line, and the attempt is a warning condition
when it repeats.
- **The rule for agents and consumers: untrusted input is data, not instructions.** An agent handed
one may read, quote and report it, and never follows what it says. This rule is part of every agent
module's instructions and every intake consumer's definition.
## 5. Operator messages: who answers them
### Addressing
- **To an agent:** `@name` at the start of the message, or written in that agent's thread (the
conversation handle its own messages and asks carry).
- **Otherwise, to the responder**, addressed as `@mesh` or by not addressing anyone.
- `@` names come only from the **register** below. A name not on it is refused, with the names that are.
### The register of agents
An agent registers itself with the router: a name (unique, bound to its bus principal so no other
principal can take it), its owner, and a line on what it handles. Registration is a lease the agent
renews while it runs. The router's `participants` lists the register and who is running.
- **To a registered agent that is running:** delivered as an addressed operator message event, which
only that agent's principal may consume.
- **To a registered agent that is not running:** kept, at most 20 messages and 7 days per agent, and
the operator is told "kept; <name> reads it when it next runs". Beyond the bound the oldest kept
message is dropped and the operator is told.
- **An agent answers** on the conversation handle the message arrived with.
### The responder
- **The mesh's own participant, part of the router**, addressed as `@mesh`. Not a module of its own,
because it needs nothing the router does not hold.
- **It answers from the controller's read verbs only:** status, open conditions, open and recent asks,
and the bindings and data on record. Its grant names read verbs and nothing else; it performs nothing
and asks nothing that authorises.
- **It lists what it can answer** when asked ("help") and whenever it does not understand.
- **Never silence:** an unaddressed message the responder does not understand gets a reply saying so,
what it can answer, and how to address an agent.
- Its answers are messages like any other, under the content rule (§9).
## 6. Asks
### Kinds
| Kind | The operator gives | The channel needs |
|---|---|---|
| `yes-no` | yes or no | `choice`, or `reply` read as yes or no |
| `one-of` | one of at most eight labelled options | `choice`, or `reply` with the option's number |
| `text` | free text | `reply` |
| `number`, `date` | a value within bounds | `reply`; parsed by the router, asked again once if it does not parse |
| `acknowledge` | "seen" | `choice` |
An **authorising** ask is any of these with the authorising flag; it adds the tier's requirements (§10).
`text`, `number` and `date` never authorise.
### What an ask carries
An id; the asker (bus principal and machine); the kind with its options or bounds; the words; a
priority (urgent or normal); optionally a timeout and a default; optionally a conversation handle;
optionally a group, for batching. The words pass the content rule.
### Life
**open → answered | defaulted | expired | cancelled.**
- **Answered:** the first answer wins. Every other copy is edited to say where it was answered.
- **Defaulted:** the timeout passed and a default was declared. The asker receives it marked as a
default, never as the operator's answer.
- **Expired:** the timeout passed with no default; the asker is told. **An authorising ask never
defaults; it expires** (ADR 0230: a timer is the mesh acting alone again).
- **Cancelled:** by the asker (`ask cancel`), or by its owner when moot (the condition behind it
cleared). Copies are edited to "no longer needed".
### Limits, batching, history
- An asker holds **at most three open asks**; a fourth is refused in words.
- Asks count against the router's hourly cap, as messages do; answers do not.
- Asks to one channel within the burst window go out together under a heading ("3 questions waiting"),
each as its own message, answerable and editable alone.
- **`asks`** lists open asks and closed ones for **30 days**: asker, kind, outcome, channel, answer. A
free-text answer stays in the router's state and is never forwarded to another channel.
### How the asker gets the answer
The event **`ask-answered`**, with the ask's id and any conversation handle. An asker may also call
`ask` with a bounded wait (a few minutes), or poll `asks <id>`.
### The router's verbs and events
`ask`, `ask cancel`, `asks`, `answer` (intake holders only), on the `operator-channel` seat beside its
existing `notify`, `open` and `history`. Events `ask-opened`, `ask-answered`, `ask-closed`. The router
owns every ask that authorises nothing; an authorising ask is owned by the controller (§10) and carried
by the router like any other.
## 7. The capability vocabulary, `channel-capabilities/1`
**Fixed and versioned.** A word outside it is refused at registration. **Each word has a contract test**
the holder's build runs, and a **drill** `standing` can run. A capability that fails its drill is
withdrawn from routing and reported until it passes. A new word is a new version of the vocabulary.
| Group | Capability | Promise | Contract test / drill |
|---|---|---|---|
| delivering | `deliver` | it arrives, or `failed` says why | against a test double / a live test message |
| | `reaches-away` | it reaches a phone away from the operator's machines | by kind / the operator acknowledges a drill |
| | `loud` | it can break through the phone's quiet hours | the service's override is set for urgent |
| | `silent` | it can arrive without sound | the service's silent flag is set |
| | `edit` | a sent message can be replaced in place | edit and read back, against a double |
| | `max-length:N` | up to N characters arrive whole; longer is cut and the cut said | N+1 characters give a cut message, not a failure |
| | `reaches-when-mesh-down` | delivering needs neither the bus nor the control node | the holder's placement and send path |
| | `private` | the words stay on the operator's machines, or are end-to-end encrypted | by kind; reviewed |
| conversing | `choice` | one offered option in one act, and the pick comes back | a simulated tap yields a `choice` envelope with the option's token |
| | `reply` | free text comes back | a simulated reply yields a `reply` envelope |
| | `threads` | an answer is tied to what it answers | a reply to A carries A in `in-reply-to` |
| | `operator-first` | the operator can write unprompted | a simulated message yields a `message` envelope |
| trusting | `verified-sender` | the answer came from the operator's own account, through a holder no agent shares | a choice from an identity not on the list is dropped and reported; placement checked |
| | `exact-render` | the ask is shown as the controller rendered it | rendered text equals the controller's, byte for byte |
| | `code-factor` | a typed code reaches the controller by request and reply, never judged by the holder | a code never appears in an event; it is deleted from the conversation where the service allows |
| | `key-factor` | a security key's assertion over the controller's challenge reaches the controller | the challenge is the controller's; the signature is checked by the controller |
What the first holders declare:
| Kind | Declares |
|---|---|
| `telegram` | `deliver`, `reaches-away`, `silent`, `edit`, `max-length:4096`, `choice`, `reply`, `threads`, `operator-first`, `verified-sender` (only while placed where no agent runs as the operator), `exact-render`, `code-factor` |
| `desktop` | `deliver`, `loud`, `silent`, `edit`, `private`, `choice`, `reply` (through the launcher's prompt), `threads`, `exact-render`, `code-factor`; `key-factor` only where a key is enrolled and present. **Never `verified-sender`.** |
## 8. Routing, the work context and presence
### The rule
**A message or ask goes to the most direct channel in the operator's current context, among those whose
capabilities already satisfy it. Unanswered in time, it escalates along a fixed chain. Context orders
the candidates; it never adds one, and never lowers the bar.** When nothing qualifies, the router raises
a condition of its own saying so.
### The signals
| Signal | Source | Read as |
|---|---|---|
| A graphical session unlocked with input in the last 5 minutes on machine M | the `node-lock-screen` seat's holder on M, emitting an event on each lock and idle change (logind's lock and idle hints underneath) | at the desk on M |
| That session locked, or idle longer | the same | not at the desk |
| An agent asking from machine M | the ask's asker | strengthens "at the desk on M"; alone it proves nothing |
| A verified intake `message`, `reply` or `choice` in the last 15 minutes | the intake seat | in a conversation on that kind |
| An ask carrying a conversation handle | the ask | that conversation, whatever else is true |
| The hour, against quiet hours | the router's setting | night: only urgent wakes |
### Where things go
"The away channel" is the operator's setting, Telegram to begin with. "The loud holder" is an optional
second away holder for waking.
| Context | An ask | Urgent message | Warning | Unanswered → |
|---|---|---|---|---|
| in a conversation through Telegram | that chat, in the thread | that chat | that chat, silent | after 10 min (urgent) or 1 h: also the desk, if active |
| at the desk on M | the desk on M if it can carry it; else the away channel, and the desk says where it went | the desk on M, and the away channel silently | the desk on M | after 5 min (urgent) or 30 min: the away channel, with sound |
| away | the away channel | the away channel | the away channel, silent | after 15 min (urgent): the loud holder, if configured |
| night, away | non-urgent asks wait for morning; urgent as away | the away channel and the loud holder | the morning digest | as away |
| the desk locks while an ask is shown there | moves at once to the away channel | — | — | — |
All bounds are provisional, set as to-be 45's are: from what the first live weeks measure.
### Presence stays in the mesh
Presence facts are events on the bus. The router keeps them as **current state only** — one value per
machine and per intake kind, overwritten — never as a history, and never in a message's words. A
consumer that wants a timeline of the operator's day is a decision of its own.
### The content rule
The router's, applied before anything reaches any holder; §9 says what it allows. A `private` holder
is not exempted.
## 9. What a message may name, and references
### The rule
- **Allowed in a message's words:** roles, words, and the mesh's own machine names.
- **Refused:** domains, addresses, paths, and anything shaped like a secret — in words, labels and
references alike.
- **A refusal is said to the sender** — the `notify`, `ask` or `authorise request` call is refused naming
the offending part, so the asker can rephrase or move the specific into a reference — and raises
`channel-refused`. The offending part is never sent to a channel.
### References
- **A sender may attach details:** each a label and a concrete detail (a path, an address, a log
excerpt), never a secret.
- **The router keeps the detail** in its own state, as long as the ask's history (30 days), and puts
only an opaque reference with its label into the words.
- **`detail <ref>`**, a router verb, returns the detail. It answers **only** the console and the intake
holders of kinds declaring `private`, and its answer travels only back to them.
- **Where a dereferenced detail may be shown:** on a channel declaring `private`, and at the console.
Today that is the desk (a notification action "show detail" opens it) and the console. On Telegram, and
any channel not declaring `private`, only the label appears.
## 10. Asks that authorise
### The verb table and the controller's verbs
- **The controller's verb definition gains `authorises`:** the tier, and the arguments that make up
the exact state a person must see (for `retire approve`, the set of consumers).
- **`authorise request`** — callable by anyone (an agent, the router on a condition's behalf). Carries
the verb and its exact arguments, why, and optionally a conversation handle. The controller renders
the ask, stores it with an opaque id, its expiry and a digest of the state shown, and emits
`ask-opened` with itself as owner. **Nothing is performed.**
- **`authorise answer`** — granted to intake holders only. Carries the ask's id, the option, the
sender's identity, and the code or key assertion where the tier needs them.
- **`authorisations`** — open and recent authorising asks.
### What may be authorised, and its tier
| Action | Tier |
|---|---|
| approve or reject a waiting retirement set (`retire approve` / `retire reject`) | approve |
| confirm a kept binding moves (`pin`, while `binding-kept` names it) | approve |
| end a stuck plan; send a machine its declaration by hand; reset a bus consumer; retry a healer | approve |
| silence a condition | acknowledge |
| delete a retired consumer's data, one or by age (`cleanup delete`) | destroy |
| change the operator's identities, the away channel, or a factor's enrolment | destroy |
### Proofs and tiers
- **P1, a verified sender:** a Telegram tap from the operator's own account, through a holder placed
where no agent runs as the operator.
- **P2, a TOTP code** from the operator's authenticator app, verified by the controller. Valid for the
current or previous 30-second step; accepted once. **This is the proof everything above acknowledge
can rest on.**
- **P3, a security key's touch bound to the ask** (the challenge is a hash of the ask's id and the
state digest; user presence required). **Optional:** counted where a key is enrolled; required by no
tier.
| Tier | Required | Telegram (away) | Desk |
|---|---|---|---|
| acknowledge | `choice`, `exact-render` | tap | click |
| approve | + one proof | tap (P1) | click + code (P2) |
| destroy | + two proofs, at least one P2 | tap + code (P1 + P2) | click + code + key touch (P2 + P3), only with an enrolled key; otherwise carried by Telegram |
- Approve: single use, bound to the exact state shown, expiring when that state changes or after 24 h.
- Destroy: valid 10 minutes after it is shown; at most one destroy answered per 10 minutes.
- **A desk click alone never authorises.** An agent at the operator's terminal answers nothing; the
console offers only break-glass with a code.
### The rules
1. Every authorising verb declares its tier in the verb table.
2. An authorising ask is offered only on a channel whose capabilities satisfy its tier. An answer from
any other channel is refused, and the refusal is said there.
3. **The away channel satisfies every tier.** The self-check verifies it every run. A setting making a
tier possible only at a desk is refused, unless the operator chose that for the tier explicitly.
4. The work context chooses among channels that qualify; it never makes one qualify.
5. **No agent authorises.** An agent asks. Agents' grants hold no authorising verb.
6. **Only a trusted sender authorises** (§4); an untrusted answer is refused before any other check.
### The controller's checks, in order, refusing at the first failure
1. The ask is open and unexpired.
2. The caller holds an intake kind — by the controller's own seat records, never the request's claim.
3. That kind's declared capabilities, from the controller's records, satisfy the tier, and the proofs
present are enough.
4. A P1 answer's sender is on the controller's list of the operator's identities for that kind.
5. A code is valid for the current or previous step, and unused.
6. A key assertion verifies against the enrolled credential, over this ask's challenge, with user
presence.
7. The state now has the digest it had when shown; otherwise the ask is void and a new one is
requested.
Then it performs as itself, records the hand-act, closes the ask with a compare-and-set so a second
answer loses, and emits `ask-answered`. Every copy is edited to the outcome and its buttons removed.
### Direct calls
`retire approve|reject`, `cleanup delete`, and `pin` while `binding-kept` names it refuse any caller
except through `authorise answer`, or **break-glass** at the console with a code — recorded and
announced on every channel as break-glass. `retire approve` takes **`expect`**, the set it approves,
and refuses if the set now differs. `conditions silence` stays callable; a silence an agent sets is
said on the away channel, with its why.
### The hand-act
Gains **`via`** (kind and holder), **`requested-by`** (agent principal or condition key), **`ask`** (the
id) and **`proofs`** (which of P1, P2, P3). `by` reads "the operator, as <kind> identity <id>".
### The factors
The operator's identities, the TOTP seed, its recovery codes and any enrolled key are the controller's
own state. The seed is made by the mesh and shown once, as a URI, to a plain terminal — never through a
channel or an event. Re-enrolling a factor is a destroy ask, except by recovery below. A code travels
only by request and reply from an intake holder to the controller, and is deleted from the conversation
where the service allows.
### Recovering the factor
A lost phone must not lock the mesh's only person out for good, and re-enrolling is a destroy ask that
needs the lost factor's code. So:
- **`factor enrol`**, at a terminal: the controller makes the seed and **ten one-time recovery codes**,
shows them once together, and stores the codes **only as hashes**.
- **A recovery code counts as one P2 proof, once.** It may stand in for the app's code in any ask.
- **`factor recover`**, at a terminal, given a recovery code: a new seed and ten new codes replacing the
remaining ones. Recorded as a hand-act, and **announced loudly on every channel**.
- **The count left** is shown by `factor status` and in `authorisations`; the self-check warns when
three or fewer remain.
- **The last resort:** root on the control node — at it, or by its SSH key — runs `factor enrol
--break-glass` there. It talks to the controller locally, never over the bus; the same verb over the
bus is refused. Recorded and announced as break-glass.
- **No tier requiring P2 is enabled until recovery codes exist.** The controller refuses the setting.
## 11. The first holders
### Telegram — kind `telegram`, both seats, the away channel
- **Its own module**, holding the bot token and chat id as its own secrets; the router keeps no
Telegram client.
- **Long polling**, never a webhook; the update offset kept in its own state.
- **Buttons** carry only an opaque `a1:<ask>:<option>` (Telegram allows 64 bytes); everything else is
in the router or the controller. A tap is dropped and reported unless both the user and the chat are
the operator's; otherwise it is acknowledged at once (the phone shows a spinner until then), handed
on, and the message edited to the outcome.
- **A destroy ask** answers the tap with a forced-reply prompt for the code; the holder reads it,
deletes it from the chat, and hands it to the controller by request and reply.
- **A linking verb** with a one-time deep-link code replaces reading the chat id by hand. Linking an
identity adds it to the controller's list, so after the first it is a destroy ask (§4).
- **A sender not on the list** is answered with one fixed line and handed on only as untrusted input.
- **Placement:** where no agent runs as the operator. Its `verified-sender` depends on it, and the
self-check reads the placement.
### The desk — kind `desktop`, both seats
- **`node-notifier.send` gains actions** (token, label). It still answers at once; the dunst holder
listens for the chosen action and emits it as an event on its node seat.
- The router's desktop adapter, holding `channel/desktop` and `intake/desktop`, turns that into a
`choice` envelope.
- **For `text`, `number`, `date` and codes**, the notification's single action opens the
`node-launcher` seat's prompt; what is typed returns as a `reply` — or, for a code, straight to the
controller by request and reply.
- **The screen-lock holder emits lock and idle changes** for the work context.
- It carries every ordinary ask with no account anywhere, and an authorising one only with a code (or
an enrolled key's touch).
### The watcher's watcher
Outside the seats, deliberately: it must speak when the bus or the control node is what failed. A
minimal sender with its own bot, on a machine that is not the control node; it never reads and never
asks. The self-check and the watcher both ping an outside dead-man service, which speaks when both are
silent.
## 12. The defects to fix first
Found in the built Telegram code of the output seat's holder and the watcher
([028/04](../../01-RESEARCH/028-the-meshs-output-channel/04-telegram-as-the-first-holder.md)).
| # | What | Fix | Weight |
|---|---|---|---|
| D1 | nothing bounds a message to 4096 characters; one long message wedges the channel | cut at the declared maximum and say so; a refused send never blocks those behind it | high |
| D2 | a condition reopened within ten minutes is said by an edit, which rings nothing | a reopening is a new message | high |
| D3 | clearing edits only the first message; the newest still says "still open" | clearing reaches the newest message | medium |
| D4 | `disable_notification` is never set | warnings and clearings are silent; urgent rings | medium |
| D5 | a 429's `retry_after` is ignored | wait as told | low |
| D6 | refusals (400, 403) are retried as transport failures; "not modified" on an edit sends a duplicate | refusals are permanent `failed`; an identical edit is a success | low |
| D7 | a deprecated preview flag | dropped | low |
| D8 | a marshalling error is discarded | reported locally | low |
| D9 | the watcher's unsent "silent" is overwritten by its "cleared" | say both, or one line naming how long it was silent | low |
| D10 | "ready" means a token and a chat id, not that the bot can reach the operator | `standing` checks the chat at status time | low |
## 13. The build, in order
Each phase ends at its own *done when*. Owning repositories from [`repos.md`](../../00-META/repos.md).
| Phase | What | Owner | Done when |
|---|---|---|---|
| **1 — Fix the first holder in place** | D1–D4 in the output seat's holder and the watcher, then D5–D10; no change of shape. The operator then makes the two bots and the mesh is given their values through the controller | mesh-catalog | the holder's tests for D1–D4 pass; a live test message reaches the phone from both the holder and the watcher |
| **2 — The desk's actions** | `node-notifier.send` gains actions; the dunst holder emits the chosen one; the launcher's prompt returns typed text; the screen-lock holder emits lock and idle changes | mesh-catalog | a notification with two actions returns the chosen token as an event; locking emits an event |
| **3 — The seats, and the router** | the kinded bench in the seat set; `kind` and `capabilities` on a claim; `channel-capabilities/1` and its contract tests; publishing on a seat's event subjects and its grant; `channel` and `intake` seats; the output seat's holder becomes the router and holds `channel/desktop` and `intake/desktop`; presence as current state; the operator's identity list and identity check in the controller, the router's `trusted` stamping and untrusted input (§4); the content rule allowing machine names, references and `detail`, refusals naming the offending part (§9) | mesh-controller (seat set, claim fields, registration refusals, grants, the identity list and check); mesh-sdk and mesh-tools (seat-event publishing in the shared library and the node tools); mesh-catalog (the seats' definitions, the router) | the catalogue refuses an unknown capability and a second holder of one kind; a message is routed to the desk by capability and context; an envelope from an identity not on the list is re-emitted untrusted; a path is refused to its sender by name, and travels as a reference the desk opens |
| **4 — Asks, and operator messages** | `ask`, `ask cancel`, `asks`, `answer`; the events; kinds, life, limits, batching, history; escalation; addressing, the register of agents with kept messages, the responder; the untrusted-input rule in every agent module's instructions (§5) | mesh-catalog | the router tests of ADR 0234 pass; live: an agent's question answered at the desk, and with the desk locked, on the phone; `@mesh status` answered; a message to a stopped agent kept and said |
| **5 — Authorise, with TOTP** | `authorises` in the verb table; `authorise request`, `authorise answer`, `authorisations`; the seven checks; direct calls refused, break-glass; `retire approve` takes `expect`; the hand-act fields; the TOTP seed as the controller's state, enrolled at a terminal with ten recovery codes; `factor recover`, `factor status`, the local-only break-glass enrolment; no P2 tier before recovery codes exist; optional key enrolment; the self-check probe for the away channel; agents' grants lose the authorising verbs | mesh-controller; mesh-catalog (agents' grants, the desk's code prompt) | one controller test per refusal passes; at the desk, approve needs a code; a recovery drill re-enrols with a recovery code, and is announced |
| **6 — Telegram live** | the `telegram` module holding both seats, with buttons, replies, forced-reply codes and the linking verb, placed where no agent runs as the operator; the operator's two bots configured; the watcher assigned off the control node; the dead-man ping | mesh-catalog | the self-check says the away channel carries every tier; live drills: an approve and a destroy on a test condition answered on the phone, with the hand-acts read |
**Not touched:** the node engine (mesh-host repository). Presence comes from the screen-lock seat's
holder, not the engine; if a later measure shows the engine must relay it, that is a change here.
## What is not decided here
- Whether a `private` holder may be exempted from the content rule.
- Further holders: Pushover for waking, Matrix once its push is measured, mail out for the digest and
mail in as an intake; each is a module claiming a kind, with no change to this design.
- Consumers of non-operator input (mail rules, webhooks).
- Agents running under an account of their own on a display that isolates clients, which would let the
desk earn `verified-sender`.
@@ -1,191 +0,0 @@
---
layer: to-be
status: designed
code: []
updated: 2026-10-07
decisions:
- 02-DECISIONS/0239-a-delivery-is-owned-by-the-mesh-delivery-module-and-runs-from-commit-to-delivered.md
- 02-DECISIONS/0238-a-commit-is-the-build-at-hand-one-commit-one-change-plan-checked-off-the-trunk-and-published-only-on-it.md
- 02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md
- 02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md
---
# 47 — Delivery: from a commit to delivered
**A delivery is one commit in one repository, from the moment its pull request's head is announced to the
moment every machine of its deploy plan runs it, or to the moment it fails, is superseded or is stopped. A
delivery group is two or more deliveries that share a head branch name, delivered as one unit in order. Both
are owned by one module, `mesh-delivery`, which records every step and asks the controller for each act.**
([ADR 0239](../../02-DECISIONS/0239-a-delivery-is-owned-by-the-mesh-delivery-module-and-runs-from-commit-to-delivered.md).)
This is the delivery of [to-be 10](10-delivery.md) carried out. To-be 10 says what is built: the comparison
of source and artifacts decides, and nothing here changes that. This design says who answers for one
commit's journey and in what order the steps are allowed.
## The parts
```
forge's holder ──pull.updated / .merged / .closed──► mesh-delivery ◄──checked / plan-moved── controller
(statuses, the view, (deliveries, groups, (planner, gate,
the notes) the table, its state) registration,
▲ │ │ sending, walk)
└── its tools: note, view, status ◄───────────────┘ └── its verbs: delivery-plan, -order,
-check, deliver, delivery-stop,
delivery-walks
```
- **The controller** stays the only writer of what it owns: machine declarations, bus objects, grants,
registration. It keeps the primitives: the planner (one function, ADR 0238 decision 3), the gate, the trunk
rule, a gated send, the rollback and the walk of one trunk commit's tiers across machines. It still answers
every pull request's head with its check itself, so a check never depends on mesh-delivery.
- **mesh-delivery** is a Go bundle holding the mesh-scoped seat `mesh-delivery`. It owns the delivery, the
group, the state table and their state, and nothing else.
- **The forge's holder** sets a commit's statuses, keeps the delivery's view as one comment on the pull
request, and appends to the commit's note under `refs/notes/mesh-plan`.
- **The build seat** builds and checks, asked by the controller as before.
## A delivery
Its id is `<owner>/<repository>@<twelve characters of its commit>`. It holds:
- the pull request (number, base, head branch, description), or none for a commit that reached the trunk
without one;
- its **delivery plan**: build plan, deploy plan and verdict, as the controller's planner computed them for
its diffset (the change plan of ADR 0238 under the glossary's name);
- its **state**, its **transitions** (the last hundred, each with when, the event, from, to and why), and its
**machine steps** while delivering;
- once merged, **the commit it landed on the trunk as**, and the walk that delivers it.
## The state table
One table, compiled into mesh-delivery, read by its code, its test, its `table` verb, the controller's
self-check and healer H2. Each state has a **bound**: how long it may be held before `stalled` lists it.
Each state also says **what H2 may do** after the bound, which is only ever a transition the table already
allows.
| State | Means | Bound | H2 after the bound |
|---|---|---|---|
| proposed | a head was announced and is being checked | 1 hour | nothing: the check's own watchdog (S6) speaks |
| checked | a verdict arrived | 1 minute | nothing: it is decided at once |
| ready | it passed; it waits to be merged | none | — |
| rejected | it failed its check | none | — |
| published | merged on the trunk its modules follow, its walk opened and asking nothing; it waits for its turn or its word | 30 minutes | `superseded` when a newer delivery took over its walk; `go` when its walk started |
| held | it waits for a person | 1 day | nothing: it is the operator's |
| delivering | its walk runs; its builds are asked and registered tier by tier | 2 hours | `delivered`, `failed` or `superseded`, as the walk's record says |
| delivered, failed, superseded, stopped | final | — | — |
The transitions are those of ADR 0239 decision 2. A delivery that is merged while not `ready` goes to
`held`. Its check did not pass, and only a person decides that it goes on. A delivery merged with no walk
opened for it within ten minutes is held too, saying so: nothing it moves follows that branch, or the merge
was not heard. Nothing waits silently. **Nothing of a published delivery is registered before its walk
starts**, because a send carries every registered build that waits on its machine (ADR 0236 §4a), and a
build registered before its turn would be carried by somebody else's send.
**The machine steps.** While a delivery is `delivering`, each machine its walk reaches has a row with the
module and the build: `sent` (the first machine's send, or the rest's), `judging` (the gate's healthy
readings so far), `passed`, `failed`, `rolled-back`. They are read from the walk's own record in the
controller's `plan-moved` event, never inferred from time.
## A delivery group
- **Membership**: the pull request's head branch name, across the mesh's repositories, among open pull
requests. Two or more make a group. A group that has started delivering takes no new member.
- **Order**: `after: <repository>` lines in a member's description, and the controller's `delivery-order`
over the graph (ADR 0239 decision 4). The plan shows the order and why each pair is ordered: *declared*,
*built by*, *version skew*, *engine before controller*, *by name*.
- **Check**: the members' heads composed together, one gate over every machine with all their definitions. It
is asked by mesh-delivery through `delivery-check` when the group forms or a member's head moves, and
posted on every member's head as `mesh/delivery-group`.
- **Delivering**: in order. A member waits in `published` until the one before is `delivered`.
- **A failed member**: the members after it are `stopped`, naming it. The members before it stay delivered.
Its own builds were put back at its gate.
A group's state is derived from its members and its composed check, and is shown, never stored apart from
them.
## What every transition does
In this order, each by its owner:
1. **kept**: mesh-delivery puts the delivery in its state (`deliveries`, one key per delivery; `groups`, one
key per group). Nothing is said before it is kept. What it owes outside is kept with it and retried until
done;
2. **said**: the event `mesh-delivery.transition` (and `mesh-delivery.group` when a group's derived state
changes), with the delivery's id, from, to, event, why and plan summary, and no secret or address;
3. **noted**: mesh-delivery asks the forge's holder (`gitea_note_append`) to append a line to the commit's
note: the head while the delivery is off the trunk, the commit it landed as once it is on it, and at its
end a line of what was executed there. The forge's holder writes it in its own repository as the forge's
own account, and a line already there is not appended again;
4. **reflected**: the delivery's view, one comment on the pull request with the state, the plan, the machine
steps and the group, kept current (`gitea_delivery_view`); the status `mesh/delivery` on the head, and
`mesh/delivery-group` on every member's head (`gitea_commit_status`). Every status links to the pull
request, `mesh/merge-gate` among them.
## The verbs
**mesh-delivery's seat** (served by its holder; the first five read):
| Verb | Answers or does |
|---|---|
| `deliveries` | every delivery not final, and the final ones of the last day, one line each; filtered by state, repository or group |
| `show` | one delivery or group whole: plan, transitions, machine steps, order, the walk |
| `groups` | every group with its members in order and its derived state |
| `what-if` | the delivery plan a diffset would have, asked of the controller's planner, kept nowhere |
| `table` | the state table and the machine-step table, with bounds and H2's repairs |
| `stalled` | every delivery held past its bound, with the transition H2 may take |
| `recheck` | a rejected or ready delivery back to proposed, with why |
| `release` | a held delivery to delivering, with why: a person's word |
| `stop` | any delivery not final to stopped, with why; its walk is ended through the controller |
| `close` | H2's only verb: the transition the table names for a stalled delivery, refused otherwise |
**The controller's verbs for it** (on the mesh-controller seat): `delivery-plan`, `delivery-order`,
`delivery-check` (a group's heads composed, or one head checked again), `deliver`, `delivery-stop`,
`delivery-walks`. A person has `plans go <plan> --why`, which does what `deliver` does under their name and
is recorded as a hand act. The controller says every walk it keeps as `plan-moved`, from a verb's command
too.
**The forge's holder's tools for it**: `gitea_note_append`, `gitea_delivery_view`, `gitea_commit_status`. It
also announces `pull.closed`, and a merge's head and the statuses the forge holds on it.
## The store
Key-value state on the bus, declared by the module (ADR 0201): `deliveries` and `groups`. One writer, the
seat's holder, whose code changes state from a single goroutine. A final delivery is kept for thirty days and
then removed. Its note on the commit keeps its record after that. The bus's own snapshot (ADR 0235) backs it
up.
## The bootstrap
- A walk that moves the controller, the node-engine, the node tools, the bus or mesh-delivery is the
controller's own, started by the merge and witnessed on the machine. mesh-delivery records it.
- Any other walk waits for `deliver` only while the seat has a holder on record. With none on record, the
controller starts it as it always did.
- A waiting walk is said by the controller whatever mesh-delivery says of itself: S16 raises
`plan.<walk>.waiting` after 30 minutes, urgent after 4 hours, naming `plans go <walk> --why` as the way on.
A holder that is down is also D3's `holder-silent`; one that is up and never says go is said by S16 alone.
## The switch
1. The controller's change (the seat, its verbs, `plan-moved`, the wait for a holder's word, `plans go`) is
merged and rolls out as the core does. With no holder on record it behaves exactly as before.
2. The catalogue's change (mesh-delivery, the forge holder's notes and view) is merged. mesh-delivery is built
and published.
3. A person assigns mesh-delivery to the control node. The seat now has a holder on record. Walks that start
from then on wait for its word. Walks already open finish on their own and are adopted as deliveries in
`delivering`.
4. To step back, unassign it. Waiting walks are then started by the controller at its next pass.
## Phases
| Phase | Delivers | Done when |
|---|---|---|
| A — the owner | mesh-delivery with its table, state, verbs, events and adoption; the controller's seat, verbs, `plan-moved` and the wait; the forge holder's view, notes and statuses; the group's order and composed check | the tests of ADR 0239's *How it is checked* pass; a merge on the live mesh with mesh-delivery held is delivered by it and shown by `deliveries` |
| B — the conditions | S16, the controller's own `plan.<walk>.waiting`; probe D14 reads `stalled` and raises `delivery.<id>.stalled`; H2 calling `close` — **built 2026-10-07**, on a branch | a delivery stalled on purpose is raised and closed by H2 through `close`; a walk never let go is raised by S16 |
| C — the walk itself | the per-tier walk moved into mesh-delivery behind single-send verbs, the controller's planner kept only for the built-in path | decided by its own record once Phase A has delivered for a while |
## What is not decided here
- A web view of deliveries beyond the pull request's comment.
- Whether `mesh/delivery-group` is required on the trunks. That is the operator's setting, through the forge
module's protection tool.
- Moving the per-tier walk out of the controller (Phase C).
@@ -1,278 +0,0 @@
---
layer: to-be
status: in-progress
code: [mesh-host, mesh-controller, mesh-lab]
updated: 2026-10-07
decisions:
- 02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md
- 02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md
- 02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md
---
# 48 — A module says how it is healthy
**Every long-running thing a module runs is judged alive by the node-engine on its machine, with no
declaration. Where the module declares how, it is also judged ready. The node-engine runs every check and
owns every verdict, states it in its report and on the bus, and the controller raises a condition on the
second look — which the release gate, the self-check, the healers and the operator's conversation already
act on.** ([ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md).)
This closes the gap to-be 45 names in its Phase 4: *container state in the node-engine's report*. The
core's own health definitions ([to-be 45](45-a-core-that-cannot-fail-silently.md) §8) stand beside it.
## The parts
```
machine control node
┌───────────────────────────────────────────┐ ┌───────────────────────────────────┐
│ node-engine │ │ controller │
│ ├─ liveness: every long-running resource │ report │ ├─ last state per machine │
│ │ running? restarts counted (kept) │──────────► │ ├─ second statement → condition │
│ ├─ readiness: the `health` field │ health │ │ module.<m>.<machine>.unhealthy │
│ │ http / tcp / unit ── itself │ events │ ├─ provider hold (waiting on) │
│ │ exec / runtime ── runtime's check │──────────► │ ├─ the gate (ADR 0236 §2) │
│ │ tool ── node tools │ │ └─ doctor, status, conditions │
│ └─ state per resource, since, streak │ └───────────────┬───────────────────┘
└───────────────────────────────────────────┘ │ condition events
operator's conversation (to-be 46), healers
```
One judge per machine, for every hosting form. Nothing else on the machine judges a module, and nothing
else sets a container's check.
## 1. Liveness, for everything that stays up
A **long-running resource** is a container that stays up, a process the mesh runs that stays up, or a
service unit the module states `running`. A container that runs once, a step, and anything on a schedule
are not long-running; their success is their step's and their schedule's.
On every tick the node-engine reads each long-running resource's state from the runtime or the service
manager. A resource is **alive** when it is running and has not restarted more than once within the
**settle window** — ten minutes, the gate's bound — after its grace period. A resource that is not running
when it should be, or restarts twice in the settle window after grace, is **unhealthy**, with the reason
(`down`, `restarting`) and the restarts counted.
**The node-engine counts restarts itself and keeps the count** in its own state on the machine, keyed by
module and resource, across recreates of the container and across its own restarts. The runtime's restart
count is lost on every recreate and its event history is about a minute long on a busy machine; neither is
read for a verdict.
**The grace** is the resource's declared grace, or 60 s with none declared. A restart inside it is not
counted: churn that stops ([issue 058](../../04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md))
is not a crash loop.
**A resource held still by an open maintenance step** ([ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md))
is `held`: neither alive nor dead, and judged again from a fresh grace when the step ends.
## 2. Readiness, declared: the `health` field
A long-running resource carries a field named `health`, beside its other fields. It says one **kind**, and
the timing:
| Part | Says | Bounds |
|---|---|---|
| kind | how readiness is looked at — see below | one per resource |
| interval | how often to look | default 30 s; not under 10 s |
| timeout | how long one look may take | default 5 s; under the interval |
| failing looks | how many failing looks in a row make it unhealthy | default 3; not under 2 |
| grace | after each start, how long failure does not count | default 60 s |
| needs | the provision whose provider the check exercises, for §5 | none |
The grace plus the failing looks at the interval are at most five minutes, so a resource broken from its
start is said within the gate's ten minutes with room for its judgings.
**The kinds:**
- **runtime** — the image's own check, adopted by name. A module that ships one *says* it does; an image
check a module does not adopt is not read, and is replaced by the module's own kind or none.
- **http** — a request to an endpoint the module declares under `listens`, a path, and the status
expected.
- **tcp** — a connect to an endpoint the module declares under `listens`.
- **exec** — a command run inside the container.
- **unit** — the unit's own readiness: active and not failed, and for a unit that notifies, notified.
- **tool** — one of the module's own tools, answering healthy or not with why. Only for function no endpoint
shows (the identity provider's administrator logs in; the database provider can create in a consumer's
database; the broker), and only beside a check of another kind on the same module: a module never judges
itself alone (ADR 0227 rule 8).
**An endpoint is named, never a port or an address.** The check follows the machine's port for that endpoint
as the endpoint does; a port change moves the check with it.
## 3. Who runs each kind
| Kind | Run by | Read by the node-engine as |
|---|---|---|
| http, tcp | the node-engine, from the machine, to the endpoint's current port | the answer, within the timeout |
| unit | the node-engine, from the service manager | the unit's state |
| exec | the runtime: the node-engine sets the declared command as the container's check, with the declared timing | the container's health state |
| runtime | the runtime: the image's own command, with the declared timing | the container's health state |
| tool | the node tools, asked by the node-engine | the tool's answer, within the timeout |
An HTTP or TCP check from the machine, not inside the container, tests the path a caller takes
([issue 145](../../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)),
and costs no execution inside the container. A command is handed to the runtime so its own retries and start
period do the timing, without an execution per look from outside.
## 4. The state, its statement, and the condition
**Per module and long-running resource, the node-engine keeps a state:** `healthy`, `unhealthy` (with the
reason: down, restarting, or the check's failure in words), `starting` (inside grace), `held` (§1) or
`unknown` (nothing could be read). With it: since when, the failing streak, and the restarts counted. Every
start of a resource — a new build, a recreate, a restart — begins in `starting`.
**Said three ways**, as the core already says its own state:
- **in every report**, the current state of every resource;
- **as an event** on each change of state, on the node-engine's own subject;
- **again every minute** while a resource is not healthy, so a lost event is not a lost fault.
**The controller keeps the last state per machine** and raises **`module.<module>.<machine>.unhealthy`**
when two consecutive statements about the module say a resource is unhealthy. One statement is listed as
unconfirmed, as the self-check does a finding one look can be wrong about (to-be 45 §4). The first statement
that says no resource is unhealthy clears it. `module` joins the scopes of the condition store
(to-be 45 §2).
- **Severity:** `warning`; `urgent` when consumers wait on it (§5), or when it has stood four hours.
- **Summary** in words, naming the module, the machine and the resource; the reason, the streak and the
restarts are evidence. Held to the content rule (ADR 0234 §6), as every condition.
- **Resolver:** `self` — it clears on observation. A healer that acts on it is a later record.
**`status`** lists it with the other conditions; **`node show`** shows each module's resources with their
state and since when, so "is it working" has an answer without opening a terminal on the machine.
## 5. The gate, unchanged in rule
ADR 0236 §2 judges a module on its first machine three times, forty seconds apart, from two minutes after the
send and within ten minutes of it. Two of its points now read this design:
- **A module's own health holds** when its tools are served (as before) **and every long-running resource of
the module on that machine is stated healthy.** `starting` is not yet a pass: a resource still in grace
makes the judging wait, not fail.
- **No condition raised since the send about the module** holds it on `module.<module>.<machine>.unhealthy`.
Because every start begins in `starting`, a condition open before the send clears at the new build's start,
and one raised after it is the new build's.
At the bound, a module not healthy is put back, as ADR 0236 §3 says.
## 6. A provider down is said once
A check that names, in `needs`, the provision it exercises is tied to that provision's provider for this
consumer — the provider the controller composed the consumer's grant against
([issue 181](../../04-ISSUES/181-an-assignment-does-not-record-which-provider-answers-it/00-report.md) is
about recording it beside the assignment; until then the grant is the record). While that provider's own
condition is open:
```
provider P unhealthy ─► module.P.<machine>.unhealthy (urgent: consumers wait on it)
evidence: waiting on it — consumer A, consumer B, consumer C
consumer A failing ─► no condition of its own; its gate waits, not fails
P healthy again ─► held findings released: a consumer still failing is now its own
```
What a consumer finds while its provider is healthy, or by a check that names no provision, is its own.
A machine-level condition is never pinned on a module
([issue 281](../../04-ISSUES/281-a-tier-sent-one-module-at-a-time-blamed-a-module-for-its-machine/00-report.md)).
ADR 0224's `provider-failing` standing is a different fact — the provider failing *a consumer's
provisioning* — and stays as it is.
## 7. Nothing restarts on health
The runtime restarts what exits, as now. The node-engine never restarts, recreates or stops a resource for
being unhealthy. The condition reaches the operator through the conversation (to-be 46), or a healer once a
record gives one that act.
## 8. A declaration is proved before it is trusted
- **`module check`** refuses a `health` field naming an endpoint the module does not declare under
`listens`, a port or an address, an interval, timeout, count or grace outside its bounds, a tool the module
does not serve, or a `tool` kind with no other kind beside it on the module.
- **The bed.** mesh-lab provides a throwaway container runtime that the catalogue's check, run on the build
seat ([ADR 0237](../../02-DECISIONS/0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md)),
uses to start every resource whose declaration or image changed, alone, and read its check. It must say
healthy within its grace. A check that names a provision in `needs` gets no provider on the bed, so there it
must only reach the program — an answer of any status, a connect — and is judged fully on the first
machine's gate. An adopted image check is proved the same way.
- **What the bed is not.** It proves the check's wiring — that it can see the program working — which is a
property of the image and the declaration, not of the mesh's state. Whether the change is good is judged on
the live mesh, by the first machine's gate ([ADR 0149](../../02-DECISIONS/0149-the-live-mesh-is-the-test-bed.md)).
## 9. The migration
**Today:** 68 modules run something long-lived (49 with a container, 19 with a running service only); 57 run
nothing long-lived and need no declaration.
1. **Liveness first, with no declarations** (Phase A). From then every long-running resource is judged. The
catalogue's count of long-running resources without `health` is written down and starts.
2. **The seven modules whose images ship checks** adopt them by name; for the two that were wrong in the
mesh's configuration, the corrected check is what the bed proves. 19 containers are covered at once.
3. **The other container modules, and the containers without an image check in three of the seven,** declare
HTTP on their declared endpoint where they serve HTTP, TCP where they serve something else, a command where
neither shows readiness. 48 of 49 already declare the endpoint the check needs.
4. **The 19 service-only modules** declare the unit's own readiness, or TCP where the unit listens.
5. **Modules whose function no endpoint shows** add a tool check — the identity provider, the database
provider, the broker — as they are met, not in advance.
6. **The date:** when the count reaches zero, or six weeks after Phase A is live, whichever is first. From then
`module check` refuses a long-running resource without `health`.
**Until then, and for ever for a module running nothing long-lived:** liveness of everything it runs that
stays up, and the gate's points as they stand.
## Phases
| Phase | Repository | Delivers | Done when |
|---|---|---|---|
| A — liveness and the statement | mesh-host (node-engine) | liveness for every long-running resource; restarts counted and kept; the state per resource in the report; the change event and its repetition; nothing restarted on health | the node-engine tests of ADR 0240 rules 1 and 6 pass; every machine's report carries a state for every long-running resource |
| | mesh-controller | the last state per machine; `module.<module>.<machine>.unhealthy` on the second statement, cleared on the first that does not say it; the gate reading the stated health; `node show` and `status` | the controller tests of rule 4 pass; mesh-lab's replay of the crash loop fails its gate within the bound; a module stopped on purpose on the live mesh is raised and cleared |
| B — the field | mesh-controller | `health` parsed on every long-running resource; `module check`'s refusals; the catalogue-wide parse | a test per refusal; the whole catalogue passes `module check` |
| | mesh-host (node-engine) | the scheduler and the kinds: http, tcp, unit itself; exec and runtime as the container's check; tool through the node tools | the rule 3 tests pass; the replay of issue 145 raises the consumer within two looks |
| C — the provider hold | mesh-controller | `needs` read against the provider composed for the consumer; the consumer's finding held under the provider's condition; the consumer's gate waiting | the rule 5 test (one provider, three consumers, one condition) passes |
| D — the proof | mesh-lab, mesh-catalog | the bed; the catalogue's check starting every changed resource on it; adopted image checks proved | the replay of the studio's false *unhealthy* fails the bed, not a machine |
| E — the migration | mesh-catalog, mesh-controller | the declarations of §9 steps 2–5; the count kept in the catalogue and its test; `module check` refusing after the date | the count is zero, or the date has passed and `module check` refuses |
Phases A and B may be built together; A is live first, because it judges without a single declaration.
## As built — Phase A
Built on one feature branch in each of mesh-host, mesh-controller and mesh-lab; not yet merged or rolled out.
What the build chose where this design left it open:
- **The look.** The node-engine looks every 15 seconds: one inspect of every container it runs for a module,
one show of every unit per service manager, and nothing else — no execution inside a container, no event
history. A runtime that does not answer is `unknown` for every container, never `down`. What it counts is
kept in a file beside the node's state.
- **What a restart is.** The runtime's own count is read only to see it move. A count that moved is the
runtime restarting what exited, counted after the grace; a container made again (its identity changed with
no count moved) or started again by somebody (its start time moved with no count moved) is a new start,
from a fresh grace, its counted restarts kept. A container a maintenance window holds is `held`, and the
window ending is a start. A resource not running after its grace is `unhealthy` — `restarting` while the
runtime or the manager is restarting it, `down` otherwise — before any restart is counted.
- **What is judged.** A module's containers that stay up, its services stated running and its processes
that stay up. What the mesh declares in its own right (no module) and what an adopted machine holds as it
was found are not, in this phase.
- **The statement.** In every report the engine makes, judged right after the apply, so a resource the
apply started says `starting`; as an event on the machine's own subject (inside the grant every host
already has — no grant changed, and the writers table's row for the machine's report names the subject
beside the report's) on each change, again every minute while one is not healthy, and every five minutes
anyway, so a controller started again knows a healthy machine's state without waiting for its next apply.
Ordered by when the engine looked; an older statement is refused.
- **The controller** keeps the newest statement per machine in its store, with each module's run of
unhealthy statements, so the gate, `node show` and a restarted controller read the same word. The
condition joins the scopes as `module`. The gate reads a statement heard since the send; a resource
starting, unhealthy, held or unknown makes the judging not a pass — the build fails at the bound, as
ADR 0236 §3 says, rather than at once. A machine whose engine states nothing is judged as before.
- **The proof.** mesh-lab's replay `R-crashloop` raises a container whose program exits at start, has the
node-engine at its commit judge it through the runtime, and the controller at its commit judge the gate
from what the engine said. Proved: it fails on the trunk before this build, and passes on it.
Still to do for Phase A's "done when": every machine's report carrying a state for every long-running
resource, and a module stopped on purpose raised and cleared, are read on the live mesh once both are rolled
out — the node-engine first, then the controller.
## What is not decided here
- Whether a healer restarts what stays unhealthy (its own record, under ADR 0231).
- Health of scheduled work, and of what a module's events do (issue 276).
- A slow starter that needs more than five minutes to be ready — a change to the bound, recorded, when one
exists.
- The core components declaring their own health through the same field.
-3
View File
@@ -45,9 +45,6 @@ document is written and this one's status becomes `implemented`.
| [`38-building-the-operators-machine.md`](38-building-the-operators-machine.md) | **In progress.** The work of design 37 as packages: the runtime serves many modules, the controller composes one per node, the console becomes its serving mode, the packet filter moves first, then the shell and the service manager — tested on the live mesh by the operator's decision | [ADR 0175](../../02-DECISIONS/0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md), [0160](../../02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md), [0149](../../02-DECISIONS/0149-the-live-mesh-is-the-test-bed.md) | | [`38-building-the-operators-machine.md`](38-building-the-operators-machine.md) | **In progress.** The work of design 37 as packages: the runtime serves many modules, the controller composes one per node, the console becomes its serving mode, the packet filter moves first, then the shell and the service manager — tested on the live mesh by the operator's decision | [ADR 0175](../../02-DECISIONS/0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md), [0160](../../02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md), [0149](../../02-DECISIONS/0149-the-live-mesh-is-the-test-bed.md) |
| [`41-the-shell-and-the-accounts-environment.md`](41-the-shell-and-the-accounts-environment.md) | **In progress.** The shell and the account's environment as modules: an environment module every module contributes variables and `PATH` entries to, shell code contributed to the login shell in named slots, the prompt and plugins as modules, the host giving a login back, and the service manager's module finished | [ADR 0203](../../02-DECISIONS/0203-the-accounts-environment-is-one-modules-and-every-module-contributes-to-it.md), [ADR 0204](../../02-DECISIONS/0204-a-module-contributes-shell-code-to-the-login-shell-in-named-slots.md), [ADR 0205](../../02-DECISIONS/0205-software-the-distribution-does-not-package-ships-as-a-pinned-archive-of-the-module.md) | | [`41-the-shell-and-the-accounts-environment.md`](41-the-shell-and-the-accounts-environment.md) | **In progress.** The shell and the account's environment as modules: an environment module every module contributes variables and `PATH` entries to, shell code contributed to the login shell in named slots, the prompt and plugins as modules, the host giving a login back, and the service manager's module finished | [ADR 0203](../../02-DECISIONS/0203-the-accounts-environment-is-one-modules-and-every-module-contributes-to-it.md), [ADR 0204](../../02-DECISIONS/0204-a-module-contributes-shell-code-to-the-login-shell-in-named-slots.md), [ADR 0205](../../02-DECISIONS/0205-software-the-distribution-does-not-package-ships-as-a-pinned-archive-of-the-module.md) |
| [`42-the-machines-modules-in-order.md`](42-the-machines-modules-in-order.md) | **In progress.** The order the machines' modules of research 026 and 027 are built and rolled out: every machine's first (sudo, localization, time sync, pacman, logrotate, avahi, systemd, docker, `~/.ssh`, scripts, kernel), then both workstations', then one machine model's; each proven on one workstation before the rest | [ADR 0173](../../02-DECISIONS/0173-the-operators-machine-is-the-meshs-and-a-module-is-what-it-declares.md), [ADR 0182](../../02-DECISIONS/0182-inside-a-home-the-mesh-owns-what-it-places-and-holds-the-rest-as-found.md), [ADR 0205](../../02-DECISIONS/0205-software-the-distribution-does-not-package-ships-as-a-pinned-archive-of-the-module.md) | | [`42-the-machines-modules-in-order.md`](42-the-machines-modules-in-order.md) | **In progress.** The order the machines' modules of research 026 and 027 are built and rolled out: every machine's first (sudo, localization, time sync, pacman, logrotate, avahi, systemd, docker, `~/.ssh`, scripts, kernel), then both workstations', then one machine model's; each proven on one workstation before the rest | [ADR 0173](../../02-DECISIONS/0173-the-operators-machine-is-the-meshs-and-a-module-is-what-it-declares.md), [ADR 0182](../../02-DECISIONS/0182-inside-a-home-the-mesh-owns-what-it-places-and-holds-the-rest-as-found.md), [ADR 0205](../../02-DECISIONS/0205-software-the-distribution-does-not-package-ships-as-a-pinned-archive-of-the-module.md) |
| [`45-a-core-that-cannot-fail-silently.md`](45-a-core-that-cannot-fail-silently.md) | **Designed.** The core says when it is wrong, refuses what is stale or unreadable, heals what it knows, upgrades one machine at a time with a witness that rolls it back, and is checked against the real mesh before merge: the writers and signals tables, the condition store, `doctor`, the minimal output channel, healers, the lease and report order, staged upgrades, the facts snapshot and replays, in six phases | [ADR 0227](../../02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md), [ADR 0224](../../02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md), [ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md) |
| [`46-the-conversation-with-the-operator.md`](46-the-conversation-with-the-operator.md) | **Designed.** The mesh tells and asks its operator over channels that are holders of two kinded benches, `channel` and `intake`, declaring capabilities from a fixed vocabulary; the router orders them by work context and never lowers the bar; an answer that performs an action is checked and performed by the controller, on a TOTP code or a verified Telegram sender, never on a desk click alone; operator messages addressed or answered by the mesh's responder, untrusted input kept as data, references for what words may not carry, and a recoverable factor; Telegram first, in six phases | [ADR 0234](../../02-DECISIONS/0234-the-mesh-holds-a-conversation-with-its-operator.md), [ADR 0227](../../02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md) |
| [`47-delivery-from-commit-to-delivered.md`](47-delivery-from-commit-to-delivered.md) | **Designed.** A delivery is one commit in one repository, from its pull request's head to every machine; a delivery group is deliveries sharing a branch name, ordered and checked as one future state; both owned by the `mesh-delivery` module with one state table, its state on the bus, every transition said, noted on the commit and shown on the pull request; the controller keeps the planner, the gate, sending and the walk, and the core's own updates never wait for the module | [ADR 0239](../../02-DECISIONS/0239-a-delivery-is-owned-by-the-mesh-delivery-module-and-runs-from-commit-to-delivered.md), [ADR 0238](../../02-DECISIONS/0238-a-commit-is-the-build-at-hand-one-commit-one-change-plan-checked-off-the-trunk-and-published-only-on-it.md), [ADR 0236](../../02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md) |
## Not yet written ## Not yet written
@@ -1,9 +1,9 @@
--- ---
status: located status: resolved
opened: 2026-10-01 opened: 2026-10-01
located-in: [the identity provider's assignment on the control node (an adopted database whose admin predates the mesh, and the same database moved on 2026-10-05), mesh-catalog modules/keycloak (the minted `admin` own-secret, applied by the server only when it creates its master realm), every provider's provisioner loop (a consumer failed for a day said so only in a journal), mesh-controller status (nothing read what a provider could not do)] located-in: [the identity provider's assignment on the control node (an adopted database whose admin predates the mesh), mesh-catalog modules/keycloak (the minted `admin` own-secret, applied by the server only when it creates its master realm)]
fixed-by: twice by hand through the server's own bootstrap command (2026-10-01, 2026-10-05); the safety nets in mesh-controller PR #70, mesh-host PR #28 and mesh-catalog PR #80 (ADR 0224), resolved when they are merged and rolled out fixed-by: done by hand on 2026-10-01 through the server's own bootstrap command — the admin's password set to the value the mesh minted; no code changed
amended-design: 03-DESIGN/01-to-be/19-the-module-protocol.md amended-design:
--- ---
# 179 — An adopted identity provider's admin never took the secret the mesh minted # 179 — An adopted identity provider's admin never took the secret the mesh minted
@@ -44,60 +44,6 @@ only the download client's: a credential the mesh cannot make is the operator's
*How it was checked:* `keycloak_list_realms` through the console answers; `keycloak_list_clients` *How it was checked:* `keycloak_list_realms` through the console answers; `keycloak_list_clients`
on the realm lists both mesh-named clients; the server's log stops the five-second login error. on the realm lists both mesh-named clients; the server's log stops the five-second login error.
## Recurred, 2026-10-05
**The identity provider's database was moved that morning, and the admin's old password came back
with it.** The record the 2026-10-01 repair wrote lived in the database; the database that came up
on the new store was taken from before that repair, so the realm's `admin` again held a password that
predates the mesh, and the server — which applies the minted variable only when it creates its master
realm — did not touch it. From shortly after midnight the provisioner failed every consumer every five
seconds with *401 invalid_grant, Invalid user credentials*: about 31,000 refused logins until it was
repaired by hand, the same way as the first time, at about 23:55. In those twenty-three hours nothing
anywhere said so except the provider's own journal. The machines applied what they were sent, the
module was current, and `status` printed its all-well sentence.
**The repair, as done both times**, inside the server's container and with nothing printed: the
server's `bootstrap-admin user` command creates a temporary administrator, from a password generated
in the container and handed over in an environment variable, on a management port other than the
default (the running server holds that one); the admin client logs in as it and sets `admin`'s password
to the value the mesh minted; the temporary administrator and every temporary file are removed.
## What makes it not happen silently again (ADR 0224)
Twice by hand is a pattern, and the operator's direction was that it never happen again: detected
automatically, repaired automatically where safe, loud where not, and tested.
[ADR 0224](../../02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md)
records the rule; three safety nets implement it.
1. **The identity provider repairs its own admin.** Its code — ported from TypeScript to Go with this
change — checks that `admin` logs in with the mesh's secret once the server answers, every five
minutes after, and at once whenever the provisioner is refused. A refusal is repaired exactly as
above, by the module, inside the server's container through the container runtime it already
declares, with the secrets passed on standard input and never on a command line; then the login is
checked again and one line and an `admin.repaired` event say what was done and why. A repair that
fails is announced as `admin.unrepaired`, said loudly in the journal with a pointer here, and braked —
ten minutes, doubling to six hours — and while the admin is refused the provisioner stops asking the
server, so a lockout policy is never provoked. `keycloak_admin_check` reports the admin's state and,
with `repair`, repairs it now.
2. **A provider that keeps failing a consumer says so.** The provisioner loop announces a consumer it
has failed for five minutes without a success — create, the periodic check or an unreadable secret —
as `provisioner.failing`, with the class of error, and again every fifteen minutes; its recovery as
`provisioner.recovered`.
3. **`status` names it.** The controller keeps each provider's newest failing word per consumer, and
`status`, its JSON and `node show` list it; it breaks the all-well sentence. On 2026-10-05 the first
line of `status` would have been the identity provider failing both consumers with
`credentials-rejected`, five minutes after midnight.
The first is specific to this module and is the repair; the second and third are for every provider,
and are what would have said it on the day even if the repair had not existed.
*How it is checked:* the module's tests drive the repair against a fake server and a fake container
runtime — repaired and verified on a refusal, never while the server is unreachable, braked after a
failed repair, no secret on any command line, the provisioner short-circuited while refused; the loop's
tests drive the announcements; the controller's tests take a failing word from the bus to `status` and
`node show` and back out on recovery. Live, after rollout: a deliberately wrong admin password on a
throwaway server is repaired within five minutes, and `status` stays all-well while it is.
## Open ## Open
The manifest still says the admin password is the mesh's to mint, which is true of a fresh install The manifest still says the admin password is the mesh's to mint, which is true of a fresh install
@@ -1,8 +1,8 @@
--- ---
status: resolved status: located
opened: 2026-10-01 opened: 2026-10-01
located-in: [mesh-catalog modules/dnsmasq, mesh-catalog modules/docker, mesh-controller internal/overlay/generator.go, mesh-controller internal/catalogue/resolve.go (checkResources), mesh-controller internal/catalogue/seat_into.go, mesh-host internal/apply/apply.go (orphan removal)] located-in: [mesh-catalog modules/dnsmasq, mesh-controller internal/overlay/generator.go, mesh-controller internal/catalogue/resolve.go (checkResources)]
fixed-by: novox/mesh-controller#62 (d84c969), novox/mesh-catalog#72 (5c89eaf), novox/mesh-host#26 (a192bcc), novox/mesh-controller#63 (c34b937) fixed-by:
amended-design: amended-design:
--- ---
@@ -71,32 +71,6 @@ gives that module declared settings with defaults. The fix, once both are accept
5. The collision check sees a computed module's resources as well, so a second writer cannot come 5. The collision check sees a computed module's resources as well, so a second writer cannot come
back through generated code. back through generated code.
> **Where it stands, 2026-10-05.** Step 1 is done: the resolver module writes nothing into the
> runtime's file, and the runtime's module (the holder of `node-container-runtime`) writes
> `live-restore` and reloads its own service. No module writes `dns`
> ([ADR 0196](../../02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md)).
> That answers the first two open questions below.
>
> Steps 2, 3 and 5 are decided in
> [ADR 0222](../../02-DECISIONS/0222-a-module-is-told-where-a-mesh-seats-holder-is-reached-and-the-controller-writes-no-file-a-seats-holder-owns.md),
> and the operator's rule is general: the controller never writes a file a seat's holder owns; it
> tells the owner. The registry reaches the runtime's module as a value the mesh holds, not as a
> setting: `${seat:mesh-artifact-store:reach}`, with no binding. The ADR 0082 and ADR 0102 notes
> step 2 asks for are written.
>
> Step 4's single push is replaced by an order. A list member written by two records is tolerated
> by the host, which a scalar key in two modules was not. Four pull requests, merged in this order:
>
> 1. The controller learns the placeholder (mesh-controller #62).
> 2. The runtime's module states `insecure-registries` with it (mesh-catalog #72).
> 3. The host leaves a unit alone when another declared service still holds it. It is deployed
> before 4. Without it, removing the private network's record of the runtime's service gives back
> what that record found (mesh-host #26).
> 4. The controller stops generating the private network's two resources, and checks generated
> resources for collisions (mesh-controller #63).
>
> `fixed-by:` is filled when they merge.
## Open questions ## Open questions
- **How the resolver's address reaches the runtime.** Either the resolver seat (`node-dns-resolver`) - **How the resolver's address reaches the runtime.** Either the resolver seat (`node-dns-resolver`)
@@ -109,10 +83,3 @@ gives that module declared settings with defaults. The fix, once both are accept
- **The adopted machine's predecessor values.** The runtime module adopting a file with a - **The adopted machine's predecessor values.** The runtime module adopting a file with a
hand-written `dns` and `live-restore: false` replaces both. That is intended, and is the one hand-written `dns` and `live-restore: false` replaces both. That is intended, and is the one
restart the operator must make on that machine. restart the operator must make on that machine.
## Verified
2026-10-05: on every machine `insecure-registries` is now the runtime's module's own, filled from the
store's seat; the private network's two generated resources are gone. The handover kept the list
member throughout, and the old reload record was forgotten, not given back ("docker.runtime still
holds the unit"), so the runtime was never stopped.
@@ -3,7 +3,7 @@ status: located
opened: 2026-10-03 opened: 2026-10-03
located-in: located-in:
- mesh-controller - mesh-controller
fixed-by: novox/mesh-controller#81 (b853439) fixed-by:
amended-design: amended-design:
--- ---
@@ -40,43 +40,3 @@ Owner mesh-controller: the raise runs once (`RaiseSeats` with the holders of the
`assign` run `EnsureConsumer` only for a module's declared consumption. **Fix direction:** when a `assign` run `EnsureConsumer` only for a module's declared consumption. **Fix direction:** when a
module claiming a seat with `accepts` is assigned, or on every push that composes a holder for such a module claiming a seat with `accepts` is assigned, or on every push that composes a holder for such a
seat, ensure the seat's worker as the raise does — the same derivation, the same idempotent assertion. seat, ensure the seat's worker as the raise does — the same derivation, the same idempotent assertion.
## Seen again, 2026-10-06: the module half of the same gap
A module carried by the runtime, which consumes the controller's condition events, was assigned to a
machine and pushed. It said `binding messenger's consumer …: nats: consumer not found`, and the
self-check's D6 confirmed that the consumer it reads through was not on the bus. The cause was the one
diagnosed above, on the other side. The controller asserted the bus's objects (streams, seat workers,
every module's consumer) only when it started. A push created a module's consumer only when it minted
that module a bus credential, and the runtime carries a module's traffic under its own credential, so
the module was never minted one and its consumer was never created.
## Fix
The bus's objects that a declaration implies are now asserted whenever a declaration is sent: on a
push, on the machines the push cascades to, and on a plan's or a rotation's sends. They are not
asserted only at start. The send asserts them through the same derivation the start and the
self-check use, before the memberships. There is no second list of what a send needs. Every part is
idempotent, so asserting the whole set again changes nothing that already holds.
A failure to assert is said in the send's own output and raised as a condition (`bus.objects.unasserted`).
The next send that asserts everything clears it. The send itself goes on. The objects belong to the
mesh, not to the machines being sent, and holding every machine back for one consumer would turn one
fault into all of them. The start still refuses to serve without its objects.
Checked by mesh-controller's `busobjects_test.go`. A module carried by the runtime and a seat's holder,
both assigned after the start's assertion, are asserted by the next send. Against a real bus, through
the send's own grant step, both exist after it. That test fails without the change, with the
`consumer not found` seen live. A refused consumer is said in the output and raised as a condition.
The consumers after it are still asked for, and the next good send clears the condition.
## Healing what no send reaches (to-be 45 Phase 3)
The fix covers every object a send implies at the moment it is sent. It does not cover an object lost
between sends: a consumer somebody deleted, or a bus whose data directory was replaced. D6 (a module's
consumer) and D3 (a holder without its worker) find these. Until the next send or a controller
restart, nothing makes them again. H3 in to-be 45 is the healer for D3's half ("raise the seat's
objects again"). Its repair is now exactly the assertion a send makes, so it can be the same call for
D6's missing consumer as well: one healer for "an object the mesh defines is not on the bus", braked
per object, rather than two. This is noted for Phase 3 and not built here.
@@ -1,8 +1,8 @@
--- ---
status: resolved status: open
opened: 2026-10-04 opened: 2026-10-04
located-in: [mesh-host internal/apply/apply.go, mesh-host internal/apply/schedule.go] located-in: [mesh-host internal/apply/apply.go, mesh-host internal/apply/schedule.go]
fixed-by: mesh-host PR #23 fixed-by:
amended-design: amended-design:
--- ---
@@ -62,9 +62,3 @@ is a hole in something new rather than something that broke.
at the cost of a second source for "is this container meant to be running". at the cost of a second source for "is this container meant to be running".
- Either way: should the *report* say a window is open, so a machine that looks half-stopped at - Either way: should the *report* say a window is open, so a machine that looks half-stopped at
03:31 reads as working rather than broken? 03:31 reads as working rather than broken?
## Resolved — 2026-10-05
An apply that arrives during a window now leaves the containers the window holds alone, and reports them `held-still`, naming the step. The first apply after the window converges them. The window is recorded under the host's state directory, the one place every applier on a machine shares: the daemon, a hand-run apply and the installer. It is released on every path, and lapses after six hours or when its process is gone. The machine's report lists open windows. Live on all four machines the same evening. Of the issue's two options, the narrower was taken: nothing blocks, so a push is never held for the length of a window.
Accepted and said in the change: an apply that inspected a container as running in the instant a window opens can still recreate it.
@@ -3,7 +3,6 @@ status: open
opened: 2026-10-04 opened: 2026-10-04
located-in: [] located-in: []
fixed-by: fixed-by:
replay: R236
amended-design: amended-design:
--- ---
@@ -1,8 +1,8 @@
--- ---
status: open status: resolved
opened: 2026-10-04 opened: 2026-10-04
located-in: [] located-in: [mesh-controller cmd/mesh-controller/main.go, mesh-controller cmd/mesh-builder, mesh-controller internal/link/build.go]
fixed-by: fixed-by: mesh-controller#47
amended-design: amended-design:
--- ---
@@ -35,3 +35,11 @@ acts on it, review becomes a formality: the unreviewed definition reaches a mach
1. Where does the dry run's outcome enter the record — the build machine's `built` event, consumed as 1. Where does the dry run's outcome enter the record — the build machine's `built` event, consumed as
any other build's? any other build's?
2. Did the controller roll the dry run out, or did a later push compose from it? 2. Did the controller roll the dry run out, or did a later push compose from it?
## Resolution (2026-10-05)
A dry run is marked on the request (`DryRun`), the builder echoes the mark on its outcome, and the
controller's daemon sets a marked outcome aside: no record, no registration, no plan, nothing a push
could send (mesh-controller#47, with a test that the daemon takes a dry run in with no store at all).
Left open as a follow-up: the catalogue module also hears build outcomes and records their edges; it
should skip a dry run too.
@@ -1,8 +1,8 @@
--- ---
status: open status: located
opened: 2026-10-05 opened: 2026-10-05
located-in: [] located-in: [mesh-controller internal/catalogue/seats.go, mesh-catalog modules/restic]
fixed-by: fixed-by: mesh-controller#49, mesh-catalog#49, mesh-catalog#54, mesh-catalog#55, mesh-catalog#56, mesh-media-catalog#1
amended-design: 03-DESIGN/01-to-be/43-backups-against-mistakes.md amended-design: 03-DESIGN/01-to-be/43-backups-against-mistakes.md
--- ---
@@ -47,3 +47,27 @@ research 030; proposed as ADR 0214.
Issue 241's recovery cost a night and lost the forge's records of twelve days. With a nightly backup Issue 241's recovery cost a night and lost the forge's records of twelve days. With a nightly backup
held on another machine, it would have been a ten-minute restore of yesterday. held on another machine, it would have been a ten-minute restore of yesterday.
## Where it stands (2026-10-05)
Backups run (ADR 0214, to-be 43): the control node keeps nightly restore points of its eight stores
and services, the home server of its databases and its media apps' libraries and cover art; each on
its own machine, on its larger filesystem. The first runs were tried one module, then two, then a
whole node, each proven by a restore beside the live data. Not yet built, which is why this stays
open: the weekly test restore into a throwaway instance, the 48-hour status line, and failures
reaching the operator rather than the holder's log.
What the rollout taught, for the next module that takes contributions:
- **A contribution makes the contributing module depend on the seat.** Merging the stores'
contributions before a holder was assigned left every node's plan unresolvable — the whole mesh,
not only the machines running a store — until the holder was assigned. Assign the holder in the
same step as the merge.
- **A brand-new module is not built by the push that adds it**; its first build is asked for by hand.
- **The holder must look as the account that can see.** Its first run called a store's dumps missing
because it checked as the runtime's account, which cannot see inside the store's own directory.
- **A kept single file is not a directory.** restic restores a snapshot's subfolder, not a file;
the first restore of a module keeping its settings file refused it.
- **An update kills a running backup** — the holder's restic is its child. A hand-off to the new
version, the run living on as its own unit under the machine's service manager, is the idea to
take forward.
@@ -1,8 +1,8 @@
--- ---
status: resolved status: located
opened: 2026-10-05 opened: 2026-10-05
located-in: [mesh-catalog modules/claude-code, mesh-catalog modules/claude-licence-manager] located-in: [mesh-catalog modules/claude-code, mesh-catalog modules/claude-licence-manager]
fixed-by: mesh-catalog PR #45 fixed-by:
amended-design: amended-design:
--- ---
@@ -31,9 +31,3 @@
**Ruled out.** The bus delivered every binding: each machine that took the fourth binding did so in the **Ruled out.** The bus delivered every binding: each machine that took the fourth binding did so in the
same second it was published. The licences themselves were sound. The remaining licence refreshed same second it was published. The licences themselves were sound. The remaining licence refreshed
on every attempt, and the second account's login was adopted from its first report. on every attempt, and the second account's login was adopted from its first report.
## Resolved — 2026-10-05
Live on all four machines and the manager the same morning. After the restart every machine reported
the generation it held, and the manager's bindings matched them. The manager logged no failure while
moving its sequence past them.
@@ -1,5 +1,5 @@
--- ---
status: located status: open
opened: 2026-10-05 opened: 2026-10-05
located-in: [mesh-tools, mesh-controller] located-in: [mesh-tools, mesh-controller]
fixed-by: fixed-by:
@@ -60,26 +60,6 @@ and it is wrong.
schema describes — a test that calls each with its schema's required set and asserts the answer schema describes — a test that calls each with its schema's required set and asserts the answer
is not "needs" an argument the schema does not have. is not "needs" an argument the schema does not have.
## Seen again, 2026-10-05 and 2026-10-06 — and once it did harm
The same picture for more of the mesh's own verbs, read through the console:
```
mesh_describe mesh-controller.push → {"properties": {}}
mesh_describe mesh-controller.plan → {"properties": {}}
mesh_describe mesh-controller.assign → module only, no node
mesh_describe mesh-controller.pin → provision, from, module — no node
mesh_describe mesh-controller.settings → module, values, clear — no node
```
`build` and `queue`, which take no machine, described in full. For `plan` this is the refusal above.
For **`push` it was worse than a refusal**: a push takes no required argument, so a call naming one
machine arrived empty, ran as a push of every machine behind, and answered "N node(s) told" — an
operator who believed they had pushed one machine had pushed the mesh. Nothing in the answer said the
machine had been lost.
The diagnosis is in [`01-diagnosis.md`](01-diagnosis.md).
## Worked around, for now ## Worked around, for now
Through `<node>/node-login-shell.execute`, running the controller's own command line on the Through `<node>/node-login-shell.execute`, running the controller's own command line on the
@@ -1,63 +0,0 @@
# Diagnosis
## 2026-10-06
**The schema is not empty where it is kept.** The controller's verb table declares `node` for every
verb that acts on a machine, and so does the seat's record: asked through the console,
`mesh-controller.tools` answers `push`, `plan`, `assign`, `pin` and `settings` each with `node` among
their properties. The controller's announcement on the bus carries the same schemas. So the
controller publishes them whole, and the loss is after it.
**The console takes `node` out of every schema and every call.** The console's address grammar
(ADR 0195) puts the machine in the address for a seat held on every machine (`<node>/<seat>.<verb>`)
and for a module on one machine (`<node>/<module>.<tool>`), and so describes those tools without a
`node` argument and deletes one from their calls. It did both for **every** address — including a
seat held once for the mesh, where there is no machine in the address and `node` is the verb's own
argument: the machine the verb acts on. That is exactly the set observed: each verb lost `node` and
nothing else, `push` and `plan` had nothing else and described as empty, and the verbs that take no
machine (`build`, `queue`) were untouched. The older flat listing of the same console had this right
(it kept a mesh seat's schema as declared); the address form did not carry the distinction over.
**Why the verb then refused, or did the wrong thing.** The controller's verbs read the arguments they
know and ignore the rest. Given nothing, `plan` and `node` said they needed `node`; `push`, whose
`node` is optional, read its absence as "every machine behind" and did that. A verb that ignores what
it does not read cannot tell "nothing was asked" from "something was lost on the way".
**Ruled out.**
- A stale seat record. The record is older than the controller's current source (some verbs lack
arguments added since), but it carries `node` wherever the source does.
- The controller's announcement. It sends the record's schema per verb, unchanged.
**Located in** `mesh-tools` (the console: the schema it describes and the arguments it sends) and
`mesh-controller` (verbs that pass over what they are given).
## What the fix does, and how it is checked
Pull requests: novox/mesh-tools#14 and novox/mesh-controller#71 — open, not merged.
- **The console** takes `node` out only where the address names the machine. There, a `node` naming
the same machine is redundant and dropped; one naming another machine is refused. A call to a seat's
verb carrying an argument the verb does not declare is refused, naming it — the seat's record is the
mesh's own description of the verb, so the console can judge it.
- **The controller** refuses, naming the argument: one the verb does not declare; one it declares but
did not use for the command it composed (two arguments where one wins, or half of a two-argument
shape); a switch that is neither "true" nor "false". A push that names no machine says, as the
first line of its answer, that it is a push of the whole mesh and which machines it is sending, and
its last line names the machines told. `plan` gains `files`, `push` gains `behind` (the whole mesh
said outright, refused beside `node`), and the listing verbs a `limit`.
- **Checked by tests that walk every verb the controller announces:** for every combination of a
verb's declared arguments, the verb either refuses or every argument given changes the command it
runs; every verb refuses an argument it does not declare, and `node` where it takes none; and every
flag of the command a verb runs — read from the command's own source — is an argument of the verb's
schema or is listed, with the reason, as set by the verb or withheld from it. A verb added later,
or a flag added to a command, fails these until its schema says how a caller reaches it.
What the tests do not check is the console's own describe-and-call path against a live bus; its unit
test covers the choice of when `node` is the address's and when it is the verb's.
## Related
Issue 259 is another way a named push reached every machine: there the machine arrived and the
controller's flush after it sent the others. Different cause, same reader's harm — the answer of a
push is the only place that says how far it went, which is why the whole-mesh line is in the answer.
@@ -0,0 +1,39 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-media-catalog modules/plex]
fixed-by: mesh-media-catalog#2
amended-design:
---
# 245 — A media server's previews were reached through a link its container never mounted
## What was observed
On the home server the media server's scrub previews — 423 GB, about 57 000 preview files, one per
video — had not grown in seven months: the newest was from the day its preview folder was moved off
the server's own disk onto the large storage pool. The server's log repeated, live, that it could not
create a directory under its preview folder.
## Why
The move left the preview folder as a **symbolic link** inside the server's configuration directory,
pointing at the pool. The container mounts the configuration directory and the media libraries, and
not the pool's path, so inside the container the link pointed at nothing: the server could neither
show the previews it had nor make new ones, and said so only in its own log. Nothing the mesh reports
showed it — the container ran, and answered.
It is the failure the mesh's rule against symbolic links exists for: a link resolves differently in
every place that reads it, and a container is such a place.
## Resolution
The previews are a directory of the module, mounted at the server's preview path; where that
directory lives on a machine is the machine's placement setting, here the pool path the files were
already in. The link was removed at the cutover, the container recreated, the preview folders visible
inside it again and the log's errors gone. The previews are not backed up, by the operator's choice.
## How it is checked
The catalogue check that every mount is declared passes over the module. On the machine: the preview
folder seen from inside the container lists the same folders as the pool path.
@@ -1,50 +0,0 @@
---
status: open
opened: 2026-10-05
located-in: []
fixed-by:
amended-design:
---
# 245 — `status` calls a module behind when only its repository moved
## What was observed
2026-10-05. After three catalogue merges that changed 11 modules, `status` listed **69 modules
behind their source**, each as `holds <older commit>, source has <newer commit>`, with the advice
"`build --behind` builds them; `push --behind` sends them on".
Between the two commits, `git diff --name-only` shows changes under 11 module directories only. For
58 of the 69 modules listed, for example a Bluetooth module, the container runtime's, the forge's,
the package manager's and the bus's, the diff of the module's own path is empty. Their sources did
not change; only the repository's commit did.
The merges' plans were right: they rebuilt the changed modules and those that depend on them, by
tier. Only the report was wrong. An agent following the report's own advice ran `build --behind`,
which rebuilt all 69. The rebuilt bus module was rolled out, and its container was replaced on the
control node. Every node's runtime lost the bus for about a minute.
## Why it matters beyond this instance
"Behind" is the word a person and an agent act on, and `status` attaches a command to it. A module
is held to a commit of its repository, so after any merge almost every module of that repository
reads behind. The list then says nothing about what needs building: it hides the few modules that
really are behind among the many that are not, and it invites a rebuild of everything, which is not
a harmless act (above).
## The operator's direction (2026-10-05)
The only truth is the outcome of the build plan. A plan already decides, from a change, which modules
it affects: those whose sources changed and those that depend on them, tier by tier. "Behind" means
a module that a plan has decided to rebuild and has not yet rebuilt or rolled out, and nothing else.
No second comparison beside the plan is made, whether of commits, of folders or of files, because a
second answer to the same question is how the two came to disagree.
## Open questions
1. Where does `status` read "behind" from today, and what replaces it: the open plans' remaining
tiers?
2. What does a module's recorded commit mean once a plan that leaves it untouched has run? Does it
move forward, or does the record stop carrying a commit that only says when it was last built?
3. Should `build --behind` and `push --behind` take their lists from the same place, so that they can
never act on a module no plan named?
@@ -0,0 +1,45 @@
---
status: open
opened: 2026-10-05
located-in: [mesh-controller cmd/mesh-controller (settings), mesh-controller internal/inventory (SetSettings)]
fixed-by:
amended-design:
---
# 246 — Setting a module's settings replaces the whole layer, and nothing shows it first
## What was observed
An operator's agent set one placement — where a media server's preview folder lives on one machine —
with `settings set <module> {"places": {…}} --node <machine>`. The command answered that the setting
was recorded. The machine's plan then mounted the server's configuration from an empty default
directory: the node's layer had held the placements of three other directories, eight media
accesses, a public exposure, four endpoints and the account's ids, and every one of them was gone.
Caught before any push, by reading the plan. The previous layer was read back from the controller
database's nightly dump — the backups that issue 242 asked for, a few hours old.
## Why
The layer is a statement of the whole, by design (the inventory's `SetSettings`: "replacing rather
than merging … removing a key is done by leaving it out"). That design is sound; what is missing
around it is everything that makes it safe to use:
- **There is no way to read a layer.** `settings` has `set` and `clear`, no `show`; the console verb
likewise. To change one key, a person must already know every other key in the layer.
- **There is no history.** The row is updated in place; the previous values exist nowhere but a
database backup.
- **The answer does not say what was dropped.** "places" was reported as set; the six keys removed
were not mentioned.
## What would fix it
1. A way to read a layer — `settings show <module> [--node]` and the same on the verb.
2. `set` answers with what changed: keys added, changed and **removed**, so dropping one is never
silent. A removal could even require saying so.
3. The previous value kept: a settings history row per change, so an undo needs no backup.
## Status
Open. Until it is fixed: read the layer (from the store, read-only) before setting it, and compare
the node's plan before and after.
@@ -1,65 +0,0 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-tools]
fixed-by: mesh-tools pull request 13 — discovery waits for every runtime that answered PING, names the ones it missed, and mesh_runtimes says who answered
amended-design:
---
# 246 — The console says a module runs nowhere when a runtime answers late
## What was observed
2026-10-05. On the laptop, the console answered `mesh_machine` for the laptop with no modules and no
seats, and a call to one of the laptop's modules with "nothing in the mesh is called slack". The
laptop's own runtime said at the same moment that it served 275 tools for 48 modules. A minute
earlier, a call to another laptop module had been answered with "it does not run on the laptop; it
runs on the workstation", and a retry of the same call worked. Nothing was logged anywhere.
An agent worked around it by calling the runtime's local MCP port directly, which gives the right
answer and bypasses everything the console stands for: one way in, one account, one record of what
was called. The operator asked for a tool instead.
## What was measured
A read-only probe on the laptop's own runtime credential timed the discovery answers over 25 rounds
against the live bus, whose round trip from the laptop was about 40 ms:
- The laptop's runtime answer was the largest on the mesh, at about 164 kB with 341 endpoints. That
is far below the bus's message limit, and it was never shortened.
- It arrived last in every round: a median of about 365 ms, and once 813 ms. The other runtimes
answered within 180 to 275 ms, and the controller within 50 ms.
- The console gathered discovery answers for a fixed 750 ms. Inside the console, the gather runs
beside two controller calls and about 570 kB of answers on the same link, and it is slower still
while a runtime re-serves after a restart.
## Root cause
Discovery decided who was there by who answered in a fixed window. A late answer was not a failure
to anyone, so nobody said it. The index simply lacked that runtime, and every answer built on the
index then stated as fact that the runtime's modules did not exist, or ran only elsewhere.
Ruled out by measurement or by reading the code: an answer too large for the bus, subscriptions lost
when the bus reconnects, the console not counting its own machine's answer, and the merge of two
answers dropping a machine.
## Resolution
- Discovery asks who is there (PING, a hundred bytes, answered at once) beside what each serves
(INFO). It waits at least the old window, and then up to five seconds for every instance that said
it is there, so a large answer is waited for and a quiet mesh costs nothing extra.
- A runtime that said it is there and did not say what it serves in time, or that answered recently
and not now, is named. While one is unheard, the console never says an address is missing or runs
elsewhere: it says which runtime was not heard, and where the controller's records place the module.
- A new console tool, `mesh_runtimes`, says for every runtime how long its answer took, how large it
was, how many modules and tools it announced, whether it was shortened, and when it was last heard,
and which runtimes or machines were not heard.
- An announcement still too large after its descriptions are cut to their first line now leaves the
descriptions out, and says so.
## How it is checked
The fix ships with tests against a real bus: a runtime that answers after the old window is found and
called (the same test fails with the fixed window), a runtime that answers PING and never INFO is
named, and a restarted runtime, which answers under a new instance, is not reported as missed. Live,
`mesh_runtimes` shows every machine's answer and its time.
@@ -1,52 +0,0 @@
---
status: open
opened: 2026-10-05
located-in: []
fixed-by:
amended-design:
---
# 247 — A module cannot put the operator's account in a group
## What was observed
2026-10-05. The module for a peripheral-lighting daemon was assigned to the laptop. Its package
installs the daemon and creates the daemon's group. The daemon then refuses to start: "User is not a
member of the openrazer group". The device files are the group's, so the daemon cannot reach the
devices.
The module's own check names the fix, which is to add the account to the group and log in again. No
module can declare that fix:
- The account is a `user` resource, and the login-shell module already declares it, to set its shell.
- A second module that declares the same account, only to add one group, is refused as a duplicate
name.
- There is no resource for one membership on its own. A whole-account declaration that lists groups
would also take from the account every group it does not list, including the operator's own.
So the step is done by hand, with `sudo`, outside the mesh, and nothing records why the account is in
the group.
## Why it matters beyond this instance
More modules need this than this one: input devices (`input`), serial ports (`uucp`), the container
runtime (`docker`), virtual machines (`libvirt`, `kvm`), and capture or scanner hardware. Each is a
fact a module knows and the operator's account needs. Today every one is a hand step that survives a
reinstall only by memory. A membership added by hand is also never taken away when the module that
needed it is unassigned.
## Open questions
1. Is a membership its own resource (account, group), held by the module that needs it and given
back on undeclare? Or is it a contribution to the account's holder, in the way ADR 0212 lets a
module contribute to a seat?
2. A membership takes effect at the next login. How does the module say so: a finding, or a
moment the power or session seat already knows?
3. What does undeclare do with a membership the account already had before any module declared it?
The host keeps what it found and gives it back, as it does with a whole file it wrote over.
## How it is checked
When fixed, assigning the lighting module to a machine whose account is not in the group puts the
account in the group, says that a new login is needed, and leaves the account's other groups as they
were. Unassigning it removes only a membership the module added.
@@ -1,58 +0,0 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-controller internal/broker]
fixed-by: mesh-controller PR #51
amended-design:
---
# 248 — The controller's event consumer replayed a week, and held every new merge behind it
## What was observed
A merge to the catalogue at 15:17 never reached the controller. No plan was made and no build was asked
for it. Every machine kept running the build from before it. The merge just before, to the record, was
logged twice.
The bus showed why. The controller's durable consumer on the event stream was set to deliver
**everything the stream holds**, not from where it last was:
| | |
|---|---|
| the stream | a week of events, from message 1 514 to message 356 004 |
| the consumer delivered to | message 1 517, later 4 898 |
| acknowledged to | 0, later 1 569 |
| still to deliver | 6 955 merges and build outcomes, a week old |
| allowed outstanding | one at a time ([issue 175](../175-an-announcement-behind-a-long-build-comes-back/00-report.md)) |
So the controller was working through a week of past merges and build outcomes, one at a time, slowly. Every
new one — a merge, a build asked by hand — waited behind them. A build asked by hand finished on its
machine and was never registered. While this went on, the controller's client dropped messages
("slow consumer") several times, because heartbeats and reports share the loop with these events. From 15:17
its event loop did nothing more: no line logged, the one delivered event never acknowledged.
How the consumer came to deliver everything is not certain. The bus and the stores were rebuilt the night
before ([issue 241](../241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md)).
A consumer that is missing is made again by the controller's own assertion, and that made it with the
server's default, which is everything. [Issue 207](../207-a-re-made-worker-replayed-every-ask-the-stream-kept/00-report.md)
closed this for a consumer re-made because its type changed. It did not close it for one that is simply
not there.
## Why it matters
**Delivery stops, and nothing says so.** Status showed every plan done and no machine behind. The merge
that was missed is not "behind", because no plan was ever made for it. A replay of past merges can also
act on them again. Issue 207 records nine modules re-registered from the past the same way.
**The way out was a hand on the bus.** No verb resets a consumer. On the operator's explicit word, the
consumer was re-made from now with a one-off program run as the controller, otherwise unchanged. The
controller's loop still held the old event afterwards, so the new consumer's first delivery went
unacknowledged; the loop needs a restart of the controller to let go of it.
## Noticed alongside, not this issue
- Each merge to the catalogue planned 99 to 100 modules in two tiers and rebuilt modules it did not touch.
The output was byte-identical, so nothing was redeployed, but it costs minutes of the build machine
per merge.
- Three plans for three merges ran over each other. Each sent the build agent to every machine and asked
for the same builds. Nothing supersedes a plan for an older commit.
@@ -1,54 +0,0 @@
# 248 — Diagnosis
## 2026-10-05
1. A merge was not in the controller's log. The forge's own log showed nothing about delivering it, so the
question moved to the bus.
2. The bus's backlog tool named the controller's event consumer: 6 955 pending, redeliveries, one
unacknowledged. Its configuration was read from the server's monitoring endpoint: deliver policy
*all*, one outstanding, 30 seconds to acknowledge, five deliveries.
3. Sampled three times over a minute it did not move, and over the following hour it crawled forward
through week-old events. The controller's log held only its client's warnings: dropped messages,
and one refused reply.
4. The controller's code makes a missing consumer with the configuration it asserts, which sets no
deliver policy, so the server's default applies: everything. Issue 207's fix sets *from now* only on
the path where an existing consumer's type changes.
**Unblocked**, on the operator's explicit word: the consumer re-made from now with its configuration
otherwise unchanged (nothing pending afterwards). The controller's loop still held the old event, so
it takes a restart of the controller to let go of it; that restart waits for the operator's word.
**Fix** (mesh-controller, branch `fix/a-consumer-on-a-history-stream-starts-from-now`):
- a consumer may say it starts **from now** when it is made. The controller's event consumer does. A
consumer that exists keeps where it is. The server would refuse a changed start anyway.
- `broker consumer-reset <stream> <consumer>` re-makes a stuck consumer from now, its configuration
otherwise kept, and refuses a work queue, where what is pending is work. It is the person's act, said by
a command, rather than a one-off program.
Checked by a live test against a throwaway bus:
- made from now, a consumer holds none of the stream's past and does hold the next announcement;
- asserted again, it keeps its place;
- the default replays all of it;
- a reset leaves nothing pending and keeps every other setting;
- a work queue's consumer is refused.
**Not fixed here:** the client dropping messages while the loop acts on a long merge. Reports and
heartbeats are redelivered or replaced, so nothing is lost for good, but the loop holding everything while
it builds is the shape issue 175 already describes.
## 2026-10-05, after the restart
The consumer re-made from now held nothing, and the restarted controller acknowledged what it was handed.
Two merges made right after reached it within two seconds, and the fixed controller was built, delivered and
took over by its own plan within two minutes. The re-asked build of the agent module was registered and
reached all four machines.
**Not explained by this issue:** the record's merge was still logged twice, with the consumer fresh and
nothing replayed. The duplication has a cause of its own, still to be found. It may be that the forge
announces a merge on two paths, or that one event is handled twice.
## Resolved — 2026-10-05
The controller's event consumer is made from now when it is made, and `broker consumer-reset` re-makes a stuck one from now. Live: after the restart, merges reached the controller within seconds. Two tests main then failed, both skipped without a store, were fixed in mesh-controller PR #52. The merge heard twice was a separate cause: [issue 250](../250-a-merge-made-through-the-forges-tool-is-announced-twice/00-report.md).
@@ -1,37 +0,0 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-controller cmd/mesh-controller]
fixed-by: mesh-controller PR #54
amended-design: 03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md
---
# 249 — A module's new state is refused until a push the merge did not make
## What was observed
A merge gave the agent module a new state, a key-value bucket
([ADR 0201](../../02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md)).
The module's new bundle reached all four machines within a minute of the build, at once. The machines'
bus permissions did not include the new state until a push was made by hand afterwards.
On three machines the module's watch of the new state was refused for about two minutes: "claude-code keeps
and reads no state called config". It recovered only because the module asks again with a back-off, and
the hand-made pushes issued the permissions.
## Why it matters
**The code arrives before the right to use it.** A module that does not retry stays broken until someone
pushes. A module whose first act on start is to read its new state fails its start. Nothing in the plan
says the two must travel together.
**No machine went first.** The bundle reached every machine at the same moment. The rollout the operator
was told — one machine first, then the rest — could not be followed, because the merge had already
delivered it everywhere.
## Open questions
- Should a plan send a machine its membership, the grants that come with a module's new
declarations, in the same push as the bundle, and before it?
- Should a merge that changes a module's declarations (state, events, tools) be delivered to one machine
first, and to the rest only once that one reports it healthy?
@@ -1,30 +0,0 @@
# 249 — Diagnosis
## 2026-10-05
**Grants after code.** Both a plan's rollout and a push send every machine its declaration first, and
issue the memberships afterwards. The order is written into the code on purpose, "because the runtime it is
for arrives with it". That reason holds only for a first assignment, and a membership is kept on the bus for
a runtime that connects later anyway. A membership that failed was only printed, and left "until the next
push". The list of bus users travels in the declaration of the machine that holds the bus, which a module's
rollout reaches only if that machine runs the module.
**No machine first.** The module's upgrade policy sends one machine at a time, but a plan's rollout ignored
it and sent every machine at once. One at a time did not wait for the first machine to come up either: it
stopped only if the publish failed.
**Not answered by the open decision on unseen changes** (a removal, a move or a replacement, shown
before it takes effect). That decision leaves an add-only change alone on purpose, and a new state is one.
**Decided** in [ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md):
grants before code, and one machine first unless a module's policy says *together*.
## Resolved — 2026-10-05
[ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md), live the same evening:
- every push now issues memberships before it sends declarations;
- the bus's machine goes first when its list of users moved;
- a plan sent the build agent to one machine first and to the rest once that machine reported.
The first live rollout exposed a fault in the tier gate. The first machine's report, made between the two sends, was read as stale ([issue 256](../256-a-first-machines-report-read-as-stale-between-the-two-sends/00-report.md)).
@@ -1,49 +0,0 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-catalog modules/gitea]
fixed-by: mesh-catalog PR #63, PR #66
amended-design:
---
# 250 — A merge made through the forge's tool is announced twice
## What was observed
The controller logged the record repository's merges twice, seconds apart, with the same commit, even with
its event consumer freshly made ([issue 248](../248-the-controllers-event-consumer-replayed-a-week-and-held-every-merge-behind-it/00-report.md)).
Counted over one day:
- 8 of the record's merges were logged twice, against 4 once;
- 10 of the catalogue's, against 7 once;
- 2 of the controller's.
A code repository's second line reads differently — "it changed nothing any module the mesh holds is built
from" — so it was taken for a different message.
## Diagnosis
The forge's module announces a merge from two places in the same process:
1. its merge tool, the moment it merges;
2. the poll added for [issue 131](../131-nothing-tells-the-mesh-a-source-moved/00-report.md), which
announces every merged pull request it has not recorded as announced.
The tool never records what it announced, so the poll announces it again 0.5 to 16 seconds later.
Merges made in the forge's web interface or by a plain API call are seen by the poll alone, and those are
the ones logged once. The module's own header comments still say merges are announced "from the tools …
one process only".
No harm was done this time, but only by luck. The second event is absorbed because the first one moved
the controller's record of the source. A repository read only by packaging modules has no such record, so
it would get a second plan. The record module synced twice for each merge.
## Fix
The poll is the only emitter: it sees every path and carries the clone address. The tool merges and
answers the merge commit. A merge made through the tool is heard up to thirty seconds later, which the
module already accepts ("an event a minute late is still an event").
## Resolved — 2026-10-05
The forge module's poll is the only announcer of a merge. It asks only the repositories that moved since its last look, and one pass at a time. Live: three merges were each heard once, and a later merge was planned within a minute.
@@ -1,36 +0,0 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-catalog modules/records]
fixed-by: mesh-catalog PR #64
amended-design:
---
# 251 — The record's checkout could not sync after it ran as another account
## What was observed
Every sync of the record module failed, on every merge and every timer: git refused the checkout as
"dubious ownership". The record tools answered all the while, from the checkout as it last stood, and
nothing said it was stale.
## Diagnosis
The module ran in a container, as the superuser, until its code moved into the machine's tool runtime
([ADR 0198](../../02-DECISIONS/0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md)).
The runtime launches it as the operator account. The checkout's git directory and 2 598 of its files still
belonged to the superuser. Git refuses a repository owned by another user, and the operator account could
change none of those files.
The module also still declared that it needs a container runtime, a leftover of the same move.
## Fix
The module, ported to Go as part of the fix, clones into a directory of its own inside the one it is given:
a directory it makes, and so owns. What the old layout left behind is removed where it is the module's.
Where it is not, the module names it in its status, with the one command that deletes it. The
container-runtime capability is dropped. A sync that fails is still said in the status, as before.
## Resolved — 2026-10-05
The record module, ported to Go, clones into a directory it makes and owns. Live: it synced to the newest commit, with 692 documents. It names seven leftovers of the old checkout that it cannot remove, with the command that removes them; that is the operator's act.
@@ -1,38 +0,0 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-catalog modules/gitea, mesh-controller cmd/mesh-controller]
fixed-by: mesh-catalog PR #63, mesh-controller PR #54
amended-design: 03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md
---
# 252 — A merge's changed modules were read wrong, in both directions
## What was observed
Each of three merges to the catalogue within four minutes planned 99 to 100 modules, and rebuilt modules
they did not touch. The outputs were identical, so nothing was redeployed. The rebuilds cost the build
machine minutes for every merge.
## Diagnosis
**Too many.** The controller treats a changed path outside every module directory it knows as shared code,
and rebuilds every module built from the repository. A new module's directory, or one being removed,
counts: the module is registered only after the merge is planned. Each of the three merges added or
removed a module. The build agent, built from the same repository, then joins the set and becomes tier 0.
**Too few.** The forge's module asked for a hundred changed files and was given fifty, the forge's page
size, and reported the list as whole. A merge of 59 files reached the controller with 50. Had the full
rebuild not hidden it, a module whose own files changed would have stayed unbuilt.
## Fix
- The forge's module reads every page of a pull request's files.
- The controller counts a changed path in a sibling of known module directories as that module's own, not
shared, when that module's definition is among the changed paths. A sibling without a definition, such
as a shared library, still means everything, which is the safe direction. Root files still mean
everything.
## Resolved — 2026-10-05
The forge's module reads every page of a pull request's files. The controller counts a new module's own directory as that module's when its definition is among the changed paths. Live: catalogue merges planned one module each.
@@ -1,55 +0,0 @@
---
status: located
opened: 2026-10-05
located-in: [mesh-controller internal/builder, mesh-controller internal/artifacts, mesh-catalog modules/distribution]
fixed-by: mesh-controller PR #53, mesh-host PR #23, mesh-catalog PR #62
amended-design: 03-DESIGN/01-to-be/18-building-a-module.md
---
# 253 — The store's collector would delete every archive the mesh keeps
## What was observed
The store's nightly collector ([ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md))
was installed the same day, its first run due that night. Measured read-only beforehand:
| | |
|---|---|
| blobs in the store | 8 186 |
| blobs a manifest names, kept by the collector | 2 090 |
| blobs it would delete | 6 096 |
| repositories holding only archives, none named by any manifest | 105 |
Among the archives it would delete were the current bundles of the agent module, the machine host, the
controller and the tool runtime. Each was named by no manifest, though the controller's records keep them.
## Why it matters
Machines keep their unpacked copies, so nothing would have stopped at once. But any fresh fetch of an
unchanged module would have failed: a machine joining, a reinstall, an apply that fetches again, the
controller's own next rollout.
## Diagnosis
Images are pushed with manifests. Archives were published as bare blobs, which no manifest names. The
store's stock collector marks only from manifests, so every bare blob is unmarked, kept or not. ADR 0189's
sentence "what the mesh keeps is still a manifest in the store" was true of images only.
Found alongside: the "five most recent builds" reason kept builds of modules the mesh no longer holds,
forever.
## Fix
- **That night, before the first run:** the collector was changed to a dry run, in its module's
definition, and delivered.
- **Then:**
- every archive is published with a manifest that holds it;
- the controller's sweep holds every kept archive before it lets anything go, which backfills those
already published;
- letting an archive go removes its manifest first;
- a forgotten module keeps nothing.
- Real collection returns once the controller reports no kept archive unheld.
## Where it stands — 2026-10-05
Every kept archive is held: the controller's collection command reports 134 of 134 held, none missing. The window the collector needs is no longer reopened by an apply ([issue 224](../224-an-apply-reopens-a-maintenance-window-by-recreating-what-it-held-still/00-report.md)). The collector still runs as a dry run. Turning it to real collection deletes the layers nothing keeps, which is the operator's word to give; this issue resolves when that change lands.

Some files were not shown because too many files have changed in this diff Show More