Merge pull request 'ADR 0235: the bus is backed up by its own snapshot of each stream' (#144) from feat/bus-snapshot into main
mesh/delivery held for a person: merged without a passing check: only a person decides that it goes on

This commit was merged in pull request #144.
This commit is contained in:
2026-10-06 16:24:54 +00:00
6 changed files with 224 additions and 2 deletions
@@ -213,6 +213,15 @@ somewhere that has it, which is not built; mailu's queue is valuable though tran
and the operator's keys are valuable, and are backed up only on machines that hold `node-backup`
(today, neither workstation does).
> **Progressive insight — 2026-10-06.** The first of these items is resolved, and was a fact about what
> was built rather than part of the decision. The record said the bus's streams "are copied as live
> files" and that a consistent copy "is not built". It now is: the nats module's streams item is
> protected by a dump — `backup: {dump, into}`, this record's own mechanism — that runs a snapshot
> program built into the bus's image, through the server's snapshot API, under the bus module's own
> account granted that API and nothing else; the live store is no longer copied, and a restore builds
> a new store beside the live one ([ADR 0235](0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md)).
> The bus's row in the table above stands: its streams and buckets are valuable, written all the time.
## Consequences
- **The first run of the self-check records what every machine holds**, and the first night backs up
@@ -0,0 +1,163 @@
---
topic: what runs on it
status: accepted
date: 2026-10-06
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md
---
# 235. The bus is backed up by its own snapshot of each stream, taken under the bus module's account
## Context
[ADR 0233](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md)
made every module declare its data and derived the night's backup from it. It left one item flagged
for the operator: **the bus's streams were copied as live files.** The bus keeps everything it holds in
JetStream — every stream, and every key-value bucket, which is a stream too: conditions and their
history, calls, the hand-act log, the controller's lease, assignments, events, every module's state.
The backup holder read the store's directory as it stood while the server wrote it. A copy taken that
way can hold a block half-written, or an index from before the block it indexes, and may not restore;
nothing would say so until the day it was needed. The operator asked for a safe snapshot the same day.
What was true when this was decided, read from the bus on the control node without changing it:
- 19 streams, 141,703 messages, 117 MB: the events stream 97 MB (139,322 messages), the calls
bucket 17.6 MB, the machines' declarations 1.6 MB, the rest under 200 KB each. All on disk; none in
memory.
- The server's own way to hand a stream out whole is the **snapshot API**: a request names a subject
to deliver to; the server answers with the stream's configuration and its state, then sends an
archive in chunks, each carrying a reply subject the server waits on past an 8 MiB window, and ends
with an empty message. The server goes on taking writes throughout; nothing is paused or
reconfigured; only a second snapshot of the same stream at once is refused. The client library the
mesh uses offers no helper for it; the operator's command-line tool and the server implement it.
- **Only the controller could reach the JetStream API** ([design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md)
§3–§4): it is the one writer of stream definitions, and its grant is the whole API.
- The bus's image is the upstream server and an entrypoint; nothing in it could take a snapshot. Its
module declared no account on the bus.
- A dump is a shell command the backup holder runs as root before it copies the item the dump writes
into ([to-be 43](../03-DESIGN/01-to-be/43-backups-against-mistakes.md)); the stores' dumps run
`docker exec` into their own container. A dump can name the module's directories and nothing else —
no machine path (ADR 0112), and no reference to where a bundle is unpacked.
## Considered Options
1. **Keep copying the live files.** Rejected: it is the thing that may not restore, and a backup that
may not restore is worse than none, because it is believed.
2. **Stop the bus for a cold copy each night.** Rejected: every machine, every tool call and every
hand act depends on the one server; an outage a night to make a copy the server can give while
running is a cost with nothing bought.
3. **The controller takes the snapshots** — it already holds the whole JetStream API — on a schedule
before the night, into its own state directory. Rejected: it puts backup work in the core, whose
rule is that it does as little as it can and fails loudly (ADR 0227); the bus's protection would
depend on the controller's machine and schedule, and a dump of one module would be a file of
another; and the reading would be done under the widest authority on the bus.
4. **A dedicated snapshot principal with an issuance path of its own.** Rejected: a second way to mint
and seal a bus credential, when `module issue` already mints a module's account and seals it to the
machine.
5. **The program in the nats module's tools bundle**, run by the node's runtime. Rejected: a dump
cannot name where a bundle is unpacked, and a restore needs the server itself, which the bundle does
not carry.
6. **The program in the bus's own image, run by the backup holder's dump under the bus module's own
account, granted the snapshot API and nothing else.** Chosen.
## Decision
**1. The bus's streams are protected by a dump, and the live store is no longer copied.** The nats
module's `jetstream` item says `backup: {dump, into: snapshots}`. The dump runs, in the bus's own
container, a program built into the bus's image, `mesh-nats-snapshot`, with the module's credential on
its standard input; it writes one archive to standard output into the module's `snapshots` directory
(rebuildable), renamed into place only when whole. The holder backs up that directory. **No transition
period copies both**: the restore points of the live files taken before stay in the repository for
the rotation's two weeks, two months and half a year, so the old copies are at hand until they age
out, while the snapshot protects the bus from its first night.
**2. The module holding `mesh-broker` is the bus, and its account may snapshot the bus and do nothing
else.** When it declares an own secret named `broker`, the controller composes its user
`<machine>.<module>` with exactly: the stream names, one stream's information, the snapshot request for
any stream, the acknowledgement subjects the server puts on the chunks, and its own inbox. Nothing that
defines, changes, purges or writes a stream — the writers table (to-be 45 §1) holds that at
composition, the same check that refuses any second writer. It answers nothing. A bus module that
declares anything else to say or hear on the bus is refused by `module check`, because it would be
granted nothing. Its tools stay the node runtime's to serve. **The controller's own grant, and so the
installer's first user list, are unchanged.**
**3. A snapshot is the server's, taken gently.** One stream at a time, in name order, with a pause
between two; each chunk acknowledged as it arrives so the server's window keeps moving; a bound per
stream (ten minutes) and for the whole (thirty). A memory stream holds nothing across a restart and
cannot be snapshotted; it is listed as skipped with why. Any stream failing fails the whole night for
the bus: a partial copy where a whole one is expected is the silent failure this exists to prevent.
The archive is a tar: first a **manifest** — when it was taken and how long it took, the server's
version, and per stream its subjects, messages, bytes, first and last sequence, consumers, and the size
and SHA-256 of each file beside it — then, per stream, the server's configuration and state and the
server's own archive of it, consumers included, in the layout the operator's command-line tool also
reads.
**What a snapshot promises, as measured, not assumed:** every message up to the last sequence the
manifest gives for a stream, exactly. The server states a stream when the snapshot starts and reads its
blocks a moment later, so a stream written during its snapshot carries some of what arrived in that
moment as well — the first test against a server being written every two milliseconds restored to
sequence 6025 where the manifest said 6024, and at a hundred megabytes to 90075 where it said 90024,
with some messages of that tail present and some not. A restore ends at or after the manifest's
sequence, and one that ends before it is refused.
**4. Restoring never touches the live bus.** A stream is restored only where it does not exist, and on
the live bus every stream exists. So the same program builds a **new store beside the live one**: it
starts the bus's own server binary, from the bus's own image, on loopback, with the bus's one account,
restores every stream, holds each to the manifest, and stops it. Swapping the new store in for the
live one, with the bus stopped, is a person's act and a planned bus step (to-be 45). The steps are in
to-be 43 and the module's README.
**5. Nothing new watches it.** A failed snapshot is a failed night for the bus, and its item's last good
backup ages into the existing `backup-stale` condition, read by D13; the holder measures the snapshots
directory like any item. Each night's runtime and size are in the manifest, and the archive's size is
the holder's measurement of the item.
## Consequences
- **The bus's image changes, so the bus restarts once**: the image now carries the program (built
from a Go toolchain declared beside the server's base). That restart is a planned bus step, done by a
person, not a side effect of a push.
- **Rollout order**: the controller first (it composes the new user once the module declares an
account; until then nothing changes); then the catalogue; then `module issue nats --node <the bus's
machine>` **before** that machine's next push — a push refuses a module whose declared account was
never issued, and says so; then the push, as the planned bus step. The first night after it is the
first snapshot.
- **The size of a night**: at most the streams' bytes — 117 MB today, less once compressed. A local run
of the program at that size (105 MiB, 90,000 messages, written to throughout) took half a second for
the snapshot and a third of a second to restore the largest stream. The archive's compression is
per block, so a night's restore point shares what did not move with the night before; the events
stream ages out a week at a time, so most of it moves within the week.
- A restore is a person's work of four steps, not a verb. A verb that swaps a store would stop the bus
from a tool call; that stays a hand act until the planned bus step (to-be 45) is a verb itself.
- A stream's messages written during its snapshot may be partly there; nothing the mesh keeps depends
on the last few milliseconds of a night.
## How it is checked
| What | Checked by |
|---|---|
| a snapshot of every stream and bucket of a server being written to, restored into a fresh server and into a new store served by a third: every message to the manifest's sequence identical, deletes still deleted, buckets and an object store working, consumers where they were; writes during the snapshot all accepted | mesh-catalog `modules/nats/snapshot/live_test.go` (`TestASnapshotOfALiveBusRestoresToTheSameContent`, throwaway `nats:2.11`) |
| nothing restored over a live stream; a damaged archive refused | the same test |
| the manifest's dump, exactly, in the module's image with the module's configuration; a refused night fails and keeps the last snapshot | `modules/nats/snapshot/image_test.go` |
| the composed grant is the snapshot API and an inbox; it writes nothing (the writers table); only the broker seat's holder gets it | mesh-controller `internal/broker/snapshot_test.go`, `internal/inventory/busrecords_test.go`, the composition golden |
| that grant, composed, against a real server: a whole snapshot taken; every write and every wider subscription refused | `internal/broker/snapshot_live_test.go` (`TestTheComposedSnapshotUserCanSnapshotAndCannotWrite`) |
| a bus module declaring more on the bus is refused; the catalogue's bus is protected by the snapshot | `internal/catalogue/bus_snapshot_test.go` |
| the installer's first user list still matches the controller's | `TestTheInstallersFirstUserListIsWhatTheControllerWouldCompose` |
| live, after rollout | `node-backup.backed-up` shows the bus's item with a recent good backup; `doctor` passes D13; `nats_users` shows the bus module's user with four grants |
## References
- [ADR 0233](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md)
— its uncertain item, resolved here.
- [ADR 0214](0214-backups-guard-against-mistakes-and-stay-on-the-machine.md),
[to-be 43](../03-DESIGN/01-to-be/43-backups-against-mistakes.md) — the night, the dump, restoring
beside.
- [Design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md) §3–§4 — the controller remains the one
writer of stream definitions.
- [To-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) — the writers table; the bus
as a planned step.
- mesh-catalog `modules/nats` (`snapshot/`, `Dockerfile`, `module.json`, `README.md`); mesh-controller
`internal/broker` (`BusSnapshotGrants`), `internal/inventory/busrecords.go`,
`internal/catalogue/manifest.go`.
+1
View File
@@ -333,6 +333,7 @@ python3 00-META/checks/index.py fail if stale
- **0228** — [A value given by hand lives only until its module's first good start](0228-a-value-given-by-hand-lives-only-until-its-modules-first-good-start.md)
- **0232** — [A binding to a consumer's data moves only by a person](0232-a-binding-to-a-consumers-data-moves-only-by-a-person.md)
- **0233** — [A module declares the data it holds, and the mesh protects and watches it from that declaration](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md)
- **0235** — [The bus is backed up by its own snapshot of each stream, taken under the bus module's account](0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md)
### How it is built
+9 -1
View File
@@ -8,8 +8,9 @@ code:
- mesh-catalog modules/nats (to be written)
- mesh-sdk src (the protocol's NATS binding, step 3)
- mesh-tools node-tools/internal/bus (a module's state, ADR 0201)
updated: 2026-10-04
updated: 2026-10-06
decisions:
- 02-DECISIONS/0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md
- 02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md
- 02-DECISIONS/0167-a-membership-carries-what-its-module-receives-and-who-the-mesh-is.md
- 02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md
@@ -175,6 +176,13 @@ call is a timeout the caller already handles.
Streams and consumers are objects the controller creates at genesis and asserts on start; a module
declares nothing about them. The controller is the only writer of stream definitions.
*Added 2026-10-06, [ADR 0235](../../02-DECISIONS/0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md):*
one other user reads the streams whole, and writes none of them. The module that is the bus — the one
holding `mesh-broker` — has an account of its own whose only grant is the snapshot API: the stream
names, a stream's information, the snapshot request and its acknowledgements, and its own inbox. The
night's backup takes every stream through it ([to-be 43](43-backups-against-mistakes.md)). The
controller stays the only writer of stream definitions; the writers table holds that at composition.
### The store window, and what moving it into the server changes
The guarantee ([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)) is
@@ -5,8 +5,11 @@ code:
- mesh-controller: internal/catalogue/seats.go (node-backup), internal/catalogue/seat_contributions.go (BackupSeat, CheckBackup, a contribution's directories)
- mesh-catalog: modules/restic (the holder; it measures the declared data and deletes a retired item, ADR 0233), every module's `data` section
- mesh-controller: internal/catalogue/data.go (the lines derived from a module's data, ADR 0233)
- mesh-catalog: modules/nats/snapshot (the bus's snapshot and its restore, ADR 0235)
- mesh-controller: internal/broker (the bus module's snapshot-only account, ADR 0235)
updated: 2026-10-06
decisions:
- 02-DECISIONS/0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md
- 02-DECISIONS/0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md
- 02-DECISIONS/0214-backups-guard-against-mistakes-and-stay-on-the-machine.md
- 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md
@@ -72,6 +75,37 @@ has not seen, so the object store's first night costs its full size and later ni
changed. Then the rotation prunes to 14 daily, 8 weekly and 6 monthly snapshots. A dump that fails
fails the night for that module only; the others are still taken.
## The bus: the server's own snapshot, never its files
*Added 2026-10-06, [ADR 0235](../../02-DECISIONS/0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md).*
The bus keeps everything it holds in streams — its key-value buckets are streams too — and its store
is written all the time, so a copy of the store's files may not restore. Its streams are protected by a
dump like a database's: in the bus's own container, a program built into the bus's image asks the
server for each stream through the server's snapshot API, one at a time, acknowledging each chunk so
the server's flow control keeps moving; the server goes on taking writes, and no stream is paused or
reconfigured. It runs under the bus module's own account, which may take snapshots and do nothing
else. It writes one archive into the module's snapshots directory, which the night copies: a manifest
— when, how long, and per stream its messages, bytes, sequences, consumers and the checksum of each
file — then each stream's configuration and the server's archive of it. The live store is not copied.
Every message up to the last sequence the manifest gives for a stream is in the archive, exactly; a
stream written during its snapshot may carry part of what arrived while its blocks were read, after
that sequence. A failure fails the bus's night, keeps the previous archive, and is said as
`backup-stale` like any other.
**Restoring the bus** is done beside it, never over it — a stream is restored only where it does not
exist, and on the live bus every one exists:
1. `restore` the bus module's snapshots from a named night, beside the live ones;
2. check the archive against its manifest (`mesh-nats-snapshot verify`);
3. build a new store from it with the bus's own image: the program starts the bus's server on loopback
with the bus's one account, restores every stream, holds each to the manifest and stops it;
4. swap the new store in for the live one with the bus stopped — a person's act and a planned bus step
(to-be 45), in one command line, because the machine's node-engine starts a stopped container again
at its next reconcile. The store kept aside is removed by a person once the bus is seen whole.
The module's README has the commands.
## The verbs
On the seat, for a person or an agent:
@@ -92,6 +126,9 @@ a manifest. Not off-site: a lost machine loses its backups with its data, by the
## Proving it
- A night that fails, or does not run, reaches the operator's output channel, naming the module.
- The bus's snapshot is proven by restoring it, in tests against the bus's own release: a server
written to throughout its snapshot, restored into a fresh server and into a new store served by a
third, compared message by message (ADR 0235).
- Weekly, the holder checks the repository's integrity and restores the newest dump of one database,
in rotation, into a throwaway instance with no network, comparing table row counts with the live
database.
@@ -408,7 +408,11 @@ controller: the streams are snapshotted; a `bus-maintenance` condition is open f
the bus is replaced; afterwards D6, D7 and the round trip must pass, or the step is reported failed and
the snapshot is the way back. A step whose new version cannot be reverted (the bus's 2.10 → 2.11 is one)
says so before it starts and runs only on a person's explicit word, recorded as a hand act. Whether the
bus becomes a cluster that can be upgraded live is left to its own effort.
bus becomes a cluster that can be upgraded live is left to its own effort. *2026-10-06:* the snapshot
and the way back from it exist — the bus image's own snapshot program, the same one the night's backup
runs, and a restore that builds a new store beside the live one for a person to swap in
([ADR 0235](../../02-DECISIONS/0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md),
[to-be 43](43-backups-against-mistakes.md)).
## 9. Before merge: facts and replays (rule 9)