Research 017: a mesh that heals itself

The operator's wish written as intended behaviour for the NATS bus:
every loop compares against what is, repairs by the ordinary path, never
destroys, and raises a condition for what it cannot fix. What is done
before NATS is limited to what survives the move.
This commit is contained in:
jochen
2026-09-26 00:54:27 +02:00
parent 56669ee23b
commit b5512bed32
3 changed files with 178 additions and 0 deletions
@@ -0,0 +1,47 @@
---
status: active
initiated: 2026-09-26
touches:
- 00-META/mission.md
- 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0010-delivery.md
- 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
- 03-DESIGN/01-to-be/06-the-controller.md
- 03-DESIGN/01-to-be/09-the-node-lifecycle.md
- 03-DESIGN/00-as-is/09-interfaces-and-observability.md
---
# 017 — A mesh that heals itself
**What.** The behaviour the operator wants: a mesh that runs itself. It notices what is wrong,
repairs what it can, and hands what it cannot repair to someone who can, with the reason. This effort
writes that wish down as intended behaviour, designed for the bus the mesh is moving to
([ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md): NATS). It also records what can be done
pragmatically before that move.
**Why.** The mission is *a mesh that controls itself* ([mission](../../00-META/mission.md)). The
mesh can tell whether it is up. It cannot tell whether it is right. The as-is page on observability says so
([as-is 09](../../03-DESIGN/00-as-is/09-interfaces-and-observability.md)). To-be 06 names an
`observability` context in the controller and leaves its store undecided. Nothing routes a condition
the mesh cannot fix to anyone. The cost is measurable: **46 of the 116 issue reports in this
repository describe a failure that was silent.** A mesh that heals itself is, first, a mesh that stops
failing silently.
**What it touches.** The controller's observability context, the node lifecycle's liveness, delivery
([ADR 0010](../../02-DECISIONS/0010-delivery.md), [ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)),
the provisioner harness, and rotation, which is proposed alongside to-be 27 as ADR 0114.
**Documents.**
- [01 — The intended behaviour](01-the-intended-behaviour.md): the wish, as principles and as how
the mesh behaves once the bus is NATS.
- [02 — Now, pragmatically](02-now-pragmatically.md): what is done before NATS, why it does not
build anything the move would throw away, and what has been done already.
**Next.** Two measurements this effort owes before it can graduate:
1. **Every loop in the mesh**: what it converges, and whether it compares against observed state or
against its own memory. Issue 120 found the provisioner harness trusting memory. The same pattern is
expected elsewhere.
2. **The 46 silent failures, classified**: a missing observation, a loop trusting memory, or a missing
escalation. That shows which mechanism removes the most of them.
@@ -0,0 +1,85 @@
# 01 — The intended behaviour
The operator's wish, written as behaviour: what a person or an agent sees the mesh doing. This is a
target to design toward, not a design. Every part of it is to be decided through a record before it
is built.
## Principles
**1. Every loop compares what should be with what is, never with what it did.** Desired state is the
mesh's: assignments, requirements, seats. Observed state is read from the thing itself: the container,
the backend, the node. A loop that compares against its own memory of what it applied is blind to
anything that changed behind its back. That is issue 120, and it is the pattern this whole effort is
written against.
**2. Healing is the ordinary path run again, never a second path.** Repairing a lost login is
provisioning it. Repairing a dead container is converging the node. Repairing a stale declaration is
delivering it. A repair that needs its own code is a second way of doing something, which is exactly
what the mesh is removing everywhere else.
**3. A repair never destroys.** Healing may recreate, re-provision, re-deliver and restart. It may
never delete a consumer's data, retire a credential someone still uses, or pick a winner between two
contradictory states. Where the only repair is destructive, it is escalated.
**4. Nothing fails silently.** Every condition the mesh cannot repair within its budget becomes
visible. It is named, it says since when, why, and who can resolve it. It is visible until it is
resolved, and resolved by observation, not by someone clicking it away.
**5. What the mesh cannot fix goes to an agent.** Per the mission, an agent may be human or not. A
condition that needs judgement is handed to one, as work, with what the mesh knows. It is not handed
over as a notification that someone may or may not read.
**6. Correctness, not only liveness.** A running process that authenticates with a dead credential,
serves an old version, or routes nowhere is not healthy. What a provision's contract promises is what
is checked: the credential authenticates, the route answers, the version is the declared one.
## The loop, everywhere
Every part of the mesh that owns something runs the same loop:
1. **know** what should be true: from assignments, requirements and seats;
2. **observe** what is true: from the thing itself, on its own cadence;
3. **repair** the difference by running the ordinary path again, within a budget of attempts and
time;
4. **raise** a *condition* when the budget is spent or the only repair is destructive;
5. **clear** the condition when observation shows it resolved.
A **condition** is a durable fact about something the mesh owns, such as a node, an assignment, a
provision, a seat or a rotation: what is wrong, since when, the evidence, what was tried, and who can
resolve it. Conditions are the one thing a person or an agent looks at to know whether the mesh is
right. `status` is the list of open conditions. When it is empty, the mesh is right, not just up.
## On NATS
[ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md) moves the bus to NATS, and NATS makes most of
this cheaper, because observation becomes something every component publishes rather than something
a central process polls.
| the wish needs | on NATS |
|---|---|
| every component says it is alive | a heartbeat on a subject per node and assignment; silence past its interval is a condition, and nobody polls |
| every component says what it observed | observations published on subjects (`mesh.observed.<node>.<assignment>`, for instance), consumed by whoever owns the comparison |
| the last known state survives restarts | a JetStream key-value bucket of observed state per owner; the provisioner's "what I applied" and a rotation's step live there, not in memory |
| conditions are durable and watchable | conditions as entries in a key-value bucket, watched by anyone who cares: a surface, an agent, the controller |
| the bus itself is observed | the server's advisories (a consumer exceeding its deliveries, a slow consumer, a client disconnecting) and its monitoring endpoint become observations like any other |
| a repair is retried, not lost | JetStream redelivery with delay, which is the same mechanism [ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)'s guarantee moves to |
| work handed to an agent | a condition that needs judgement published as a task on a subject an agent's queue group consumes |
**Who compares.** Each owner compares its own: the host for its node's containers and files, a
provisioner for its backend, the vault for rotations, the controller for delivery and seats. The
controller's observability context does not repair anything. It holds conditions, their history,
and the view across the mesh. It notices what no owner can see about itself: an owner gone silent.
## What stays human
Some repairs need the operator's key, and the mesh says so rather than pretending otherwise:
re-raising the vault or the broker, and recovering a node's identity. These are conditions too, with
the procedure named, and they are the only ones that can never clear themselves.
## Open
- The budgets: how many attempts, over how long, per kind of repair.
- How a condition that needs judgement reaches an agent, and how the agent's action is recorded.
- Where the observability context stores history (to-be 06 left it open; volume argues against the
relational store).
- Which correctness probe each provision's contract offers, and how often it runs.
@@ -0,0 +1,46 @@
# 02 — Now, pragmatically
The intended behaviour lands on NATS. The bus moves after the migration's core
([ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md)). Until then, work toward it is chosen by one
test:
**Does it survive the move?** A change to what a loop compares against, or to what an adapter can
tell about its backend, survives, because it is independent of the bus. A new AMQP queue for health
reports, a poller written against the broker's management API, or a condition store built on the
current broker does not survive, and is not built.
## Done
**The provisioner asks the backend, not memory** (issue 120). The SDK's harness gained an optional
`holds` on the adapter, asked for every applied consumer every minute. A consumer the backend no
longer holds is provisioned again. Being unable to ask is not treated as loss. The cache module
implements it first, because its server keeps its users in memory and forgets them all on a restart.
That was verified against a real server: a restart erases every consumer's user, and `holds` answers
correctly for absent, present, wrong-password, disabled and deleted.
Changes: mesh-sdk PR #7 (0.1.1) and mesh-catalog PR #84.
This is principle 1 applied to one loop. It survives the move unchanged. On NATS, the harness's
record of what it applied moves from memory into a key-value bucket, and `holds` stays as it is.
## Next, in order of silent failures removed
1. **`holds` for the other credential providers.** Each backend can answer whether a login exists
with the mesh's password without changing anything. Where a backend cannot check a password without
logging in, logging in is the check.
2. **The harness's other blind spot.** A consumer that goes away while its provisioner is down is never
removed. The fix is the same principle in reverse: list what the backend holds, and compare it with
what the mesh asks for. Removal stays subject to ADR 0114's rule that it never follows from a login
changing.
3. **`status` reports what owners already know.** The host knows which containers it recreated and
why. A rotation knows who it waits on. Delivery knows what is outstanding. Surfacing those as
conditions in the existing `status` needs no new transport. It is the shape the NATS condition store
will hold.
4. **The loop inventory and the classification** in [00](00-overview.md). They decide what comes after
these three.
## Not now
- Heartbeats, observation subjects, key-value state, advisories: all NATS, all after the move.
- Handing conditions to agents: designed with NATS, where a task on a subject is native.
- Choosing the observability store: decided when there is something to store, which is after the
move.