Files
hq/01-RESEARCH/017-a-mesh-that-heals-itself/01-the-intended-behaviour.md
jochen b5512bed32 Research 017: a mesh that heals itself
The operator's wish written as intended behaviour for the NATS bus:
every loop compares against what is, repairs by the ordinary path, never
destroys, and raises a condition for what it cannot fix. What is done
before NATS is limited to what survives the move.
2026-09-26 00:54:27 +02:00

5.4 KiB

01 — The intended behaviour

The operator's wish, written as behaviour: what a person or an agent sees the mesh doing. This is a target to design toward, not a design. Every part of it is to be decided through a record before it is built.

Principles

1. Every loop compares what should be with what is, never with what it did. Desired state is the mesh's: assignments, requirements, seats. Observed state is read from the thing itself: the container, the backend, the node. A loop that compares against its own memory of what it applied is blind to anything that changed behind its back. That is issue 120, and it is the pattern this whole effort is written against.

2. Healing is the ordinary path run again, never a second path. Repairing a lost login is provisioning it. Repairing a dead container is converging the node. Repairing a stale declaration is delivering it. A repair that needs its own code is a second way of doing something, which is exactly what the mesh is removing everywhere else.

3. A repair never destroys. Healing may recreate, re-provision, re-deliver and restart. It may never delete a consumer's data, retire a credential someone still uses, or pick a winner between two contradictory states. Where the only repair is destructive, it is escalated.

4. Nothing fails silently. Every condition the mesh cannot repair within its budget becomes visible. It is named, it says since when, why, and who can resolve it. It is visible until it is resolved, and resolved by observation, not by someone clicking it away.

5. What the mesh cannot fix goes to an agent. Per the mission, an agent may be human or not. A condition that needs judgement is handed to one, as work, with what the mesh knows. It is not handed over as a notification that someone may or may not read.

6. Correctness, not only liveness. A running process that authenticates with a dead credential, serves an old version, or routes nowhere is not healthy. What a provision's contract promises is what is checked: the credential authenticates, the route answers, the version is the declared one.

The loop, everywhere

Every part of the mesh that owns something runs the same loop:

  1. know what should be true: from assignments, requirements and seats;
  2. observe what is true: from the thing itself, on its own cadence;
  3. repair the difference by running the ordinary path again, within a budget of attempts and time;
  4. raise a condition when the budget is spent or the only repair is destructive;
  5. clear the condition when observation shows it resolved.

A condition is a durable fact about something the mesh owns, such as a node, an assignment, a provision, a seat or a rotation: what is wrong, since when, the evidence, what was tried, and who can resolve it. Conditions are the one thing a person or an agent looks at to know whether the mesh is right. status is the list of open conditions. When it is empty, the mesh is right, not just up.

On NATS

ADR 0106 moves the bus to NATS, and NATS makes most of this cheaper, because observation becomes something every component publishes rather than something a central process polls.

the wish needs on NATS
every component says it is alive a heartbeat on a subject per node and assignment; silence past its interval is a condition, and nobody polls
every component says what it observed observations published on subjects (mesh.observed.<node>.<assignment>, for instance), consumed by whoever owns the comparison
the last known state survives restarts a JetStream key-value bucket of observed state per owner; the provisioner's "what I applied" and a rotation's step live there, not in memory
conditions are durable and watchable conditions as entries in a key-value bucket, watched by anyone who cares: a surface, an agent, the controller
the bus itself is observed the server's advisories (a consumer exceeding its deliveries, a slow consumer, a client disconnecting) and its monitoring endpoint become observations like any other
a repair is retried, not lost JetStream redelivery with delay, which is the same mechanism ADR 0083's guarantee moves to
work handed to an agent a condition that needs judgement published as a task on a subject an agent's queue group consumes

Who compares. Each owner compares its own: the host for its node's containers and files, a provisioner for its backend, the vault for rotations, the controller for delivery and seats. The controller's observability context does not repair anything. It holds conditions, their history, and the view across the mesh. It notices what no owner can see about itself: an owner gone silent.

What stays human

Some repairs need the operator's key, and the mesh says so rather than pretending otherwise: re-raising the vault or the broker, and recovering a node's identity. These are conditions too, with the procedure named, and they are the only ones that can never clear themselves.

Open

  • The budgets: how many attempts, over how long, per kind of repair.
  • How a condition that needs judgement reaches an agent, and how the agent's action is recorded.
  • Where the observability context stores history (to-be 06 left it open; volume argues against the relational store).
  • Which correctness probe each provision's contract offers, and how often it runs.