Files
hq/01-RESEARCH/017-a-mesh-that-heals-itself/02-now-pragmatically.md
T
jochen b5512bed32 Research 017: a mesh that heals itself
The operator's wish written as intended behaviour for the NATS bus:
every loop compares against what is, repairs by the ordinary path, never
destroys, and raises a condition for what it cannot fix. What is done
before NATS is limited to what survives the move.
2026-09-26 00:54:27 +02:00

2.7 KiB

02 — Now, pragmatically

The intended behaviour lands on NATS. The bus moves after the migration's core (ADR 0106). Until then, work toward it is chosen by one test:

Does it survive the move? A change to what a loop compares against, or to what an adapter can tell about its backend, survives, because it is independent of the bus. A new AMQP queue for health reports, a poller written against the broker's management API, or a condition store built on the current broker does not survive, and is not built.

Done

The provisioner asks the backend, not memory (issue 120). The SDK's harness gained an optional holds on the adapter, asked for every applied consumer every minute. A consumer the backend no longer holds is provisioned again. Being unable to ask is not treated as loss. The cache module implements it first, because its server keeps its users in memory and forgets them all on a restart. That was verified against a real server: a restart erases every consumer's user, and holds answers correctly for absent, present, wrong-password, disabled and deleted. Changes: mesh-sdk PR #7 (0.1.1) and mesh-catalog PR #84.

This is principle 1 applied to one loop. It survives the move unchanged. On NATS, the harness's record of what it applied moves from memory into a key-value bucket, and holds stays as it is.

Next, in order of silent failures removed

  1. holds for the other credential providers. Each backend can answer whether a login exists with the mesh's password without changing anything. Where a backend cannot check a password without logging in, logging in is the check.
  2. The harness's other blind spot. A consumer that goes away while its provisioner is down is never removed. The fix is the same principle in reverse: list what the backend holds, and compare it with what the mesh asks for. Removal stays subject to ADR 0114's rule that it never follows from a login changing.
  3. status reports what owners already know. The host knows which containers it recreated and why. A rotation knows who it waits on. Delivery knows what is outstanding. Surfacing those as conditions in the existing status needs no new transport. It is the shape the NATS condition store will hold.
  4. The loop inventory and the classification in 00. They decide what comes after these three.

Not now

  • Heartbeats, observation subjects, key-value state, advisories: all NATS, all after the move.
  • Handing conditions to agents: designed with NATS, where a task on a subject is native.
  • Choosing the observability store: decided when there is something to store, which is after the move.