Files
hq/03-DESIGN/01-to-be/06-the-controller.md
jschoubben 90e4a368dc ADR 0081: a decision nothing cites is not yet in the chain
Decisions were the one link the cycle checks skipped, and measuring found 19 of 70 records
orphaned — the credential flow and the module-runtime cluster among them, which is how a
stale premise about a settled decision survived in working memory. cycle.py now refuses an
accepted record nothing cites; the 19 got true homes (design frontmatter, the playbook that
implements 0021, META for the process records). The overview names the practice: spec-driven
development with provenance.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:36:33 +02:00

13 KiB

layer, status, code, updated, decisions
layer status code updated decisions
to-be in-progress
mesh-controller
2026-08-31
02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md
02-DECISIONS/0005-the-node-host.md
02-DECISIONS/0006-the-substrate-and-the-control-plane.md
02-DECISIONS/0008-a-context-owns-its-store.md
02-DECISIONS/0019-how-this-repository-works.md
02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md

The controller

Tier 2. The term appears seventy-nine times across this repository and was defined nowhere, which is how-we-build §5 failing on this repository's own vocabulary.

This document defines it. It does not design the contexts inside it; those are open in research 006.

The definition

The controller is everything that needs to know about more than one node.

That is the whole test, and it is not arbitrary — it follows from ADR 0005. The host applies and does not decide because deciding needs knowledge the machine does not have. So the line falls exactly there:

Question Whose
write this file, with this content, with this mode the host — one machine
which nodes should run the store the controller — needs every node
is this unit running the host — one machine
which peers belong in this node's overlay the controller — needs every node
what does this machine have installed the host reports; the controller records
has this node been unreachable for a week the controller — nobody else is watching

A useful consequence: anything a single machine could answer alone is not the control plane's. If it needs no second node, putting it here is a mistake, and the tier rule will not catch it because the dependency direction is still correct.

What is inside it

Seven contexts and one interface (ADR 0006) — each one earning its place by the test above rather than by being ours:

needs to know about more than one node because
inventory nodes, modules, assignments, versions that is the mesh-wide fact
config settings, secrets, and deriving them onto nodes it derives onto nodes
connectivity overlay, resolution, exposure, filtering, certificates — specified in full in 08-connectivity.md who peers with whom; which node is reachable
provisioning resource grants between modules consumer and provider may be on different nodes
delivery source to artifact to node it targets nodes
observability health, logs, metrics, alerts unreachable for a week is nobody else's to notice
identity agents, humans, services, authorisation credentials follow an agent's node bindings and modality
api the one interface every surface speaks to — it is an interface, not a context

What is deliberately not here. work, knowledge and stream are mesh-hosted applications — first-party, shipped with everything else, and running on the mesh the way anything else does. A task does not need to know a node exists, and being ours does not make something infrastructure. ai is folded into config: a provider licence is an ordinary grant.

record is an open question rather than an eighth entry. Contexts integrate through it (ADR 0008), which makes it load-bearing, and research 006 leaves where it lives unresolved — putting it in the foundation risks recreating the circularity the tier design just removed. Listing it here would settle by naming what has not been settled by arguing.

One of the seven is built. inventory owns a database of that name and holds the node records; the rest do not exist. What it takes to run any of them — the language, and what must already be running before it starts — is ADR 0006, which also records where the build stopped and why: at identity, because what a node presents to prove who it is is not decided anywhere, and a migration is the most expensive place in this system to guess.

These are contexts, not services. They are separate in the sense that matters — each owns its own store, and they integrate through the record rather than by reading one another (how-we-build §4). They are not separate deployables, and research 011 records why that constraint is load-bearing: a single surface can compose them only while there is one interface in front of them.

Nothing outside a context touches its store

The question this answers: can a node write to the registry database? No — and not "only through one node", which is the weaker arrangement it might be mistaken for.

No node holds a credential to any controller store, for writing or for reading.

That is not a new rule here; it is four already taken, and it is worth seeing them together because each one alone reads like a detail:

ADR 0005 the host never queries the mesh database
ADR 0004 a node holds its own identity and nothing else — the shared database credential every node carries today is the exposure this exists to remove
ADR 0008 a context is granted only what it exclusively owns: no shared writes, no read-only roles
ADR 0006 there is no single mesh database, and nothing reads one

So how does anything get in

Over the broker, as a message; the owning context writes.

node ──event/report──► broker ──► the context that owns that data ──► its own store

A node reports what it applied, what it holds, and that it is alive. It states; it does not write. The difference is the whole security boundary: a node that can write cannot be prevented from writing anything, and a node that can only state has its blast radius bounded by what the message vocabulary can say.

Reads work the same way in reverse — a node is told, in declarations. It never asks.

Who actually consumes, and who writes

The controller is the consumer. There is one of it, and the context that owns the data does the write.

node ──► broker ──► the controller, consuming
                          ├─ a node reported what it applied   ─► inventory   writes the registry
                          ├─ a node reported health            ─► observability  writes its own store
                          └─ a grant was requested             ─► provisioning   writes its own store

Seven contexts, one deployable — they are not separate services, so this is one process consuming and dispatching internally, not seven consumers racing. Each context then writes only the store it exclusively owns (ADR 0008).

One consumer is a property worth having, not just a consequence of ADR 0006. The as-is records that two consumers accidentally sharing one queue silently split the traffic between them, each receiving half of what it expects — which has happened, between a module's daemon and its capability server. With one consumer that class of fault cannot arise.

And the broker is the buffer while the controller is down. Nodes go on publishing; messages queue; the controller drains them when it returns. That is what makes ADR 0006's single controller tolerable — an outage delays the mesh's knowledge rather than losing it.

With one consequence that must be bounded before it is discovered: a queue with no limit grows until the broker's disk is full, and the broker is the one component every node depends on. Queues carrying node reports need a maximum length or a message lifetime, and losing the oldest health report is obviously right where losing the oldest declaration acknowledgement is not — so the bound is per queue and is not decided here.

On volume, which is the real worry underneath

Most high-frequency writes are not registry writes, and that is the first thing to check before designing for throughput. The registry is inventory's store: nodes, modules, assignments, versions. Those change when somebody changes something.

Logs, metrics and health checks belong to observability, which owns a different store (ADR 0008). Sending them to the registry would be exactly the shared-schema mistake 0045 exists to stop, arriving through the back door marked performance.

That leaves one genuine funnel: every context's writes go through the process that owns it, and ADR 0006 says there is one of it. For a mesh of a handful of machines this is not a scaling problem, and it should not be solved by giving nodes database credentials — that trades a bounded problem for an unbounded one. If it ever binds, the answers are at the consumer: batch, apply backpressure, or move the highest volume stream out of a relational store entirely.

What observability actually stores its data in is not decided, and it is the one place where volume genuinely argues against a relational store.

What it is not

  • Not the thing that changes machines. It decides; the host applies. It never reaches into a node except through the host.
  • Not a surface. Tier 3 is how people and agents reach it. It has one interface; the surfaces are what speak to that interface.
  • Not the foundation. It runs on tier 1 — PostgreSQL, LavinMQ, an OCI registry (ADR 0006) — and cannot start without them, which is what makes them a lower tier.
  • Not privileged on a node. It has no more access to a machine than the declaration vocabulary allows (ADR 0004).

It is also a consumer

The property that makes tier 2 unlike the others: the controller has requirements of its own. It needs a PostgreSQL database, an AMQP virtual host, and a bucket — the same things any module needs, granted the same way.

That is the circularity the tiers exist to resolve rather than hide: the controller cannot provision its own database, because it is not running yet. So its store is raised from the bundle the host carries, before there is a controller to ask (ADR 0004, research 011).

Its virtual host and its bucket are not in the bundle — by the time they are wanted there is a controller to grant them. Whether the bus must come first is open, and it turns on whether these contexts talk to each other over it.

Where it runs

On nodes, like anything else. It is not a place outside the mesh; it is modules the mesh hosts, assigned to nodes by the same mechanism as everything else.

One node runs it, and nothing takes over (ADR 0006). The node is assigned, never elected — no promotion, no quorum, no split brain.

That is sound rather than merely cheap, because the design already tolerates the controller being absent by construction: a node reconciles from its own store (ADR 0005) and never needed to ask anybody to hold the state it was last given. So the controller being down is not a new failure mode — it is ADR 0004's ordinary disconnected situation, happening to every node at once. What is lost is change, not operation.

The honest half: this node is a single point of failure, recovery is restore rather than failover, and certificate renewal is the clock — an outage outlasting a renewal window expires every public name.

Open

  • The contexts themselves. Decided — seven, by ADR 0006. What remains open is narrower and named there: where the record lives, which research 006 leaves unresolved because the foundation is the one place it must not go.
  • How far it may be split. One deployable today. Splitting a context out costs the single interface a surface depends on (research 011).
  • How many run, and what a node does without one. Resolved by ADR 0006. What remains is measurement: nothing reports how long the controller has been unreachable, or how close a certificate is to expiry — both needed for restore-not-failover to be a plan rather than a hope.
  • What the interface is. One interface is stated; its shape, and whether it is request, subscription or both, is not (research 011).