Files
hq/04-ISSUES/051-the-mesh-cannot-update-what-it-depends-on/00-report.md
jschoubben 36d9b38a0d An issue is open, diagnosing, located, resolved or wontfix — nothing else
The playbook, the README and the status skill knew five statuses; the cycle check
knew a sixth, 'fixed', and not 'wontfix'. Eleven issues sat in the sixth for weeks
with their fixes shipped, one step short of closed. They are resolved; the check
refuses the word from now on and accepts the one the playbook allows.
2026-09-21 17:38:17 +02:00

8.0 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
resolved 2026-09-14
mesh-host
mesh-catalog
mesh-host 56124c3; mesh-catalog 5e4dc37; mesh-lab 440e265 03-DESIGN/01-to-be/07-the-foundation.md

051 — The mesh can update everything except what it depends on

Symptom

A mesh raised from bare metal runs its store and its broker from a bundle the installer wrote once, at install time, with the images pinned in it. Nothing can change them afterwards. There is no build for them, no version for them to be behind, no upgrade … roll-out, and no way for the mesh to report that either is out of date — because the mesh holds no record of them as modules at all.

Every other thing on that machine has all of it.

Observed on a one-node mesh, where the duplication makes it plain — the same image, twice:

mesh-store      <the postgres image>     the control plane's own records
postgres        <the same digest>        the postgres module's server
mesh-postgres   <mesh-built runtime>     that module's provisioner and tools

One of those two servers can be rebuilt from source and rolled out. The other cannot, and it is the one holding the mesh's inventory, identity and licences.

Why this matters

It is exactly backwards. The store and the broker are the components every other thing depends on, so they are the ones whose security updates matter most, and they are the only ones the mesh has no mechanism to deliver. A release that reaches every module's database does not reach the one the control plane keeps its own records in.

It is invisible rather than reported. status says what is behind its source. The substrate cannot be behind anything, because the mesh does not know it exists as something with a source. So a mesh with a year-old broker reports itself entirely current, which is worse than reporting a problem.

The pattern that would fix it already exists and is proven. The control plane is carried in, raised, and then adopted as an ordinary module pinned to the image that is running — the pivot in ADR 0067. Nothing about that pattern is specific to the control plane. Applied to the store and the broker it would make the substrate a moment rather than a kind: how a mesh starts, not what it permanently is.

It is also why the same image runs twice. A store that cannot be a module cannot provide postgres-database, so a module wanting a database needs a second server. Adoption removes the duplication as a side effect, and the control plane already reaches its three contexts through three separate credentials — which is the shape of a consumer, not an owner.

And the module it becomes is postgres, not store. The naming rule settles it (ADR 0040): where a consumer speaks a protocol, the interface is the protocol, and "database" is not a capability. The control plane's own queries use distinct on and on conflict, so the coupling is to postgres and a store module would promise a swap that fails the first time anybody tries it. The broker collapses the same way with a different outcome — amqp is a protocol that several implementations speak, so amqp is a legitimate provision and lavinmq is one provider of it.

So adoption is not only an upgrade path. It is two rows of a mesh's module list becoming one, twice.

And it is one SERVER each, not two

The broker duplication is running today, not just latent. The lavinmq module raises its own server container and points its provisioner at http://lavinmq:15672 — a second LavinMQ, separate from the substrate's mesh-broker. A mesh with the module assigned runs both.

This is not how LavinMQ is meant to be used, and the module's own provisioner says so: a consumer is given a vhost named for its login, isolated from every other consumer's by the vhost boundary — "the exact analog of postgres's database-per-login". One server hosts the mesh's own control traffic on the / vhost and every consumer's broker as a vhost beside it. Two servers is the same mistake as two postgres containers, wearing AMQP.

So adoption means the lavinmq module does not run a server of its own. Its server IS the substrate broker, adopted; the module contributes the provisioner, the tools, the event consumer and the run-once bootstrap that configures it — all against the one broker. Same for postgres: one server, the control plane's records in their databases and every module's database beside them.

The provisioner already assumes this — it creates a vhost, not a broker — so the change is removing the second server, not building a new isolation model. What has to be designed is only the adoption itself: raising the one broker at genesis because nothing else can, then holding it as the module, over the broker it is.

What makes this harder than it looks

The recursion is real, not incidental. The control plane learns what modules exist by reading its store. A store that is a module is a record inside the thing it is holding up. Adoption is what resolves it — the store is raised by the installer because nothing else can, and only afterwards becomes something the mesh has a record of — but the order of operations during an upgrade needs stating, not assuming: a machine being sent a new store while the control plane is reading from it is not a rollout, it is an outage.

And the broker carries the rollout itself. A declaration reaches a machine over the broker. A broker upgrade is the mesh asking the machine to replace the thing the request arrived on.

Open questions

  • Is adoption the right mechanism, extending ADR 0067 to the substrate, or should the substrate stay outside the module system and gain its own narrower update path?
  • What does a store upgrade look like when the control plane is mid-read? Drain, quiesce, or a window where the mesh accepts that it cannot be asked anything.
  • What does a broker upgrade look like when the instruction travels over it? A machine that is told to replace its broker has to complete the work without being able to report progress.
  • Does adoption also mean one server instead of two by default, with a separate one for the control plane as a choice for those who want the isolation — which is assigning a different provider, a mechanism that already exists?
  • What reports it today? Whatever is decided, status should be able to say the substrate is behind. Today it cannot form the sentence.

How this would be checked

Rule Checked by
The mesh can say its substrate is out of date The store's source moves and status reports it behind, the way it does for any module.
The mesh can deliver a substrate update A store or broker is upgraded on a running mesh and the control plane is answering afterwards, with its records intact.
Nothing runs twice without a reason A mesh with one machine runs one postgres unless somebody asked for two.

Resolution

The foundation's store and broker are adopted in place as the ordinary postgres and lavinmq modules (WBS Phase 3). Each module declares the container the foundation raised — the same name, image, ports, volumes and args — so the applier reconciles it rather than raising a second server; mesh-host's InstallStore (phase3.go) carries in the genesis credentials the mesh cannot invent. The servers bind mesh-wide so consumers reach them, and their provisioners run host-networked. An upgrade of each is proven through its stated window — the store's a pool reconnect, the broker's the harder case of recreating the bus the push travels over — and status now reports both as ordinary modules that can be behind their source. A mesh built from bare metal runs one postgres and one lavinmq, proven 22/22 in the one-node lab. Two follow-ups are tracked: 054 (the pre-filter exposure window) and 055 (multi-node reachability).