ADR 0083 and issues 057/058/060/063/064 — the P1 sweep #54

Merged
jschoubben merged 5 commits from issue/057-058-one-push-patient-runtime into main 2026-09-20 10:53:25 +00:00
10 changed files with 238 additions and 8 deletions
@@ -0,0 +1,59 @@
---
topic: the mesh
status: accepted
date: 2026-09-18
deciders: jochen
reconstructed: false
extends: 0010-delivery.md
---
# 83. One push leaves the mesh consistent
## Context
A provision is minted while composing the *consumer's* node; the provider's grant list is a pure
read of secrets already issued from it. So assigning a cross-node consumer and pushing its node
produced a consumer that retried forever against a provider that had never heard of it, until the
provider's node was pushed a second time — an action with no signal to take, documented nowhere,
and invisible whenever consumer and provider share a machine
([issue 057](../04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md)).
Two remedies were on the table: **cascade** — a push also delivers to the machines its compose
changed — or **report** — a push says "now push the provider" and leaves the act to the operator.
## Decision
A push finishes what it starts: after composing and sending the named node, the controller flushes
every *other* machine that is now behind — whose declaration differs from what it was last sent —
by name, in the push's own output, converging over a bounded number of rounds (a flushed send may
itself mint).
Behind is measured against what a machine was last *sent*, not against a before/after snapshot of
this push. The mint that makes a provider behind happens when the consumer is assigned or its
account issued — before `push` runs at all — so by push time the provider already differs from
what it holds, with no in-command delta to detect. The only durable signal is "what it should be"
versus "what it last received", which is the same comparison `push --behind` already makes.
A machine behind for an unrelated reason is flushed by this too, and that is correct rather than a
cost: a named push that knew a machine was behind and left it so would be the very silence this
decision removes. The narrower reading — flush only what this push provably changed — was
rejected because it cannot see a mint that a prior command performed, which is precisely the 057
case.
Reporting alone was rejected because it converts a derived fact the controller already holds into
an operator obligation, and an obligation enforced by nothing is issue 057 restated. The
declaration is computed from the whole mesh; delivering a mesh that is knowingly inconsistent and
merely saying so would make "push succeeded" mean less than it says.
## Consequences
- One push is sufficient for a cross-node consumer: the provider's grants arrive from the same
act that minted the provision. The undocumented rule "push the provider node too" ceases to
exist rather than becoming documentation.
- A named push delivers to every machine that is behind, not only the one named — each named in
the output, never silent. `push --behind` remains the way to reconcile the mesh without naming
a node; a named push now carries the same guarantee for the machines its work touched and any
others already waiting.
- How this is checked: the built-store-cross-node bed registers a cross-node consumer, pushes
only the consumer's node, and asserts the provider minted its vhost — the workaround push is
removed, so a regression fails the bed.
+1
View File
@@ -85,6 +85,7 @@ python3 00-META/checks/index.py fail if stale
- **0002** — [Nodes communicate over a message broker, not over HTTP](0002-nodes-communicate-over-a-broker.md)
- **0003** — [An agent is a persistent employee, not an instance of a pool](0003-agents-are-persistent-employees.md)
- **0077** — [The parts are named controller, foundation, node — not control plane, substrate, master](0077-the-controller-and-the-foundation.md)
- **0083** — [One push leaves the mesh consistent](0083-one-push-leaves-the-mesh-consistent.md)
### Its tiers, from the bottom up
@@ -1,9 +1,10 @@
---
status: open
status: located
opened: 2026-09-17
located-in: []
located-in:
- mesh-controller
fixed-by:
amended-design:
amended-design: 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
---
# 057 — A cross-node consumer is provisioned only when the provider is pushed again
@@ -0,0 +1,16 @@
# Diagnosis — 2026-09-18
The mechanism was already understood when the issue was opened: a provision secret is minted as a
side-effect of composing the consumer's node, and the provider's grant list is a pure read of
secrets already issued — so a provider composed before the consumer existed, and never again, is
blind to it. The open question was the remedy: cascade the push, or report the obligation.
Decided as [ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md): a push
finishes what it starts. The controller captures the digest of what every machine should be
before composing the named node, recomputes after, and delivers to exactly the machines whose
declaration changed because of this push — named in the output, bounded rounds, converging.
Machines behind for unrelated reasons stay the business of `push --behind`.
**Located in:** mesh-controller (the push command). Checked by the built-store-cross-node bed,
which now pushes only the consumer's node and asserts the provider minted the vhost — the
workaround push is removed, so a regression fails the bed.
@@ -1,7 +1,8 @@
---
status: open
status: located
opened: 2026-09-17
located-in: []
located-in:
- mesh-tools
fixed-by:
amended-design:
---
@@ -0,0 +1,17 @@
# Diagnosis — 2026-09-18
1. The exit is the shared runtime's: serve mode connected the broker once and treated any failure
as fatal, so "the overlay tunnel came up a moment after the container" — the normal case on a
joined node — became an exit, a container-runtime restart, and a visible crash-loop.
2. The failing step is the certificate pre-fetch that pins the broker (a 15s TLS probe), but the
shape is general: any reachability failure at startup had the same consequence.
3. Answering the report's open questions: the shared runtime now retries the broker connection
in-process with capped backoff (2s doubling to 30s), aloud, indefinitely — "how long before it
is fatal" is answered *never*, because the failure modes that do not heal are not reachability:
a pinned-certificate mismatch still refuses immediately (an impostor does not become the
broker by being asked again), and configuration errors still exit at once. One-shot commands
(emit, invoke, run) still fail fast — a person is waiting.
4. Checked by the built-store-cross-node bed: the joined consumer must report a container-runtime
restart count of zero.
**Located in:** mesh-tools (the shared runtime's serve mode).
@@ -1,8 +1,10 @@
---
status: open
status: resolved
opened: 2026-09-17
located-in: []
fixed-by:
located-in:
- mesh-catalog
- mesh-controller
fixed-by: mesh-catalog feat/the-mesh-builds-its-catalogue (965c58f); proven by the built-store-cross-node bed and direct builds
amended-design:
---
@@ -0,0 +1,32 @@
# Diagnosis — 2026-09-18
The gap was mechanical for most of the catalogue and structural for a few.
**The mechanical part (fixed).** 44 of 70 modules now carry a `build` section and a Dockerfile
mirroring the eight that already had one: two named bases (build and runtime), the module's own
source compiled against the sdk in the base, and every serve-time entrypoint named in the runtime
image so tools serve, events flow, and a provider's provisioner reconciles in the same process
(the convention hq issue 061 settled — postgres, gitea and lavinmq were retrofitted to it from
running their provisioner as a separate `args` command). Proven end to end: redis, keycloak,
mongodb, mosquitto and openai-consumer build through the mesh's own builder from their own
directory; the built-store-cross-node bed builds and runs the retrofitted lavinmq and postgres.
**The structural part (deferred, each with an owner).**
- Three modules need an artifact the mesh build environment cannot fetch — an npm package
(model-usage: `pg`; anthropic-manager: `tweetnacl`) or a binary from a public image (minio:
`mc`). The build's npm and container runtime reach the mesh's own registry, not the public one;
apt works, npm and image pulls do not. This is its own question — [issue 064](../064-a-mesh-build-cannot-fetch-a-modules-external-dependencies/00-report.md).
- route-proxy's build context is the mesh-controller repository, not its own module directory (it
ships the packaging of a proxy whose canonical source lives with the contract). The build
section names a Dockerfile relative to the module; a cross-repository context is a shape it
cannot yet express. Deferred pending a design decision on cross-repo build inputs.
The remaining modules are exempt rather than unbuilt: the foundation trio (builder,
mesh-controller, distribution — built by the foundation's own path, and distribution cannot build
through the store it provides, ADR 0072/issue 029), and the upstream-image-only modules that run
a stock image and carry no code of their own.
**Located in:** mesh-catalog (the manifests and Dockerfiles), with the convention enforced by
mesh-controller's builder. Resolved for the mechanical majority; the two structural remainders are
tracked as issue 064 and the route-proxy cross-repo deferral.
@@ -0,0 +1,50 @@
---
status: resolved
opened: 2026-09-18
located-in:
- mesh-controller
fixed-by: mesh-controller fix/one-push-is-enough (25e42b3); gated by the built-store-cross-node bed
amended-design:
---
# 063 — The foundation's ports are accepted on input but not forwarded
## Symptom
A joined node reaches the mesh's broker once, then — after the broker restarts — can never reach
it again, and so never receives another declaration. Every check between them passes: the port is
open, the broker is up, the overlay is established.
Observed on the built-store-cross-node bed, intermittently. A joined node's host connected to the
broker, received its first declaration, and worked. When the foundation's own broker was adopted —
which restarts it — the node's link dropped with `CONNECTION_FORCED - Broker shutdown`, then
`connection refused`, then `i/o timeout` for ever. The node was left with no declaration on disk
at all, and every downstream assertion (here, the registry trust) failed as a consequence of a
node that had received nothing.
The bed passed on the runs where the broker did not happen to restart after the joined node first
connected, which is why three green runs preceded the red one on identical code.
## Why it matters beyond the instance
The mesh firewall opens the foundation's ports — the broker above all — in the **input** chain,
from anywhere, so a node can enrol before it has an overlay address. But the broker is a published
container port: a cross-node packet to it is redirected by the runtime and handled in the
**forward** chain, which never carried a rule for the foundation. The first connection survived
only on its conntrack `established` entry; a broker restart dropped the entry, and the next dial
hit the forward chain's default drop.
This is [issue 047](../047-the-firewall-does-not-cover-published-container-ports/00-report.md)'s
lesson — a firewall with no forward coverage says nothing about container ports — reappearing for
the one port the whole mesh depends on, because the foundation is not a module and was added to
input alone. A rule that works only until the thing it governs restarts is worse than none: it
passes every test written before the restart.
## Open questions
- Is the broker the only foundation port that resolves to a container, or should every foundation
port be assumed to be published and forwarded on principle? (The fix forwards all of them, on
that principle.)
- Should a bed assert reachability *after* a deliberate broker restart, so "reachable until it
bounces" can never again read as "reachable"?
@@ -0,0 +1,51 @@
---
status: located
opened: 2026-09-18
located-in:
- mesh-catalog
- mesh-controller
fixed-by:
amended-design:
---
# 064 — A mesh build cannot fetch a module's external dependencies
## Symptom
Three modules cannot be built by the mesh's own builder, each because the build needs an artifact
from a public registry the build environment does not reach:
- A module whose runtime installs an npm package its code imports (`npm install` in the
Dockerfile) fails with a 404 — the build's npm is pointed at the mesh's own registry, which does
not carry the public package.
- A module whose runtime copies a binary out of a public image (`COPY --from=vendor/tool:latest`)
fails with `pull access denied` — the build's container runtime reaches the mesh's registry, not
the public one.
apt-based installs in the same Dockerfiles succeed, so the build has ordinary internet: it is npm
and image pulls specifically that are redirected to the mesh's registry. Observed while giving the
catalogue its build sections (hq issue 060): 44 of 70 modules build from their own source with no
external fetch; these three need one and stop.
## Why it matters beyond the instance
The delivery design says a module is built by the mesh from a repository and a path (ADR 0069),
compiled against the sdk and tool runtime the base images already carry. That holds for a module
whose only dependencies are the base's — but a module with a third-party runtime dependency, or a
tool binary from a vendor image, has no sanctioned way to bring it into a mesh build. The
workstation build script got away with it by running on a host with public npm and Docker Hub;
retiring that shortcut (the no-fake principle) exposes that the mesh has no answer.
Left unanswered, "the mesh builds its own catalogue" is true only for modules that happen to have
no external dependency, and which those are is invisible until each is built.
## Open questions
- Should a module's third-party npm dependencies be published to the mesh's own registry as part
of building it (a dependency is an artifact like any other), or should the build environment
proxy the public registry read-only?
- Is a vendor binary (`mc`, and others like it) the same question as an npm dependency, or does it
want a distinct answer — a module declaring an external image as a build input the mesh
pre-fetches and pins, the way it pins its bases?
- Should a manifest that declares a build whose Dockerfile fetches from a public registry be
refused at `module add`, so the gap is caught at registration rather than at build?