Genesis is a pivot, public routing is name-agnostic, and five issues the fake registry was hiding #32

Merged
jschoubben merged 19 commits from design/bootstrap-is-a-pivot into main 2026-09-11 19:56:50 +00:00
12 changed files with 689 additions and 5 deletions
@@ -1,5 +1,5 @@
---
topic: model access
topic: what runs on it
status: accepted
date: 2026-09-06
deciders: jochen
@@ -1,5 +1,5 @@
---
topic: model access
topic: what runs on it
status: accepted
date: 2026-09-07
deciders: jochen
@@ -0,0 +1,115 @@
---
topic: the tiers
status: accepted
date: 2026-09-09
deciders: jochen
reconstructed: false
extends: 0007-connectivity.md
---
# 66. Public routing is name-agnostic, its names are resolved inside the mesh, and an internal authority can certify them
## Context
**[ADR 0007](0007-connectivity.md) and [connectivity §3](../03-DESIGN/01-to-be/08-connectivity.md)
made a public route a grant: a workload that must be reachable requires a route, the proxy provides
it, the consumer contributes the name it wants and the port it listens on.** What was never pinned
is **what that name is** — and building a whole mesh in the lab showed the gap costs more than it
looks.
**The catalogue shipped each route as a full domain.** A module that needed a public name carried
that name, in full, as a literal in its manifest. Running the same catalogue against a different
domain — a lab standing in for production, or a second operator's mesh — meant overriding that
literal on every routed module, per node. The mesh was, in effect, carrying a **map of names to
services**: the one thing it should never hold, because a name is the operator's choice (one runs
the forge at `git`, another at `code`) and the domain is the node's, and neither is the mesh's to
know.
**And a second gap surfaced the moment an internal issuer tried to certify those names.**
[Connectivity §5](../03-DESIGN/01-to-be/08-connectivity.md) already states the issuer must be
configurable and that the lab runs its own ACME authority. With that authority wired to the proxy,
issuance still could not complete: the authority accepted the order and offered a challenge, then
**could not connect to the validation target.** Nothing inside the mesh resolved the public route
name. The mesh publishes each `<node>.internal` name into every container, but not the public names
the proxy serves — so a validator living in the mesh had no address to reach, and a name the mesh
cannot resolve is a name it cannot have certified.
**The two are one problem.** A name the mesh can *compose* from parts it is given, and *propagate*
to whoever needs to resolve it, is exactly a name it can also have *certified* — and the reverse:
without the composition and the propagation, neither the routing nor the certificate is the
operator's to move between meshes.
## Considered Options
**1. Keep the full domain in the manifest, override per node.** The status quo. It works, and it is
wrong in the specific way this repository cares about: the catalogue holds a domain map, lab and
production differ by an override on every routed module rather than one fact, and a module manifest
names something — the public domain — that belongs to the node, not the module. An unowned name in
the wrong place is the shape of a leak.
**2. A module declares a label; the node declares its public domain; the mesh composes.** The route
contribution carries a subdomain the operator chose, the node carries its public domain as
node-level configuration, and the mesh joins `<label>.<public-domain>` and grants exactly that. The
mesh interprets nothing. Lab-versus-production becomes one node setting. Chosen.
**3. For certification, issue only from a publicly reachable node against a public authority.**
This is already true for public meshes and stays true. It is not an option for a lab or an
internal-only mesh: there is no public authority to answer, and no public reachability to validate
against. An internal issuer is required there — and an internal issuer must be able to *validate*,
which it cannot do unless the routed name resolves and is reachable **inside** the mesh. So the
resolution gap is not optional to close; it is what makes an internal authority possible at all.
## Decision
**The mesh core holds no map of hostnames, subdomains or domains.** A route contribution carries a
**label** (the subdomain) chosen by the module's operator. A node contributes its **public domain**
as node-level configuration. The mesh composes `<label>.<public-domain>`, grants exactly that name,
and never interprets what it means. One operator's forge at `git.example.tld` and another's at
`code.other.example` are the same module with two facts supplied around it.
**When the proxy is granted a name, the mesh publishes that name → the node that serves it into
internal resolution, mesh-wide** — the same mechanism, and the same "given by the mesh, not chosen
by a module," that already writes `<node>.internal` into every declared container
([connectivity §2](../03-DESIGN/01-to-be/08-connectivity.md)). The mesh propagates the names it was
told to serve. It still knows nothing about what any of them mean.
**An internal authority certifies those names by the same path a public one would.** The proxy is
pointed at whichever issuer the mesh names — a public ACME authority, or an internal one — and
trusts that issuer's root; nothing else about issuance changes. The internal authority validates by
reaching the routed name, which the clause above has just made resolvable inside the mesh. So the
three are one decision: **compose the name, propagate it, certify it** — each is meaningless without
the one before it.
## Consequences
- **Lab-versus-production is one node-level `public-domain` setting**, not an override on every
routed module. The same catalogue runs against any domain.
- **The manifest layer needs composition it does not yet have.** Today a route name is stored as a
literal, with no interpolation of a node's domain into a module's label. Until that exists, the
composed name is produced by a per-node settings override — a stopgap that reproduces option 1's
per-module cost and is explicitly *not* the design.
- **An internal issuer depends on route-name resolution.** Its challenge validates against the
routed name; without that name in internal resolution, issuance for it cannot complete inside the
mesh. The lab found this as a live failure, not a theory.
- **Nothing about the public path changes.** A publicly reachable node issuing a public name from a
public authority is untouched; this widens the same shape to names and meshes that are not public.
**How each is checked** — an unenforced rule is indistinguishable from a wrong one:
- **Name-agnostic:** the same catalogue resolves against two different public domains by changing
one node setting and nothing else; and no module manifest contains a full public domain. A
manifest that pins an FQDN is the smell the check looks for.
- **Resolution:** a request to a routed name, made from inside the mesh, reaches the workload that
serves it — and, the sharper check, issuance for that name against the internal authority
completes, which it cannot unless the validator resolved and reached the target.
- **Internal authority:** a TLS handshake to a routed name verifies against the internal root and
nothing else, the same shape §5 already uses for internal node-to-node names.
## References
- [ADR 0007 — connectivity](0007-connectivity.md), which made a route a grant and named exposure,
resolution and certificates as one context.
- [ADR 0009 — modules and the graph](0009-modules-and-the-graph.md), the provide/require/contribute
vocabulary a route and an authority both use.
- [Connectivity design §2, §3, §5](../03-DESIGN/01-to-be/08-connectivity.md), amended alongside this
record.
+129
View File
@@ -0,0 +1,129 @@
---
topic: the tiers
status: accepted
date: 2026-09-10
deciders: jochen
reconstructed: false
extends: 0006-the-substrate-and-the-control-plane.md
---
# 67. Genesis is a pivot: a temporary control plane installs the registry that makes it permanent
## Context
**The mesh builds its own modules into its own registry, and there is no public registry for them
— that is the point, not an omission.** A module is cloned from the forge, built, and published
where the mesh can move, replace and back it up. Nothing about that arrangement wants a copy of the
mesh's code hosted by somebody else.
**Every image must be pinned by digest** ([ADR 0006](0006-the-substrate-and-the-control-plane.md)),
and the reasoning is exactly right: a bundle is applied where no mesh exists to check anything
against anything, so what it names must be exact.
Those two sentences are individually correct and together they close a door. The digest a pin means
is a *manifest* digest, and a manifest digest is **assigned by a registry when something is pushed
to it**. The control plane's image is built from source and pushed nowhere, so it has no such
digest, so it cannot be named — and a registry cannot be installed without a control plane to
install it. That is not a pin. It is a dependency the pinning rule created by accident.
**It went unnoticed because the lab hid it.** The lab raised a disposable registry, stocked it from
a workstation, and rewrote every image reference to point at it — so the lab bootstrapped along a
path no real machine has. A first node in the lab always worked, and a first node anywhere else had
no path at all. Every bootstrap fault found this year was found late for the same reason: **the
install procedure existed only as a test fixture**, and a fixture is free to invent what it needs.
## Considered Options
**1. Publish the mesh's own images to a public registry.** Rejected on the premise: there is no
public registry for the mesh's modules and there is not meant to be. It would also make raising a
mesh depend on somebody continuing to host its code, which is the dependency the whole arrangement
exists to remove.
**2. Build the control plane from source on the first machine.** Rejected. A bare machine would
need a toolchain and a working tree — and worse, the source lives in a forge **that runs on the
mesh**. A total rebuild would then need the mesh it is rebuilding. Acceptable for adding a node to
a healthy mesh; useless for the case that matters.
**3. Keep a disposable registry as an install step.** Rejected. It exists in no production, and
concealing this problem is precisely what it has been doing.
**4. Pivot through a temporary control plane.** Chosen.
## Decision
**An image may be named by the digest of its own configuration.** A bare `sha256:…` names an image
the machine already holds — content-addressed, immutable, unforgeable, and requiring nothing to
have served it. It satisfies what the pinning rule asks for; the rule simply never contemplated an
image that no registry had ever seen. It is legal exactly where nothing could have served one.
**Genesis is a pivot**, in this order:
1. the installer **carries the control-plane image** and loads it onto the machine
2. a **temporary** control plane is raised from it, named by that image's own digest
3. the **registry module is installed** — its image is upstream and it is never built, which
[`04-ISSUES/029`](../04-ISSUES/029-the-artifact-store-cannot-be-delivered-by-the-artifact-store/00-report.md)
already settled: a module that provides the artifact store cannot be delivered through it
4. the control-plane image is **pushed into the mesh's own registry**, which assigns it a manifest
digest — the first one it has ever had
5. the control plane is **reinstalled as an ordinary module** pinned to that digest
**The host performs the replacement, not the control plane.** Tier 0 outlives tier 2: the control
plane composes a declaration naming the registry-pinned image, and the host applies it and recreates
the container. Nothing is asked to replace itself while running, and the control plane is stateless
— what it knows is in the store.
**The installer is tier 0, and a separate program from the host.** Bootstrapping is by hand and
changes the machine, which is tier 0's definition. But the host states that it *connects to nothing
and listens on nothing*, and that claim is what makes the one thing running forever on every
machine auditable. An installer connects to plenty. Same tier, same delivery, different program.
## Consequences
- **The control plane stops being a special case.** It becomes an ordinary module with an ordinary
image in the mesh's own registry — so the mesh can build and roll out **its own upgrades**, which
is what a mesh that runs itself was always reaching for.
- **The bundle's job shrinks** to raising a temporary control plane exactly once.
- **The source builds the installer; it does not run it.** Cloning moves to a release machine, where
a forge being available is an ordinary working assumption, and leaves the disaster-recovery path
where it very much is not.
- **The lab's disposable registry is deleted.** The lab bootstraps by running the same program a
bare machine runs — the only arrangement in which the installer cannot quietly drift out of truth
again.
- **Two public images remain at genesis** — the store and the broker. An air-gapped install would
embed those too, at a much larger artifact; that is a build variant, not a different design.
- **A control-plane module manifest must exist**, and did not.
- **The registry may require nothing.**
[`04-ISSUES/029`](../04-ISSUES/029-the-artifact-store-cannot-be-delivered-by-the-artifact-store/00-report.md)
states that a module providing the artifact store may not *build* artifacts, because there is
nowhere to put them until it runs. The pivot shows that is the narrow case of a wider rule: **it
may not require anything the store is needed to deliver.** Found the hard way — an unrelated
change gave the registry a public name and, with it, a route requirement. At genesis nothing
provides a route, and nothing can, because the routing stack needs images and images need the
store. The same cycle, re-entered through a door the existing wording did not cover.
- **The handover is the sharp edge.** For one moment the bundle and the module both describe the
same container, and the host tracks what it owns. If a safe handover is not expressible with what
exists, the install **stops before it** and says what is missing. A machine left without a control
plane cannot be fixed remotely, so a partial install that halts cleanly is the better outcome.
**How each is checked** — an unenforced rule is indistinguishable from a wrong one:
- **Naming by its own digest:** a machine that can reach no registry at all raises a control plane.
- **The pivot completed:** after installing, the running control plane's image is pinned by a digest
**the mesh's own registry assigned** — not by an image id. If it is still the image id, the pivot
did not happen and the mesh cannot upgrade itself.
- **No fiction left in the lab:** the scenario declares no registry machine, and the bed bootstraps
through the installer rather than around it.
- **The handover:** a machine whose control plane has been replaced still has one, and it answers.
- **The registry requires nothing:** its manifest is resolvable on a mesh that has no other module
in it. A requirement added to it later is caught where it is written, rather than by a genesis
that cannot complete — which is how this one was found.
## References
- [ADR 0006 — the substrate and the control plane](0006-the-substrate-and-the-control-plane.md),
which pinned images by digest and named what a first node fetches.
- [ADR 0005 — the node host](0005-the-node-host.md), which makes tier 0 the one thing installed by
hand and the only thing that changes a machine — the property this keeps true by shipping the
installer beside the host rather than inside it.
- [`04-ISSUES/029`](../04-ISSUES/029-the-artifact-store-cannot-be-delivered-by-the-artifact-store/00-report.md),
the same cycle one layer down, and the rule that a registry module is named and never built.
+5
View File
@@ -98,6 +98,8 @@ python3 00-META/checks/index.py fail if stale
- **0031** — [The control plane authenticates nobody, so identity is a module](0031-the-control-plane-authenticates-nobody.md)
- **0033** — [The substrate is a store and a broker](0033-the-substrate-is-a-store-and-a-broker.md)
- **0036** — [Bootstrap ends at a usable mesh, and the first credential comes from a person](0036-bootstrap-ends-at-a-usable-mesh.md)
- **0066** — [Public routing is name-agnostic, its names are resolved inside the mesh, and an internal authority can certify them](0066-public-routing-is-name-agnostic.md)
- **0067** — [Genesis is a pivot: a temporary control plane installs the registry that makes it permanent](0067-genesis-is-a-pivot.md)
### What runs on them, and how it gets there
@@ -121,6 +123,9 @@ python3 00-META/checks/index.py fail if stale
- **0050** — [Model access is vendor-agnostic, and a vendor is an adapter](0050-model-access-is-vendor-agnostic.md)
- **0051** — [Shared data is the operator's, and a module is granted access to it](0051-shared-data-is-the-operators.md)
- **0052** — [An init step is a container run once to completion, gating what follows](0052-a-step-that-runs-once-before-a-container.md)
- **0053** — [A scheduled step is a container run on a recurring schedule](0053-a-step-that-runs-on-a-schedule.md)
- **0054** — [Model usage is a vendor-neutral record, produced by the adapter, at two grains](0054-model-usage-is-recorded-at-two-grains.md)
- **0055** — [Model access is answered by a licence, or by a node that hosts the model](0055-model-access-is-answered-by-a-licence-or-a-node.md)
### How it is built
+18 -2
View File
@@ -2,7 +2,7 @@
layer: to-be
status: in-progress
code: [mesh-lab]
updated: 2026-08-31
updated: 2026-09-11
decisions:
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0010-delivery.md
@@ -48,6 +48,7 @@ present, but that the machine can actually do the work:
| **the pool the lab will use is on that driver** | a pool exists, and is the slow kind — the failure that has no symptom |
| hardware virtualisation is present | machines will be emulated and unusably slow |
| an image can be fetched or is cached | the first raise will fail late instead of early |
| **a machine on an uplink reaches something real**, by fetching it — not by reading a route or a policy | forwarding is being dropped by something else on the workstation; names still resolve, and every pull hangs |
Each check states **why it matters**, in the terms of what it costs. *"The pool uses the `dir`
driver"* means nothing to someone who does not already know it means seventy-six times slower
@@ -101,7 +102,22 @@ In order, on a machine with nothing:
existed. Observed directly: a shell whose process tree predated the grant could not reach
the daemon while a fresh lookup showed the membership present. The bootstrap has to say so,
or the first thing a person meets is a permission error that looks like a broken install.
5. **Verification**, as above, before anything is raised.
5. **A path out to the internet for the machines that need one.** A scenario's own segments are
isolated on purpose, but a machine that fetches what it starts from is attached to an uplink
the daemon translates. That is the daemon's business and it does it — and then a **container
runtime on the same workstation sets the kernel's forwarding policy to drop**, which the
daemon's own accept rules do not override, because both are consulted and a drop anywhere is
the answer.
The result is the sharpest instance of *available is not adequate* yet: the machines get
addresses, they resolve names — the daemon's resolver is on the bridge, so that half works —
and every packet to anything real is discarded. Nothing is misconfigured, nothing logs, and
the failure presents as *every image pull hangs*. A workstation that runs containers is the
ordinary case, so this is a prerequisite rather than a quirk: forwarding must be permitted for
the lab's own bridges, and it must be **verified by reaching something**, never by reading a
setting.
6. **Verification**, as above, before anything is raised.
## Open
+71 -1
View File
@@ -7,7 +7,7 @@ code:
- mesh-control internal/identity/authority.go
- mesh-host internal/identity/serving.go
- mesh-host internal/apply (the service that reflects a rule set)
updated: 2026-08-31
updated: 2026-09-09
decisions:
- 02-DECISIONS/0005-the-node-host.md
- 02-DECISIONS/0004-a-node-and-how-it-joins.md
@@ -16,6 +16,7 @@ decisions:
- 02-DECISIONS/0004-a-node-and-how-it-joins.md
- 02-DECISIONS/0007-connectivity.md
- 02-DECISIONS/0006-the-substrate-and-the-control-plane.md
- 02-DECISIONS/0066-public-routing-is-name-agnostic.md
---
# Connectivity
@@ -343,6 +344,27 @@ that module and nothing else.
the argument for the table in ADR 0009 being a table: the pattern is only obvious once seen, and
the cost of not seeing it is inventing a mechanism that already exists.
### And the public names a proxy serves must resolve in the mesh too
*2026-09-09, found by an internal certificate authority that could not issue.* The mesh writes every
`<node>.internal` name into every declared container and treats the public names a proxy serves as a
separate matter — *what routes it once it arrives is a proxy's, and stays separate*, above. That
holds for a client dialling by internal name. It does not hold for anything **inside** the mesh that
must reach a public name, and the first such thing to appear was the internal issuer of §5.
**An issuer validates by connecting to the name it is certifying.** Asked for a certificate for a
routed public name, the internal authority accepted the order, offered a challenge, and then could
not connect: nothing in the mesh resolved that name, so the challenge had no target. A name the mesh
can reach from the outside but cannot resolve from the inside is a name it cannot certify with an
authority of its own.
**So a granted route is published into internal resolution as well** — the routed name to the node
that serves it, mesh-wide, by the same mechanism that writes the node names. It is *given by the
mesh, not chosen by a module*, for the same reason the node names are: a module listing the routes
would go stale the day one changes. The mesh propagates the names it was told to serve and still
knows nothing about what they mean
([ADR 0066](../../02-DECISIONS/0066-public-routing-is-name-agnostic.md)).
## 3 — Exposure
Settled by [ADR 0007](../../02-DECISIONS/0007-connectivity.md); summarised here because
@@ -389,6 +411,27 @@ the mesh to tell them apart.
returning the workload's own answer, then by unassigning the module and requiring the same request
to stop working.*
### The name is a label, not a domain
*2026-09-09, from running the whole mesh in the lab.* A route contribution carried its public name
in full — the forge as `git.example.tld`, spelt out in the module. Pointing the same catalogue at a
different domain — a lab standing in for production, or a second operator's mesh — meant rewriting
that name on every routed module. The mesh was holding a **map of names to services**, which is the
one thing it must not: a public name is two facts owned by two different places, and neither is the
module's manifest.
**A module contributes a label; the node contributes its public domain; the mesh composes.** The
operator chooses where the forge lives — `git`, or `code` — and that is the module's to say. The
domain is the node's, set once. The mesh joins them and grants `<label>.<public-domain>`,
interpreting neither half. Moving a mesh to another domain is one node setting, not an edit per
module.
**Today the name is still a literal, and that is the gap.** There is no interpolation of a node's
domain into a module's label, so the composition is done by a per-node override — which reproduces
exactly the per-module cost it is meant to remove. The design is the composition; the override is a
stopgap until the manifest layer can carry a label and a domain separately
([ADR 0066](../../02-DECISIONS/0066-public-routing-is-name-agnostic.md)).
## 4 — Filtering
**Derived from what is assigned here, and from the overlay's shape** — a node's open ports are a
@@ -539,6 +582,24 @@ generated, the other verifies against the mesh's authority and nothing else. Eve
passed while the server could not start — the key was present, the certificate was valid, and
nothing read either the way a server would.*
### An internal issuer, pointed at and trusted
*2026-09-09, from wiring one to the proxy in the lab.* §5 above asks for two things this build
leaned on: the issuer must be configurable, and the lab runs its own ACME authority rather than
collapsing the split. Wiring the proxy to that authority is the whole of it — the proxy is told
**which** issuer to use and given that issuer's **root** to trust, and every other step of issuance
is unchanged. **The same code path certifies against an internal authority as against a public one;
only the issuer differs.** That is what makes trusted certificates possible for a mesh whose names
the public internet cannot resolve.
**And it does not work until the routed name resolves inside the mesh** — the §2 finding above,
arriving here because this is what needed it. The authority's challenge reaches the routed name only
once that name is in internal resolution; a public authority is handed that dependency by public
DNS, and an internal one has to be handed it by the mesh. *Checked by a handshake to a routed name
that verifies against the internal root and nothing else — which cannot succeed unless the issuer
first reached the name to certify it*
([ADR 0066](../../02-DECISIONS/0066-public-routing-is-name-agnostic.md)).
## What this removes
The list is worth having in one place, because it is most of the argument:
@@ -572,3 +633,12 @@ The list is worth having in one place, because it is most of the argument:
it expressible; nothing here says the overlay or the resolver handle it.
- **Reporting declared-versus-observed.** ADR 0007 makes the disagreement detectable and does not
say who looks or what they are told.
- **Composing a route name from a label and a node's domain.**
[ADR 0066](../../02-DECISIONS/0066-public-routing-is-name-agnostic.md) makes the public name the
operator's to move between meshes, but the manifest layer still stores it as a literal — so today
the composition is a per-node override rather than the design. The interpolation that would let a
module carry a label and a node carry the domain, and the mesh join them, does not yet exist.
- **Publishing route names into internal resolution.** The same ADR requires a granted route to be
resolvable inside the mesh, not only routable from outside it; the mechanism that writes
`<node>.internal` into containers does not yet also write the routed names, which is why an
internal issuer cannot currently validate one without a hand-placed entry.
@@ -0,0 +1,91 @@
---
status: resolved
opened: 2026-09-09
located-in: [mesh-control]
fixed-by: mesh-control — a same-node provider is announced at the port it is published on
amended-design:
---
# 038 — A provider is announced at a name its port is not bound to
## Symptom
A module that provides a `from: mesh` provision (observed with the database provider) is
announced to its consumers, by [issue 018](../018-a-provider-on-the-same-machine-was-never-announced/00-report.md)'s
fix, at the node's private-network name — the binding a consumer reads carries
`at: <node>.internal` and `serves.port: 5432`.
But the provider's container port is **published bound to loopback** (`127.0.0.1:<assigned>`),
not to the address `<node>.internal` resolves to. So every consumer dials the announced
`<node>.internal:5432`, which resolves to the node's private-network address, where **nothing is
listening** — the port is open only on `127.0.0.1`.
Observed on a node hosting the provider and several consumers:
- The consumer's binding file says `"at": "<node>.internal"`, `"serves": { "port": 5432 }`.
- Inside a consumer container, that name resolves to the node's private-network address.
- A connection test from the node: the private-network address on port 5432 is **CLOSED**; only
`127.0.0.1` (on the assigned host port) is OPEN.
- Consumers that touch the database only lazily serve a landing page and *look* healthy; consumers
that require the database at startup crash-loop — one with "acquisition timeout while waiting for
a new connection", another connecting and then timing out on its first query.
- The provider itself is healthy: a direct client on loopback answers instantly, few connections,
no locks.
## Why this matters
The announced address and the actual listener disagree, so the binding is a promise the mesh does
not keep. It is not one module's misconfiguration: it is the port-publishing step choosing a bind
address that does not match the `at` the resolver hands consumers, so it fails the same way for
**every** `from: mesh` provider with an off-node-reachable consumer — and, on a single node, for
same-node consumers too.
It hides well. The provider is up, the credential is correct, the database exists, a manual client
works — every part a person checks in isolation passes. Only a consumer that must use the provision
before it can serve anything reveals it, and it reveals it as a timeout, which reads as slowness or
load rather than "the address was never listening". A mesh that co-locates a provider with its
consumers (the ordinary small-mesh case) is exactly where it bites.
It also blocks anything that must *reach* a routed/served name from inside the mesh, not just
application traffic — see the internal-CA validation dependency noted in the connectivity design.
## Diagnosis
The symptom's first reading — "published on loopback" — was **partly a red herring**. Two things
were tangled:
1. **The real, current-code defect is a served-*port* mismatch, not a bind address.** A bare
`ports: ["5432"]` is assigned a host port and published as `"15432:5432"` — no bind IP, so on
**all interfaces**, reachable at the node's private-network address. But the port a consumer is
*told* is only re-derived from the assignment on the **cross-node** path. The **same-node** paths
(the resolver's `servedHere`, and the `here()` fallback) settle their served facts *while
resolving* — before the host port is assigned — so they carry the **declared** port (5432), not
the **assigned** one (15432). A co-located consumer is therefore announced
`<node>.internal:5432` while the provider is published on `<node>.internal:15432`, and dials a
port nothing listens on. Cross-node consumers were always fine, which is why it read as "the
small-mesh case."
2. **The `127.0.0.1:15432` seen in the running lab was a stale build.** Current code's publish step
binds all interfaces; the running instance was raised from a mesh-control predating the ADR 0038
publish rewrite. The substrate's own store *is* deliberately `127.0.0.1:5432` (a private store
must not be exposed) — correct, and not this bug.
## Fixed by
`mesh-control` branch `fix/same-node-provider-announced-port` (`c147a26`): after the host port is
assigned, same-node needs (and the `here()` fallback) are redirected through the same
provision→module→assigned-port lookup the cross-node path already uses, so a co-located consumer is
announced the port that is actually published. Idempotent (keyed by the declared port). Regression
test `TestASameNodeProviderIsAnnouncedAtThePortItIsPublishedOn` asserts the announced port equals
the published host port for a co-located provider/consumer — the next assertion after 018's (which
only checked a binding file exists); verified failing without the change.
*Not yet merged, and the running lab is additionally stale — proving it end-to-end there needs
mesh-control rebuilt and the affected consumer containers recreated.*
## Noted, not taken
Binding the assigned port to the node's private-network address specifically (rather than all
interfaces) would be defence-in-depth and would make the publish address match `at` by construction
— but it is a larger behavioural change entangled with the unenforced firewall scope
([003](../003-firewall-scope-is-read-by-no-code/00-report.md)), so it is left as an option.
@@ -0,0 +1,78 @@
---
status: resolved
opened: 2026-09-10
located-in: [mesh-catalog]
fixed-by: mesh-catalog — the operator's own images are pinned by digest
amended-design:
---
# 039 — The lab's registry was pinning what the catalogue left unpinned
## Symptom
Nine container images across seven modules name their image by **tag** — the shape
`<registry>/<org>/<name>:latest` — rather than by digest. They are the operator's own application
images, the ones built from their own source.
[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md) requires a digest, and
the host refuses a tag by name: *"image %q is not pinned. Write it as name@sha256:… — a tag moves,
and a bundle that pinned a tag would not be pinned."*
**Every bed passed anyway, for as long as the lab has existed.** The lab raised a registry of its
own, pushed every image into it, and rewrote every reference in every manifest to the digest **that
registry had just assigned**. So a manifest naming a tag arrived at a machine naming a digest. The
rewriting was doing the pinning.
It surfaced only when the lab's registry was deleted — the modules now carry their tags all the way
to the machine, and fail there, which is the correct behaviour finally being reachable.
## Why this matters
**A rule enforced by scenery is not enforced.** The host's refusal is right and has never once
fired in a bed, because nothing unpinned could reach it. The check exists, the tests are green, and
the property they appear to defend was being supplied by the test harness — which is the same shape
as [003](../003-firewall-scope-is-read-by-no-code/00-report.md), one layer further out: there the
rule was read by no code, here it is read by code that never saw a violation.
**And these are the worst images for it to be true of.** A tag republished on every build is the
one reference that genuinely moves. A machine reconciling against an unchanged declaration can
change what it runs, with nothing in the declaration or the mesh's records saying anything did —
which is precisely the failure pinning exists to prevent, aimed at the images that change most
often.
**The general form is the part worth keeping.** Any invariant the lab happens to satisfy
incidentally is an invariant no bed tests. The harness was not merely serving images; it was
quietly supplying a property of the system under test, and nothing said so.
## Fixed, and what the fix does not settle
Each of the nine references now names the digest its tag resolved to, read from the registry that
serves them. Every manifest in the catalogue is pinned; the host's refusal has nothing left to
catch, and the check that would have caught this — *no manifest names a tag* — now passes on
content rather than on a harness's rewriting.
**It is a stopgap and should be read as one.** A digest written into a repository is wrong the
moment anybody rebuilds, which is exactly the argument
[`12-a-module-repository`](../../03-DESIGN/01-to-be/12-a-module-repository.md) makes for two
documents — the repository naming *artifacts*, the mesh holding *digests*. Until something builds
and publishes, a digest that goes stale is still better than a tag that moves without telling
anyone: stale fails loudly at the pull, and a moved tag changes what a machine runs while every
record says nothing changed.
So the open questions below stand. What closed is the immediate fault; what remains is the reason
it was possible.
## Open questions
- What should a manifest name for an image the operator builds themselves? A digest changes on
every build, so a manifest carrying one is wrong the moment anybody commits — which is the
argument [`12-a-module-repository`](../../03-DESIGN/01-to-be/12-a-module-repository.md) already
makes for *two* documents, the repository's naming artifacts and the mesh's naming digests. Is
this simply the build pipeline's absence showing, rather than seven mistakes?
- Until that pipeline exists, what names these images — and is a tag with a loud warning better or
worse than a digest that is stale by construction?
- How is *"nothing reaches a machine unpinned"* checked anywhere other than on the machine that
refuses it? A check that only ever runs at the last possible moment, on the one path a harness
was rewriting, is a check nobody can see failing.
- Which other invariants is the lab supplying rather than testing? This one was found by deleting
the thing that supplied it. That is not a repeatable technique.
@@ -0,0 +1,58 @@
---
status: open
opened: 2026-09-10
located-in: []
fixed-by:
amended-design:
---
# 040 — The only description of how a mesh is stood up is a test
## Symptom
There is no installer. The one complete, executable account of how a mesh comes into existence is
an **integration test** in the lab repository: it builds the images, applies the substrate, enrols
every node, places the overlay, seeds the operator's secrets, registers modules, assigns them,
pushes, and waits for convergence.
Nothing else does this. `packaging/` installs and supervises the **host agent** on one machine.
[`04-lab-installation`](../../03-DESIGN/01-to-be/04-lab-installation.md) covers the **lab's own**
prerequisites — a virtualisation daemon, storage, a pool. Neither installs a mesh, and no design
section describes doing so.
## Why this matters
**A test fixture is allowed to invent what it needs, and this one did.** It raised an image
registry that exists in no production, stocked it from a workstation, and rewrote every image
reference in every manifest to point at it. That single convenience concealed at least two separate
faults for as long as the lab has existed:
- modules naming a tag where a digest is required, never caught because the harness was assigning
the digests ([039](../039-the-lab-registry-was-silently-pinning-unpinned-modules/00-report.md))
- the bootstrap itself having **no path at all** on a machine that is not the lab: the mesh's own
control-plane image exists in no registry, so nothing could name it, and the fixture's registry
was the only reason that never surfaced
**So the bed being green said nothing about whether a mesh could be installed.** Both are the same
fault: the procedure and the reality drift, and the drift is invisible precisely because the thing
that would notice is the thing doing the inventing.
**And a person cannot run it.** The procedure is expressed as assertions in a test runner, in a
repository whose purpose is to raise disposable virtual machines. Somebody standing up a real first
node has no artefact to use — which is why standing one up has been "a person driving the CLI",
the loop [ADR 0010](../../02-DECISIONS/0010-delivery.md) removed everywhere else and left here.
## Open questions
- Where should the procedure live so that **the lab uses it rather than reimplementing it**? The
lab must remain able to raise machines and inject faults; what it should not be able to do is
describe installation differently from the thing an operator runs.
- What is the smallest arrangement that makes drift *impossible* rather than merely discouraged —
the bed invoking the installer, or something generating one from the other? A second description
that is merely kept in step is the arrangement that just failed.
- Which parts of what the test does are genuinely lab-only — documentation addresses, names that
resolve nowhere, injected faults — and which are the mesh's own and belong in an installer? The
distinction has never been drawn, and everything landed on the lab side of it by default.
- How is *"the installer is what installed this"* checked afterwards, on a mesh that is already
running? A mesh cannot say how it was raised, so nothing can currently contradict a claim that it
was raised the supported way.
@@ -0,0 +1,58 @@
---
status: open
opened: 2026-09-10
located-in: []
fixed-by:
amended-design:
---
# 041 — A credential the mesh took care to seal ends up in the process environment
## Symptom
The mesh generates a module's own secret, **seals it to the machine and discards the plaintext** —
it cannot read the value back even if asked. The host unseals it into a file the module names, at
`0600`.
Then the module hands it to its container as an environment variable, and the runtime puts it
where anything on that machine that can talk to the runtime can read it: `docker inspect` prints
it, and `/proc/<pid>/environ` holds it for the life of the process.
Observed while making the control plane an ordinary module. Its **store** connections were moved to
files, read by a `…_FILE` variable naming the path. Its **broker** credentials have no such variable,
so they are still delivered through an env-file — which the runtime turns into exactly the
environment above. Same credential handling, same machine, two different exposures, decided by
whether the program that reads it happens to accept a path.
## Why this matters
**The care taken elsewhere is what makes this stand out.** Sealing to a machine and discarding the
plaintext is expensive and deliberate: it exists so that a credential is readable only where it is
used. Handing that same value to the runtime as an environment variable gives it back to anything
that can run `inspect` — and `inspect` is a routine operation. It lands in support output, in
captured logs, in a screenshot of a terminal, and in any tooling that dumps container state.
**It is not a module's mistake.** Nothing in the manifest format is being misused: placing a secret
into an env-file with `${secret:…}` is a supported shape and other modules use it. So each module is
correct on its own, and the property — *a sealed credential is not readable by everything on the
machine* — holds or fails per variable, by accident of what each program accepts.
**And the two halves now disagree inside one module.** The control plane reads its store connection
from a file and its broker credential from the environment. A reader cannot tell from the manifest
which secrets are protected from `inspect` and which are not, because the manifest looks the same
either way.
## Open questions
- Should every program the mesh runs accept a path for anything secret — a `…_FILE` twin as a
convention rather than a thing each program decides? That is a small change in several programs
and a large one in what the manifest can promise.
- Should the manifest layer **refuse** `${secret:…}` inside a container's `env`, or inside an
`env-file`, once a path-shaped alternative exists? A rule nothing enforces is the shape this
repository keeps finding.
- Is there a case where the environment is genuinely the only channel — a program that cannot be
changed and reads no file? If so, what should the mesh say about that module, out loud, rather
than treating it as equivalent?
- What is the actual reach of the exposure on a node — which identities can talk to the container
runtime, and is that set smaller than "anything running as the operator"? The answer decides
whether this is a hardening item or something sharper.
@@ -0,0 +1,64 @@
---
status: open
opened: 2026-09-11
located-in: []
fixed-by:
amended-design:
---
# 042 — Nothing gives a node an account for a registry
## Symptom
A module names an image in a registry that requires authentication. The machine cannot fetch it:
```
pull access denied … authorization failed: no basic auth credentials
```
Nothing in the mesh delivers a registry credential to a node. There is no provision for it, no
field in a manifest for it, and no step in enrolment that establishes one.
Observed when the lab's own image registry was deleted and machines were given a real path to the
internet for the first time. Public images fetched normally; the operator's own — in their own
registry — could not be fetched at all.
## Why this matters
**Delivery ends at a node pulling an image, and this is the last step of it.** The whole design
from source to artifact to machine assumes the final fetch succeeds. It succeeded until now for
one reason: everything was either public, or came from a registry the harness ran with no
authentication at all.
**And it applies to the mesh's own store, not only to somebody else's.** The registry the mesh runs
is reached over the private network and today asks for nothing. That is a decision — *anything that
can reach it can read every artifact the mesh holds, and push to it if pushing is open* — but it
has never been written down as one. It reads as an absence, which is the kind of thing that stays
true by accident until it is exploited.
**The mesh already knows how to do this.** A provision grants a consumer a credential, minted by
the mesh, sealed to the machine, with the plaintext discarded. An account on a registry is exactly
that shape: the artifact store is already a provision, and a node that must pull from it is already
a consumer of it. What is missing is not a mechanism but the recognition that pulling is a use of
the store, not a thing that happens beneath it.
**It is also a bootstrap question**, which is where it will bite first: the machine that installs a
mesh pulls the control plane's image before the mesh exists to grant anything
([ADR 0067](../../02-DECISIONS/0067-genesis-is-a-pivot.md)). Whatever the answer is, it has to
survive a moment when there is nobody to ask.
## Open questions
- **Should the ability to pull be granted like anything else** — the store mints a per-node account,
seals it to the machine, and the host writes whatever the runtime reads? That would make a node's
reach into the artifact store follow the same rules as its reach into a database.
- **What should the mesh's own registry require?** If the answer is "nothing, inside the private
network", that is defensible and must be *recorded* as the decision it is, with what it assumes
about who is on that network.
- **What authenticates a push?** Reading and writing are not the same grant, and a builder that may
publish is a much stronger thing than a node that may fetch.
- **What does an operator's own external registry do here?** It is not the mesh's to grant accounts
on. Is the credential operator-supplied, like any other secret the mesh cannot invent — and if so,
which module owns it, given that every module pulling from there needs it?
- **How does any of this work before the mesh exists?** The bootstrap fetches images with no mesh to
grant a credential, so the first pull is either public, anonymous, or carried.