Merge branch 'issue/021-provider-port-published-on-loopback' into design/bootstrap-is-a-pivot

# Conflicts:
#	03-DESIGN/01-to-be/04-lab-installation.md
This commit is contained in:
2026-09-11 00:45:59 +02:00
188 changed files with 14552 additions and 2968 deletions
@@ -1,8 +1,8 @@
---
status: open
status: resolved
opened: 2026-08-22
located-in: []
fixed-by:
located-in: [mesh-host]
fixed-by: mesh-host — a package is read back from the package database after installing
amended-design:
---
@@ -31,14 +31,14 @@ exists to catch — a step that failed, reported success, and left the next step
state that was never produced.
It is also a direct violation of a decision already taken and recorded:
[ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) says a step that fails must fail the
[ADR 0010](../../02-DECISIONS/0010-delivery.md) says a step that fails must fail the
job. That record notes the rule is applied instance by instance and enforced by no mechanism.
This is an instance where it was never applied.
## Evidence
- Observed 2026-08-22 while declaring the virtualisation package required by
[ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md).
[ADR 0016](../../02-DECISIONS/0016-the-lab.md).
- A fix is written and open as a pull request, unmerged since 2026-08-20.
## Open questions
@@ -47,3 +47,18 @@ This is an instance where it was never applied.
- Is this specific to package installation, or does the surrounding stage swallow every
non-zero exit?
- The fix has been open for two days. What is the review path for a change of this class?
## How it is answered
*2026-08-31.* **The host reads the package database back after installing**, and refuses when it
does not have the package:
> `<name> was installed without error and the package database does not have it`
That is the general rule this issue is one instance of, and the host applies it to everything it
does: a command exiting zero says a transaction was *accepted*, not that the machine changed. The
same read-back is why a container that starts and immediately dies fails an apply, and why a
service asked to run is checked rather than assumed.
HAL keeps the fault until its provisioning is switched off. Fixing it there would mean
implementing the read-back twice, in the system being replaced.
@@ -1,8 +1,8 @@
---
status: open
status: resolved
opened: 2026-08-22
located-in: []
fixed-by:
located-in: [mesh-host]
fixed-by: mesh-host — a stale package index is named rather than reported as a failed install
amended-design:
---
@@ -38,3 +38,29 @@ the job is green.
versions — or is an index sync part of the install step?
- A partial sync is unsafe on the platform in use; a full upgrade is the only sanctioned fix.
Does that make index freshness a scheduled node concern rather than a pipeline one?
## How it is answered
*2026-08-31. It was present in the replacement too, which is why this is a fix rather than a note
saying the new mesh does not have it.*
**The failure is named.** A machine asking for a version the mirrors have replaced now says so, and
says what fixes it — a full upgrade of the machine.
**It is deliberately not fixed by synchronising.** `pacman -Sy <pkg>` installs a package built
against libraries the machine does not have: a partial upgrade, unsupported on this distribution,
which surfaces much later as something apparently unrelated. That is a decision about the whole
machine, and a host that made it silently while applying one resource would be taking a large
decision in a small place.
So the host distinguishes the two cases and leaves the decision where it belongs. **A declaration
that is wrong and a machine that is out of date fail identically otherwise, and they are fixed in
completely different places.**
**And the package manager's own words were being thrown away** — the output was read into `_`, so
the 404s that name the cause never reached anybody. Whatever it said is now part of the failure,
which is the rule everywhere else here and was not being followed in the one place where the reason
exists only in the output.
*Checked by a stale-index failure being named as one, an ordinary missing package not being, a
single mirror timing out not being, and a successful install still saying nothing.*
@@ -1,9 +1,9 @@
---
status: open
status: resolved
opened: 2026-08-22
located-in: []
fixed-by:
amended-design:
located-in: [mesh-control]
fixed-by: mesh-control — a machine's filtering is computed from what it was assigned
amended-design: 03-DESIGN/01-to-be/08-connectivity.md
---
# 003 — A firewall rule's `scope:` is read by no code
@@ -38,3 +38,29 @@ any check.
- Were the five declarations intended to restrict something that is currently open? Each needs
checking against what the node actually exposes — the declaration cannot be trusted either
way.
## How it is answered
*2026-08-31.* Both halves, in Novox Mesh. HAL keeps the fault until its provisioning is switched
off, which is what this issue is now waiting on rather than a fix of its own — patching `scope:`
into something that works would mean implementing it twice, in the system being replaced.
**The unknown key.** A manifest is parsed strictly: an unknown key is refused with the key named,
the discipline the host's declaration parser has always had. `scope:` would not survive being
written today, and neither would a misspelling of anything else. This is the general fix — the
issue's own observation was that *any* invented key behaved this way, and that the one instance
was found by reading rather than by any check.
**The rule that restricts nothing.** `scope:` is not reimplemented. A module says what it listens
on and **who may reach it**, and saying from where is required rather than defaulted: a rule with
no source is open, and must say so rather than appear to restrict something. A machine's whole
rule set is then derived from every module assigned to it — so there is no second list to keep in
step, which is the condition that let the first one drift out of use unnoticed.
**And it is enforced, which is the part that makes this different from before.** The mesh renders
the rule set; a service on the node is declared to reflect that file, so replacing it restarts
what loads it. Proven in the lab against two real ports on a real machine: the declared one
answers from another machine, the undeclared one does not, and removing the module that wanted the
port closes it with nobody editing a rule.
The design is [`03-DESIGN/01-to-be/08-connectivity.md`](../../03-DESIGN/01-to-be/08-connectivity.md) §4.
@@ -1,8 +1,8 @@
---
status: open
status: resolved
opened: 2026-08-22
located-in: []
fixed-by:
located-in: [hal]
fixed-by: hal — ACME_CA_SERVER selects the authority, and defaults to staging
amended-design:
---
@@ -21,7 +21,7 @@ recoverable by retrying — it removes the ability to issue a certificate anyone
The consequence lands hardest on exactly the work most likely to iterate: standing up a new
node, changing how names resolve, or testing the lab's certificate authority split
([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)).
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
## Evidence
@@ -35,3 +35,41 @@ node, changing how names resolve, or testing the lab's certificate authority spl
- The lab issues its own certificates and so does not consume public quota at all. Does that
make this a problem only for experiments run outside the lab, and therefore an argument for
running them inside it?
## Resolution
*2026-08-31.* **The authority is now selectable, and the default is staging.**
Confirmed first: the resolver declared no `caServer` at all, so the client fell to its built-in
default of the public authority's production endpoint. There was no setting to change — not a
setting set wrongly.
`ACME_CA_SERVER` now names the authority, and **defaults to staging**. That answers the first
open question in the affirmative: it is a node property, and the node that serves real traffic is
the one that states so.
**Why staging is the default rather than the safe-looking alternative.** Defaulting to production
and documenting the override would leave the safe path depending on somebody remembering to opt
out of it — on exactly the work most likely to iterate. That is the same fault as
[005](../005-pipeline-test-harness-unbuildable/00-report.md): *remembering is not a mechanism.*
A staging certificate is trusted by no browser, so the mistake announces itself in the first
request rather than a fortnight later when the quota is gone. **The failure that is loud and
immediate is the cheaper one**, and quota exhaustion is neither.
**The second open question is answered too, and it is not the whole answer.** The lab issues its
own certificates and consumes no public quota, so experiments belong there. But "run it in the
lab" is advice, and the nodes this issue is about are the ones outside it — the default is what
protects those.
### The rollout is ordered, and the order is the dangerous part
Both public-serving nodes were checked: neither set the variable. Applying the change without
pinning them first would re-issue their public certificates from an untrusted authority and break
TLS for every hosted service — **this issue's own failure mode, arriving through its fix.** Pin
first, then merge. The commit carries the exact commands.
## Deliberately not done
**Nothing was changed on a running node.** Pinning the public nodes and regenerating their
environment restarts the reverse proxy that fronts every hosted service, and that is an operator's
decision rather than a fix's.
@@ -1,9 +1,9 @@
---
status: open
status: resolved
opened: 2026-08-22
located-in: [hal]
fixed-by:
amended-design:
located-in: [hal, mesh-lab]
fixed-by: mesh-lab — a run leaves a receipt, and the receipt says what it covered
amended-design: 03-DESIGN/01-to-be/01-end-to-end-testing.md
---
# 005 — The end-to-end pipeline harness has not built since the workspace was removed
@@ -29,7 +29,7 @@ coverage was assumed, not checked.
## Evidence
- The workspace was removed by pull request #240 on 2026-06-04
([ADR 0007](../../02-DECISIONS/0007-no-npm-workspace.md)).
([ADR 0014](../../02-DECISIONS/0014-no-npm-workspace.md)).
- The harness has not built since that date.
- Recorded in the knowledge base as a standing entry, not as a fixed incident.
@@ -45,3 +45,52 @@ unmentioned: until the lab exists, this is the coverage the pipeline is presumed
- Repair, or retire in favour of the lab? Leaving it in the repository unbuilt is the one
option that keeps the false impression of coverage.
- Was anything relying on it, or had it already stopped running before the workspace removal?
## Resolution
*2026-08-31.* **Retired in favour of the lab, and the reason it went unnoticed was fixed
separately from the harness itself.**
The old harness is not repaired. What replaced it is the end-to-end suite on a lab mesh, which
raises real machines and proves the pipeline against them. That answers the first open question.
The second finding is the one worth keeping. *Nothing runs it, and nothing reports that nothing
runs it* is not a fact about that harness — it is a fact about **any** suite too expensive to run
on every push, and the lab suite is exactly that: it needs a machine with a hypervisor, so it runs
when somebody remembers. **Remembering is not a mechanism**, and the replacement inherited the
fault it was replacing.
So three things now hold, each checked by a test that was confirmed to fail without it:
- **A run leaves a receipt** — when it ran, what passed, and the commit each repository was at.
Kept outside version control, because the question is *has this machine run it*, and a receipt in
git would be a claim about everybody's machine made by whoever committed last.
- **The receipt can be judged, and says why it does not count.** Old, failed, taken against code
the repositories have since moved past, or a run that never raised a machine — each reads
differently, and only the last of those is new. **A receipt that says nothing about something is
not a receipt that clears it.**
- **The artifacts are rebuilt by the run, not beside it.** The suite consumes three artifacts from
two repositories. They were rebuilt by hand, from memory, and a rename that needed two of them
got one — leaving a binary eleven hours old refusing a field the mesh had just renamed, found by
a full run. That step now lives in the repository rather than in a terminal history.
### What this issue taught twice
**The fix reintroduced the fault, in miniature, and the second time was caught by running it.**
The suite takes paths, so it can be pointed at one quick unit file — and the receipt from that run
was, at first, indistinguishable from a receipt for the real thing. A green record standing for a
run that raised no machines is this issue's own symptom, rebuilt inside its remedy. The receipt now
records what it ran, and a run that did not include the end-to-end file is not coverage.
Separately, the code that decides *no receipt rather than a guessed one* — the rule that keeps the
record meaning something — was first written where no test could reach it. Writing "0 failed"
because nothing said otherwise is how a green record comes to mean nothing.
And the counting itself **passed every test while reading nothing**: the test runner colours its
summary even into a pipe, so the anchored pattern never matched, and the fixtures it was checked
against were output that had been imagined rather than captured. **A fixture that agrees with the
mistake proves the mistake.** It is now checked against the runner's real bytes.
Each of these was found by running the thing, not by reading it — which is the same argument this
issue makes about the pipeline.
@@ -1,9 +1,9 @@
---
status: open
status: located
opened: 2026-08-23
located-in: [hal, hq]
fixed-by:
amended-design:
amended-design: 02-DECISIONS/0025-the-design-record-is-read-not-copied.md
---
# 006 — This repository is not indexed into the knowledge base, and the claim that it is holds up a decision
@@ -59,7 +59,7 @@ checked it — including in the same commit that wrote the rule.
## Proposed direction — Nox is the search
*Added 2026-08-23.* Rather than syncing these documents into the knowledge base, **Nox
([ADR 0027](../../02-DECISIONS/0027-the-product-is-novox-mesh.md)) works from within this
([ADR 0019](../../02-DECISIONS/0019-how-this-repository-works.md)) works from within this
repository and holds its knowledge directly.** Retrieval becomes an agent reading the source,
not a copy living in a second store.
@@ -88,3 +88,70 @@ consults Nox — the answer is yes and the original promise holds.
That is a design question for Nox, not a defect in this repository, and it should be settled
before ADR 0019 is treated as answered.
## Where this stands
*2026-08-31. Re-checked, and deliberately not closed.*
**The indexing still does not exist.** Two searches today, against both the symptom-indexed
memory and the structured archive, using a decision record's full title and a distinctive phrase
from a design document: no results, no partial match, no stale copy. The symptom in this report
is unchanged.
**But the part that made it an issue is gone.** This report's argument was that the claim was
*load-bearing* — that a decision rested on a mechanism nobody had checked. It no longer rests on
it. The README now names the gap in the place the claim used to sit, and says it is left standing
rather than quietly reworded. The decision record that separates this repository does not invoke
indexing at all; its reasoning is cadence, reviewers, and scope, none of which depend on it.
So what remains is not a false claim. It is an unbuilt capability and an open design question,
and those are different things.
### What was done
**A signpost, in the knowledge base, pointing here** — what lives in this repository, which
folders hold what, and when to come looking rather than search there. Explicitly a pointer and
not a copy: a derived copy drifts, and the enforced copy wins while the reasoned one quietly
stops being true.
**It was tested, and it half works.** A search for *design records, decisions, repository* returns
it. A search phrased the way somebody would actually ask — *why is the mesh built this way* —
returns nothing, because the store matches terms rather than meaning.
That is this report's own distinction, confirmed by measurement rather than argued: **a signpost
is reachable, it is not surfacing.** Someone who suspects the answer exists will now find it.
Someone debugging an error, with no reason to think this repository knows anything about their
symptom, still will not.
### Why it stays open
The question this report narrows to is unchanged and unanswered:
> When a symptom is searched and the answer happens to live in a design document or a decision
> record here, does the searcher find it without already suspecting it exists?
Today: **no.** Closing this means choosing between a one-way sync into the knowledge base and an
agent that reads this repository and contributes to a symptom search — and that is a decision
about how the knowledge system works, not a defect to be fixed quietly.
**Marking it resolved while the indexing does not exist would be the failure this repository was
created to name**, one folder away from where it names it.
## The direction is decided
*2026-08-31.* **The agent reads this repository; nothing is copied.** Recorded as
[ADR 0025](../../02-DECISIONS/0025-the-design-record-is-read-not-copied.md), which also amends
what [ADR 0019](../../02-DECISIONS/0019-how-this-repository-works.md) promised: these documents
will not be *indexed*, they will be *read*, and the search consults the agent so its answers
appear beside ordinary results.
A sync was the option that works with what exists today, and it was rejected on the one ground
this repository can least afford: it makes a second copy, and *the copy that is searched quietly
stops matching the copy that is edited*.
**So the open question above is answered, and this report stays open on the build.** What closes
it is the check ADR 0025 names — search the mesh's memory for a phrase that appears only in a
design document here, and get it back. That check fails today by design.
**What stands until then** is the signpost, and the honest description of it: reachable, not
surfacing.
@@ -44,7 +44,7 @@ The distance between the two is the same one the delivery layer already has a na
## Why it matters now
This is the first requirement of the lab
([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)), which is
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)), which is
phase 0 of the entire migration. The first capability the new work depends on is present,
declared, and unusable — and would have stayed unusable silently.
@@ -0,0 +1,102 @@
---
status: resolved
opened: 2026-08-28
located-in: [hal]
fixed-by: hal — the script says it is manual, because it is
amended-design:
---
# 008 — The documented automatic node rescue does not exist
## Symptom
The mesh's documentation describes an automatic node rescue: a node that fails is recovered
without anybody intervening. **Nothing implements it.**
Found incidentally while investigating supervision
([research 003](../../01-RESEARCH/003-service-supervision/00-overview.md)), which counted what
actually supervises what:
- **no unit declares `OnFailure=`**, so nothing runs when a unit gives up;
- **nothing calls the rescue script on a timer**, so it runs only when a person runs it.
The script exists. The thing that would invoke it does not.
## Why this is worse than having no rescue
A rescue nobody wrote is a gap somebody can see. A rescue that is *documented* and absent is a
gap nobody looks for, because the documentation says it is covered — and it is read exactly when
a node has failed and somebody is deciding whether to intervene.
This is `how-we-build` §5 in its most expensive form: *an unenforced rule is indistinguishable
from a wrong one, and costs more, because people believe it.* Here the belief is that a failed
node recovers itself.
## Scope
**The as-is only.** The design being built has a different answer:
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) puts recovery
in a launcher that supervises the host, and that recovery is tested — 32 assertions, each
confirmed to fail when the behaviour is removed.
So this issue is about the mesh that runs **now**, and it has two possible resolutions rather
than one:
1. **Implement it** — an `OnFailure=` and a timer — if node rescue is wanted before the new host
reaches the fleet.
2. **Delete the documentation** — and say plainly that a failed node needs a person, which is
what is true today.
**Either is honest. Leaving it as it is, is not.** The choice turns on how far away the new host
is, which is a scheduling question rather than a technical one.
## What it would take to be sure
Read back rather than assumed
([ADR 0018](../../02-DECISIONS/0018-a-picture-is-read-from-what-runs.md)): list every unit on a
node and grep for `OnFailure=`; list every timer and check what each one calls. The finding above
came from reading the repository, and confirming it against a running node is the difference
between *no unit declares this* and *no unit in the source declares this*.
## Resolution
*2026-08-31.* **Resolved the second way: the documentation now says what is true.**
### Read back from running nodes, and the finding sharpened
This report was written from the repository. Checked against three running nodes, as the section
above asks — and one of its own claims was wrong in a way that matters:
| Claim | Verified |
|---|---|
| no unit declares `OnFailure=` | **true**, and in *both* scopes — of 604 system units, the only two are stock (`snapd`, `local-fs.target`); across ~30 user units, none |
| nothing calls the rescue script on a timer | **the timer exists.** `hal-health.timer` fires every 30 minutes |
| the rescue is not actually triggered | **true** — the *deployed* `hal-health.sh`, not only the repository's copy, contains no call to it |
**The trigger exists and does not do the thing the script says it does.** That is worse than the
absence this report described, because it survives a halfway check: somebody verifying "is there a
health timer?" finds one, and stops.
The precise falsehood was a single line in the rescue script — *triggered automatically by
`hal-health.timer` when the mesh service is failed.* It is now the opposite sentence, and says
what does trigger it.
**Two further claims were found and narrowed.** Documentation in two places called the mesh
*self-healing* — once for peers refreshing a tool cache, once for nodes auto-updating on a repo
change. Both behaviours are real and neither is healing. **A phrase that overstates by a category
is read in the crisis it describes**, and "the mesh is self-healing" is exactly the sentence that
stops somebody intervening.
### Why not the first way
Implementing it was the other honest option, and it was not taken. The replacement host already
supervises recovery and that recovery is tested; adding unattended download-and-restart behaviour
to the fleet it is replacing is a change with real failure modes of its own — a rescue loop that
thrashes is worse than a node that waits.
**This is the scheduling judgement this report said the choice turned on, and it is recorded
rather than assumed.** If rescue is wanted on the current fleet before the new host reaches it,
that is a decision to take deliberately: an `OnFailure=` and a call from the health script, both
small, neither free.
Until then the documentation is true, which is the part that was costing something.
@@ -0,0 +1,114 @@
---
status: fixed
opened: 2026-08-28
located-in: [mesh-lab, mesh-host]
fixed-by:
- "mesh-lab: a registry raised inside the scenario. Verified in a sealed machine — all four shapes applied with the image pinned by digest, idempotent, read back from the machine."
amended-design:
---
# 009 — A digest-pinned image cannot be placed in the lab, so `container` cannot be tested there
## Symptom
Two accepted decisions collide, and the collision makes one resource shape untestable.
- **[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)** pins images by
digest, and the host **refuses** an image reference that is not pinned:
```
resource "store": image "alpine:3.20" is not pinned. Write it as name@sha256:...
```
- **The lab cannot place a digest-pinned image.** A sealed scenario cannot reach a registry, so
the lab exports an image from the workstation and loads it in the machine — and that loses the
digest.
So a `container` resource is refused by the host if it names a tag, and unusable if it names a
digest. **There is no declaration the lab can currently raise that exercises the shape.**
## What was measured
Not inferred. `docker save alpine@sha256:d9e8…` produces an archive with **no repo tag**, because
a repo digest exists only for an image a registry served. Loading it says:
```
Loaded image ID: sha256:63f227… (not "Loaded image: alpine:3.20")
```
and `docker images` then lists nothing — the image is there but dangling. A container declaring
that digest therefore falls through to the registry:
```
Unable to find image 'alpine@sha256:d9e8…' locally
dial tcp: lookup registry-1.docker.io: no such host
```
which is correct behaviour on a machine with no route out.
## What is not affected
Everything else placed in the same sealed machine works, and was verified there:
| shape | |
|---|---|
| `package` | applied, idempotent |
| `service` incl. `boot: enabled` | applied, read back as `enabled` |
| `action` | ran, verified |
| `container` | **blocked by this issue** |
## Why it matters more than one shape
The container shape is the substrate. Every step of raising a mesh past the container runtime is
a container ([`07-the-substrate.md`](../../03-DESIGN/01-to-be/07-the-substrate.md)), so the
bootstrap cannot be tested end-to-end until this is resolved — which is the thing the lab exists
for.
## How it was fixed
*2026-08-29.* A registry inside the scenario, as below — and it turned out to be the shape the
resolution predicted rather than a compromise on it.
A scenario declares `images:` by tag. The lab stocks a registry **on the workstation**, where
there is a network, then raises one **inside the scenario** as scenery and serves them from it.
What a declaration pins is reported when the scenario is raised, because the digest belongs to
that registry and is not knowable before it exists.
**The digests are the lab registry's own, and that is correct rather than a workaround.** What
[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md) requires is a
reference that is exact and cannot move. A digest this registry assigned is both.
Verified in a machine confirmed to have no route out: `package`, `service` including boot state,
a `container` pinned by digest, and an `action` inside that container — applied, idempotent on
re-apply, and read back from the machine rather than from the apply's own report.
**One fault is worth keeping**, because it is this repository's own subject arriving in the
tooling built to catch it. The read-back checked that the registry's catalog endpoint answered,
by looking for the substring `repositories` — which `{"repositories":[]}` also contains. So it
**passed on a registry holding nothing**, and the failure surfaced much later as a container that
could not be pulled, a long way from its cause. It now asks for each image's manifest **by
digest**, which is what a machine actually does.
## The shape of a resolution
**A registry inside the scenario**, on its public segment, that machines pull from. That is not a
workaround: it is what the real mesh does — [ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)
names an OCI registry as substrate, and every node after the first pulls from the mesh's own.
Testing against a registry is testing the real path rather than a stand-in for it.
It also removes the lab's export-and-push mechanism rather than fixing it, which is the better
outcome: pushing image tarballs over the hypervisor was always a lab-only invention.
**Not decided here**, because it is design rather than repair: where the registry runs, whether
it is scenery like the router ([ADR 0016](../../02-DECISIONS/0016-the-lab.md))
or a placed artifact, and how images get into it.
## Incidental, and already fixed
The lab's own check on the load was too weak: it matched `"Loaded image"`, which is a prefix of
both `Loaded image:` and `Loaded image ID:`. So a load that produced an unusable dangling image
**reported success**, and the failure surfaced later as a container that would not start. It now
matches `Loaded image:` exactly and says what the runtime actually said.
That is this repository's own subject arriving in its own tooling: a check that passes on the
wrong thing is worse than no check, because it moves the failure away from its cause.
@@ -0,0 +1,104 @@
---
status: fixed
opened: 2026-08-29
located-in: [mesh-host, mesh-control]
fixed-by:
- "mesh-host: the store records where each resource came from — carried or declared — and each origin removes only its own. Verified in the lab on the exact scenario that caused this: the substrate survived, and a later declaration still removed what it had itself declared."
amended-design:
---
# 010 — The first declaration a node receives destroys the substrate it raised
## Symptom
A first node was raised from its carried bundle: container runtime, store, two context databases,
their schemas, broker, control plane. Eleven resources, all running. It then enrolled against the
control plane on its own machine, held its link open, and was sent a declaration naming two
resources — a directory and a file.
Both were applied correctly. And **every container on the machine was removed**: the store, the
broker, and the control plane that had sent the declaration. The link died mid-sentence with
`the link closed: Exception (501) Reason: "EOF"`, because the broker carrying it had just been
torn down by the message it carried.
Afterwards `mesh-host owned` listed two resources. The mesh had deleted itself.
## What is actually wrong
Nothing in the code is behaving incorrectly. `apply` removes what the store holds and the incoming
declaration does not name, which is what reconciliation means — the declaration is the desired
state, not a patch, and anything else would make it impossible to remove a resource by omission.
**The fault is that the carried bundle and mesh declarations share one store.** The host cannot
tell "this machine raised this for itself before there was a mesh" from "the mesh told this
machine to have this", so the second overwrites the first completely.
That is invisible until the two meet, which happens exactly once per mesh: on the first node,
after enrolment, at the moment the control plane first speaks.
## Why it matters more than a footgun
**The first node is the only node where the substrate is not the mesh's doing.** Every other node
receives everything it runs from the control plane, so a complete declaration is complete by
construction. The first node raised its own substrate from a file it carried, and the control
plane has never been told about it — so the control plane cannot include it in a declaration even
if it wanted to.
So the first node is left in a state no other node is in, and the ordinary path destroys it.
## What is not the answer
- **Making the control plane send the substrate back.** It does not know what the bundle contained
and should not: the bundle exists precisely because there was no control plane yet.
- **Making apply stop removing orphans.** Removal by omission is how a declaration says *stop
running this*, and losing it costs the property that a node converges on what it was told rather
than accumulating.
- **Special-casing the first node.** [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)
is explicit that its specialness lasts two commands, and this would extend it for ever.
## The fix
**The store records where each resource came from** — `carried` or `declared` — and each origin
removes only its own. A declaration removes what the mesh previously declared and never what the
bundle raised; reconciling the bundle removes what the bundle previously raised and never what the
mesh assigned.
State written before the field existed reads as `carried`, because everything a host had applied
at that point came from its bundle — there was no other way to tell it anything. Guessing the
other way would have the first upgrade remove the substrate, which is this fault arriving through
the change that fixes it.
**Verified on the scenario that caused it.** A first node raised eleven resources, enrolled, and
was sent the same two-resource declaration. Both applied; the store, the broker and the control
plane were still running afterwards. A second declaration dropping one resource removed that
resource and nothing else, so removal by omission still works — which is the property that had to
survive the fix.
**What remains open** is what happens when the mesh eventually declares the substrate, which it
must, or the substrate can never be upgraded. Two sources claiming one container is the ambiguity
this issue is made of, narrowed rather than removed: it can no longer happen by accident, and
nothing yet says what it means when it happens on purpose.
## How it was found
In the lab, on a sealed machine, by doing the ordinary thing: raise a first node, enrol it, and
tell it something. It was not a test of this — it was the first end-to-end run of the link, and
this fell out of it.
The declaration was two lines and destroyed a working mesh in under a second, which is worth
holding on to: this is not an edge case reached by trying, it is the first thing that happens.
## Two more faults found while fixing it
Both of the same shape, and worth recording because the shape is the point.
**A report published to an unbound routing key vanishes.** The control plane bound `enrol` and not
`report`, so nodes announced what they had applied into a void — the broker accepted each message,
found no queue for it, and dropped it. The publisher was told nothing. Reports are now published
`mandatory`, so anything unroutable comes back and is said out loud, and the binding covers every
key a node may publish.
**A publish failure was being swallowed.** `publishReport` discarded its error, so a node that
could not tell the mesh what it had done looked identical to one that had. That is the fault this
repository keeps cataloguing, written by hand into the newest code in it.
@@ -0,0 +1,85 @@
---
status: fixed
opened: 2026-08-30
located-in: [mesh-host]
fixed-by: mesh-host — apply attempts every resource and reports every failure
amended-design:
---
# 011 — One broken module stops every module after it, for ever
## Symptom
A machine was assigned a module declaring a package that does not exist. Every later push to that
machine applied **nothing at all**, and kept doing so.
Found while proving something else. A test assigned a deliberately-impossible module to a machine
to check that the mesh reports a failure — which it does. A later test on the same machine then
failed, and the evidence said why:
```
applied 0 and failed: applying "impossible.nothing":
installing a-package-that-does-not-exist: target not found
0 resource(s) were applied before this and remain
```
The broker's queues were **empty**, so the declaration had been delivered and read. The machine
simply stopped at the first failing resource and never reached the rest.
## Why it matters more than one machine
- **A machine with one bad module and nine good ones runs none of the nine**, and the mesh reports
"failed" without saying that the rest were never attempted.
- **It cannot be recovered by retrying.** Anything that re-pushes to machines that are behind — the
obvious next feature — would retry a permanent failure for ever and make no progress on
everything else.
- **The order is not the operator's.** Which module is "first" is an accident of resolution, so
which nine modules a broken one blocks is unpredictable.
## What was there, and what it rested on
The behaviour had a test asserting it: *nothing after the failure ran*. Its comment cites
[ADR 0010](../../02-DECISIONS/0010-delivery.md).
**That record does not decide this.** What it says is that a failed *job* stops and names its step
while a reconciler retries forever, as an argument about pipelines against reconcilers. It says
nothing about whether one resource failing should prevent the next from being attempted. The
citation was doing more work than the record supports.
## The fix, and the argument that was on the other side
**Everything is attempted, and every failure is reported.**
The case for stopping is that a resource may depend on an earlier one — a service on the file it
reads. That is real, and it survives: such a service fails its own check and is reported. This host
reads back after every write precisely so a thing that did not work is caught rather than assumed,
so attempting it produces *more* information than skipping it.
**What is unchanged:** a declaration that cannot be parsed is still refused whole, and nothing is
applied. That is a different thing — *this machine could not do it* against *this was never a
declaration* — and they are fixed in different places.
### And one shape still stops what follows, which the first fix got wrong
**An action does.** The first version of this fix continued past everything, and the next lab run
failed at the bootstrap: the store did not answer in three minutes and then said *the database
system is shutting down*. Carrying on past the readiness gate had started the broker and the
control plane against a machine that was not ready, and on a small machine that is how a database
still initialising has its memory taken away.
**An action is the only shape whose purpose is to make something true *before* the next thing needs
it** — which is why it is the only one with a `verify`. The bootstrap is a row of them: the store
answers, then its databases exist, then their schemas, then the broker. Everything else is
independent state: a package that will not install has nothing to do with a file on the other side
of the declaration.
So the rule is: **a failed action stops what follows; nothing else does.** Both faults are fixed by
it, and the report says which happened — *these things failed* and *these things failed and the
rest was never tried* are different machines.
## How this is checked
`internal/apply`, three tests: an apply with a resource that cannot succeed still applies the ones
after it; every failure is counted, not just the first; and a failed **action** stops what follows
and says so. Each confirmed to fail when its behaviour is removed — including the last, which fails
if actions stop being treated as gates *and* if everything is treated as one.
@@ -0,0 +1,88 @@
---
status: resolved
opened: 2026-08-30
located-in: [mesh-lab]
fixed-by: mesh-lab — scenario machines stay at 1 GiB, and the scenario now stocks seven images
amended-design:
---
# 012 — Raising a scenario machine's memory stops the substrate coming up
*Renamed after the first diagnosis turned out to be wrong. What that was, and how it was wrong, is
below — it is the more useful half of this report.*
## Symptom
Adding a fourth image to the two-machine scenario made the bootstrap fail every time. The store
container was created, and its readiness check then failed for the full three minutes with
**no output at all**:
```
failed store-ready (in mesh-store: ... pg_isready ... exit 1):
the action ran without error and its own verify still fails: docker exited 1:
```
Empty after the colon. The check runs `pg_isready` inside the container and prints the store's own
last lines when it gives up; producing nothing means **the container was not running**, which is a
different fault from a database that is slow to start.
Three images: the scenario raises, both machines join, and nine assertions pass. Four: it never
gets past the store. Reverting the fourth image restores it.
## The first diagnosis was wrong, and this is why it is worth writing down
**Two things changed at once.** A fourth image was added to the scenario, and — reasoning that a
machine running a database, a broker and the control plane at once is genuinely small — scenario
machines were raised from 1 GiB to 2 GiB. The bootstrap then failed every time, and the fourth
image was blamed.
**Removing the image did not fix it. Removing the memory increase did.** Nine assertions pass again
with the extra memory reverted, on a scenario with three images. So the cause is the memory change:
three machines at 2 GiB, on a host also running other work, contend enough that the store container
does not come up at all.
**The lesson is the ordinary one and it still caught me:** two changes went in together, the failure
was attributed to the plausible one, and an issue was written recording the wrong cause. What found
it was reverting to the exact last-known-good state rather than reverting the suspicious change.
**What remains untested** is whether a fourth image alone is fine. It probably is. Nothing has
measured it, and the honest state of this issue is that the thing it was opened about was never
demonstrated.
## What was worth keeping
**A better diagnostic.** The readiness check now prints what it saw before giving up. That is what
showed the output was empty, which is what said *the container is not running* rather than *the
database is slow* — and which will make the next occurrence of this a diagnosis instead of a
retry.
## What this blocks
The mesh running its **own artifact store** — a registry as a module — needs a registry image on
the machine so the module can mirror one, which is the fourth image. The module is written and its
manifest is accepted; what has not been proven is a machine assigned it serving artifacts to
another machine.
## How this will be checked
A scenario raised with four images comes up and passes the assertions that three do — which is
what was never actually established. Until then the artifact-store test is not in the shared
scenario, with a note saying where it went and why.
## Resolved
*2026-08-31.* The condition this report set was *a scenario raised with four images comes up and
passes the assertions that three do*. The scenario now stocks **seven** and has raised cleanly
many times over, with the machines at 1 GiB where the wrong diagnosis had put them at 2.
So both halves are settled. **The memory increase was the cause** — reverted, and never
reintroduced. **A fourth image was never the problem**, which this report said had not been
demonstrated either way, and now has been: three more were added on top of it, and the artifact
store the issue said was blocked is proven in the shared scenario rather than kept out of it.
**The diagnostic that came out of it is what remains valuable.** The readiness check prints what it
saw before giving up, which is what turned *the database is slow* into *the container is not
running*. It has since caught a different fault of the same shape — an action succeeding into a
state its own verify rejects
([04-ISSUES/017](../017-an-action-succeeded-into-a-state-its-verify-rejects/00-report.md)) — which
is the argument for keeping a good diagnostic after the incident that prompted it is gone.
@@ -0,0 +1,56 @@
---
status: resolved
opened: 2026-08-31
located-in: [mesh-control]
fixed-by: mesh-control — what the mesh computes is applied before what the module declared
amended-design:
---
# 013 — A file the mesh computes arrives after the service that needs it
## Symptom
Everything the control plane computes for a module — a certificate, a sealed credential, a bound
file, a rule set — was placed **after** that module's own resources in the declaration. The host
applies resources in the order it is given and
[does not sort](../../02-DECISIONS/0005-the-node-host.md), so a service or container declared in a
manifest was applied **before** the file it depends on existed.
On the first apply the service starts against a missing file and fails. The next reconcile finds
the file there and starts it.
## Why this matters
**It repairs itself, which is why nothing caught it.** A fault that is gone by the second attempt
is worse than one that persists: what gets remembered is that the thing works, and the failed
first apply is read as a machine that was briefly slow. The mesh reports a failure, then reports
success, and nobody looks again.
It was also invisible to every test that existed, because none of them combined the two halves.
Modules with computed files declared no service; modules with a service needed no computed file.
The fault lived exactly in the gap between two repositories' assumptions — the control plane
deciding an order, the host promising not to change it — which is the shape this folder exists for.
Found by reading, while writing the first module that has both: a firewall whose service must
reflect a rule set the mesh computes.
## Evidence
`internal/catalogue/declaration.go` built each module's resource list as
`append(module's own, computed...)` in six places — certificate, authority, needs, secrets, grants,
bindings. `internal/apply/apply.go` iterates `d.Resources` in order, and
`internal/declaration/declaration_test.go` states the rule directly: *order is stated, not derived.
The host must not sort.*
## What was done
The computed resources are assembled first and the module's own resources follow. Nothing the mesh
computes is derived from a module's resources, so the order is unconditionally right rather than a
heuristic — there is no case where a module's resource must precede a file the mesh made for it.
Merged after the computed-resources branch, which replaces a module's resources wholesale and
would otherwise discard everything the mesh had made for it.
*Checked by a module declaring a service that reflects a rule set, asserting the rule set is first;
and by a module whose resources are computed elsewhere, asserting its credential survives and is
still first — the case the merge point exists for.*
@@ -0,0 +1,48 @@
---
status: resolved
opened: 2026-08-31
located-in: [mesh-host]
fixed-by: mesh-host — a node's serving key is stored in the format a server reads
amended-design:
---
# 014 — A node's serving key was present, correct, and unusable
## Symptom
A node generates the key it serves TLS with, the mesh certifies the public half, and the
certificate arrives on the machine as an ordinary file. Everything about that worked. But the host
stored the private half in its own encoding — base64 of the raw key — and **nothing that serves
TLS can read it**: not a web server's `ssl_certificate_key`, not Go's `LoadX509KeyPair`, not
`openssl s_server -key`.
The file was there, owned by root, mode 0600, holding the right key. The certificate beside it was
valid and chained to the mesh's authority. The server would not start.
## Why this matters
**Every check that reads the file passes.** The key exists, the certificate exists, the mesh
recorded the public half, the machine reports the declaration applied. The failure surfaces only
when something connects — the worst place to find out, and the place the certificate work was
specifically designed to move away from.
It is the same shape as [013](../013-a-file-arrives-after-the-service-that-needs-it/00-report.md)
and worth naming as a class: **two halves of one mechanism designed separately, each correct
about its own half.** The control plane issues PEM because that is what a certificate is. The host
stored the key in whatever was convenient, because nothing in the host reads it back — the whole
point of the file is that *something else* does, and that something else was not in view.
**The generalisation:** where a file exists so a third party can read it, the format is not an
implementation detail of whoever writes it. It is the interface, and it needs a check that reads
it the way that third party will.
## What was done
PKCS#8 PEM, which is what every TLS server reads. A key in the old encoding is refused **by name**
rather than reported as corrupt — it is intact, and the remedy is to enrol again, which is a
different action from repairing a damaged file.
*Checked by writing a key, decoding the file as PEM, parsing it as PKCS#8, and asserting it is the
same key — and, in the lab, by a real handshake from a second machine that verifies against the
mesh's authority and nothing else. A key that parses is not a key a server can use, which is why
the lab check connects.*
@@ -0,0 +1,44 @@
---
status: resolved
opened: 2026-08-31
located-in: [mesh-lab]
fixed-by: mesh-lab — a command with no marker is a failure, not a success
amended-design:
---
# 015 — A command that said nothing was read as having succeeded
## Symptom
The end-to-end harness runs a command on a machine and reads its exit status from a marker it
appends to the output. When the marker was absent, the parse produced `Number("")`, which is `0`,
and **the command was reported as having succeeded.**
The marker went missing whenever a command contained a heredoc. Everything was wrapped on a single
line — `<cmd> 2>&1; echo "__exit=$?"` — so a heredoc's terminator line became
`MARKER 2>&1; echo "__exit=$?"`, matched nothing, and the heredoc consumed the rest of the script,
the marker included.
## Why this matters
**This is the harness lying in the one direction a harness must never lie.** Everything else in
this repository is arranged around the principle that absence must never be indistinguishable from
success — the host says so about a service that does not exist, the builder says so about a build
that failed, the control plane says so about an empty list. The thing that checks all of that had
the fault itself.
Its reach is every heredoc in the suite, which is how large files are written to machines: a
substrate bundle, a certificate authority, a listener script. Each was written with trailing
junk from the swallowed wrapper, each reported success, and each happened to be tolerated by
whatever read it — until one was a Python script, which did not run, and the test failed on its own
setup. **That failure read exactly like the thing being tested working.**
## What was done
`exec 2>&1` on its own first line, so nothing is appended to the command's last line and a heredoc
terminates where it says it does. A missing marker is now a failure, returning whatever was said
so the reason is visible rather than inferred.
*Checked by the firewall test, whose listener is written with a heredoc: it could not have started
before this, and the assertion that it is reachable before any rule set exists is what makes the
rest of that test mean anything.*
@@ -0,0 +1,42 @@
---
status: resolved
opened: 2026-08-31
located-in: [mesh-host]
fixed-by: mesh-host — something after the declaration is refused whole
amended-design:
---
# 016 — Anything after the declaration in a file was ignored
## Symptom
A JSON decoder reads one value and stops. The host's declaration parser used one, so a file
holding a declaration **followed by anything at all** — a stray line, a second declaration, the
tail of a truncated rewrite — parsed as the first value and the rest was never looked at.
The machine applied something, reported success, and what it applied was not what the file said.
## Why this matters
The host refuses a partial declaration everywhere else, in these words: *a host that applied the
parts it understood would leave a machine that looks configured and is not.* This was the same
fault in its quietest form — not a part left out, but a part never seen, with nothing anywhere
saying so.
**It was live and invisible.** The end-to-end harness had been appending a line to the substrate
bundle by accident ([015](../015-a-command-with-no-answer-was-read-as-success/00-report.md)), and
every bootstrap in every run applied a bundle with a junk line on the end. Nothing failed, so
nothing was looked at, and the corruption was found only by tracing a different bug backwards.
**That is the argument for refusing rather than tolerating.** A file with something after it is
more likely to be damaged than deliberate: an interrupted write, two files concatenated, a
generator that emitted twice. Applying the first value is applying something nobody wrote.
## What was done
After decoding, the parser requires end of input. Trailing whitespace is not "something after it"
— refusing that would make every file an editor writes unusable.
*Checked by a valid declaration with a stray line after it, with a second declaration after it,
and with garbage after it, each refused; and by the same declaration with trailing blank lines,
accepted.*
@@ -0,0 +1,44 @@
---
status: resolved
opened: 2026-08-31
located-in: [mesh-host]
fixed-by: mesh-host — the store's readiness is checked over TCP, not the socket
amended-design:
---
# 017 — An action succeeded into a state its own verify rejects
## Symptom
The substrate's `store-ready` action waits for the store to answer, then the host runs the
action's `verify` to read back that it worked. Intermittently the host reported:
> the action ran without error and its own verify still fails
Both statements were true. The action waited on the **unix socket**; its verify checked the same
way, a moment later, and found nothing. Roughly one bootstrap in three.
## Why this matters
**The action and its verify were asking different questions without appearing to.** While the
store initialises it runs a temporary server on the socket only, then stops it and starts the real
one. The action's loop saw the temporary server and exited happy; the verify landed in the gap
between the two.
So an action can **succeed into a state its own verify rejects** — and when it does, the host's
report is accurate and useless. It says the command worked and the read-back did not, which is
exactly what the mechanism is for, and names nothing a person can act on. It read as a slow
machine, and the remedy people reach for is a longer timeout, which cannot help.
**The general rule, which the host's design should carry:** an action's verify is the *definition*
of what the action is for. If the action's own waiting decides it is done by a different test than
the verify uses, the two can disagree — and the disagreement surfaces as an intermittent failure
in the one place designed to catch silent success.
## What was done
Both check the store over TCP, which the init phase deliberately does not open — so neither can
mistake the temporary server for the real one, and neither can be satisfied while the other is not.
*Checked by the bootstrap itself, which is where it failed: this action gates everything after it,
so a mesh coming up at all is the check.*
@@ -0,0 +1,51 @@
---
status: resolved
opened: 2026-08-31
located-in: [mesh-control]
fixed-by: mesh-control — something answered on this machine is still bound
amended-design:
---
# 018 — A provider on the same machine was never announced to its consumer
## Symptom
A module that `binds` a provision receives a file naming where the provider is and what it said a
consumer must know. When the provision turned out to be answered by another module **on the same
machine**, no file was written at all.
A build machine sharing a node with the registry it pushes to therefore started, connected to the
broker, and looped: *cannot read what the mesh said about the artifact store: no such file or
directory.*
## Why this matters
It was deliberate, and the reasoning is in the code: *a file saying "it is on this node" would be a
fact nobody needs and one more thing to keep true.* That is **right about the location and wrong
about everything beside it.** A binding also carries the provider's `serves` block — the port —
and a consumer cannot invent that whether the provider is next door or on the same disk.
**The failure names nothing.** Every part a person would check was correct: the module resolved,
the machine applied it, the container ran, the credential was delivered and worked. The one file
that did not exist was one the module never asked for by name — it asked for a *provision*, and
the mesh silently decided the answer needed no writing down. The error is about a path, and the
cause is a decision three layers away.
**And it only appears when two modules land on one node**, which is the ordinary case in a small
mesh and the rare case in a large one. It would have been found in production.
## What was done
The binding is written for a local provider too, **when the provider said something a consumer
must know**. That keeps the original intent exactly where it was right: a shell is answered here
and there is genuinely nothing to say about it; a registry is answered here and the port is still
unguessable.
The address is this machine's own name on the private network, or loopback when it has none — a
machine off the network still reaches itself, and a name nothing resolves is worse than an address
that always works.
*Checked by resolving a node holding both a provider and its consumer and asserting the binding
carries the port and the address; by the same on a machine with no private network, asserting
loopback; and by the pre-existing check that a provision with nothing to say still writes nothing,
which is the half that was right.*
@@ -0,0 +1,59 @@
---
status: resolved
opened: 2026-08-31
located-in: [mesh-control]
fixed-by: mesh-control — the resolver module's claims about machines are checked on a machine
amended-design:
---
# 019 — A comment asserting a fact about a machine, which nothing checked
## Symptom
The resolver module carried two statements about the machine it runs on. Both read as reasoned,
both were in prose beside the setting they justified, and **both were wrong**:
| it said | the machine said |
|---|---|
| `127.0.0.54` is free — "not `.53`, that is systemd-resolved's" | systemd-resolved holds **both**; `.54` is its proxy stub. dnsmasq could not create the socket and never started |
| it takes only `127.0.0.55` | listening on a loopback address takes the rest of loopback with it, `127.0.0.1` included |
A third statement in the same file was true and incomplete in a way that mattered as much: the
config read `/etc/resolv.conf` for upstreams without saying so, and the module that points a
machine at the mesh writes *this resolver's own address* into that file. So its upstream was
itself. Its receive queue filled with 15KB of queries and every lookup on the machine hung.
## Why this matters
**The module had unit tests, and they all passed.** They checked that it names an address, that it
reads what the mesh writes, that it restarts when that changes, and that the two asking modules
point where it answers. Every one of those was true while the daemon could not start at all.
That is not a gap in those tests. It is what a unit test *is*: it confirms the assertion was made,
never that it is true of any machine. **Only a machine knows which of its addresses are spare, or
what a daemon does with a file when it starts.**
This is [04-ISSUES/003](../003-firewall-scope-is-read-by-no-code/00-report.md) in prose rather
than in a manifest key. There, five manifests carried a `scope:` that read as a restriction and
restricted nothing. Here, a comment read as a reasoned choice of address and chose a taken one. In
both cases *an unenforced rule is indistinguishable from a wrong one, and costs more, because
people believe it* — and a comment is the least enforced rule there is.
**It cost three full lab cycles**, at fifteen minutes each, because each one revealed exactly one
of the three faults.
## What was done
**The wrong statements are corrected, and the correction says what it now knows rather than
asserting a new comfort.** `.55` is written down as *a convention, not a reservation*: if a future
systemd takes it, one line changes. What the module takes is what its claim already said — the
machine's DNS port — rather than a promise about one address.
**The unit tests hold what a machine has told us.** They assert the module does not take `.53`,
`.54` or `127.0.0.1`, and that it does not read resolv.conf for upstreams. A unit test cannot
discover those facts; it can refuse to forget them.
**And the order changed.** A module that asserts something about machines is proven on a machine
*before* its assertions are believed — the lab test written first, not last. Written here because
the cost of the old order is measurable: three cycles, forty-five minutes, for a module whose
mesh-side half was correct from the start.
@@ -0,0 +1,94 @@
---
status: open
opened: 2026-08-31
located-in: [mesh-control, mesh-lab]
fixed-by:
amended-design:
---
# 020 — A certificate is issued and never collected
## Symptom
Against a real ACME server in the lab, the proxy orders a certificate for a name the mesh routes,
the challenge is answered, the authority **issues the certificate** — and the proxy never obtains
it. Every TLS handshake then fails, and the order is retried indefinitely.
The client's error, once per attempt:
```
http: TLS handshake error: Post "": unsupported protocol scheme ""
```
A POST to an empty URL: the certificate's location, on an order the authority considers valid.
## What is proven, and it is most of it
Read from the authority's own log rather than inferred:
```
Starting 3 validations
authz … set VALID by completed challenge …
POST /finalize-order/ → Order … is fully authorized. Processing finalization
Issued certificate serial 3ef142939115ee88
```
**The hard half works.** The order is created, the HTTP-01 challenge is answered *at the name being
certified* on port 80 through the proxy itself, the authorisation goes valid, finalisation is
accepted, and a certificate is issued. Across one run the authority issued **two** certificates and
accepted finalise **three** times — the client reaches issuance every attempt and fails at the same
step after it.
**And the policy that guards the quota is proven too.** The second assertion in the same file
passes: no certificate is ordered for a name nothing routes, so a scan cannot spend an account's
rate limit.
## What is not known
**Whether this happens against a real authority at all.** Everything above is against Pebble, which
exists to be a test server. The failure is in the last hop between one client and one server, and
may say nothing about behaviour against a public authority.
## Ruled out
| | |
|---|---|
| the directory | fetched and complete — `newAccount`, `newNonce`, `newOrder`, `revokeCert` all present |
| the authority's API certificate | covers `127.0.0.1`; the bundle is named explicitly and verification is not skipped |
| a hand-written server config | suspected, and wrong. Replacing it with the server's **own** default config, changing only the challenge port, gives the identical error |
| the finalize URL being empty | the authority logs finalisation being accepted |
| the challenge path | the authorisation goes valid |
| the server version | pinned 2.5.0 behaves exactly as `latest`, so the draft profiles extension is not it |
## Where it might be
- **The order's `certificate` field is absent when the client reads it.** The client waits for the
order to become valid and only then fetches, so an empty location on a valid order is the
remaining shape.
- ~~**A moving tag was used.**~~ **Ruled out.** `pebble:latest` advertises a draft *profiles*
extension, so a pinned 2.5.0 was tried: **identical failure**. The scenario now pins it anyway,
which it should have from the start.
## Why this is filed rather than pursued
**The mesh-side behaviour is proven and the remainder is interop between two libraries.** Continuing
would be several more twenty-minute lab runs against a server that is not the one production uses,
to chase a defect that may not exist there.
**What the mesh needed to show, it showed**: a name it routes gets a certificate ordered from a
configured authority, and a name it does not route gets nothing. The configuration is right, the
challenge path is right, and issuance happens.
## What would close it
- The same scenario against a different ACME implementation — a second server, or a real staging
endpoint from a machine that can reach one. **If it passes there, this is a Pebble interop
detail and the issue closes with that recorded.**
- Or the client's request captured on the wire, showing what the order actually contained when the
location was read.
## Evidence
- `mesh-lab test/integration/certificates.test.ts`, scenario `a-public-name`
- One assertion passes (no certificate for an unrouted name); one fails (a routed name is never
served).
@@ -0,0 +1,90 @@
---
status: fixed
opened: 2026-09-01
located-in: [mesh-control]
fixed-by: mesh-control df62bb5
amended-design:
---
# 021 — A consumer on the provider's machine is given no credential
## Symptom
A module that requires something answered **on the same machine** resolves cleanly and is given
**no credential at all**. Two modules, zero needs:
```
postgres provides postgres-database, grants /var/lib/postgres/grants
keycloak requires postgres-database, secrets /var/lib/keycloak/database.env
→ modules: 2, needs: 0
```
Nothing is refused and nothing is reported. The consumer's `secrets:` path is simply never
written, and whatever reads it fails later, somewhere else.
## Where it comes from
The world a node resolves against is **every other node**:
```go
for _, n := range nodes {
if n.Name == exclude { continue }
```
So a provider on the same machine is never a `Provider` in `world.Offered`, never becomes a
`Needed`, and the credential loop — which walks `resolved.Needs` — has nothing to walk. Every step
is individually reasonable and the sum is a silent gap.
## Why it was not noticed
**Everything proven so far was cross-machine.** The lab's provisioner scenarios put the consumer on
one node and the provider on another, which is the interesting case for a *mesh* and the rare case
in practice. The first module to want a database on its own machine was the first real one.
The postgres provisioner even records the assumption in passing — *"Node is empty for a module on
this machine, which is asking for something local and is not this provisioner's business"* — which
reads as a deliberate exclusion of local consumers.
## Why the assumption is wrong
It holds for a process on the machine reaching a unix socket, where the operating system can vouch
for who is calling. **It does not hold for containers**, which is how nearly everything runs here: a
module's containers reach a provider's containers over TCP on a shared network, and the database
asks for a password exactly as it would from another machine.
**The machine is not a trust boundary once both sides are containers.** Treating it as one gives
the most common arrangement — a service and its database on one node — the weakest handling.
## What it is not
Not the same as [`020`](../020-a-certificate-is-issued-and-never-collected/00-report.md) or a
provisioner defect. The provisioner never sees these consumers because the mesh never records
them as consumers.
## What a fix has to keep
- **A local consumer still appears in the provider's grants**, so its provisioner creates the role
or bucket or client, exactly as for a remote one.
- **The credential is still sealed**, to the one node that is both ends. The mesh holding a
readable secret for local consumers would be a hole opened for convenience.
- **Refusing must stay refusing.** A requirement nothing answers is still refused; this is about a
requirement that *was* answered.
## Fixed
A requirement answered on this machine is still a requirement. Resolution now records a need for
it, so a credential is made, the provider is told who asked, and the consumer's file is written —
the same as if the two were on different machines.
The reasoning that made it a gap is now written where it was assumed: the machine is not a trust
boundary once both ends are containers, and treating it as one gave the commonest arrangement of
all — a service and its database on one node — the weakest handling.
Two later issues came out of the same mistaken instinct and are worth reading together:
[`022`](../022-one-credential-per-node-per-provision-not-per-module/00-report.md), where the
machine was treated as an *identity* rather than a boundary, and
[`023`](../023-a-consumer-cannot-build-a-connection-string/00-report.md), where the consumer was
given a password and never told the name to present with it.
*Closed 2026-09-01. The fix landed the same day and this record was left open by oversight — the
code and the tests were in place for hours while the record still said `located`.*
@@ -0,0 +1,113 @@
---
status: fixed
opened: 2026-09-01
located-in: [mesh-control]
fixed-by: mesh-control 0af3ea1
amended-design:
---
# 022 — A credential belongs to a node and a provision, so a second consumer refuses
## Symptom
A node running more than one module that wants the same provision **cannot be planned at all**:
```
anchor has 3 modules asking for "postgres-database" and they would share one
credential: gitea, keycloak, umami
```
**And that is only the loud half.** The consuming node does not refuse at all. Three modules
wanting one database produce **one** need:
```
modules=3 needs=1
name=postgres-database from=anchor for=gitea
```
So the first module gets a credential, the other two get no file at all, and each starts and fails
to authenticate with nothing anywhere saying why — the shape of
[`021`](../021-a-consumer-on-the-providers-machine-is-given-no-credential/00-report.md), on a
different axis. The refusal that reads like a decision is on the provider; the silence is on the
consumer.
The refusal is correct about what it says. They *would* share one credential, and sharing one is
worse than refusing — a login that opens three databases is not three credentials. But the
arrangement being refused is the ordinary one. **The node this mesh exists to take over runs
eight modules against one database server.**
## Where it comes from
A credential is keyed by *(provision, consumer node, provider node)*:
```go
func (i *Inventory) SecretFor(ctx context.Context, name, consumer, provider string) (Secret, error)
```
`consumer` is a **node**. Everything downstream inherits that granularity: `Grant.Consumer` is a
node, the grant's file is named after a node, and the provisioner names the role it creates after
one — `role := mark + c.Node`.
So the refusal in `ContributionsTo` is not a check that found a problem. It is the only honest
thing that function can do, given a key that cannot tell two consumers apart.
## Why it was not noticed
**Every scenario so far had one consumer per node.** That is the natural shape of a small test —
a consumer here, a provider there — and it is the shape of every lab scenario written to date. A
node with two modules wanting a database is not an edge case discovered by fuzzing; it is what a
real machine looks like, and nothing had modelled a real machine yet.
The refusal also reads as deliberate. It names the modules, explains the consequence, and refuses
rather than picking — the house rule everywhere else. It looks like a decision. It is a limit.
## Why the granularity is wrong
The same argument that closed [`021`](../021-a-consumer-on-the-providers-machine-is-given-no-credential/00-report.md).
There, the machine was treated as a trust boundary and containers made that untrue. Here, the
machine is treated as an *identity* — as though "who is asking" is answered by naming a host.
**Two modules on one node are as separate as two on different nodes.** They run as different
containers, on different networks, with different data. A key that cannot distinguish them means
the mesh cannot express the thing it is for.
It also silently weakens what the provisioner does. `mesh_<node>` is one role. Had the refusal not
been there, gitea's login would have opened keycloak's database — and nothing anywhere would have
said so, because from the provisioner's side it created exactly what it was asked to create.
## What a fix has to keep
- **The refusal, where it is still right.** Two modules wanting one provision must not silently
share a credential. After a fix they do not share one, so there is nothing to refuse — but a
genuine collision must still refuse rather than pick.
- **A credential per consuming module**, sealed to the node that holds it. Both facts are needed:
the module is who it is for, the node is what it is sealed to.
- **The provisioner names what it creates after the module**, so a login is traceable to the thing
using it, and so withdrawing one consumer does not remove another's.
- **Withdrawal still works.** A module unassigned must lose its login while the others keep theirs
— which is precisely what one role per node cannot do.
- **Existing single-consumer nodes keep working**, since that is every scenario that exists.
## Scope
This crosses the control plane, the grant file naming, and every provisioner that names something
after `Consumer`. It is not a local fix, and it is the last thing between the current state and a
node that looks like a real one.
## Fixed
Needs fan out per consuming module in one place, after the resolution walk. The credential's key
gains the consuming module, the grant file is named after both halves of the consumer, needs are
matched by provision *and* module, and the provisioners name the role and the access key after the
module rather than the machine. The refusal is gone because there is nothing left to refuse.
Existing credentials are discarded rather than backfilled: they cannot say which module they were
for, and one is remade and delivered to both ends on the next push, so it costs one rotation.
A guard was added for PostgreSQL's 63-byte identifier limit, which truncates with a notice rather
than an error — two consumers whose role names agree that far would otherwise become one login,
which is this same fault at a length nobody would think to test.
What it did **not** fix is [`023`](../023-a-consumer-cannot-build-a-connection-string/00-report.md):
a consumer now receives its own password and still cannot build a connection string, because the
user name is the provisioner's invention and the bound values cannot reach a configuration file.
@@ -0,0 +1,111 @@
---
status: fixed
opened: 2026-09-01
located-in: [mesh-control]
fixed-by: mesh-control 122680b
amended-design:
---
# 023 — A consumer is given every part of a connection except the two it cannot invent
## Symptom
A module that requires a database is now given its password in whatever shape its configuration
needs ([`022`](../022-one-credential-per-node-per-provision-not-per-module/00-report.md) and the
sealed-placeholder work). It still cannot connect, because a password is not a connection.
What it is given is a **binding**, as JSON:
```
provision postgres-database
from the node providing it
at that node's address on the private network
serves what the provider said a consumer must know — the port
```
What it needs, to write `KC_DB_URL` or `GITEA__database__USER`, is the host, the port, the
database name and **the user name**. Two of those are missing, for two different reasons.
## The user name is nobody's to say
The provisioner invents it — `mesh_<node>_<module>` — and nothing else in the mesh knows that
string. The control plane does not record it, the binding does not carry it, and the consumer
cannot derive it without hard-coding another module's naming convention.
So the one identifier a consumer must present in order to authenticate is the one thing no part
of the mesh will tell it. It works today only because nothing has yet had to write a connection
string; every proof so far stopped at "the credential arrived".
## The values cannot reach the file that needs them
The binding is a JSON document. The consumers are containers reading `KEY=value`, or a program
reading a YAML file, or one reading an attribute inside a different JSON document. A sealed secret
can now be placed inside any of those — the module writes the file with a hole in it and the host
fills the hole on the machine. **The bound values have no such route**, so the half of the
connection that is not secret is the half that cannot be delivered.
This asymmetry is backwards. The secret is the hard case, because the mesh must not be able to
read it. The host and port are ordinary facts the mesh knows in the clear, and they are the ones
stuck in a document nothing can read.
## Why it was not noticed
Every provider so far has been reached by a **provisioner**, a program written for the job, which
reads the JSON because it was built to. The first consumers to need a plain configuration file
were the first real applications. The binding was designed for the program and then handed to the
application.
There is also a stale comment saying a binding *"carries no credential: the mesh has no way to
issue one yet"*. That stopped being true when [`021`](../021-a-consumer-on-the-providers-machine-is-given-no-credential/00-report.md)
was fixed.
## What a fix has to settle
- **Who names the role.** Either the mesh records what the provisioner will create, or the
consumer contributes the name it wants and the provisioner uses it. The second is more in
keeping with the rest — a consumer already contributes the database name it wants — and it
removes an invented convention rather than documenting one.
- **How a bound value reaches a file.** The symmetric answer to the sealed placeholder, and
simpler: these values are not secret, so the control plane can put them in before sending and
the host learns nothing new.
- **That it stays name-agnostic.** The control plane must not learn what a `postgres-database`
is. What the keys mean is agreed by the requirement's name
([ADR 0027](../../02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md)), so
whatever is added has to work for a bucket and a mail relay without being told about either.
## Blocked on this
Keycloak, Gitea, Mailu and MinIO all have manifests that parse and resolve, and none of them can
start. This is what stands between the module set and a running one.
## Fixed
Both halves had the same cause: **the mesh knew something and did not say it.**
**Who a consumer is, said once.** `ConsumerIdentity(node, module)` is one derivation, sent to the
provider in its grant and to the consumer in its binding — so the two agree by construction rather
than by two conventions that happened to match on the day they were written. The provisioners now
use the name they are given and **refuse to invent one** if the mesh says nothing: falling back to
a name of their own would create a login the consumer could never guess, and everything would
report success. They also refuse a name that does not carry the mesh's prefix, because that prefix
is how withdrawal finds what it made.
**Bound values reach the file that needs them.** `${bound:provision:key}` is the symmetric twin of
the sealed placeholder and simpler: these values are not secret, so the control plane fills them
in before sending, and the host gains no field and learns no format. `at`, `as` and `from` are
true of any provision; every other key comes from what the provider said it *serves*, so the
control plane still learns nothing about what a `postgres-database` is.
Keycloak and Gitea now produce complete connection strings — asserted from the manifests on disk,
checking that every part is filled, that no placeholder survives as a value, and that the password
is still a hole only the host can close.
### Also found, and separate
The lab run that was meant to prove this failed in a way that looked like the fix being wrong: a
rotation test could not authenticate against a real database. The cause was that the suite
rebuilt the control plane's image and not the provisioner's, so a run with an image built that
minute used a provisioner built the day before. Fixed in `mesh-lab`; it is
[`005`](../005-pipeline-test-harness-unbuildable/00-report.md)'s family — a rebuild covering
most of what a run uses is worse than one covering none, because the run that follows it is
believed.
@@ -0,0 +1,106 @@
---
status: fixed
opened: 2026-09-01
located-in: [mesh-lab]
fixed-by: mesh-lab 3503ad9
amended-design:
---
# 024 — A run stalls before the host is placed, and says nothing while it does
## Symptom
The first scenario of the end-to-end suite raises its machines and then stops. Observed twice on
2026-09-01, both times after the rebuild step grew:
- **Once mid-run**, during the rotation test: the suite had reported thirteen passes, then the
process ended with no summary, no failure and no receipt.
- **Once from the start**: `a bare machine becomes a mesh` ran for **35 minutes** against a
measured 4.5, produced no output at all, and was still running when it was stopped by hand.
Both machines were `RUNNING` throughout. The second one was interrogated directly: the anchor VM
answered, and had **no `mesh-host` log and no containers** — so the run had not reached placing
the host. It was stuck earlier, in stocking the scenario's registry.
## What is not the cause
- **Not memory.** 84 GiB available, no OOM in the kernel log.
- **Not the daemon.** `incus exec` into the stalled machine answered immediately.
- **Not the changes under test.** The credential work is applied after the host is placed, and the
host was never placed.
## What changed just before — *and it was not the cause*
The rebuild step had gone from two artifacts to six, and every image is pushed into the scenario's
registry, which looked like where the stall sat. That was written down as a coincidence rather
than a diagnosis, and it is as well: **stocking takes 34 seconds and always did.** Timed directly,
eight images, before anything was changed.
The suspicion was the ordinary kind — the thing that changed most recently looks guilty — and the
thing that changed had nothing to do with it.
## Why it matters more than a slow test
**A stall is indistinguishable from work.** The suite prints nothing between starting a scenario
and finishing its first test, so four and a half minutes and thirty-five look identical from the
outside — and the operator's only recourse is to guess, which is precisely how a workstation was
left unbootable in August by killing a package manager that was working.
The first stall is worse: the process ended *silently* after thirteen passes. No summary, no
receipt, nothing that says the run was cut short. A run that stops without saying so is a run
somebody may believe.
## What a fix has to give
- **Progress while stocking**, so a long step is visibly a long step. Bytes moved, images pushed,
anything that changes.
- **A receipt when a run is cut short**, saying how far it got. `lastrun` already refuses to guess;
what is missing is it being written at all when the process dies mid-run.
- **A stated timeout on stocking**, so a stall ends as a failure with a reason rather than as a
process somebody eventually kills.
## The cause
**The registry machine was addressed by hand and every other machine was not.**
Machines get a systemd-networkd unit with a static `Address=`, so networkd finishes configuring
the link and reports it `configured`. The registry instead ran `ip addr add` inline. An address
put on a link that way leaves networkd still waiting to configure something it was never told
about, so the link sits at `configuring` — and `systemd-networkd-wait-online` has
`TimeoutStartUSec=infinity`.
So `network-online.target` is never reached, and **everything ordered after it never starts.** On
these machines that is Docker. `docker load` then blocks on a socket whose daemon is queued behind
a target that will never come, and the three bounded timeouts around it — save, push, load — stack
to thirty-five minutes.
Measured on one scenario, before and after:
| | before | after |
|---|---|---|
| the registry's link | `configuring` | `configured` |
| `docker.service` | inactive, 5 jobs pending | active, no jobs |
| the raise | never finished | **87.5 s** |
These machines have **no DHCP by design** — a scenario is a closed address space and the
declaration owns the addresses — so nothing was ever going to complete that wait.
## Fixed
- **The registry is addressed the way every other machine is**, through the same helper.
- **Placing an image waits for the container runtime** and refuses after 120s, naming what systemd
is still waiting on. A stall becomes a failure that says why.
- **The end-to-end test passes `onProgress`.** The raise reported every step and the test threw it
away, which is why thirty-five minutes of silence and four minutes of silence looked the same.
The suite then ran to completion: **23 of 24**, the one failure a check of its own that flagged
`/var/lib/mesh/builder/broker` as a credential because `/` is in the base64 alphabet. Fixed with
it.
## Also learned, at some cost
**A redirected log lags.** Node block-buffers stdout when it is a file, so `> run.log` sits
unchanged for minutes while the run is fine. That was read as a stall twice — the second time
immediately after the real fix, where a buffering artifact argues the fix did not work. The
machines answer instantly and are the source of truth. *"I cannot see progress" is not evidence of
no progress.*
@@ -0,0 +1,121 @@
---
status: located
opened: 2026-09-01
located-in: [mesh-control, mesh-host]
fixed-by: partly — mesh-control ee3cc1b
amended-design:
---
# 025 — A module must pin a digest, and nothing produces one
## Symptom
Every image reference in every example module is **sixty-four zeros**:
```
gitea@sha256:0000000000000000000000000000000000000000000000000000000000000000
```
Eighteen of them, across five modules. Each one parses, resolves, and composes into a declaration
a host accepts. None of them could ever start: the machine would reach `docker pull` and stop.
This is why those modules are *written* and not *running*, and it was not visible from any check
because every check passes.
## Why nothing caught it
The host validates the **shape** of a reference and nothing else — that it is `name@sha256:` plus
sixty-four hexadecimal characters. Sixty-four zeros satisfies that exactly.
That check is not wrong. A host cannot verify a digest exists without reaching a registry, and
reaching a registry is precisely what the design refuses to make it do
([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)). The host is the last
place that could catch this and the wrong place to try.
## The actual gap
**A manifest must carry a digest, and nothing in the system produces one.**
- Images the mesh builds are fine: the bundle writes the digest down *after* building, which is
the whole reason the bundle exists in that shape.
- Images from anywhere else — a forge, a mail system, a database — have no path at all. Somebody
has to look up what `gitea:1.22` points at today and paste it in, and nothing re-checks it.
So the design is coherent about *pinning* and silent about *where a pin comes from*. A person
writing a module is asked for something they cannot reasonably produce by hand, and given a
placeholder shape that passes every gate.
## What a fix has to keep
- **The host still refuses a tag.** A digest is what makes a declaration exact, and that must not
soften. The fix belongs where a module is added or built, not on the machine.
- **A person writes a tag; the mesh holds a digest.** A manifest in a repository naming
`gitea:1.22` is readable and reviewable; the mesh resolving that against a registry once, and
recording the answer, is what makes it exact. The module table already records this shape for
source repositories — where it came from, the branch followed, the commit read — and an image is
the same question asked of a registry.
- **Re-resolving is a decision, not a side effect.** A tag that moves must not silently change what
a machine runs. Whatever resolves it records both, so *this pin is behind its tag* is a question
the mesh can answer rather than something discovered on a restart.
## Cheaply, now
An all-zero digest is a placeholder and never a real image. Refusing it costs three lines and
would have caught all eighteen the day they were written. It does not fix the gap; it stops the
gap being invisible.
## What this blocks
Every module that names a third-party image, which is every module that is not the mesh itself.
The forge and the mail system are otherwise ready to run.
## Half of it is done
**A placeholder can no longer reach a machine.** The refusal sits where a declaration is composed,
not where a manifest is parsed — a file awaiting a pin is legitimate, and the design already says
so for artifacts the mesh builds. Composing is the last moment before a machine sees it.
**The examples now pin images that exist.** Twelve third-party digests were resolved against their
registries without pulling anything, which is also the mechanism the rest of this issue needs:
`docker manifest inspect --verbose` answers *what does this tag point at* in about a second.
Two faults came free, and both had been invisible for the same reason as the digests: the mail
system's seven images named repositories that **do not exist** — it publishes to a different
registry entirely — and one of the seven had been renamed upstream. Nothing that only checks the
shape of a reference could ever have found either.
## The mechanism already existed, and this issue was wrong about that
**Corrected 2026-09-01, the same day.** This was filed saying nothing turns a tag into a digest.
That is false, and the answer had been designed and built before any of it was written.
A module does not name an image at all. It names an **artifact**, and declares where that artifact
comes from:
```
build.artifacts: [{ name: "gitea", kind: "upstream", from: "gitea/gitea:1.22" }]
resources: [{ id: "server", type: "container", artifact: "gitea", … }]
```
`kind: upstream` means *an image somebody else built, mirrored into the mesh's own registry and
pinned by the digest it lands with*. The builder produces it; the manifest the mesh holds is
derived, with `artifact` replaced by the real reference and the key removed, because the host has
never heard of that word. A resource naming an artifact nothing produced is refused.
So the two-document split this issue described as the shape of a fix **is the design**, and it
covers both cases it said were unsolved: an image the mesh builds, and an image somebody else
built. Mirroring also removes something worse than a stale pin — every machine needing a route to
a public registry, and a tag a stranger can move.
**What was actually wrong was the examples.** They hard-coded image references instead of naming
artifacts, so they inherited a problem the design does not have. Pinning twelve of them by hand
was treating the symptom, and left the reference pointing at a public registry rather than the
mesh's own.
## What is still open
**The examples should name upstream artifacts** rather than carry hand-pinned digests. That is the
remaining work, and it is a rewrite of five manifests rather than a mechanism to build.
The refusal added here stays: a placeholder must not reach a machine whatever the reason it is
there.
@@ -0,0 +1,96 @@
---
status: located
opened: 2026-09-01
located-in: [mesh-control]
fixed-by: partly — mesh-control 53eb000, withdrawn in 83c6a2f
amended-design:
---
# 026 — The data directories are mounted and never declared
## Symptom
Four modules mount **fourteen host paths** that no resource in those modules declares:
```
gitea /services/gitea/gitea
postgres /services/postgres/db-data
minio /services/minio/data/data1-1
mailu eleven more, including the mail spool and the admin database
```
Each is a bind mount on a container. None is a `directory` resource. The mesh has never heard of
any of them.
## What that costs
**They are created by the container runtime, as root.** A bind mount whose source does not exist
is created for you, owned by root, with whatever mode the runtime picks. So `owner` and `mode` —
which exist precisely so a module can say who its data belongs to — are silently not applied to
the only directories that hold data.
**The protection that exists for exactly this does not reach them.** A directory the mesh declared
and no longer wants is *kept*, not removed, when it holds anything the mesh did not put there
([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)). That rule is the
answer to *what happens to my data when a module goes away*, and it is written in terms of
declared directories. **An undeclared one is not protected by it, because the mesh does not know
it is there.**
So the single rule guarding against data loss covers the configuration directories, which are
cheap to lose, and not the data directories, which are the reason the rule exists.
## Where it came from
These manifests were written by reading the arrangement being replaced and carrying its
`docker-compose` files across — service, image, ports, volumes, environment — into the new
manifest's container shape. That shape can express all of it, which is what made the
transliteration feel like progress.
**A container shape that can express a compose file will be filled in like a compose file.** The
mesh's model is larger than that: a directory is a thing the mesh owns, with an owner and a mode
and a rule about what happens when it is no longer wanted. A volume line borrowed from compose
declares none of it, and nothing complains, because a bind mount source is a string.
## What a fix has to keep
- **Every host path a container mounts is declared.** If a module wants a directory on the
machine, it says so, with who owns it and what mode — and gets the removal rule with it.
- **The check is mechanical.** A person comparing volumes against declared directories by hand is
the process that produced this. It is a few lines against the manifest and belongs beside the
other manifest checks.
- **Not by inventing directories at apply time.** The host creating what a mount needs would make
the mesh's ownership of a directory depend on which resource mentioned it first, and would put
the same undeclared path back a layer down.
## Not yet answered
**Where a module's data should live at all.** These paths were inherited whole from the
arrangement being replaced, which put everything under one directory per service. Whether that is
right here is a separate question, and a bigger one — it decides what a person backs up, and what
survives a module being removed.
## Half fixed
**All fourteen are declared**, across the forge, the mail system, the store and the object store —
each mount now resolves to a `directory` or to a file the module already names.
**Enforcing it was tried and withdrawn**, and the withdrawal is the interesting half. A refusal
for any mount no resource declares refuses the **builder**, which mounts the container runtime's
socket. That socket is not the builder's data. It does not belong to the module, it already exists,
and declaring it as one of the module's own directories would be a lie that the host would act on.
So the rule is right about data and wrong about everything else, because the manifest cannot
currently say which a path is. Two kinds of mount are spelled identically:
- **the directory my data lives in** — created if absent, owned by the module, protected by
[ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)
- **a machine facility I was granted** — a socket, a device; it exists, the machine owns it, and
the module is being given access to it
`capabilities` is the closest existing thing to the second and names no paths. Inventing a field
to separate them is a design decision, so it is recorded here rather than made to get a check
green.
**Until then the manifests are right by coincidence**, which is the state this issue was opened
about. What is kept is a check that every real manifest still parses — worth nothing against this
fault, and the reason the next attempt finds out in a second rather than in a fifteen-minute run.
@@ -0,0 +1,65 @@
---
status: located
opened: 2026-09-01
located-in: [mesh-host, mesh-control]
fixed-by:
amended-design:
---
# 027 — A container cannot follow a file, and a rotated credential is the case
## Symptom
A service can say `restart-on`: *these files changed, so I must be restarted*. A container cannot.
It is not in the shape, and the host refuses a declaration that tries.
So a container reading its password from a file keeps the password it started with, for ever.
Nothing reports anything: the file is right, the container is up, every check passes.
## Why this is the same fault the mechanism exists for
`restart-on` is written against exactly this, in the host's own words:
> a running service does not re-read its configuration. Replace the file, find the service already
> running, do nothing, and the machine keeps behaving the way it did before — while every check
> passes, because the file is right and the service is up.
Every word applies to a container, and more so. **Nearly everything the mesh runs is a container**
— a database, a forge, a mail system — and a credential arrives as a file it reads at start.
## What it costs, concretely
**Rotation does not reach a container.** Rotating a credential replaces the file on the machine and
tells the provider to accept the new one. The provider is a program that reconciles, so it takes
the change. The consumer is usually a container, so it does not. The two ends then hold different
passwords, which is the fault the whole design is arranged to prevent
([ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) records it costing
two days).
The end-to-end test that proves rotation works uses a consumer that reads the file on each
attempt, so it does not meet this.
## Why it was not noticed
The gap is invisible from the control plane. A manifest carrying `restart-on` on a container
composes into a declaration without complaint and is refused on the machine, so the only way to
learn is to run one — which is how it was found, after nine of them had shipped across seven
modules.
## What a fix has to keep
- **Declared state, not a command.** `restart-on` is deliberately not *restart this*: it says the
running thing must reflect these files, and the host works out that it does not. Whatever
containers get must keep that shape, because the link may not carry an action
([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)).
- **Recreate, not restart, where that is the honest verb.** A container's environment is fixed at
creation. If what changed is an `env-file`, restarting the container is not enough — it has to be
made again. That is a different act from a service reload and should not be described as one.
- **It must not fire on every reconcile.** A container that is recreated whenever the host looks at
it is worse than one that never follows the file.
## The near alternative, and why it is not enough
A module can avoid the problem by having its program read the file on each use rather than at
start. That works for something written for this mesh, and is not available for a database, a
forge or a mail system — which is the whole population this is about.
@@ -0,0 +1,99 @@
---
status: fixed
opened: 2026-09-01
located-in: [mesh-control, mesh-host]
fixed-by: mesh-control 1f5b70a, 41f7c51; mesh-host b91342a
amended-design: 02-DECISIONS/0038-the-mesh-assigns-the-port.md
---
# 028 — Two things want one port, and nothing says so until the machine
## Symptom
The database module cannot start on a machine that runs the control plane:
```
Bind for 127.0.0.1:5432 failed: port is already allocated
```
The mesh keeps its own store on that machine, from the bundle, and it holds 5432. The module
publishes 5432 too. Everything up to the machine is content: it resolves, it composes, it is
pushed, and it is applied — the container is simply the one resource that fails.
## Why nothing catches it
**The substrate is not a module.** It arrives from the bundle a host carries, before there is a
mesh to ask. So the control plane has never heard of `mesh-store` and does not know it holds a
port. Resolution can compare modules against each other and cannot compare a module against the
thing the mesh is built on.
**And nothing compares modules against each other either.** A port is exclusive on a machine in
exactly the way a claim is — one seat, one display server, one artifact store — and the mesh has a
mechanism for that, which ports do not use. Two modules both publishing 5432 would meet the same
wall, one machine later.
## It has been met before, and worked around
The end-to-end test that exercises a real database publishes `5433:5432` rather than `5432:5432`.
The workaround is right there, inline, with no note saying why — which is how a constraint becomes
folklore.
## What a fix has to settle
- **Whether a module should publish to the machine at all.** Consumers reach a provider by the
machine's address and the port it *serves*, so publishing is what makes that true. An alternative
is that they reach it on the module's own network by name, and nothing is published — which
changes what `serves` means and is a larger decision than it looks.
- **Where the substrate's ports are written down.** Whatever compares them needs to know what the
bundle holds. The bundle is a list of pinned references; what those containers bind is not in it.
- **What a refusal should say.** *5432 is held by the mesh's own store on this machine* is a useful
sentence. *Port is already allocated*, arriving from a container runtime three layers down, is
not.
## Not the same as a firewall rule
`listens` already says which ports a module accepts on, and filtering is computed from it. That is
about what may reach a port from elsewhere. This is about two things on one machine wanting to own
the same one, which `listens` does not model and could not answer.
## Answered in principle
[ADR 0038](../../02-DECISIONS/0038-the-mesh-assigns-the-port.md), proposed the same day: **the mesh
assigns the machine-side port and a module does not care.** A module cannot choose well, because it
is written once and assigned anywhere — any number it picks is a guess about a machine it has never
seen.
A port fixed by its protocol — mail on 25, submission on 587 — becomes a **claim**, which is the
mechanism the mesh already has for what is singular on a machine. Two modules wanting 25 is the
same shape as two wanting the seat, and earns the same refusal at assignment rather than at apply.
The record also names what this issue missed: the same number is written **three times** in every
module — once for the rule set, once for what a consumer is told, once for what the runtime
publishes — and nothing checks that they agree. A module whose `serves` and whose container
disagreed would hand every consumer a port that answers nothing.
## Fixed
**The mesh assigns the machine-side port**, from a high unprivileged range, recorded per machine
and module and kept once chosen. A module says the port its software uses, once, in `listens`. The
container's mapping, the rule set, and what a consumer is told are all derived from the assignment
— so the three copies that agreed only because one person wrote them are now one fact.
**A port the protocol fixes says so**, and is then a claim: one holder per machine, and the second
refused by name at assignment rather than by a container runtime at apply.
**And the machine says what it already holds.** This was the half that made the issue: the
substrate is not a module, so nothing in the mesh had heard of the store or the broker. The host
already distinguished what it carried from what the mesh sent — that distinction exists so the two
never remove each other — and now records what each resource binds and reports the carried ones.
The allocator treats those as taken.
What the declaration binds, not what is open: a machine's open ports are a moving target, and
assigning around those would mean a port that was free when it was asked for and taken when it was
used.
## What it does not settle
The question underneath, unchanged: **whether a module should publish to the machine at all**.
Assignment makes publishing safe without making it necessary, and consumers reaching a provider on
the module's own network by name would make the question moot for anything inside the mesh.
@@ -0,0 +1,104 @@
---
status: fixed
opened: 2026-09-01
located-in: [mesh-control]
fixed-by: mesh-control be62f49; mesh-lab f85dbb0
amended-design:
---
# 029 — The artifact store cannot be delivered by the artifact store
## Symptom
A mesh that has just bootstrapped cannot install a registry. The module describing one resolves,
composes and pushes; the build never completes, because there is nowhere to put what it builds.
It has never been seen, because the lab always has a registry standing before the mesh asks for
one, and so does any mesh built on a machine that already had one.
## What it actually blocks
Not "a registry cannot be installed" — **a mesh cannot hold its own modules.**
The artifact store is one store for everything a module ships: images, and archives, which are
directories from a module's repository packed and pushed as content-addressed blobs. So it is the
mesh's module catalogue in artefact form, and the same role a separate object store plays in the
arrangement being replaced.
Until it exists, a module can be described and resolved but nothing it carries can be kept
anywhere. A mesh that has just bootstrapped can therefore run only what its bundle already holds.
## The cycle
Three facts, each correct on its own:
- **A registry module's image is mirrored in.** `kind: upstream` pulls the reference the module
names and pushes it under a name of the mesh's own, so what a machine fetches is pinned by a
digest this mesh assigned rather than by a tag somebody else can move.
- **The builder publishes to the artifact store**, learned from its `artifact-store` binding —
the same binding any consumer of any provision gets.
- **The builder refuses to run without one**: *"a built artifact nobody can fetch is not built."*
So installing the thing that provides `artifact-store` requires something that provides
`artifact-store`.
## Why the design already answers it
The substrate record asks of each candidate *can it grant itself the thing it provides?* The store
cannot create its own database; the broker cannot create its own virtual host; and the registry
**cannot grant itself a repository**. That is why the registry is substrate by role.
The same sentence answers this. A module that provides the artifact store cannot be delivered
through the artifact store, so its image is not the mesh's to mirror — it is named directly and
pulled from upstream exactly once, which is what the bundle already does for the three images a
first node starts from.
## The fix
**The registry module names its image, and is never built.** A container naming
`registry@sha256:…` needs no builder, no binding and no store. Every module after it mirrors
normally, into the registry now running.
Nothing new is required: naming an image directly is what most modules do.
## What it costs, and what it does not
The machine running the first registry needs to reach a public registry once, to pull that one
image by digest. That is already true of a first node, which fetches three images the same way
before a mesh exists.
It does not weaken pinning. A digest is exact wherever it came from; what mirroring adds is that
the mesh keeps its own copy and does not depend on a tag somebody else controls. For one image, on
one machine, once, the bundle already accepts that trade and says why.
## The constraint this puts on a registry module
**A module providing `artifact-store` may not build artifacts of its own** — not its image, and
not a user interface or a tool server shipped beside it. There is nowhere to put them until it is
running.
A registry that wants more than the upstream image is therefore two modules: one that provides the
store and only names an image, and an ordinary module beside it that builds whatever else and
mirrors it in the normal way. That is a real limit on how a registry module can be written, and it
should be said out loud rather than discovered.
## How it is checked
A manifest that both provides `artifact-store` and declares built artifacts is refused where it is
written, naming the cycle. Otherwise the fault surfaces as a build that never returns, on a mesh
too new to have anybody watching it.
## Fixed
**The registry module names its image and is added, not built.** The lab's artifact-store test now
walks the only path open to a real first mesh — the manifest goes in directly, no builder and no
store involved — and passes.
**And the cycle is refused where it is written.** A manifest that provides `artifact-store` and
also declares built artifacts is refused at parse, naming the cycle: building publishes to the
store, so it asks the mesh to put an artifact into the thing that artifact is needed to create.
The provision name became a constant for the rule to turn on.
What surfaced it in the lab is worth keeping: the test had always *built* the registry module and
passed, because the scenario's stand-in for the public registry was standing there to receive the
push. A prop that quietly covers for the thing under test is how a bootstrap hole stays invisible.
@@ -0,0 +1,56 @@
---
status: fixed
opened: 2026-09-01
located-in: [mesh-control]
fixed-by: mesh-control 38d4e77
amended-design:
---
# 030 — Asking what a machine should be re-signed its certificate
## Symptom
Every machine carrying a certificate reported as *waiting* — not running what the mesh would send
it — for ever. Pushed seconds ago, already behind again. Nothing wrong, nothing failed, nothing
quiet; only a comparison that never came out equal.
It surfaced as the one red test in four consecutive runs, and wore three other faults' clothes
first: a test racing the apply it asserted on, a status command that wrote to the database it was
reading, a machine starved at its default size. Each was real; each was fixed; the symptom stayed.
## Cause
The mesh signs a certificate for a machine's internal name as part of composing its declaration —
and signed **anew on every composition**. Same authority, same key, same name, same validity
window; a fresh random serial each time, because that is what signing does. So the declaration
composed to answer *is this machine current* differed from the declaration sent by exactly one
serial number, every time, deterministically.
The comparison is a digest, so one changed byte is as unequal as a different world.
## How it was found, which is the lesson
Not by deduction — deduction produced the three wrong theories above. The suite was run once with
its scenario kept standing, and the standing mesh was asked twice: `plan`, `plan`, diff. Two
answers seconds apart, identical to the byte but for one serial, in the certificate file. The
diff had one line where four theories had none.
A verdict machine that can be kept and interrogated is worth more than the verdict.
## The rule it broke, third find of its kind
**Issued once and kept.** The port had it, the secret had it, the certificate did not — composed
fresh on every asking, by the same code that holds the other two still. And like
[`028`](../028-two-things-want-one-port-and-nothing-says-so/00-report.md)'s ReleasePorts, the
keeping was designed and never wired: the serving-key migration added a column *"and what was
issued for it"*, and nothing wrote it.
A kept certificate stands while the name, the key and the clock agree. A node that rejoined with
a new key or changed its name gets a fresh signing exactly as if nothing were kept; so does one
whose certificate is into its last tenth of life.
## Verified
Live, on the kept mesh, before any suite run: one push with the fixed binary and the machine
settled; the second machine likewise; then the mesh's own sentence — all doing what they were
told, running what the mesh would send them.
@@ -0,0 +1,42 @@
---
status: open
opened: 2026-09-02
located-in: [mesh-host]
fixed-by:
amended-design:
---
# 031 — A machine becomes each thing it was told, in turn
## Symptom
A machine that was pushed several declarations in quick succession applies every one of them,
oldest first, at the better part of a minute each. Under the lab's suite — twenty-odd pushes in
fifteen minutes — the anchor machine ran minutes behind the newest declaration, and a test that
waited for it honestly timed out while the machine was busy becoming things nobody wanted any
more.
Only visible since caught-up became an equality: each report now names the declaration it applied,
and the reports arriving were about ever-older ones. Before that, the same backlog hid inside
timestamp comparisons that happened to pass.
## Why it is wrong, and why it is also right
Each declaration is complete — the whole machine, not a delta — so applying an old one is never
*incorrect*, only wasted: the machine converges to a state the mesh has already moved past, then
does it again. The queue keeps a disconnected machine's instructions safe, which is right. What
is wrong is only the order of consumption: **a machine asked to be five successive things should
become the last one.**
## The shape of a fix
On waking with a queue, drain it and apply only the newest declaration; acknowledge the
superseded ones without applying them. Whether a superseded declaration deserves a report — and
what its outcome should be called — is the real question for the link's vocabulary: silence reads
as a machine that ignored an instruction, and "applied" would be a lie.
## What it costs today
Nothing on a real mesh at rest: pushes are far apart. It costs the lab about a doubling of one
test's wait, and it will cost a real mesh exactly when things are busiest — a flurry of changes is
when a machine can least afford to replay history.
@@ -0,0 +1,135 @@
---
status: resolved
opened: 2026-09-04
located-in: [mesh-sdk, mesh-catalog]
fixed-by: mesh-sdk src/provisioner rework + redis/postgres/minio/umami adapters (ADR 0048)
amended-design: 0048-a-provider-creates-the-credential-the-mesh-minted.md
---
# A provider's provisioner seals with a key the mesh has no way to deliver — and does not need to
## What was observed
Building the vertical slice for the module runtime (the module runs its own code as its own
process under its own account), a **provider** module — one that stands up a per-consumer
resource and hands back a credential — was assigned to a node and run as a broker-bound
runtime. The runtime hosts the module's provisioner (the sdk's `runProvisioner`), and the
harness opens by reading a **seal key** from `$MESH_SEAL_KEY`, failing immediately without
one. Every credential it produces for a consumer is sealed to that key with the sdk's
symmetric `seal()` (AES-256-GCM, `mesh-sdk/src/primitives/index.ts`) before being written.
Nothing in the mesh sets `$MESH_SEAL_KEY`. It is read in exactly two places in the sdk and
set nowhere — no manifest, no control-plane code, no host code. So a provider runtime, as
delivered, aborts at start-up. The slice proved the mechanism only by setting a lab-local key
in the manifest by hand.
## What a trace of the credential path turned up
The seal key is not a missing delivery. **The whole symmetric-seal provisioner is orphaned,
and it duplicates — badly — a job the mesh already does.**
- `runProvisioner` reads request files named `*.grant.json`. **Nothing writes those.**
- It writes sealed credential files named `<consumer>.<resource>.credential`. **Nothing reads
those** — not the host, not the control plane. The host reports applied-resource digests
upward and never ships credentials; the control plane has no reference to that filename.
- No consumer ever calls the symmetric `unseal()`. Consumers receive **plaintext**.
Meanwhile the mesh already carries a provider→consumer credential across nodes, with **no
shared key anywhere**:
- The control plane mints the password once (`secrets.Make`) and seals it **twice,
asymmetrically** — `ForConsumer` to the consumer node's X25519 public key, `ForProvider` to
the provider node's (`mesh-control/internal/secrets/seal.go`, `mesh-host/internal/identity/
sealing.go`, NaCl box).
- Each host opens its own copy with its own private key on the machine; the plaintext exists
only for the length of one function call (`mesh-host/internal/apply/apply.go`, the
`${secret:name}` substitution — ADR 0024's "the host is the only thing that ever holds
both").
- `serves` carries no credential and says so; `receives`/`bound` tell each side *where* its
sealed secret is, never the value.
The two models also **contradict** each other. The sdk's `seal()` comment says the key is "a
per-node passphrase the host holds"; the host holds no such passphrase — it holds an X25519
private key, and the control plane's own code refuses a shared symmetric key on principle:
"a key both ends hold is a key the mesh would have to distribute, which is this problem again
one level down" (`secrets/seal.go`). A symmetric `MESH_SEAL_KEY` shared between a provider
node and a consumer node is exactly the thing the mesh was built not to have.
And the provisioner's model is wrong in a second way: its adapter **generates its own
password** (`generatePassword()`) and creates the resource with it — a different password from
the one the mesh mints and hands the consumer. Even with a seal key delivered, a consumer
would authenticate with the mesh's password against a resource created with the provisioner's.
## Why it matters beyond this instance
This is not a four-module problem. The provider contract lives in **one place** — the sdk's
`runProvisioner(resource, adapter)` harness — and every provider is built on it. Four exist
today (redis, postgres, minio, umami); a mesh of any size ends up with many. Whatever the
provisioner harness does, every present and future provider inherits, so the orphaned
symmetric seal is a fault stamped into the interface, not into four adapters. That also sets
the cost of getting it wrong: a contract N providers depend on is N migrations to change
later, which is the argument for settling it deliberately now rather than patching around it.
As written, each provider carries a provisioner that cannot start (no key), and that, if it
did, would create resources with a password it invented — a *different* password from the one
the mesh minted and handed the consumer — and seal them for a reader that does not exist. The
rule the design states, "a consumer receives a sealed credential and unseals it," is enforced
by nothing: no consumer unseals, and no shared key exists to unseal with.
## The mesh already does this — confirmed
The premise the fix rests on is not a hope; it is in the control plane today. For a served
interface, `Inventory.SecretFor` mints one password per (consumer, provider) pair via
`secrets.Make`, sealing it to **both** node keys — `ForConsumer` and `ForProvider`.
`SecretsFrom(provider)` is documented as "every credential a provider node was issued, so it
can be told what to create," and `grantsFor` (plan.go) hands the provider node one `Grant` per
consumer carrying `Sealed: ForProvider`. The provider receives, at the path its `receives`
names, one `Contribution` per consumer: the login to create (`As`, derived by the mesh so both
ends agree — 04-ISSUES/023), the consumer's address (`At`) and requested `Values`, and
`Secret`, the file holding that consumer's password sealed to this provider and unsealed by
its host. Everything the provisioner needs is delivered. It reads the wrong files
(`*.grant.json`, which nothing writes) and invents a password instead of reading the one in
`Secret`.
## The fix this points to
A **one-place contract change in the sdk harness**, plus re-pointing today's adapters at it —
not per-provider surgery, and inherited correctly by every provider after them:
- `runProvisioner` reconciles the mesh-delivered `receives` contributions (not `*.grant.json`):
for each consumer, create the resource under the login `As` with the password read from the
delivered `Secret` file, for its `Values`; withdraw the login when a consumer leaves the file.
- The adapter stops generating a password and stops returning a credential — it is handed the
name and the password and only makes the resource exist. Roughly `create({as, password,
values})` / `remove({as})`, no return.
- `sealKey`, `seal()`, `writeSealedCredential`, `MESH_SEAL_KEY`, and the `.credential` file
leave entirely; the consumer already receives its copy through the mesh's own channel.
This is proposed as ADR 0048, which defines the corrected provider contract, for ratification.
## Resolution
ADR 0048 was accepted and implemented on the branches this issue is fixed by:
- `mesh-sdk` `src/provisioner/index.ts` now reconciles the mesh's `receives` contributions and,
per consumer, reads the mesh-minted password from the file the host unsealed, calling the
adapter to create the resource under the mesh's login. `$MESH_SEAL_KEY`, the symmetric seal,
`writeSealedCredential`, and the `*.grant.json` / `*.credential` files are gone. The symmetric
`seal()`/`unseal()` primitive had no other caller and was removed.
- The four adapters (redis, postgres, minio, umami) were re-pointed at the new contract —
`create({ as, password, values })` / `remove({ as })`, returning nothing. minio's client gained
a secret-key argument so it sets the mesh's secret rather than generating one.
- Proven in the mesh-lab: `provider-uses-mesh-credential` is green — redis creates the consumer's
login with the password the mesh minted, a client authenticates as that consumer and gets PONG,
with no seal key set anywhere.
Two things were carved out deliberately, neither blocking:
- **Data provisions are a separate shape.** umami's `analytics` returns a `siteId` umami
*generates*, not a secret the mesh mints, and a contract that returns nothing cannot hand that
back. ADR 0048 is scoped to credential provisions and says so; the provider→consumer return
path for generated data is left to a separate decision. umami compiles and reconciles under the
new harness; only that return is unaddressed, and it never had the seal-key fault.
- **Teardown beyond "remove the login"** — an object store's leftover data — is each adapter's to
name (minio leaves a non-empty bucket for an operator rather than deleting a consumer's data),
not the harness's.
@@ -0,0 +1,67 @@
---
status: open
opened: 2026-09-04
located-in: []
fixed-by:
amended-design:
---
# Changing a module's settings does not restart its runtime — config is stale until recreated
## What was observed
Rolling the module runtime out to the catalogue (the runtime that serves a module's tools and
runs its events under the module's own account), each tools+events module receives its
configuration the way the design intends: a mergeable config file the module declares, into
which the assignment's settings are merged. The runtime container mounts that file and reads
it once at start-up, when it builds its API client.
The design for settings says a config file a module owns can be changed **without editing
it** — a person states an intention, the file is regenerated, and the change takes effect.
The decision that config is the assignment's, not the manifest's, is explicitly so that
configuration can be updated *on the fly* and managed from a dashboard.
For a runtime delivered as a **container**, that last part does not hold. When settings
change, the control plane re-renders the config file on the node — but the runtime container
is only ever recreated when its **spec** changes, and the spec is image, name, env, ports,
volumes and args. The *content* of a mounted file is not part of it. So the file on disk
updates and the process that already read it keeps the value it read at start-up. The new
configuration does not take effect until something changes the container's spec, or it is
recreated by hand.
A **service** resource has `restart-on`, which names the resources whose change forces a
restart — exactly this problem, already solved, for units. A **container** resource has no
equivalent field, and the apply path for containers never consults the set of resources that
changed this pass. So the one kind of resource that hosts a module's runtime is the kind that
cannot say "restart me when my config changes."
The effect is quiet, which is the worst part: setting a value appears to succeed (the file is
correct on disk), and the running tools keep answering with the old configuration, or keep
failing to load because the value that would fix them is present but unread.
## Why it matters beyond this instance
Every tools+events module converted to the runtime model now takes its URL and credentials
this way, so this is not one module's quirk — it is the config path for the whole catalogue.
The gap turns the headline promise of the settings design ("change it without editing it, on
the fly") into "change it, then recreate the container by hand," which is the manual step the
design existed to remove. And because the file is genuinely updated, nothing surfaces the
staleness; a dashboard that set the value would report success while the mesh kept doing the
old thing.
Config set **before** the runtime first starts (settings, then assign, then push) does work —
the file is right when the process reads it. So the gap is specifically about *updates* to an
already-running runtime, which is precisely the case the "on the fly" promise is about.
## Open questions
- Should a `container` gain `restart-on`, mirroring the service field, so a module can point
it at its config resource?
- Or should the apply path recreate a container when a file it mounts changed this pass —
making mounted-file content behave like part of the spec, without a new field to declare?
- Should the config file's content (or a hash of it) fold into the container spec, so an
ordinary spec-diff already catches it? That restarts on every change with no new mechanism,
at the cost of a spec that is no longer only the container's own declaration.
- Is a restart even the right primitive for a runtime that could instead watch its config
file and rebuild its clients in place — and if so, is that each module's job or the
runtime host's?
@@ -0,0 +1,77 @@
---
status: resolved
opened: 2026-09-05
located-in: [mesh-control, mesh-catalog]
fixed-by: ADR 0049 (a slug for the login) + a shorter minted secret (mesh-control)
amended-design: 0049-a-consumers-identity-fits-the-tightest-backend.md
---
# The mesh's derived login does not fit every backend's identity rules — S3 rejects it
## What was observed
Proving the provider/consumer contract per backend (ADR 0048), redis and postgres passed: a
consumer authenticated against the provider with the login the mesh derived and the password
the mesh minted. **minio failed**, and not on the credential — on the *name*:
```
mc: <ERROR> Unable to add a new service account. The access key is invalid.
(access key length should be between 3 and 20).
```
The mesh derives a consumer's login as `mesh_<node>_<module>` — here `mesh_anchor_bucketuser`,
22 characters. That is a valid postgres role and a valid redis ACL user, so those providers
create it verbatim. S3 access keys are capped at **20 characters**, so minio refuses to create
the service account under it, and the provisioner retries forever while the consumer, holding
that same too-long access key, could never present it either.
## Why it matters beyond this instance
ADR 0048 says a provider creates *exactly* the login the mesh derived, so that the two ends
agree by construction — the mesh hands the same name to the provider (to create) and the
consumer (to present). That only holds if the derived name is one every provider can accept.
It is not: the mesh's `as` is a single format with no knowledge of a backend's identity rules,
and S3's are stricter than a database's. Any provider whose backend constrains identifiers more
tightly than postgres — a length cap, a charset, a required prefix — inherits this, and the
failure lands at provision time, per consumer, as an infinite retry rather than a refusal at
assignment.
This also shows the seam is real, not cosmetic: `as` is doing two jobs — a stable per-consumer
identity the two ends must agree on, and a literal identifier a specific backend must accept —
and those are not always the same string.
## The shape of a fix (open, not decided)
- **Constrain the derivation** so `as` is broadly acceptable — short (≤ 20), a conservative
charset, deterministic. This keeps "the provider creates exactly what the mesh derived" true
everywhere, at the cost of a less legible name, and it is a mesh-wide identity change (every
provider that already created the longer name would see it change).
- **Let a provider map `as` to a backend-valid identifier** it derives the same way on create
and on the consumer's behalf — but the consumer is generic and cannot run minio's mapping, so
this only works if the mapped identifier is *delivered back* to the consumer. That is the
data-provision return path this era keeps meeting (umami's siteId, cloudflare's record) and
does not yet have.
- **Declare the constraint on the interface** (`s3-bucket` states its identifier bounds) and
have the mesh derive within them — the most honest, the most work.
## Open questions
- Is `as` meant to be human-legible, or is a short opaque token acceptable — i.e., can the
derivation simply be shortened without anyone minding?
- Do redis/postgres actually want the long name, or did it only survive because they are
permissive? If nothing needs it long, the cheap fix is to cap it.
- Does this fold into the same decision as the data-provision return path, or is it separate?
## Resolution
Accepted **ADR 0049** (option E): a module declares an optional short `slug`, and the mesh derives
`mesh_<node>_<slug|name>`, bounded by the tightest backend (an S3 access key's 20) and refused at
assignment — naming the slug as the remedy — when it still would not fit. The minio grant e2e proved
it: `bucketuser` declares `slug: bkt`, so its access key `mesh_anchor_bkt` (15) is accepted where
`mesh_anchor_bucketuser` (22) was refused.
Proving that surfaced a **second S3 length constraint on the same credential** — the secret. The
mesh minted a 43-character password (32 random bytes, base64url), and an S3 secret key is 8–40. Fixed
in `mesh-control` `internal/secrets/seal.go` by minting 30 bytes → exactly 40 characters (240 bits,
ample), which fits S3 and every other backend. Both halves of an S3 credential — the access key
(login) and the secret key (password) — now fit the tightest backend, by the same rule.
@@ -0,0 +1,48 @@
---
status: open
opened: 2026-09-02
located-in: []
fixed-by:
amended-design:
---
# 035 — Reconciling a seed file wipes what grew in it
## The symptom, as observed
Found by review of the catalogue examples (2026-09-02), not by an outage — the outage is the
part the design permits to be silent.
The cache module declares its access-control file as an ordinary file resource with fixed,
empty content. The program that consumes the file requires it to exist at startup, which is
why the manifest declares it at all. But the same file is the one the provisioner writes
consumer users into, and the one the running program persists ACL changes back to.
A declaration is complete for what the host owns, and the host reconciles what is declared
([ADR 0010](../../02-DECISIONS/0010-delivery.md)).
So every apply that revisits this resource restores the declared content — empty — behind the
running program. Every consumer credential granted since the last apply is removed, the apply
reports success, and nothing anywhere says a grant vanished.
## Why it matters beyond the instance
The manifest needed *the file to exist before first start*, and the only vocabulary available
was *the file has this content, forever*. Those are different intentions, and the gap between
them is generic: any resource that a module seeds and something else then legitimately mutates
— an ACL file, a bootstrap configuration a program rewrites, an htpasswd a provisioner appends
to — has the same two owners and the same silent loss on reconcile.
It is also the mirror image of the boundary ADR 0010 draws so carefully on the *removal* side:
the host never removes what it did not create, but it happily overwrites what it *did* create,
even when what grew inside since is somebody else's work the mesh asked for.
## Open questions
- Is the missing thing a create-once file semantic ("present with this content if absent,
untouched otherwise"), or is the real fault that two owners share one file — and the
provisioner, not the declaration, should own it entirely, with first-start ordering solved
some other way?
- ADR 0010 treats every added resource type as a security artefact. Does a create-once
semantic widen what a compromised control plane can express, or narrow it?
- Are there other seeded-then-mutated files already in the catalogue that this failure is
waiting inside?
@@ -0,0 +1,51 @@
---
status: located
opened: 2026-09-02
located-in: [mesh-control, mesh-catalog, mesh-host]
fixed-by: 02-DECISIONS/0051-shared-data-is-the-operators.md
amended-design: 02-DECISIONS/0051-shared-data-is-the-operators.md
---
# 036 — Six modules own what they must share
## The symptom, as observed
Found by review of the catalogue examples (2026-09-02). The media stack is several modules —
a library server, the acquisition managers, a download client and their satellites — and each
of them declares the same library and download directories as its own resources.
The resolver refuses two modules that declare one path on one node, with no exemption for
identical content and no merge. That rule is right in general: two owners of one path is the
class of fault this repository keeps recording. But sharing those directories on one machine
is the entire point of this stack — the download client and the managers must see the same
downloads, the library server must see the same libraries. So the set, as written, refuses
its own only sensible assignment.
No test co-resolves any two of them, which is why the manifests pass today. The first machine
to be assigned the stack together is where the refusal would have surfaced.
## Why it matters beyond the instance
The manifests can express *a directory I own* and nothing else, so a directory that is the
shared workspace of several modules was written six times as six private ones. The intention
— several modules, one filesystem contract between them — has no vocabulary, and this is not
a media-stack peculiarity: any pipeline of modules handing files to each other on one machine
(an ingest directory, a spool, a drop folder) hits the same wall.
It is also a fork in the design the catalogue has otherwise avoided: the fix could be a new
owning module the others depend on, a shared-resource concept in the manifest, or a statement
that co-located file handoff is not a thing the mesh supports and these modules are one
module. Each answer changes what a module *is*, which is why this is an issue and not a patch.
## Open questions
- Is the unit wrong — is a stack that must share a filesystem one module with several
containers, the way the mail module already is?
- If it stays several modules: does one of them own the directories and the rest require
them, and is *requiring a directory from a neighbour* a provision, a claim, or a third
thing?
- The duplicate-path rule protects against genuinely rivalrous owners. Whatever expresses
sharing must not weaken it for the cases where refusal is the right answer — what
distinguishes the two, machine-checkably?
- The mesh's own rule is that a rule states how it is checked: whichever shape is chosen,
what test co-resolves the stack so this class of refusal is caught before a machine is?
@@ -0,0 +1,57 @@
---
status: located
opened: 2026-09-05
located-in: [mesh-control, mesh-host, mesh-catalog]
fixed-by: 02-DECISIONS/0052-a-step-that-runs-once-before-a-container.md
amended-design: 02-DECISIONS/0052-a-step-that-runs-once-before-a-container.md
---
# 037 — A module cannot run its own code at a lifecycle phase
## The symptom, as observed
Found while converting the catalogue (2026-09-05), across several modules at once. A module can
declare *things that exist* — a directory, a file with fixed content, a network, a container — but
it cannot declare *a step that runs* at a defined point in its own lifecycle. Three converted
modules need exactly that and have nowhere to put it:
- **mosquitto.** Its Dynamic Security plugin will not start unless `dynamic-security.json` already
contains an admin client *before the broker's first start* — the broker loads the plugin at
boot. Seeding it is a run-once step that must happen after the file resource exists and before
the container starts. The vocabulary has no "before first start."
- **The database providers (postgres/mongodb/mssql).** First-boot seeding works today only because
the *image* happens to do it from an env var. Anything the mesh itself must run once against the
server — a schema migration, an extension enable, a health gate before the module is announced
ready — has no home.
- The seed-then-mutate family already recorded in [035](../035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md)
is the same shape seen from the *content* side; this is it seen from the *timing* side.
## Why it matters beyond the instance
This is not a defect in a module — it is a **capability the module system does not yet offer.** A
real class of modules needs to run their own code at points in the build/install/run lifecycle:
seed-before-start, migrate, post-start health-gate, pre-remove drain. The declarative resource
model deliberately describes *state*, not *steps*, and that is right for what it covers; the gap is
that some modules genuinely have a step.
**Prior art, and its warning.** An earlier mesh had exactly this as a feature: event-driven
**hooks** that ran custom code at phases of the build/publish/deploy pipeline. It was powerful and
it was **complex to set up and flaky** — which is the real content of this record. The need is not
in question; the cost of the obvious answer is. Whatever shape this takes must not reproduce that
fragility, or it will be worse than the gap.
## Open questions
- Is the right unit narrow — a **run-once / init resource** ("run this once, here, in the
lifecycle") — or general — a **per-phase lifecycle hook** on a module, and if so which phases
(build / publish / install / pre-start / post-start / pre-remove)?
- Where does a hook's code run — in the module's own runtime container under its scoped account
(ADR 0043/0047), so it inherits the same isolation as its tools and events? Or is some of it the
host's, before a container exists?
- How is a step made **idempotent and reconcilable** so a re-apply does not re-run it
destructively — the same discipline the resource model gets for free and a step does not?
- What is the smallest version that unblocks the three modules above without rebuilding the old
flaky hook engine? Is "seed-before-first-start" alone enough for now, with the general case
deferred?
- A rule states how it is checked: whatever shape is chosen, what lab scenario proves a hook runs
exactly once, at the right phase, and converges on re-apply?
@@ -0,0 +1,91 @@
---
status: resolved
opened: 2026-09-09
located-in: [mesh-control]
fixed-by: mesh-control — a same-node provider is announced at the port it is published on
amended-design:
---
# 038 — A provider is announced at a name its port is not bound to
## Symptom
A module that provides a `from: mesh` provision (observed with the database provider) is
announced to its consumers, by [issue 018](../018-a-provider-on-the-same-machine-was-never-announced/00-report.md)'s
fix, at the node's private-network name — the binding a consumer reads carries
`at: <node>.internal` and `serves.port: 5432`.
But the provider's container port is **published bound to loopback** (`127.0.0.1:<assigned>`),
not to the address `<node>.internal` resolves to. So every consumer dials the announced
`<node>.internal:5432`, which resolves to the node's private-network address, where **nothing is
listening** — the port is open only on `127.0.0.1`.
Observed on a node hosting the provider and several consumers:
- The consumer's binding file says `"at": "<node>.internal"`, `"serves": { "port": 5432 }`.
- Inside a consumer container, that name resolves to the node's private-network address.
- A connection test from the node: the private-network address on port 5432 is **CLOSED**; only
`127.0.0.1` (on the assigned host port) is OPEN.
- Consumers that touch the database only lazily serve a landing page and *look* healthy; consumers
that require the database at startup crash-loop — one with "acquisition timeout while waiting for
a new connection", another connecting and then timing out on its first query.
- The provider itself is healthy: a direct client on loopback answers instantly, few connections,
no locks.
## Why this matters
The announced address and the actual listener disagree, so the binding is a promise the mesh does
not keep. It is not one module's misconfiguration: it is the port-publishing step choosing a bind
address that does not match the `at` the resolver hands consumers, so it fails the same way for
**every** `from: mesh` provider with an off-node-reachable consumer — and, on a single node, for
same-node consumers too.
It hides well. The provider is up, the credential is correct, the database exists, a manual client
works — every part a person checks in isolation passes. Only a consumer that must use the provision
before it can serve anything reveals it, and it reveals it as a timeout, which reads as slowness or
load rather than "the address was never listening". A mesh that co-locates a provider with its
consumers (the ordinary small-mesh case) is exactly where it bites.
It also blocks anything that must *reach* a routed/served name from inside the mesh, not just
application traffic — see the internal-CA validation dependency noted in the connectivity design.
## Diagnosis
The symptom's first reading — "published on loopback" — was **partly a red herring**. Two things
were tangled:
1. **The real, current-code defect is a served-*port* mismatch, not a bind address.** A bare
`ports: ["5432"]` is assigned a host port and published as `"15432:5432"` — no bind IP, so on
**all interfaces**, reachable at the node's private-network address. But the port a consumer is
*told* is only re-derived from the assignment on the **cross-node** path. The **same-node** paths
(the resolver's `servedHere`, and the `here()` fallback) settle their served facts *while
resolving* — before the host port is assigned — so they carry the **declared** port (5432), not
the **assigned** one (15432). A co-located consumer is therefore announced
`<node>.internal:5432` while the provider is published on `<node>.internal:15432`, and dials a
port nothing listens on. Cross-node consumers were always fine, which is why it read as "the
small-mesh case."
2. **The `127.0.0.1:15432` seen in the running lab was a stale build.** Current code's publish step
binds all interfaces; the running instance was raised from a mesh-control predating the ADR 0038
publish rewrite. The substrate's own store *is* deliberately `127.0.0.1:5432` (a private store
must not be exposed) — correct, and not this bug.
## Fixed by
`mesh-control` branch `fix/same-node-provider-announced-port` (`c147a26`): after the host port is
assigned, same-node needs (and the `here()` fallback) are redirected through the same
provision→module→assigned-port lookup the cross-node path already uses, so a co-located consumer is
announced the port that is actually published. Idempotent (keyed by the declared port). Regression
test `TestASameNodeProviderIsAnnouncedAtThePortItIsPublishedOn` asserts the announced port equals
the published host port for a co-located provider/consumer — the next assertion after 018's (which
only checked a binding file exists); verified failing without the change.
*Not yet merged, and the running lab is additionally stale — proving it end-to-end there needs
mesh-control rebuilt and the affected consumer containers recreated.*
## Noted, not taken
Binding the assigned port to the node's private-network address specifically (rather than all
interfaces) would be defence-in-depth and would make the publish address match `at` by construction
— but it is a larger behavioural change entangled with the unenforced firewall scope
([003](../003-firewall-scope-is-read-by-no-code/00-report.md)), so it is left as an option.