Merge branch 'issue/021-provider-port-published-on-loopback' into design/bootstrap-is-a-pivot
# Conflicts: # 03-DESIGN/01-to-be/04-lab-installation.md
This commit is contained in:
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: open
|
||||
status: resolved
|
||||
opened: 2026-08-22
|
||||
located-in: []
|
||||
fixed-by:
|
||||
located-in: [mesh-host]
|
||||
fixed-by: mesh-host — a package is read back from the package database after installing
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -31,14 +31,14 @@ exists to catch — a step that failed, reported success, and left the next step
|
||||
state that was never produced.
|
||||
|
||||
It is also a direct violation of a decision already taken and recorded:
|
||||
[ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) says a step that fails must fail the
|
||||
[ADR 0010](../../02-DECISIONS/0010-delivery.md) says a step that fails must fail the
|
||||
job. That record notes the rule is applied instance by instance and enforced by no mechanism.
|
||||
This is an instance where it was never applied.
|
||||
|
||||
## Evidence
|
||||
|
||||
- Observed 2026-08-22 while declaring the virtualisation package required by
|
||||
[ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md).
|
||||
[ADR 0016](../../02-DECISIONS/0016-the-lab.md).
|
||||
- A fix is written and open as a pull request, unmerged since 2026-08-20.
|
||||
|
||||
## Open questions
|
||||
@@ -47,3 +47,18 @@ This is an instance where it was never applied.
|
||||
- Is this specific to package installation, or does the surrounding stage swallow every
|
||||
non-zero exit?
|
||||
- The fix has been open for two days. What is the review path for a change of this class?
|
||||
|
||||
## How it is answered
|
||||
|
||||
*2026-08-31.* **The host reads the package database back after installing**, and refuses when it
|
||||
does not have the package:
|
||||
|
||||
> `<name> was installed without error and the package database does not have it`
|
||||
|
||||
That is the general rule this issue is one instance of, and the host applies it to everything it
|
||||
does: a command exiting zero says a transaction was *accepted*, not that the machine changed. The
|
||||
same read-back is why a container that starts and immediately dies fails an apply, and why a
|
||||
service asked to run is checked rather than assumed.
|
||||
|
||||
HAL keeps the fault until its provisioning is switched off. Fixing it there would mean
|
||||
implementing the read-back twice, in the system being replaced.
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: open
|
||||
status: resolved
|
||||
opened: 2026-08-22
|
||||
located-in: []
|
||||
fixed-by:
|
||||
located-in: [mesh-host]
|
||||
fixed-by: mesh-host — a stale package index is named rather than reported as a failed install
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -38,3 +38,29 @@ the job is green.
|
||||
versions — or is an index sync part of the install step?
|
||||
- A partial sync is unsafe on the platform in use; a full upgrade is the only sanctioned fix.
|
||||
Does that make index freshness a scheduled node concern rather than a pipeline one?
|
||||
|
||||
## How it is answered
|
||||
|
||||
*2026-08-31. It was present in the replacement too, which is why this is a fix rather than a note
|
||||
saying the new mesh does not have it.*
|
||||
|
||||
**The failure is named.** A machine asking for a version the mirrors have replaced now says so, and
|
||||
says what fixes it — a full upgrade of the machine.
|
||||
|
||||
**It is deliberately not fixed by synchronising.** `pacman -Sy <pkg>` installs a package built
|
||||
against libraries the machine does not have: a partial upgrade, unsupported on this distribution,
|
||||
which surfaces much later as something apparently unrelated. That is a decision about the whole
|
||||
machine, and a host that made it silently while applying one resource would be taking a large
|
||||
decision in a small place.
|
||||
|
||||
So the host distinguishes the two cases and leaves the decision where it belongs. **A declaration
|
||||
that is wrong and a machine that is out of date fail identically otherwise, and they are fixed in
|
||||
completely different places.**
|
||||
|
||||
**And the package manager's own words were being thrown away** — the output was read into `_`, so
|
||||
the 404s that name the cause never reached anybody. Whatever it said is now part of the failure,
|
||||
which is the rule everywhere else here and was not being followed in the one place where the reason
|
||||
exists only in the output.
|
||||
|
||||
*Checked by a stale-index failure being named as one, an ordinary missing package not being, a
|
||||
single mirror timing out not being, and a successful install still saying nothing.*
|
||||
|
||||
@@ -1,9 +1,9 @@
|
||||
---
|
||||
status: open
|
||||
status: resolved
|
||||
opened: 2026-08-22
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
located-in: [mesh-control]
|
||||
fixed-by: mesh-control — a machine's filtering is computed from what it was assigned
|
||||
amended-design: 03-DESIGN/01-to-be/08-connectivity.md
|
||||
---
|
||||
|
||||
# 003 — A firewall rule's `scope:` is read by no code
|
||||
@@ -38,3 +38,29 @@ any check.
|
||||
- Were the five declarations intended to restrict something that is currently open? Each needs
|
||||
checking against what the node actually exposes — the declaration cannot be trusted either
|
||||
way.
|
||||
|
||||
## How it is answered
|
||||
|
||||
*2026-08-31.* Both halves, in Novox Mesh. HAL keeps the fault until its provisioning is switched
|
||||
off, which is what this issue is now waiting on rather than a fix of its own — patching `scope:`
|
||||
into something that works would mean implementing it twice, in the system being replaced.
|
||||
|
||||
**The unknown key.** A manifest is parsed strictly: an unknown key is refused with the key named,
|
||||
the discipline the host's declaration parser has always had. `scope:` would not survive being
|
||||
written today, and neither would a misspelling of anything else. This is the general fix — the
|
||||
issue's own observation was that *any* invented key behaved this way, and that the one instance
|
||||
was found by reading rather than by any check.
|
||||
|
||||
**The rule that restricts nothing.** `scope:` is not reimplemented. A module says what it listens
|
||||
on and **who may reach it**, and saying from where is required rather than defaulted: a rule with
|
||||
no source is open, and must say so rather than appear to restrict something. A machine's whole
|
||||
rule set is then derived from every module assigned to it — so there is no second list to keep in
|
||||
step, which is the condition that let the first one drift out of use unnoticed.
|
||||
|
||||
**And it is enforced, which is the part that makes this different from before.** The mesh renders
|
||||
the rule set; a service on the node is declared to reflect that file, so replacing it restarts
|
||||
what loads it. Proven in the lab against two real ports on a real machine: the declared one
|
||||
answers from another machine, the undeclared one does not, and removing the module that wanted the
|
||||
port closes it with nobody editing a rule.
|
||||
|
||||
The design is [`03-DESIGN/01-to-be/08-connectivity.md`](../../03-DESIGN/01-to-be/08-connectivity.md) §4.
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: open
|
||||
status: resolved
|
||||
opened: 2026-08-22
|
||||
located-in: []
|
||||
fixed-by:
|
||||
located-in: [hal]
|
||||
fixed-by: hal — ACME_CA_SERVER selects the authority, and defaults to staging
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -21,7 +21,7 @@ recoverable by retrying — it removes the ability to issue a certificate anyone
|
||||
|
||||
The consequence lands hardest on exactly the work most likely to iterate: standing up a new
|
||||
node, changing how names resolve, or testing the lab's certificate authority split
|
||||
([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)).
|
||||
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
|
||||
|
||||
## Evidence
|
||||
|
||||
@@ -35,3 +35,41 @@ node, changing how names resolve, or testing the lab's certificate authority spl
|
||||
- The lab issues its own certificates and so does not consume public quota at all. Does that
|
||||
make this a problem only for experiments run outside the lab, and therefore an argument for
|
||||
running them inside it?
|
||||
|
||||
## Resolution
|
||||
|
||||
*2026-08-31.* **The authority is now selectable, and the default is staging.**
|
||||
|
||||
Confirmed first: the resolver declared no `caServer` at all, so the client fell to its built-in
|
||||
default of the public authority's production endpoint. There was no setting to change — not a
|
||||
setting set wrongly.
|
||||
|
||||
`ACME_CA_SERVER` now names the authority, and **defaults to staging**. That answers the first
|
||||
open question in the affirmative: it is a node property, and the node that serves real traffic is
|
||||
the one that states so.
|
||||
|
||||
**Why staging is the default rather than the safe-looking alternative.** Defaulting to production
|
||||
and documenting the override would leave the safe path depending on somebody remembering to opt
|
||||
out of it — on exactly the work most likely to iterate. That is the same fault as
|
||||
[005](../005-pipeline-test-harness-unbuildable/00-report.md): *remembering is not a mechanism.*
|
||||
A staging certificate is trusted by no browser, so the mistake announces itself in the first
|
||||
request rather than a fortnight later when the quota is gone. **The failure that is loud and
|
||||
immediate is the cheaper one**, and quota exhaustion is neither.
|
||||
|
||||
**The second open question is answered too, and it is not the whole answer.** The lab issues its
|
||||
own certificates and consumes no public quota, so experiments belong there. But "run it in the
|
||||
lab" is advice, and the nodes this issue is about are the ones outside it — the default is what
|
||||
protects those.
|
||||
|
||||
### The rollout is ordered, and the order is the dangerous part
|
||||
|
||||
Both public-serving nodes were checked: neither set the variable. Applying the change without
|
||||
pinning them first would re-issue their public certificates from an untrusted authority and break
|
||||
TLS for every hosted service — **this issue's own failure mode, arriving through its fix.** Pin
|
||||
first, then merge. The commit carries the exact commands.
|
||||
|
||||
## Deliberately not done
|
||||
|
||||
**Nothing was changed on a running node.** Pinning the public nodes and regenerating their
|
||||
environment restarts the reverse proxy that fronts every hosted service, and that is an operator's
|
||||
decision rather than a fix's.
|
||||
|
||||
@@ -1,9 +1,9 @@
|
||||
---
|
||||
status: open
|
||||
status: resolved
|
||||
opened: 2026-08-22
|
||||
located-in: [hal]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
located-in: [hal, mesh-lab]
|
||||
fixed-by: mesh-lab — a run leaves a receipt, and the receipt says what it covered
|
||||
amended-design: 03-DESIGN/01-to-be/01-end-to-end-testing.md
|
||||
---
|
||||
|
||||
# 005 — The end-to-end pipeline harness has not built since the workspace was removed
|
||||
@@ -29,7 +29,7 @@ coverage was assumed, not checked.
|
||||
## Evidence
|
||||
|
||||
- The workspace was removed by pull request #240 on 2026-06-04
|
||||
([ADR 0007](../../02-DECISIONS/0007-no-npm-workspace.md)).
|
||||
([ADR 0014](../../02-DECISIONS/0014-no-npm-workspace.md)).
|
||||
- The harness has not built since that date.
|
||||
- Recorded in the knowledge base as a standing entry, not as a fixed incident.
|
||||
|
||||
@@ -45,3 +45,52 @@ unmentioned: until the lab exists, this is the coverage the pipeline is presumed
|
||||
- Repair, or retire in favour of the lab? Leaving it in the repository unbuilt is the one
|
||||
option that keeps the false impression of coverage.
|
||||
- Was anything relying on it, or had it already stopped running before the workspace removal?
|
||||
|
||||
## Resolution
|
||||
|
||||
*2026-08-31.* **Retired in favour of the lab, and the reason it went unnoticed was fixed
|
||||
separately from the harness itself.**
|
||||
|
||||
The old harness is not repaired. What replaced it is the end-to-end suite on a lab mesh, which
|
||||
raises real machines and proves the pipeline against them. That answers the first open question.
|
||||
|
||||
The second finding is the one worth keeping. *Nothing runs it, and nothing reports that nothing
|
||||
runs it* is not a fact about that harness — it is a fact about **any** suite too expensive to run
|
||||
on every push, and the lab suite is exactly that: it needs a machine with a hypervisor, so it runs
|
||||
when somebody remembers. **Remembering is not a mechanism**, and the replacement inherited the
|
||||
fault it was replacing.
|
||||
|
||||
So three things now hold, each checked by a test that was confirmed to fail without it:
|
||||
|
||||
- **A run leaves a receipt** — when it ran, what passed, and the commit each repository was at.
|
||||
Kept outside version control, because the question is *has this machine run it*, and a receipt in
|
||||
git would be a claim about everybody's machine made by whoever committed last.
|
||||
- **The receipt can be judged, and says why it does not count.** Old, failed, taken against code
|
||||
the repositories have since moved past, or a run that never raised a machine — each reads
|
||||
differently, and only the last of those is new. **A receipt that says nothing about something is
|
||||
not a receipt that clears it.**
|
||||
- **The artifacts are rebuilt by the run, not beside it.** The suite consumes three artifacts from
|
||||
two repositories. They were rebuilt by hand, from memory, and a rename that needed two of them
|
||||
got one — leaving a binary eleven hours old refusing a field the mesh had just renamed, found by
|
||||
a full run. That step now lives in the repository rather than in a terminal history.
|
||||
|
||||
### What this issue taught twice
|
||||
|
||||
**The fix reintroduced the fault, in miniature, and the second time was caught by running it.**
|
||||
|
||||
The suite takes paths, so it can be pointed at one quick unit file — and the receipt from that run
|
||||
was, at first, indistinguishable from a receipt for the real thing. A green record standing for a
|
||||
run that raised no machines is this issue's own symptom, rebuilt inside its remedy. The receipt now
|
||||
records what it ran, and a run that did not include the end-to-end file is not coverage.
|
||||
|
||||
Separately, the code that decides *no receipt rather than a guessed one* — the rule that keeps the
|
||||
record meaning something — was first written where no test could reach it. Writing "0 failed"
|
||||
because nothing said otherwise is how a green record comes to mean nothing.
|
||||
|
||||
And the counting itself **passed every test while reading nothing**: the test runner colours its
|
||||
summary even into a pipe, so the anchored pattern never matched, and the fixtures it was checked
|
||||
against were output that had been imagined rather than captured. **A fixture that agrees with the
|
||||
mistake proves the mistake.** It is now checked against the runner's real bytes.
|
||||
|
||||
Each of these was found by running the thing, not by reading it — which is the same argument this
|
||||
issue makes about the pipeline.
|
||||
|
||||
@@ -1,9 +1,9 @@
|
||||
---
|
||||
status: open
|
||||
status: located
|
||||
opened: 2026-08-23
|
||||
located-in: [hal, hq]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
amended-design: 02-DECISIONS/0025-the-design-record-is-read-not-copied.md
|
||||
---
|
||||
|
||||
# 006 — This repository is not indexed into the knowledge base, and the claim that it is holds up a decision
|
||||
@@ -59,7 +59,7 @@ checked it — including in the same commit that wrote the rule.
|
||||
## Proposed direction — Nox is the search
|
||||
|
||||
*Added 2026-08-23.* Rather than syncing these documents into the knowledge base, **Nox
|
||||
([ADR 0027](../../02-DECISIONS/0027-the-product-is-novox-mesh.md)) works from within this
|
||||
([ADR 0019](../../02-DECISIONS/0019-how-this-repository-works.md)) works from within this
|
||||
repository and holds its knowledge directly.** Retrieval becomes an agent reading the source,
|
||||
not a copy living in a second store.
|
||||
|
||||
@@ -88,3 +88,70 @@ consults Nox — the answer is yes and the original promise holds.
|
||||
|
||||
That is a design question for Nox, not a defect in this repository, and it should be settled
|
||||
before ADR 0019 is treated as answered.
|
||||
|
||||
## Where this stands
|
||||
|
||||
*2026-08-31. Re-checked, and deliberately not closed.*
|
||||
|
||||
**The indexing still does not exist.** Two searches today, against both the symptom-indexed
|
||||
memory and the structured archive, using a decision record's full title and a distinctive phrase
|
||||
from a design document: no results, no partial match, no stale copy. The symptom in this report
|
||||
is unchanged.
|
||||
|
||||
**But the part that made it an issue is gone.** This report's argument was that the claim was
|
||||
*load-bearing* — that a decision rested on a mechanism nobody had checked. It no longer rests on
|
||||
it. The README now names the gap in the place the claim used to sit, and says it is left standing
|
||||
rather than quietly reworded. The decision record that separates this repository does not invoke
|
||||
indexing at all; its reasoning is cadence, reviewers, and scope, none of which depend on it.
|
||||
|
||||
So what remains is not a false claim. It is an unbuilt capability and an open design question,
|
||||
and those are different things.
|
||||
|
||||
### What was done
|
||||
|
||||
**A signpost, in the knowledge base, pointing here** — what lives in this repository, which
|
||||
folders hold what, and when to come looking rather than search there. Explicitly a pointer and
|
||||
not a copy: a derived copy drifts, and the enforced copy wins while the reasoned one quietly
|
||||
stops being true.
|
||||
|
||||
**It was tested, and it half works.** A search for *design records, decisions, repository* returns
|
||||
it. A search phrased the way somebody would actually ask — *why is the mesh built this way* —
|
||||
returns nothing, because the store matches terms rather than meaning.
|
||||
|
||||
That is this report's own distinction, confirmed by measurement rather than argued: **a signpost
|
||||
is reachable, it is not surfacing.** Someone who suspects the answer exists will now find it.
|
||||
Someone debugging an error, with no reason to think this repository knows anything about their
|
||||
symptom, still will not.
|
||||
|
||||
### Why it stays open
|
||||
|
||||
The question this report narrows to is unchanged and unanswered:
|
||||
|
||||
> When a symptom is searched and the answer happens to live in a design document or a decision
|
||||
> record here, does the searcher find it without already suspecting it exists?
|
||||
|
||||
Today: **no.** Closing this means choosing between a one-way sync into the knowledge base and an
|
||||
agent that reads this repository and contributes to a symptom search — and that is a decision
|
||||
about how the knowledge system works, not a defect to be fixed quietly.
|
||||
|
||||
**Marking it resolved while the indexing does not exist would be the failure this repository was
|
||||
created to name**, one folder away from where it names it.
|
||||
|
||||
## The direction is decided
|
||||
|
||||
*2026-08-31.* **The agent reads this repository; nothing is copied.** Recorded as
|
||||
[ADR 0025](../../02-DECISIONS/0025-the-design-record-is-read-not-copied.md), which also amends
|
||||
what [ADR 0019](../../02-DECISIONS/0019-how-this-repository-works.md) promised: these documents
|
||||
will not be *indexed*, they will be *read*, and the search consults the agent so its answers
|
||||
appear beside ordinary results.
|
||||
|
||||
A sync was the option that works with what exists today, and it was rejected on the one ground
|
||||
this repository can least afford: it makes a second copy, and *the copy that is searched quietly
|
||||
stops matching the copy that is edited*.
|
||||
|
||||
**So the open question above is answered, and this report stays open on the build.** What closes
|
||||
it is the check ADR 0025 names — search the mesh's memory for a phrase that appears only in a
|
||||
design document here, and get it back. That check fails today by design.
|
||||
|
||||
**What stands until then** is the signpost, and the honest description of it: reachable, not
|
||||
surfacing.
|
||||
|
||||
@@ -44,7 +44,7 @@ The distance between the two is the same one the delivery layer already has a na
|
||||
## Why it matters now
|
||||
|
||||
This is the first requirement of the lab
|
||||
([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)), which is
|
||||
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)), which is
|
||||
phase 0 of the entire migration. The first capability the new work depends on is present,
|
||||
declared, and unusable — and would have stayed unusable silently.
|
||||
|
||||
|
||||
@@ -0,0 +1,102 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-08-28
|
||||
located-in: [hal]
|
||||
fixed-by: hal — the script says it is manual, because it is
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 008 — The documented automatic node rescue does not exist
|
||||
|
||||
## Symptom
|
||||
|
||||
The mesh's documentation describes an automatic node rescue: a node that fails is recovered
|
||||
without anybody intervening. **Nothing implements it.**
|
||||
|
||||
Found incidentally while investigating supervision
|
||||
([research 003](../../01-RESEARCH/003-service-supervision/00-overview.md)), which counted what
|
||||
actually supervises what:
|
||||
|
||||
- **no unit declares `OnFailure=`**, so nothing runs when a unit gives up;
|
||||
- **nothing calls the rescue script on a timer**, so it runs only when a person runs it.
|
||||
|
||||
The script exists. The thing that would invoke it does not.
|
||||
|
||||
## Why this is worse than having no rescue
|
||||
|
||||
A rescue nobody wrote is a gap somebody can see. A rescue that is *documented* and absent is a
|
||||
gap nobody looks for, because the documentation says it is covered — and it is read exactly when
|
||||
a node has failed and somebody is deciding whether to intervene.
|
||||
|
||||
This is `how-we-build` §5 in its most expensive form: *an unenforced rule is indistinguishable
|
||||
from a wrong one, and costs more, because people believe it.* Here the belief is that a failed
|
||||
node recovers itself.
|
||||
|
||||
## Scope
|
||||
|
||||
**The as-is only.** The design being built has a different answer:
|
||||
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) puts recovery
|
||||
in a launcher that supervises the host, and that recovery is tested — 32 assertions, each
|
||||
confirmed to fail when the behaviour is removed.
|
||||
|
||||
So this issue is about the mesh that runs **now**, and it has two possible resolutions rather
|
||||
than one:
|
||||
|
||||
1. **Implement it** — an `OnFailure=` and a timer — if node rescue is wanted before the new host
|
||||
reaches the fleet.
|
||||
2. **Delete the documentation** — and say plainly that a failed node needs a person, which is
|
||||
what is true today.
|
||||
|
||||
**Either is honest. Leaving it as it is, is not.** The choice turns on how far away the new host
|
||||
is, which is a scheduling question rather than a technical one.
|
||||
|
||||
## What it would take to be sure
|
||||
|
||||
Read back rather than assumed
|
||||
([ADR 0018](../../02-DECISIONS/0018-a-picture-is-read-from-what-runs.md)): list every unit on a
|
||||
node and grep for `OnFailure=`; list every timer and check what each one calls. The finding above
|
||||
came from reading the repository, and confirming it against a running node is the difference
|
||||
between *no unit declares this* and *no unit in the source declares this*.
|
||||
|
||||
## Resolution
|
||||
|
||||
*2026-08-31.* **Resolved the second way: the documentation now says what is true.**
|
||||
|
||||
### Read back from running nodes, and the finding sharpened
|
||||
|
||||
This report was written from the repository. Checked against three running nodes, as the section
|
||||
above asks — and one of its own claims was wrong in a way that matters:
|
||||
|
||||
| Claim | Verified |
|
||||
|---|---|
|
||||
| no unit declares `OnFailure=` | **true**, and in *both* scopes — of 604 system units, the only two are stock (`snapd`, `local-fs.target`); across ~30 user units, none |
|
||||
| nothing calls the rescue script on a timer | **the timer exists.** `hal-health.timer` fires every 30 minutes |
|
||||
| the rescue is not actually triggered | **true** — the *deployed* `hal-health.sh`, not only the repository's copy, contains no call to it |
|
||||
|
||||
**The trigger exists and does not do the thing the script says it does.** That is worse than the
|
||||
absence this report described, because it survives a halfway check: somebody verifying "is there a
|
||||
health timer?" finds one, and stops.
|
||||
|
||||
The precise falsehood was a single line in the rescue script — *triggered automatically by
|
||||
`hal-health.timer` when the mesh service is failed.* It is now the opposite sentence, and says
|
||||
what does trigger it.
|
||||
|
||||
**Two further claims were found and narrowed.** Documentation in two places called the mesh
|
||||
*self-healing* — once for peers refreshing a tool cache, once for nodes auto-updating on a repo
|
||||
change. Both behaviours are real and neither is healing. **A phrase that overstates by a category
|
||||
is read in the crisis it describes**, and "the mesh is self-healing" is exactly the sentence that
|
||||
stops somebody intervening.
|
||||
|
||||
### Why not the first way
|
||||
|
||||
Implementing it was the other honest option, and it was not taken. The replacement host already
|
||||
supervises recovery and that recovery is tested; adding unattended download-and-restart behaviour
|
||||
to the fleet it is replacing is a change with real failure modes of its own — a rescue loop that
|
||||
thrashes is worse than a node that waits.
|
||||
|
||||
**This is the scheduling judgement this report said the choice turned on, and it is recorded
|
||||
rather than assumed.** If rescue is wanted on the current fleet before the new host reaches it,
|
||||
that is a decision to take deliberately: an `OnFailure=` and a call from the health script, both
|
||||
small, neither free.
|
||||
|
||||
Until then the documentation is true, which is the part that was costing something.
|
||||
@@ -0,0 +1,114 @@
|
||||
---
|
||||
status: fixed
|
||||
opened: 2026-08-28
|
||||
located-in: [mesh-lab, mesh-host]
|
||||
fixed-by:
|
||||
- "mesh-lab: a registry raised inside the scenario. Verified in a sealed machine — all four shapes applied with the image pinned by digest, idempotent, read back from the machine."
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 009 — A digest-pinned image cannot be placed in the lab, so `container` cannot be tested there
|
||||
|
||||
## Symptom
|
||||
|
||||
Two accepted decisions collide, and the collision makes one resource shape untestable.
|
||||
|
||||
- **[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)** pins images by
|
||||
digest, and the host **refuses** an image reference that is not pinned:
|
||||
|
||||
```
|
||||
resource "store": image "alpine:3.20" is not pinned. Write it as name@sha256:...
|
||||
```
|
||||
|
||||
- **The lab cannot place a digest-pinned image.** A sealed scenario cannot reach a registry, so
|
||||
the lab exports an image from the workstation and loads it in the machine — and that loses the
|
||||
digest.
|
||||
|
||||
So a `container` resource is refused by the host if it names a tag, and unusable if it names a
|
||||
digest. **There is no declaration the lab can currently raise that exercises the shape.**
|
||||
|
||||
## What was measured
|
||||
|
||||
Not inferred. `docker save alpine@sha256:d9e8…` produces an archive with **no repo tag**, because
|
||||
a repo digest exists only for an image a registry served. Loading it says:
|
||||
|
||||
```
|
||||
Loaded image ID: sha256:63f227… (not "Loaded image: alpine:3.20")
|
||||
```
|
||||
|
||||
and `docker images` then lists nothing — the image is there but dangling. A container declaring
|
||||
that digest therefore falls through to the registry:
|
||||
|
||||
```
|
||||
Unable to find image 'alpine@sha256:d9e8…' locally
|
||||
dial tcp: lookup registry-1.docker.io: no such host
|
||||
```
|
||||
|
||||
which is correct behaviour on a machine with no route out.
|
||||
|
||||
## What is not affected
|
||||
|
||||
Everything else placed in the same sealed machine works, and was verified there:
|
||||
|
||||
| shape | |
|
||||
|---|---|
|
||||
| `package` | applied, idempotent |
|
||||
| `service` incl. `boot: enabled` | applied, read back as `enabled` |
|
||||
| `action` | ran, verified |
|
||||
| `container` | **blocked by this issue** |
|
||||
|
||||
## Why it matters more than one shape
|
||||
|
||||
The container shape is the substrate. Every step of raising a mesh past the container runtime is
|
||||
a container ([`07-the-substrate.md`](../../03-DESIGN/01-to-be/07-the-substrate.md)), so the
|
||||
bootstrap cannot be tested end-to-end until this is resolved — which is the thing the lab exists
|
||||
for.
|
||||
|
||||
## How it was fixed
|
||||
|
||||
*2026-08-29.* A registry inside the scenario, as below — and it turned out to be the shape the
|
||||
resolution predicted rather than a compromise on it.
|
||||
|
||||
A scenario declares `images:` by tag. The lab stocks a registry **on the workstation**, where
|
||||
there is a network, then raises one **inside the scenario** as scenery and serves them from it.
|
||||
What a declaration pins is reported when the scenario is raised, because the digest belongs to
|
||||
that registry and is not knowable before it exists.
|
||||
|
||||
**The digests are the lab registry's own, and that is correct rather than a workaround.** What
|
||||
[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md) requires is a
|
||||
reference that is exact and cannot move. A digest this registry assigned is both.
|
||||
|
||||
Verified in a machine confirmed to have no route out: `package`, `service` including boot state,
|
||||
a `container` pinned by digest, and an `action` inside that container — applied, idempotent on
|
||||
re-apply, and read back from the machine rather than from the apply's own report.
|
||||
|
||||
**One fault is worth keeping**, because it is this repository's own subject arriving in the
|
||||
tooling built to catch it. The read-back checked that the registry's catalog endpoint answered,
|
||||
by looking for the substring `repositories` — which `{"repositories":[]}` also contains. So it
|
||||
**passed on a registry holding nothing**, and the failure surfaced much later as a container that
|
||||
could not be pulled, a long way from its cause. It now asks for each image's manifest **by
|
||||
digest**, which is what a machine actually does.
|
||||
|
||||
## The shape of a resolution
|
||||
|
||||
**A registry inside the scenario**, on its public segment, that machines pull from. That is not a
|
||||
workaround: it is what the real mesh does — [ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)
|
||||
names an OCI registry as substrate, and every node after the first pulls from the mesh's own.
|
||||
Testing against a registry is testing the real path rather than a stand-in for it.
|
||||
|
||||
It also removes the lab's export-and-push mechanism rather than fixing it, which is the better
|
||||
outcome: pushing image tarballs over the hypervisor was always a lab-only invention.
|
||||
|
||||
**Not decided here**, because it is design rather than repair: where the registry runs, whether
|
||||
it is scenery like the router ([ADR 0016](../../02-DECISIONS/0016-the-lab.md))
|
||||
or a placed artifact, and how images get into it.
|
||||
|
||||
## Incidental, and already fixed
|
||||
|
||||
The lab's own check on the load was too weak: it matched `"Loaded image"`, which is a prefix of
|
||||
both `Loaded image:` and `Loaded image ID:`. So a load that produced an unusable dangling image
|
||||
**reported success**, and the failure surfaced later as a container that would not start. It now
|
||||
matches `Loaded image:` exactly and says what the runtime actually said.
|
||||
|
||||
That is this repository's own subject arriving in its own tooling: a check that passes on the
|
||||
wrong thing is worse than no check, because it moves the failure away from its cause.
|
||||
@@ -0,0 +1,104 @@
|
||||
---
|
||||
status: fixed
|
||||
opened: 2026-08-29
|
||||
located-in: [mesh-host, mesh-control]
|
||||
fixed-by:
|
||||
- "mesh-host: the store records where each resource came from — carried or declared — and each origin removes only its own. Verified in the lab on the exact scenario that caused this: the substrate survived, and a later declaration still removed what it had itself declared."
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 010 — The first declaration a node receives destroys the substrate it raised
|
||||
|
||||
## Symptom
|
||||
|
||||
A first node was raised from its carried bundle: container runtime, store, two context databases,
|
||||
their schemas, broker, control plane. Eleven resources, all running. It then enrolled against the
|
||||
control plane on its own machine, held its link open, and was sent a declaration naming two
|
||||
resources — a directory and a file.
|
||||
|
||||
Both were applied correctly. And **every container on the machine was removed**: the store, the
|
||||
broker, and the control plane that had sent the declaration. The link died mid-sentence with
|
||||
`the link closed: Exception (501) Reason: "EOF"`, because the broker carrying it had just been
|
||||
torn down by the message it carried.
|
||||
|
||||
Afterwards `mesh-host owned` listed two resources. The mesh had deleted itself.
|
||||
|
||||
## What is actually wrong
|
||||
|
||||
Nothing in the code is behaving incorrectly. `apply` removes what the store holds and the incoming
|
||||
declaration does not name, which is what reconciliation means — the declaration is the desired
|
||||
state, not a patch, and anything else would make it impossible to remove a resource by omission.
|
||||
|
||||
**The fault is that the carried bundle and mesh declarations share one store.** The host cannot
|
||||
tell "this machine raised this for itself before there was a mesh" from "the mesh told this
|
||||
machine to have this", so the second overwrites the first completely.
|
||||
|
||||
That is invisible until the two meet, which happens exactly once per mesh: on the first node,
|
||||
after enrolment, at the moment the control plane first speaks.
|
||||
|
||||
## Why it matters more than a footgun
|
||||
|
||||
**The first node is the only node where the substrate is not the mesh's doing.** Every other node
|
||||
receives everything it runs from the control plane, so a complete declaration is complete by
|
||||
construction. The first node raised its own substrate from a file it carried, and the control
|
||||
plane has never been told about it — so the control plane cannot include it in a declaration even
|
||||
if it wanted to.
|
||||
|
||||
So the first node is left in a state no other node is in, and the ordinary path destroys it.
|
||||
|
||||
## What is not the answer
|
||||
|
||||
- **Making the control plane send the substrate back.** It does not know what the bundle contained
|
||||
and should not: the bundle exists precisely because there was no control plane yet.
|
||||
- **Making apply stop removing orphans.** Removal by omission is how a declaration says *stop
|
||||
running this*, and losing it costs the property that a node converges on what it was told rather
|
||||
than accumulating.
|
||||
- **Special-casing the first node.** [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)
|
||||
is explicit that its specialness lasts two commands, and this would extend it for ever.
|
||||
|
||||
## The fix
|
||||
|
||||
**The store records where each resource came from** — `carried` or `declared` — and each origin
|
||||
removes only its own. A declaration removes what the mesh previously declared and never what the
|
||||
bundle raised; reconciling the bundle removes what the bundle previously raised and never what the
|
||||
mesh assigned.
|
||||
|
||||
State written before the field existed reads as `carried`, because everything a host had applied
|
||||
at that point came from its bundle — there was no other way to tell it anything. Guessing the
|
||||
other way would have the first upgrade remove the substrate, which is this fault arriving through
|
||||
the change that fixes it.
|
||||
|
||||
**Verified on the scenario that caused it.** A first node raised eleven resources, enrolled, and
|
||||
was sent the same two-resource declaration. Both applied; the store, the broker and the control
|
||||
plane were still running afterwards. A second declaration dropping one resource removed that
|
||||
resource and nothing else, so removal by omission still works — which is the property that had to
|
||||
survive the fix.
|
||||
|
||||
**What remains open** is what happens when the mesh eventually declares the substrate, which it
|
||||
must, or the substrate can never be upgraded. Two sources claiming one container is the ambiguity
|
||||
this issue is made of, narrowed rather than removed: it can no longer happen by accident, and
|
||||
nothing yet says what it means when it happens on purpose.
|
||||
|
||||
## How it was found
|
||||
|
||||
In the lab, on a sealed machine, by doing the ordinary thing: raise a first node, enrol it, and
|
||||
tell it something. It was not a test of this — it was the first end-to-end run of the link, and
|
||||
this fell out of it.
|
||||
|
||||
The declaration was two lines and destroyed a working mesh in under a second, which is worth
|
||||
holding on to: this is not an edge case reached by trying, it is the first thing that happens.
|
||||
|
||||
|
||||
## Two more faults found while fixing it
|
||||
|
||||
Both of the same shape, and worth recording because the shape is the point.
|
||||
|
||||
**A report published to an unbound routing key vanishes.** The control plane bound `enrol` and not
|
||||
`report`, so nodes announced what they had applied into a void — the broker accepted each message,
|
||||
found no queue for it, and dropped it. The publisher was told nothing. Reports are now published
|
||||
`mandatory`, so anything unroutable comes back and is said out loud, and the binding covers every
|
||||
key a node may publish.
|
||||
|
||||
**A publish failure was being swallowed.** `publishReport` discarded its error, so a node that
|
||||
could not tell the mesh what it had done looked identical to one that had. That is the fault this
|
||||
repository keeps cataloguing, written by hand into the newest code in it.
|
||||
@@ -0,0 +1,85 @@
|
||||
---
|
||||
status: fixed
|
||||
opened: 2026-08-30
|
||||
located-in: [mesh-host]
|
||||
fixed-by: mesh-host — apply attempts every resource and reports every failure
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 011 — One broken module stops every module after it, for ever
|
||||
|
||||
## Symptom
|
||||
|
||||
A machine was assigned a module declaring a package that does not exist. Every later push to that
|
||||
machine applied **nothing at all**, and kept doing so.
|
||||
|
||||
Found while proving something else. A test assigned a deliberately-impossible module to a machine
|
||||
to check that the mesh reports a failure — which it does. A later test on the same machine then
|
||||
failed, and the evidence said why:
|
||||
|
||||
```
|
||||
applied 0 and failed: applying "impossible.nothing":
|
||||
installing a-package-that-does-not-exist: target not found
|
||||
0 resource(s) were applied before this and remain
|
||||
```
|
||||
|
||||
The broker's queues were **empty**, so the declaration had been delivered and read. The machine
|
||||
simply stopped at the first failing resource and never reached the rest.
|
||||
|
||||
## Why it matters more than one machine
|
||||
|
||||
- **A machine with one bad module and nine good ones runs none of the nine**, and the mesh reports
|
||||
"failed" without saying that the rest were never attempted.
|
||||
- **It cannot be recovered by retrying.** Anything that re-pushes to machines that are behind — the
|
||||
obvious next feature — would retry a permanent failure for ever and make no progress on
|
||||
everything else.
|
||||
- **The order is not the operator's.** Which module is "first" is an accident of resolution, so
|
||||
which nine modules a broken one blocks is unpredictable.
|
||||
|
||||
## What was there, and what it rested on
|
||||
|
||||
The behaviour had a test asserting it: *nothing after the failure ran*. Its comment cites
|
||||
[ADR 0010](../../02-DECISIONS/0010-delivery.md).
|
||||
|
||||
**That record does not decide this.** What it says is that a failed *job* stops and names its step
|
||||
while a reconciler retries forever, as an argument about pipelines against reconcilers. It says
|
||||
nothing about whether one resource failing should prevent the next from being attempted. The
|
||||
citation was doing more work than the record supports.
|
||||
|
||||
## The fix, and the argument that was on the other side
|
||||
|
||||
**Everything is attempted, and every failure is reported.**
|
||||
|
||||
The case for stopping is that a resource may depend on an earlier one — a service on the file it
|
||||
reads. That is real, and it survives: such a service fails its own check and is reported. This host
|
||||
reads back after every write precisely so a thing that did not work is caught rather than assumed,
|
||||
so attempting it produces *more* information than skipping it.
|
||||
|
||||
**What is unchanged:** a declaration that cannot be parsed is still refused whole, and nothing is
|
||||
applied. That is a different thing — *this machine could not do it* against *this was never a
|
||||
declaration* — and they are fixed in different places.
|
||||
|
||||
### And one shape still stops what follows, which the first fix got wrong
|
||||
|
||||
**An action does.** The first version of this fix continued past everything, and the next lab run
|
||||
failed at the bootstrap: the store did not answer in three minutes and then said *the database
|
||||
system is shutting down*. Carrying on past the readiness gate had started the broker and the
|
||||
control plane against a machine that was not ready, and on a small machine that is how a database
|
||||
still initialising has its memory taken away.
|
||||
|
||||
**An action is the only shape whose purpose is to make something true *before* the next thing needs
|
||||
it** — which is why it is the only one with a `verify`. The bootstrap is a row of them: the store
|
||||
answers, then its databases exist, then their schemas, then the broker. Everything else is
|
||||
independent state: a package that will not install has nothing to do with a file on the other side
|
||||
of the declaration.
|
||||
|
||||
So the rule is: **a failed action stops what follows; nothing else does.** Both faults are fixed by
|
||||
it, and the report says which happened — *these things failed* and *these things failed and the
|
||||
rest was never tried* are different machines.
|
||||
|
||||
## How this is checked
|
||||
|
||||
`internal/apply`, three tests: an apply with a resource that cannot succeed still applies the ones
|
||||
after it; every failure is counted, not just the first; and a failed **action** stops what follows
|
||||
and says so. Each confirmed to fail when its behaviour is removed — including the last, which fails
|
||||
if actions stop being treated as gates *and* if everything is treated as one.
|
||||
@@ -0,0 +1,88 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-08-30
|
||||
located-in: [mesh-lab]
|
||||
fixed-by: mesh-lab — scenario machines stay at 1 GiB, and the scenario now stocks seven images
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 012 — Raising a scenario machine's memory stops the substrate coming up
|
||||
|
||||
*Renamed after the first diagnosis turned out to be wrong. What that was, and how it was wrong, is
|
||||
below — it is the more useful half of this report.*
|
||||
|
||||
## Symptom
|
||||
|
||||
Adding a fourth image to the two-machine scenario made the bootstrap fail every time. The store
|
||||
container was created, and its readiness check then failed for the full three minutes with
|
||||
**no output at all**:
|
||||
|
||||
```
|
||||
failed store-ready (in mesh-store: ... pg_isready ... exit 1):
|
||||
the action ran without error and its own verify still fails: docker exited 1:
|
||||
```
|
||||
|
||||
Empty after the colon. The check runs `pg_isready` inside the container and prints the store's own
|
||||
last lines when it gives up; producing nothing means **the container was not running**, which is a
|
||||
different fault from a database that is slow to start.
|
||||
|
||||
Three images: the scenario raises, both machines join, and nine assertions pass. Four: it never
|
||||
gets past the store. Reverting the fourth image restores it.
|
||||
|
||||
## The first diagnosis was wrong, and this is why it is worth writing down
|
||||
|
||||
**Two things changed at once.** A fourth image was added to the scenario, and — reasoning that a
|
||||
machine running a database, a broker and the control plane at once is genuinely small — scenario
|
||||
machines were raised from 1 GiB to 2 GiB. The bootstrap then failed every time, and the fourth
|
||||
image was blamed.
|
||||
|
||||
**Removing the image did not fix it. Removing the memory increase did.** Nine assertions pass again
|
||||
with the extra memory reverted, on a scenario with three images. So the cause is the memory change:
|
||||
three machines at 2 GiB, on a host also running other work, contend enough that the store container
|
||||
does not come up at all.
|
||||
|
||||
**The lesson is the ordinary one and it still caught me:** two changes went in together, the failure
|
||||
was attributed to the plausible one, and an issue was written recording the wrong cause. What found
|
||||
it was reverting to the exact last-known-good state rather than reverting the suspicious change.
|
||||
|
||||
**What remains untested** is whether a fourth image alone is fine. It probably is. Nothing has
|
||||
measured it, and the honest state of this issue is that the thing it was opened about was never
|
||||
demonstrated.
|
||||
|
||||
## What was worth keeping
|
||||
|
||||
**A better diagnostic.** The readiness check now prints what it saw before giving up. That is what
|
||||
showed the output was empty, which is what said *the container is not running* rather than *the
|
||||
database is slow* — and which will make the next occurrence of this a diagnosis instead of a
|
||||
retry.
|
||||
|
||||
## What this blocks
|
||||
|
||||
The mesh running its **own artifact store** — a registry as a module — needs a registry image on
|
||||
the machine so the module can mirror one, which is the fourth image. The module is written and its
|
||||
manifest is accepted; what has not been proven is a machine assigned it serving artifacts to
|
||||
another machine.
|
||||
|
||||
## How this will be checked
|
||||
|
||||
A scenario raised with four images comes up and passes the assertions that three do — which is
|
||||
what was never actually established. Until then the artifact-store test is not in the shared
|
||||
scenario, with a note saying where it went and why.
|
||||
|
||||
## Resolved
|
||||
|
||||
*2026-08-31.* The condition this report set was *a scenario raised with four images comes up and
|
||||
passes the assertions that three do*. The scenario now stocks **seven** and has raised cleanly
|
||||
many times over, with the machines at 1 GiB where the wrong diagnosis had put them at 2.
|
||||
|
||||
So both halves are settled. **The memory increase was the cause** — reverted, and never
|
||||
reintroduced. **A fourth image was never the problem**, which this report said had not been
|
||||
demonstrated either way, and now has been: three more were added on top of it, and the artifact
|
||||
store the issue said was blocked is proven in the shared scenario rather than kept out of it.
|
||||
|
||||
**The diagnostic that came out of it is what remains valuable.** The readiness check prints what it
|
||||
saw before giving up, which is what turned *the database is slow* into *the container is not
|
||||
running*. It has since caught a different fault of the same shape — an action succeeding into a
|
||||
state its own verify rejects
|
||||
([04-ISSUES/017](../017-an-action-succeeded-into-a-state-its-verify-rejects/00-report.md)) — which
|
||||
is the argument for keeping a good diagnostic after the incident that prompted it is gone.
|
||||
@@ -0,0 +1,56 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-08-31
|
||||
located-in: [mesh-control]
|
||||
fixed-by: mesh-control — what the mesh computes is applied before what the module declared
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 013 — A file the mesh computes arrives after the service that needs it
|
||||
|
||||
## Symptom
|
||||
|
||||
Everything the control plane computes for a module — a certificate, a sealed credential, a bound
|
||||
file, a rule set — was placed **after** that module's own resources in the declaration. The host
|
||||
applies resources in the order it is given and
|
||||
[does not sort](../../02-DECISIONS/0005-the-node-host.md), so a service or container declared in a
|
||||
manifest was applied **before** the file it depends on existed.
|
||||
|
||||
On the first apply the service starts against a missing file and fails. The next reconcile finds
|
||||
the file there and starts it.
|
||||
|
||||
## Why this matters
|
||||
|
||||
**It repairs itself, which is why nothing caught it.** A fault that is gone by the second attempt
|
||||
is worse than one that persists: what gets remembered is that the thing works, and the failed
|
||||
first apply is read as a machine that was briefly slow. The mesh reports a failure, then reports
|
||||
success, and nobody looks again.
|
||||
|
||||
It was also invisible to every test that existed, because none of them combined the two halves.
|
||||
Modules with computed files declared no service; modules with a service needed no computed file.
|
||||
The fault lived exactly in the gap between two repositories' assumptions — the control plane
|
||||
deciding an order, the host promising not to change it — which is the shape this folder exists for.
|
||||
|
||||
Found by reading, while writing the first module that has both: a firewall whose service must
|
||||
reflect a rule set the mesh computes.
|
||||
|
||||
## Evidence
|
||||
|
||||
`internal/catalogue/declaration.go` built each module's resource list as
|
||||
`append(module's own, computed...)` in six places — certificate, authority, needs, secrets, grants,
|
||||
bindings. `internal/apply/apply.go` iterates `d.Resources` in order, and
|
||||
`internal/declaration/declaration_test.go` states the rule directly: *order is stated, not derived.
|
||||
The host must not sort.*
|
||||
|
||||
## What was done
|
||||
|
||||
The computed resources are assembled first and the module's own resources follow. Nothing the mesh
|
||||
computes is derived from a module's resources, so the order is unconditionally right rather than a
|
||||
heuristic — there is no case where a module's resource must precede a file the mesh made for it.
|
||||
|
||||
Merged after the computed-resources branch, which replaces a module's resources wholesale and
|
||||
would otherwise discard everything the mesh had made for it.
|
||||
|
||||
*Checked by a module declaring a service that reflects a rule set, asserting the rule set is first;
|
||||
and by a module whose resources are computed elsewhere, asserting its credential survives and is
|
||||
still first — the case the merge point exists for.*
|
||||
@@ -0,0 +1,48 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-08-31
|
||||
located-in: [mesh-host]
|
||||
fixed-by: mesh-host — a node's serving key is stored in the format a server reads
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 014 — A node's serving key was present, correct, and unusable
|
||||
|
||||
## Symptom
|
||||
|
||||
A node generates the key it serves TLS with, the mesh certifies the public half, and the
|
||||
certificate arrives on the machine as an ordinary file. Everything about that worked. But the host
|
||||
stored the private half in its own encoding — base64 of the raw key — and **nothing that serves
|
||||
TLS can read it**: not a web server's `ssl_certificate_key`, not Go's `LoadX509KeyPair`, not
|
||||
`openssl s_server -key`.
|
||||
|
||||
The file was there, owned by root, mode 0600, holding the right key. The certificate beside it was
|
||||
valid and chained to the mesh's authority. The server would not start.
|
||||
|
||||
## Why this matters
|
||||
|
||||
**Every check that reads the file passes.** The key exists, the certificate exists, the mesh
|
||||
recorded the public half, the machine reports the declaration applied. The failure surfaces only
|
||||
when something connects — the worst place to find out, and the place the certificate work was
|
||||
specifically designed to move away from.
|
||||
|
||||
It is the same shape as [013](../013-a-file-arrives-after-the-service-that-needs-it/00-report.md)
|
||||
and worth naming as a class: **two halves of one mechanism designed separately, each correct
|
||||
about its own half.** The control plane issues PEM because that is what a certificate is. The host
|
||||
stored the key in whatever was convenient, because nothing in the host reads it back — the whole
|
||||
point of the file is that *something else* does, and that something else was not in view.
|
||||
|
||||
**The generalisation:** where a file exists so a third party can read it, the format is not an
|
||||
implementation detail of whoever writes it. It is the interface, and it needs a check that reads
|
||||
it the way that third party will.
|
||||
|
||||
## What was done
|
||||
|
||||
PKCS#8 PEM, which is what every TLS server reads. A key in the old encoding is refused **by name**
|
||||
rather than reported as corrupt — it is intact, and the remedy is to enrol again, which is a
|
||||
different action from repairing a damaged file.
|
||||
|
||||
*Checked by writing a key, decoding the file as PEM, parsing it as PKCS#8, and asserting it is the
|
||||
same key — and, in the lab, by a real handshake from a second machine that verifies against the
|
||||
mesh's authority and nothing else. A key that parses is not a key a server can use, which is why
|
||||
the lab check connects.*
|
||||
@@ -0,0 +1,44 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-08-31
|
||||
located-in: [mesh-lab]
|
||||
fixed-by: mesh-lab — a command with no marker is a failure, not a success
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 015 — A command that said nothing was read as having succeeded
|
||||
|
||||
## Symptom
|
||||
|
||||
The end-to-end harness runs a command on a machine and reads its exit status from a marker it
|
||||
appends to the output. When the marker was absent, the parse produced `Number("")`, which is `0`,
|
||||
and **the command was reported as having succeeded.**
|
||||
|
||||
The marker went missing whenever a command contained a heredoc. Everything was wrapped on a single
|
||||
line — `<cmd> 2>&1; echo "__exit=$?"` — so a heredoc's terminator line became
|
||||
`MARKER 2>&1; echo "__exit=$?"`, matched nothing, and the heredoc consumed the rest of the script,
|
||||
the marker included.
|
||||
|
||||
## Why this matters
|
||||
|
||||
**This is the harness lying in the one direction a harness must never lie.** Everything else in
|
||||
this repository is arranged around the principle that absence must never be indistinguishable from
|
||||
success — the host says so about a service that does not exist, the builder says so about a build
|
||||
that failed, the control plane says so about an empty list. The thing that checks all of that had
|
||||
the fault itself.
|
||||
|
||||
Its reach is every heredoc in the suite, which is how large files are written to machines: a
|
||||
substrate bundle, a certificate authority, a listener script. Each was written with trailing
|
||||
junk from the swallowed wrapper, each reported success, and each happened to be tolerated by
|
||||
whatever read it — until one was a Python script, which did not run, and the test failed on its own
|
||||
setup. **That failure read exactly like the thing being tested working.**
|
||||
|
||||
## What was done
|
||||
|
||||
`exec 2>&1` on its own first line, so nothing is appended to the command's last line and a heredoc
|
||||
terminates where it says it does. A missing marker is now a failure, returning whatever was said
|
||||
so the reason is visible rather than inferred.
|
||||
|
||||
*Checked by the firewall test, whose listener is written with a heredoc: it could not have started
|
||||
before this, and the assertion that it is reachable before any rule set exists is what makes the
|
||||
rest of that test mean anything.*
|
||||
@@ -0,0 +1,42 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-08-31
|
||||
located-in: [mesh-host]
|
||||
fixed-by: mesh-host — something after the declaration is refused whole
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 016 — Anything after the declaration in a file was ignored
|
||||
|
||||
## Symptom
|
||||
|
||||
A JSON decoder reads one value and stops. The host's declaration parser used one, so a file
|
||||
holding a declaration **followed by anything at all** — a stray line, a second declaration, the
|
||||
tail of a truncated rewrite — parsed as the first value and the rest was never looked at.
|
||||
|
||||
The machine applied something, reported success, and what it applied was not what the file said.
|
||||
|
||||
## Why this matters
|
||||
|
||||
The host refuses a partial declaration everywhere else, in these words: *a host that applied the
|
||||
parts it understood would leave a machine that looks configured and is not.* This was the same
|
||||
fault in its quietest form — not a part left out, but a part never seen, with nothing anywhere
|
||||
saying so.
|
||||
|
||||
**It was live and invisible.** The end-to-end harness had been appending a line to the substrate
|
||||
bundle by accident ([015](../015-a-command-with-no-answer-was-read-as-success/00-report.md)), and
|
||||
every bootstrap in every run applied a bundle with a junk line on the end. Nothing failed, so
|
||||
nothing was looked at, and the corruption was found only by tracing a different bug backwards.
|
||||
|
||||
**That is the argument for refusing rather than tolerating.** A file with something after it is
|
||||
more likely to be damaged than deliberate: an interrupted write, two files concatenated, a
|
||||
generator that emitted twice. Applying the first value is applying something nobody wrote.
|
||||
|
||||
## What was done
|
||||
|
||||
After decoding, the parser requires end of input. Trailing whitespace is not "something after it"
|
||||
— refusing that would make every file an editor writes unusable.
|
||||
|
||||
*Checked by a valid declaration with a stray line after it, with a second declaration after it,
|
||||
and with garbage after it, each refused; and by the same declaration with trailing blank lines,
|
||||
accepted.*
|
||||
@@ -0,0 +1,44 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-08-31
|
||||
located-in: [mesh-host]
|
||||
fixed-by: mesh-host — the store's readiness is checked over TCP, not the socket
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 017 — An action succeeded into a state its own verify rejects
|
||||
|
||||
## Symptom
|
||||
|
||||
The substrate's `store-ready` action waits for the store to answer, then the host runs the
|
||||
action's `verify` to read back that it worked. Intermittently the host reported:
|
||||
|
||||
> the action ran without error and its own verify still fails
|
||||
|
||||
Both statements were true. The action waited on the **unix socket**; its verify checked the same
|
||||
way, a moment later, and found nothing. Roughly one bootstrap in three.
|
||||
|
||||
## Why this matters
|
||||
|
||||
**The action and its verify were asking different questions without appearing to.** While the
|
||||
store initialises it runs a temporary server on the socket only, then stops it and starts the real
|
||||
one. The action's loop saw the temporary server and exited happy; the verify landed in the gap
|
||||
between the two.
|
||||
|
||||
So an action can **succeed into a state its own verify rejects** — and when it does, the host's
|
||||
report is accurate and useless. It says the command worked and the read-back did not, which is
|
||||
exactly what the mechanism is for, and names nothing a person can act on. It read as a slow
|
||||
machine, and the remedy people reach for is a longer timeout, which cannot help.
|
||||
|
||||
**The general rule, which the host's design should carry:** an action's verify is the *definition*
|
||||
of what the action is for. If the action's own waiting decides it is done by a different test than
|
||||
the verify uses, the two can disagree — and the disagreement surfaces as an intermittent failure
|
||||
in the one place designed to catch silent success.
|
||||
|
||||
## What was done
|
||||
|
||||
Both check the store over TCP, which the init phase deliberately does not open — so neither can
|
||||
mistake the temporary server for the real one, and neither can be satisfied while the other is not.
|
||||
|
||||
*Checked by the bootstrap itself, which is where it failed: this action gates everything after it,
|
||||
so a mesh coming up at all is the check.*
|
||||
@@ -0,0 +1,51 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-08-31
|
||||
located-in: [mesh-control]
|
||||
fixed-by: mesh-control — something answered on this machine is still bound
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 018 — A provider on the same machine was never announced to its consumer
|
||||
|
||||
## Symptom
|
||||
|
||||
A module that `binds` a provision receives a file naming where the provider is and what it said a
|
||||
consumer must know. When the provision turned out to be answered by another module **on the same
|
||||
machine**, no file was written at all.
|
||||
|
||||
A build machine sharing a node with the registry it pushes to therefore started, connected to the
|
||||
broker, and looped: *cannot read what the mesh said about the artifact store: no such file or
|
||||
directory.*
|
||||
|
||||
## Why this matters
|
||||
|
||||
It was deliberate, and the reasoning is in the code: *a file saying "it is on this node" would be a
|
||||
fact nobody needs and one more thing to keep true.* That is **right about the location and wrong
|
||||
about everything beside it.** A binding also carries the provider's `serves` block — the port —
|
||||
and a consumer cannot invent that whether the provider is next door or on the same disk.
|
||||
|
||||
**The failure names nothing.** Every part a person would check was correct: the module resolved,
|
||||
the machine applied it, the container ran, the credential was delivered and worked. The one file
|
||||
that did not exist was one the module never asked for by name — it asked for a *provision*, and
|
||||
the mesh silently decided the answer needed no writing down. The error is about a path, and the
|
||||
cause is a decision three layers away.
|
||||
|
||||
**And it only appears when two modules land on one node**, which is the ordinary case in a small
|
||||
mesh and the rare case in a large one. It would have been found in production.
|
||||
|
||||
## What was done
|
||||
|
||||
The binding is written for a local provider too, **when the provider said something a consumer
|
||||
must know**. That keeps the original intent exactly where it was right: a shell is answered here
|
||||
and there is genuinely nothing to say about it; a registry is answered here and the port is still
|
||||
unguessable.
|
||||
|
||||
The address is this machine's own name on the private network, or loopback when it has none — a
|
||||
machine off the network still reaches itself, and a name nothing resolves is worse than an address
|
||||
that always works.
|
||||
|
||||
*Checked by resolving a node holding both a provider and its consumer and asserting the binding
|
||||
carries the port and the address; by the same on a machine with no private network, asserting
|
||||
loopback; and by the pre-existing check that a provision with nothing to say still writes nothing,
|
||||
which is the half that was right.*
|
||||
@@ -0,0 +1,59 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-08-31
|
||||
located-in: [mesh-control]
|
||||
fixed-by: mesh-control — the resolver module's claims about machines are checked on a machine
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 019 — A comment asserting a fact about a machine, which nothing checked
|
||||
|
||||
## Symptom
|
||||
|
||||
The resolver module carried two statements about the machine it runs on. Both read as reasoned,
|
||||
both were in prose beside the setting they justified, and **both were wrong**:
|
||||
|
||||
| it said | the machine said |
|
||||
|---|---|
|
||||
| `127.0.0.54` is free — "not `.53`, that is systemd-resolved's" | systemd-resolved holds **both**; `.54` is its proxy stub. dnsmasq could not create the socket and never started |
|
||||
| it takes only `127.0.0.55` | listening on a loopback address takes the rest of loopback with it, `127.0.0.1` included |
|
||||
|
||||
A third statement in the same file was true and incomplete in a way that mattered as much: the
|
||||
config read `/etc/resolv.conf` for upstreams without saying so, and the module that points a
|
||||
machine at the mesh writes *this resolver's own address* into that file. So its upstream was
|
||||
itself. Its receive queue filled with 15KB of queries and every lookup on the machine hung.
|
||||
|
||||
## Why this matters
|
||||
|
||||
**The module had unit tests, and they all passed.** They checked that it names an address, that it
|
||||
reads what the mesh writes, that it restarts when that changes, and that the two asking modules
|
||||
point where it answers. Every one of those was true while the daemon could not start at all.
|
||||
|
||||
That is not a gap in those tests. It is what a unit test *is*: it confirms the assertion was made,
|
||||
never that it is true of any machine. **Only a machine knows which of its addresses are spare, or
|
||||
what a daemon does with a file when it starts.**
|
||||
|
||||
This is [04-ISSUES/003](../003-firewall-scope-is-read-by-no-code/00-report.md) in prose rather
|
||||
than in a manifest key. There, five manifests carried a `scope:` that read as a restriction and
|
||||
restricted nothing. Here, a comment read as a reasoned choice of address and chose a taken one. In
|
||||
both cases *an unenforced rule is indistinguishable from a wrong one, and costs more, because
|
||||
people believe it* — and a comment is the least enforced rule there is.
|
||||
|
||||
**It cost three full lab cycles**, at fifteen minutes each, because each one revealed exactly one
|
||||
of the three faults.
|
||||
|
||||
## What was done
|
||||
|
||||
**The wrong statements are corrected, and the correction says what it now knows rather than
|
||||
asserting a new comfort.** `.55` is written down as *a convention, not a reservation*: if a future
|
||||
systemd takes it, one line changes. What the module takes is what its claim already said — the
|
||||
machine's DNS port — rather than a promise about one address.
|
||||
|
||||
**The unit tests hold what a machine has told us.** They assert the module does not take `.53`,
|
||||
`.54` or `127.0.0.1`, and that it does not read resolv.conf for upstreams. A unit test cannot
|
||||
discover those facts; it can refuse to forget them.
|
||||
|
||||
**And the order changed.** A module that asserts something about machines is proven on a machine
|
||||
*before* its assertions are believed — the lab test written first, not last. Written here because
|
||||
the cost of the old order is measurable: three cycles, forty-five minutes, for a module whose
|
||||
mesh-side half was correct from the start.
|
||||
@@ -0,0 +1,94 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-08-31
|
||||
located-in: [mesh-control, mesh-lab]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 020 — A certificate is issued and never collected
|
||||
|
||||
## Symptom
|
||||
|
||||
Against a real ACME server in the lab, the proxy orders a certificate for a name the mesh routes,
|
||||
the challenge is answered, the authority **issues the certificate** — and the proxy never obtains
|
||||
it. Every TLS handshake then fails, and the order is retried indefinitely.
|
||||
|
||||
The client's error, once per attempt:
|
||||
|
||||
```
|
||||
http: TLS handshake error: Post "": unsupported protocol scheme ""
|
||||
```
|
||||
|
||||
A POST to an empty URL: the certificate's location, on an order the authority considers valid.
|
||||
|
||||
## What is proven, and it is most of it
|
||||
|
||||
Read from the authority's own log rather than inferred:
|
||||
|
||||
```
|
||||
Starting 3 validations
|
||||
authz … set VALID by completed challenge …
|
||||
POST /finalize-order/ → Order … is fully authorized. Processing finalization
|
||||
Issued certificate serial 3ef142939115ee88
|
||||
```
|
||||
|
||||
**The hard half works.** The order is created, the HTTP-01 challenge is answered *at the name being
|
||||
certified* on port 80 through the proxy itself, the authorisation goes valid, finalisation is
|
||||
accepted, and a certificate is issued. Across one run the authority issued **two** certificates and
|
||||
accepted finalise **three** times — the client reaches issuance every attempt and fails at the same
|
||||
step after it.
|
||||
|
||||
**And the policy that guards the quota is proven too.** The second assertion in the same file
|
||||
passes: no certificate is ordered for a name nothing routes, so a scan cannot spend an account's
|
||||
rate limit.
|
||||
|
||||
## What is not known
|
||||
|
||||
**Whether this happens against a real authority at all.** Everything above is against Pebble, which
|
||||
exists to be a test server. The failure is in the last hop between one client and one server, and
|
||||
may say nothing about behaviour against a public authority.
|
||||
|
||||
## Ruled out
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| the directory | fetched and complete — `newAccount`, `newNonce`, `newOrder`, `revokeCert` all present |
|
||||
| the authority's API certificate | covers `127.0.0.1`; the bundle is named explicitly and verification is not skipped |
|
||||
| a hand-written server config | suspected, and wrong. Replacing it with the server's **own** default config, changing only the challenge port, gives the identical error |
|
||||
| the finalize URL being empty | the authority logs finalisation being accepted |
|
||||
| the challenge path | the authorisation goes valid |
|
||||
| the server version | pinned 2.5.0 behaves exactly as `latest`, so the draft profiles extension is not it |
|
||||
|
||||
## Where it might be
|
||||
|
||||
- **The order's `certificate` field is absent when the client reads it.** The client waits for the
|
||||
order to become valid and only then fetches, so an empty location on a valid order is the
|
||||
remaining shape.
|
||||
- ~~**A moving tag was used.**~~ **Ruled out.** `pebble:latest` advertises a draft *profiles*
|
||||
extension, so a pinned 2.5.0 was tried: **identical failure**. The scenario now pins it anyway,
|
||||
which it should have from the start.
|
||||
|
||||
## Why this is filed rather than pursued
|
||||
|
||||
**The mesh-side behaviour is proven and the remainder is interop between two libraries.** Continuing
|
||||
would be several more twenty-minute lab runs against a server that is not the one production uses,
|
||||
to chase a defect that may not exist there.
|
||||
|
||||
**What the mesh needed to show, it showed**: a name it routes gets a certificate ordered from a
|
||||
configured authority, and a name it does not route gets nothing. The configuration is right, the
|
||||
challenge path is right, and issuance happens.
|
||||
|
||||
## What would close it
|
||||
|
||||
- The same scenario against a different ACME implementation — a second server, or a real staging
|
||||
endpoint from a machine that can reach one. **If it passes there, this is a Pebble interop
|
||||
detail and the issue closes with that recorded.**
|
||||
- Or the client's request captured on the wire, showing what the order actually contained when the
|
||||
location was read.
|
||||
|
||||
## Evidence
|
||||
|
||||
- `mesh-lab test/integration/certificates.test.ts`, scenario `a-public-name`
|
||||
- One assertion passes (no certificate for an unrouted name); one fails (a routed name is never
|
||||
served).
|
||||
@@ -0,0 +1,90 @@
|
||||
---
|
||||
status: fixed
|
||||
opened: 2026-09-01
|
||||
located-in: [mesh-control]
|
||||
fixed-by: mesh-control df62bb5
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 021 — A consumer on the provider's machine is given no credential
|
||||
|
||||
## Symptom
|
||||
|
||||
A module that requires something answered **on the same machine** resolves cleanly and is given
|
||||
**no credential at all**. Two modules, zero needs:
|
||||
|
||||
```
|
||||
postgres provides postgres-database, grants /var/lib/postgres/grants
|
||||
keycloak requires postgres-database, secrets /var/lib/keycloak/database.env
|
||||
→ modules: 2, needs: 0
|
||||
```
|
||||
|
||||
Nothing is refused and nothing is reported. The consumer's `secrets:` path is simply never
|
||||
written, and whatever reads it fails later, somewhere else.
|
||||
|
||||
## Where it comes from
|
||||
|
||||
The world a node resolves against is **every other node**:
|
||||
|
||||
```go
|
||||
for _, n := range nodes {
|
||||
if n.Name == exclude { continue }
|
||||
```
|
||||
|
||||
So a provider on the same machine is never a `Provider` in `world.Offered`, never becomes a
|
||||
`Needed`, and the credential loop — which walks `resolved.Needs` — has nothing to walk. Every step
|
||||
is individually reasonable and the sum is a silent gap.
|
||||
|
||||
## Why it was not noticed
|
||||
|
||||
**Everything proven so far was cross-machine.** The lab's provisioner scenarios put the consumer on
|
||||
one node and the provider on another, which is the interesting case for a *mesh* and the rare case
|
||||
in practice. The first module to want a database on its own machine was the first real one.
|
||||
|
||||
The postgres provisioner even records the assumption in passing — *"Node is empty for a module on
|
||||
this machine, which is asking for something local and is not this provisioner's business"* — which
|
||||
reads as a deliberate exclusion of local consumers.
|
||||
|
||||
## Why the assumption is wrong
|
||||
|
||||
It holds for a process on the machine reaching a unix socket, where the operating system can vouch
|
||||
for who is calling. **It does not hold for containers**, which is how nearly everything runs here: a
|
||||
module's containers reach a provider's containers over TCP on a shared network, and the database
|
||||
asks for a password exactly as it would from another machine.
|
||||
|
||||
**The machine is not a trust boundary once both sides are containers.** Treating it as one gives
|
||||
the most common arrangement — a service and its database on one node — the weakest handling.
|
||||
|
||||
## What it is not
|
||||
|
||||
Not the same as [`020`](../020-a-certificate-is-issued-and-never-collected/00-report.md) or a
|
||||
provisioner defect. The provisioner never sees these consumers because the mesh never records
|
||||
them as consumers.
|
||||
|
||||
## What a fix has to keep
|
||||
|
||||
- **A local consumer still appears in the provider's grants**, so its provisioner creates the role
|
||||
or bucket or client, exactly as for a remote one.
|
||||
- **The credential is still sealed**, to the one node that is both ends. The mesh holding a
|
||||
readable secret for local consumers would be a hole opened for convenience.
|
||||
- **Refusing must stay refusing.** A requirement nothing answers is still refused; this is about a
|
||||
requirement that *was* answered.
|
||||
|
||||
## Fixed
|
||||
|
||||
A requirement answered on this machine is still a requirement. Resolution now records a need for
|
||||
it, so a credential is made, the provider is told who asked, and the consumer's file is written —
|
||||
the same as if the two were on different machines.
|
||||
|
||||
The reasoning that made it a gap is now written where it was assumed: the machine is not a trust
|
||||
boundary once both ends are containers, and treating it as one gave the commonest arrangement of
|
||||
all — a service and its database on one node — the weakest handling.
|
||||
|
||||
Two later issues came out of the same mistaken instinct and are worth reading together:
|
||||
[`022`](../022-one-credential-per-node-per-provision-not-per-module/00-report.md), where the
|
||||
machine was treated as an *identity* rather than a boundary, and
|
||||
[`023`](../023-a-consumer-cannot-build-a-connection-string/00-report.md), where the consumer was
|
||||
given a password and never told the name to present with it.
|
||||
|
||||
*Closed 2026-09-01. The fix landed the same day and this record was left open by oversight — the
|
||||
code and the tests were in place for hours while the record still said `located`.*
|
||||
@@ -0,0 +1,113 @@
|
||||
---
|
||||
status: fixed
|
||||
opened: 2026-09-01
|
||||
located-in: [mesh-control]
|
||||
fixed-by: mesh-control 0af3ea1
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 022 — A credential belongs to a node and a provision, so a second consumer refuses
|
||||
|
||||
## Symptom
|
||||
|
||||
A node running more than one module that wants the same provision **cannot be planned at all**:
|
||||
|
||||
```
|
||||
anchor has 3 modules asking for "postgres-database" and they would share one
|
||||
credential: gitea, keycloak, umami
|
||||
```
|
||||
|
||||
**And that is only the loud half.** The consuming node does not refuse at all. Three modules
|
||||
wanting one database produce **one** need:
|
||||
|
||||
```
|
||||
modules=3 needs=1
|
||||
name=postgres-database from=anchor for=gitea
|
||||
```
|
||||
|
||||
So the first module gets a credential, the other two get no file at all, and each starts and fails
|
||||
to authenticate with nothing anywhere saying why — the shape of
|
||||
[`021`](../021-a-consumer-on-the-providers-machine-is-given-no-credential/00-report.md), on a
|
||||
different axis. The refusal that reads like a decision is on the provider; the silence is on the
|
||||
consumer.
|
||||
|
||||
The refusal is correct about what it says. They *would* share one credential, and sharing one is
|
||||
worse than refusing — a login that opens three databases is not three credentials. But the
|
||||
arrangement being refused is the ordinary one. **The node this mesh exists to take over runs
|
||||
eight modules against one database server.**
|
||||
|
||||
## Where it comes from
|
||||
|
||||
A credential is keyed by *(provision, consumer node, provider node)*:
|
||||
|
||||
```go
|
||||
func (i *Inventory) SecretFor(ctx context.Context, name, consumer, provider string) (Secret, error)
|
||||
```
|
||||
|
||||
`consumer` is a **node**. Everything downstream inherits that granularity: `Grant.Consumer` is a
|
||||
node, the grant's file is named after a node, and the provisioner names the role it creates after
|
||||
one — `role := mark + c.Node`.
|
||||
|
||||
So the refusal in `ContributionsTo` is not a check that found a problem. It is the only honest
|
||||
thing that function can do, given a key that cannot tell two consumers apart.
|
||||
|
||||
## Why it was not noticed
|
||||
|
||||
**Every scenario so far had one consumer per node.** That is the natural shape of a small test —
|
||||
a consumer here, a provider there — and it is the shape of every lab scenario written to date. A
|
||||
node with two modules wanting a database is not an edge case discovered by fuzzing; it is what a
|
||||
real machine looks like, and nothing had modelled a real machine yet.
|
||||
|
||||
The refusal also reads as deliberate. It names the modules, explains the consequence, and refuses
|
||||
rather than picking — the house rule everywhere else. It looks like a decision. It is a limit.
|
||||
|
||||
## Why the granularity is wrong
|
||||
|
||||
The same argument that closed [`021`](../021-a-consumer-on-the-providers-machine-is-given-no-credential/00-report.md).
|
||||
There, the machine was treated as a trust boundary and containers made that untrue. Here, the
|
||||
machine is treated as an *identity* — as though "who is asking" is answered by naming a host.
|
||||
|
||||
**Two modules on one node are as separate as two on different nodes.** They run as different
|
||||
containers, on different networks, with different data. A key that cannot distinguish them means
|
||||
the mesh cannot express the thing it is for.
|
||||
|
||||
It also silently weakens what the provisioner does. `mesh_<node>` is one role. Had the refusal not
|
||||
been there, gitea's login would have opened keycloak's database — and nothing anywhere would have
|
||||
said so, because from the provisioner's side it created exactly what it was asked to create.
|
||||
|
||||
## What a fix has to keep
|
||||
|
||||
- **The refusal, where it is still right.** Two modules wanting one provision must not silently
|
||||
share a credential. After a fix they do not share one, so there is nothing to refuse — but a
|
||||
genuine collision must still refuse rather than pick.
|
||||
- **A credential per consuming module**, sealed to the node that holds it. Both facts are needed:
|
||||
the module is who it is for, the node is what it is sealed to.
|
||||
- **The provisioner names what it creates after the module**, so a login is traceable to the thing
|
||||
using it, and so withdrawing one consumer does not remove another's.
|
||||
- **Withdrawal still works.** A module unassigned must lose its login while the others keep theirs
|
||||
— which is precisely what one role per node cannot do.
|
||||
- **Existing single-consumer nodes keep working**, since that is every scenario that exists.
|
||||
|
||||
## Scope
|
||||
|
||||
This crosses the control plane, the grant file naming, and every provisioner that names something
|
||||
after `Consumer`. It is not a local fix, and it is the last thing between the current state and a
|
||||
node that looks like a real one.
|
||||
|
||||
## Fixed
|
||||
|
||||
Needs fan out per consuming module in one place, after the resolution walk. The credential's key
|
||||
gains the consuming module, the grant file is named after both halves of the consumer, needs are
|
||||
matched by provision *and* module, and the provisioners name the role and the access key after the
|
||||
module rather than the machine. The refusal is gone because there is nothing left to refuse.
|
||||
|
||||
Existing credentials are discarded rather than backfilled: they cannot say which module they were
|
||||
for, and one is remade and delivered to both ends on the next push, so it costs one rotation.
|
||||
|
||||
A guard was added for PostgreSQL's 63-byte identifier limit, which truncates with a notice rather
|
||||
than an error — two consumers whose role names agree that far would otherwise become one login,
|
||||
which is this same fault at a length nobody would think to test.
|
||||
|
||||
What it did **not** fix is [`023`](../023-a-consumer-cannot-build-a-connection-string/00-report.md):
|
||||
a consumer now receives its own password and still cannot build a connection string, because the
|
||||
user name is the provisioner's invention and the bound values cannot reach a configuration file.
|
||||
@@ -0,0 +1,111 @@
|
||||
---
|
||||
status: fixed
|
||||
opened: 2026-09-01
|
||||
located-in: [mesh-control]
|
||||
fixed-by: mesh-control 122680b
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 023 — A consumer is given every part of a connection except the two it cannot invent
|
||||
|
||||
## Symptom
|
||||
|
||||
A module that requires a database is now given its password in whatever shape its configuration
|
||||
needs ([`022`](../022-one-credential-per-node-per-provision-not-per-module/00-report.md) and the
|
||||
sealed-placeholder work). It still cannot connect, because a password is not a connection.
|
||||
|
||||
What it is given is a **binding**, as JSON:
|
||||
|
||||
```
|
||||
provision postgres-database
|
||||
from the node providing it
|
||||
at that node's address on the private network
|
||||
serves what the provider said a consumer must know — the port
|
||||
```
|
||||
|
||||
What it needs, to write `KC_DB_URL` or `GITEA__database__USER`, is the host, the port, the
|
||||
database name and **the user name**. Two of those are missing, for two different reasons.
|
||||
|
||||
## The user name is nobody's to say
|
||||
|
||||
The provisioner invents it — `mesh_<node>_<module>` — and nothing else in the mesh knows that
|
||||
string. The control plane does not record it, the binding does not carry it, and the consumer
|
||||
cannot derive it without hard-coding another module's naming convention.
|
||||
|
||||
So the one identifier a consumer must present in order to authenticate is the one thing no part
|
||||
of the mesh will tell it. It works today only because nothing has yet had to write a connection
|
||||
string; every proof so far stopped at "the credential arrived".
|
||||
|
||||
## The values cannot reach the file that needs them
|
||||
|
||||
The binding is a JSON document. The consumers are containers reading `KEY=value`, or a program
|
||||
reading a YAML file, or one reading an attribute inside a different JSON document. A sealed secret
|
||||
can now be placed inside any of those — the module writes the file with a hole in it and the host
|
||||
fills the hole on the machine. **The bound values have no such route**, so the half of the
|
||||
connection that is not secret is the half that cannot be delivered.
|
||||
|
||||
This asymmetry is backwards. The secret is the hard case, because the mesh must not be able to
|
||||
read it. The host and port are ordinary facts the mesh knows in the clear, and they are the ones
|
||||
stuck in a document nothing can read.
|
||||
|
||||
## Why it was not noticed
|
||||
|
||||
Every provider so far has been reached by a **provisioner**, a program written for the job, which
|
||||
reads the JSON because it was built to. The first consumers to need a plain configuration file
|
||||
were the first real applications. The binding was designed for the program and then handed to the
|
||||
application.
|
||||
|
||||
There is also a stale comment saying a binding *"carries no credential: the mesh has no way to
|
||||
issue one yet"*. That stopped being true when [`021`](../021-a-consumer-on-the-providers-machine-is-given-no-credential/00-report.md)
|
||||
was fixed.
|
||||
|
||||
## What a fix has to settle
|
||||
|
||||
- **Who names the role.** Either the mesh records what the provisioner will create, or the
|
||||
consumer contributes the name it wants and the provisioner uses it. The second is more in
|
||||
keeping with the rest — a consumer already contributes the database name it wants — and it
|
||||
removes an invented convention rather than documenting one.
|
||||
- **How a bound value reaches a file.** The symmetric answer to the sealed placeholder, and
|
||||
simpler: these values are not secret, so the control plane can put them in before sending and
|
||||
the host learns nothing new.
|
||||
- **That it stays name-agnostic.** The control plane must not learn what a `postgres-database`
|
||||
is. What the keys mean is agreed by the requirement's name
|
||||
([ADR 0027](../../02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md)), so
|
||||
whatever is added has to work for a bucket and a mail relay without being told about either.
|
||||
|
||||
## Blocked on this
|
||||
|
||||
Keycloak, Gitea, Mailu and MinIO all have manifests that parse and resolve, and none of them can
|
||||
start. This is what stands between the module set and a running one.
|
||||
|
||||
## Fixed
|
||||
|
||||
Both halves had the same cause: **the mesh knew something and did not say it.**
|
||||
|
||||
**Who a consumer is, said once.** `ConsumerIdentity(node, module)` is one derivation, sent to the
|
||||
provider in its grant and to the consumer in its binding — so the two agree by construction rather
|
||||
than by two conventions that happened to match on the day they were written. The provisioners now
|
||||
use the name they are given and **refuse to invent one** if the mesh says nothing: falling back to
|
||||
a name of their own would create a login the consumer could never guess, and everything would
|
||||
report success. They also refuse a name that does not carry the mesh's prefix, because that prefix
|
||||
is how withdrawal finds what it made.
|
||||
|
||||
**Bound values reach the file that needs them.** `${bound:provision:key}` is the symmetric twin of
|
||||
the sealed placeholder and simpler: these values are not secret, so the control plane fills them
|
||||
in before sending, and the host gains no field and learns no format. `at`, `as` and `from` are
|
||||
true of any provision; every other key comes from what the provider said it *serves*, so the
|
||||
control plane still learns nothing about what a `postgres-database` is.
|
||||
|
||||
Keycloak and Gitea now produce complete connection strings — asserted from the manifests on disk,
|
||||
checking that every part is filled, that no placeholder survives as a value, and that the password
|
||||
is still a hole only the host can close.
|
||||
|
||||
### Also found, and separate
|
||||
|
||||
The lab run that was meant to prove this failed in a way that looked like the fix being wrong: a
|
||||
rotation test could not authenticate against a real database. The cause was that the suite
|
||||
rebuilt the control plane's image and not the provisioner's, so a run with an image built that
|
||||
minute used a provisioner built the day before. Fixed in `mesh-lab`; it is
|
||||
[`005`](../005-pipeline-test-harness-unbuildable/00-report.md)'s family — a rebuild covering
|
||||
most of what a run uses is worse than one covering none, because the run that follows it is
|
||||
believed.
|
||||
@@ -0,0 +1,106 @@
|
||||
---
|
||||
status: fixed
|
||||
opened: 2026-09-01
|
||||
located-in: [mesh-lab]
|
||||
fixed-by: mesh-lab 3503ad9
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 024 — A run stalls before the host is placed, and says nothing while it does
|
||||
|
||||
## Symptom
|
||||
|
||||
The first scenario of the end-to-end suite raises its machines and then stops. Observed twice on
|
||||
2026-09-01, both times after the rebuild step grew:
|
||||
|
||||
- **Once mid-run**, during the rotation test: the suite had reported thirteen passes, then the
|
||||
process ended with no summary, no failure and no receipt.
|
||||
- **Once from the start**: `a bare machine becomes a mesh` ran for **35 minutes** against a
|
||||
measured 4.5, produced no output at all, and was still running when it was stopped by hand.
|
||||
|
||||
Both machines were `RUNNING` throughout. The second one was interrogated directly: the anchor VM
|
||||
answered, and had **no `mesh-host` log and no containers** — so the run had not reached placing
|
||||
the host. It was stuck earlier, in stocking the scenario's registry.
|
||||
|
||||
## What is not the cause
|
||||
|
||||
- **Not memory.** 84 GiB available, no OOM in the kernel log.
|
||||
- **Not the daemon.** `incus exec` into the stalled machine answered immediately.
|
||||
- **Not the changes under test.** The credential work is applied after the host is placed, and the
|
||||
host was never placed.
|
||||
|
||||
## What changed just before — *and it was not the cause*
|
||||
|
||||
The rebuild step had gone from two artifacts to six, and every image is pushed into the scenario's
|
||||
registry, which looked like where the stall sat. That was written down as a coincidence rather
|
||||
than a diagnosis, and it is as well: **stocking takes 34 seconds and always did.** Timed directly,
|
||||
eight images, before anything was changed.
|
||||
|
||||
The suspicion was the ordinary kind — the thing that changed most recently looks guilty — and the
|
||||
thing that changed had nothing to do with it.
|
||||
|
||||
## Why it matters more than a slow test
|
||||
|
||||
**A stall is indistinguishable from work.** The suite prints nothing between starting a scenario
|
||||
and finishing its first test, so four and a half minutes and thirty-five look identical from the
|
||||
outside — and the operator's only recourse is to guess, which is precisely how a workstation was
|
||||
left unbootable in August by killing a package manager that was working.
|
||||
|
||||
The first stall is worse: the process ended *silently* after thirteen passes. No summary, no
|
||||
receipt, nothing that says the run was cut short. A run that stops without saying so is a run
|
||||
somebody may believe.
|
||||
|
||||
## What a fix has to give
|
||||
|
||||
- **Progress while stocking**, so a long step is visibly a long step. Bytes moved, images pushed,
|
||||
anything that changes.
|
||||
- **A receipt when a run is cut short**, saying how far it got. `lastrun` already refuses to guess;
|
||||
what is missing is it being written at all when the process dies mid-run.
|
||||
- **A stated timeout on stocking**, so a stall ends as a failure with a reason rather than as a
|
||||
process somebody eventually kills.
|
||||
|
||||
## The cause
|
||||
|
||||
**The registry machine was addressed by hand and every other machine was not.**
|
||||
|
||||
Machines get a systemd-networkd unit with a static `Address=`, so networkd finishes configuring
|
||||
the link and reports it `configured`. The registry instead ran `ip addr add` inline. An address
|
||||
put on a link that way leaves networkd still waiting to configure something it was never told
|
||||
about, so the link sits at `configuring` — and `systemd-networkd-wait-online` has
|
||||
`TimeoutStartUSec=infinity`.
|
||||
|
||||
So `network-online.target` is never reached, and **everything ordered after it never starts.** On
|
||||
these machines that is Docker. `docker load` then blocks on a socket whose daemon is queued behind
|
||||
a target that will never come, and the three bounded timeouts around it — save, push, load — stack
|
||||
to thirty-five minutes.
|
||||
|
||||
Measured on one scenario, before and after:
|
||||
|
||||
| | before | after |
|
||||
|---|---|---|
|
||||
| the registry's link | `configuring` | `configured` |
|
||||
| `docker.service` | inactive, 5 jobs pending | active, no jobs |
|
||||
| the raise | never finished | **87.5 s** |
|
||||
|
||||
These machines have **no DHCP by design** — a scenario is a closed address space and the
|
||||
declaration owns the addresses — so nothing was ever going to complete that wait.
|
||||
|
||||
## Fixed
|
||||
|
||||
- **The registry is addressed the way every other machine is**, through the same helper.
|
||||
- **Placing an image waits for the container runtime** and refuses after 120s, naming what systemd
|
||||
is still waiting on. A stall becomes a failure that says why.
|
||||
- **The end-to-end test passes `onProgress`.** The raise reported every step and the test threw it
|
||||
away, which is why thirty-five minutes of silence and four minutes of silence looked the same.
|
||||
|
||||
The suite then ran to completion: **23 of 24**, the one failure a check of its own that flagged
|
||||
`/var/lib/mesh/builder/broker` as a credential because `/` is in the base64 alphabet. Fixed with
|
||||
it.
|
||||
|
||||
## Also learned, at some cost
|
||||
|
||||
**A redirected log lags.** Node block-buffers stdout when it is a file, so `> run.log` sits
|
||||
unchanged for minutes while the run is fine. That was read as a stall twice — the second time
|
||||
immediately after the real fix, where a buffering artifact argues the fix did not work. The
|
||||
machines answer instantly and are the source of truth. *"I cannot see progress" is not evidence of
|
||||
no progress.*
|
||||
@@ -0,0 +1,121 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-09-01
|
||||
located-in: [mesh-control, mesh-host]
|
||||
fixed-by: partly — mesh-control ee3cc1b
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 025 — A module must pin a digest, and nothing produces one
|
||||
|
||||
## Symptom
|
||||
|
||||
Every image reference in every example module is **sixty-four zeros**:
|
||||
|
||||
```
|
||||
gitea@sha256:0000000000000000000000000000000000000000000000000000000000000000
|
||||
```
|
||||
|
||||
Eighteen of them, across five modules. Each one parses, resolves, and composes into a declaration
|
||||
a host accepts. None of them could ever start: the machine would reach `docker pull` and stop.
|
||||
|
||||
This is why those modules are *written* and not *running*, and it was not visible from any check
|
||||
because every check passes.
|
||||
|
||||
## Why nothing caught it
|
||||
|
||||
The host validates the **shape** of a reference and nothing else — that it is `name@sha256:` plus
|
||||
sixty-four hexadecimal characters. Sixty-four zeros satisfies that exactly.
|
||||
|
||||
That check is not wrong. A host cannot verify a digest exists without reaching a registry, and
|
||||
reaching a registry is precisely what the design refuses to make it do
|
||||
([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)). The host is the last
|
||||
place that could catch this and the wrong place to try.
|
||||
|
||||
## The actual gap
|
||||
|
||||
**A manifest must carry a digest, and nothing in the system produces one.**
|
||||
|
||||
- Images the mesh builds are fine: the bundle writes the digest down *after* building, which is
|
||||
the whole reason the bundle exists in that shape.
|
||||
- Images from anywhere else — a forge, a mail system, a database — have no path at all. Somebody
|
||||
has to look up what `gitea:1.22` points at today and paste it in, and nothing re-checks it.
|
||||
|
||||
So the design is coherent about *pinning* and silent about *where a pin comes from*. A person
|
||||
writing a module is asked for something they cannot reasonably produce by hand, and given a
|
||||
placeholder shape that passes every gate.
|
||||
|
||||
## What a fix has to keep
|
||||
|
||||
- **The host still refuses a tag.** A digest is what makes a declaration exact, and that must not
|
||||
soften. The fix belongs where a module is added or built, not on the machine.
|
||||
- **A person writes a tag; the mesh holds a digest.** A manifest in a repository naming
|
||||
`gitea:1.22` is readable and reviewable; the mesh resolving that against a registry once, and
|
||||
recording the answer, is what makes it exact. The module table already records this shape for
|
||||
source repositories — where it came from, the branch followed, the commit read — and an image is
|
||||
the same question asked of a registry.
|
||||
- **Re-resolving is a decision, not a side effect.** A tag that moves must not silently change what
|
||||
a machine runs. Whatever resolves it records both, so *this pin is behind its tag* is a question
|
||||
the mesh can answer rather than something discovered on a restart.
|
||||
|
||||
## Cheaply, now
|
||||
|
||||
An all-zero digest is a placeholder and never a real image. Refusing it costs three lines and
|
||||
would have caught all eighteen the day they were written. It does not fix the gap; it stops the
|
||||
gap being invisible.
|
||||
|
||||
## What this blocks
|
||||
|
||||
Every module that names a third-party image, which is every module that is not the mesh itself.
|
||||
The forge and the mail system are otherwise ready to run.
|
||||
|
||||
## Half of it is done
|
||||
|
||||
**A placeholder can no longer reach a machine.** The refusal sits where a declaration is composed,
|
||||
not where a manifest is parsed — a file awaiting a pin is legitimate, and the design already says
|
||||
so for artifacts the mesh builds. Composing is the last moment before a machine sees it.
|
||||
|
||||
**The examples now pin images that exist.** Twelve third-party digests were resolved against their
|
||||
registries without pulling anything, which is also the mechanism the rest of this issue needs:
|
||||
`docker manifest inspect --verbose` answers *what does this tag point at* in about a second.
|
||||
|
||||
Two faults came free, and both had been invisible for the same reason as the digests: the mail
|
||||
system's seven images named repositories that **do not exist** — it publishes to a different
|
||||
registry entirely — and one of the seven had been renamed upstream. Nothing that only checks the
|
||||
shape of a reference could ever have found either.
|
||||
|
||||
## The mechanism already existed, and this issue was wrong about that
|
||||
|
||||
**Corrected 2026-09-01, the same day.** This was filed saying nothing turns a tag into a digest.
|
||||
That is false, and the answer had been designed and built before any of it was written.
|
||||
|
||||
A module does not name an image at all. It names an **artifact**, and declares where that artifact
|
||||
comes from:
|
||||
|
||||
```
|
||||
build.artifacts: [{ name: "gitea", kind: "upstream", from: "gitea/gitea:1.22" }]
|
||||
resources: [{ id: "server", type: "container", artifact: "gitea", … }]
|
||||
```
|
||||
|
||||
`kind: upstream` means *an image somebody else built, mirrored into the mesh's own registry and
|
||||
pinned by the digest it lands with*. The builder produces it; the manifest the mesh holds is
|
||||
derived, with `artifact` replaced by the real reference and the key removed, because the host has
|
||||
never heard of that word. A resource naming an artifact nothing produced is refused.
|
||||
|
||||
So the two-document split this issue described as the shape of a fix **is the design**, and it
|
||||
covers both cases it said were unsolved: an image the mesh builds, and an image somebody else
|
||||
built. Mirroring also removes something worse than a stale pin — every machine needing a route to
|
||||
a public registry, and a tag a stranger can move.
|
||||
|
||||
**What was actually wrong was the examples.** They hard-coded image references instead of naming
|
||||
artifacts, so they inherited a problem the design does not have. Pinning twelve of them by hand
|
||||
was treating the symptom, and left the reference pointing at a public registry rather than the
|
||||
mesh's own.
|
||||
|
||||
## What is still open
|
||||
|
||||
**The examples should name upstream artifacts** rather than carry hand-pinned digests. That is the
|
||||
remaining work, and it is a rewrite of five manifests rather than a mechanism to build.
|
||||
|
||||
The refusal added here stays: a placeholder must not reach a machine whatever the reason it is
|
||||
there.
|
||||
@@ -0,0 +1,96 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-09-01
|
||||
located-in: [mesh-control]
|
||||
fixed-by: partly — mesh-control 53eb000, withdrawn in 83c6a2f
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 026 — The data directories are mounted and never declared
|
||||
|
||||
## Symptom
|
||||
|
||||
Four modules mount **fourteen host paths** that no resource in those modules declares:
|
||||
|
||||
```
|
||||
gitea /services/gitea/gitea
|
||||
postgres /services/postgres/db-data
|
||||
minio /services/minio/data/data1-1
|
||||
mailu eleven more, including the mail spool and the admin database
|
||||
```
|
||||
|
||||
Each is a bind mount on a container. None is a `directory` resource. The mesh has never heard of
|
||||
any of them.
|
||||
|
||||
## What that costs
|
||||
|
||||
**They are created by the container runtime, as root.** A bind mount whose source does not exist
|
||||
is created for you, owned by root, with whatever mode the runtime picks. So `owner` and `mode` —
|
||||
which exist precisely so a module can say who its data belongs to — are silently not applied to
|
||||
the only directories that hold data.
|
||||
|
||||
**The protection that exists for exactly this does not reach them.** A directory the mesh declared
|
||||
and no longer wants is *kept*, not removed, when it holds anything the mesh did not put there
|
||||
([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)). That rule is the
|
||||
answer to *what happens to my data when a module goes away*, and it is written in terms of
|
||||
declared directories. **An undeclared one is not protected by it, because the mesh does not know
|
||||
it is there.**
|
||||
|
||||
So the single rule guarding against data loss covers the configuration directories, which are
|
||||
cheap to lose, and not the data directories, which are the reason the rule exists.
|
||||
|
||||
## Where it came from
|
||||
|
||||
These manifests were written by reading the arrangement being replaced and carrying its
|
||||
`docker-compose` files across — service, image, ports, volumes, environment — into the new
|
||||
manifest's container shape. That shape can express all of it, which is what made the
|
||||
transliteration feel like progress.
|
||||
|
||||
**A container shape that can express a compose file will be filled in like a compose file.** The
|
||||
mesh's model is larger than that: a directory is a thing the mesh owns, with an owner and a mode
|
||||
and a rule about what happens when it is no longer wanted. A volume line borrowed from compose
|
||||
declares none of it, and nothing complains, because a bind mount source is a string.
|
||||
|
||||
## What a fix has to keep
|
||||
|
||||
- **Every host path a container mounts is declared.** If a module wants a directory on the
|
||||
machine, it says so, with who owns it and what mode — and gets the removal rule with it.
|
||||
- **The check is mechanical.** A person comparing volumes against declared directories by hand is
|
||||
the process that produced this. It is a few lines against the manifest and belongs beside the
|
||||
other manifest checks.
|
||||
- **Not by inventing directories at apply time.** The host creating what a mount needs would make
|
||||
the mesh's ownership of a directory depend on which resource mentioned it first, and would put
|
||||
the same undeclared path back a layer down.
|
||||
|
||||
## Not yet answered
|
||||
|
||||
**Where a module's data should live at all.** These paths were inherited whole from the
|
||||
arrangement being replaced, which put everything under one directory per service. Whether that is
|
||||
right here is a separate question, and a bigger one — it decides what a person backs up, and what
|
||||
survives a module being removed.
|
||||
|
||||
## Half fixed
|
||||
|
||||
**All fourteen are declared**, across the forge, the mail system, the store and the object store —
|
||||
each mount now resolves to a `directory` or to a file the module already names.
|
||||
|
||||
**Enforcing it was tried and withdrawn**, and the withdrawal is the interesting half. A refusal
|
||||
for any mount no resource declares refuses the **builder**, which mounts the container runtime's
|
||||
socket. That socket is not the builder's data. It does not belong to the module, it already exists,
|
||||
and declaring it as one of the module's own directories would be a lie that the host would act on.
|
||||
|
||||
So the rule is right about data and wrong about everything else, because the manifest cannot
|
||||
currently say which a path is. Two kinds of mount are spelled identically:
|
||||
|
||||
- **the directory my data lives in** — created if absent, owned by the module, protected by
|
||||
[ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)
|
||||
- **a machine facility I was granted** — a socket, a device; it exists, the machine owns it, and
|
||||
the module is being given access to it
|
||||
|
||||
`capabilities` is the closest existing thing to the second and names no paths. Inventing a field
|
||||
to separate them is a design decision, so it is recorded here rather than made to get a check
|
||||
green.
|
||||
|
||||
**Until then the manifests are right by coincidence**, which is the state this issue was opened
|
||||
about. What is kept is a check that every real manifest still parses — worth nothing against this
|
||||
fault, and the reason the next attempt finds out in a second rather than in a fifteen-minute run.
|
||||
@@ -0,0 +1,65 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-09-01
|
||||
located-in: [mesh-host, mesh-control]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 027 — A container cannot follow a file, and a rotated credential is the case
|
||||
|
||||
## Symptom
|
||||
|
||||
A service can say `restart-on`: *these files changed, so I must be restarted*. A container cannot.
|
||||
It is not in the shape, and the host refuses a declaration that tries.
|
||||
|
||||
So a container reading its password from a file keeps the password it started with, for ever.
|
||||
Nothing reports anything: the file is right, the container is up, every check passes.
|
||||
|
||||
## Why this is the same fault the mechanism exists for
|
||||
|
||||
`restart-on` is written against exactly this, in the host's own words:
|
||||
|
||||
> a running service does not re-read its configuration. Replace the file, find the service already
|
||||
> running, do nothing, and the machine keeps behaving the way it did before — while every check
|
||||
> passes, because the file is right and the service is up.
|
||||
|
||||
Every word applies to a container, and more so. **Nearly everything the mesh runs is a container**
|
||||
— a database, a forge, a mail system — and a credential arrives as a file it reads at start.
|
||||
|
||||
## What it costs, concretely
|
||||
|
||||
**Rotation does not reach a container.** Rotating a credential replaces the file on the machine and
|
||||
tells the provider to accept the new one. The provider is a program that reconciles, so it takes
|
||||
the change. The consumer is usually a container, so it does not. The two ends then hold different
|
||||
passwords, which is the fault the whole design is arranged to prevent
|
||||
([ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) records it costing
|
||||
two days).
|
||||
|
||||
The end-to-end test that proves rotation works uses a consumer that reads the file on each
|
||||
attempt, so it does not meet this.
|
||||
|
||||
## Why it was not noticed
|
||||
|
||||
The gap is invisible from the control plane. A manifest carrying `restart-on` on a container
|
||||
composes into a declaration without complaint and is refused on the machine, so the only way to
|
||||
learn is to run one — which is how it was found, after nine of them had shipped across seven
|
||||
modules.
|
||||
|
||||
## What a fix has to keep
|
||||
|
||||
- **Declared state, not a command.** `restart-on` is deliberately not *restart this*: it says the
|
||||
running thing must reflect these files, and the host works out that it does not. Whatever
|
||||
containers get must keep that shape, because the link may not carry an action
|
||||
([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)).
|
||||
- **Recreate, not restart, where that is the honest verb.** A container's environment is fixed at
|
||||
creation. If what changed is an `env-file`, restarting the container is not enough — it has to be
|
||||
made again. That is a different act from a service reload and should not be described as one.
|
||||
- **It must not fire on every reconcile.** A container that is recreated whenever the host looks at
|
||||
it is worse than one that never follows the file.
|
||||
|
||||
## The near alternative, and why it is not enough
|
||||
|
||||
A module can avoid the problem by having its program read the file on each use rather than at
|
||||
start. That works for something written for this mesh, and is not available for a database, a
|
||||
forge or a mail system — which is the whole population this is about.
|
||||
@@ -0,0 +1,99 @@
|
||||
---
|
||||
status: fixed
|
||||
opened: 2026-09-01
|
||||
located-in: [mesh-control, mesh-host]
|
||||
fixed-by: mesh-control 1f5b70a, 41f7c51; mesh-host b91342a
|
||||
amended-design: 02-DECISIONS/0038-the-mesh-assigns-the-port.md
|
||||
---
|
||||
|
||||
# 028 — Two things want one port, and nothing says so until the machine
|
||||
|
||||
## Symptom
|
||||
|
||||
The database module cannot start on a machine that runs the control plane:
|
||||
|
||||
```
|
||||
Bind for 127.0.0.1:5432 failed: port is already allocated
|
||||
```
|
||||
|
||||
The mesh keeps its own store on that machine, from the bundle, and it holds 5432. The module
|
||||
publishes 5432 too. Everything up to the machine is content: it resolves, it composes, it is
|
||||
pushed, and it is applied — the container is simply the one resource that fails.
|
||||
|
||||
## Why nothing catches it
|
||||
|
||||
**The substrate is not a module.** It arrives from the bundle a host carries, before there is a
|
||||
mesh to ask. So the control plane has never heard of `mesh-store` and does not know it holds a
|
||||
port. Resolution can compare modules against each other and cannot compare a module against the
|
||||
thing the mesh is built on.
|
||||
|
||||
**And nothing compares modules against each other either.** A port is exclusive on a machine in
|
||||
exactly the way a claim is — one seat, one display server, one artifact store — and the mesh has a
|
||||
mechanism for that, which ports do not use. Two modules both publishing 5432 would meet the same
|
||||
wall, one machine later.
|
||||
|
||||
## It has been met before, and worked around
|
||||
|
||||
The end-to-end test that exercises a real database publishes `5433:5432` rather than `5432:5432`.
|
||||
The workaround is right there, inline, with no note saying why — which is how a constraint becomes
|
||||
folklore.
|
||||
|
||||
## What a fix has to settle
|
||||
|
||||
- **Whether a module should publish to the machine at all.** Consumers reach a provider by the
|
||||
machine's address and the port it *serves*, so publishing is what makes that true. An alternative
|
||||
is that they reach it on the module's own network by name, and nothing is published — which
|
||||
changes what `serves` means and is a larger decision than it looks.
|
||||
- **Where the substrate's ports are written down.** Whatever compares them needs to know what the
|
||||
bundle holds. The bundle is a list of pinned references; what those containers bind is not in it.
|
||||
- **What a refusal should say.** *5432 is held by the mesh's own store on this machine* is a useful
|
||||
sentence. *Port is already allocated*, arriving from a container runtime three layers down, is
|
||||
not.
|
||||
|
||||
## Not the same as a firewall rule
|
||||
|
||||
`listens` already says which ports a module accepts on, and filtering is computed from it. That is
|
||||
about what may reach a port from elsewhere. This is about two things on one machine wanting to own
|
||||
the same one, which `listens` does not model and could not answer.
|
||||
|
||||
## Answered in principle
|
||||
|
||||
[ADR 0038](../../02-DECISIONS/0038-the-mesh-assigns-the-port.md), proposed the same day: **the mesh
|
||||
assigns the machine-side port and a module does not care.** A module cannot choose well, because it
|
||||
is written once and assigned anywhere — any number it picks is a guess about a machine it has never
|
||||
seen.
|
||||
|
||||
A port fixed by its protocol — mail on 25, submission on 587 — becomes a **claim**, which is the
|
||||
mechanism the mesh already has for what is singular on a machine. Two modules wanting 25 is the
|
||||
same shape as two wanting the seat, and earns the same refusal at assignment rather than at apply.
|
||||
|
||||
The record also names what this issue missed: the same number is written **three times** in every
|
||||
module — once for the rule set, once for what a consumer is told, once for what the runtime
|
||||
publishes — and nothing checks that they agree. A module whose `serves` and whose container
|
||||
disagreed would hand every consumer a port that answers nothing.
|
||||
|
||||
## Fixed
|
||||
|
||||
**The mesh assigns the machine-side port**, from a high unprivileged range, recorded per machine
|
||||
and module and kept once chosen. A module says the port its software uses, once, in `listens`. The
|
||||
container's mapping, the rule set, and what a consumer is told are all derived from the assignment
|
||||
— so the three copies that agreed only because one person wrote them are now one fact.
|
||||
|
||||
**A port the protocol fixes says so**, and is then a claim: one holder per machine, and the second
|
||||
refused by name at assignment rather than by a container runtime at apply.
|
||||
|
||||
**And the machine says what it already holds.** This was the half that made the issue: the
|
||||
substrate is not a module, so nothing in the mesh had heard of the store or the broker. The host
|
||||
already distinguished what it carried from what the mesh sent — that distinction exists so the two
|
||||
never remove each other — and now records what each resource binds and reports the carried ones.
|
||||
The allocator treats those as taken.
|
||||
|
||||
What the declaration binds, not what is open: a machine's open ports are a moving target, and
|
||||
assigning around those would mean a port that was free when it was asked for and taken when it was
|
||||
used.
|
||||
|
||||
## What it does not settle
|
||||
|
||||
The question underneath, unchanged: **whether a module should publish to the machine at all**.
|
||||
Assignment makes publishing safe without making it necessary, and consumers reaching a provider on
|
||||
the module's own network by name would make the question moot for anything inside the mesh.
|
||||
+104
@@ -0,0 +1,104 @@
|
||||
---
|
||||
status: fixed
|
||||
opened: 2026-09-01
|
||||
located-in: [mesh-control]
|
||||
fixed-by: mesh-control be62f49; mesh-lab f85dbb0
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 029 — The artifact store cannot be delivered by the artifact store
|
||||
|
||||
## Symptom
|
||||
|
||||
A mesh that has just bootstrapped cannot install a registry. The module describing one resolves,
|
||||
composes and pushes; the build never completes, because there is nowhere to put what it builds.
|
||||
|
||||
It has never been seen, because the lab always has a registry standing before the mesh asks for
|
||||
one, and so does any mesh built on a machine that already had one.
|
||||
|
||||
## What it actually blocks
|
||||
|
||||
Not "a registry cannot be installed" — **a mesh cannot hold its own modules.**
|
||||
|
||||
The artifact store is one store for everything a module ships: images, and archives, which are
|
||||
directories from a module's repository packed and pushed as content-addressed blobs. So it is the
|
||||
mesh's module catalogue in artefact form, and the same role a separate object store plays in the
|
||||
arrangement being replaced.
|
||||
|
||||
Until it exists, a module can be described and resolved but nothing it carries can be kept
|
||||
anywhere. A mesh that has just bootstrapped can therefore run only what its bundle already holds.
|
||||
|
||||
## The cycle
|
||||
|
||||
Three facts, each correct on its own:
|
||||
|
||||
- **A registry module's image is mirrored in.** `kind: upstream` pulls the reference the module
|
||||
names and pushes it under a name of the mesh's own, so what a machine fetches is pinned by a
|
||||
digest this mesh assigned rather than by a tag somebody else can move.
|
||||
- **The builder publishes to the artifact store**, learned from its `artifact-store` binding —
|
||||
the same binding any consumer of any provision gets.
|
||||
- **The builder refuses to run without one**: *"a built artifact nobody can fetch is not built."*
|
||||
|
||||
So installing the thing that provides `artifact-store` requires something that provides
|
||||
`artifact-store`.
|
||||
|
||||
## Why the design already answers it
|
||||
|
||||
The substrate record asks of each candidate *can it grant itself the thing it provides?* The store
|
||||
cannot create its own database; the broker cannot create its own virtual host; and the registry
|
||||
**cannot grant itself a repository**. That is why the registry is substrate by role.
|
||||
|
||||
The same sentence answers this. A module that provides the artifact store cannot be delivered
|
||||
through the artifact store, so its image is not the mesh's to mirror — it is named directly and
|
||||
pulled from upstream exactly once, which is what the bundle already does for the three images a
|
||||
first node starts from.
|
||||
|
||||
## The fix
|
||||
|
||||
**The registry module names its image, and is never built.** A container naming
|
||||
`registry@sha256:…` needs no builder, no binding and no store. Every module after it mirrors
|
||||
normally, into the registry now running.
|
||||
|
||||
Nothing new is required: naming an image directly is what most modules do.
|
||||
|
||||
## What it costs, and what it does not
|
||||
|
||||
The machine running the first registry needs to reach a public registry once, to pull that one
|
||||
image by digest. That is already true of a first node, which fetches three images the same way
|
||||
before a mesh exists.
|
||||
|
||||
It does not weaken pinning. A digest is exact wherever it came from; what mirroring adds is that
|
||||
the mesh keeps its own copy and does not depend on a tag somebody else controls. For one image, on
|
||||
one machine, once, the bundle already accepts that trade and says why.
|
||||
|
||||
## The constraint this puts on a registry module
|
||||
|
||||
**A module providing `artifact-store` may not build artifacts of its own** — not its image, and
|
||||
not a user interface or a tool server shipped beside it. There is nowhere to put them until it is
|
||||
running.
|
||||
|
||||
A registry that wants more than the upstream image is therefore two modules: one that provides the
|
||||
store and only names an image, and an ordinary module beside it that builds whatever else and
|
||||
mirrors it in the normal way. That is a real limit on how a registry module can be written, and it
|
||||
should be said out loud rather than discovered.
|
||||
|
||||
## How it is checked
|
||||
|
||||
A manifest that both provides `artifact-store` and declares built artifacts is refused where it is
|
||||
written, naming the cycle. Otherwise the fault surfaces as a build that never returns, on a mesh
|
||||
too new to have anybody watching it.
|
||||
|
||||
## Fixed
|
||||
|
||||
**The registry module names its image and is added, not built.** The lab's artifact-store test now
|
||||
walks the only path open to a real first mesh — the manifest goes in directly, no builder and no
|
||||
store involved — and passes.
|
||||
|
||||
**And the cycle is refused where it is written.** A manifest that provides `artifact-store` and
|
||||
also declares built artifacts is refused at parse, naming the cycle: building publishes to the
|
||||
store, so it asks the mesh to put an artifact into the thing that artifact is needed to create.
|
||||
The provision name became a constant for the rule to turn on.
|
||||
|
||||
What surfaced it in the lab is worth keeping: the test had always *built* the registry module and
|
||||
passed, because the scenario's stand-in for the public registry was standing there to receive the
|
||||
push. A prop that quietly covers for the thing under test is how a bootstrap hole stays invisible.
|
||||
@@ -0,0 +1,56 @@
|
||||
---
|
||||
status: fixed
|
||||
opened: 2026-09-01
|
||||
located-in: [mesh-control]
|
||||
fixed-by: mesh-control 38d4e77
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 030 — Asking what a machine should be re-signed its certificate
|
||||
|
||||
## Symptom
|
||||
|
||||
Every machine carrying a certificate reported as *waiting* — not running what the mesh would send
|
||||
it — for ever. Pushed seconds ago, already behind again. Nothing wrong, nothing failed, nothing
|
||||
quiet; only a comparison that never came out equal.
|
||||
|
||||
It surfaced as the one red test in four consecutive runs, and wore three other faults' clothes
|
||||
first: a test racing the apply it asserted on, a status command that wrote to the database it was
|
||||
reading, a machine starved at its default size. Each was real; each was fixed; the symptom stayed.
|
||||
|
||||
## Cause
|
||||
|
||||
The mesh signs a certificate for a machine's internal name as part of composing its declaration —
|
||||
and signed **anew on every composition**. Same authority, same key, same name, same validity
|
||||
window; a fresh random serial each time, because that is what signing does. So the declaration
|
||||
composed to answer *is this machine current* differed from the declaration sent by exactly one
|
||||
serial number, every time, deterministically.
|
||||
|
||||
The comparison is a digest, so one changed byte is as unequal as a different world.
|
||||
|
||||
## How it was found, which is the lesson
|
||||
|
||||
Not by deduction — deduction produced the three wrong theories above. The suite was run once with
|
||||
its scenario kept standing, and the standing mesh was asked twice: `plan`, `plan`, diff. Two
|
||||
answers seconds apart, identical to the byte but for one serial, in the certificate file. The
|
||||
diff had one line where four theories had none.
|
||||
|
||||
A verdict machine that can be kept and interrogated is worth more than the verdict.
|
||||
|
||||
## The rule it broke, third find of its kind
|
||||
|
||||
**Issued once and kept.** The port had it, the secret had it, the certificate did not — composed
|
||||
fresh on every asking, by the same code that holds the other two still. And like
|
||||
[`028`](../028-two-things-want-one-port-and-nothing-says-so/00-report.md)'s ReleasePorts, the
|
||||
keeping was designed and never wired: the serving-key migration added a column *"and what was
|
||||
issued for it"*, and nothing wrote it.
|
||||
|
||||
A kept certificate stands while the name, the key and the clock agree. A node that rejoined with
|
||||
a new key or changed its name gets a fresh signing exactly as if nothing were kept; so does one
|
||||
whose certificate is into its last tenth of life.
|
||||
|
||||
## Verified
|
||||
|
||||
Live, on the kept mesh, before any suite run: one push with the fixed binary and the machine
|
||||
settled; the second machine likewise; then the mesh's own sentence — all doing what they were
|
||||
told, running what the mesh would send them.
|
||||
@@ -0,0 +1,42 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-02
|
||||
located-in: [mesh-host]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 031 — A machine becomes each thing it was told, in turn
|
||||
|
||||
## Symptom
|
||||
|
||||
A machine that was pushed several declarations in quick succession applies every one of them,
|
||||
oldest first, at the better part of a minute each. Under the lab's suite — twenty-odd pushes in
|
||||
fifteen minutes — the anchor machine ran minutes behind the newest declaration, and a test that
|
||||
waited for it honestly timed out while the machine was busy becoming things nobody wanted any
|
||||
more.
|
||||
|
||||
Only visible since caught-up became an equality: each report now names the declaration it applied,
|
||||
and the reports arriving were about ever-older ones. Before that, the same backlog hid inside
|
||||
timestamp comparisons that happened to pass.
|
||||
|
||||
## Why it is wrong, and why it is also right
|
||||
|
||||
Each declaration is complete — the whole machine, not a delta — so applying an old one is never
|
||||
*incorrect*, only wasted: the machine converges to a state the mesh has already moved past, then
|
||||
does it again. The queue keeps a disconnected machine's instructions safe, which is right. What
|
||||
is wrong is only the order of consumption: **a machine asked to be five successive things should
|
||||
become the last one.**
|
||||
|
||||
## The shape of a fix
|
||||
|
||||
On waking with a queue, drain it and apply only the newest declaration; acknowledge the
|
||||
superseded ones without applying them. Whether a superseded declaration deserves a report — and
|
||||
what its outcome should be called — is the real question for the link's vocabulary: silence reads
|
||||
as a machine that ignored an instruction, and "applied" would be a lie.
|
||||
|
||||
## What it costs today
|
||||
|
||||
Nothing on a real mesh at rest: pushes are far apart. It costs the lab about a doubling of one
|
||||
test's wait, and it will cost a real mesh exactly when things are busiest — a flurry of changes is
|
||||
when a machine can least afford to replay history.
|
||||
@@ -0,0 +1,135 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-04
|
||||
located-in: [mesh-sdk, mesh-catalog]
|
||||
fixed-by: mesh-sdk src/provisioner rework + redis/postgres/minio/umami adapters (ADR 0048)
|
||||
amended-design: 0048-a-provider-creates-the-credential-the-mesh-minted.md
|
||||
---
|
||||
|
||||
# A provider's provisioner seals with a key the mesh has no way to deliver — and does not need to
|
||||
|
||||
## What was observed
|
||||
|
||||
Building the vertical slice for the module runtime (the module runs its own code as its own
|
||||
process under its own account), a **provider** module — one that stands up a per-consumer
|
||||
resource and hands back a credential — was assigned to a node and run as a broker-bound
|
||||
runtime. The runtime hosts the module's provisioner (the sdk's `runProvisioner`), and the
|
||||
harness opens by reading a **seal key** from `$MESH_SEAL_KEY`, failing immediately without
|
||||
one. Every credential it produces for a consumer is sealed to that key with the sdk's
|
||||
symmetric `seal()` (AES-256-GCM, `mesh-sdk/src/primitives/index.ts`) before being written.
|
||||
|
||||
Nothing in the mesh sets `$MESH_SEAL_KEY`. It is read in exactly two places in the sdk and
|
||||
set nowhere — no manifest, no control-plane code, no host code. So a provider runtime, as
|
||||
delivered, aborts at start-up. The slice proved the mechanism only by setting a lab-local key
|
||||
in the manifest by hand.
|
||||
|
||||
## What a trace of the credential path turned up
|
||||
|
||||
The seal key is not a missing delivery. **The whole symmetric-seal provisioner is orphaned,
|
||||
and it duplicates — badly — a job the mesh already does.**
|
||||
|
||||
- `runProvisioner` reads request files named `*.grant.json`. **Nothing writes those.**
|
||||
- It writes sealed credential files named `<consumer>.<resource>.credential`. **Nothing reads
|
||||
those** — not the host, not the control plane. The host reports applied-resource digests
|
||||
upward and never ships credentials; the control plane has no reference to that filename.
|
||||
- No consumer ever calls the symmetric `unseal()`. Consumers receive **plaintext**.
|
||||
|
||||
Meanwhile the mesh already carries a provider→consumer credential across nodes, with **no
|
||||
shared key anywhere**:
|
||||
|
||||
- The control plane mints the password once (`secrets.Make`) and seals it **twice,
|
||||
asymmetrically** — `ForConsumer` to the consumer node's X25519 public key, `ForProvider` to
|
||||
the provider node's (`mesh-control/internal/secrets/seal.go`, `mesh-host/internal/identity/
|
||||
sealing.go`, NaCl box).
|
||||
- Each host opens its own copy with its own private key on the machine; the plaintext exists
|
||||
only for the length of one function call (`mesh-host/internal/apply/apply.go`, the
|
||||
`${secret:name}` substitution — ADR 0024's "the host is the only thing that ever holds
|
||||
both").
|
||||
- `serves` carries no credential and says so; `receives`/`bound` tell each side *where* its
|
||||
sealed secret is, never the value.
|
||||
|
||||
The two models also **contradict** each other. The sdk's `seal()` comment says the key is "a
|
||||
per-node passphrase the host holds"; the host holds no such passphrase — it holds an X25519
|
||||
private key, and the control plane's own code refuses a shared symmetric key on principle:
|
||||
"a key both ends hold is a key the mesh would have to distribute, which is this problem again
|
||||
one level down" (`secrets/seal.go`). A symmetric `MESH_SEAL_KEY` shared between a provider
|
||||
node and a consumer node is exactly the thing the mesh was built not to have.
|
||||
|
||||
And the provisioner's model is wrong in a second way: its adapter **generates its own
|
||||
password** (`generatePassword()`) and creates the resource with it — a different password from
|
||||
the one the mesh mints and hands the consumer. Even with a seal key delivered, a consumer
|
||||
would authenticate with the mesh's password against a resource created with the provisioner's.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
This is not a four-module problem. The provider contract lives in **one place** — the sdk's
|
||||
`runProvisioner(resource, adapter)` harness — and every provider is built on it. Four exist
|
||||
today (redis, postgres, minio, umami); a mesh of any size ends up with many. Whatever the
|
||||
provisioner harness does, every present and future provider inherits, so the orphaned
|
||||
symmetric seal is a fault stamped into the interface, not into four adapters. That also sets
|
||||
the cost of getting it wrong: a contract N providers depend on is N migrations to change
|
||||
later, which is the argument for settling it deliberately now rather than patching around it.
|
||||
|
||||
As written, each provider carries a provisioner that cannot start (no key), and that, if it
|
||||
did, would create resources with a password it invented — a *different* password from the one
|
||||
the mesh minted and handed the consumer — and seal them for a reader that does not exist. The
|
||||
rule the design states, "a consumer receives a sealed credential and unseals it," is enforced
|
||||
by nothing: no consumer unseals, and no shared key exists to unseal with.
|
||||
|
||||
## The mesh already does this — confirmed
|
||||
|
||||
The premise the fix rests on is not a hope; it is in the control plane today. For a served
|
||||
interface, `Inventory.SecretFor` mints one password per (consumer, provider) pair via
|
||||
`secrets.Make`, sealing it to **both** node keys — `ForConsumer` and `ForProvider`.
|
||||
`SecretsFrom(provider)` is documented as "every credential a provider node was issued, so it
|
||||
can be told what to create," and `grantsFor` (plan.go) hands the provider node one `Grant` per
|
||||
consumer carrying `Sealed: ForProvider`. The provider receives, at the path its `receives`
|
||||
names, one `Contribution` per consumer: the login to create (`As`, derived by the mesh so both
|
||||
ends agree — 04-ISSUES/023), the consumer's address (`At`) and requested `Values`, and
|
||||
`Secret`, the file holding that consumer's password sealed to this provider and unsealed by
|
||||
its host. Everything the provisioner needs is delivered. It reads the wrong files
|
||||
(`*.grant.json`, which nothing writes) and invents a password instead of reading the one in
|
||||
`Secret`.
|
||||
|
||||
## The fix this points to
|
||||
|
||||
A **one-place contract change in the sdk harness**, plus re-pointing today's adapters at it —
|
||||
not per-provider surgery, and inherited correctly by every provider after them:
|
||||
|
||||
- `runProvisioner` reconciles the mesh-delivered `receives` contributions (not `*.grant.json`):
|
||||
for each consumer, create the resource under the login `As` with the password read from the
|
||||
delivered `Secret` file, for its `Values`; withdraw the login when a consumer leaves the file.
|
||||
- The adapter stops generating a password and stops returning a credential — it is handed the
|
||||
name and the password and only makes the resource exist. Roughly `create({as, password,
|
||||
values})` / `remove({as})`, no return.
|
||||
- `sealKey`, `seal()`, `writeSealedCredential`, `MESH_SEAL_KEY`, and the `.credential` file
|
||||
leave entirely; the consumer already receives its copy through the mesh's own channel.
|
||||
|
||||
This is proposed as ADR 0048, which defines the corrected provider contract, for ratification.
|
||||
|
||||
## Resolution
|
||||
|
||||
ADR 0048 was accepted and implemented on the branches this issue is fixed by:
|
||||
|
||||
- `mesh-sdk` `src/provisioner/index.ts` now reconciles the mesh's `receives` contributions and,
|
||||
per consumer, reads the mesh-minted password from the file the host unsealed, calling the
|
||||
adapter to create the resource under the mesh's login. `$MESH_SEAL_KEY`, the symmetric seal,
|
||||
`writeSealedCredential`, and the `*.grant.json` / `*.credential` files are gone. The symmetric
|
||||
`seal()`/`unseal()` primitive had no other caller and was removed.
|
||||
- The four adapters (redis, postgres, minio, umami) were re-pointed at the new contract —
|
||||
`create({ as, password, values })` / `remove({ as })`, returning nothing. minio's client gained
|
||||
a secret-key argument so it sets the mesh's secret rather than generating one.
|
||||
- Proven in the mesh-lab: `provider-uses-mesh-credential` is green — redis creates the consumer's
|
||||
login with the password the mesh minted, a client authenticates as that consumer and gets PONG,
|
||||
with no seal key set anywhere.
|
||||
|
||||
Two things were carved out deliberately, neither blocking:
|
||||
|
||||
- **Data provisions are a separate shape.** umami's `analytics` returns a `siteId` umami
|
||||
*generates*, not a secret the mesh mints, and a contract that returns nothing cannot hand that
|
||||
back. ADR 0048 is scoped to credential provisions and says so; the provider→consumer return
|
||||
path for generated data is left to a separate decision. umami compiles and reconciles under the
|
||||
new harness; only that return is unaddressed, and it never had the seal-key fault.
|
||||
- **Teardown beyond "remove the login"** — an object store's leftover data — is each adapter's to
|
||||
name (minio leaves a non-empty bucket for an operator rather than deleting a consumer's data),
|
||||
not the harness's.
|
||||
@@ -0,0 +1,67 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-04
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# Changing a module's settings does not restart its runtime — config is stale until recreated
|
||||
|
||||
## What was observed
|
||||
|
||||
Rolling the module runtime out to the catalogue (the runtime that serves a module's tools and
|
||||
runs its events under the module's own account), each tools+events module receives its
|
||||
configuration the way the design intends: a mergeable config file the module declares, into
|
||||
which the assignment's settings are merged. The runtime container mounts that file and reads
|
||||
it once at start-up, when it builds its API client.
|
||||
|
||||
The design for settings says a config file a module owns can be changed **without editing
|
||||
it** — a person states an intention, the file is regenerated, and the change takes effect.
|
||||
The decision that config is the assignment's, not the manifest's, is explicitly so that
|
||||
configuration can be updated *on the fly* and managed from a dashboard.
|
||||
|
||||
For a runtime delivered as a **container**, that last part does not hold. When settings
|
||||
change, the control plane re-renders the config file on the node — but the runtime container
|
||||
is only ever recreated when its **spec** changes, and the spec is image, name, env, ports,
|
||||
volumes and args. The *content* of a mounted file is not part of it. So the file on disk
|
||||
updates and the process that already read it keeps the value it read at start-up. The new
|
||||
configuration does not take effect until something changes the container's spec, or it is
|
||||
recreated by hand.
|
||||
|
||||
A **service** resource has `restart-on`, which names the resources whose change forces a
|
||||
restart — exactly this problem, already solved, for units. A **container** resource has no
|
||||
equivalent field, and the apply path for containers never consults the set of resources that
|
||||
changed this pass. So the one kind of resource that hosts a module's runtime is the kind that
|
||||
cannot say "restart me when my config changes."
|
||||
|
||||
The effect is quiet, which is the worst part: setting a value appears to succeed (the file is
|
||||
correct on disk), and the running tools keep answering with the old configuration, or keep
|
||||
failing to load because the value that would fix them is present but unread.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
Every tools+events module converted to the runtime model now takes its URL and credentials
|
||||
this way, so this is not one module's quirk — it is the config path for the whole catalogue.
|
||||
The gap turns the headline promise of the settings design ("change it without editing it, on
|
||||
the fly") into "change it, then recreate the container by hand," which is the manual step the
|
||||
design existed to remove. And because the file is genuinely updated, nothing surfaces the
|
||||
staleness; a dashboard that set the value would report success while the mesh kept doing the
|
||||
old thing.
|
||||
|
||||
Config set **before** the runtime first starts (settings, then assign, then push) does work —
|
||||
the file is right when the process reads it. So the gap is specifically about *updates* to an
|
||||
already-running runtime, which is precisely the case the "on the fly" promise is about.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should a `container` gain `restart-on`, mirroring the service field, so a module can point
|
||||
it at its config resource?
|
||||
- Or should the apply path recreate a container when a file it mounts changed this pass —
|
||||
making mounted-file content behave like part of the spec, without a new field to declare?
|
||||
- Should the config file's content (or a hash of it) fold into the container spec, so an
|
||||
ordinary spec-diff already catches it? That restarts on every change with no new mechanism,
|
||||
at the cost of a spec that is no longer only the container's own declaration.
|
||||
- Is a restart even the right primitive for a runtime that could instead watch its config
|
||||
file and rebuild its clients in place — and if so, is that each module's job or the
|
||||
runtime host's?
|
||||
@@ -0,0 +1,77 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-05
|
||||
located-in: [mesh-control, mesh-catalog]
|
||||
fixed-by: ADR 0049 (a slug for the login) + a shorter minted secret (mesh-control)
|
||||
amended-design: 0049-a-consumers-identity-fits-the-tightest-backend.md
|
||||
---
|
||||
|
||||
# The mesh's derived login does not fit every backend's identity rules — S3 rejects it
|
||||
|
||||
## What was observed
|
||||
|
||||
Proving the provider/consumer contract per backend (ADR 0048), redis and postgres passed: a
|
||||
consumer authenticated against the provider with the login the mesh derived and the password
|
||||
the mesh minted. **minio failed**, and not on the credential — on the *name*:
|
||||
|
||||
```
|
||||
mc: <ERROR> Unable to add a new service account. The access key is invalid.
|
||||
(access key length should be between 3 and 20).
|
||||
```
|
||||
|
||||
The mesh derives a consumer's login as `mesh_<node>_<module>` — here `mesh_anchor_bucketuser`,
|
||||
22 characters. That is a valid postgres role and a valid redis ACL user, so those providers
|
||||
create it verbatim. S3 access keys are capped at **20 characters**, so minio refuses to create
|
||||
the service account under it, and the provisioner retries forever while the consumer, holding
|
||||
that same too-long access key, could never present it either.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
ADR 0048 says a provider creates *exactly* the login the mesh derived, so that the two ends
|
||||
agree by construction — the mesh hands the same name to the provider (to create) and the
|
||||
consumer (to present). That only holds if the derived name is one every provider can accept.
|
||||
It is not: the mesh's `as` is a single format with no knowledge of a backend's identity rules,
|
||||
and S3's are stricter than a database's. Any provider whose backend constrains identifiers more
|
||||
tightly than postgres — a length cap, a charset, a required prefix — inherits this, and the
|
||||
failure lands at provision time, per consumer, as an infinite retry rather than a refusal at
|
||||
assignment.
|
||||
|
||||
This also shows the seam is real, not cosmetic: `as` is doing two jobs — a stable per-consumer
|
||||
identity the two ends must agree on, and a literal identifier a specific backend must accept —
|
||||
and those are not always the same string.
|
||||
|
||||
## The shape of a fix (open, not decided)
|
||||
|
||||
- **Constrain the derivation** so `as` is broadly acceptable — short (≤ 20), a conservative
|
||||
charset, deterministic. This keeps "the provider creates exactly what the mesh derived" true
|
||||
everywhere, at the cost of a less legible name, and it is a mesh-wide identity change (every
|
||||
provider that already created the longer name would see it change).
|
||||
- **Let a provider map `as` to a backend-valid identifier** it derives the same way on create
|
||||
and on the consumer's behalf — but the consumer is generic and cannot run minio's mapping, so
|
||||
this only works if the mapped identifier is *delivered back* to the consumer. That is the
|
||||
data-provision return path this era keeps meeting (umami's siteId, cloudflare's record) and
|
||||
does not yet have.
|
||||
- **Declare the constraint on the interface** (`s3-bucket` states its identifier bounds) and
|
||||
have the mesh derive within them — the most honest, the most work.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Is `as` meant to be human-legible, or is a short opaque token acceptable — i.e., can the
|
||||
derivation simply be shortened without anyone minding?
|
||||
- Do redis/postgres actually want the long name, or did it only survive because they are
|
||||
permissive? If nothing needs it long, the cheap fix is to cap it.
|
||||
- Does this fold into the same decision as the data-provision return path, or is it separate?
|
||||
|
||||
## Resolution
|
||||
|
||||
Accepted **ADR 0049** (option E): a module declares an optional short `slug`, and the mesh derives
|
||||
`mesh_<node>_<slug|name>`, bounded by the tightest backend (an S3 access key's 20) and refused at
|
||||
assignment — naming the slug as the remedy — when it still would not fit. The minio grant e2e proved
|
||||
it: `bucketuser` declares `slug: bkt`, so its access key `mesh_anchor_bkt` (15) is accepted where
|
||||
`mesh_anchor_bucketuser` (22) was refused.
|
||||
|
||||
Proving that surfaced a **second S3 length constraint on the same credential** — the secret. The
|
||||
mesh minted a 43-character password (32 random bytes, base64url), and an S3 secret key is 8–40. Fixed
|
||||
in `mesh-control` `internal/secrets/seal.go` by minting 30 bytes → exactly 40 characters (240 bits,
|
||||
ample), which fits S3 and every other backend. Both halves of an S3 credential — the access key
|
||||
(login) and the secret key (password) — now fit the tightest backend, by the same rule.
|
||||
@@ -0,0 +1,48 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-02
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 035 — Reconciling a seed file wipes what grew in it
|
||||
|
||||
## The symptom, as observed
|
||||
|
||||
Found by review of the catalogue examples (2026-09-02), not by an outage — the outage is the
|
||||
part the design permits to be silent.
|
||||
|
||||
The cache module declares its access-control file as an ordinary file resource with fixed,
|
||||
empty content. The program that consumes the file requires it to exist at startup, which is
|
||||
why the manifest declares it at all. But the same file is the one the provisioner writes
|
||||
consumer users into, and the one the running program persists ACL changes back to.
|
||||
|
||||
A declaration is complete for what the host owns, and the host reconciles what is declared
|
||||
([ADR 0010](../../02-DECISIONS/0010-delivery.md)).
|
||||
So every apply that revisits this resource restores the declared content — empty — behind the
|
||||
running program. Every consumer credential granted since the last apply is removed, the apply
|
||||
reports success, and nothing anywhere says a grant vanished.
|
||||
|
||||
## Why it matters beyond the instance
|
||||
|
||||
The manifest needed *the file to exist before first start*, and the only vocabulary available
|
||||
was *the file has this content, forever*. Those are different intentions, and the gap between
|
||||
them is generic: any resource that a module seeds and something else then legitimately mutates
|
||||
— an ACL file, a bootstrap configuration a program rewrites, an htpasswd a provisioner appends
|
||||
to — has the same two owners and the same silent loss on reconcile.
|
||||
|
||||
It is also the mirror image of the boundary ADR 0010 draws so carefully on the *removal* side:
|
||||
the host never removes what it did not create, but it happily overwrites what it *did* create,
|
||||
even when what grew inside since is somebody else's work the mesh asked for.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Is the missing thing a create-once file semantic ("present with this content if absent,
|
||||
untouched otherwise"), or is the real fault that two owners share one file — and the
|
||||
provisioner, not the declaration, should own it entirely, with first-start ordering solved
|
||||
some other way?
|
||||
- ADR 0010 treats every added resource type as a security artefact. Does a create-once
|
||||
semantic widen what a compromised control plane can express, or narrow it?
|
||||
- Are there other seeded-then-mutated files already in the catalogue that this failure is
|
||||
waiting inside?
|
||||
@@ -0,0 +1,51 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-09-02
|
||||
located-in: [mesh-control, mesh-catalog, mesh-host]
|
||||
fixed-by: 02-DECISIONS/0051-shared-data-is-the-operators.md
|
||||
amended-design: 02-DECISIONS/0051-shared-data-is-the-operators.md
|
||||
---
|
||||
|
||||
# 036 — Six modules own what they must share
|
||||
|
||||
## The symptom, as observed
|
||||
|
||||
Found by review of the catalogue examples (2026-09-02). The media stack is several modules —
|
||||
a library server, the acquisition managers, a download client and their satellites — and each
|
||||
of them declares the same library and download directories as its own resources.
|
||||
|
||||
The resolver refuses two modules that declare one path on one node, with no exemption for
|
||||
identical content and no merge. That rule is right in general: two owners of one path is the
|
||||
class of fault this repository keeps recording. But sharing those directories on one machine
|
||||
is the entire point of this stack — the download client and the managers must see the same
|
||||
downloads, the library server must see the same libraries. So the set, as written, refuses
|
||||
its own only sensible assignment.
|
||||
|
||||
No test co-resolves any two of them, which is why the manifests pass today. The first machine
|
||||
to be assigned the stack together is where the refusal would have surfaced.
|
||||
|
||||
## Why it matters beyond the instance
|
||||
|
||||
The manifests can express *a directory I own* and nothing else, so a directory that is the
|
||||
shared workspace of several modules was written six times as six private ones. The intention
|
||||
— several modules, one filesystem contract between them — has no vocabulary, and this is not
|
||||
a media-stack peculiarity: any pipeline of modules handing files to each other on one machine
|
||||
(an ingest directory, a spool, a drop folder) hits the same wall.
|
||||
|
||||
It is also a fork in the design the catalogue has otherwise avoided: the fix could be a new
|
||||
owning module the others depend on, a shared-resource concept in the manifest, or a statement
|
||||
that co-located file handoff is not a thing the mesh supports and these modules are one
|
||||
module. Each answer changes what a module *is*, which is why this is an issue and not a patch.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Is the unit wrong — is a stack that must share a filesystem one module with several
|
||||
containers, the way the mail module already is?
|
||||
- If it stays several modules: does one of them own the directories and the rest require
|
||||
them, and is *requiring a directory from a neighbour* a provision, a claim, or a third
|
||||
thing?
|
||||
- The duplicate-path rule protects against genuinely rivalrous owners. Whatever expresses
|
||||
sharing must not weaken it for the cases where refusal is the right answer — what
|
||||
distinguishes the two, machine-checkably?
|
||||
- The mesh's own rule is that a rule states how it is checked: whichever shape is chosen,
|
||||
what test co-resolves the stack so this class of refusal is caught before a machine is?
|
||||
@@ -0,0 +1,57 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-09-05
|
||||
located-in: [mesh-control, mesh-host, mesh-catalog]
|
||||
fixed-by: 02-DECISIONS/0052-a-step-that-runs-once-before-a-container.md
|
||||
amended-design: 02-DECISIONS/0052-a-step-that-runs-once-before-a-container.md
|
||||
---
|
||||
|
||||
# 037 — A module cannot run its own code at a lifecycle phase
|
||||
|
||||
## The symptom, as observed
|
||||
|
||||
Found while converting the catalogue (2026-09-05), across several modules at once. A module can
|
||||
declare *things that exist* — a directory, a file with fixed content, a network, a container — but
|
||||
it cannot declare *a step that runs* at a defined point in its own lifecycle. Three converted
|
||||
modules need exactly that and have nowhere to put it:
|
||||
|
||||
- **mosquitto.** Its Dynamic Security plugin will not start unless `dynamic-security.json` already
|
||||
contains an admin client *before the broker's first start* — the broker loads the plugin at
|
||||
boot. Seeding it is a run-once step that must happen after the file resource exists and before
|
||||
the container starts. The vocabulary has no "before first start."
|
||||
- **The database providers (postgres/mongodb/mssql).** First-boot seeding works today only because
|
||||
the *image* happens to do it from an env var. Anything the mesh itself must run once against the
|
||||
server — a schema migration, an extension enable, a health gate before the module is announced
|
||||
ready — has no home.
|
||||
- The seed-then-mutate family already recorded in [035](../035-reconciling-a-seed-file-wipes-what-grew-in-it/00-report.md)
|
||||
is the same shape seen from the *content* side; this is it seen from the *timing* side.
|
||||
|
||||
## Why it matters beyond the instance
|
||||
|
||||
This is not a defect in a module — it is a **capability the module system does not yet offer.** A
|
||||
real class of modules needs to run their own code at points in the build/install/run lifecycle:
|
||||
seed-before-start, migrate, post-start health-gate, pre-remove drain. The declarative resource
|
||||
model deliberately describes *state*, not *steps*, and that is right for what it covers; the gap is
|
||||
that some modules genuinely have a step.
|
||||
|
||||
**Prior art, and its warning.** An earlier mesh had exactly this as a feature: event-driven
|
||||
**hooks** that ran custom code at phases of the build/publish/deploy pipeline. It was powerful and
|
||||
it was **complex to set up and flaky** — which is the real content of this record. The need is not
|
||||
in question; the cost of the obvious answer is. Whatever shape this takes must not reproduce that
|
||||
fragility, or it will be worse than the gap.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Is the right unit narrow — a **run-once / init resource** ("run this once, here, in the
|
||||
lifecycle") — or general — a **per-phase lifecycle hook** on a module, and if so which phases
|
||||
(build / publish / install / pre-start / post-start / pre-remove)?
|
||||
- Where does a hook's code run — in the module's own runtime container under its scoped account
|
||||
(ADR 0043/0047), so it inherits the same isolation as its tools and events? Or is some of it the
|
||||
host's, before a container exists?
|
||||
- How is a step made **idempotent and reconcilable** so a re-apply does not re-run it
|
||||
destructively — the same discipline the resource model gets for free and a step does not?
|
||||
- What is the smallest version that unblocks the three modules above without rebuilding the old
|
||||
flaky hook engine? Is "seed-before-first-start" alone enough for now, with the general case
|
||||
deferred?
|
||||
- A rule states how it is checked: whatever shape is chosen, what lab scenario proves a hook runs
|
||||
exactly once, at the right phase, and converges on re-apply?
|
||||
@@ -0,0 +1,91 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-09
|
||||
located-in: [mesh-control]
|
||||
fixed-by: mesh-control — a same-node provider is announced at the port it is published on
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 038 — A provider is announced at a name its port is not bound to
|
||||
|
||||
## Symptom
|
||||
|
||||
A module that provides a `from: mesh` provision (observed with the database provider) is
|
||||
announced to its consumers, by [issue 018](../018-a-provider-on-the-same-machine-was-never-announced/00-report.md)'s
|
||||
fix, at the node's private-network name — the binding a consumer reads carries
|
||||
`at: <node>.internal` and `serves.port: 5432`.
|
||||
|
||||
But the provider's container port is **published bound to loopback** (`127.0.0.1:<assigned>`),
|
||||
not to the address `<node>.internal` resolves to. So every consumer dials the announced
|
||||
`<node>.internal:5432`, which resolves to the node's private-network address, where **nothing is
|
||||
listening** — the port is open only on `127.0.0.1`.
|
||||
|
||||
Observed on a node hosting the provider and several consumers:
|
||||
|
||||
- The consumer's binding file says `"at": "<node>.internal"`, `"serves": { "port": 5432 }`.
|
||||
- Inside a consumer container, that name resolves to the node's private-network address.
|
||||
- A connection test from the node: the private-network address on port 5432 is **CLOSED**; only
|
||||
`127.0.0.1` (on the assigned host port) is OPEN.
|
||||
- Consumers that touch the database only lazily serve a landing page and *look* healthy; consumers
|
||||
that require the database at startup crash-loop — one with "acquisition timeout while waiting for
|
||||
a new connection", another connecting and then timing out on its first query.
|
||||
- The provider itself is healthy: a direct client on loopback answers instantly, few connections,
|
||||
no locks.
|
||||
|
||||
## Why this matters
|
||||
|
||||
The announced address and the actual listener disagree, so the binding is a promise the mesh does
|
||||
not keep. It is not one module's misconfiguration: it is the port-publishing step choosing a bind
|
||||
address that does not match the `at` the resolver hands consumers, so it fails the same way for
|
||||
**every** `from: mesh` provider with an off-node-reachable consumer — and, on a single node, for
|
||||
same-node consumers too.
|
||||
|
||||
It hides well. The provider is up, the credential is correct, the database exists, a manual client
|
||||
works — every part a person checks in isolation passes. Only a consumer that must use the provision
|
||||
before it can serve anything reveals it, and it reveals it as a timeout, which reads as slowness or
|
||||
load rather than "the address was never listening". A mesh that co-locates a provider with its
|
||||
consumers (the ordinary small-mesh case) is exactly where it bites.
|
||||
|
||||
It also blocks anything that must *reach* a routed/served name from inside the mesh, not just
|
||||
application traffic — see the internal-CA validation dependency noted in the connectivity design.
|
||||
|
||||
## Diagnosis
|
||||
|
||||
The symptom's first reading — "published on loopback" — was **partly a red herring**. Two things
|
||||
were tangled:
|
||||
|
||||
1. **The real, current-code defect is a served-*port* mismatch, not a bind address.** A bare
|
||||
`ports: ["5432"]` is assigned a host port and published as `"15432:5432"` — no bind IP, so on
|
||||
**all interfaces**, reachable at the node's private-network address. But the port a consumer is
|
||||
*told* is only re-derived from the assignment on the **cross-node** path. The **same-node** paths
|
||||
(the resolver's `servedHere`, and the `here()` fallback) settle their served facts *while
|
||||
resolving* — before the host port is assigned — so they carry the **declared** port (5432), not
|
||||
the **assigned** one (15432). A co-located consumer is therefore announced
|
||||
`<node>.internal:5432` while the provider is published on `<node>.internal:15432`, and dials a
|
||||
port nothing listens on. Cross-node consumers were always fine, which is why it read as "the
|
||||
small-mesh case."
|
||||
|
||||
2. **The `127.0.0.1:15432` seen in the running lab was a stale build.** Current code's publish step
|
||||
binds all interfaces; the running instance was raised from a mesh-control predating the ADR 0038
|
||||
publish rewrite. The substrate's own store *is* deliberately `127.0.0.1:5432` (a private store
|
||||
must not be exposed) — correct, and not this bug.
|
||||
|
||||
## Fixed by
|
||||
|
||||
`mesh-control` branch `fix/same-node-provider-announced-port` (`c147a26`): after the host port is
|
||||
assigned, same-node needs (and the `here()` fallback) are redirected through the same
|
||||
provision→module→assigned-port lookup the cross-node path already uses, so a co-located consumer is
|
||||
announced the port that is actually published. Idempotent (keyed by the declared port). Regression
|
||||
test `TestASameNodeProviderIsAnnouncedAtThePortItIsPublishedOn` asserts the announced port equals
|
||||
the published host port for a co-located provider/consumer — the next assertion after 018's (which
|
||||
only checked a binding file exists); verified failing without the change.
|
||||
|
||||
*Not yet merged, and the running lab is additionally stale — proving it end-to-end there needs
|
||||
mesh-control rebuilt and the affected consumer containers recreated.*
|
||||
|
||||
## Noted, not taken
|
||||
|
||||
Binding the assigned port to the node's private-network address specifically (rather than all
|
||||
interfaces) would be defence-in-depth and would make the publish address match `at` by construction
|
||||
— but it is a larger behavioural change entangled with the unenforced firewall scope
|
||||
([003](../003-firewall-scope-is-read-by-no-code/00-report.md)), so it is left as an option.
|
||||
Reference in New Issue
Block a user