Compare commits
26
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
83791f0921 | ||
|
|
bf4d4e2e7b | ||
|
|
ec42ee0846 | ||
|
|
5a3dee9e9e | ||
|
|
47909c6b71 | ||
|
|
7e5edab8da | ||
|
|
3649f82204 | ||
|
|
b400ce2c54 | ||
|
|
995c8cb266 | ||
|
|
f89aef992d | ||
|
|
b967ef7be3 | ||
|
|
e417906241 | ||
|
|
b1bf895688 | ||
|
|
0d64677c70 | ||
|
|
04205c1dc8 | ||
|
|
9dc49cd831 | ||
|
|
66b413a076 | ||
|
|
fcba05fed9 | ||
|
|
18f37c25b2 | ||
|
|
3c2b4fc6b6 | ||
|
|
72eaf52867 | ||
|
|
741625e725 | ||
|
|
c3730b9a23 | ||
|
|
3a59099c81 | ||
|
|
3bd6f34de3 | ||
|
|
ec8676c225 |
@@ -54,4 +54,6 @@ whose failure has never been observed is a guess about its own correctness.
|
||||
The development cycle, checked ([ADR 0080](../../02-DECISIONS/0080-the-development-cycle-is-checked.md)):
|
||||
a to-be design names a decision, an in-progress/implemented design names its owning code, a
|
||||
located/fixed issue names its owner, a fixed/resolved issue says what fixed it, a graduated
|
||||
research overview says what it became. `python3 00-META/checks/cycle.py`
|
||||
research overview says what it became, and no two issue records share a number (issue 155 — the
|
||||
number is how a record is cited, and `main` lags every open pull request, so two people reading it
|
||||
allocate the same one). `python3 00-META/checks/cycle.py`
|
||||
|
||||
+23
-1
@@ -14,7 +14,8 @@ What is enforced:
|
||||
its owning code (`code:`) -- no development without a design that says where.
|
||||
issues a known `status:`; once `located`, `located-in:` names the owner;
|
||||
once `resolved`, `fixed-by:` says what fixed it (prose counts --
|
||||
"nothing, the capability existed" is an answer).
|
||||
"nothing, the capability existed" is an answer). And no two records share a
|
||||
number -- the number is how a record is cited.
|
||||
research a known `status:`; a `graduated` overview says what it `became:`, and every
|
||||
target it names exists.
|
||||
decisions every accepted record is REACHABLE from the cycle: cited by a design doc's
|
||||
@@ -109,6 +110,27 @@ def main():
|
||||
"without a design that says where" % status)
|
||||
|
||||
# ---- issues ------------------------------------------------------------------------
|
||||
# Two records may not share a number. Numbers are taken as "next free after main", and work
|
||||
# sits on unmerged branches for days -- so two people reading the same main allocate the same
|
||||
# number, and nothing said so. It happened twice in one evening between two machines, and the
|
||||
# second collision landed on main with all three checks passing (issue 155). An issue number is
|
||||
# how every other record cites this one; two records answering to it means a pointer that
|
||||
# resolves to whichever the reader happened to open.
|
||||
seen = {}
|
||||
for folder in sorted(glob.glob(os.path.join(ROOT, "04-ISSUES", "*", ""))):
|
||||
name = os.path.basename(os.path.normpath(folder))
|
||||
number = name.split("-", 1)[0]
|
||||
if not number.isdigit():
|
||||
continue
|
||||
if number in seen:
|
||||
bad(os.path.join("04-ISSUES", name),
|
||||
"is numbered %s, and so is %s -- an issue number is how it is cited, and two "
|
||||
"records answering to one means a citation that resolves to whichever the reader "
|
||||
"opened. Take the next free number across main AND every open pull request"
|
||||
% (number, seen[number]))
|
||||
else:
|
||||
seen[number] = name
|
||||
|
||||
for path in sorted(glob.glob(os.path.join(ROOT, "04-ISSUES", "*", "00-report.md"))):
|
||||
front = frontmatter(path)
|
||||
if front is None:
|
||||
|
||||
@@ -21,7 +21,12 @@ incident someone must **clear**.
|
||||
|
||||
## Steps
|
||||
|
||||
1. Take the next free number. Create `04-ISSUES/NNN-short-name/00-report.md`:
|
||||
1. Take the next free number — **across `main` and every open pull request**, not `main` alone.
|
||||
Work sits on unmerged branches for days, so two people both reading `main` allocate the same
|
||||
number; it happened twice in one hour between two machines, and the second collision reached
|
||||
`main` with every check passing (issue 155). `cycle.py` now refuses two records sharing a number,
|
||||
which catches a collision but does not prevent one. Create
|
||||
`04-ISSUES/NNN-short-name/00-report.md`:
|
||||
|
||||
```yaml
|
||||
---
|
||||
@@ -42,6 +47,14 @@ incident someone must **clear**.
|
||||
## Rules
|
||||
|
||||
- Closed issues are never deleted — they are the mesh's symptom-to-component memory.
|
||||
- `fixed-by:` names something that will still exist: a commit or a pull request, never a branch. A
|
||||
branch is deleted when it merges, so a branch name there is a pointer that resolves to nothing by
|
||||
the time anybody follows it.
|
||||
- A fix that turns out to have broken something else is written back into the record that asked for
|
||||
it, pointing at the new issue. Somebody arriving at a record to learn why the code is the way it
|
||||
is must not have to already know there was a sequel.
|
||||
- Renumbering a collision happens once, in the branch that lands last. Renumbering a branch whose
|
||||
author is still pushing only moves the race.
|
||||
- An issue whose answer is a general lesson should also be written to the knowledge base, so
|
||||
the next person searching a symptom finds it. Both, not either.
|
||||
- `status: wontfix` is legitimate and requires a sentence saying why.
|
||||
|
||||
@@ -26,6 +26,14 @@ runtime is per-module, not per-node, and treating the audit-logger as special le
|
||||
modules' code with nothing to run it: the conversion produced tools and events that, as it stands,
|
||||
never execute.
|
||||
|
||||
> **The hosting form is settled elsewhere — 2026-09-30.** Where this record says "a container", read
|
||||
> [ADR 0150](0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md): a module's own
|
||||
> code runs as supervised processes under this record's one account. Nothing else here changes — the
|
||||
> per-module runtime, the per-tool key and the single scoped account are the argument this record made
|
||||
> and they are why 0150 goes the way it does. The note is here because two design documents chose the
|
||||
> other form without knowing this record existed
|
||||
> ([issue 117](../04-ISSUES/117-a-modules-own-code-is-a-container-and-a-process/00-report.md)).
|
||||
|
||||
## Decision
|
||||
|
||||
### A module with tools or events runs a process of its own
|
||||
|
||||
@@ -1,14 +1,21 @@
|
||||
---
|
||||
topic: building it
|
||||
status: proposed
|
||||
status: superseded
|
||||
date: 2026-09-12
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 0016-the-lab.md
|
||||
superseded-by: 02-DECISIONS/0149-the-live-mesh-is-the-test-bed.md
|
||||
---
|
||||
|
||||
# 68. The lab takes requests, one at a time, and runs each from its own copy
|
||||
|
||||
> **Superseded — 2026-09-30, by [ADR 0149](0149-the-live-mesh-is-the-test-bed.md).** Never built. The
|
||||
> live mesh became the test bed, because the faults that cost the most are faults of a mesh that
|
||||
> already exists — bound consumers, containers made against an older roster, an adopted machine — and a
|
||||
> bed is by construction a mesh that does not. The rule worth keeping from below is that a run reads a
|
||||
> copy that is not anybody's working tree.
|
||||
|
||||
## Context
|
||||
|
||||
**The lab is exclusive hardware, and today a person holds it.** Raising a scenario takes over
|
||||
|
||||
@@ -35,6 +35,16 @@ The development cycle is enforced mechanically, to the extent frontmatter can ca
|
||||
capability existed" is an answer).
|
||||
- **No silent graduation** — a `graduated` research overview says what it `became:`, and the
|
||||
targets exist.
|
||||
- **No two records answering to one number** — added 2026-09-30; see the insight below.
|
||||
|
||||
> **Progressive insight — 2026-09-30.** The list above named four things `cycle.py` enforces, and
|
||||
> now names five. Nothing enforced that two issue records hold different numbers: two machines
|
||||
> filing issues within one hour both read `main`, both took "the next free number", and collided
|
||||
> twice — the second collision reaching `main` with `records.py`, `cycle.py` and `index.py` all
|
||||
> reporting success (04-ISSUES/155). A number is how every other record cites one, so two records
|
||||
> answering to it is a citation that resolves to whichever folder the reader opened. `cycle.py`
|
||||
> refuses it now. The decision here stands exactly as written: this is one more thing frontmatter
|
||||
> and file names can carry, found by its absence rather than by reasoning.
|
||||
|
||||
[`00-META/checks/cycle.py`](../00-META/checks/cycle.py) refuses violations, beside `records.py`
|
||||
and `index.py`; all three run before any merge here. What frontmatter cannot see — that code work
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
topic: what runs on it
|
||||
status: proposed
|
||||
status: accepted
|
||||
date: 2026-09-26
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
@@ -189,6 +189,23 @@ On acceptance, each of these is amended by this record, not edited:
|
||||
credential is no longer all-or-nothing with a window. It overlaps, with each step confirmed.
|
||||
A single-party credential is staged, not replaced.
|
||||
|
||||
## Accepted, 2026-09-30, and not scheduled
|
||||
|
||||
Accepted as written. The separation it draws — a consumer's *resource* and a *credential that reaches
|
||||
it* are different things with different lifecycles — is the part that had to be settled, because the
|
||||
alternative is what the record was written against: retiring a credential taking the data it reached
|
||||
with it. That is a data-loss shape, and a record that names it should not sit unresolved while the
|
||||
code that could hit it is being written.
|
||||
|
||||
**It is not built, and accepting it does not schedule it.** The SDK's provisioner adapter is still
|
||||
`create` / `remove` / `holds` rather than the four operations above, and no provider implements the
|
||||
two-credential rotation. Accepted-and-not-built is an ordinary state here — 0141 and 0142 are both in
|
||||
it — and it is the honest one: leaving this `proposed` made it invisible to anyone reading what the
|
||||
mesh has decided, while changing nothing about what runs.
|
||||
|
||||
The work it implies belongs with the provisioner contract, beside
|
||||
[issue 124](../04-ISSUES/124-a-consumer-cannot-be-told-what-its-provider-derived/00-report.md).
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Every credential provider's adapter changes**, in two steps. The first separates *retire a
|
||||
|
||||
@@ -0,0 +1,159 @@
|
||||
---
|
||||
topic: the tiers
|
||||
status: accepted
|
||||
date: 2026-09-30
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 02-DECISIONS/0066-public-routing-is-name-agnostic.md
|
||||
---
|
||||
|
||||
# 148. The mesh's names are resolved, not copied into every container
|
||||
|
||||
## Context
|
||||
|
||||
The mesh gives every container it declares the whole roster of mesh names as entries written into
|
||||
the container's own hosts file at creation
|
||||
([design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md),
|
||||
[ADR 0066](0066-public-routing-is-name-agnostic.md)). A container takes those entries once and never
|
||||
looks again.
|
||||
|
||||
Three issues are the same fact arriving three times.
|
||||
|
||||
**A container keeps the address it was made with.** Adopting the predecessor's tunnel moved the hub's
|
||||
private address; the declaration followed it within one push and nothing on the machine did. The
|
||||
forge's container held the old address, lost its database, reported healthy while its existing
|
||||
connections lasted, and then the public name went down
|
||||
([issue 109](../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md)).
|
||||
|
||||
**The same fault, four days later, undetected for five days.** One container had restarted 2286 times
|
||||
against a database it could no longer find, while the mesh reported the machine as doing what it was
|
||||
told. Inside it, `novox.internal` was an address that had not existed for five days
|
||||
([issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md)). Forty-eight
|
||||
other containers were current, none of them corrected — each had been recreated for some other
|
||||
reason and picked up the roster on the way.
|
||||
|
||||
135 was fixed by putting the roster into the digest the host compares a container against, so a
|
||||
container whose names moved is recreated like one whose image moved. **That made the roster part of
|
||||
every container's identity**, which is the third arrival:
|
||||
|
||||
**One name moving replaces every container in the mesh.** Migrating one small module on one machine
|
||||
took four routine actions; each changed the roster, and each replaced every container on the control
|
||||
node — its own store, the registry, the edge proxy, the forge, the directory, mail. The control plane
|
||||
was unreachable twice while its own store came back through crash recovery
|
||||
([issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). None of
|
||||
the replaced containers had anything to do with the module being migrated, or with its machine.
|
||||
|
||||
The blast radius of a name is now every container that carries the list, which is all of them. The
|
||||
node-by-node migration ahead adds names one module at a time — on one machine alone that is around
|
||||
twenty-five — and each would be a full restart of every service on the hub.
|
||||
|
||||
## Considered Options
|
||||
|
||||
**1. Keep the roster in every container and accept the churn.** Rejected. It is not a cost that can
|
||||
be paid down: the mesh gets more names as it grows, and every name costs a restart of everything.
|
||||
A rollback costs another.
|
||||
|
||||
**2. Scope each container's entries to the names it actually binds.** A container is given the names
|
||||
of the things it declared a requirement on, so a name's blast radius is its consumers. Tidy, needs no
|
||||
new mechanism, and keeps 135's guarantee exactly.
|
||||
|
||||
Rejected, and this is the close one. It contradicts the standing intent that **anything on the mesh
|
||||
can call anything on it** — three cases, same machine, the private network, the public network, and
|
||||
no fourth. Scoping resolution to declared couplings makes a name reachable only where the mesh was
|
||||
told in advance that it would be wanted, and a person debugging inside a container would find names
|
||||
missing that exist everywhere else on the machine. It also leaves the roster in the digest, so the
|
||||
churn returns the moment a widely-bound name moves — smaller, not gone.
|
||||
|
||||
**3. Resolve at lookup time through the machine's resolver, and copy nothing.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**A container resolves the mesh's names through its machine's resolver, at the moment it asks. No
|
||||
mesh name and no mesh address is written into a container, and none is part of a container's
|
||||
identity.**
|
||||
|
||||
The three consequences that make this worth doing:
|
||||
|
||||
- **Staleness stops being possible**, rather than being detected. 109 and 135 are not bugs that were
|
||||
fixed; they are a shape that no longer exists. A name that moves is answered differently by the next
|
||||
lookup, in every container, with nothing recreated and nothing restarted.
|
||||
- **A name's blast radius becomes nothing.** Assigning a module on one machine does not touch a
|
||||
container on another.
|
||||
- **Anything can still call anything**, which option 2 gave up. The resolver answers every mesh name to
|
||||
every asker on the machine, exactly as it answers the machine itself.
|
||||
|
||||
**The resolver is a machine-level process, not a container** — one of the modules that is not a
|
||||
container at all — so a container depending on it is not the circularity it would be if the mesh's
|
||||
own store had to resolve a name through something the store's own runtime had to start first.
|
||||
|
||||
**What a module declares for itself is untouched.** Entries a manifest asks for are the module's own,
|
||||
stay in the container, and stay in its identity: they are part of what the module *is*, they do not
|
||||
move when the mesh's roster does, and the mesh does not know what they mean.
|
||||
|
||||
**The machine's own roster file is untouched.** It is a file, rewritten in place, read by processes and
|
||||
people; nothing restarts when it changes. It is only the *copy into each container* that this ends.
|
||||
|
||||
### The order this lands in, which is not a preference
|
||||
|
||||
**Nothing may stop copying names until resolution works from a container.** Removing the copy first
|
||||
reintroduces 109 and 135 — silently, and on a live mesh, which is exactly how both were found.
|
||||
|
||||
1. **A container on any network can reach the resolver.** Today a container on the runtime's default
|
||||
network asks from an address the converged filter drops, so it has no DNS at all
|
||||
([issue 110](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md));
|
||||
and on two of four machines the resolver binds loopback only, so the runtime hands containers a
|
||||
public resolver instead. Both are prerequisites, not related work.
|
||||
2. **The runtime is told which resolver to use, per machine, as a file** — not per container as a
|
||||
creation-time argument, or the resolver's address is back in every container's identity and the
|
||||
problem has only got smaller.
|
||||
3. **Then, and only then, the roster leaves the declaration and the digest.**
|
||||
|
||||
Until step 3 the mesh keeps copying, and keeps comparing. 151 stays open until step 3 lands; it is not
|
||||
closed by this record, only answered by it.
|
||||
|
||||
## How this is checked
|
||||
|
||||
- **A container resolves a name that moved, without being recreated.** Move a name the mesh serves;
|
||||
from a container that was running before the move and has not been touched since, the name answers
|
||||
with the new address. This is the one 109 and 135 would both have failed.
|
||||
- **A name's blast radius is nothing.** Add a routed name on one machine; no container on any other
|
||||
machine is recreated. The apply report on each machine says nothing changed. This is 151.
|
||||
- **Anything calls anything.** From a container on any machine, every `<node>.internal` name and every
|
||||
routed name the mesh serves resolves — including names the module never declared a requirement on,
|
||||
which is the guarantee option 2 would have given up.
|
||||
- **On every network the runtime offers.** The first three hold for a container on the runtime's
|
||||
default network as well as one on a declared network, because the default network is the case that
|
||||
has no DNS today.
|
||||
- **No mesh name is in a container's spec.** A test asserts the digest a host computes for a container
|
||||
does not move when the mesh's roster does, and does move when the module's own declared entries do.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **The resolver becomes load-bearing for every container**, where before it was load-bearing for the
|
||||
machine. This is a real cost and is accepted: a resolver that is down is a machine that cannot
|
||||
resolve, which is already true of the machine itself, and is a smaller event than a roster change
|
||||
destroying and recreating every container on the machine.
|
||||
- **Design 08's "a file rather than a resolver" no longer describes containers.** It was written when
|
||||
the mesh had no resolver and it gave the right answer then. The reasoning it rested on — every Linux
|
||||
has a hosts file, no package needed — was already overtaken by names a hosts file cannot express:
|
||||
service names and wildcards under `<node>.internal`, which is why the resolver was built.
|
||||
- **ADR 0066's mesh-wide propagation is kept and its mechanism changes.** A routed name still reaches
|
||||
every asker in the mesh; it reaches them through the resolver rather than by being written into each
|
||||
container. The consequence 0066 records — that an internal issuer's challenge needs the routed name
|
||||
resolvable inside the mesh — holds unchanged and by the same means the machine already uses.
|
||||
- **Issue 110 stops being a container-DNS inconvenience and becomes a prerequisite** for the mesh not
|
||||
restarting itself whenever it learns a name.
|
||||
- **A container started by hand gets the mesh's names too**, where before only declared containers did.
|
||||
Design 08 drew that boundary deliberately, on the grounds that reaching into every container is what
|
||||
a nameserver would be for. This record accepts that consequence rather than working around it: a
|
||||
person debugging in a hand-started container resolving the same names as everything else is the
|
||||
behaviour worth having, and it is what "anything can call anything" means.
|
||||
|
||||
## References
|
||||
|
||||
- [issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md) — one name replaces every container; the question this answers
|
||||
- [issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md) — the roster put into the digest
|
||||
- [issue 109](../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md) — the first arrival
|
||||
- [issue 110](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md) — the prerequisite
|
||||
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — routed names propagate mesh-wide; extended here
|
||||
- [design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md) — the file-not-resolver reasoning this narrows
|
||||
@@ -0,0 +1,106 @@
|
||||
---
|
||||
topic: building it
|
||||
status: accepted
|
||||
date: 2026-09-30
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
supersedes: 02-DECISIONS/0068-the-lab-takes-requests.md
|
||||
---
|
||||
|
||||
# 149. The live mesh is the test bed
|
||||
|
||||
## Context
|
||||
|
||||
[ADR 0068](0068-the-lab-takes-requests.md) proposed that the lab accept queued requests
|
||||
— a bed and a commit — answer them one at a time from a copy it owns, and expose that through tools so
|
||||
an agent could start a run and come back to it. It has been `proposed` since 2026-09-12 and nothing was
|
||||
built.
|
||||
|
||||
What happened instead is that the mesh became the thing under test. It runs on four machines; every
|
||||
fault worth finding in the last month was found on them, and none was found in a bed:
|
||||
|
||||
- a container holding an address that had not existed for five days, on the control node
|
||||
([issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md));
|
||||
- a machine reading healthy for eleven hours while no module could reach another
|
||||
([issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md));
|
||||
- one name replacing every container on the hub
|
||||
([issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md));
|
||||
- a consumer assertion that is correct on a mesh being raised and fatal on one that is running
|
||||
([issue 156](../04-ISSUES/156-moving-a-consumers-delivery-subject-stops-the-control-plane/00-report.md)).
|
||||
|
||||
The last is the one that settles it. That change was exercised on the raise path — which is what a bed
|
||||
*is* — and the raise path is the only path on which the fault cannot appear. A bed raises a mesh; it
|
||||
does not have a mesh that has been running for weeks, with consumers already bound, containers created
|
||||
against an older roster, and an adopted machine carrying a predecessor's configuration. **The faults
|
||||
that cost the most were all faults of a mesh that already exists**, and a bed is by construction a mesh
|
||||
that does not.
|
||||
|
||||
Lab runs are also expensive in a way that changed the behaviour around them: each costs a build and
|
||||
several minutes, so they were batched, and a batched test is one whose result arrives after the next
|
||||
three changes were already written.
|
||||
|
||||
## Considered Options
|
||||
|
||||
**1. Build 0068 as proposed.** Rejected. It answers a question nobody is asking: the bottleneck was
|
||||
never that a person had to sit at the lab, it was that a bed cannot hold the state the faults live in.
|
||||
Queueing and tooling a mechanism that finds the wrong class of fault faster is not an improvement.
|
||||
|
||||
**2. Leave 0068 `proposed`.** Rejected, and it is why this record exists rather than nothing. A record
|
||||
that contradicts current practice and sits unresolved is worse than either answer: it reads as intent
|
||||
to anyone who finds it, and the practice it contradicts is written down nowhere but a handoff note.
|
||||
|
||||
**3. Record that the live mesh is the test bed, and supersede 0068.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**A change is verified against the mesh that is running.** Not because a bed would be unwelcome, but
|
||||
because the state that breaks things is state a bed does not have: containers made against an older
|
||||
roster, consumers already bound, an adopted machine, a store with weeks of history.
|
||||
|
||||
**A change that can only be exercised on the raise path is not verified.** If the only test available
|
||||
raises a fresh mesh, the record says so, and says which case was therefore not covered. The words
|
||||
"exercised on a fresh mesh" are a statement about coverage, not a pass.
|
||||
|
||||
**The lab is not retired**, and [ADR 0016](0016-the-lab.md) stands. It remains the place to raise a
|
||||
mesh from bare, which is the one thing the live mesh cannot be asked to do and the one thing a bed does
|
||||
better than anything else. What this record removes is the lab as the *default* answer to "is this
|
||||
change good", and with it 0068's queue, tools and request protocol.
|
||||
|
||||
**Accuracy over a green run.** An honest failure on the live mesh beats a pass in a bed that could not
|
||||
have failed — and a change that is risky on the running mesh is a reason to make the change smaller,
|
||||
not a reason to test it somewhere it cannot break.
|
||||
|
||||
**What 0068 got right is kept as a rule, not a mechanism:** a run reads a copy that is not anybody's
|
||||
working tree. Every run of the lab that mattered was pinned to a checkout rather than a worktree, and
|
||||
the ones that were not produced results about code nobody had written down.
|
||||
|
||||
## How this is checked
|
||||
|
||||
- **A record that says a change was verified says on what.** Where it was a fresh mesh, it says which
|
||||
case is uncovered. This is the clause that would have caught 156: its change was verified, honestly,
|
||||
on the only path where it works.
|
||||
- **The lab is not in the path of a merge.** No check, playbook or handoff requires a bed to have run.
|
||||
- **0068 is unreachable as intent.** Its status is `superseded` and it names this record, so a reader
|
||||
arriving at the queue design finds out immediately that it was not built and why.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **A fault can be introduced on the machines that serve.** This is the cost, it is real, and it was
|
||||
paid twice in one evening — a control plane crash-looping for half an hour, and every container on the
|
||||
hub recreated five times. Both were found in minutes because they were live, and both would have
|
||||
passed a bed.
|
||||
- **There is no pre-merge gate beyond the repositories' own suites.** `make check` and the three hq
|
||||
checks are what stands between a change and the machines, which raises what those suites are worth
|
||||
and makes a test that cannot fail a genuine defect rather than an untidiness.
|
||||
- **Raising a mesh from bare is now the lab's whole job**, and is exercised deliberately rather than
|
||||
as a side effect of testing something else. The foundation work
|
||||
([issue 146](../04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md))
|
||||
is that job, and it is also the proof that the mesh can make another of itself.
|
||||
- **An agent cannot hand a run to a queue and come back**, which 0068 would have given. In practice it
|
||||
watches a push and reads the machines, which is what happened anyway.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR 0068](0068-the-lab-takes-requests.md) — superseded by this
|
||||
- [ADR 0016](0016-the-lab.md) — the lab, which stands
|
||||
- [issue 156](../04-ISSUES/156-moving-a-consumers-delivery-subject-stops-the-control-plane/00-report.md) — correct on the raise path, fatal on a running mesh
|
||||
+113
@@ -0,0 +1,113 @@
|
||||
---
|
||||
topic: what runs on it
|
||||
status: accepted
|
||||
date: 2026-09-30
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 02-DECISIONS/0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md
|
||||
---
|
||||
|
||||
# 150. A module's own code runs as supervised processes under the module's one account
|
||||
|
||||
## Context
|
||||
|
||||
The repository answers "what runs a module's own code" two ways and reconciles them nowhere
|
||||
([issue 117](../04-ISSUES/117-a-modules-own-code-is-a-container-and-a-process/00-report.md)).
|
||||
|
||||
[ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) is accepted and
|
||||
says **a container** — "the tool runtime carrying that module's compiled code" — and "one module, one
|
||||
process, one account". Two `proposed` design documents say a **`process`** resource running an argv,
|
||||
supervised by the machine, and one of them declares *four* of them for a single module and presents
|
||||
four as the point. Neither design document names 0047 in its `decisions:`, and the string `process`
|
||||
as a resource type appears in no decision record at all. The thing as built is the container.
|
||||
|
||||
Two things have happened since 0047 was written that bear on it directly.
|
||||
|
||||
[**ADR 0142**](0142-the-mesh-delivers-its-own-components-as-binaries.md) decided that the mesh's own
|
||||
components are binaries on the machine rather than container images, and
|
||||
[issue 114](../04-ISSUES/114-should-the-controller-be-a-container-or-a-process/00-report.md) was closed
|
||||
by it. That settled the mesh's components and deliberately said nothing about a module's.
|
||||
|
||||
And the standing definition of a module hardened: **a module is software that delivers one or more
|
||||
services, and a module is not a container.** It may deliver them as a container, an installed package
|
||||
with a unit, a binary, or configuration files; 61 of 73 happen to use a container and 11 do not,
|
||||
including the resolver, sshd and fail2ban. A rule that a module's *own code* must be a container makes
|
||||
the one kind of module the mesh writes itself the only kind that has no choice.
|
||||
|
||||
## Considered Options
|
||||
|
||||
**1. Hold 0047 as written: a container.** Rejected. Its own reasoning does not require one. What 0047
|
||||
argued for was a runtime **per module** rather than one for the whole node, because a node-wide runtime
|
||||
could not hold a per-module broker account and per-module runtimes competing on one tool key would each
|
||||
be handed calls for tools they do not have. A supervised unit per module satisfies that argument
|
||||
exactly — it is per module, and a unit runs as an account. The container was the mechanism to hand, not
|
||||
the conclusion.
|
||||
|
||||
**2. Let each design document choose.** Rejected; that is the present state and it is what issue 117
|
||||
reports. A module author reading the guide writes four processes; a module author reading the record
|
||||
writes a container; nothing tells either that the other exists.
|
||||
|
||||
**3. Settle the hosting form as a supervised process, and settle the count separately.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**A module's own code runs as one or more supervised processes on the machine, under the module's single
|
||||
account.** Where [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) says
|
||||
"a container, the tool runtime carrying that module's compiled code", read this record. Everything else
|
||||
0047 decided stands untouched: a tool is served on its own key, only the module that serves it answers,
|
||||
and the module's account is scoped to exactly its tool keys.
|
||||
|
||||
**The invariant is the account, not the process count.** 0047's "one module, one process, one account"
|
||||
carried its weight in the last clause. Its stated worry about a second process was "not a second one to
|
||||
scope and seal" — a second *identity* to grant, seal a secret to, and scope on the bus. Several
|
||||
processes sharing the module's one account create no second identity, so nothing further is scoped or
|
||||
sealed, and a module may therefore declare as many as its work has shapes: events, tools, a
|
||||
provisioner, a scheduled ingest. **What a module may not have is two accounts.**
|
||||
|
||||
**A module that delivers its service as a container still does.** This record is about the code the
|
||||
module itself carries — its tools, its events, its provisioner — and not about the software it delivers.
|
||||
A module wrapping a third-party image wraps a third-party image.
|
||||
|
||||
**Why supervised by the machine rather than by the mesh:** it is the same answer ADR 0142 gave for the
|
||||
mesh's own components, for the same reason. A unit the machine restarts needs no image, no registry
|
||||
pull and no runtime to be up before the mesh's own code can run — which matters most for exactly the
|
||||
modules whose code the mesh cannot start any other way.
|
||||
|
||||
## How this is checked
|
||||
|
||||
- **No design document describes a hosting form for a module's own code without citing this record.**
|
||||
Designs 18 and 20 name it in `decisions:`; this is the gap issue 117's third point reports, and
|
||||
`cycle.py` already enforces that a to-be design names its decisions.
|
||||
- **A module declaring several processes resolves to one account.** A test composes a module with more
|
||||
than one process resource and asserts the mesh mints exactly one broker account for it, scoped to that
|
||||
module's tool keys and nothing else — which is 0047's invariant stated as an assertion rather than a
|
||||
sentence.
|
||||
- **A module's own code does not require the container runtime.** A machine with no container runtime
|
||||
can still run a module whose code is its own, which is the claim that separates this from option 1 and
|
||||
is checkable on a machine that has one by asserting the declaration names no image for it.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **The sidecar port stops being needed.** [ADR 0029](0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md)
|
||||
records that "anything that is a service plus a sidecar currently has to publish a port to talk to
|
||||
itself", and the host's `network` shape exists partly for it. A process beside the service on the same
|
||||
machine reaches it without publishing anything, so that pressure goes.
|
||||
- **Something must supervise, and it is the machine.** This adds a unit per module's code to what the
|
||||
host writes and owns. The mesh already writes and owns units — `nftables` proves a module can write one
|
||||
and run it — so the mechanism exists; the count grows.
|
||||
- **A module's code is delivered, not pulled**, which puts it behind the same gap as the host's own
|
||||
delivery ([ADR 0141](0141-the-host-delivers-its-own-successor.md), not built): nothing yet delivers a
|
||||
version of a module's binary to a machine. A container's code arrives by `docker pull`, and this does
|
||||
not. **This is the cost of the decision and it is not paid**; until delivery exists, a module whose code
|
||||
is its own is a module somebody places by hand.
|
||||
- **Issue 117 is answered and its three disagreements close differently:** container-or-unit is decided
|
||||
here; one-process-or-several is decided here as several under one account; and whether the record was
|
||||
consulted is fixed by designs 18 and 20 naming this one.
|
||||
|
||||
## References
|
||||
|
||||
- [issue 117](../04-ISSUES/117-a-modules-own-code-is-a-container-and-a-process/00-report.md) — the contradiction this answers
|
||||
- [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) — extended; its "a container" clause is settled here
|
||||
- [ADR 0142](0142-the-mesh-delivers-its-own-components-as-binaries.md) — the same answer for the mesh's own components
|
||||
- [ADR 0141](0141-the-host-delivers-its-own-successor.md) — the delivery this depends on and which is not built
|
||||
- [ADR 0029](0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md) — the sidecar port this relieves
|
||||
@@ -174,6 +174,7 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0108** — [A route carries the policy applied to a request, and names a secret rather than holding one](0108-a-route-carries-the-policy-applied-to-a-request.md)
|
||||
- **0109** — [A package registry seat is one per ecosystem, not one for all of them](0109-a-package-registry-seat-is-one-per-ecosystem.md)
|
||||
- **0126** — [A module declares its own seats; the mesh reserves its own](0126-a-module-declares-its-own-seats.md)
|
||||
- **0148** — [The mesh's names are resolved, not copied into every container](0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
|
||||
|
||||
### What runs on them, and how it gets there
|
||||
|
||||
@@ -208,7 +209,7 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0110** — [A seat is held by one assignment, from a closed set, and it may deliver a provision](0110-a-seat-is-a-module-assignment-from-a-closed-set.md)
|
||||
- **0112** — [A module definition names no node, no mesh and no path: everything it needs is a requirement the mesh resolves](0112-a-module-definition-names-no-node-mesh-or-path.md)
|
||||
- **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md)
|
||||
- **0114** — [A credential two parties hold rotates over two credentials; one a single party holds rotates in place, staged; and retiring a credential never removes what it reached](0114-a-shared-credential-rotates-over-two-credentials.md) *(proposed)*
|
||||
- **0114** — [A credential two parties hold rotates over two credentials; one a single party holds rotates in place, staged; and retiring a credential never removes what it reached](0114-a-shared-credential-rotates-over-two-credentials.md)
|
||||
- **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md)
|
||||
- **0117** — [A machine's uplink is a seat: the mesh configures the manager, never the link](0117-a-machines-uplink-is-a-seat.md)
|
||||
- **0118** — [Undeclaring removes what the mesh made, and gives a unit back the state it was found in](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md)
|
||||
@@ -228,6 +229,7 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0145** — [A module checks what the mesh claims is reachable, and it checks itself](0145-a-module-checks-what-the-mesh-claims-is-reachable.md) *(superseded)*
|
||||
- **0146** — [Connectivity is checked by name, per hosting form, with a valid certificate](0146-connectivity-is-checked-by-name-per-hosting-form.md)
|
||||
- **0147** — [A module anchors the mesh's authority on a machine, and takes it away again](0147-a-module-anchors-the-meshs-authority.md)
|
||||
- **0150** — [A module's own code runs as supervised processes under the module's one account](0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md)
|
||||
|
||||
### How it is built
|
||||
|
||||
@@ -239,7 +241,7 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0016** — [The lab](0016-the-lab.md)
|
||||
- **0037** — [Where a module lives](0037-where-a-module-lives.md)
|
||||
- **0039** — [What the SDK holds, and what it refuses](0039-what-the-sdk-holds-and-refuses.md)
|
||||
- **0068** — [The lab takes requests, one at a time, and runs each from its own copy](0068-the-lab-takes-requests.md) *(proposed)*
|
||||
- **0068** — [The lab takes requests, one at a time, and runs each from its own copy](0068-the-lab-takes-requests.md) *(superseded)*
|
||||
- **0069** — [A module is a repository and a path within it](0069-a-module-is-a-repository-and-a-path.md)
|
||||
- **0076** — [The SDK is a published package, and the toolchain resolves it by version](0076-the-sdk-is-a-published-package.md)
|
||||
- **0082** — [The registry is reached by name, and the overlay is its security](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)
|
||||
@@ -248,6 +250,7 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0097** — [A vendor image is a declared build input, and a recipe fetches nothing undeclared](0097-a-vendor-image-is-a-declared-build-input.md)
|
||||
- **0107** — [Persistent data is a directory bind, never a named volume](0107-persistent-data-is-a-directory-bind-never-a-named-volume.md)
|
||||
- **0111** — [A build source is on the mesh's git seat, or it is an external repository](0111-a-build-source-is-on-the-git-seat-or-external.md)
|
||||
- **0149** — [The live mesh is the test bed](0149-the-live-mesh-is-the-test-bed.md)
|
||||
|
||||
### How it is checked
|
||||
|
||||
|
||||
@@ -7,8 +7,9 @@ code:
|
||||
- mesh-controller internal/identity/authority.go
|
||||
- mesh-host internal/identity/serving.go
|
||||
- mesh-host internal/apply (the service that reflects a rule set)
|
||||
updated: 2026-09-29
|
||||
updated: 2026-09-30
|
||||
decisions:
|
||||
- 02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md
|
||||
- 02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md
|
||||
- 02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md
|
||||
- 02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md
|
||||
@@ -289,6 +290,30 @@ hosts file by the runtime. That extends the file decision rather than overturnin
|
||||
mesh and not chosen by a module: a module that listed the machines would go stale the day one
|
||||
joins, and a module that did not would be one whose containers cannot reach anything by name.
|
||||
|
||||
**Superseded for containers by [ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
|
||||
(2026-09-30).** Everything above still describes the machine's own roster file, which is how it works
|
||||
and how it will keep working. It no longer describes containers.
|
||||
|
||||
Copying the roster into each container made the roster part of each container's identity, so one name
|
||||
moving replaced every container in the mesh — a module assigned on one machine restarted the store, the
|
||||
registry, the edge and mail on another
|
||||
([issue 151](../../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). And it
|
||||
did not stop the staleness it was meant to: a copy taken at creation is stale the moment the roster
|
||||
moves, twice found as a container holding an address that had not existed for days
|
||||
([issues 109](../../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md)
|
||||
and [135](../../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md)).
|
||||
|
||||
**A container resolves the mesh's names through its machine's resolver, at the moment it asks, and
|
||||
nothing is copied.** The resolver is a machine-level process rather than a container, so nothing
|
||||
circular is being asked for. This is gated on a container being able to reach the resolver from any of
|
||||
the runtime's networks, which it cannot today
|
||||
([issue 110](../../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)) —
|
||||
until that lands the mesh keeps copying and keeps comparing, and the order is stated in the record.
|
||||
|
||||
The paragraph below states the old boundary, and 0148 deliberately gives it up: a container somebody
|
||||
started by hand resolves the same names as everything else, because the resolver answers the machine,
|
||||
not a list of containers.
|
||||
|
||||
**The boundary, which is deliberate and worth stating:** *declared* containers. A container
|
||||
somebody starts by hand is not the mesh's to configure, and reaching into every container on a
|
||||
machine — declared or not — is what a nameserver in `resolv.conf` would be for.
|
||||
|
||||
@@ -7,6 +7,7 @@ code:
|
||||
- mesh-catalog modules/builder
|
||||
updated: 2026-09-29
|
||||
decisions:
|
||||
- 02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md
|
||||
- 02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md
|
||||
- 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md
|
||||
- 02-DECISIONS/0097-a-vendor-image-is-a-declared-build-input.md
|
||||
|
||||
@@ -7,6 +7,7 @@ code:
|
||||
- mesh-sdk src
|
||||
updated: 2026-09-21
|
||||
decisions:
|
||||
- 02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md
|
||||
- 02-DECISIONS/0099-a-step-that-runs-once-names-what-it-reads.md
|
||||
- 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md
|
||||
- 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md
|
||||
|
||||
+16
@@ -51,3 +51,19 @@ knows that is what the rule means.
|
||||
the runtime's default one? That is a stronger rule and would have prevented 109 as well.
|
||||
- What checks it? A converged bed with a container on the default network resolving a mesh name is
|
||||
the missing assertion; nothing in the resolver's own beds covers the filter.
|
||||
|
||||
## What now depends on this (2026-09-30)
|
||||
|
||||
This stopped being a container-DNS inconvenience.
|
||||
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) decides
|
||||
that a container resolves the mesh's names rather than being given a copy of them, which is what stops
|
||||
one name moving from replacing every container in the mesh
|
||||
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)) and what makes a
|
||||
stale address impossible rather than merely noticed
|
||||
([issues 109](../109-a-container-keeps-the-address-it-was-made-with/00-report.md)
|
||||
and [135](../135-a-containers-mesh-names-are-not-compared/00-report.md)).
|
||||
|
||||
**That decision cannot land until this one does**, and not partly: a container on the runtime's default
|
||||
network is the case with no DNS at all, and it is the case the mesh's own forge runs in. Two of four
|
||||
machines also bind the resolver to loopback only, so the runtime hands their containers a public
|
||||
resolver. Both halves are this issue.
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: located
|
||||
status: resolved
|
||||
opened: 2026-09-25
|
||||
located-in: [hq, mesh-catalog modules/showcase, mesh-sdk src/tools/index.ts, mesh-tools]
|
||||
fixed-by:
|
||||
fixed-by: hq ADR 0150 — a module's own code runs as supervised processes under the module's one account; designs 18 and 20 now cite it, and ADR 0047 carries a dated note pointing at it
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -99,3 +99,27 @@ what the sidecar is has nowhere in the design layer to look, which is how this w
|
||||
is mechanically checkable: the resource types a design doc names are a closed set, and every
|
||||
member of it either appears in a decision or does not. Whether that check is worth writing is
|
||||
part of this issue, not settled by it.
|
||||
|
||||
## Answered (2026-09-30)
|
||||
|
||||
[ADR 0150](../../02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md)
|
||||
settles all three disagreements, and the design documents win two of them:
|
||||
|
||||
1. **Container or unit — a supervised process.** ADR 0047's argument never required a container. It
|
||||
argued for a runtime *per module*, because a node-wide one could not hold a per-module account and
|
||||
per-module runtimes on one tool key would be handed calls for tools they do not have. A unit per
|
||||
module satisfies that exactly, and a unit runs as an account. The container was the mechanism to
|
||||
hand. [ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md) had
|
||||
already gone the same way for the mesh's own components, and a module is not a container.
|
||||
2. **One process or several — several, under one account.** 0047's "one module, one process, one
|
||||
account" carried its weight in the last clause; its stated worry was "not a second one to scope and
|
||||
seal", which is about a second *identity*. Processes sharing the module's one account create none.
|
||||
What a module may not have is two accounts.
|
||||
3. **Whether the record was consulted — fixed rather than answered.** Designs 18 and 20 now name 0150
|
||||
in `decisions:`, and 0047 carries a dated note saying where its hosting form was settled, so neither
|
||||
door leads to the wrong answer any more.
|
||||
|
||||
**What this does not fix.** A module's code delivered as a binary is behind the same gap as the host's
|
||||
own ([ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md), accepted and not
|
||||
built): a container's code arrives by `docker pull` and this does not, so until delivery exists such a
|
||||
module is one somebody places by hand. 0150 records that as the cost of the decision, unpaid.
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
---
|
||||
status: located
|
||||
status: resolved
|
||||
opened: 2026-09-28
|
||||
located-in:
|
||||
- mesh-controller internal/catalogue/manifest.go
|
||||
@@ -7,7 +7,7 @@ located-in:
|
||||
- mesh-controller internal/catalogue/declaration.go
|
||||
- mesh-controller examples/route-proxy
|
||||
- mesh-catalog (every routed module manifest)
|
||||
fixed-by:
|
||||
fixed-by: mesh-controller bdf965d (a module names its endpoints) and c68d3a7 (an assignment configures an endpoint as one thing) — the filter, the proxy's names and both authorities now read one statement
|
||||
amended-design: 03-DESIGN/01-to-be/08-connectivity.md
|
||||
---
|
||||
|
||||
|
||||
@@ -0,0 +1,47 @@
|
||||
# Resolution
|
||||
|
||||
*2026-09-29.*
|
||||
|
||||
**Built, and this record did not say so.** The issue was written on 2026-09-28 and answered the same
|
||||
week by two commits in `mesh-controller`; nothing came back to close it, so the mesh's own account of
|
||||
itself said for a day that reach was declared nowhere while the code read it in three places.
|
||||
|
||||
- `bdf965d` — *a module names its endpoints, and a route names the one it serves*. `listens[].name`
|
||||
is the endpoint; a route contribution names the endpoint rather than repeating a port.
|
||||
- `c68d3a7` — *an assignment configures an endpoint as one thing*. The `endpoints` settings key, per
|
||||
node, by endpoint name: `{"endpoints": {"ssh": {"port": 20134, "reach": "public"}}}` — port, label
|
||||
and reach in one block, which is what [ADR 0138](../../02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md)
|
||||
asked for and what [ADR 0046](../../02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md)
|
||||
said configuration is.
|
||||
|
||||
## The three readers, which is what the issue was about
|
||||
|
||||
The complaint was that the per-node source override had exactly one caller. It now has three, and
|
||||
they are the three mechanisms reach was decided to settle at once:
|
||||
|
||||
| reader | what it does with it |
|
||||
|---|---|
|
||||
| the filter | `Reaches` turns each endpoint's reach into the rule for its machine port |
|
||||
| the proxy's names | `composeName` composes the public name, the internal name, or both — and a name nobody asked for is not composed |
|
||||
| the authorities | the proxy certifies only names it was actually given, each from its own authority, through two host policies rather than one |
|
||||
|
||||
**A routed endpoint keeps the manifest's port**, which is ADR 0138's own insight and older than it
|
||||
([ADR 0045](../../02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)): the
|
||||
proxy is how it is reached, so `public` there asks for a public *name*, not an open port.
|
||||
|
||||
**An endpoint that is not routed is reached and never named.** Git over ssh is that case — the one
|
||||
the issue said the model could not express — and it is now the ordinary one.
|
||||
|
||||
## How it is checked
|
||||
|
||||
`internal/catalogue/endpoints_setting_test.go`: a block says port, label and reach; a block may say
|
||||
only a reach; a name the module does not declare is refused; a reach outside the four values is
|
||||
refused; and saying the same thing twice — once in the block, once through the older per-port keys —
|
||||
is refused rather than resolved by whichever is read last. The proxy's half is `policy_test.go` and
|
||||
`authority_test.go`: a name the mesh did not send is not certified, by either authority.
|
||||
|
||||
## What is left, and it is not this
|
||||
|
||||
The older keys (`ports`, `expose`, and reach keyed by port) still work beside the block. They are
|
||||
what the block replaces, and retiring them is its own small change — not a gap in what reach can
|
||||
say.
|
||||
+12
@@ -71,3 +71,15 @@ against a suite nobody can execute.
|
||||
- The lab rewrites a bundle's registry-prefixed third-party references to upstream ones for a
|
||||
machine with an uplink (`test/integration/harness.ts`); the new bundle's bus reference needed
|
||||
that rule added, which is done and is not this issue.
|
||||
|
||||
## What one of its fixes then did to a running mesh (2026-09-30)
|
||||
|
||||
The change that stopped the doubling — putting the stream into a push consumer's delivery subject —
|
||||
is correct on a foundation being raised and fatal on a mesh that is already running: the server will
|
||||
not move that subject while a subscriber is bound, and a node is bound to its declaration consumer
|
||||
the whole time it is up. The control plane crash-looped on the first build that carried it.
|
||||
|
||||
Recorded and fixed as [issue 156](../156-moving-a-consumers-delivery-subject-stops-the-control-plane/00-report.md).
|
||||
Noted here because this record is where somebody will arrive when reading why the subject carries the
|
||||
stream at all, and the answer is incomplete without it: **the raise path was the only one exercised,
|
||||
and it is the one path on which nothing is bound.**
|
||||
|
||||
@@ -0,0 +1,52 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-29
|
||||
located-in:
|
||||
- mesh-controller internal/catalogue/declaration.go (contributions are emitted for an assigned module, taken or not)
|
||||
fixed-by:
|
||||
---
|
||||
|
||||
# 150 — A route is contributed before its module is taken
|
||||
|
||||
## What was observed
|
||||
|
||||
ace is adopted and runs `route-adapter` beside the predecessor's traefik (ADR 0104). The first web
|
||||
module migrated there was searxng. `assign ace searxng` + `push ace` — the step that is supposed to
|
||||
change nothing on an adopted machine (ADR 0100, hq 125: assign holds, take replaces) — took
|
||||
`searxng.zurag.be` down:
|
||||
|
||||
```
|
||||
https://searxng.zurag.be/ → 502
|
||||
```
|
||||
|
||||
for about five minutes, until the module was unassigned again.
|
||||
|
||||
## Why
|
||||
|
||||
Assign held everything it found on the machine — the predecessor's `searxng` container, its
|
||||
directories — exactly as designed. But the module's **route contribution** is not a resource on the
|
||||
machine, so nothing held it: it reached `route-adapter` at once, which wrote
|
||||
`mesh-searxng.zurag.be.yml` into traefik's file provider pointing at the mesh's assigned machine port
|
||||
(`http://ace.internal:20000`) — where nothing listened, because the container that would was held.
|
||||
traefik's file router for the name then won over the predecessor's docker-label router for the same
|
||||
name, and the name served a dead backend.
|
||||
|
||||
On this occasion the window was lengthened by an unrelated failure (the machine's docker address
|
||||
pools were exhausted, so the module's network could not be created), but the fault does not depend on
|
||||
it: **between assign and take, every routed module's public name points at a backend the mesh has
|
||||
deliberately not started.** The runbook's §6 step 2 ("check the preparation — HAL's service still
|
||||
serves") is false for every routed module on a node running route-adapter.
|
||||
|
||||
## What the operator did
|
||||
|
||||
Rolled back per runbook (unassign before take), then re-ran as assign → push → wait for the node to
|
||||
report its holds → take → push, keeping the window to the ~10 s between the two pushes. A workaround,
|
||||
not a fix: a module whose data must move between assign and take (runbook §6 step 3) cannot shrink
|
||||
its window this way.
|
||||
|
||||
## What would be right
|
||||
|
||||
A contribution from a module that is assigned but **not taken** on an adopted node should be held like
|
||||
the module's resources are — withheld from the provider until take — or the provider should be told
|
||||
the contributor is held and leave the predecessor's route alone. Either keeps "assign changes nothing"
|
||||
true for routed modules.
|
||||
@@ -0,0 +1,94 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-29
|
||||
located-in:
|
||||
- mesh-host internal/apply/apply.go (containerSpecReading hashes every `host` entry)
|
||||
- mesh-controller internal/catalogue/declaration.go (withMeshNames gives every container the mesh's names)
|
||||
fixed-by:
|
||||
---
|
||||
|
||||
# 151 — A new name recreates every container in the mesh
|
||||
|
||||
## What was observed
|
||||
|
||||
Migrating one small module on ace (searxng) took four routine controller actions: `node public-domain
|
||||
ace zurag.be`, `assign ace searxng` + push, `unassign ace searxng` + push (a rollback, see
|
||||
[issue 150](../150-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take
|
||||
+ push again. Each push to ace also pushed novox ("this push left g14, novox, shanks behind … sending it
|
||||
too"). novox's host then **replaced every container it runs, twice**:
|
||||
|
||||
```
|
||||
22:46 … updated postgres.server (mesh-store): replaced; a container's configuration is fixed when it is created
|
||||
22:46 … updated distribution.store (mesh-registry): replaced; …
|
||||
22:46 … updated route-proxy.server (route-proxy): recreated …
|
||||
22:51 … updated postgres.server (mesh-store): replaced; …
|
||||
22:52 … updated gitea.server (gitea): replaced; …
|
||||
22:52–22:55 mesh-vault, builder, all of mailu, keycloak, minio, mongodb, invoicing, photos, umami, …
|
||||
```
|
||||
|
||||
Replacements per minute on novox, 22:46–22:55: 12, 8, 13, 5, 6, 8, 8, 8, 8, 4. The control plane was
|
||||
unreachable twice while its own store came back through crash recovery
|
||||
(`FATAL: the database system is starting up`), and dependants (umami) crash-looped until it did. None
|
||||
of the replaced containers belonged to the module being migrated, or to ace.
|
||||
|
||||
## Why (confirmed part)
|
||||
|
||||
Every container the mesh runs is given the mesh's names as `--add-host` entries
|
||||
(`withMeshNames`), and the host puts every entry into the container's spec digest — deliberately,
|
||||
since [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md): a container left alone when the roster moved kept a five-day-old
|
||||
address and restarted 2286 times. A running container cannot have its hosts changed, so a changed
|
||||
digest means a replace.
|
||||
|
||||
The consequence is that **the roster is part of every container everywhere**: anything that adds,
|
||||
removes or moves one name — a node's public domain, a routed name, an unassign — replaces every
|
||||
container on every machine that carries the list. On the hub that includes the control plane's store,
|
||||
the registry, the edge and mail.
|
||||
|
||||
## Not yet established
|
||||
|
||||
Which name moved in each pass. gitea's hosts after the fact list the `.internal` names and novox's
|
||||
internally routed `*.novox.be` names, and **not** `searxng.zurag.be` — so it is not simply "a routed
|
||||
name was added". Two passes suggest the set changed at assign and changed back at unassign; the
|
||||
declarations before and after would say, and nothing on the machine records the previous one.
|
||||
|
||||
## Why it matters now
|
||||
|
||||
The node-by-node migration adds names one module at a time. On ace alone that is ~25 web modules; if
|
||||
each assignment moves the roster, each is a full restart of every hub service, and a rollback is
|
||||
another. The migration is paused on this.
|
||||
|
||||
## What has since been ruled out as a cause (2026-09-30)
|
||||
|
||||
The roster was also moving for a reason that was not a name changing at all: a node whose plan would
|
||||
not compose had its routed names silently dropped from the roster handed to every machine, so a
|
||||
briefly unreachable store withdrew and restored a name on alternating passes. That is
|
||||
[issue 152](../152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md), and it is
|
||||
fixed — it is what made the churn on the control node repeat every six minutes for twenty-five
|
||||
minutes rather than once per operator action.
|
||||
|
||||
**It does not close this record.** A roster that changes for a real reason — a machine joining, a
|
||||
module assigned, a public domain set — still replaces every container on every machine that carries
|
||||
the list. 152 removed the false reasons; the question below is still open.
|
||||
|
||||
## What would be right (for diagnosis)
|
||||
|
||||
Keep 135's guarantee (no container runs with a stale address) without making the roster part of every
|
||||
container's identity — e.g. resolve mesh names through a resolver the container asks at lookup time
|
||||
rather than baked entries, or scope each container's entries to the names it actually binds.
|
||||
|
||||
## Answered (2026-09-30): the first of those two
|
||||
|
||||
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) takes
|
||||
the resolver, not the scoping. Scoping was the close call and was rejected for contradicting the
|
||||
standing intent that anything on the mesh can call anything on it: it would make a name reachable only
|
||||
where the mesh was told in advance it would be wanted, and it would leave the roster in the digest, so
|
||||
the churn returns whenever a widely-bound name moves.
|
||||
|
||||
Resolving at lookup time makes staleness impossible rather than detected — 109 and 135 stop being bugs
|
||||
that were fixed and become a shape that does not exist — and makes a name's blast radius nothing.
|
||||
|
||||
**This record stays open**, because the record answers it and the code does not. Nothing may stop
|
||||
copying names until a container can reach the resolver from any of the runtime's networks
|
||||
([issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)),
|
||||
which it cannot on two of four machines today. Removing the copy first reintroduces 109 and 135
|
||||
silently, on a live mesh, which is how both were found. The order is in the record.
|
||||
@@ -0,0 +1,124 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-29
|
||||
located-in:
|
||||
- mesh-controller cmd/mesh-controller/plan.go (routeNamesInTheMesh skips a node whose plan will not compose)
|
||||
fixed-by: mesh-controller 6c5dfd0 (PR 147)
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 152 — A node whose plan will not compose silently removes its names from every machine
|
||||
|
||||
## What was observed
|
||||
|
||||
For twenty-five minutes, and for seventeen of them after the last operator action, the control-node's
|
||||
host applied all 327 of its resources every ~6.5 minutes without pause, replacing every container on
|
||||
the machine each time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail,
|
||||
and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine
|
||||
sat near 8. Nothing was converging: each pass ended and the next began four seconds later.
|
||||
|
||||
The two machines carrying no containers were not churning. They were only knocked off the bus each
|
||||
time the control node re-created it, reconnected, re-heard the same declaration and applied it again
|
||||
as a no-op.
|
||||
|
||||
## Why: the roster alternates between two values, and it is part of every container
|
||||
|
||||
Two consecutive declarations were compared by reading the `--add-host` entries of four containers
|
||||
the moment each pass created them:
|
||||
|
||||
```
|
||||
23:02 ace.internal drive.novox.be g14.internal keycloak.novox.be novox.internal
|
||||
office.novox.be portainer.novox.be shanks.internal umami.novox.be (9 names)
|
||||
23:07 … the same nine, and searxng.zurag.be (10 names)
|
||||
```
|
||||
|
||||
One routed name — belonging to a module on another machine entirely — leaves the roster and comes
|
||||
back. Because the roster is part of every container's spec digest ([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)),
|
||||
each flip is a different identity for every container on the machine, and a running container cannot
|
||||
have its hosts changed. So every flip replaces all of them.
|
||||
|
||||
**What makes it flip is a swallowed error.** `routeNamesInTheMesh` composes every node's plan to
|
||||
find the names it serves, and when one will not compose it moves on:
|
||||
|
||||
```
|
||||
plan, settings, err := planFor(ctx, open, n.Name)
|
||||
if err != nil {
|
||||
continue
|
||||
}
|
||||
```
|
||||
|
||||
A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does
|
||||
not know", but "the mesh states these names do not exist", to every machine at once.
|
||||
|
||||
## Why it sustains itself
|
||||
|
||||
The loop closes through the control plane's own database:
|
||||
|
||||
1. An apply replaces `mesh-store` — the store the control plane reads — by removing the container,
|
||||
so postgres comes back through crash recovery.
|
||||
2. While it recovers it refuses connections: `FATAL: the database system is not yet accepting
|
||||
connections / Consistent recovery state has not been yet reached` (observed, 21:08:39 UTC, every
|
||||
pass).
|
||||
3. `planFor` for the other machine fails against that store. `routeNamesInTheMesh` swallows it and
|
||||
drops its routed name.
|
||||
4. The roster changed, so all 327 resources differ, so all are replaced — including `mesh-store`,
|
||||
and including the bus, which is why the host also cannot report: `applied, and could not tell the
|
||||
mesh: reporting: nats: connection closed`.
|
||||
5. Back to 1.
|
||||
|
||||
Every pass destroys the evidence the next pass needs to decide it has nothing to do, and nothing
|
||||
outside the machine has to be wrong for it to continue.
|
||||
|
||||
**It is metastable, not permanent.** It ran from 22:46 to 23:11 — five full replacements of every
|
||||
container on the machine — and then stopped on its own, when one pass happened to read the store
|
||||
during a window it was up, composed the same roster twice running, and found nothing to do. Load fell
|
||||
from 7.8 to 1.7 and the machine returned to its five-minute idle tick.
|
||||
|
||||
That it ends by luck is the point, not a mitigation. The exit condition is a race the mesh does not
|
||||
control, the operator cannot see, and nothing reports; the same four actions on a slower machine, or
|
||||
a larger store, would not have found it. An outage that clears itself after twenty-five minutes and
|
||||
five restarts of the forge, the directory and mail is not a smaller fault than one that does not — it
|
||||
is the same fault, harder to catch.
|
||||
|
||||
## What it is not
|
||||
|
||||
- Not the operator's four actions on the other machine. Those explain the first passes
|
||||
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)); they were
|
||||
finished seventeen minutes and three full passes before these measurements.
|
||||
- Not a file that keeps changing. `/etc/hosts`, the bus's account list and the vault's export were
|
||||
hashed across passes and are byte-identical, and the directory reported `mode 755 to 700` every
|
||||
pass while already being `700`. Those resources are **misreported as changed** and are worth their
|
||||
own question, but they are not what moves a container's identity.
|
||||
- Not the lost report alone. A report that cannot be delivered explains a re-apply; it does not
|
||||
explain a re-apply that finds 327 differences.
|
||||
|
||||
## Why it matters beyond this outage
|
||||
|
||||
The same `continue` makes every routed name in the mesh conditional on every node's plan composing at
|
||||
the moment any machine is pushed to. One unreachable or half-migrated machine is enough to withdraw
|
||||
its names from everywhere — and the withdrawal is indistinguishable, on the receiving machine, from
|
||||
the operator having removed them.
|
||||
|
||||
The codebase already states the rule this breaks, forty lines away, about the same kind of lookup:
|
||||
|
||||
> A lookup failure is an error, never "not found": collapsing the two composed a declaration without
|
||||
> the trust whenever the inventory hiccuped, delivered by a push that reported success.
|
||||
|
||||
## How it was fixed, and how the fix is checked
|
||||
|
||||
`planFor` now marks the two failures that really are a statement about the node — its set not
|
||||
composing, and a setting that reaches nothing — and the three gatherers pass over those and only
|
||||
those. Every other failure is raised, naming the machine and the read.
|
||||
|
||||
Five tests hold it: a set that cannot compose is marked as the node's own; a store that cannot be
|
||||
read is *not*; one incoherent node still does not cost the rest their names; a roster is never
|
||||
returned beside an error; and the raised failure names what could not be read.
|
||||
|
||||
The two sibling gatherers were audited and fixed the same way — the grant composer, which would have
|
||||
withheld a consumer's credential, and the private-network membership, which would have taken a
|
||||
machine off the overlay. Three other `planFor` callers were audited and left alone: they refuse or
|
||||
report rather than silently withdraw, which is the safe direction.
|
||||
|
||||
**This does not close [issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md).**
|
||||
A roster that changes for a real reason still replaces every container in the mesh. This removes the
|
||||
false reasons; whether the roster belongs in a container's identity at all is that record's question.
|
||||
@@ -0,0 +1,57 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-29
|
||||
located-in:
|
||||
- mesh-controller internal/catalogue/dir_into.go (dirsFor: a stated path or <data root>/<module>/<id>, nothing else)
|
||||
- mesh-controller (accesses: the path is the manifest's literal)
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 153 — An adopted machine's data cannot be placed where it is
|
||||
|
||||
## What was observed
|
||||
|
||||
Preparing ace's media modules (plex, sonarr, radarr, lidarr, bazarr, nzbget, qbittorrent, bookshelf)
|
||||
for migration. ace is adopted; its data is where the predecessor put it and **must stay there**:
|
||||
|
||||
- the library and download spool: `/storage/media/*`, `/storage/downloads` — a separate ZFS pool,
|
||||
~40 TB, the operator's shared data (ADR 0051);
|
||||
- plex's own state: `/mnt/plex/{config,data,temp}` — 133 GB on a second disk;
|
||||
- large configuration directories held in place: lidarr 46 GB, radarr 17 GB, sonarr 3.2 GB.
|
||||
|
||||
The catalogue's manifests name `/services/media/*` (as `accesses`) and `/services/<m>/config` (as
|
||||
owned directories), which is novox's layout, not ace's, and not a value a definition may carry.
|
||||
|
||||
## What was decided, and what exists
|
||||
|
||||
[ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md) (accepted) says
|
||||
exactly what is needed:
|
||||
|
||||
> *Where* it is on the machine is the assignment's. A node has a default layout, and an assignment may
|
||||
> place a directory elsewhere: on a second disk, or where an adopted machine's data already is.
|
||||
|
||||
> [ADR 0051]: an access keeps its shape and its semantics; its path moves from the definition to the
|
||||
> assignment.
|
||||
|
||||
What the control plane implements (`dirsFor`): a directory is either a path the manifest states, or
|
||||
`<data root>/<module>/<id>` under the node's one data root. There is **no per-assignment placement of
|
||||
one directory**, and an `access` path is the manifest's literal — no setting reaches either.
|
||||
|
||||
## Consequence
|
||||
|
||||
Every module whose data an adopted machine already holds somewhere other than the default layout can
|
||||
only be migrated by (a) writing the machine's path into the manifest — which 0112 forbids and which
|
||||
is wrong on the next machine — or (b) moving the data into the placed layout in a window. (b) is
|
||||
acceptable for a 40 MB configuration and impossible for a 40 TB library the operator has ruled must
|
||||
never be moved, copied or re-owned.
|
||||
|
||||
The same gap covers ownership: the predecessor runs ace's media stack as `1001:2000`; a manifest's
|
||||
`owner` is one value for every machine.
|
||||
|
||||
## What would be right
|
||||
|
||||
The two assignment halves 0112 decided: a setting that places a declared directory (by id) at a given
|
||||
path on this node, and a setting that says where an access's data is — both validated like
|
||||
`endpoints` (unknown ids refused), and an access placed by the assignment still never created,
|
||||
chowned or removed.
|
||||
@@ -0,0 +1,42 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-29
|
||||
located-in:
|
||||
- hq 02-DECISIONS/0138 (reach: internal | public | both)
|
||||
- mesh-controller internal/catalogue/filtering.go (Reaches)
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 154 — A machine's own network is not a reach
|
||||
|
||||
## What was observed
|
||||
|
||||
Preparing ace's modules. ace sits on a home network (192.168.1.0/24) behind a router, and several of
|
||||
its services are reached **from that network by devices that will never be mesh machines**:
|
||||
|
||||
- mosquitto `1883` — an IoT light switch (`sonoff-office-light-switch`) and home-assistant;
|
||||
- unifi `8080`/`3478 udp`/`10001 udp` — the access points' inform, STUN and discovery;
|
||||
- plex `32400` — LAN streaming clients (three connected at survey time);
|
||||
- home-assistant `8123`, and the resolver on the LAN address.
|
||||
|
||||
ADR 0138 gives an endpoint's reach as `internal` (the private overlay), `public` (anywhere) or `both`.
|
||||
None of them says *this machine's own network*. The predecessor could: its unifi manifest opened
|
||||
inform/STUN/discovery `from: 192.168.0.0/16, 10.0.0.0/8, 172.16.0.0/12`.
|
||||
|
||||
## Consequence
|
||||
|
||||
The only reach that includes a LAN device is `public`. While ace is adopted that is harmless — its
|
||||
own firewall stays and admits the LAN — and behind NAT "anywhere" happens to mean the LAN. But:
|
||||
|
||||
- it states the wrong thing: an operator reading `reach: public` on an IoT broker believes it is on
|
||||
the internet, and a router port-forward added later for something else makes it so;
|
||||
- at `converge ace`, the mesh's filter is the sum of what it listens on (ADR 0045). An endpoint left
|
||||
`internal` cuts every LAN device off at the flip; one set `public` opens it to the internet on any
|
||||
machine with a public address.
|
||||
|
||||
## What would be right (for diagnosis)
|
||||
|
||||
A reach — or a source — that means the networks the machine is directly attached to (its uplink's
|
||||
subnets, as the machine reports them), so a LAN-only service is declared as exactly that and the
|
||||
filter can admit it without admitting the internet.
|
||||
@@ -0,0 +1,58 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-29
|
||||
located-in: [hq 00-META/checks/cycle.py]
|
||||
fixed-by: hq f89aef9 (PR 192)
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 155 — Two records may share a number, and every check passes
|
||||
|
||||
## What was observed
|
||||
|
||||
On 2026-09-29 two machines opened issues against this repository within the same hour. Both read
|
||||
`main` correctly and both took "the next free number", and they collided twice:
|
||||
|
||||
| | one machine opened | the other had already used |
|
||||
|---|---|---|
|
||||
| first | 147, 148 | 147, 148 on an unmerged branch |
|
||||
| second | 149, 150 | 149 on an unmerged branch, 150 from renumbering the first collision |
|
||||
|
||||
The first collision was reconciled by hand before merging. The second was **merged into `main`**, and
|
||||
`records.py`, `cycle.py` and `index.py` all reported success over a tree holding
|
||||
`149-a-declaration-that-shrinks-to-empty` beside `149-an-adopted-machines-data-cannot-be-placed-where-it-is`,
|
||||
and two folders numbered 150.
|
||||
|
||||
## Why
|
||||
|
||||
The number is allocated as `max(main) + 1`, and `main` lags every open pull request — seven of them
|
||||
that evening. Two readers of the same `main` therefore compute the same next number, and neither is
|
||||
doing anything wrong. The existing reconciliation precedent (a second record numbered 127 became 149)
|
||||
assumed a single writer, which stopped being true when a second machine began filing its own findings.
|
||||
|
||||
## Why it matters
|
||||
|
||||
An issue number is how every other record cites this one — `fixed-by:`, `located-in:`, a decision
|
||||
record's consequence, a commit message. Two records answering to one number is a citation that
|
||||
resolves to whichever folder the reader happened to open, and the failure is silent on both sides:
|
||||
the citer is not wrong, and the cited record exists.
|
||||
|
||||
It is also exactly the class this repository says it does not permit — a rule (`00-META/process/03-issues.md`:
|
||||
"take the next free number") enforced by nothing.
|
||||
|
||||
## How it was fixed, and how the fix is checked
|
||||
|
||||
`cycle.py` now refuses a tree in which two issue folders share a leading number, and names both.
|
||||
Proven by adding a duplicate and watching it fail, then removing it and watching it pass.
|
||||
|
||||
The colliding records were renumbered 153 and 154, in the branch that landed last — renumbering a
|
||||
branch whose author is still pushing only moves the race.
|
||||
|
||||
**The check catches the collision; it does not prevent it.** Allocating a number still needs the open
|
||||
pull requests read as well as `main`. That is a habit the check now backstops rather than one it
|
||||
replaces, and playbook [03](../../00-META/process/03-issues.md) now says so at the step where the
|
||||
number is taken.
|
||||
|
||||
[ADR 0080](../../02-DECISIONS/0080-the-development-cycle-is-checked.md) enumerates what `cycle.py`
|
||||
enforces and named four things; it carries a progressive insight naming the fifth. The decision
|
||||
stands — this is one more thing frontmatter and file names can carry, found by its absence.
|
||||
+98
@@ -0,0 +1,98 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-29
|
||||
located-in:
|
||||
- mesh-controller internal/broker/jetstream.go (EnsureConsumer)
|
||||
fixed-by: mesh-controller e7da39d (PR 148)
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 156 — Moving a consumer's delivery subject stops the control plane, and only on a mesh that is running
|
||||
|
||||
## What was observed
|
||||
|
||||
The control plane crash-looped, every restart ending the same way:
|
||||
|
||||
```
|
||||
mesh-controller: asserting how novox hears its declaration:
|
||||
bringing consumer novox on NODES to match: nats: consumer name already in use
|
||||
```
|
||||
|
||||
It came up on the first build of the controller in eight hours. The machines kept running what they
|
||||
already held — this stops the mesh being *changed*, not the services it placed — and nothing could be
|
||||
pushed, no report was consumed and no enrolment answered, for as long as it lasted.
|
||||
|
||||
## Why
|
||||
|
||||
[Issue 146](../146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md) put the
|
||||
stream into a push consumer's delivery subject, because one process holding two consumers of the same
|
||||
name on two streams was given one subject and acted on every message twice.
|
||||
|
||||
**The server will not move a push consumer's delivery subject while a subscriber is bound to it.** It
|
||||
refuses with `consumer name already in use` — a message about the name, for a conflict about the
|
||||
subject, which is why the trail starts in the wrong place.
|
||||
|
||||
A node is bound to its declaration consumer the whole time it is up. That *is* a node listening for
|
||||
what it should be. So every node consumer in a mesh that is running is one the assertion cannot bring
|
||||
to match — and the assertion happens before the controller serves, so it never serves.
|
||||
|
||||
The controller's own two consumers moved without trouble, and are on the new subject in the live mesh.
|
||||
It asserts them before it subscribes, so nothing was bound.
|
||||
|
||||
## Why nothing caught it
|
||||
|
||||
The change was exercised on a mesh being raised, where every consumer is created rather than updated
|
||||
and nothing is bound to any of them. On that path the code is correct. The test that would have caught
|
||||
it needs a mesh that is already running: an existing consumer, a subscriber still attached, and then
|
||||
the assertion.
|
||||
|
||||
Reproduced exactly that way before the fix — same server version, same stream shape, same consumer —
|
||||
and it fails with the same words as the machine did. An earlier version of the same test unsubscribed
|
||||
first and passed against the code that was crash-looping on the control node.
|
||||
|
||||
## What it is not
|
||||
|
||||
- Not the change it shipped beside. The merge that triggered this build carried
|
||||
[issue 152](../152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md)'s fix and three
|
||||
other commits that had never been deployed; this one is 146's.
|
||||
- Not a version difference. The test server and the mesh's broker are both nats-server v2.10.29.
|
||||
|
||||
## How it was fixed
|
||||
|
||||
The consumer that works is kept, and the assertion says so instead of failing.
|
||||
|
||||
**Not deleted and re-made.** Re-making moves the subject, and a holder may not yet be allowed to
|
||||
subscribe to the new one: the wider grant travels in the bus's user list, which the control plane
|
||||
composes and a machine applies minutes later. On the live mesh the nodes are granted `_DELIVER.<node>`
|
||||
and not `_DELIVER.<node>.>` — re-making would have silenced every machine in the mesh, which is worse
|
||||
than the collision it was fixing and far harder to undo. That was the first fix written here, and the
|
||||
permission is the reason it was not shipped.
|
||||
|
||||
**Not fatal**, which is what 146's change intended and did not do: the bare subject still delivers, and
|
||||
collides only where one holder has two consumers of one name. A node has one.
|
||||
|
||||
## How the fix is checked
|
||||
|
||||
Two tests against a real server: a consumer with a subscriber bound keeps its subject, is reported,
|
||||
and still delivers to that subscriber; a consumer with nothing bound moves, so 146's fix still applies
|
||||
where the collision actually was.
|
||||
|
||||
## What is left
|
||||
|
||||
The node consumers stay on the bare subject, which is correct and not tidy. Nothing is wrong while
|
||||
they do not move: one consumer per name per stream cannot collide with itself.
|
||||
|
||||
The controller reports each one it kept, and did, on the start that fixed this — four node consumers
|
||||
and the build machine's worker, which is bound the same way and was not anticipated here:
|
||||
|
||||
```
|
||||
consumer novox on NODES still delivers to "_DELIVER.novox" and not "_DELIVER.novox.NODES":
|
||||
nats: consumer name already in use. It keeps working; the subject moves on an assertion
|
||||
made while nothing is bound to it
|
||||
```
|
||||
|
||||
**The wider grant has since landed** (2026-09-30, measured on the mesh's own broker config): every
|
||||
node is now allowed `_DELIVER.<node>.>` as well as the bare subject. That was the thing missing when
|
||||
this was diagnosed, and it is why re-making the consumers then would have silenced every machine.
|
||||
What remains is only the second half — an assertion made while each node is detached from its
|
||||
consumer — and that is its own piece of work, not a side effect of a restart.
|
||||
Reference in New Issue
Block a user