Compare commits
36
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
83791f0921 | ||
|
|
bf4d4e2e7b | ||
|
|
ec42ee0846 | ||
|
|
5a3dee9e9e | ||
|
|
47909c6b71 | ||
|
|
7e5edab8da | ||
|
|
3649f82204 | ||
|
|
b400ce2c54 | ||
|
|
995c8cb266 | ||
|
|
f89aef992d | ||
|
|
b967ef7be3 | ||
|
|
e417906241 | ||
|
|
b1bf895688 | ||
|
|
0d64677c70 | ||
|
|
04205c1dc8 | ||
|
|
9dc49cd831 | ||
|
|
66b413a076 | ||
|
|
fcba05fed9 | ||
|
|
18f37c25b2 | ||
|
|
3c2b4fc6b6 | ||
|
|
72eaf52867 | ||
|
|
741625e725 | ||
|
|
c3730b9a23 | ||
|
|
3bd6f34de3 | ||
|
|
2b5119ecd2 | ||
|
|
96bdffa9bc | ||
|
|
4eb16f1028 | ||
|
|
d199de40db | ||
|
|
9a1dc4665c | ||
|
|
14be8576f8 | ||
|
|
ec8676c225 | ||
|
|
3c0f7082e6 | ||
|
|
0dd00e88b6 | ||
|
|
eef54917ec | ||
|
|
e9b1010bc0 | ||
|
|
f6ed3545b7 |
@@ -54,4 +54,6 @@ whose failure has never been observed is a guess about its own correctness.
|
||||
The development cycle, checked ([ADR 0080](../../02-DECISIONS/0080-the-development-cycle-is-checked.md)):
|
||||
a to-be design names a decision, an in-progress/implemented design names its owning code, a
|
||||
located/fixed issue names its owner, a fixed/resolved issue says what fixed it, a graduated
|
||||
research overview says what it became. `python3 00-META/checks/cycle.py`
|
||||
research overview says what it became, and no two issue records share a number (issue 155 — the
|
||||
number is how a record is cited, and `main` lags every open pull request, so two people reading it
|
||||
allocate the same one). `python3 00-META/checks/cycle.py`
|
||||
|
||||
+23
-1
@@ -14,7 +14,8 @@ What is enforced:
|
||||
its owning code (`code:`) -- no development without a design that says where.
|
||||
issues a known `status:`; once `located`, `located-in:` names the owner;
|
||||
once `resolved`, `fixed-by:` says what fixed it (prose counts --
|
||||
"nothing, the capability existed" is an answer).
|
||||
"nothing, the capability existed" is an answer). And no two records share a
|
||||
number -- the number is how a record is cited.
|
||||
research a known `status:`; a `graduated` overview says what it `became:`, and every
|
||||
target it names exists.
|
||||
decisions every accepted record is REACHABLE from the cycle: cited by a design doc's
|
||||
@@ -109,6 +110,27 @@ def main():
|
||||
"without a design that says where" % status)
|
||||
|
||||
# ---- issues ------------------------------------------------------------------------
|
||||
# Two records may not share a number. Numbers are taken as "next free after main", and work
|
||||
# sits on unmerged branches for days -- so two people reading the same main allocate the same
|
||||
# number, and nothing said so. It happened twice in one evening between two machines, and the
|
||||
# second collision landed on main with all three checks passing (issue 155). An issue number is
|
||||
# how every other record cites this one; two records answering to it means a pointer that
|
||||
# resolves to whichever the reader happened to open.
|
||||
seen = {}
|
||||
for folder in sorted(glob.glob(os.path.join(ROOT, "04-ISSUES", "*", ""))):
|
||||
name = os.path.basename(os.path.normpath(folder))
|
||||
number = name.split("-", 1)[0]
|
||||
if not number.isdigit():
|
||||
continue
|
||||
if number in seen:
|
||||
bad(os.path.join("04-ISSUES", name),
|
||||
"is numbered %s, and so is %s -- an issue number is how it is cited, and two "
|
||||
"records answering to one means a citation that resolves to whichever the reader "
|
||||
"opened. Take the next free number across main AND every open pull request"
|
||||
% (number, seen[number]))
|
||||
else:
|
||||
seen[number] = name
|
||||
|
||||
for path in sorted(glob.glob(os.path.join(ROOT, "04-ISSUES", "*", "00-report.md"))):
|
||||
front = frontmatter(path)
|
||||
if front is None:
|
||||
|
||||
@@ -21,7 +21,12 @@ incident someone must **clear**.
|
||||
|
||||
## Steps
|
||||
|
||||
1. Take the next free number. Create `04-ISSUES/NNN-short-name/00-report.md`:
|
||||
1. Take the next free number — **across `main` and every open pull request**, not `main` alone.
|
||||
Work sits on unmerged branches for days, so two people both reading `main` allocate the same
|
||||
number; it happened twice in one hour between two machines, and the second collision reached
|
||||
`main` with every check passing (issue 155). `cycle.py` now refuses two records sharing a number,
|
||||
which catches a collision but does not prevent one. Create
|
||||
`04-ISSUES/NNN-short-name/00-report.md`:
|
||||
|
||||
```yaml
|
||||
---
|
||||
@@ -42,6 +47,14 @@ incident someone must **clear**.
|
||||
## Rules
|
||||
|
||||
- Closed issues are never deleted — they are the mesh's symptom-to-component memory.
|
||||
- `fixed-by:` names something that will still exist: a commit or a pull request, never a branch. A
|
||||
branch is deleted when it merges, so a branch name there is a pointer that resolves to nothing by
|
||||
the time anybody follows it.
|
||||
- A fix that turns out to have broken something else is written back into the record that asked for
|
||||
it, pointing at the new issue. Somebody arriving at a record to learn why the code is the way it
|
||||
is must not have to already know there was a sequel.
|
||||
- Renumbering a collision happens once, in the branch that lands last. Renumbering a branch whose
|
||||
author is still pushing only moves the race.
|
||||
- An issue whose answer is a general lesson should also be written to the knowledge base, so
|
||||
the next person searching a symptom finds it. Both, not either.
|
||||
- `status: wontfix` is legitimate and requires a sentence saying why.
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
topic: building it
|
||||
status: proposed
|
||||
status: accepted
|
||||
date: 2026-09-01
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
@@ -101,3 +101,19 @@ the digest down after building.
|
||||
**Whether kind 4 deserves a module at all.** Thirty-five descriptions that say *install this and
|
||||
write these files* may be better as one module with settings than as thirty-five modules. Left
|
||||
open deliberately; it is a question about the shape of the catalogue, not about whether to have one.
|
||||
|
||||
## Accepted, 2026-09-29, against what was built
|
||||
|
||||
*Marked in a grooming pass: the mesh was built to this record and the record still said `proposed`.*
|
||||
|
||||
The proposal is the arrangement that exists. `mesh-catalog` holds descriptions of software we did
|
||||
not write and the programs that provision it, and holds neither the mesh's own components nor an
|
||||
application's own module. The mesh's list of modules is a table in the control plane, filled by
|
||||
`module add`, and every module records the source it came from with the commit it was read at.
|
||||
|
||||
**One half is not built: `module check` as a command on the control plane's binary.** A manifest is
|
||||
still validated by a test that reaches into the control plane's internals — which works for this
|
||||
catalogue and gives nothing at all to somebody describing their own application in their own
|
||||
repository, which this record says is the case that matters most. That is
|
||||
[issue 148](../04-ISSUES/148-a-manifest-outside-this-catalogue-has-no-check/00-report.md).
|
||||
|
||||
|
||||
@@ -26,6 +26,14 @@ runtime is per-module, not per-node, and treating the audit-logger as special le
|
||||
modules' code with nothing to run it: the conversion produced tools and events that, as it stands,
|
||||
never execute.
|
||||
|
||||
> **The hosting form is settled elsewhere — 2026-09-30.** Where this record says "a container", read
|
||||
> [ADR 0150](0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md): a module's own
|
||||
> code runs as supervised processes under this record's one account. Nothing else here changes — the
|
||||
> per-module runtime, the per-tool key and the single scoped account are the argument this record made
|
||||
> and they are why 0150 goes the way it does. The note is here because two design documents chose the
|
||||
> other form without knowing this record existed
|
||||
> ([issue 117](../04-ISSUES/117-a-modules-own-code-is-a-container-and-a-process/00-report.md)).
|
||||
|
||||
## Decision
|
||||
|
||||
### A module with tools or events runs a process of its own
|
||||
|
||||
@@ -1,14 +1,21 @@
|
||||
---
|
||||
topic: building it
|
||||
status: proposed
|
||||
status: superseded
|
||||
date: 2026-09-12
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 0016-the-lab.md
|
||||
superseded-by: 02-DECISIONS/0149-the-live-mesh-is-the-test-bed.md
|
||||
---
|
||||
|
||||
# 68. The lab takes requests, one at a time, and runs each from its own copy
|
||||
|
||||
> **Superseded — 2026-09-30, by [ADR 0149](0149-the-live-mesh-is-the-test-bed.md).** Never built. The
|
||||
> live mesh became the test bed, because the faults that cost the most are faults of a mesh that
|
||||
> already exists — bound consumers, containers made against an older roster, an adopted machine — and a
|
||||
> bed is by construction a mesh that does not. The rule worth keeping from below is that a run reads a
|
||||
> copy that is not anybody's working tree.
|
||||
|
||||
## Context
|
||||
|
||||
**The lab is exclusive hardware, and today a person holds it.** Raising a scenario takes over
|
||||
|
||||
@@ -35,6 +35,16 @@ The development cycle is enforced mechanically, to the extent frontmatter can ca
|
||||
capability existed" is an answer).
|
||||
- **No silent graduation** — a `graduated` research overview says what it `became:`, and the
|
||||
targets exist.
|
||||
- **No two records answering to one number** — added 2026-09-30; see the insight below.
|
||||
|
||||
> **Progressive insight — 2026-09-30.** The list above named four things `cycle.py` enforces, and
|
||||
> now names five. Nothing enforced that two issue records hold different numbers: two machines
|
||||
> filing issues within one hour both read `main`, both took "the next free number", and collided
|
||||
> twice — the second collision reaching `main` with `records.py`, `cycle.py` and `index.py` all
|
||||
> reporting success (04-ISSUES/155). A number is how every other record cites one, so two records
|
||||
> answering to it is a citation that resolves to whichever folder the reader opened. `cycle.py`
|
||||
> refuses it now. The decision here stands exactly as written: this is one more thing frontmatter
|
||||
> and file names can carry, found by its absence rather than by reasoning.
|
||||
|
||||
[`00-META/checks/cycle.py`](../00-META/checks/cycle.py) refuses violations, beside `records.py`
|
||||
and `index.py`; all three run before any merge here. What frontmatter cannot see — that code work
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
topic: what runs on it
|
||||
status: proposed
|
||||
status: accepted
|
||||
date: 2026-09-25
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
@@ -275,3 +275,11 @@ On acceptance, each of these is superseded or amended by this record, not edited
|
||||
everything a module needs is a requirement
|
||||
- [Issue 095](../04-ISSUES/095-a-module-assigned-after-genesis-has-no-broker-account/00-report.md),
|
||||
[issue 103](../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md): what fails today
|
||||
|
||||
## Accepted, 2026-09-29, against what was built
|
||||
|
||||
*Marked in a grooming pass.* The vault is a module providing `secret` at mesh scope, and six
|
||||
modules in the catalogue require it — so a shared secret is a requirement answered by the vault,
|
||||
which is what this record asks for. Private keys are still made where they are used and never
|
||||
travel, which is the other half and was never in question.
|
||||
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
topic: what runs on it
|
||||
status: proposed
|
||||
status: accepted
|
||||
date: 2026-09-26
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
@@ -189,6 +189,23 @@ On acceptance, each of these is amended by this record, not edited:
|
||||
credential is no longer all-or-nothing with a window. It overlaps, with each step confirmed.
|
||||
A single-party credential is staged, not replaced.
|
||||
|
||||
## Accepted, 2026-09-30, and not scheduled
|
||||
|
||||
Accepted as written. The separation it draws — a consumer's *resource* and a *credential that reaches
|
||||
it* are different things with different lifecycles — is the part that had to be settled, because the
|
||||
alternative is what the record was written against: retiring a credential taking the data it reached
|
||||
with it. That is a data-loss shape, and a record that names it should not sit unresolved while the
|
||||
code that could hit it is being written.
|
||||
|
||||
**It is not built, and accepting it does not schedule it.** The SDK's provisioner adapter is still
|
||||
`create` / `remove` / `holds` rather than the four operations above, and no provider implements the
|
||||
two-credential rotation. Accepted-and-not-built is an ordinary state here — 0141 and 0142 are both in
|
||||
it — and it is the honest one: leaving this `proposed` made it invisible to anyone reading what the
|
||||
mesh has decided, while changing nothing about what runs.
|
||||
|
||||
The work it implies belongs with the provisioner contract, beside
|
||||
[issue 124](../04-ISSUES/124-a-consumer-cannot-be-told-what-its-provider-derived/00-report.md).
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Every credential provider's adapter changes**, in two steps. The first separates *retire a
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
topic: what runs on it
|
||||
status: proposed
|
||||
status: accepted
|
||||
date: 2026-09-26
|
||||
deciders: jochen
|
||||
extends: 0112-a-module-definition-names-no-node-mesh-or-path.md
|
||||
@@ -47,3 +47,10 @@ other boundary already is: the module name.
|
||||
node runs one of each (ADR 0115)" — instead of failing on whichever name collides first.
|
||||
- Multi-tenant asks are answered in the catalogue (a second module definition), not in the
|
||||
control plane.
|
||||
|
||||
## Accepted, 2026-09-29, against what was built
|
||||
|
||||
*Marked in a grooming pass.* The rule is enforced where it cannot be forgotten: `assignment`'s
|
||||
primary key is `(node, module)`, so a second assignment of one module to one machine is not a thing
|
||||
the mesh can hold. The record read `proposed` while the schema had already settled it.
|
||||
|
||||
|
||||
@@ -12,7 +12,7 @@ extends: 0102-the-mesh-writes-into-a-shared-file-never-over-it.md
|
||||
## Context
|
||||
|
||||
When a resource stops being declared — its module unassigned, the node sent a
|
||||
deliberately-empty declaration ([issue 127](../04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)),
|
||||
deliberately-empty declaration ([issue 149](../04-ISSUES/149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)),
|
||||
or a new catalogue version renaming its id — the host undoes it. The host's own code states
|
||||
the rule it means to follow: **it removes what it made and leaves what it merely configured.**
|
||||
For almost every resource it does exactly that:
|
||||
|
||||
@@ -0,0 +1,159 @@
|
||||
---
|
||||
topic: the tiers
|
||||
status: accepted
|
||||
date: 2026-09-30
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 02-DECISIONS/0066-public-routing-is-name-agnostic.md
|
||||
---
|
||||
|
||||
# 148. The mesh's names are resolved, not copied into every container
|
||||
|
||||
## Context
|
||||
|
||||
The mesh gives every container it declares the whole roster of mesh names as entries written into
|
||||
the container's own hosts file at creation
|
||||
([design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md),
|
||||
[ADR 0066](0066-public-routing-is-name-agnostic.md)). A container takes those entries once and never
|
||||
looks again.
|
||||
|
||||
Three issues are the same fact arriving three times.
|
||||
|
||||
**A container keeps the address it was made with.** Adopting the predecessor's tunnel moved the hub's
|
||||
private address; the declaration followed it within one push and nothing on the machine did. The
|
||||
forge's container held the old address, lost its database, reported healthy while its existing
|
||||
connections lasted, and then the public name went down
|
||||
([issue 109](../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md)).
|
||||
|
||||
**The same fault, four days later, undetected for five days.** One container had restarted 2286 times
|
||||
against a database it could no longer find, while the mesh reported the machine as doing what it was
|
||||
told. Inside it, `novox.internal` was an address that had not existed for five days
|
||||
([issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md)). Forty-eight
|
||||
other containers were current, none of them corrected — each had been recreated for some other
|
||||
reason and picked up the roster on the way.
|
||||
|
||||
135 was fixed by putting the roster into the digest the host compares a container against, so a
|
||||
container whose names moved is recreated like one whose image moved. **That made the roster part of
|
||||
every container's identity**, which is the third arrival:
|
||||
|
||||
**One name moving replaces every container in the mesh.** Migrating one small module on one machine
|
||||
took four routine actions; each changed the roster, and each replaced every container on the control
|
||||
node — its own store, the registry, the edge proxy, the forge, the directory, mail. The control plane
|
||||
was unreachable twice while its own store came back through crash recovery
|
||||
([issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). None of
|
||||
the replaced containers had anything to do with the module being migrated, or with its machine.
|
||||
|
||||
The blast radius of a name is now every container that carries the list, which is all of them. The
|
||||
node-by-node migration ahead adds names one module at a time — on one machine alone that is around
|
||||
twenty-five — and each would be a full restart of every service on the hub.
|
||||
|
||||
## Considered Options
|
||||
|
||||
**1. Keep the roster in every container and accept the churn.** Rejected. It is not a cost that can
|
||||
be paid down: the mesh gets more names as it grows, and every name costs a restart of everything.
|
||||
A rollback costs another.
|
||||
|
||||
**2. Scope each container's entries to the names it actually binds.** A container is given the names
|
||||
of the things it declared a requirement on, so a name's blast radius is its consumers. Tidy, needs no
|
||||
new mechanism, and keeps 135's guarantee exactly.
|
||||
|
||||
Rejected, and this is the close one. It contradicts the standing intent that **anything on the mesh
|
||||
can call anything on it** — three cases, same machine, the private network, the public network, and
|
||||
no fourth. Scoping resolution to declared couplings makes a name reachable only where the mesh was
|
||||
told in advance that it would be wanted, and a person debugging inside a container would find names
|
||||
missing that exist everywhere else on the machine. It also leaves the roster in the digest, so the
|
||||
churn returns the moment a widely-bound name moves — smaller, not gone.
|
||||
|
||||
**3. Resolve at lookup time through the machine's resolver, and copy nothing.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**A container resolves the mesh's names through its machine's resolver, at the moment it asks. No
|
||||
mesh name and no mesh address is written into a container, and none is part of a container's
|
||||
identity.**
|
||||
|
||||
The three consequences that make this worth doing:
|
||||
|
||||
- **Staleness stops being possible**, rather than being detected. 109 and 135 are not bugs that were
|
||||
fixed; they are a shape that no longer exists. A name that moves is answered differently by the next
|
||||
lookup, in every container, with nothing recreated and nothing restarted.
|
||||
- **A name's blast radius becomes nothing.** Assigning a module on one machine does not touch a
|
||||
container on another.
|
||||
- **Anything can still call anything**, which option 2 gave up. The resolver answers every mesh name to
|
||||
every asker on the machine, exactly as it answers the machine itself.
|
||||
|
||||
**The resolver is a machine-level process, not a container** — one of the modules that is not a
|
||||
container at all — so a container depending on it is not the circularity it would be if the mesh's
|
||||
own store had to resolve a name through something the store's own runtime had to start first.
|
||||
|
||||
**What a module declares for itself is untouched.** Entries a manifest asks for are the module's own,
|
||||
stay in the container, and stay in its identity: they are part of what the module *is*, they do not
|
||||
move when the mesh's roster does, and the mesh does not know what they mean.
|
||||
|
||||
**The machine's own roster file is untouched.** It is a file, rewritten in place, read by processes and
|
||||
people; nothing restarts when it changes. It is only the *copy into each container* that this ends.
|
||||
|
||||
### The order this lands in, which is not a preference
|
||||
|
||||
**Nothing may stop copying names until resolution works from a container.** Removing the copy first
|
||||
reintroduces 109 and 135 — silently, and on a live mesh, which is exactly how both were found.
|
||||
|
||||
1. **A container on any network can reach the resolver.** Today a container on the runtime's default
|
||||
network asks from an address the converged filter drops, so it has no DNS at all
|
||||
([issue 110](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md));
|
||||
and on two of four machines the resolver binds loopback only, so the runtime hands containers a
|
||||
public resolver instead. Both are prerequisites, not related work.
|
||||
2. **The runtime is told which resolver to use, per machine, as a file** — not per container as a
|
||||
creation-time argument, or the resolver's address is back in every container's identity and the
|
||||
problem has only got smaller.
|
||||
3. **Then, and only then, the roster leaves the declaration and the digest.**
|
||||
|
||||
Until step 3 the mesh keeps copying, and keeps comparing. 151 stays open until step 3 lands; it is not
|
||||
closed by this record, only answered by it.
|
||||
|
||||
## How this is checked
|
||||
|
||||
- **A container resolves a name that moved, without being recreated.** Move a name the mesh serves;
|
||||
from a container that was running before the move and has not been touched since, the name answers
|
||||
with the new address. This is the one 109 and 135 would both have failed.
|
||||
- **A name's blast radius is nothing.** Add a routed name on one machine; no container on any other
|
||||
machine is recreated. The apply report on each machine says nothing changed. This is 151.
|
||||
- **Anything calls anything.** From a container on any machine, every `<node>.internal` name and every
|
||||
routed name the mesh serves resolves — including names the module never declared a requirement on,
|
||||
which is the guarantee option 2 would have given up.
|
||||
- **On every network the runtime offers.** The first three hold for a container on the runtime's
|
||||
default network as well as one on a declared network, because the default network is the case that
|
||||
has no DNS today.
|
||||
- **No mesh name is in a container's spec.** A test asserts the digest a host computes for a container
|
||||
does not move when the mesh's roster does, and does move when the module's own declared entries do.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **The resolver becomes load-bearing for every container**, where before it was load-bearing for the
|
||||
machine. This is a real cost and is accepted: a resolver that is down is a machine that cannot
|
||||
resolve, which is already true of the machine itself, and is a smaller event than a roster change
|
||||
destroying and recreating every container on the machine.
|
||||
- **Design 08's "a file rather than a resolver" no longer describes containers.** It was written when
|
||||
the mesh had no resolver and it gave the right answer then. The reasoning it rested on — every Linux
|
||||
has a hosts file, no package needed — was already overtaken by names a hosts file cannot express:
|
||||
service names and wildcards under `<node>.internal`, which is why the resolver was built.
|
||||
- **ADR 0066's mesh-wide propagation is kept and its mechanism changes.** A routed name still reaches
|
||||
every asker in the mesh; it reaches them through the resolver rather than by being written into each
|
||||
container. The consequence 0066 records — that an internal issuer's challenge needs the routed name
|
||||
resolvable inside the mesh — holds unchanged and by the same means the machine already uses.
|
||||
- **Issue 110 stops being a container-DNS inconvenience and becomes a prerequisite** for the mesh not
|
||||
restarting itself whenever it learns a name.
|
||||
- **A container started by hand gets the mesh's names too**, where before only declared containers did.
|
||||
Design 08 drew that boundary deliberately, on the grounds that reaching into every container is what
|
||||
a nameserver would be for. This record accepts that consequence rather than working around it: a
|
||||
person debugging in a hand-started container resolving the same names as everything else is the
|
||||
behaviour worth having, and it is what "anything can call anything" means.
|
||||
|
||||
## References
|
||||
|
||||
- [issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md) — one name replaces every container; the question this answers
|
||||
- [issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md) — the roster put into the digest
|
||||
- [issue 109](../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md) — the first arrival
|
||||
- [issue 110](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md) — the prerequisite
|
||||
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — routed names propagate mesh-wide; extended here
|
||||
- [design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md) — the file-not-resolver reasoning this narrows
|
||||
@@ -0,0 +1,106 @@
|
||||
---
|
||||
topic: building it
|
||||
status: accepted
|
||||
date: 2026-09-30
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
supersedes: 02-DECISIONS/0068-the-lab-takes-requests.md
|
||||
---
|
||||
|
||||
# 149. The live mesh is the test bed
|
||||
|
||||
## Context
|
||||
|
||||
[ADR 0068](0068-the-lab-takes-requests.md) proposed that the lab accept queued requests
|
||||
— a bed and a commit — answer them one at a time from a copy it owns, and expose that through tools so
|
||||
an agent could start a run and come back to it. It has been `proposed` since 2026-09-12 and nothing was
|
||||
built.
|
||||
|
||||
What happened instead is that the mesh became the thing under test. It runs on four machines; every
|
||||
fault worth finding in the last month was found on them, and none was found in a bed:
|
||||
|
||||
- a container holding an address that had not existed for five days, on the control node
|
||||
([issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md));
|
||||
- a machine reading healthy for eleven hours while no module could reach another
|
||||
([issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md));
|
||||
- one name replacing every container on the hub
|
||||
([issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md));
|
||||
- a consumer assertion that is correct on a mesh being raised and fatal on one that is running
|
||||
([issue 156](../04-ISSUES/156-moving-a-consumers-delivery-subject-stops-the-control-plane/00-report.md)).
|
||||
|
||||
The last is the one that settles it. That change was exercised on the raise path — which is what a bed
|
||||
*is* — and the raise path is the only path on which the fault cannot appear. A bed raises a mesh; it
|
||||
does not have a mesh that has been running for weeks, with consumers already bound, containers created
|
||||
against an older roster, and an adopted machine carrying a predecessor's configuration. **The faults
|
||||
that cost the most were all faults of a mesh that already exists**, and a bed is by construction a mesh
|
||||
that does not.
|
||||
|
||||
Lab runs are also expensive in a way that changed the behaviour around them: each costs a build and
|
||||
several minutes, so they were batched, and a batched test is one whose result arrives after the next
|
||||
three changes were already written.
|
||||
|
||||
## Considered Options
|
||||
|
||||
**1. Build 0068 as proposed.** Rejected. It answers a question nobody is asking: the bottleneck was
|
||||
never that a person had to sit at the lab, it was that a bed cannot hold the state the faults live in.
|
||||
Queueing and tooling a mechanism that finds the wrong class of fault faster is not an improvement.
|
||||
|
||||
**2. Leave 0068 `proposed`.** Rejected, and it is why this record exists rather than nothing. A record
|
||||
that contradicts current practice and sits unresolved is worse than either answer: it reads as intent
|
||||
to anyone who finds it, and the practice it contradicts is written down nowhere but a handoff note.
|
||||
|
||||
**3. Record that the live mesh is the test bed, and supersede 0068.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**A change is verified against the mesh that is running.** Not because a bed would be unwelcome, but
|
||||
because the state that breaks things is state a bed does not have: containers made against an older
|
||||
roster, consumers already bound, an adopted machine, a store with weeks of history.
|
||||
|
||||
**A change that can only be exercised on the raise path is not verified.** If the only test available
|
||||
raises a fresh mesh, the record says so, and says which case was therefore not covered. The words
|
||||
"exercised on a fresh mesh" are a statement about coverage, not a pass.
|
||||
|
||||
**The lab is not retired**, and [ADR 0016](0016-the-lab.md) stands. It remains the place to raise a
|
||||
mesh from bare, which is the one thing the live mesh cannot be asked to do and the one thing a bed does
|
||||
better than anything else. What this record removes is the lab as the *default* answer to "is this
|
||||
change good", and with it 0068's queue, tools and request protocol.
|
||||
|
||||
**Accuracy over a green run.** An honest failure on the live mesh beats a pass in a bed that could not
|
||||
have failed — and a change that is risky on the running mesh is a reason to make the change smaller,
|
||||
not a reason to test it somewhere it cannot break.
|
||||
|
||||
**What 0068 got right is kept as a rule, not a mechanism:** a run reads a copy that is not anybody's
|
||||
working tree. Every run of the lab that mattered was pinned to a checkout rather than a worktree, and
|
||||
the ones that were not produced results about code nobody had written down.
|
||||
|
||||
## How this is checked
|
||||
|
||||
- **A record that says a change was verified says on what.** Where it was a fresh mesh, it says which
|
||||
case is uncovered. This is the clause that would have caught 156: its change was verified, honestly,
|
||||
on the only path where it works.
|
||||
- **The lab is not in the path of a merge.** No check, playbook or handoff requires a bed to have run.
|
||||
- **0068 is unreachable as intent.** Its status is `superseded` and it names this record, so a reader
|
||||
arriving at the queue design finds out immediately that it was not built and why.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **A fault can be introduced on the machines that serve.** This is the cost, it is real, and it was
|
||||
paid twice in one evening — a control plane crash-looping for half an hour, and every container on the
|
||||
hub recreated five times. Both were found in minutes because they were live, and both would have
|
||||
passed a bed.
|
||||
- **There is no pre-merge gate beyond the repositories' own suites.** `make check` and the three hq
|
||||
checks are what stands between a change and the machines, which raises what those suites are worth
|
||||
and makes a test that cannot fail a genuine defect rather than an untidiness.
|
||||
- **Raising a mesh from bare is now the lab's whole job**, and is exercised deliberately rather than
|
||||
as a side effect of testing something else. The foundation work
|
||||
([issue 146](../04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md))
|
||||
is that job, and it is also the proof that the mesh can make another of itself.
|
||||
- **An agent cannot hand a run to a queue and come back**, which 0068 would have given. In practice it
|
||||
watches a push and reads the machines, which is what happened anyway.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR 0068](0068-the-lab-takes-requests.md) — superseded by this
|
||||
- [ADR 0016](0016-the-lab.md) — the lab, which stands
|
||||
- [issue 156](../04-ISSUES/156-moving-a-consumers-delivery-subject-stops-the-control-plane/00-report.md) — correct on the raise path, fatal on a running mesh
|
||||
+113
@@ -0,0 +1,113 @@
|
||||
---
|
||||
topic: what runs on it
|
||||
status: accepted
|
||||
date: 2026-09-30
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 02-DECISIONS/0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md
|
||||
---
|
||||
|
||||
# 150. A module's own code runs as supervised processes under the module's one account
|
||||
|
||||
## Context
|
||||
|
||||
The repository answers "what runs a module's own code" two ways and reconciles them nowhere
|
||||
([issue 117](../04-ISSUES/117-a-modules-own-code-is-a-container-and-a-process/00-report.md)).
|
||||
|
||||
[ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) is accepted and
|
||||
says **a container** — "the tool runtime carrying that module's compiled code" — and "one module, one
|
||||
process, one account". Two `proposed` design documents say a **`process`** resource running an argv,
|
||||
supervised by the machine, and one of them declares *four* of them for a single module and presents
|
||||
four as the point. Neither design document names 0047 in its `decisions:`, and the string `process`
|
||||
as a resource type appears in no decision record at all. The thing as built is the container.
|
||||
|
||||
Two things have happened since 0047 was written that bear on it directly.
|
||||
|
||||
[**ADR 0142**](0142-the-mesh-delivers-its-own-components-as-binaries.md) decided that the mesh's own
|
||||
components are binaries on the machine rather than container images, and
|
||||
[issue 114](../04-ISSUES/114-should-the-controller-be-a-container-or-a-process/00-report.md) was closed
|
||||
by it. That settled the mesh's components and deliberately said nothing about a module's.
|
||||
|
||||
And the standing definition of a module hardened: **a module is software that delivers one or more
|
||||
services, and a module is not a container.** It may deliver them as a container, an installed package
|
||||
with a unit, a binary, or configuration files; 61 of 73 happen to use a container and 11 do not,
|
||||
including the resolver, sshd and fail2ban. A rule that a module's *own code* must be a container makes
|
||||
the one kind of module the mesh writes itself the only kind that has no choice.
|
||||
|
||||
## Considered Options
|
||||
|
||||
**1. Hold 0047 as written: a container.** Rejected. Its own reasoning does not require one. What 0047
|
||||
argued for was a runtime **per module** rather than one for the whole node, because a node-wide runtime
|
||||
could not hold a per-module broker account and per-module runtimes competing on one tool key would each
|
||||
be handed calls for tools they do not have. A supervised unit per module satisfies that argument
|
||||
exactly — it is per module, and a unit runs as an account. The container was the mechanism to hand, not
|
||||
the conclusion.
|
||||
|
||||
**2. Let each design document choose.** Rejected; that is the present state and it is what issue 117
|
||||
reports. A module author reading the guide writes four processes; a module author reading the record
|
||||
writes a container; nothing tells either that the other exists.
|
||||
|
||||
**3. Settle the hosting form as a supervised process, and settle the count separately.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**A module's own code runs as one or more supervised processes on the machine, under the module's single
|
||||
account.** Where [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) says
|
||||
"a container, the tool runtime carrying that module's compiled code", read this record. Everything else
|
||||
0047 decided stands untouched: a tool is served on its own key, only the module that serves it answers,
|
||||
and the module's account is scoped to exactly its tool keys.
|
||||
|
||||
**The invariant is the account, not the process count.** 0047's "one module, one process, one account"
|
||||
carried its weight in the last clause. Its stated worry about a second process was "not a second one to
|
||||
scope and seal" — a second *identity* to grant, seal a secret to, and scope on the bus. Several
|
||||
processes sharing the module's one account create no second identity, so nothing further is scoped or
|
||||
sealed, and a module may therefore declare as many as its work has shapes: events, tools, a
|
||||
provisioner, a scheduled ingest. **What a module may not have is two accounts.**
|
||||
|
||||
**A module that delivers its service as a container still does.** This record is about the code the
|
||||
module itself carries — its tools, its events, its provisioner — and not about the software it delivers.
|
||||
A module wrapping a third-party image wraps a third-party image.
|
||||
|
||||
**Why supervised by the machine rather than by the mesh:** it is the same answer ADR 0142 gave for the
|
||||
mesh's own components, for the same reason. A unit the machine restarts needs no image, no registry
|
||||
pull and no runtime to be up before the mesh's own code can run — which matters most for exactly the
|
||||
modules whose code the mesh cannot start any other way.
|
||||
|
||||
## How this is checked
|
||||
|
||||
- **No design document describes a hosting form for a module's own code without citing this record.**
|
||||
Designs 18 and 20 name it in `decisions:`; this is the gap issue 117's third point reports, and
|
||||
`cycle.py` already enforces that a to-be design names its decisions.
|
||||
- **A module declaring several processes resolves to one account.** A test composes a module with more
|
||||
than one process resource and asserts the mesh mints exactly one broker account for it, scoped to that
|
||||
module's tool keys and nothing else — which is 0047's invariant stated as an assertion rather than a
|
||||
sentence.
|
||||
- **A module's own code does not require the container runtime.** A machine with no container runtime
|
||||
can still run a module whose code is its own, which is the claim that separates this from option 1 and
|
||||
is checkable on a machine that has one by asserting the declaration names no image for it.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **The sidecar port stops being needed.** [ADR 0029](0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md)
|
||||
records that "anything that is a service plus a sidecar currently has to publish a port to talk to
|
||||
itself", and the host's `network` shape exists partly for it. A process beside the service on the same
|
||||
machine reaches it without publishing anything, so that pressure goes.
|
||||
- **Something must supervise, and it is the machine.** This adds a unit per module's code to what the
|
||||
host writes and owns. The mesh already writes and owns units — `nftables` proves a module can write one
|
||||
and run it — so the mechanism exists; the count grows.
|
||||
- **A module's code is delivered, not pulled**, which puts it behind the same gap as the host's own
|
||||
delivery ([ADR 0141](0141-the-host-delivers-its-own-successor.md), not built): nothing yet delivers a
|
||||
version of a module's binary to a machine. A container's code arrives by `docker pull`, and this does
|
||||
not. **This is the cost of the decision and it is not paid**; until delivery exists, a module whose code
|
||||
is its own is a module somebody places by hand.
|
||||
- **Issue 117 is answered and its three disagreements close differently:** container-or-unit is decided
|
||||
here; one-process-or-several is decided here as several under one account; and whether the record was
|
||||
consulted is fixed by designs 18 and 20 naming this one.
|
||||
|
||||
## References
|
||||
|
||||
- [issue 117](../04-ISSUES/117-a-modules-own-code-is-a-container-and-a-process/00-report.md) — the contradiction this answers
|
||||
- [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) — extended; its "a container" clause is settled here
|
||||
- [ADR 0142](0142-the-mesh-delivers-its-own-components-as-binaries.md) — the same answer for the mesh's own components
|
||||
- [ADR 0141](0141-the-host-delivers-its-own-successor.md) — the delivery this depends on and which is not built
|
||||
- [ADR 0029](0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md) — the sidecar port this relieves
|
||||
@@ -174,6 +174,7 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0108** — [A route carries the policy applied to a request, and names a secret rather than holding one](0108-a-route-carries-the-policy-applied-to-a-request.md)
|
||||
- **0109** — [A package registry seat is one per ecosystem, not one for all of them](0109-a-package-registry-seat-is-one-per-ecosystem.md)
|
||||
- **0126** — [A module declares its own seats; the mesh reserves its own](0126-a-module-declares-its-own-seats.md)
|
||||
- **0148** — [The mesh's names are resolved, not copied into every container](0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
|
||||
|
||||
### What runs on them, and how it gets there
|
||||
|
||||
@@ -207,9 +208,9 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0099** — [A step that runs once names what it reads, and runs again when it changed](0099-a-step-that-runs-once-names-what-it-reads.md)
|
||||
- **0110** — [A seat is held by one assignment, from a closed set, and it may deliver a provision](0110-a-seat-is-a-module-assignment-from-a-closed-set.md)
|
||||
- **0112** — [A module definition names no node, no mesh and no path: everything it needs is a requirement the mesh resolves](0112-a-module-definition-names-no-node-mesh-or-path.md)
|
||||
- **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md) *(proposed)*
|
||||
- **0114** — [A credential two parties hold rotates over two credentials; one a single party holds rotates in place, staged; and retiring a credential never removes what it reached](0114-a-shared-credential-rotates-over-two-credentials.md) *(proposed)*
|
||||
- **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md) *(proposed)*
|
||||
- **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md)
|
||||
- **0114** — [A credential two parties hold rotates over two credentials; one a single party holds rotates in place, staged; and retiring a credential never removes what it reached](0114-a-shared-credential-rotates-over-two-credentials.md)
|
||||
- **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md)
|
||||
- **0117** — [A machine's uplink is a seat: the mesh configures the manager, never the link](0117-a-machines-uplink-is-a-seat.md)
|
||||
- **0118** — [Undeclaring removes what the mesh made, and gives a unit back the state it was found in](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md)
|
||||
- **0120** — [A roster fact carries its format as a template: the mesh owns the data, the module owns the format](0120-a-roster-fact-carries-its-format-as-a-template.md)
|
||||
@@ -228,6 +229,7 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0145** — [A module checks what the mesh claims is reachable, and it checks itself](0145-a-module-checks-what-the-mesh-claims-is-reachable.md) *(superseded)*
|
||||
- **0146** — [Connectivity is checked by name, per hosting form, with a valid certificate](0146-connectivity-is-checked-by-name-per-hosting-form.md)
|
||||
- **0147** — [A module anchors the mesh's authority on a machine, and takes it away again](0147-a-module-anchors-the-meshs-authority.md)
|
||||
- **0150** — [A module's own code runs as supervised processes under the module's one account](0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md)
|
||||
|
||||
### How it is built
|
||||
|
||||
@@ -237,9 +239,9 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0014** — [No workspace — each module is a standalone package consuming published dependencies](0014-no-npm-workspace.md)
|
||||
- **0015** — [Applications live in their own repository; the monorepo is for the mesh](0015-applications-live-in-their-own-repository.md)
|
||||
- **0016** — [The lab](0016-the-lab.md)
|
||||
- **0037** — [Where a module lives](0037-where-a-module-lives.md) *(proposed)*
|
||||
- **0037** — [Where a module lives](0037-where-a-module-lives.md)
|
||||
- **0039** — [What the SDK holds, and what it refuses](0039-what-the-sdk-holds-and-refuses.md)
|
||||
- **0068** — [The lab takes requests, one at a time, and runs each from its own copy](0068-the-lab-takes-requests.md) *(proposed)*
|
||||
- **0068** — [The lab takes requests, one at a time, and runs each from its own copy](0068-the-lab-takes-requests.md) *(superseded)*
|
||||
- **0069** — [A module is a repository and a path within it](0069-a-module-is-a-repository-and-a-path.md)
|
||||
- **0076** — [The SDK is a published package, and the toolchain resolves it by version](0076-the-sdk-is-a-published-package.md)
|
||||
- **0082** — [The registry is reached by name, and the overlay is its security](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)
|
||||
@@ -248,6 +250,7 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0097** — [A vendor image is a declared build input, and a recipe fetches nothing undeclared](0097-a-vendor-image-is-a-declared-build-input.md)
|
||||
- **0107** — [Persistent data is a directory bind, never a named volume](0107-persistent-data-is-a-directory-bind-never-a-named-volume.md)
|
||||
- **0111** — [A build source is on the mesh's git seat, or it is an external repository](0111-a-build-source-is-on-the-git-seat-or-external.md)
|
||||
- **0149** — [The live mesh is the test bed](0149-the-live-mesh-is-the-test-bed.md)
|
||||
|
||||
### How it is checked
|
||||
|
||||
|
||||
@@ -7,8 +7,9 @@ code:
|
||||
- mesh-controller internal/identity/authority.go
|
||||
- mesh-host internal/identity/serving.go
|
||||
- mesh-host internal/apply (the service that reflects a rule set)
|
||||
updated: 2026-09-29
|
||||
updated: 2026-09-30
|
||||
decisions:
|
||||
- 02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md
|
||||
- 02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md
|
||||
- 02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md
|
||||
- 02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md
|
||||
@@ -289,6 +290,30 @@ hosts file by the runtime. That extends the file decision rather than overturnin
|
||||
mesh and not chosen by a module: a module that listed the machines would go stale the day one
|
||||
joins, and a module that did not would be one whose containers cannot reach anything by name.
|
||||
|
||||
**Superseded for containers by [ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
|
||||
(2026-09-30).** Everything above still describes the machine's own roster file, which is how it works
|
||||
and how it will keep working. It no longer describes containers.
|
||||
|
||||
Copying the roster into each container made the roster part of each container's identity, so one name
|
||||
moving replaced every container in the mesh — a module assigned on one machine restarted the store, the
|
||||
registry, the edge and mail on another
|
||||
([issue 151](../../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). And it
|
||||
did not stop the staleness it was meant to: a copy taken at creation is stale the moment the roster
|
||||
moves, twice found as a container holding an address that had not existed for days
|
||||
([issues 109](../../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md)
|
||||
and [135](../../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md)).
|
||||
|
||||
**A container resolves the mesh's names through its machine's resolver, at the moment it asks, and
|
||||
nothing is copied.** The resolver is a machine-level process rather than a container, so nothing
|
||||
circular is being asked for. This is gated on a container being able to reach the resolver from any of
|
||||
the runtime's networks, which it cannot today
|
||||
([issue 110](../../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)) —
|
||||
until that lands the mesh keeps copying and keeps comparing, and the order is stated in the record.
|
||||
|
||||
The paragraph below states the old boundary, and 0148 deliberately gives it up: a container somebody
|
||||
started by hand resolves the same names as everything else, because the resolver answers the machine,
|
||||
not a list of containers.
|
||||
|
||||
**The boundary, which is deliberate and worth stating:** *declared* containers. A container
|
||||
somebody starts by hand is not the mesh's to configure, and reaching into every container on a
|
||||
machine — declared or not — is what a nameserver in `resolv.conf` would be for.
|
||||
|
||||
@@ -7,6 +7,7 @@ code:
|
||||
- mesh-catalog modules/builder
|
||||
updated: 2026-09-29
|
||||
decisions:
|
||||
- 02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md
|
||||
- 02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md
|
||||
- 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md
|
||||
- 02-DECISIONS/0097-a-vendor-image-is-a-declared-build-input.md
|
||||
|
||||
@@ -7,6 +7,7 @@ code:
|
||||
- mesh-sdk src
|
||||
updated: 2026-09-21
|
||||
decisions:
|
||||
- 02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md
|
||||
- 02-DECISIONS/0099-a-step-that-runs-once-names-what-it-reads.md
|
||||
- 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md
|
||||
- 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md
|
||||
|
||||
@@ -155,3 +155,17 @@ design document here, and get it back. That check fails today by design.
|
||||
|
||||
**What stands until then** is the signpost, and the honest description of it: reachable, not
|
||||
surfacing.
|
||||
|
||||
## Where this stands, 2026-09-29
|
||||
|
||||
*Added in a grooming pass.* The knowledge base this record is about is the **predecessor's**, and it
|
||||
is no longer reachable from anything: the surface that answered `recall_search` speaks the transport
|
||||
the mesh removed at the cut-over
|
||||
([issue 147](../147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md)).
|
||||
|
||||
So the sentence in `README.md` that this record catches — *these documents are still indexed into
|
||||
the knowledge base* — is now wrong twice over: nothing indexed them, and there is nothing to index
|
||||
them into. The record stays open, and its answer is no longer "index this repository somewhere"; it
|
||||
is whatever the mesh grows as its own knowledge surface, if it grows one. Until then the honest fix
|
||||
is the README, which should stop claiming a property nothing provides.
|
||||
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: open
|
||||
status: resolved
|
||||
opened: 2026-09-22
|
||||
located-in: []
|
||||
fixed-by:
|
||||
located-in: [mesh-catalog modules/gitea, mesh-controller internal/catalogue/declaration.go]
|
||||
fixed-by: mesh-controller 7352c84, merged in #46 — a module is told its port in a container's environment too, as a file already was
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -39,3 +39,9 @@ the assignment happens to differ.
|
||||
- Should composition refuse an environment value that names a port the module does not fix, the
|
||||
way it refuses other claims a module cannot make?
|
||||
- Which other modules write their own address, with a port, into their environment?
|
||||
|
||||
## Closed
|
||||
|
||||
*2026-09-29, in a grooming pass rather than by whoever fixed it.* The forge's address follows a moved port the same way every other reader does. Found by
|
||||
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
|
||||
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: open
|
||||
status: resolved
|
||||
opened: 2026-09-22
|
||||
located-in: []
|
||||
fixed-by:
|
||||
located-in: [mesh-controller internal/catalogue/declaration.go, mesh-catalog (every routed module)]
|
||||
fixed-by: mesh-controller bdf965d (a route names the endpoint it serves) with `portOfEndpoint` and `AtPublishedPort` — the contribution carries the endpoint's declared port and the machine-side redirection is applied to it like any other
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -45,3 +45,9 @@ precisely because the predecessor holds the usual one.
|
||||
keeping the mapping out of rendered configuration?
|
||||
- What should refuse a declaration whose contributed route names a port nothing on that node
|
||||
listens on?
|
||||
|
||||
## Closed
|
||||
|
||||
*2026-09-29, in a grooming pass rather than by whoever fixed it.* A route names an endpoint rather than a port, and the redirection that turns a declared port into the published one is applied to contributions too. Found by
|
||||
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
|
||||
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
|
||||
|
||||
+16
@@ -51,3 +51,19 @@ knows that is what the rule means.
|
||||
the runtime's default one? That is a stronger rule and would have prevented 109 as well.
|
||||
- What checks it? A converged bed with a container on the default network resolving a mesh name is
|
||||
the missing assertion; nothing in the resolver's own beds covers the filter.
|
||||
|
||||
## What now depends on this (2026-09-30)
|
||||
|
||||
This stopped being a container-DNS inconvenience.
|
||||
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) decides
|
||||
that a container resolves the mesh's names rather than being given a copy of them, which is what stops
|
||||
one name moving from replacing every container in the mesh
|
||||
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)) and what makes a
|
||||
stale address impossible rather than merely noticed
|
||||
([issues 109](../109-a-container-keeps-the-address-it-was-made-with/00-report.md)
|
||||
and [135](../135-a-containers-mesh-names-are-not-compared/00-report.md)).
|
||||
|
||||
**That decision cannot land until this one does**, and not partly: a container on the runtime's default
|
||||
network is the case with no DNS at all, and it is the case the mesh's own forge runs in. Two of four
|
||||
machines also bind the resolver to loopback only, so the runtime hands their containers a public
|
||||
resolver. Both halves are this issue.
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: open
|
||||
status: resolved
|
||||
opened: 2026-09-24
|
||||
located-in: [mesh-controller module.json, mesh-host internal/apply]
|
||||
fixed-by:
|
||||
fixed-by: ADR 0142 — the mesh's own components are delivered as binaries on the machine, so the controller is a process; the delivery itself is issue 142
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -84,3 +84,21 @@ restart and run-to-completion semantics — so this would not need host-side wor
|
||||
|
||||
The two do not collapse into one. The controller is not a code-carrying sidecar, and `network: host`
|
||||
is what makes the asymmetry visible here and nowhere else.
|
||||
|
||||
## Answered
|
||||
|
||||
*2026-09-29, in a grooming pass.* This asked a question rather than reporting a defect, and the
|
||||
question was taken: [ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md)
|
||||
decides that the mesh's own components — the host, the controller, the catalogue, the builder, the
|
||||
vault — are **binaries on the machine**, delivered by the mechanism
|
||||
[ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md) built, and that
|
||||
third-party software (the store, the registry, the broker) stays a container because an image is the
|
||||
right way to carry somebody else's build.
|
||||
|
||||
So the operating experience this record was written from — every mutating command reached through
|
||||
`docker exec mesh-controller` — is answered, and answered against the container.
|
||||
|
||||
**The delivery is a separate matter and is not this record's.** Step 1 of it is built and no
|
||||
component travels yet; that is
|
||||
[issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md).
|
||||
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: located
|
||||
status: resolved
|
||||
opened: 2026-09-25
|
||||
located-in: [hq, mesh-catalog modules/showcase, mesh-sdk src/tools/index.ts, mesh-tools]
|
||||
fixed-by:
|
||||
fixed-by: hq ADR 0150 — a module's own code runs as supervised processes under the module's one account; designs 18 and 20 now cite it, and ADR 0047 carries a dated note pointing at it
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -99,3 +99,27 @@ what the sidecar is has nowhere in the design layer to look, which is how this w
|
||||
is mechanically checkable: the resource types a design doc names are a closed set, and every
|
||||
member of it either appears in a decision or does not. Whether that check is worth writing is
|
||||
part of this issue, not settled by it.
|
||||
|
||||
## Answered (2026-09-30)
|
||||
|
||||
[ADR 0150](../../02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md)
|
||||
settles all three disagreements, and the design documents win two of them:
|
||||
|
||||
1. **Container or unit — a supervised process.** ADR 0047's argument never required a container. It
|
||||
argued for a runtime *per module*, because a node-wide one could not hold a per-module account and
|
||||
per-module runtimes on one tool key would be handed calls for tools they do not have. A unit per
|
||||
module satisfies that exactly, and a unit runs as an account. The container was the mechanism to
|
||||
hand. [ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md) had
|
||||
already gone the same way for the mesh's own components, and a module is not a container.
|
||||
2. **One process or several — several, under one account.** 0047's "one module, one process, one
|
||||
account" carried its weight in the last clause; its stated worry was "not a second one to scope and
|
||||
seal", which is about a second *identity*. Processes sharing the module's one account create none.
|
||||
What a module may not have is two accounts.
|
||||
3. **Whether the record was consulted — fixed rather than answered.** Designs 18 and 20 now name 0150
|
||||
in `decisions:`, and 0047 carries a dated note saying where its hosting form was settled, so neither
|
||||
door leads to the wrong answer any more.
|
||||
|
||||
**What this does not fix.** A module's code delivered as a binary is behind the same gap as the host's
|
||||
own ([ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md), accepted and not
|
||||
built): a container's code arrives by `docker pull` and this does not, so until delivery exists such a
|
||||
module is one somebody places by hand. 0150 records that as the cost of the decision, unpaid.
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: located
|
||||
status: resolved
|
||||
opened: 2026-09-26
|
||||
located-in: [mesh-sdk src/provisioner, mesh-catalog modules/redis]
|
||||
fixed-by:
|
||||
fixed-by: mesh-catalog bbda88c, merged in #84 — every credential provider says whether it still holds a consumer
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -60,3 +60,9 @@ checks it after the first pass.
|
||||
instance and leaves the gap for the others.
|
||||
- Where does the record of what was applied live, if not in memory? ADR 0114, still
|
||||
proposed, puts rotation state with the vault. The same place may answer this.
|
||||
|
||||
## Closed
|
||||
|
||||
*2026-09-29, in a grooming pass rather than by whoever fixed it.* A provisioner asks the backend what is there rather than trusting what it remembers doing. Found by
|
||||
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
|
||||
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
|
||||
|
||||
@@ -1,5 +1,6 @@
|
||||
---
|
||||
status: located
|
||||
status: resolved
|
||||
fixed-by: mesh-host 1cb8953 and fdc768c — the mesh writes into a marked block of a text file instead of over it
|
||||
opened: 2026-09-26
|
||||
located-in: [mesh-controller internal/catalogue/facts.go, mesh-host internal/apply]
|
||||
---
|
||||
@@ -71,3 +72,9 @@ private network loses that name too.
|
||||
- The host's file resource supports `into: "json"` only; anything else is a whole write.
|
||||
- `node show <node>` on the adopted workstation: `holds file /etc/hosts
|
||||
mesh-wireguard.fact-node-names`, original kept.
|
||||
|
||||
## Closed
|
||||
|
||||
*2026-09-29, in a grooming pass rather than by whoever fixed it.* A shared hosts file keeps every line that is not the mesh's. Found by
|
||||
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
|
||||
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
|
||||
|
||||
@@ -30,3 +30,16 @@ network, installs it as a trust anchor, refreshes the machine's bundles, and —
|
||||
unassigned stops its unit, and stopping the unit is what undoes it — takes both away again.
|
||||
|
||||
The owner is therefore `mesh-catalog`, module `ca-trust`, and nothing in the control plane.
|
||||
|
||||
## The module exists, and this stays open until a machine holds it
|
||||
|
||||
*2026-09-29.* `ca-trust` is in the catalogue and merged
|
||||
([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md)), and what it renders
|
||||
is checked in the control plane's own suite: the script fetches from the authority it was bound to,
|
||||
and the unit runs it both ways.
|
||||
|
||||
**No machine has been assigned it, and nothing has verified a name because of it.** The bed written
|
||||
for that cannot run ([issue 146](../146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md)),
|
||||
and the live mesh has not been given the module. So the symptom this record opened on — every
|
||||
internal name failing verification on every machine — is still true everywhere, and the record stays
|
||||
`located` until it is not. Closing it on a module that exists would be closing it on an intention.
|
||||
|
||||
@@ -1,5 +1,6 @@
|
||||
---
|
||||
status: located
|
||||
status: resolved
|
||||
fixed-by: mesh-host 3112c88 — undeclaring gives a unit back the state it was found in, and removes only a process the mesh made
|
||||
opened: 2026-09-27
|
||||
located-in: [mesh-host internal/apply/apply.go (remove)]
|
||||
amended-design: 02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md
|
||||
@@ -15,7 +16,7 @@ found that the host's `remove` path stops every `service` resource that is no lo
|
||||
delete". `store.Orphans` matches by id alone. So any of these stops the unit:
|
||||
|
||||
- the module is unassigned — by mistake, or to switch it for another;
|
||||
- the node is sent a deliberately-empty declaration ([issue 127](../127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md));
|
||||
- the node is sent a deliberately-empty declaration ([issue 149](../149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md));
|
||||
- a later catalogue version renames the resource's `id`.
|
||||
|
||||
That is right for a service the mesh brought into being. It is wrong for a unit the mesh
|
||||
@@ -64,3 +65,9 @@ something to settle in passing.
|
||||
|
||||
The unassign preview is partly answered — the host's plan names each unit it will stop — and the
|
||||
controller's side is left open.
|
||||
|
||||
## Closed
|
||||
|
||||
*2026-09-29, in a grooming pass rather than by whoever fixed it.* Undeclaring no longer stops a unit the mesh only reloaded or only kept running. Found by
|
||||
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
|
||||
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
---
|
||||
status: located
|
||||
status: resolved
|
||||
opened: 2026-09-28
|
||||
located-in:
|
||||
- mesh-controller internal/catalogue/manifest.go
|
||||
@@ -7,7 +7,7 @@ located-in:
|
||||
- mesh-controller internal/catalogue/declaration.go
|
||||
- mesh-controller examples/route-proxy
|
||||
- mesh-catalog (every routed module manifest)
|
||||
fixed-by:
|
||||
fixed-by: mesh-controller bdf965d (a module names its endpoints) and c68d3a7 (an assignment configures an endpoint as one thing) — the filter, the proxy's names and both authorities now read one statement
|
||||
amended-design: 03-DESIGN/01-to-be/08-connectivity.md
|
||||
---
|
||||
|
||||
|
||||
@@ -0,0 +1,47 @@
|
||||
# Resolution
|
||||
|
||||
*2026-09-29.*
|
||||
|
||||
**Built, and this record did not say so.** The issue was written on 2026-09-28 and answered the same
|
||||
week by two commits in `mesh-controller`; nothing came back to close it, so the mesh's own account of
|
||||
itself said for a day that reach was declared nowhere while the code read it in three places.
|
||||
|
||||
- `bdf965d` — *a module names its endpoints, and a route names the one it serves*. `listens[].name`
|
||||
is the endpoint; a route contribution names the endpoint rather than repeating a port.
|
||||
- `c68d3a7` — *an assignment configures an endpoint as one thing*. The `endpoints` settings key, per
|
||||
node, by endpoint name: `{"endpoints": {"ssh": {"port": 20134, "reach": "public"}}}` — port, label
|
||||
and reach in one block, which is what [ADR 0138](../../02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md)
|
||||
asked for and what [ADR 0046](../../02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md)
|
||||
said configuration is.
|
||||
|
||||
## The three readers, which is what the issue was about
|
||||
|
||||
The complaint was that the per-node source override had exactly one caller. It now has three, and
|
||||
they are the three mechanisms reach was decided to settle at once:
|
||||
|
||||
| reader | what it does with it |
|
||||
|---|---|
|
||||
| the filter | `Reaches` turns each endpoint's reach into the rule for its machine port |
|
||||
| the proxy's names | `composeName` composes the public name, the internal name, or both — and a name nobody asked for is not composed |
|
||||
| the authorities | the proxy certifies only names it was actually given, each from its own authority, through two host policies rather than one |
|
||||
|
||||
**A routed endpoint keeps the manifest's port**, which is ADR 0138's own insight and older than it
|
||||
([ADR 0045](../../02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)): the
|
||||
proxy is how it is reached, so `public` there asks for a public *name*, not an open port.
|
||||
|
||||
**An endpoint that is not routed is reached and never named.** Git over ssh is that case — the one
|
||||
the issue said the model could not express — and it is now the ordinary one.
|
||||
|
||||
## How it is checked
|
||||
|
||||
`internal/catalogue/endpoints_setting_test.go`: a block says port, label and reach; a block may say
|
||||
only a reach; a name the module does not declare is refused; a reach outside the four values is
|
||||
refused; and saying the same thing twice — once in the block, once through the older per-port keys —
|
||||
is refused rather than resolved by whichever is read last. The proxy's half is `policy_test.go` and
|
||||
`authority_test.go`: a name the mesh did not send is not certified, by either authority.
|
||||
|
||||
## What is left, and it is not this
|
||||
|
||||
The older keys (`ports`, `expose`, and reach keyed by port) still work beside the block. They are
|
||||
what the block replaces, and retiring them is its own small change — not a gap in what reach can
|
||||
say.
|
||||
+12
@@ -71,3 +71,15 @@ against a suite nobody can execute.
|
||||
- The lab rewrites a bundle's registry-prefixed third-party references to upstream ones for a
|
||||
machine with an uplink (`test/integration/harness.ts`); the new bundle's bus reference needed
|
||||
that rule added, which is done and is not this issue.
|
||||
|
||||
## What one of its fixes then did to a running mesh (2026-09-30)
|
||||
|
||||
The change that stopped the doubling — putting the stream into a push consumer's delivery subject —
|
||||
is correct on a foundation being raised and fatal on a mesh that is already running: the server will
|
||||
not move that subject while a subscriber is bound, and a node is bound to its declaration consumer
|
||||
the whole time it is up. The control plane crash-looped on the first build that carried it.
|
||||
|
||||
Recorded and fixed as [issue 156](../156-moving-a-consumers-delivery-subject-stops-the-control-plane/00-report.md).
|
||||
Noted here because this record is where somebody will arrive when reading why the subject carries the
|
||||
stream at all, and the answer is incomplete without it: **the raise path was the only one exercised,
|
||||
and it is the one path on which nothing is bound.**
|
||||
|
||||
+88
@@ -82,12 +82,100 @@ managing, or genesis carries a user list that includes the first node's enrolmen
|
||||
takes over from there. Both are decisions, not patches, and both belong to the genesis step that was
|
||||
deliberately left until last.
|
||||
|
||||
## 5 — the composed user list has to be placed by hand at genesis *(fixed)*
|
||||
|
||||
The account a token is the password of is **not recorded at all**: the composer names an enrolment
|
||||
user for every machine with a live token, nothing minted a credential for it, and the composition
|
||||
left it out as a user with no password. The comment above the issuing code already claimed
|
||||
otherwise — *"the account is created before the token is handed over"* — which is how it went
|
||||
unnoticed. Issuing a token now records that account, with the token's own secret as its password,
|
||||
because that is the string the machine will present.
|
||||
|
||||
Placing it is the other half. The list reaches the machine running the bus in that machine's
|
||||
declaration, which a machine that has not enrolled does not get, so at genesis it cannot arrive
|
||||
that way. **The control plane composes and says what it composed** — `broker accounts`, to standard
|
||||
output — and whoever is raising the machine writes it beside the bus's configuration and makes the
|
||||
server re-read it. Twice, because two accounts come into existence at different moments: the
|
||||
enrolment when the token is issued, and the machine's own when it enrols. A control plane that
|
||||
wrote the file itself would have to know where the bus keeps its configuration and how to make it
|
||||
reload, which is the module's knowledge and is what the module takes over on the first push.
|
||||
|
||||
With that, **a first node enrols against the bus it just raised** — measured, from bare, in the
|
||||
lab.
|
||||
|
||||
## 6 — and is enrolled twice, keeping a credential the mesh has replaced *(fixed)*
|
||||
|
||||
```
|
||||
mesh-controller: enrolled anchor
|
||||
mesh-controller: enrolled anchor (the same second)
|
||||
```
|
||||
|
||||
One `enrol` on the machine, two enrolments in the control plane. Each mints the node a fresh bus
|
||||
password and returns it; the machine keeps the answer to the first, and the mesh keeps the hash of
|
||||
the second. The machine then reconnects for ever as a user whose password the mesh rotated out from
|
||||
under it — *authentication error - User "node.anchor"* on the bus, `Authorization Violation` in the
|
||||
host's log, and a node that never reports.
|
||||
|
||||
What is ruled out: the host asking twice — it asks again only when the mesh says *try again*, and
|
||||
a refused attempt is not logged as an enrolment. Redelivery by the consumer — there is one
|
||||
consumer, its acknowledgement window is thirty seconds, and the handler is quick.
|
||||
|
||||
What is left: the client re-publishing when an acknowledgement is slow, which is what its defaults
|
||||
do. That was addressed by giving the publish a message id derived from its own bytes, so the stream
|
||||
discards the copy — **and the duplicate survived it**, so either the id is not reaching the stream
|
||||
or the second copy is not a copy. This is where the trail stops.
|
||||
|
||||
Worth saying plainly: **the mint is the fragile part, not the delivery.** An enrolment answered
|
||||
twice is survivable if the answer is the same both times, and it cannot be — the mesh keeps only
|
||||
the hash, so a second answer is necessarily a different credential.
|
||||
|
||||
**Found, and it is not about enrolment at all** *(fixed)*. The bus's own counters settled it: one
|
||||
message published, one held in the stream, one delivery, nothing redelivered — and the controller
|
||||
enrolled the machine twice. So the handler ran twice on one delivery.
|
||||
|
||||
A push consumer delivers onto an ordinary subject, and **everything subscribed to that subject gets
|
||||
a copy**. The controller holds a consumer called `controller` on CONTROL and another called
|
||||
`controller` on EVENTS, and the delivery subject was derived from the consumer's name alone — so
|
||||
both were `_DELIVER.controller`, the one process held both subscriptions, and every message from
|
||||
either stream was acted on twice.
|
||||
|
||||
Enrolment is where it drew blood, because enrolling twice mints twice and the second credential
|
||||
replaces the first. But it applied to **every report and every event the controller follows**, and
|
||||
it is the kind of fault that leaves no trace: nothing is redelivered, no counter is wrong, the work
|
||||
simply happens twice. The comment in the receiving code about a merge that ran the whole catalogue
|
||||
five times over on 2026-09-28 is the same shape seen from the other end.
|
||||
|
||||
The stream is in the delivery subject now, because the pair is what identifies a consumer — the
|
||||
server scopes a durable's name to its stream, and this subject was the one place that scoping was
|
||||
dropped. A subscriber's permission gains the same shape, keeping the bare name so an existing
|
||||
consumer keeps working until the controller's next assertion moves it.
|
||||
|
||||
*How it is checked:* the consumers the mesh asks for are asserted to deliver onto distinct subjects,
|
||||
in the controller's own suite. Against a server it would be invisible, which is the point.
|
||||
|
||||
## Where it belongs
|
||||
|
||||
`mesh-host` (the bundle and the enrolment path) and `mesh-controller` (the certificate command, and
|
||||
the composition that cannot reach the bus at genesis). Three of the four are fixed on branches; the
|
||||
fourth is the genesis work.
|
||||
|
||||
## What made it slow, and what was changed so it is not
|
||||
|
||||
Six faults behind one another, each found by raising a machine and reading what it said. What cost
|
||||
the most was not the faults:
|
||||
|
||||
- **Every bed's own instructions named the bundle that cannot work**, so the first three attempts
|
||||
ended in a control plane crash-looping on a missing bus. They name the working one now.
|
||||
- **A host binary built without its system** refuses everything it is given with *this host was
|
||||
built for ""*, which reads like a broken bundle. The lab's README says so.
|
||||
- **`make image` in the control plane had been broken for as long as its base was pinned**: the
|
||||
Dockerfile's fallback is a Go older than the module asks for, and the pipeline never saw it
|
||||
because the pipeline passes the declared base in. It reads the base from the manifest now.
|
||||
- **Leaving the machine standing is what answers the question.** Every finding above came from
|
||||
shelling in afterwards — the host's log, the bus's log, the file the bus was actually handed —
|
||||
and none from the test's own output, which says only that nothing converged. The bed takes
|
||||
`MESH_LAB_KEEP`, and the README says to reach for it first.
|
||||
|
||||
## What it cost, for the next person
|
||||
|
||||
Every lab bed still names `foundation-first-node.lock` in its own instructions, and that bundle
|
||||
|
||||
@@ -0,0 +1,75 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-09-29
|
||||
located-in: [nothing in the mesh — the surface in use is the predecessor's, installed on the workstation]
|
||||
---
|
||||
|
||||
# 147 — the operator's tools still dial the bus that was removed
|
||||
|
||||
## What was observed
|
||||
|
||||
Every tool call an operator makes against the mesh fails, on every machine, with the same answer:
|
||||
|
||||
```
|
||||
AMQP not connected — cannot reach hal/mesh@novox
|
||||
AMQP not connected — cannot reach hal/mesh@shanks
|
||||
```
|
||||
|
||||
The mesh moved to one bus and the previous transport was deleted
|
||||
([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md), cut over
|
||||
2026-09-28). The surface an operator drives the mesh through — the tool bridge their assistant
|
||||
speaks to — still opens an AMQP connection, so it cannot reach anything. Not one node answers,
|
||||
including the machine the operator is sitting at.
|
||||
|
||||
**What that leaves.** The mesh itself is healthy: nodes are current, modules run, the bus carries
|
||||
the mesh's own traffic. What is gone is the way a person asks it anything. Every question — what a
|
||||
node runs, what is assigned, what a module's settings are — has to be asked by opening a shell on
|
||||
the machine and running the control plane's binary inside its container, which is precisely the
|
||||
path the tool surface exists to remove, and which nothing checks, records or permits.
|
||||
|
||||
It also silently changes how work gets done: an assistant told to use the mesh's tools finds them
|
||||
dead, falls back to `ssh` and `docker exec`, and the operator discovers the fallback rather than
|
||||
the fault.
|
||||
|
||||
## Why this is here and not a note in the knowledge base
|
||||
|
||||
The rule is that everything happens over the mesh's bus. A surface that cannot reach that bus is
|
||||
that rule enforced by nothing — and the mesh reports itself healthy throughout, because what broke
|
||||
is not a module, a node or a provision but the thing standing outside asking them questions.
|
||||
|
||||
The same cut-over removed the assistant's long-term memory (`recall`, `memorize`) for the same
|
||||
reason, which is why lessons from the last two days were written into this repository by hand.
|
||||
|
||||
## What would have prevented it
|
||||
|
||||
- **The tool surface as a consumer of the bus like any other**, so moving the bus moves it — rather
|
||||
than a separate bridge with its own connection settings that nothing resolves.
|
||||
- **A check that a tool call reaches a node**, run where the mesh's other checks run. Every check
|
||||
the mesh makes today is about the relationship between the mesh and a machine; none asks whether
|
||||
a person can ask it anything ([issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
|
||||
the same shape one level out).
|
||||
|
||||
## Diagnosed at once, because the answer was in the configuration
|
||||
|
||||
**The surface is not the mesh's.** The tool server the operator's assistant speaks to is the
|
||||
predecessor's brain, installed on the workstation and started as a local process, with the
|
||||
predecessor's broker URL — `amqp://…@<the control node>` — written into the assistant's own
|
||||
configuration. Nothing about it is a module: it has no manifest, no assignment, no seat, no account
|
||||
on the mesh's bus, and the mesh has never known it exists.
|
||||
|
||||
So nothing regressed. The mesh removed a transport that this program still dials, and the program
|
||||
was never part of the mesh to be moved. **The mesh has a tool model** — `tools` in a manifest, the
|
||||
calls a seat accepts, the subjects a module answers on, and a control-plane command that asks one —
|
||||
and **nothing publishes an operator-facing surface onto it.** What an operator uses is the thing
|
||||
that came before, kept alive by a URL in a file.
|
||||
|
||||
That is the issue, and it is larger than a broken connection: the way a person drives this mesh is
|
||||
outside the mesh.
|
||||
|
||||
## Evidence to carry into diagnosis
|
||||
|
||||
- The failure text names `hal/mesh@<node>`, so the bridge is resolving a node and then dialling the
|
||||
old transport.
|
||||
- It fails identically for the local machine, which rules out reachability and points at the
|
||||
transport alone.
|
||||
- The mesh's own traffic over the new bus is unaffected: nodes report, declarations apply.
|
||||
@@ -0,0 +1,37 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-29
|
||||
located-in: [mesh-controller cmd/mesh-controller]
|
||||
---
|
||||
|
||||
# 148 — a manifest outside this catalogue has no check
|
||||
|
||||
## What was observed
|
||||
|
||||
A module's manifest is validated by a **test** — `internal/catalogue`'s suite parses every manifest
|
||||
in the catalogue checkout beside it and fails on one it cannot resolve. That works, and it is how
|
||||
several real faults were caught before a machine saw them.
|
||||
|
||||
It is available to exactly one repository: this one. Somebody describing their own application in
|
||||
their own repository — the case
|
||||
[ADR 0037](../../02-DECISIONS/0037-where-a-module-lives.md) calls *the case that matters most* —
|
||||
has no check at all. They write a manifest, register it with a running mesh, and find out whether
|
||||
it is valid when the mesh refuses it, or later, when a machine applies something that resolved and
|
||||
should not have.
|
||||
|
||||
The same record asks for the answer: **a `module check` command on the control plane's binary**, so
|
||||
a manifest is checked by the tool rather than by a test that imports the tool's internals.
|
||||
|
||||
## What would have prevented it
|
||||
|
||||
Nothing prevents this; it was noticed and left. ADR 0037 named it on 2026-09-01 and the record sat
|
||||
`proposed` until 2026-09-29, so the missing half was never anybody's task.
|
||||
|
||||
## Evidence to carry into diagnosis
|
||||
|
||||
- `mesh-controller/internal/catalogue` — `ParseManifest` and `CatalogueProblems` are the check, and
|
||||
both are internal.
|
||||
- The catalogue-wide test is `TestEveryCatalogueManifestDeclaresWhatItMounts` and its siblings; they
|
||||
take a path from `MESH_CATALOG`, so the mechanism is already path-driven and not repository-bound.
|
||||
- `mesh-controller module add` refuses a bad manifest at registration, which is the same check far
|
||||
too late: by then it is in a running mesh's records.
|
||||
+22
-3
@@ -1,10 +1,15 @@
|
||||
---
|
||||
status: located
|
||||
status: resolved
|
||||
opened: 2026-09-27
|
||||
located-in: [mesh-controller cmd/mesh-controller/push.go]
|
||||
located-in: [mesh-controller cmd/mesh-controller/push.go, mesh-controller cmd/mesh-controller/sendable.go]
|
||||
fixed-by: mesh-controller sendable.go and push.go — an empty declaration is sent carrying `owns_nothing`, and the host refuses an empty body that does not carry it
|
||||
---
|
||||
|
||||
# A declaration that shrinks to empty is skipped, so the node keeps what it should drop
|
||||
# 149 — a declaration that shrinks to empty is skipped, so the node keeps what it should drop
|
||||
|
||||
*Opened as 127 and renumbered on 2026-09-29: two records were given that number on the same day.*
|
||||
*The other kept it, because three documents and three source files cite it by number and nothing
|
||||
cited this one but a decision and a sibling issue, both corrected with this move.*
|
||||
|
||||
## What was observed
|
||||
|
||||
@@ -37,3 +42,17 @@ mean "own nothing", which the host already applies correctly when it receives on
|
||||
|
||||
On ace, one command drops it permanently (the corrected controller never re-composes it):
|
||||
`sudo ufw delete allow 5671`. At ace's converge it would clear on its own.
|
||||
|
||||
## Closed
|
||||
|
||||
*2026-09-29, in a grooming pass.* Both halves are on `main` and both name this issue.
|
||||
|
||||
- The control plane **sends** it: a declaration that composes to no resources goes out with
|
||||
`owns_nothing`, and `push` says *sent, not skipped*.
|
||||
- The host **refuses an empty body that does not carry it**, so a truncated or mis-composed
|
||||
declaration can never be read as "own nothing" — which is the failure the fix had to avoid while
|
||||
making the empty case expressible.
|
||||
|
||||
Closed by reading the code rather than by watching a machine let go of a stray resource; the record
|
||||
says so rather than implying a run.
|
||||
|
||||
@@ -0,0 +1,52 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-29
|
||||
located-in:
|
||||
- mesh-controller internal/catalogue/declaration.go (contributions are emitted for an assigned module, taken or not)
|
||||
fixed-by:
|
||||
---
|
||||
|
||||
# 150 — A route is contributed before its module is taken
|
||||
|
||||
## What was observed
|
||||
|
||||
ace is adopted and runs `route-adapter` beside the predecessor's traefik (ADR 0104). The first web
|
||||
module migrated there was searxng. `assign ace searxng` + `push ace` — the step that is supposed to
|
||||
change nothing on an adopted machine (ADR 0100, hq 125: assign holds, take replaces) — took
|
||||
`searxng.zurag.be` down:
|
||||
|
||||
```
|
||||
https://searxng.zurag.be/ → 502
|
||||
```
|
||||
|
||||
for about five minutes, until the module was unassigned again.
|
||||
|
||||
## Why
|
||||
|
||||
Assign held everything it found on the machine — the predecessor's `searxng` container, its
|
||||
directories — exactly as designed. But the module's **route contribution** is not a resource on the
|
||||
machine, so nothing held it: it reached `route-adapter` at once, which wrote
|
||||
`mesh-searxng.zurag.be.yml` into traefik's file provider pointing at the mesh's assigned machine port
|
||||
(`http://ace.internal:20000`) — where nothing listened, because the container that would was held.
|
||||
traefik's file router for the name then won over the predecessor's docker-label router for the same
|
||||
name, and the name served a dead backend.
|
||||
|
||||
On this occasion the window was lengthened by an unrelated failure (the machine's docker address
|
||||
pools were exhausted, so the module's network could not be created), but the fault does not depend on
|
||||
it: **between assign and take, every routed module's public name points at a backend the mesh has
|
||||
deliberately not started.** The runbook's §6 step 2 ("check the preparation — HAL's service still
|
||||
serves") is false for every routed module on a node running route-adapter.
|
||||
|
||||
## What the operator did
|
||||
|
||||
Rolled back per runbook (unassign before take), then re-ran as assign → push → wait for the node to
|
||||
report its holds → take → push, keeping the window to the ~10 s between the two pushes. A workaround,
|
||||
not a fix: a module whose data must move between assign and take (runbook §6 step 3) cannot shrink
|
||||
its window this way.
|
||||
|
||||
## What would be right
|
||||
|
||||
A contribution from a module that is assigned but **not taken** on an adopted node should be held like
|
||||
the module's resources are — withheld from the provider until take — or the provider should be told
|
||||
the contributor is held and leave the predecessor's route alone. Either keeps "assign changes nothing"
|
||||
true for routed modules.
|
||||
@@ -0,0 +1,94 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-29
|
||||
located-in:
|
||||
- mesh-host internal/apply/apply.go (containerSpecReading hashes every `host` entry)
|
||||
- mesh-controller internal/catalogue/declaration.go (withMeshNames gives every container the mesh's names)
|
||||
fixed-by:
|
||||
---
|
||||
|
||||
# 151 — A new name recreates every container in the mesh
|
||||
|
||||
## What was observed
|
||||
|
||||
Migrating one small module on ace (searxng) took four routine controller actions: `node public-domain
|
||||
ace zurag.be`, `assign ace searxng` + push, `unassign ace searxng` + push (a rollback, see
|
||||
[issue 150](../150-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take
|
||||
+ push again. Each push to ace also pushed novox ("this push left g14, novox, shanks behind … sending it
|
||||
too"). novox's host then **replaced every container it runs, twice**:
|
||||
|
||||
```
|
||||
22:46 … updated postgres.server (mesh-store): replaced; a container's configuration is fixed when it is created
|
||||
22:46 … updated distribution.store (mesh-registry): replaced; …
|
||||
22:46 … updated route-proxy.server (route-proxy): recreated …
|
||||
22:51 … updated postgres.server (mesh-store): replaced; …
|
||||
22:52 … updated gitea.server (gitea): replaced; …
|
||||
22:52–22:55 mesh-vault, builder, all of mailu, keycloak, minio, mongodb, invoicing, photos, umami, …
|
||||
```
|
||||
|
||||
Replacements per minute on novox, 22:46–22:55: 12, 8, 13, 5, 6, 8, 8, 8, 8, 4. The control plane was
|
||||
unreachable twice while its own store came back through crash recovery
|
||||
(`FATAL: the database system is starting up`), and dependants (umami) crash-looped until it did. None
|
||||
of the replaced containers belonged to the module being migrated, or to ace.
|
||||
|
||||
## Why (confirmed part)
|
||||
|
||||
Every container the mesh runs is given the mesh's names as `--add-host` entries
|
||||
(`withMeshNames`), and the host puts every entry into the container's spec digest — deliberately,
|
||||
since [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md): a container left alone when the roster moved kept a five-day-old
|
||||
address and restarted 2286 times. A running container cannot have its hosts changed, so a changed
|
||||
digest means a replace.
|
||||
|
||||
The consequence is that **the roster is part of every container everywhere**: anything that adds,
|
||||
removes or moves one name — a node's public domain, a routed name, an unassign — replaces every
|
||||
container on every machine that carries the list. On the hub that includes the control plane's store,
|
||||
the registry, the edge and mail.
|
||||
|
||||
## Not yet established
|
||||
|
||||
Which name moved in each pass. gitea's hosts after the fact list the `.internal` names and novox's
|
||||
internally routed `*.novox.be` names, and **not** `searxng.zurag.be` — so it is not simply "a routed
|
||||
name was added". Two passes suggest the set changed at assign and changed back at unassign; the
|
||||
declarations before and after would say, and nothing on the machine records the previous one.
|
||||
|
||||
## Why it matters now
|
||||
|
||||
The node-by-node migration adds names one module at a time. On ace alone that is ~25 web modules; if
|
||||
each assignment moves the roster, each is a full restart of every hub service, and a rollback is
|
||||
another. The migration is paused on this.
|
||||
|
||||
## What has since been ruled out as a cause (2026-09-30)
|
||||
|
||||
The roster was also moving for a reason that was not a name changing at all: a node whose plan would
|
||||
not compose had its routed names silently dropped from the roster handed to every machine, so a
|
||||
briefly unreachable store withdrew and restored a name on alternating passes. That is
|
||||
[issue 152](../152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md), and it is
|
||||
fixed — it is what made the churn on the control node repeat every six minutes for twenty-five
|
||||
minutes rather than once per operator action.
|
||||
|
||||
**It does not close this record.** A roster that changes for a real reason — a machine joining, a
|
||||
module assigned, a public domain set — still replaces every container on every machine that carries
|
||||
the list. 152 removed the false reasons; the question below is still open.
|
||||
|
||||
## What would be right (for diagnosis)
|
||||
|
||||
Keep 135's guarantee (no container runs with a stale address) without making the roster part of every
|
||||
container's identity — e.g. resolve mesh names through a resolver the container asks at lookup time
|
||||
rather than baked entries, or scope each container's entries to the names it actually binds.
|
||||
|
||||
## Answered (2026-09-30): the first of those two
|
||||
|
||||
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) takes
|
||||
the resolver, not the scoping. Scoping was the close call and was rejected for contradicting the
|
||||
standing intent that anything on the mesh can call anything on it: it would make a name reachable only
|
||||
where the mesh was told in advance it would be wanted, and it would leave the roster in the digest, so
|
||||
the churn returns whenever a widely-bound name moves.
|
||||
|
||||
Resolving at lookup time makes staleness impossible rather than detected — 109 and 135 stop being bugs
|
||||
that were fixed and become a shape that does not exist — and makes a name's blast radius nothing.
|
||||
|
||||
**This record stays open**, because the record answers it and the code does not. Nothing may stop
|
||||
copying names until a container can reach the resolver from any of the runtime's networks
|
||||
([issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)),
|
||||
which it cannot on two of four machines today. Removing the copy first reintroduces 109 and 135
|
||||
silently, on a live mesh, which is how both were found. The order is in the record.
|
||||
@@ -0,0 +1,124 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-29
|
||||
located-in:
|
||||
- mesh-controller cmd/mesh-controller/plan.go (routeNamesInTheMesh skips a node whose plan will not compose)
|
||||
fixed-by: mesh-controller 6c5dfd0 (PR 147)
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 152 — A node whose plan will not compose silently removes its names from every machine
|
||||
|
||||
## What was observed
|
||||
|
||||
For twenty-five minutes, and for seventeen of them after the last operator action, the control-node's
|
||||
host applied all 327 of its resources every ~6.5 minutes without pause, replacing every container on
|
||||
the machine each time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail,
|
||||
and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine
|
||||
sat near 8. Nothing was converging: each pass ended and the next began four seconds later.
|
||||
|
||||
The two machines carrying no containers were not churning. They were only knocked off the bus each
|
||||
time the control node re-created it, reconnected, re-heard the same declaration and applied it again
|
||||
as a no-op.
|
||||
|
||||
## Why: the roster alternates between two values, and it is part of every container
|
||||
|
||||
Two consecutive declarations were compared by reading the `--add-host` entries of four containers
|
||||
the moment each pass created them:
|
||||
|
||||
```
|
||||
23:02 ace.internal drive.novox.be g14.internal keycloak.novox.be novox.internal
|
||||
office.novox.be portainer.novox.be shanks.internal umami.novox.be (9 names)
|
||||
23:07 … the same nine, and searxng.zurag.be (10 names)
|
||||
```
|
||||
|
||||
One routed name — belonging to a module on another machine entirely — leaves the roster and comes
|
||||
back. Because the roster is part of every container's spec digest ([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)),
|
||||
each flip is a different identity for every container on the machine, and a running container cannot
|
||||
have its hosts changed. So every flip replaces all of them.
|
||||
|
||||
**What makes it flip is a swallowed error.** `routeNamesInTheMesh` composes every node's plan to
|
||||
find the names it serves, and when one will not compose it moves on:
|
||||
|
||||
```
|
||||
plan, settings, err := planFor(ctx, open, n.Name)
|
||||
if err != nil {
|
||||
continue
|
||||
}
|
||||
```
|
||||
|
||||
A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does
|
||||
not know", but "the mesh states these names do not exist", to every machine at once.
|
||||
|
||||
## Why it sustains itself
|
||||
|
||||
The loop closes through the control plane's own database:
|
||||
|
||||
1. An apply replaces `mesh-store` — the store the control plane reads — by removing the container,
|
||||
so postgres comes back through crash recovery.
|
||||
2. While it recovers it refuses connections: `FATAL: the database system is not yet accepting
|
||||
connections / Consistent recovery state has not been yet reached` (observed, 21:08:39 UTC, every
|
||||
pass).
|
||||
3. `planFor` for the other machine fails against that store. `routeNamesInTheMesh` swallows it and
|
||||
drops its routed name.
|
||||
4. The roster changed, so all 327 resources differ, so all are replaced — including `mesh-store`,
|
||||
and including the bus, which is why the host also cannot report: `applied, and could not tell the
|
||||
mesh: reporting: nats: connection closed`.
|
||||
5. Back to 1.
|
||||
|
||||
Every pass destroys the evidence the next pass needs to decide it has nothing to do, and nothing
|
||||
outside the machine has to be wrong for it to continue.
|
||||
|
||||
**It is metastable, not permanent.** It ran from 22:46 to 23:11 — five full replacements of every
|
||||
container on the machine — and then stopped on its own, when one pass happened to read the store
|
||||
during a window it was up, composed the same roster twice running, and found nothing to do. Load fell
|
||||
from 7.8 to 1.7 and the machine returned to its five-minute idle tick.
|
||||
|
||||
That it ends by luck is the point, not a mitigation. The exit condition is a race the mesh does not
|
||||
control, the operator cannot see, and nothing reports; the same four actions on a slower machine, or
|
||||
a larger store, would not have found it. An outage that clears itself after twenty-five minutes and
|
||||
five restarts of the forge, the directory and mail is not a smaller fault than one that does not — it
|
||||
is the same fault, harder to catch.
|
||||
|
||||
## What it is not
|
||||
|
||||
- Not the operator's four actions on the other machine. Those explain the first passes
|
||||
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)); they were
|
||||
finished seventeen minutes and three full passes before these measurements.
|
||||
- Not a file that keeps changing. `/etc/hosts`, the bus's account list and the vault's export were
|
||||
hashed across passes and are byte-identical, and the directory reported `mode 755 to 700` every
|
||||
pass while already being `700`. Those resources are **misreported as changed** and are worth their
|
||||
own question, but they are not what moves a container's identity.
|
||||
- Not the lost report alone. A report that cannot be delivered explains a re-apply; it does not
|
||||
explain a re-apply that finds 327 differences.
|
||||
|
||||
## Why it matters beyond this outage
|
||||
|
||||
The same `continue` makes every routed name in the mesh conditional on every node's plan composing at
|
||||
the moment any machine is pushed to. One unreachable or half-migrated machine is enough to withdraw
|
||||
its names from everywhere — and the withdrawal is indistinguishable, on the receiving machine, from
|
||||
the operator having removed them.
|
||||
|
||||
The codebase already states the rule this breaks, forty lines away, about the same kind of lookup:
|
||||
|
||||
> A lookup failure is an error, never "not found": collapsing the two composed a declaration without
|
||||
> the trust whenever the inventory hiccuped, delivered by a push that reported success.
|
||||
|
||||
## How it was fixed, and how the fix is checked
|
||||
|
||||
`planFor` now marks the two failures that really are a statement about the node — its set not
|
||||
composing, and a setting that reaches nothing — and the three gatherers pass over those and only
|
||||
those. Every other failure is raised, naming the machine and the read.
|
||||
|
||||
Five tests hold it: a set that cannot compose is marked as the node's own; a store that cannot be
|
||||
read is *not*; one incoherent node still does not cost the rest their names; a roster is never
|
||||
returned beside an error; and the raised failure names what could not be read.
|
||||
|
||||
The two sibling gatherers were audited and fixed the same way — the grant composer, which would have
|
||||
withheld a consumer's credential, and the private-network membership, which would have taken a
|
||||
machine off the overlay. Three other `planFor` callers were audited and left alone: they refuse or
|
||||
report rather than silently withdraw, which is the safe direction.
|
||||
|
||||
**This does not close [issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md).**
|
||||
A roster that changes for a real reason still replaces every container in the mesh. This removes the
|
||||
false reasons; whether the roster belongs in a container's identity at all is that record's question.
|
||||
+1
-15
@@ -8,7 +8,7 @@ fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 149 — An adopted machine's data cannot be placed where it is
|
||||
# 153 — An adopted machine's data cannot be placed where it is
|
||||
|
||||
## What was observed
|
||||
|
||||
@@ -55,17 +55,3 @@ The two assignment halves 0112 decided: a setting that places a declared directo
|
||||
path on this node, and a setting that says where an access's data is — both validated like
|
||||
`endpoints` (unknown ids refused), and an access placed by the assignment still never created,
|
||||
chowned or removed.
|
||||
|
||||
## Addendum 2026-09-30 — whose the data is, not only where
|
||||
|
||||
The same modules need to run **as the owner of the data they access**: linuxserver images take
|
||||
`PUID`/`PGID`, and ace's library is `media:media` (`1001:2000`, group members `ace`, `n8n`, `media`).
|
||||
A manifest default is one value for every machine, and an assignment's settings do not reach a
|
||||
container's environment. Adding user and group management to the mesh would contradict ADR 0051 —
|
||||
the mesh owns nothing about the operator's data.
|
||||
|
||||
The data already says whose it is, and the host already looks at it when it confirms an access exists.
|
||||
Proposed: expose that as a fact the module can ask for, like `${port:…}` — e.g. `${access:<id>:uid}`
|
||||
and `${access:<id>:gid}`, read from the accessed path on the machine — so a media module declares
|
||||
`PUID=${access:media:uid}` and is right on every machine without the mesh creating, naming or
|
||||
chowning anyone.
|
||||
@@ -0,0 +1,42 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-29
|
||||
located-in:
|
||||
- hq 02-DECISIONS/0138 (reach: internal | public | both)
|
||||
- mesh-controller internal/catalogue/filtering.go (Reaches)
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 154 — A machine's own network is not a reach
|
||||
|
||||
## What was observed
|
||||
|
||||
Preparing ace's modules. ace sits on a home network (192.168.1.0/24) behind a router, and several of
|
||||
its services are reached **from that network by devices that will never be mesh machines**:
|
||||
|
||||
- mosquitto `1883` — an IoT light switch (`sonoff-office-light-switch`) and home-assistant;
|
||||
- unifi `8080`/`3478 udp`/`10001 udp` — the access points' inform, STUN and discovery;
|
||||
- plex `32400` — LAN streaming clients (three connected at survey time);
|
||||
- home-assistant `8123`, and the resolver on the LAN address.
|
||||
|
||||
ADR 0138 gives an endpoint's reach as `internal` (the private overlay), `public` (anywhere) or `both`.
|
||||
None of them says *this machine's own network*. The predecessor could: its unifi manifest opened
|
||||
inform/STUN/discovery `from: 192.168.0.0/16, 10.0.0.0/8, 172.16.0.0/12`.
|
||||
|
||||
## Consequence
|
||||
|
||||
The only reach that includes a LAN device is `public`. While ace is adopted that is harmless — its
|
||||
own firewall stays and admits the LAN — and behind NAT "anywhere" happens to mean the LAN. But:
|
||||
|
||||
- it states the wrong thing: an operator reading `reach: public` on an IoT broker believes it is on
|
||||
the internet, and a router port-forward added later for something else makes it so;
|
||||
- at `converge ace`, the mesh's filter is the sum of what it listens on (ADR 0045). An endpoint left
|
||||
`internal` cuts every LAN device off at the flip; one set `public` opens it to the internet on any
|
||||
machine with a public address.
|
||||
|
||||
## What would be right (for diagnosis)
|
||||
|
||||
A reach — or a source — that means the networks the machine is directly attached to (its uplink's
|
||||
subnets, as the machine reports them), so a LAN-only service is declared as exactly that and the
|
||||
filter can admit it without admitting the internet.
|
||||
@@ -0,0 +1,58 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-29
|
||||
located-in: [hq 00-META/checks/cycle.py]
|
||||
fixed-by: hq f89aef9 (PR 192)
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 155 — Two records may share a number, and every check passes
|
||||
|
||||
## What was observed
|
||||
|
||||
On 2026-09-29 two machines opened issues against this repository within the same hour. Both read
|
||||
`main` correctly and both took "the next free number", and they collided twice:
|
||||
|
||||
| | one machine opened | the other had already used |
|
||||
|---|---|---|
|
||||
| first | 147, 148 | 147, 148 on an unmerged branch |
|
||||
| second | 149, 150 | 149 on an unmerged branch, 150 from renumbering the first collision |
|
||||
|
||||
The first collision was reconciled by hand before merging. The second was **merged into `main`**, and
|
||||
`records.py`, `cycle.py` and `index.py` all reported success over a tree holding
|
||||
`149-a-declaration-that-shrinks-to-empty` beside `149-an-adopted-machines-data-cannot-be-placed-where-it-is`,
|
||||
and two folders numbered 150.
|
||||
|
||||
## Why
|
||||
|
||||
The number is allocated as `max(main) + 1`, and `main` lags every open pull request — seven of them
|
||||
that evening. Two readers of the same `main` therefore compute the same next number, and neither is
|
||||
doing anything wrong. The existing reconciliation precedent (a second record numbered 127 became 149)
|
||||
assumed a single writer, which stopped being true when a second machine began filing its own findings.
|
||||
|
||||
## Why it matters
|
||||
|
||||
An issue number is how every other record cites this one — `fixed-by:`, `located-in:`, a decision
|
||||
record's consequence, a commit message. Two records answering to one number is a citation that
|
||||
resolves to whichever folder the reader happened to open, and the failure is silent on both sides:
|
||||
the citer is not wrong, and the cited record exists.
|
||||
|
||||
It is also exactly the class this repository says it does not permit — a rule (`00-META/process/03-issues.md`:
|
||||
"take the next free number") enforced by nothing.
|
||||
|
||||
## How it was fixed, and how the fix is checked
|
||||
|
||||
`cycle.py` now refuses a tree in which two issue folders share a leading number, and names both.
|
||||
Proven by adding a duplicate and watching it fail, then removing it and watching it pass.
|
||||
|
||||
The colliding records were renumbered 153 and 154, in the branch that landed last — renumbering a
|
||||
branch whose author is still pushing only moves the race.
|
||||
|
||||
**The check catches the collision; it does not prevent it.** Allocating a number still needs the open
|
||||
pull requests read as well as `main`. That is a habit the check now backstops rather than one it
|
||||
replaces, and playbook [03](../../00-META/process/03-issues.md) now says so at the step where the
|
||||
number is taken.
|
||||
|
||||
[ADR 0080](../../02-DECISIONS/0080-the-development-cycle-is-checked.md) enumerates what `cycle.py`
|
||||
enforces and named four things; it carries a progressive insight naming the fifth. The decision
|
||||
stands — this is one more thing frontmatter and file names can carry, found by its absence.
|
||||
+98
@@ -0,0 +1,98 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-29
|
||||
located-in:
|
||||
- mesh-controller internal/broker/jetstream.go (EnsureConsumer)
|
||||
fixed-by: mesh-controller e7da39d (PR 148)
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 156 — Moving a consumer's delivery subject stops the control plane, and only on a mesh that is running
|
||||
|
||||
## What was observed
|
||||
|
||||
The control plane crash-looped, every restart ending the same way:
|
||||
|
||||
```
|
||||
mesh-controller: asserting how novox hears its declaration:
|
||||
bringing consumer novox on NODES to match: nats: consumer name already in use
|
||||
```
|
||||
|
||||
It came up on the first build of the controller in eight hours. The machines kept running what they
|
||||
already held — this stops the mesh being *changed*, not the services it placed — and nothing could be
|
||||
pushed, no report was consumed and no enrolment answered, for as long as it lasted.
|
||||
|
||||
## Why
|
||||
|
||||
[Issue 146](../146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md) put the
|
||||
stream into a push consumer's delivery subject, because one process holding two consumers of the same
|
||||
name on two streams was given one subject and acted on every message twice.
|
||||
|
||||
**The server will not move a push consumer's delivery subject while a subscriber is bound to it.** It
|
||||
refuses with `consumer name already in use` — a message about the name, for a conflict about the
|
||||
subject, which is why the trail starts in the wrong place.
|
||||
|
||||
A node is bound to its declaration consumer the whole time it is up. That *is* a node listening for
|
||||
what it should be. So every node consumer in a mesh that is running is one the assertion cannot bring
|
||||
to match — and the assertion happens before the controller serves, so it never serves.
|
||||
|
||||
The controller's own two consumers moved without trouble, and are on the new subject in the live mesh.
|
||||
It asserts them before it subscribes, so nothing was bound.
|
||||
|
||||
## Why nothing caught it
|
||||
|
||||
The change was exercised on a mesh being raised, where every consumer is created rather than updated
|
||||
and nothing is bound to any of them. On that path the code is correct. The test that would have caught
|
||||
it needs a mesh that is already running: an existing consumer, a subscriber still attached, and then
|
||||
the assertion.
|
||||
|
||||
Reproduced exactly that way before the fix — same server version, same stream shape, same consumer —
|
||||
and it fails with the same words as the machine did. An earlier version of the same test unsubscribed
|
||||
first and passed against the code that was crash-looping on the control node.
|
||||
|
||||
## What it is not
|
||||
|
||||
- Not the change it shipped beside. The merge that triggered this build carried
|
||||
[issue 152](../152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md)'s fix and three
|
||||
other commits that had never been deployed; this one is 146's.
|
||||
- Not a version difference. The test server and the mesh's broker are both nats-server v2.10.29.
|
||||
|
||||
## How it was fixed
|
||||
|
||||
The consumer that works is kept, and the assertion says so instead of failing.
|
||||
|
||||
**Not deleted and re-made.** Re-making moves the subject, and a holder may not yet be allowed to
|
||||
subscribe to the new one: the wider grant travels in the bus's user list, which the control plane
|
||||
composes and a machine applies minutes later. On the live mesh the nodes are granted `_DELIVER.<node>`
|
||||
and not `_DELIVER.<node>.>` — re-making would have silenced every machine in the mesh, which is worse
|
||||
than the collision it was fixing and far harder to undo. That was the first fix written here, and the
|
||||
permission is the reason it was not shipped.
|
||||
|
||||
**Not fatal**, which is what 146's change intended and did not do: the bare subject still delivers, and
|
||||
collides only where one holder has two consumers of one name. A node has one.
|
||||
|
||||
## How the fix is checked
|
||||
|
||||
Two tests against a real server: a consumer with a subscriber bound keeps its subject, is reported,
|
||||
and still delivers to that subscriber; a consumer with nothing bound moves, so 146's fix still applies
|
||||
where the collision actually was.
|
||||
|
||||
## What is left
|
||||
|
||||
The node consumers stay on the bare subject, which is correct and not tidy. Nothing is wrong while
|
||||
they do not move: one consumer per name per stream cannot collide with itself.
|
||||
|
||||
The controller reports each one it kept, and did, on the start that fixed this — four node consumers
|
||||
and the build machine's worker, which is bound the same way and was not anticipated here:
|
||||
|
||||
```
|
||||
consumer novox on NODES still delivers to "_DELIVER.novox" and not "_DELIVER.novox.NODES":
|
||||
nats: consumer name already in use. It keeps working; the subject moves on an assertion
|
||||
made while nothing is bound to it
|
||||
```
|
||||
|
||||
**The wider grant has since landed** (2026-09-30, measured on the mesh's own broker config): every
|
||||
node is now allowed `_DELIVER.<node>.>` as well as the bare subject. That was the thing missing when
|
||||
this was diagnosed, and it is why re-making the consumers then would have silenced every machine.
|
||||
What remains is only the second half — an assertion made while each node is detached from its
|
||||
consumer — and that is its own piece of work, not a side effect of a restart.
|
||||
Reference in New Issue
Block a user