Compare commits
1
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
9cb53cd022 |
@@ -54,6 +54,4 @@ whose failure has never been observed is a guess about its own correctness.
|
||||
The development cycle, checked ([ADR 0080](../../02-DECISIONS/0080-the-development-cycle-is-checked.md)):
|
||||
a to-be design names a decision, an in-progress/implemented design names its owning code, a
|
||||
located/fixed issue names its owner, a fixed/resolved issue says what fixed it, a graduated
|
||||
research overview says what it became, and no two issue records share a number (issue 155 — the
|
||||
number is how a record is cited, and `main` lags every open pull request, so two people reading it
|
||||
allocate the same one). `python3 00-META/checks/cycle.py`
|
||||
research overview says what it became. `python3 00-META/checks/cycle.py`
|
||||
|
||||
+1
-23
@@ -14,8 +14,7 @@ What is enforced:
|
||||
its owning code (`code:`) -- no development without a design that says where.
|
||||
issues a known `status:`; once `located`, `located-in:` names the owner;
|
||||
once `resolved`, `fixed-by:` says what fixed it (prose counts --
|
||||
"nothing, the capability existed" is an answer). And no two records share a
|
||||
number -- the number is how a record is cited.
|
||||
"nothing, the capability existed" is an answer).
|
||||
research a known `status:`; a `graduated` overview says what it `became:`, and every
|
||||
target it names exists.
|
||||
decisions every accepted record is REACHABLE from the cycle: cited by a design doc's
|
||||
@@ -110,27 +109,6 @@ def main():
|
||||
"without a design that says where" % status)
|
||||
|
||||
# ---- issues ------------------------------------------------------------------------
|
||||
# Two records may not share a number. Numbers are taken as "next free after main", and work
|
||||
# sits on unmerged branches for days -- so two people reading the same main allocate the same
|
||||
# number, and nothing said so. It happened twice in one evening between two machines, and the
|
||||
# second collision landed on main with all three checks passing (issue 155). An issue number is
|
||||
# how every other record cites this one; two records answering to it means a pointer that
|
||||
# resolves to whichever the reader happened to open.
|
||||
seen = {}
|
||||
for folder in sorted(glob.glob(os.path.join(ROOT, "04-ISSUES", "*", ""))):
|
||||
name = os.path.basename(os.path.normpath(folder))
|
||||
number = name.split("-", 1)[0]
|
||||
if not number.isdigit():
|
||||
continue
|
||||
if number in seen:
|
||||
bad(os.path.join("04-ISSUES", name),
|
||||
"is numbered %s, and so is %s -- an issue number is how it is cited, and two "
|
||||
"records answering to one means a citation that resolves to whichever the reader "
|
||||
"opened. Take the next free number across main AND every open pull request"
|
||||
% (number, seen[number]))
|
||||
else:
|
||||
seen[number] = name
|
||||
|
||||
for path in sorted(glob.glob(os.path.join(ROOT, "04-ISSUES", "*", "00-report.md"))):
|
||||
front = frontmatter(path)
|
||||
if front is None:
|
||||
|
||||
@@ -21,12 +21,7 @@ incident someone must **clear**.
|
||||
|
||||
## Steps
|
||||
|
||||
1. Take the next free number — **across `main` and every open pull request**, not `main` alone.
|
||||
Work sits on unmerged branches for days, so two people both reading `main` allocate the same
|
||||
number; it happened twice in one hour between two machines, and the second collision reached
|
||||
`main` with every check passing (issue 155). `cycle.py` now refuses two records sharing a number,
|
||||
which catches a collision but does not prevent one. Create
|
||||
`04-ISSUES/NNN-short-name/00-report.md`:
|
||||
1. Take the next free number. Create `04-ISSUES/NNN-short-name/00-report.md`:
|
||||
|
||||
```yaml
|
||||
---
|
||||
@@ -47,14 +42,6 @@ incident someone must **clear**.
|
||||
## Rules
|
||||
|
||||
- Closed issues are never deleted — they are the mesh's symptom-to-component memory.
|
||||
- `fixed-by:` names something that will still exist: a commit or a pull request, never a branch. A
|
||||
branch is deleted when it merges, so a branch name there is a pointer that resolves to nothing by
|
||||
the time anybody follows it.
|
||||
- A fix that turns out to have broken something else is written back into the record that asked for
|
||||
it, pointing at the new issue. Somebody arriving at a record to learn why the code is the way it
|
||||
is must not have to already know there was a sequel.
|
||||
- Renumbering a collision happens once, in the branch that lands last. Renumbering a branch whose
|
||||
author is still pushing only moves the race.
|
||||
- An issue whose answer is a general lesson should also be written to the knowledge base, so
|
||||
the next person searching a symptom finds it. Both, not either.
|
||||
- `status: wontfix` is legitimate and requires a sentence saying why.
|
||||
|
||||
@@ -13,14 +13,6 @@ decisions taken over three days; the reasoning is kept, the fragmentation is not
|
||||
|
||||
The environment a change is run against before it reaches real machines.
|
||||
|
||||
> **Still the lab, no longer the test bed — 2026-09-30, by [ADR 0149](0149-the-live-mesh-is-the-test-bed.md).**
|
||||
> Everything here stands. What changed is what the lab is *for*: a change is verified against the mesh
|
||||
> that is running, because the faults that cost the most are faults of a mesh that already exists —
|
||||
> bound consumers, containers made against an older roster, an adopted machine — and a bed is by
|
||||
> construction a mesh that does not. Raising a mesh from bare is now the lab's whole job, which is the
|
||||
> one thing the live mesh cannot be asked to do. 0149 also supersedes
|
||||
> [ADR 0068](0068-the-lab-takes-requests.md), which extended this one and was never built.
|
||||
|
||||
## A node in the lab is a virtual machine
|
||||
|
||||
It boots a stock Linux image, runs the real install, and becomes a node. **It is not a model of
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
topic: building it
|
||||
status: accepted
|
||||
status: proposed
|
||||
date: 2026-09-01
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
@@ -101,19 +101,3 @@ the digest down after building.
|
||||
**Whether kind 4 deserves a module at all.** Thirty-five descriptions that say *install this and
|
||||
write these files* may be better as one module with settings than as thirty-five modules. Left
|
||||
open deliberately; it is a question about the shape of the catalogue, not about whether to have one.
|
||||
|
||||
## Accepted, 2026-09-29, against what was built
|
||||
|
||||
*Marked in a grooming pass: the mesh was built to this record and the record still said `proposed`.*
|
||||
|
||||
The proposal is the arrangement that exists. `mesh-catalog` holds descriptions of software we did
|
||||
not write and the programs that provision it, and holds neither the mesh's own components nor an
|
||||
application's own module. The mesh's list of modules is a table in the control plane, filled by
|
||||
`module add`, and every module records the source it came from with the commit it was read at.
|
||||
|
||||
**One half is not built: `module check` as a command on the control plane's binary.** A manifest is
|
||||
still validated by a test that reaches into the control plane's internals — which works for this
|
||||
catalogue and gives nothing at all to somebody describing their own application in their own
|
||||
repository, which this record says is the case that matters most. That is
|
||||
[issue 148](../04-ISSUES/148-a-manifest-outside-this-catalogue-has-no-check/00-report.md).
|
||||
|
||||
|
||||
@@ -26,14 +26,6 @@ runtime is per-module, not per-node, and treating the audit-logger as special le
|
||||
modules' code with nothing to run it: the conversion produced tools and events that, as it stands,
|
||||
never execute.
|
||||
|
||||
> **The hosting form is settled elsewhere — 2026-09-30.** Where this record says "a container", read
|
||||
> [ADR 0150](0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md): a module's own
|
||||
> code runs as supervised processes under this record's one account. Nothing else here changes — the
|
||||
> per-module runtime, the per-tool key and the single scoped account are the argument this record made
|
||||
> and they are why 0150 goes the way it does. The note is here because two design documents chose the
|
||||
> other form without knowing this record existed
|
||||
> ([issue 117](../04-ISSUES/117-a-modules-own-code-is-a-container-and-a-process/00-report.md)).
|
||||
|
||||
## Decision
|
||||
|
||||
### A module with tools or events runs a process of its own
|
||||
|
||||
@@ -80,15 +80,6 @@ reaching the routed name, which the clause above has just made resolvable inside
|
||||
three are one decision: **compose the name, propagate it, certify it** — each is meaningless without
|
||||
the one before it.
|
||||
|
||||
> **The mechanism changed — 2026-09-30, by [ADR 0148](0148-the-meshs-names-are-resolved-not-copied-into-containers.md).**
|
||||
> A routed name still reaches every asker in the mesh, which is what this record decided and it stands.
|
||||
> It no longer reaches them by being written into each declared container: copying the roster in made the
|
||||
> roster part of every container's identity, so one name moving replaced every container in the mesh
|
||||
> ([issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). A
|
||||
> container resolves through its machine's resolver instead. The consequence below — that an internal
|
||||
> issuer's challenge needs the routed name resolvable inside the mesh — holds unchanged, by the means the
|
||||
> machine itself already uses.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Lab-versus-production is one node-level `public-domain` setting**, not an override on every
|
||||
|
||||
@@ -1,21 +1,14 @@
|
||||
---
|
||||
topic: building it
|
||||
status: superseded
|
||||
status: proposed
|
||||
date: 2026-09-12
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 0016-the-lab.md
|
||||
superseded-by: 02-DECISIONS/0149-the-live-mesh-is-the-test-bed.md
|
||||
---
|
||||
|
||||
# 68. The lab takes requests, one at a time, and runs each from its own copy
|
||||
|
||||
> **Superseded — 2026-09-30, by [ADR 0149](0149-the-live-mesh-is-the-test-bed.md).** Never built. The
|
||||
> live mesh became the test bed, because the faults that cost the most are faults of a mesh that
|
||||
> already exists — bound consumers, containers made against an older roster, an adopted machine — and a
|
||||
> bed is by construction a mesh that does not. The rule worth keeping from below is that a run reads a
|
||||
> copy that is not anybody's working tree.
|
||||
|
||||
## Context
|
||||
|
||||
**The lab is exclusive hardware, and today a person holds it.** Raising a scenario takes over
|
||||
|
||||
@@ -35,16 +35,6 @@ The development cycle is enforced mechanically, to the extent frontmatter can ca
|
||||
capability existed" is an answer).
|
||||
- **No silent graduation** — a `graduated` research overview says what it `became:`, and the
|
||||
targets exist.
|
||||
- **No two records answering to one number** — added 2026-09-30; see the insight below.
|
||||
|
||||
> **Progressive insight — 2026-09-30.** The list above named four things `cycle.py` enforces, and
|
||||
> now names five. Nothing enforced that two issue records hold different numbers: two machines
|
||||
> filing issues within one hour both read `main`, both took "the next free number", and collided
|
||||
> twice — the second collision reaching `main` with `records.py`, `cycle.py` and `index.py` all
|
||||
> reporting success (04-ISSUES/155). A number is how every other record cites one, so two records
|
||||
> answering to it is a citation that resolves to whichever folder the reader opened. `cycle.py`
|
||||
> refuses it now. The decision here stands exactly as written: this is one more thing frontmatter
|
||||
> and file names can carry, found by its absence rather than by reasoning.
|
||||
|
||||
[`00-META/checks/cycle.py`](../00-META/checks/cycle.py) refuses violations, beside `records.py`
|
||||
and `index.py`; all three run before any merge here. What frontmatter cannot see — that code work
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
topic: what runs on it
|
||||
status: accepted
|
||||
status: proposed
|
||||
date: 2026-09-25
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
@@ -275,11 +275,3 @@ On acceptance, each of these is superseded or amended by this record, not edited
|
||||
everything a module needs is a requirement
|
||||
- [Issue 095](../04-ISSUES/095-a-module-assigned-after-genesis-has-no-broker-account/00-report.md),
|
||||
[issue 103](../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md): what fails today
|
||||
|
||||
## Accepted, 2026-09-29, against what was built
|
||||
|
||||
*Marked in a grooming pass.* The vault is a module providing `secret` at mesh scope, and six
|
||||
modules in the catalogue require it — so a shared secret is a requirement answered by the vault,
|
||||
which is what this record asks for. Private keys are still made where they are used and never
|
||||
travel, which is the other half and was never in question.
|
||||
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
topic: what runs on it
|
||||
status: accepted
|
||||
status: proposed
|
||||
date: 2026-09-26
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
@@ -189,23 +189,6 @@ On acceptance, each of these is amended by this record, not edited:
|
||||
credential is no longer all-or-nothing with a window. It overlaps, with each step confirmed.
|
||||
A single-party credential is staged, not replaced.
|
||||
|
||||
## Accepted, 2026-09-30, and not scheduled
|
||||
|
||||
Accepted as written. The separation it draws — a consumer's *resource* and a *credential that reaches
|
||||
it* are different things with different lifecycles — is the part that had to be settled, because the
|
||||
alternative is what the record was written against: retiring a credential taking the data it reached
|
||||
with it. That is a data-loss shape, and a record that names it should not sit unresolved while the
|
||||
code that could hit it is being written.
|
||||
|
||||
**It is not built, and accepting it does not schedule it.** The SDK's provisioner adapter is still
|
||||
`create` / `remove` / `holds` rather than the four operations above, and no provider implements the
|
||||
two-credential rotation. Accepted-and-not-built is an ordinary state here — 0141 and 0142 are both in
|
||||
it — and it is the honest one: leaving this `proposed` made it invisible to anyone reading what the
|
||||
mesh has decided, while changing nothing about what runs.
|
||||
|
||||
The work it implies belongs with the provisioner contract, beside
|
||||
[issue 124](../04-ISSUES/124-a-consumer-cannot-be-told-what-its-provider-derived/00-report.md).
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Every credential provider's adapter changes**, in two steps. The first separates *retire a
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
topic: what runs on it
|
||||
status: accepted
|
||||
status: proposed
|
||||
date: 2026-09-26
|
||||
deciders: jochen
|
||||
extends: 0112-a-module-definition-names-no-node-mesh-or-path.md
|
||||
@@ -47,10 +47,3 @@ other boundary already is: the module name.
|
||||
node runs one of each (ADR 0115)" — instead of failing on whichever name collides first.
|
||||
- Multi-tenant asks are answered in the catalogue (a second module definition), not in the
|
||||
control plane.
|
||||
|
||||
## Accepted, 2026-09-29, against what was built
|
||||
|
||||
*Marked in a grooming pass.* The rule is enforced where it cannot be forgotten: `assignment`'s
|
||||
primary key is `(node, module)`, so a second assignment of one module to one machine is not a thing
|
||||
the mesh can hold. The record read `proposed` while the schema had already settled it.
|
||||
|
||||
|
||||
@@ -12,7 +12,7 @@ extends: 0102-the-mesh-writes-into-a-shared-file-never-over-it.md
|
||||
## Context
|
||||
|
||||
When a resource stops being declared — its module unassigned, the node sent a
|
||||
deliberately-empty declaration ([issue 149](../04-ISSUES/149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)),
|
||||
deliberately-empty declaration ([issue 127](../04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)),
|
||||
or a new catalogue version renaming its id — the host undoes it. The host's own code states
|
||||
the rule it means to follow: **it removes what it made and leaves what it merely configured.**
|
||||
For almost every resource it does exactly that:
|
||||
|
||||
@@ -109,23 +109,6 @@ understated as none. The remaining work is a way to build the host and a way to
|
||||
path, and until both exist nothing delivers a version and every machine takes the fallback — which is
|
||||
what every machine does today.
|
||||
|
||||
> **Progressive insight — 2026-09-30. Both of those exist now.** The paragraph above named two missing
|
||||
> things and they are built: a Go toolchain, based on a new `mesh-tools-go` module so the compiler is
|
||||
> named and not pinned, and `${version}` in any value of a resource that uses an archive or a bundle.
|
||||
> The mesh compiles its own host and publishes it to its own registry, measured — a statically linked
|
||||
> stripped binary, fetched back out and run. **The cost was larger again than this note said**: three
|
||||
> more things in the path assumed one language or one shape, and a fourth was in the base image.
|
||||
> The account is [issue 142](../04-ISSUES/142-the-host-is-the-one-thing-the-mesh-does-not-deliver/01-progress.md).
|
||||
>
|
||||
> The version in a path is the artifact's **digest**, not the commit this note's own wording would
|
||||
> suggest. Two builds of one commit are meant to be the same bytes, so a content-addressed version
|
||||
> means an unchanged build keeps the path it had; a commit-named one would move for an identical binary
|
||||
> and recreate everything reading it.
|
||||
>
|
||||
> **Still nothing delivers a version to a machine.** The host is a module and builds, and declares no
|
||||
> resources, so the bundle sits in the registry and no machine is asked to take it. That is the next
|
||||
> piece, and the decision above is unchanged by any of this.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **The host becomes a build target and a module** — a module whose resource is the next host, applied
|
||||
|
||||
@@ -113,36 +113,6 @@ the authority it was bound to, checked in the control plane's own test suite —
|
||||
that verified anything. That is a weaker thing than the paragraph above describes, and it stays
|
||||
written this way until the bed runs.
|
||||
|
||||
> **Progressive insight — 2026-09-30. It has now been run, on the live mesh rather than in the bed.**
|
||||
> The paragraph above said nothing had verified anything, and something has. The module was registered
|
||||
> from the catalogue, assigned to a workstation, and checked in the form this section prescribes — the
|
||||
> authority's own API, so the handshake needs nothing else in the mesh to be right:
|
||||
>
|
||||
> ```
|
||||
> $ curl -sS -o /dev/null -w '%{http_code}' https://<the authority>:9000/health
|
||||
> 200
|
||||
> subject=CN=Step Online CA
|
||||
> issuer=O=Mesh Internal CA, CN=Mesh Internal CA Intermediate CA
|
||||
> Verify return code: 0 (ok)
|
||||
> ```
|
||||
>
|
||||
> **Both halves.** Unassigning and pushing removed the anchor, emptied the trust store of the mesh's
|
||||
> authority, and returned the plain client to *unable to get local issuer certificate* — then assigning
|
||||
> again restored it. The negative half is what distinguishes the anchor working from something else
|
||||
> having trusted it, and it is the half nothing had ever exercised.
|
||||
>
|
||||
> **One thing this found that is not in the module.** The removal only works because the *host* removes
|
||||
> the service before the script: stopping the unit is what deletes the certificate and refreshes the
|
||||
> bundles, and it needs the script it calls to still exist. Nothing in the module states that ordering;
|
||||
> the symmetry this record claims rests on it.
|
||||
>
|
||||
> Run on the live mesh because that is where a change is verified now
|
||||
> ([ADR 0149](0149-the-live-mesh-is-the-test-bed.md)), and the bed still cannot raise a foundation. The
|
||||
> evidence is [issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/02-resolution.md).
|
||||
> Extended the same day to every converged machine — `novox`, `g14` and `shanks` each hold the anchor
|
||||
> and verify with a plain client. `ace` is excluded on purpose: it is adopted, so a module assigned
|
||||
> there is held rather than run, which is right and is not trust.
|
||||
|
||||
## Consequences
|
||||
|
||||
The predecessor's authority can be retired from a machine once this module is assigned to it,
|
||||
|
||||
@@ -1,176 +0,0 @@
|
||||
---
|
||||
topic: the tiers
|
||||
status: accepted
|
||||
date: 2026-09-30
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 02-DECISIONS/0066-public-routing-is-name-agnostic.md
|
||||
---
|
||||
|
||||
# 148. The mesh's names are resolved, not copied into every container
|
||||
|
||||
## Context
|
||||
|
||||
The mesh gives every container it declares the whole roster of mesh names as entries written into
|
||||
the container's own hosts file at creation
|
||||
([design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md),
|
||||
[ADR 0066](0066-public-routing-is-name-agnostic.md)). A container takes those entries once and never
|
||||
looks again.
|
||||
|
||||
Three issues are the same fact arriving three times.
|
||||
|
||||
**A container keeps the address it was made with.** Adopting the predecessor's tunnel moved the hub's
|
||||
private address; the declaration followed it within one push and nothing on the machine did. The
|
||||
forge's container held the old address, lost its database, reported healthy while its existing
|
||||
connections lasted, and then the public name went down
|
||||
([issue 109](../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md)).
|
||||
|
||||
**The same fault, four days later, undetected for five days.** One container had restarted 2286 times
|
||||
against a database it could no longer find, while the mesh reported the machine as doing what it was
|
||||
told. Inside it, `novox.internal` was an address that had not existed for five days
|
||||
([issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md)). Forty-eight
|
||||
other containers were current, none of them corrected — each had been recreated for some other
|
||||
reason and picked up the roster on the way.
|
||||
|
||||
135 was fixed by putting the roster into the digest the host compares a container against, so a
|
||||
container whose names moved is recreated like one whose image moved. **That made the roster part of
|
||||
every container's identity**, which is the third arrival:
|
||||
|
||||
**One name moving replaces every container in the mesh.** Migrating one small module on one machine
|
||||
took four routine actions; each changed the roster, and each replaced every container on the control
|
||||
node — its own store, the registry, the edge proxy, the forge, the directory, mail. The control plane
|
||||
was unreachable twice while its own store came back through crash recovery
|
||||
([issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). None of
|
||||
the replaced containers had anything to do with the module being migrated, or with its machine.
|
||||
|
||||
The blast radius of a name is now every container that carries the list, which is all of them. The
|
||||
node-by-node migration ahead adds names one module at a time — on one machine alone that is around
|
||||
twenty-five — and each would be a full restart of every service on the hub.
|
||||
|
||||
## Considered Options
|
||||
|
||||
**1. Keep the roster in every container and accept the churn.** Rejected. It is not a cost that can
|
||||
be paid down: the mesh gets more names as it grows, and every name costs a restart of everything.
|
||||
A rollback costs another.
|
||||
|
||||
**2. Scope each container's entries to the names it actually binds.** A container is given the names
|
||||
of the things it declared a requirement on, so a name's blast radius is its consumers. Tidy, needs no
|
||||
new mechanism, and keeps 135's guarantee exactly.
|
||||
|
||||
Rejected, and this is the close one. It contradicts the standing intent that **anything on the mesh
|
||||
can call anything on it** — three cases, same machine, the private network, the public network, and
|
||||
no fourth. Scoping resolution to declared couplings makes a name reachable only where the mesh was
|
||||
told in advance that it would be wanted, and a person debugging inside a container would find names
|
||||
missing that exist everywhere else on the machine. It also leaves the roster in the digest, so the
|
||||
churn returns the moment a widely-bound name moves — smaller, not gone.
|
||||
|
||||
**3. Resolve at lookup time through the machine's resolver, and copy nothing.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**A container resolves the mesh's names through its machine's resolver, at the moment it asks. No
|
||||
mesh name and no mesh address is written into a container, and none is part of a container's
|
||||
identity.**
|
||||
|
||||
The three consequences that make this worth doing:
|
||||
|
||||
- **Staleness stops being possible**, rather than being detected. 109 and 135 are not bugs that were
|
||||
fixed; they are a shape that no longer exists. A name that moves is answered differently by the next
|
||||
lookup, in every container, with nothing recreated and nothing restarted.
|
||||
- **A name's blast radius becomes nothing.** Assigning a module on one machine does not touch a
|
||||
container on another.
|
||||
- **Anything can still call anything**, which option 2 gave up. The resolver answers every mesh name to
|
||||
every asker on the machine, exactly as it answers the machine itself.
|
||||
|
||||
**The resolver is a machine-level process, not a container** — one of the modules that is not a
|
||||
container at all — so a container depending on it is not the circularity it would be if the mesh's
|
||||
own store had to resolve a name through something the store's own runtime had to start first.
|
||||
|
||||
**What a module declares for itself is untouched.** Entries a manifest asks for are the module's own,
|
||||
stay in the container, and stay in its identity: they are part of what the module *is*, they do not
|
||||
move when the mesh's roster does, and the mesh does not know what they mean.
|
||||
|
||||
**The machine's own roster file is untouched.** It is a file, rewritten in place, read by processes and
|
||||
people; nothing restarts when it changes. It is only the *copy into each container* that this ends.
|
||||
|
||||
### The order this lands in, which is not a preference
|
||||
|
||||
**Nothing may stop copying names until resolution works from a container.** Removing the copy first
|
||||
reintroduces 109 and 135 — silently, and on a live mesh, which is exactly how both were found.
|
||||
|
||||
1. **A container on any network can reach the resolver.** Today a container on the runtime's default
|
||||
network asks from an address the converged filter drops, so it has no DNS at all
|
||||
([issue 110](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md));
|
||||
and on two of four machines the resolver binds loopback only, so the runtime hands containers a
|
||||
public resolver instead. Both are prerequisites, not related work.
|
||||
|
||||
> **Progressive insight — 2026-09-30, later the same day. The loopback claim was wrong.** The
|
||||
> resolver bound the private address on all four machines; on two the runtime had never been told
|
||||
> to use it, and on all four the resolver discarded a query that arrived on the runtime's bridge.
|
||||
> The step stands; the facts under it were those. Both fixed the same day
|
||||
> ([issue 110's resolution](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/01-resolution.md)),
|
||||
> and step 3 landed after them.
|
||||
2. **The runtime is told which resolver to use, per machine, as a file** — not per container as a
|
||||
creation-time argument, or the resolver's address is back in every container's identity and the
|
||||
problem has only got smaller.
|
||||
3. **Then, and only then, the roster leaves the declaration and the digest.**
|
||||
|
||||
Until step 3 the mesh keeps copying, and keeps comparing. 151 stays open until step 3 lands; it is not
|
||||
closed by this record, only answered by it.
|
||||
|
||||
## How this is checked
|
||||
|
||||
- **A container resolves a name that moved, without being recreated.** Move a name the mesh serves;
|
||||
from a container that was running before the move and has not been touched since, the name answers
|
||||
with the new address. This is the one 109 and 135 would both have failed.
|
||||
- **A name's blast radius is nothing.** Add a routed name on one machine; no container on any other
|
||||
machine is recreated. The apply report on each machine says nothing changed. This is 151.
|
||||
- **Anything calls anything.** From a container on any machine, every `<node>.internal` name and every
|
||||
routed name the mesh serves resolves — including names the module never declared a requirement on,
|
||||
which is the guarantee option 2 would have given up.
|
||||
- **On every network the runtime offers.** The first three hold for a container on the runtime's
|
||||
default network as well as one on a declared network, because the default network is the case that
|
||||
has no DNS today.
|
||||
- **No mesh name is in a container's spec.** A test asserts the digest a host computes for a container
|
||||
does not move when the mesh's roster does, and does move when the module's own declared entries do.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **The resolver becomes load-bearing for every container**, where before it was load-bearing for the
|
||||
machine. This is a real cost and is accepted: a resolver that is down is a machine that cannot
|
||||
resolve, which is already true of the machine itself, and is a smaller event than a roster change
|
||||
destroying and recreating every container on the machine.
|
||||
- **Design 08's "a file rather than a resolver" no longer describes containers.** It was written when
|
||||
the mesh had no resolver and it gave the right answer then. The reasoning it rested on — every Linux
|
||||
has a hosts file, no package needed — was already overtaken by names a hosts file cannot express:
|
||||
service names and wildcards under `<node>.internal`, which is why the resolver was built.
|
||||
- **ADR 0066's mesh-wide propagation is kept and its mechanism changes.** A routed name still reaches
|
||||
every asker in the mesh; it reaches them through the resolver rather than by being written into each
|
||||
container. The consequence 0066 records — that an internal issuer's challenge needs the routed name
|
||||
resolvable inside the mesh — holds unchanged and by the same means the machine already uses.
|
||||
- **Issue 110 stops being a container-DNS inconvenience and becomes a prerequisite** for the mesh not
|
||||
restarting itself whenever it learns a name.
|
||||
- **A container that names a resolver of its own has opted out of the machine's**, and the copy this
|
||||
record removes was the only reason such a container could reach anything by a mesh name.
|
||||
|
||||
> **Progressive insight — 2026-09-30, the afternoon this landed. Found the hard way.** The mail
|
||||
> system's admin, behind Mailu's own resolver, lost its database the moment the copy went
|
||||
> ([issue 171](../04-ISSUES/171-a-modules-own-resolver-knows-no-mesh-name/00-report.md)). A `dns` on
|
||||
> a container is a decision about whether mesh names exist inside it, not a preference; the module
|
||||
> was corrected, and whether the controller should refuse the contradiction is that issue's open
|
||||
> question.
|
||||
|
||||
- **A container started by hand gets the mesh's names too**, where before only declared containers did.
|
||||
Design 08 drew that boundary deliberately, on the grounds that reaching into every container is what
|
||||
a nameserver would be for. This record accepts that consequence rather than working around it: a
|
||||
person debugging in a hand-started container resolving the same names as everything else is the
|
||||
behaviour worth having, and it is what "anything can call anything" means.
|
||||
|
||||
## References
|
||||
|
||||
- [issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md) — one name replaces every container; the question this answers
|
||||
- [issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md) — the roster put into the digest
|
||||
- [issue 109](../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md) — the first arrival
|
||||
- [issue 110](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md) — the prerequisite
|
||||
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — routed names propagate mesh-wide; extended here
|
||||
- [design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md) — the file-not-resolver reasoning this narrows
|
||||
@@ -1,106 +0,0 @@
|
||||
---
|
||||
topic: building it
|
||||
status: accepted
|
||||
date: 2026-09-30
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
supersedes: 02-DECISIONS/0068-the-lab-takes-requests.md
|
||||
---
|
||||
|
||||
# 149. The live mesh is the test bed
|
||||
|
||||
## Context
|
||||
|
||||
[ADR 0068](0068-the-lab-takes-requests.md) proposed that the lab accept queued requests
|
||||
— a bed and a commit — answer them one at a time from a copy it owns, and expose that through tools so
|
||||
an agent could start a run and come back to it. It has been `proposed` since 2026-09-12 and nothing was
|
||||
built.
|
||||
|
||||
What happened instead is that the mesh became the thing under test. It runs on four machines; every
|
||||
fault worth finding in the last month was found on them, and none was found in a bed:
|
||||
|
||||
- a container holding an address that had not existed for five days, on the control node
|
||||
([issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md));
|
||||
- a machine reading healthy for eleven hours while no module could reach another
|
||||
([issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md));
|
||||
- one name replacing every container on the hub
|
||||
([issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md));
|
||||
- a consumer assertion that is correct on a mesh being raised and fatal on one that is running
|
||||
([issue 156](../04-ISSUES/156-moving-a-consumers-delivery-subject-stops-the-control-plane/00-report.md)).
|
||||
|
||||
The last is the one that settles it. That change was exercised on the raise path — which is what a bed
|
||||
*is* — and the raise path is the only path on which the fault cannot appear. A bed raises a mesh; it
|
||||
does not have a mesh that has been running for weeks, with consumers already bound, containers created
|
||||
against an older roster, and an adopted machine carrying a predecessor's configuration. **The faults
|
||||
that cost the most were all faults of a mesh that already exists**, and a bed is by construction a mesh
|
||||
that does not.
|
||||
|
||||
Lab runs are also expensive in a way that changed the behaviour around them: each costs a build and
|
||||
several minutes, so they were batched, and a batched test is one whose result arrives after the next
|
||||
three changes were already written.
|
||||
|
||||
## Considered Options
|
||||
|
||||
**1. Build 0068 as proposed.** Rejected. It answers a question nobody is asking: the bottleneck was
|
||||
never that a person had to sit at the lab, it was that a bed cannot hold the state the faults live in.
|
||||
Queueing and tooling a mechanism that finds the wrong class of fault faster is not an improvement.
|
||||
|
||||
**2. Leave 0068 `proposed`.** Rejected, and it is why this record exists rather than nothing. A record
|
||||
that contradicts current practice and sits unresolved is worse than either answer: it reads as intent
|
||||
to anyone who finds it, and the practice it contradicts is written down nowhere but a handoff note.
|
||||
|
||||
**3. Record that the live mesh is the test bed, and supersede 0068.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**A change is verified against the mesh that is running.** Not because a bed would be unwelcome, but
|
||||
because the state that breaks things is state a bed does not have: containers made against an older
|
||||
roster, consumers already bound, an adopted machine, a store with weeks of history.
|
||||
|
||||
**A change that can only be exercised on the raise path is not verified.** If the only test available
|
||||
raises a fresh mesh, the record says so, and says which case was therefore not covered. The words
|
||||
"exercised on a fresh mesh" are a statement about coverage, not a pass.
|
||||
|
||||
**The lab is not retired**, and [ADR 0016](0016-the-lab.md) stands. It remains the place to raise a
|
||||
mesh from bare, which is the one thing the live mesh cannot be asked to do and the one thing a bed does
|
||||
better than anything else. What this record removes is the lab as the *default* answer to "is this
|
||||
change good", and with it 0068's queue, tools and request protocol.
|
||||
|
||||
**Accuracy over a green run.** An honest failure on the live mesh beats a pass in a bed that could not
|
||||
have failed — and a change that is risky on the running mesh is a reason to make the change smaller,
|
||||
not a reason to test it somewhere it cannot break.
|
||||
|
||||
**What 0068 got right is kept as a rule, not a mechanism:** a run reads a copy that is not anybody's
|
||||
working tree. Every run of the lab that mattered was pinned to a checkout rather than a worktree, and
|
||||
the ones that were not produced results about code nobody had written down.
|
||||
|
||||
## How this is checked
|
||||
|
||||
- **A record that says a change was verified says on what.** Where it was a fresh mesh, it says which
|
||||
case is uncovered. This is the clause that would have caught 156: its change was verified, honestly,
|
||||
on the only path where it works.
|
||||
- **The lab is not in the path of a merge.** No check, playbook or handoff requires a bed to have run.
|
||||
- **0068 is unreachable as intent.** Its status is `superseded` and it names this record, so a reader
|
||||
arriving at the queue design finds out immediately that it was not built and why.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **A fault can be introduced on the machines that serve.** This is the cost, it is real, and it was
|
||||
paid twice in one evening — a control plane crash-looping for half an hour, and every container on the
|
||||
hub recreated five times. Both were found in minutes because they were live, and both would have
|
||||
passed a bed.
|
||||
- **There is no pre-merge gate beyond the repositories' own suites.** `make check` and the three hq
|
||||
checks are what stands between a change and the machines, which raises what those suites are worth
|
||||
and makes a test that cannot fail a genuine defect rather than an untidiness.
|
||||
- **Raising a mesh from bare is now the lab's whole job**, and is exercised deliberately rather than
|
||||
as a side effect of testing something else. The foundation work
|
||||
([issue 146](../04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md))
|
||||
is that job, and it is also the proof that the mesh can make another of itself.
|
||||
- **An agent cannot hand a run to a queue and come back**, which 0068 would have given. In practice it
|
||||
watches a push and reads the machines, which is what happened anyway.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR 0068](0068-the-lab-takes-requests.md) — superseded by this
|
||||
- [ADR 0016](0016-the-lab.md) — the lab, which stands
|
||||
- [issue 156](../04-ISSUES/156-moving-a-consumers-delivery-subject-stops-the-control-plane/00-report.md) — correct on the raise path, fatal on a running mesh
|
||||
-113
@@ -1,113 +0,0 @@
|
||||
---
|
||||
topic: what runs on it
|
||||
status: accepted
|
||||
date: 2026-09-30
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 02-DECISIONS/0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md
|
||||
---
|
||||
|
||||
# 150. A module's own code runs as supervised processes under the module's one account
|
||||
|
||||
## Context
|
||||
|
||||
The repository answers "what runs a module's own code" two ways and reconciles them nowhere
|
||||
([issue 117](../04-ISSUES/117-a-modules-own-code-is-a-container-and-a-process/00-report.md)).
|
||||
|
||||
[ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) is accepted and
|
||||
says **a container** — "the tool runtime carrying that module's compiled code" — and "one module, one
|
||||
process, one account". Two `proposed` design documents say a **`process`** resource running an argv,
|
||||
supervised by the machine, and one of them declares *four* of them for a single module and presents
|
||||
four as the point. Neither design document names 0047 in its `decisions:`, and the string `process`
|
||||
as a resource type appears in no decision record at all. The thing as built is the container.
|
||||
|
||||
Two things have happened since 0047 was written that bear on it directly.
|
||||
|
||||
[**ADR 0142**](0142-the-mesh-delivers-its-own-components-as-binaries.md) decided that the mesh's own
|
||||
components are binaries on the machine rather than container images, and
|
||||
[issue 114](../04-ISSUES/114-should-the-controller-be-a-container-or-a-process/00-report.md) was closed
|
||||
by it. That settled the mesh's components and deliberately said nothing about a module's.
|
||||
|
||||
And the standing definition of a module hardened: **a module is software that delivers one or more
|
||||
services, and a module is not a container.** It may deliver them as a container, an installed package
|
||||
with a unit, a binary, or configuration files; 61 of 73 happen to use a container and 11 do not,
|
||||
including the resolver, sshd and fail2ban. A rule that a module's *own code* must be a container makes
|
||||
the one kind of module the mesh writes itself the only kind that has no choice.
|
||||
|
||||
## Considered Options
|
||||
|
||||
**1. Hold 0047 as written: a container.** Rejected. Its own reasoning does not require one. What 0047
|
||||
argued for was a runtime **per module** rather than one for the whole node, because a node-wide runtime
|
||||
could not hold a per-module broker account and per-module runtimes competing on one tool key would each
|
||||
be handed calls for tools they do not have. A supervised unit per module satisfies that argument
|
||||
exactly — it is per module, and a unit runs as an account. The container was the mechanism to hand, not
|
||||
the conclusion.
|
||||
|
||||
**2. Let each design document choose.** Rejected; that is the present state and it is what issue 117
|
||||
reports. A module author reading the guide writes four processes; a module author reading the record
|
||||
writes a container; nothing tells either that the other exists.
|
||||
|
||||
**3. Settle the hosting form as a supervised process, and settle the count separately.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**A module's own code runs as one or more supervised processes on the machine, under the module's single
|
||||
account.** Where [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) says
|
||||
"a container, the tool runtime carrying that module's compiled code", read this record. Everything else
|
||||
0047 decided stands untouched: a tool is served on its own key, only the module that serves it answers,
|
||||
and the module's account is scoped to exactly its tool keys.
|
||||
|
||||
**The invariant is the account, not the process count.** 0047's "one module, one process, one account"
|
||||
carried its weight in the last clause. Its stated worry about a second process was "not a second one to
|
||||
scope and seal" — a second *identity* to grant, seal a secret to, and scope on the bus. Several
|
||||
processes sharing the module's one account create no second identity, so nothing further is scoped or
|
||||
sealed, and a module may therefore declare as many as its work has shapes: events, tools, a
|
||||
provisioner, a scheduled ingest. **What a module may not have is two accounts.**
|
||||
|
||||
**A module that delivers its service as a container still does.** This record is about the code the
|
||||
module itself carries — its tools, its events, its provisioner — and not about the software it delivers.
|
||||
A module wrapping a third-party image wraps a third-party image.
|
||||
|
||||
**Why supervised by the machine rather than by the mesh:** it is the same answer ADR 0142 gave for the
|
||||
mesh's own components, for the same reason. A unit the machine restarts needs no image, no registry
|
||||
pull and no runtime to be up before the mesh's own code can run — which matters most for exactly the
|
||||
modules whose code the mesh cannot start any other way.
|
||||
|
||||
## How this is checked
|
||||
|
||||
- **No design document describes a hosting form for a module's own code without citing this record.**
|
||||
Designs 18 and 20 name it in `decisions:`; this is the gap issue 117's third point reports, and
|
||||
`cycle.py` already enforces that a to-be design names its decisions.
|
||||
- **A module declaring several processes resolves to one account.** A test composes a module with more
|
||||
than one process resource and asserts the mesh mints exactly one broker account for it, scoped to that
|
||||
module's tool keys and nothing else — which is 0047's invariant stated as an assertion rather than a
|
||||
sentence.
|
||||
- **A module's own code does not require the container runtime.** A machine with no container runtime
|
||||
can still run a module whose code is its own, which is the claim that separates this from option 1 and
|
||||
is checkable on a machine that has one by asserting the declaration names no image for it.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **The sidecar port stops being needed.** [ADR 0029](0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md)
|
||||
records that "anything that is a service plus a sidecar currently has to publish a port to talk to
|
||||
itself", and the host's `network` shape exists partly for it. A process beside the service on the same
|
||||
machine reaches it without publishing anything, so that pressure goes.
|
||||
- **Something must supervise, and it is the machine.** This adds a unit per module's code to what the
|
||||
host writes and owns. The mesh already writes and owns units — `nftables` proves a module can write one
|
||||
and run it — so the mechanism exists; the count grows.
|
||||
- **A module's code is delivered, not pulled**, which puts it behind the same gap as the host's own
|
||||
delivery ([ADR 0141](0141-the-host-delivers-its-own-successor.md), not built): nothing yet delivers a
|
||||
version of a module's binary to a machine. A container's code arrives by `docker pull`, and this does
|
||||
not. **This is the cost of the decision and it is not paid**; until delivery exists, a module whose code
|
||||
is its own is a module somebody places by hand.
|
||||
- **Issue 117 is answered and its three disagreements close differently:** container-or-unit is decided
|
||||
here; one-process-or-several is decided here as several under one account; and whether the record was
|
||||
consulted is fixed by designs 18 and 20 naming this one.
|
||||
|
||||
## References
|
||||
|
||||
- [issue 117](../04-ISSUES/117-a-modules-own-code-is-a-container-and-a-process/00-report.md) — the contradiction this answers
|
||||
- [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) — extended; its "a container" clause is settled here
|
||||
- [ADR 0142](0142-the-mesh-delivers-its-own-components-as-binaries.md) — the same answer for the mesh's own components
|
||||
- [ADR 0141](0141-the-host-delivers-its-own-successor.md) — the delivery this depends on and which is not built
|
||||
- [ADR 0029](0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md) — the sidecar port this relieves
|
||||
@@ -1,99 +0,0 @@
|
||||
---
|
||||
topic: the tiers
|
||||
status: accepted
|
||||
date: 2026-09-30
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 02-DECISIONS/0066-public-routing-is-name-agnostic.md
|
||||
---
|
||||
|
||||
# 151. A route's internal name is composed under the node that serves it
|
||||
|
||||
## Context
|
||||
|
||||
A module that requires a route is given two names from one label: a public one, `<label>.<public
|
||||
domain>`, and an internal one, `<label>.<node>.internal`
|
||||
([ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md)). Both were composed
|
||||
from the node the module runs on.
|
||||
|
||||
The two are answered differently. The public name is published into every machine's roster at the
|
||||
address of the node whose proxy serves it ([ADR 0066](0066-public-routing-is-name-agnostic.md)), so
|
||||
it reaches the proxy from anywhere in the mesh. The internal name is answered by every machine's
|
||||
resolver as *anything under a node's name goes to that node*
|
||||
([design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md)) — the node it was composed from, which is
|
||||
the consumer's. Where the proxy runs on another machine, that name sends a client to a machine with
|
||||
nothing listening, while the public name works
|
||||
([issue 139](../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md)).
|
||||
Every route on this mesh today is served beside its module, so it has not been seen; `route` is
|
||||
provided mesh-wide precisely so that stops being true.
|
||||
|
||||
Beside it, the roster gave every routed name a second entry with the mesh's suffix appended —
|
||||
`<name>.<public domain>.internal` — because it composed a full name for every entry as it does for a
|
||||
machine. That name resolved on every machine, was served by nothing, and was refused by the proxy at
|
||||
the handshake; the first three names tried while reproducing an unrelated issue were those, and the
|
||||
evidence pointed at a regression that had not happened
|
||||
([issue 157](../04-ISSUES/157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md)).
|
||||
|
||||
## Considered Options
|
||||
|
||||
**1. Keep the consumer's name and publish it at the serving node's address**, as the public name is.
|
||||
The name stays `<label>.<consumer>.internal` and an exact roster entry overrides the wildcard.
|
||||
Rejected: it makes `<x>.<node>.internal` mean *goes to that node* except when it does not, which is
|
||||
the one rule the resolver design states; it needs an entry per route where the wildcard needed none;
|
||||
and which of an exact entry and a wildcard a resolver answers first is the resolver's business, which
|
||||
the mesh deliberately does not know.
|
||||
|
||||
**2. A proxy on every machine, so the serving node is always the consumer's.** Rejected for this
|
||||
question: it is a different decision about what `route` is — a node-scoped seat with a mesh-wide
|
||||
fallback — and this mesh runs one proxy on the hub today. Whatever is decided there, a route served
|
||||
from another machine must have a name that reaches it.
|
||||
|
||||
**3. Compose the internal name under the node that serves the route.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**A route's internal name is `<label>.<serving node>.internal` — composed under the node whose proxy
|
||||
answers the route, which is the machine the request arrives at.** The public name is unchanged:
|
||||
`<label>.<public domain>` of the node the module runs on, which is where the operator put it.
|
||||
|
||||
Where the proxy runs beside the module — every route on this mesh today — the two nodes are one and
|
||||
nothing changes. Where it does not, the name says where the request goes, which is what a name under
|
||||
a node's name has always meant.
|
||||
|
||||
**A routed name has no mesh form.** The roster publishes it as itself, once, at the serving node's
|
||||
address. Only a machine has a bare name beside its full one.
|
||||
|
||||
What certifies the internal name is unchanged by this: the proxy that terminates it obtains a
|
||||
certificate from the mesh's authority for the names it is given, and it is given this one.
|
||||
|
||||
Taken on the operator's standing instruction to answer the open design questions in the work order.
|
||||
|
||||
## How this is checked
|
||||
|
||||
- **Composition.** A controller test contributes a route from a module on one node to a proxy offered
|
||||
from another, gathered the way the controller gathers a consumer's contribution for a provider on
|
||||
another machine, and asserts the internal name carries the serving node.
|
||||
- **Publication.** A controller test renders a roster with a machine and a routed name and asserts
|
||||
the routed name appears as itself, once, and never with the suffix appended.
|
||||
- **On the mesh.** After the change no machine's roster carries a `<domain>.internal` entry, and a
|
||||
route's internal name still answers from a container with a certificate from the mesh's authority.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **A route served from another machine now has a usable internal name.** The first module assigned
|
||||
that way will resolve, where before it would have resolved to the wrong machine with no error.
|
||||
- **The internal name of a route can change when its proxy moves.** A route re-homed from one proxy
|
||||
to another gets a new internal name, as the design's rule implies; clients that dialled the old one
|
||||
reach the old machine. The public name does not move with the proxy and is the stable one.
|
||||
- **The roster is one line shorter per routed name**, and a person reading a hosts file no longer
|
||||
finds names that resolve to a refusal.
|
||||
- **Issue 139's second question — a per-node route holder — is left open**, and is a decision about
|
||||
what a seat is rather than about a name.
|
||||
|
||||
## References
|
||||
|
||||
- [issue 139](../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md) — the question
|
||||
- [issue 157](../04-ISSUES/157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md) — the alias
|
||||
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — routed names propagate mesh-wide; extended here
|
||||
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — how the two names are composed and how far each reaches
|
||||
- [design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md) — anything under a node's name goes to that node
|
||||
+5
-34
@@ -48,31 +48,6 @@ form above and dated no earlier than the record's own `date:` — an unmarked ed
|
||||
violation the reviewer looks for in the diff, and a marked one is legible in the record itself.
|
||||
The git history is the backstop, not the record of intent; the note is the record of intent.
|
||||
|
||||
## A pointer back from what a record changes
|
||||
|
||||
A new record naming an old one is not enough. **Where a record changes a mechanism an older record
|
||||
states — without reversing the decision, so no supersession — the older record gets a dated note
|
||||
saying where its mechanism now lives.** A reader arrives at the old record by following a citation,
|
||||
and finds text that is still the decision and no longer the method; nothing in it says a later record
|
||||
moved the method, and the new record is not in their hands.
|
||||
|
||||
> **The mechanism changed — YYYY-MM-DD, by ADR NNNN.** What still stands, what moved,
|
||||
> and why.
|
||||
|
||||
Three examples of the shape, all found by being missed: ADR 0066 still described a routed name being
|
||||
written into every container after 0148 replaced that with resolution; ADR 0047 still said a module's
|
||||
code runs in a container after 0150 made it a supervised process; and ADR 0016 still read as though the
|
||||
lab were the test bed after 0149 said the live mesh is. Each was a citation leading to the wrong
|
||||
answer, in a record that was not wrong about anything it decided.
|
||||
|
||||
**This is not machine-checked, and it cannot be from `extends:` alone.** 102 records extend another and
|
||||
87 name a parent that does not mention them, which is correct: extending usually means building on a
|
||||
context, and a one-directional pointer is the right shape for that. What needs a note is the narrower
|
||||
case where the parent's own text has gone stale, and which case that is, is a judgement — so it is a
|
||||
rule for the author and the reviewer, and the diff is where it is caught. Making it mechanical would
|
||||
mean a record declaring the relationship in its frontmatter, which is a change to the record schema and
|
||||
has not been decided.
|
||||
|
||||
The records run in the order the decisions were taken, oldest first.
|
||||
|
||||
**Every decision is a record.** There is no ledger and no index file — if a decision is worth
|
||||
@@ -199,8 +174,6 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0108** — [A route carries the policy applied to a request, and names a secret rather than holding one](0108-a-route-carries-the-policy-applied-to-a-request.md)
|
||||
- **0109** — [A package registry seat is one per ecosystem, not one for all of them](0109-a-package-registry-seat-is-one-per-ecosystem.md)
|
||||
- **0126** — [A module declares its own seats; the mesh reserves its own](0126-a-module-declares-its-own-seats.md)
|
||||
- **0148** — [The mesh's names are resolved, not copied into every container](0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
|
||||
- **0151** — [A route's internal name is composed under the node that serves it](0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md)
|
||||
|
||||
### What runs on them, and how it gets there
|
||||
|
||||
@@ -234,9 +207,9 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0099** — [A step that runs once names what it reads, and runs again when it changed](0099-a-step-that-runs-once-names-what-it-reads.md)
|
||||
- **0110** — [A seat is held by one assignment, from a closed set, and it may deliver a provision](0110-a-seat-is-a-module-assignment-from-a-closed-set.md)
|
||||
- **0112** — [A module definition names no node, no mesh and no path: everything it needs is a requirement the mesh resolves](0112-a-module-definition-names-no-node-mesh-or-path.md)
|
||||
- **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md)
|
||||
- **0114** — [A credential two parties hold rotates over two credentials; one a single party holds rotates in place, staged; and retiring a credential never removes what it reached](0114-a-shared-credential-rotates-over-two-credentials.md)
|
||||
- **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md)
|
||||
- **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md) *(proposed)*
|
||||
- **0114** — [A credential two parties hold rotates over two credentials; one a single party holds rotates in place, staged; and retiring a credential never removes what it reached](0114-a-shared-credential-rotates-over-two-credentials.md) *(proposed)*
|
||||
- **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md) *(proposed)*
|
||||
- **0117** — [A machine's uplink is a seat: the mesh configures the manager, never the link](0117-a-machines-uplink-is-a-seat.md)
|
||||
- **0118** — [Undeclaring removes what the mesh made, and gives a unit back the state it was found in](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md)
|
||||
- **0120** — [A roster fact carries its format as a template: the mesh owns the data, the module owns the format](0120-a-roster-fact-carries-its-format-as-a-template.md)
|
||||
@@ -255,7 +228,6 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0145** — [A module checks what the mesh claims is reachable, and it checks itself](0145-a-module-checks-what-the-mesh-claims-is-reachable.md) *(superseded)*
|
||||
- **0146** — [Connectivity is checked by name, per hosting form, with a valid certificate](0146-connectivity-is-checked-by-name-per-hosting-form.md)
|
||||
- **0147** — [A module anchors the mesh's authority on a machine, and takes it away again](0147-a-module-anchors-the-meshs-authority.md)
|
||||
- **0150** — [A module's own code runs as supervised processes under the module's one account](0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md)
|
||||
|
||||
### How it is built
|
||||
|
||||
@@ -265,9 +237,9 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0014** — [No workspace — each module is a standalone package consuming published dependencies](0014-no-npm-workspace.md)
|
||||
- **0015** — [Applications live in their own repository; the monorepo is for the mesh](0015-applications-live-in-their-own-repository.md)
|
||||
- **0016** — [The lab](0016-the-lab.md)
|
||||
- **0037** — [Where a module lives](0037-where-a-module-lives.md)
|
||||
- **0037** — [Where a module lives](0037-where-a-module-lives.md) *(proposed)*
|
||||
- **0039** — [What the SDK holds, and what it refuses](0039-what-the-sdk-holds-and-refuses.md)
|
||||
- **0068** — [The lab takes requests, one at a time, and runs each from its own copy](0068-the-lab-takes-requests.md) *(superseded)*
|
||||
- **0068** — [The lab takes requests, one at a time, and runs each from its own copy](0068-the-lab-takes-requests.md) *(proposed)*
|
||||
- **0069** — [A module is a repository and a path within it](0069-a-module-is-a-repository-and-a-path.md)
|
||||
- **0076** — [The SDK is a published package, and the toolchain resolves it by version](0076-the-sdk-is-a-published-package.md)
|
||||
- **0082** — [The registry is reached by name, and the overlay is its security](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)
|
||||
@@ -276,7 +248,6 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0097** — [A vendor image is a declared build input, and a recipe fetches nothing undeclared](0097-a-vendor-image-is-a-declared-build-input.md)
|
||||
- **0107** — [Persistent data is a directory bind, never a named volume](0107-persistent-data-is-a-directory-bind-never-a-named-volume.md)
|
||||
- **0111** — [A build source is on the mesh's git seat, or it is an external repository](0111-a-build-source-is-on-the-git-seat-or-external.md)
|
||||
- **0149** — [The live mesh is the test bed](0149-the-live-mesh-is-the-test-bed.md)
|
||||
|
||||
### How it is checked
|
||||
|
||||
|
||||
@@ -7,10 +7,8 @@ code:
|
||||
- mesh-controller internal/identity/authority.go
|
||||
- mesh-host internal/identity/serving.go
|
||||
- mesh-host internal/apply (the service that reflects a rule set)
|
||||
updated: 2026-09-30
|
||||
updated: 2026-09-29
|
||||
decisions:
|
||||
- 02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md
|
||||
- 02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md
|
||||
- 02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md
|
||||
- 02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md
|
||||
- 02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md
|
||||
@@ -291,31 +289,6 @@ hosts file by the runtime. That extends the file decision rather than overturnin
|
||||
mesh and not chosen by a module: a module that listed the machines would go stale the day one
|
||||
joins, and a module that did not would be one whose containers cannot reach anything by name.
|
||||
|
||||
**Superseded for containers by [ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
|
||||
(2026-09-30).** Everything above still describes the machine's own roster file, which is how it works
|
||||
and how it will keep working. It no longer describes containers.
|
||||
|
||||
Copying the roster into each container made the roster part of each container's identity, so one name
|
||||
moving replaced every container in the mesh — a module assigned on one machine restarted the store, the
|
||||
registry, the edge and mail on another
|
||||
([issue 151](../../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). And it
|
||||
did not stop the staleness it was meant to: a copy taken at creation is stale the moment the roster
|
||||
moves, twice found as a container holding an address that had not existed for days
|
||||
([issues 109](../../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md)
|
||||
and [135](../../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md)).
|
||||
|
||||
**A container resolves the mesh's names through its machine's resolver, at the moment it asks, and
|
||||
nothing is copied.** The resolver is a machine-level process rather than a container, so nothing
|
||||
circular is being asked for. It was gated on a container being able to reach the resolver from any of
|
||||
the runtime's networks
|
||||
([issue 110](../../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)),
|
||||
and landed the day that did, 2026-09-30: the controller writes no mesh name into a container and the
|
||||
host's digest carries only what the module declared for itself.
|
||||
|
||||
The paragraph below states the old boundary, and 0148 deliberately gives it up: a container somebody
|
||||
started by hand resolves the same names as everything else, because the resolver answers the machine,
|
||||
not a list of containers.
|
||||
|
||||
**The boundary, which is deliberate and worth stating:** *declared* containers. A container
|
||||
somebody starts by hand is not the mesh's to configure, and reaching into every container on a
|
||||
machine — declared or not — is what a nameserver in `resolv.conf` would be for.
|
||||
@@ -326,14 +299,6 @@ machine — declared or not — is what a nameserver in `resolv.conf` would be f
|
||||
service, the rest is the node — so what resolves is *anything under a node's name*, going to that
|
||||
node. What routes it once it arrives is a proxy's, and stays separate.
|
||||
|
||||
*2026-09-30.* **So the node in a route's internal name is the one whose proxy answers it**
|
||||
([ADR 0151](../../02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md)).
|
||||
Composed from the node the module ran on, the name sent a client to a machine with nothing listening
|
||||
whenever the proxy ran elsewhere
|
||||
([issue 139](../../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md));
|
||||
composed from the serving node, the rule above holds without exception. The public name stays the
|
||||
module's node's, which is where the operator put it.
|
||||
|
||||
**The mesh writes the data and runs no daemon.** One wildcard per machine, from the same set that
|
||||
writes the hosts file. A resolver is third-party software and runs *on* the mesh rather than being
|
||||
*of* it: the mesh has no business shipping one, choosing which one, or knowing its configuration
|
||||
@@ -409,9 +374,7 @@ can reach from the outside but cannot resolve from the inside is a name it canno
|
||||
authority of its own.
|
||||
|
||||
**So a granted route is published into internal resolution as well** — the routed name to the node
|
||||
that serves it, mesh-wide, by the same mechanism that writes the node names — and as itself: a routed
|
||||
name has no mesh form, and the suffixed alias the roster once added beside it resolved to a refusal
|
||||
([issue 157](../../04-ISSUES/157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md)). It is *given by the
|
||||
that serves it, mesh-wide, by the same mechanism that writes the node names. It is *given by the
|
||||
mesh, not chosen by a module*, for the same reason the node names are: a module listing the routes
|
||||
would go stale the day one changes. The mesh propagates the names it was told to serve and still
|
||||
knows nothing about what they mean
|
||||
|
||||
@@ -5,9 +5,8 @@ code:
|
||||
- mesh-controller cmd/mesh-builder
|
||||
- mesh-controller internal/builder
|
||||
- mesh-catalog modules/builder
|
||||
updated: 2026-09-30
|
||||
updated: 2026-09-29
|
||||
decisions:
|
||||
- 02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md
|
||||
- 02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md
|
||||
- 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md
|
||||
- 02-DECISIONS/0097-a-vendor-image-is-a-declared-build-input.md
|
||||
|
||||
@@ -5,9 +5,8 @@ code:
|
||||
- mesh-catalog modules/showcase
|
||||
- mesh-controller internal/builder
|
||||
- mesh-sdk src
|
||||
updated: 2026-09-30
|
||||
updated: 2026-09-21
|
||||
decisions:
|
||||
- 02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md
|
||||
- 02-DECISIONS/0099-a-step-that-runs-once-names-what-it-reads.md
|
||||
- 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md
|
||||
- 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md
|
||||
|
||||
@@ -155,17 +155,3 @@ design document here, and get it back. That check fails today by design.
|
||||
|
||||
**What stands until then** is the signpost, and the honest description of it: reachable, not
|
||||
surfacing.
|
||||
|
||||
## Where this stands, 2026-09-29
|
||||
|
||||
*Added in a grooming pass.* The knowledge base this record is about is the **predecessor's**, and it
|
||||
is no longer reachable from anything: the surface that answered `recall_search` speaks the transport
|
||||
the mesh removed at the cut-over
|
||||
([issue 147](../147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md)).
|
||||
|
||||
So the sentence in `README.md` that this record catches — *these documents are still indexed into
|
||||
the knowledge base* — is now wrong twice over: nothing indexed them, and there is nothing to index
|
||||
them into. The record stays open, and its answer is no longer "index this repository somewhere"; it
|
||||
is whatever the mesh grows as its own knowledge surface, if it grows one. Until then the honest fix
|
||||
is the README, which should stop claiming a property nothing provides.
|
||||
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: resolved
|
||||
status: open
|
||||
opened: 2026-09-22
|
||||
located-in: [mesh-controller internal/link/protocol.go (the report field that was missing), internal/inventory, cmd/mesh-controller (node show and status)]
|
||||
fixed-by: mesh-controller 7683ba8, corrected by 1d9c102
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
|
||||
@@ -1,143 +0,0 @@
|
||||
# 087 — resolved: the mesh knows which host runs a machine
|
||||
|
||||
*2026-09-30.*
|
||||
|
||||
## The field existed and was thrown away on arrival
|
||||
|
||||
The machine has reported its host version since
|
||||
[ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md) — `Host` on the report, with
|
||||
a comment saying why it must be there: *"without it nothing can say a machine is behind."*
|
||||
|
||||
**The controller's own copy of the report did not have the field.** Two structs describe one message,
|
||||
one on each side of the wire, and only the sending side had it — so it unmarshalled into nothing and the
|
||||
mesh could not answer a question the machine had been answering for a week. That is the whole of this
|
||||
issue's mechanism, and it is worth stating plainly because neither side was wrong on its own.
|
||||
|
||||
## What it says now
|
||||
|
||||
`node show` names it per machine:
|
||||
|
||||
```
|
||||
last heard from here
|
||||
host 2026-09-30-0214
|
||||
```
|
||||
|
||||
`not reported — this machine has not said since the mesh began keeping it` where the mesh has not been
|
||||
told, because a machine that has not said is a different thing from a machine running nothing.
|
||||
|
||||
`status` names the machines that are behind another:
|
||||
|
||||
```
|
||||
1 machine(s) run an older host than another machine does:
|
||||
ace 2026-09-29-0113
|
||||
|
||||
the newest any machine reports is 2026-09-30-0214. A host refuses a declaration carrying a
|
||||
field it does not know, whole — so a new field reaches these machines last
|
||||
```
|
||||
|
||||
## Disagreement, not staleness, and that is deliberate
|
||||
|
||||
The open questions asked whether the controller should refuse to send a declaration a node cannot
|
||||
parse. It cannot yet, honestly: **nothing delivers a host version** (ADR 0141 is accepted and not
|
||||
built, which is [issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md)),
|
||||
so the mesh holds no canonical current version and "behind" has no fixed point to be behind.
|
||||
|
||||
What it can say truthfully is that these machines do not all run the same host, and which is newest of
|
||||
the ones it has been told about. That is the fact that matters before a declaration gains a field: **the
|
||||
oldest host in the mesh is what the mesh may send.**
|
||||
|
||||
Two deliberate refusals to guess:
|
||||
|
||||
- **A machine that has reported nothing is not called behind.** It may be running anything. `node show`
|
||||
says it has not said, per machine, which is the honest form.
|
||||
- **Versions compare as strings.** That suits the timestamps and commits this mesh uses and is wrong
|
||||
for a scheme where `10` sorts before `9`. Said in the code at the place that would have to learn,
|
||||
rather than left as a surprise.
|
||||
|
||||
## What I shipped first was wrong, and the mesh said so within the hour
|
||||
|
||||
The first version reported *"N machine(s) run an older host than another machine does"* and worked out
|
||||
which by comparing versions as strings. **A host reports its version as a commit, and commits have no
|
||||
order.**
|
||||
|
||||
On the live mesh, with the adopted machine pushed for the first time:
|
||||
|
||||
```
|
||||
3 machine(s) run an older host than another machine does:
|
||||
g14 04a27ca
|
||||
novox 04a27ca
|
||||
shanks 04a27ca
|
||||
```
|
||||
|
||||
Those three run the **newer** host — installed 09:18, against the adopted machine's 01:13. `ced54d4`
|
||||
sorts above `04a27ca` and that is all it means. An arbitrary lexicographic result, presented as a fact,
|
||||
about the one thing this record exists to make trustworthy.
|
||||
|
||||
The code carried a caveat saying versions compare as strings and that this "is enough for the timestamps
|
||||
and commits this mesh uses". That was the error, written down and not noticed: it is enough for
|
||||
timestamps and it is **meaningless** for commits, and the mesh reports commits.
|
||||
|
||||
It now reports the split and claims no ordering:
|
||||
|
||||
```
|
||||
4 machine(s) do not all run the same host:
|
||||
04a27ca g14, novox, shanks
|
||||
ced54d4 ace
|
||||
|
||||
a host refuses a declaration carrying a field it does not know, whole — so the mesh may
|
||||
send only what every one of these understands. Which of them is newer is not readable
|
||||
from a commit; that needs a version the host reports as ordered
|
||||
```
|
||||
|
||||
More useful as well as more honest — the reader sees who is on which side of the split, which is what
|
||||
decides whether a field can be sent — and it leaves the ordering where it belongs: with the host, which
|
||||
would have to report something ordered for anybody to have it.
|
||||
|
||||
**This is [issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
|
||||
arriving by my own door**, an hour after closing it: a report that confidently says the opposite of the
|
||||
truth is worse than one that says less, because it trains a reader to distrust the whole surface.
|
||||
|
||||
## Measured on the mesh, and one limitation it exposed
|
||||
|
||||
All four machines run the identical host binary — same digest, installed within eighteen seconds of each
|
||||
other — and at first only one reported its version. The other three said *not reported*, which read as a
|
||||
difference between machines where there was none.
|
||||
|
||||
**A machine states its host version only when the mesh sends it a declaration.** The report is published
|
||||
after an apply; the periodic reconcile that runs every five minutes publishes nothing, because it is the
|
||||
machine keeping itself as declared rather than answering anything. So a machine that is current and idle
|
||||
never says, and the mesh cannot distinguish that from a machine running something ancient.
|
||||
|
||||
Confirmed by pushing: before, `not reported`; after, `04a27ca` — the same version the machine that had
|
||||
been pushed already reported.
|
||||
|
||||
```
|
||||
novox host 04a27ca
|
||||
shanks host 04a27ca
|
||||
g14 host 04a27ca
|
||||
ace host not reported — this machine has not said since the mesh began keeping it
|
||||
```
|
||||
|
||||
`ace` has not been pushed since the field existed; it is adopted and parked.
|
||||
|
||||
**This is enough for the purpose and not enough for the claim.** For deciding whether a new declaration
|
||||
field is safe it is sufficient, because pushing is what the mesh is about to do anyway and the answer
|
||||
arrives with the act. For *knowing what the mesh is running*, it is not: a long-idle machine's entry is
|
||||
as old as its last push, and the honest reading of `not reported` is "nobody has asked recently" rather
|
||||
than "this machine is silent". The words say the first, which is why they are those words.
|
||||
|
||||
Making a heartbeat carry it would close the gap and is a change to what a heartbeat is — a bare word
|
||||
that the node is there, deliberately carrying nothing else. Left alone rather than widened in passing.
|
||||
|
||||
## The open questions, answered as far as they can be
|
||||
|
||||
- *Should a node report the version of its host?* It already did. The gap was the reading.
|
||||
- *Should the mesh refuse to send a field no node understands yet, or refuse per node and say so?*
|
||||
Neither, yet — refusing needs the mesh to know which fields need which version, which is the third
|
||||
question below and is not answered here. What it does is make the disagreement visible before
|
||||
somebody adds a field.
|
||||
- *Is there a general shape — a declaration saying which version of the host it needs?* Still open, and
|
||||
now cheaper to answer: the versions are recorded, so a minimum-version field on a declaration has
|
||||
something to compare against. It belongs with
|
||||
[issue 107](../107-a-declaration-carries-no-order/00-report.md), which wants to add a field and is the
|
||||
first thing this makes safe.
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: resolved
|
||||
status: open
|
||||
opened: 2026-09-22
|
||||
located-in: [mesh-catalog modules/gitea, mesh-controller internal/catalogue/declaration.go]
|
||||
fixed-by: mesh-controller 7352c84, merged in #46 — a module is told its port in a container's environment too, as a file already was
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -39,9 +39,3 @@ the assignment happens to differ.
|
||||
- Should composition refuse an environment value that names a port the module does not fix, the
|
||||
way it refuses other claims a module cannot make?
|
||||
- Which other modules write their own address, with a port, into their environment?
|
||||
|
||||
## Closed
|
||||
|
||||
*2026-09-29, in a grooming pass rather than by whoever fixed it.* The forge's address follows a moved port the same way every other reader does. Found by
|
||||
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
|
||||
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: resolved
|
||||
status: open
|
||||
opened: 2026-09-22
|
||||
located-in: [mesh-controller internal/catalogue/declaration.go, mesh-catalog (every routed module)]
|
||||
fixed-by: mesh-controller bdf965d (a route names the endpoint it serves) with `portOfEndpoint` and `AtPublishedPort` — the contribution carries the endpoint's declared port and the machine-side redirection is applied to it like any other
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -45,9 +45,3 @@ precisely because the predecessor holds the usual one.
|
||||
keeping the mapping out of rendered configuration?
|
||||
- What should refuse a declaration whose contributed route names a port nothing on that node
|
||||
listens on?
|
||||
|
||||
## Closed
|
||||
|
||||
*2026-09-29, in a grooming pass rather than by whoever fixed it.* A route names an endpoint rather than a port, and the redirection that turns a declared port into the published one is applied to contributions too. Found by
|
||||
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
|
||||
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: resolved
|
||||
status: located
|
||||
opened: 2026-09-23
|
||||
located-in: [mesh-controller internal/link, mesh-host internal/link]
|
||||
fixed-by: mesh-host PR 59 (the host refuses an older sequence and drains by it), mesh-controller PR 160 (each send is numbered under the node's hold) — measured 2026-09-30, 02-resolution.md
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -40,16 +40,3 @@ exists there and is thrown away at the wire.
|
||||
separate genesis-digest branch is needed on the host?
|
||||
- Is a sequence enough, or does a mode change deserve its own marker, so a replayed converged
|
||||
declaration is refused by mode as well as by order?
|
||||
|
||||
## What has since made this safer to do (2026-09-30)
|
||||
|
||||
Adding a `sequence` to a declaration is adding a field, and a host refuses a declaration carrying a
|
||||
field it does not know — whole. That was
|
||||
[issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md), and it is resolved: the
|
||||
mesh now records which host each machine reports and `status` names every machine running an older one
|
||||
than another does.
|
||||
|
||||
So the flag day is visible before it is walked into, which it was not when this was filed. It does not
|
||||
make the field free: **the oldest host in the mesh is still what the mesh may send**, and one machine of
|
||||
four is behind today. A sequence that an old host refuses takes that machine out of the mesh's reach
|
||||
entirely — worse than the replay it prevents, which has never been observed.
|
||||
|
||||
@@ -1,76 +0,0 @@
|
||||
# 107 — diagnosis: the fix is a flag day, and it should wait for delivery
|
||||
|
||||
*2026-09-30. Read, measured, and not built — deliberately.*
|
||||
|
||||
## The premise is confirmed
|
||||
|
||||
A host parses a declaration with unknown fields refused, and the code says why rather than leaving it
|
||||
to be inferred:
|
||||
|
||||
> `DisallowUnknownFields` is the whole point rather than strictness for its own sake: a field the host
|
||||
> does not know is a thing the control plane believes it asked for.
|
||||
|
||||
So adding `sequence` and `supersedes` is not an additive change. **Any host that has not been upgraded
|
||||
refuses the whole declaration and applies nothing** — which is exactly the behaviour that keeps a
|
||||
half-understood declaration off a machine, and exactly what makes this expensive.
|
||||
|
||||
## What has changed since this was filed
|
||||
|
||||
[Issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md) is resolved: the mesh now
|
||||
records the host version each machine reports and `status` names every machine running an older host
|
||||
than another does. The flag day is visible before it is walked into, which it was not on 2026-09-23.
|
||||
|
||||
That makes the cost measurable rather than hypothetical, and the measurement is the reason this is not
|
||||
being built today.
|
||||
|
||||
## Why it waits
|
||||
|
||||
**One machine of four runs an older host, and it cannot be upgraded.** `ace` is adopted, deliberately
|
||||
parked until the network and module-assignment work is settled, and **nothing delivers a host version at
|
||||
all** — [ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md) is accepted and not
|
||||
built, which is [issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md).
|
||||
Every machine takes a hand-placed binary.
|
||||
|
||||
So shipping the field means, in order: place a host by hand on three machines, unpark the fourth, place
|
||||
it there too, and only then turn the controller half on. A machine missed in that sequence is a machine
|
||||
the mesh cannot send anything to at all — not degraded, unreachable.
|
||||
|
||||
**And the fault it prevents has never been observed.** The record says so itself: *"Not observed;
|
||||
constructed from the code, and narrow."* It needs a backlog of more than sixteen declarations queued
|
||||
across a `converge`/`adopt` pair, or a broker slow enough to split one, and the host already applies the
|
||||
newest of a drained batch and refuses a declaration that is not the last by digest.
|
||||
|
||||
**Trading a machine's reachability for a replay nobody has seen is the wrong way round.** The right
|
||||
order is [issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md) first —
|
||||
when the mesh can deliver a host, a declaration field costs a rollout instead of an expedition — and the
|
||||
work order already puts that in its last group, as the proof that the mesh can make another of itself.
|
||||
|
||||
## What the open questions look like now
|
||||
|
||||
- *A per-node `sequence` under the controller's node hold, and `supersedes` as the previous digest?*
|
||||
Still the right shape. The controller already holds the lock and already records each send, so the
|
||||
order exists and is thrown away at the wire — unchanged since this was filed.
|
||||
- *Genesis signing its bundle as sequence zero?* Yes, and it is the cheaper half: the bundle is written
|
||||
by the host that will read it, so it has no flag day of its own.
|
||||
- *Is a sequence enough, or does a mode change deserve its own marker?* A sequence alone does not stop
|
||||
a replayed *converged* declaration reaching a node that has since been returned to adopted, which is
|
||||
the incident of issue 104 by another door and is what this record names as its real risk. It wants
|
||||
both, and the second is the one worth having first.
|
||||
- **And one this record did not ask:** should a declaration say which host version it needs? 087 makes
|
||||
that comparable for the first time, and it is the general form of the answer — a field that announces
|
||||
its own requirement, rather than a flag day per field, for ever.
|
||||
|
||||
## Status
|
||||
|
||||
Left `located`. The owner is unchanged, the shape of the fix is agreed, and the gate is
|
||||
[issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md) rather than anything
|
||||
in this record. **This is a judgement about order, not a refusal** — it is cheap to overrule, and the
|
||||
code is a day's work once a host can be delivered.
|
||||
|
||||
## The gate has opened (2026-09-30, evening)
|
||||
|
||||
The mesh delivers the host now — built by its own toolchain, published to its own registry, delivered
|
||||
over the bus and started by the launcher, on all four machines
|
||||
([issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/01-progress.md)). A declaration
|
||||
field is a build and a push, not an expedition. The order this record asked for — hosts first, then
|
||||
the controller — is now two commands and a status line that says when the first has finished.
|
||||
@@ -1,60 +0,0 @@
|
||||
# 107 — resolved: a declaration carries its order
|
||||
|
||||
*2026-09-30. Measured on the mesh.*
|
||||
|
||||
## What was done
|
||||
|
||||
**Hosts first, then the controller** — the order [issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md)
|
||||
says a new declaration field needs, and now a build and a push rather than an expedition
|
||||
([issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/01-progress.md)).
|
||||
|
||||
The host understands a `sequence` on a declaration and tolerates its absence: absent reads as "no
|
||||
order claimed", not "first", so a controller that sends none is still understood and a host that kept
|
||||
a declaration before it understood the field compares nothing. It refuses a declaration with a lower
|
||||
sequence than the one it kept, whole, and says why; and the drain that picks one declaration from a
|
||||
batch keeps the highest sequence rather than the last to arrive — which is the case the report
|
||||
constructed, a backlog drained out of order.
|
||||
|
||||
The controller numbers each send: the next number for that node, taken under the node's hold, before
|
||||
the body exists, so the number is inside what the mesh signs and a replayed older declaration cannot
|
||||
borrow a newer one's.
|
||||
|
||||
## Measured
|
||||
|
||||
```
|
||||
push shanks; push shanks
|
||||
sequence in kept declaration: 2
|
||||
node sequence
|
||||
novox 2
|
||||
shanks 2
|
||||
ace (none — not sent since numbering)
|
||||
g14 (none)
|
||||
status: nobody "not running what the mesh would send them"
|
||||
```
|
||||
|
||||
Both applies went through; neither was refused; the machine holding the earlier one accepted the later.
|
||||
|
||||
## The subtlety, which would have read every machine as behind for ever
|
||||
|
||||
The mesh decides a machine is behind by comparing the digest of what it **would** send against what it
|
||||
**did** send. A number changes the bytes. So the read-only comparison composes with the number the
|
||||
machine was *last* sent — not a fresh one — and is byte for byte what was sent when nothing else
|
||||
changed. Without that, numbering would have made `status` name all four machines as out of date on
|
||||
every reading, permanently.
|
||||
|
||||
## The open questions
|
||||
|
||||
- *A per-node `sequence` under the controller's node hold?* Yes, as described. **`supersedes` — the
|
||||
previous digest — is not added.** A strictly-greater sequence gives the ordering; a chain of digests
|
||||
would give continuity, which nothing here needs yet and which every re-composition would break.
|
||||
- *Genesis signing its bundle as sequence zero?* Zero is "no order claimed", which is what the bundle
|
||||
carries by carrying nothing. Same rule, no genesis branch.
|
||||
- *A marker for a mode change?* Not needed for the incident it guards: a replayed converged declaration
|
||||
reaching a node returned to adopted is already refused **by mode**, before this check runs.
|
||||
|
||||
## How it is checked
|
||||
|
||||
Host: an older sequence is refused, a newer or equal one is not, and no order claimed on either side
|
||||
compares nothing; the drain keeps the highest sequence, and falls back to arrival when none is claimed.
|
||||
Controller: a send carries its number inside the signed bytes, an unnumbered send is byte for byte what
|
||||
it was before, and each node's counter is one higher per send and readable for the comparison.
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: resolved
|
||||
status: located
|
||||
opened: 2026-09-24
|
||||
located-in: [mesh-controller internal/catalogue/declaration.go (every container was given the roster at creation)]
|
||||
fixed-by: mesh-controller PR 161 — no container is given a mesh name; it resolves through its machine's resolver (ADR 0148, landed 2026-09-30 once issue 110 did)
|
||||
located-in: [mesh-host internal/apply]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -59,27 +59,3 @@ container is made with are the same kind of input, read once at creation, and ar
|
||||
and is the stronger statement; it is also what the mesh's own resolver exists for.
|
||||
- Either way: what tells an operator that a container is running with an address the node no longer
|
||||
has? Nothing did.
|
||||
|
||||
## Answered at the cause (2026-09-30)
|
||||
|
||||
This was the first of three arrivals of one fact: a container is given the mesh's names when it is
|
||||
created and never looks again, so a name that moves afterwards is wrong inside it for as long as it
|
||||
runs. It arrived again as [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md),
|
||||
whose fix made the names comparable — and that fix made the roster part of every container's identity,
|
||||
which arrived as [issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md).
|
||||
|
||||
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) ends the
|
||||
copying: a container resolves through its machine's resolver at the moment it asks. The shape this
|
||||
record reports then has nowhere to occur. It is gated on
|
||||
[issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md), so
|
||||
until that lands the mesh still copies and still compares.
|
||||
|
||||
## Resolved (2026-09-30)
|
||||
|
||||
110 landed the same day ([its resolution](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/01-resolution.md)),
|
||||
and mesh-controller PR 161 then removed the copy: no container is given a mesh name or a mesh address,
|
||||
and a module's own declared entries are the only `host` lines it carries. Verified on the control-node
|
||||
after its containers were recreated once — the last time a name will do that: the forge's container
|
||||
carries no extra hosts and resolves another machine and a routed name through the machine's resolver,
|
||||
so the shape this record describes has nowhere to occur. Checked in the controller's tests: a
|
||||
container's declaration is byte-for-byte the same under a roster of one machine and a roster of three.
|
||||
|
||||
+3
-23
@@ -1,9 +1,8 @@
|
||||
---
|
||||
status: resolved
|
||||
status: open
|
||||
opened: 2026-09-24
|
||||
located-in:
|
||||
- mesh-catalog modules/dnsmasq (the runtime was never told; the resolver answered by interface)
|
||||
fixed-by: mesh-catalog PR 175 (the runtime is reloaded and keeps its containers over a restart) and PR 176 (the resolver answers by address, so a query from a bridge is admitted) — measured 2026-09-30, 01-resolution.md
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -52,22 +51,3 @@ knows that is what the rule means.
|
||||
the runtime's default one? That is a stronger rule and would have prevented 109 as well.
|
||||
- What checks it? A converged bed with a container on the default network resolving a mesh name is
|
||||
the missing assertion; nothing in the resolver's own beds covers the filter.
|
||||
|
||||
## What now depends on this (2026-09-30)
|
||||
|
||||
This stopped being a container-DNS inconvenience.
|
||||
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) decides
|
||||
that a container resolves the mesh's names rather than being given a copy of them, which is what stops
|
||||
one name moving from replacing every container in the mesh
|
||||
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)) and what makes a
|
||||
stale address impossible rather than merely noticed
|
||||
([issues 109](../109-a-container-keeps-the-address-it-was-made-with/00-report.md)
|
||||
and [135](../135-a-containers-mesh-names-are-not-compared/00-report.md)).
|
||||
|
||||
**That decision cannot land until this one does**, and not partly: a container on the runtime's default
|
||||
network is the case with no DNS at all, and it is the case the mesh's own forge runs in. Two of four
|
||||
machines also bind the resolver to loopback only, so the runtime hands their containers a public
|
||||
resolver. Both halves are this issue.
|
||||
|
||||
*Later the same day: the second half was wrong, and the first had a different cause than the one above.
|
||||
[01-resolution.md](01-resolution.md) has what was actually found.*
|
||||
|
||||
-64
@@ -1,64 +0,0 @@
|
||||
# 110 — resolved: a container on any network reaches the resolver, and is answered
|
||||
|
||||
*2026-09-30. Measured on the three converged machines; the adopted one holds its resolver module until it
|
||||
is taken and is not covered.*
|
||||
|
||||
## What was actually wrong
|
||||
|
||||
Not what the report predicted. The report named the filter: a container on the runtime's default
|
||||
network asks from a bridge address, and the converged filter admitted queries by source address only.
|
||||
That was true when it was written and was fixed before this issue was ever tested — the filter admits
|
||||
by the link a packet arrives on ([ADR 0144](../../02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md)),
|
||||
and a container's bridge is admitted whole. Tested on every machine: the query arrives, the filter
|
||||
passes it.
|
||||
|
||||
Three other things were wrong, each hiding the next.
|
||||
|
||||
**The runtime had never been told.** The resolver module writes the runtime's `dns` key into the
|
||||
runtime's own configuration file. The runtime reads that key when it starts and not on a reload, and on
|
||||
two machines the runtime predated the file — so every container they started got a public resolver, and
|
||||
`novox.internal` came back as not existing. Nothing reported this: the file was present and current,
|
||||
the resolver ran, and a name not existing is a valid answer. Fixed in mesh-catalog PR 175: the module
|
||||
also sets `live-restore` and reloads the runtime when its file changes, so the one restart the `dns` key
|
||||
needs no longer stops every container. The restart is then the operator's, once per machine; done on
|
||||
both today, with every running container kept.
|
||||
|
||||
**The resolver dropped the query.** With the runtime corrected, a container's query reached the resolver
|
||||
— and got no answer, on every machine, including the one whose runtime had been right all along. The
|
||||
socket was bound to the private address; the filter admitted the packet; dnsmasq received it and
|
||||
discarded it without a line of log. Its configuration said `interface=mesh0`, and dnsmasq admits a
|
||||
query by the interface it arrives on when told an interface: a container's query is addressed to the
|
||||
private address but arrives on the runtime's bridge, and the bridge is not `mesh0`. Fixed in mesh-catalog
|
||||
PR 176: the resolver is told the address to answer on, not the interface that carries it, and a query to
|
||||
that address is admitted whatever bridge brings it. The bridges are the runtime's to name.
|
||||
|
||||
**The report's second half was wrong.** "Two of four machines bind the resolver to loopback only" was
|
||||
an inference from the containers' behaviour, and the behaviour had the cause above. The resolver bound
|
||||
the private address on all four; nothing had asked it there.
|
||||
|
||||
## What is verified
|
||||
|
||||
From a container on the runtime's default network, started by hand and given nothing, on each of the
|
||||
three converged machines: `novox.internal` answers with the hub's private address, through the machine's
|
||||
own resolver. That is the fourth check of
|
||||
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) — "on every
|
||||
network the runtime offers" — and its first step; the record's step 2 (the runtime told per machine, as
|
||||
a file) was already how the module works. Step 3 may now begin.
|
||||
|
||||
## What checks it
|
||||
|
||||
By hand, today. Nothing in the mesh asserts that a container can resolve a mesh name: the resolver's
|
||||
own tests cover what it answers, not who can ask. The check that would have caught all three faults is
|
||||
the one the report asked for and 0148 lists — a container on the default network resolving a mesh name
|
||||
— and it is not built. It belongs with the reachability check of
|
||||
[issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
|
||||
which is parked; until then this is a thing a person verifies after touching the resolver, the filter,
|
||||
or the runtime's configuration.
|
||||
|
||||
## What this cost to find
|
||||
|
||||
The three faults produced one symptom — a container that cannot resolve — and each fix revealed the
|
||||
next. The first was found by reading the runtime's own view of its configuration rather than the file;
|
||||
the second by capturing the query on the bridge and finding it arrive and go unanswered; the third only
|
||||
by admitting the first belief was wrong. A machine that had been believed to work all day had never
|
||||
worked either.
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: resolved
|
||||
status: open
|
||||
opened: 2026-09-24
|
||||
located-in: [mesh-controller module.json, mesh-host internal/apply]
|
||||
fixed-by: ADR 0142 — the mesh's own components are delivered as binaries on the machine, so the controller is a process; the delivery itself is issue 142
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -84,21 +84,3 @@ restart and run-to-completion semantics — so this would not need host-side wor
|
||||
|
||||
The two do not collapse into one. The controller is not a code-carrying sidecar, and `network: host`
|
||||
is what makes the asymmetry visible here and nowhere else.
|
||||
|
||||
## Answered
|
||||
|
||||
*2026-09-29, in a grooming pass.* This asked a question rather than reporting a defect, and the
|
||||
question was taken: [ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md)
|
||||
decides that the mesh's own components — the host, the controller, the catalogue, the builder, the
|
||||
vault — are **binaries on the machine**, delivered by the mechanism
|
||||
[ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md) built, and that
|
||||
third-party software (the store, the registry, the broker) stays a container because an image is the
|
||||
right way to carry somebody else's build.
|
||||
|
||||
So the operating experience this record was written from — every mutating command reached through
|
||||
`docker exec mesh-controller` — is answered, and answered against the container.
|
||||
|
||||
**The delivery is a separate matter and is not this record's.** Step 1 of it is built and no
|
||||
component travels yet; that is
|
||||
[issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md).
|
||||
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: resolved
|
||||
status: located
|
||||
opened: 2026-09-25
|
||||
located-in: [hq, mesh-catalog modules/showcase, mesh-sdk src/tools/index.ts, mesh-tools]
|
||||
fixed-by: hq 83791f0 (PR 196) — ADR 0150: a module's own code runs as supervised processes under the module's one account; designs 18 and 20 now cite it, and ADR 0047 carries a dated note pointing at it
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -99,27 +99,3 @@ what the sidecar is has nowhere in the design layer to look, which is how this w
|
||||
is mechanically checkable: the resource types a design doc names are a closed set, and every
|
||||
member of it either appears in a decision or does not. Whether that check is worth writing is
|
||||
part of this issue, not settled by it.
|
||||
|
||||
## Answered (2026-09-30)
|
||||
|
||||
[ADR 0150](../../02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md)
|
||||
settles all three disagreements, and the design documents win two of them:
|
||||
|
||||
1. **Container or unit — a supervised process.** ADR 0047's argument never required a container. It
|
||||
argued for a runtime *per module*, because a node-wide one could not hold a per-module account and
|
||||
per-module runtimes on one tool key would be handed calls for tools they do not have. A unit per
|
||||
module satisfies that exactly, and a unit runs as an account. The container was the mechanism to
|
||||
hand. [ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md) had
|
||||
already gone the same way for the mesh's own components, and a module is not a container.
|
||||
2. **One process or several — several, under one account.** 0047's "one module, one process, one
|
||||
account" carried its weight in the last clause; its stated worry was "not a second one to scope and
|
||||
seal", which is about a second *identity*. Processes sharing the module's one account create none.
|
||||
What a module may not have is two accounts.
|
||||
3. **Whether the record was consulted — fixed rather than answered.** Designs 18 and 20 now name 0150
|
||||
in `decisions:`, and 0047 carries a dated note saying where its hosting form was settled, so neither
|
||||
door leads to the wrong answer any more.
|
||||
|
||||
**What this does not fix.** A module's code delivered as a binary is behind the same gap as the host's
|
||||
own ([ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md), accepted and not
|
||||
built): a container's code arrives by `docker pull` and this does not, so until delivery exists such a
|
||||
module is one somebody places by hand. 0150 records that as the cost of the decision, unpaid.
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: resolved
|
||||
status: located
|
||||
opened: 2026-09-25
|
||||
located-in: [mesh-host internal/apply (the spec comparison, via issue 135) — not mesh-catalog modules/umami, which this record first named and which was never at fault]
|
||||
fixed-by: mesh-host e82789a (issue 135's fix) — recreating the container with a current roster ended it; verified live 2026-09-30
|
||||
located-in: [mesh-catalog modules/umami]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
|
||||
@@ -1,79 +0,0 @@
|
||||
# 118 — resolved: it was issue 135, and it is over
|
||||
|
||||
*Verified on the machine, 2026-09-30.*
|
||||
|
||||
## It no longer happens
|
||||
|
||||
```
|
||||
$ docker inspect -f '{{.RestartCount}}' umami
|
||||
0
|
||||
$ docker logs umami --tail 12
|
||||
26 migrations found in prisma/migrations
|
||||
No pending migrations to apply.
|
||||
✓ Database is up to date.
|
||||
▲ Next.js 16.3.4
|
||||
✓ Ready in 0ms
|
||||
$ curl -o /dev/null -w '%{http_code}' https://umami.novox.be/
|
||||
200
|
||||
```
|
||||
|
||||
No restart loop, no `prisma.$queryRaw()` timeout, and the public name that had answered `502` since
|
||||
2026-09-25 answers `200`. The raw query that could not complete now runs twenty-six migrations and
|
||||
reports the store up to date.
|
||||
|
||||
## What it was
|
||||
|
||||
**The same fault as [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md), which
|
||||
was diagnosed three days later without either record noticing the other.** 135's container is this
|
||||
one, named in its own evidence table:
|
||||
|
||||
```
|
||||
umami created 2026-09-23 novox.internal:10.42.0.1
|
||||
mesh-catalog created today novox.internal:10.10.0.1
|
||||
```
|
||||
|
||||
The mesh's overlay range had moved. Umami had been created before the move and held the store's name
|
||||
at an address that no longer existed, while every container made after the move held the current one.
|
||||
That is why the dial appeared to succeed and the first real query timed out, and why the same query
|
||||
from the same network with the same credential answered in milliseconds — **what differed was the
|
||||
name, not the path, not the credential and not the store.**
|
||||
|
||||
Today it holds `novox.internal:10.10.0.1`.
|
||||
|
||||
## Why this record did not find it
|
||||
|
||||
The report's own reasoning is worth keeping as a warning, because it is careful and it is wrong:
|
||||
|
||||
> A dial that succeeds and a query that times out, from a container on one network to a store on
|
||||
> another, has the shape of a path-MTU or conntrack fault (large response packets dropped after the
|
||||
> small handshake ones pass), or of the store accepting the TCP connection while the backend it
|
||||
> proxies for is wedged.
|
||||
|
||||
Both are good hypotheses about a network path. Neither is the answer, and the record also names the
|
||||
move that would have found it — *"what is known to differ for umami against every working consumer of
|
||||
the same store tonight is nothing yet — that comparison is the first move"* — and then did not make it.
|
||||
Comparing umami's hosts entries against any container created that week would have shown a five-day-old
|
||||
address in one field.
|
||||
|
||||
**A stale name presents as a network fault.** That is the lesson, and it is the reason
|
||||
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) stops
|
||||
copying names into containers at all: not because detecting staleness is hard, but because it disguises
|
||||
itself as something else for five days while every check reports success.
|
||||
|
||||
## What actually ended it
|
||||
|
||||
Issue 135's fix — `mesh-host` e82789a, *a container's mesh names are part of what it is* — put the
|
||||
roster into the spec digest the host compares, so a container whose names moved is recreated like one
|
||||
whose image moved. That recreated umami with a current roster and ended this.
|
||||
|
||||
That fix has since been superseded in turn, by 0148, because making the roster part of every
|
||||
container's identity meant one name moving replaced every container in the mesh
|
||||
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). So this record
|
||||
closes on a fix that is itself on the way out — which does not make it less closed, and is worth saying
|
||||
plainly rather than leaving a reader to find out.
|
||||
|
||||
## Not carried forward
|
||||
|
||||
The `502` had one other contributor worth recording as ruled out: a stale duplicate Traefik router for
|
||||
this name, hand-authored in the predecessor's era beside the mesh-written one, removed on 2026-09-25
|
||||
and — as the report says — changing nothing. It was not the cause and it is gone.
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: resolved
|
||||
status: located
|
||||
opened: 2026-09-26
|
||||
located-in: [mesh-sdk src/provisioner, mesh-catalog modules/redis]
|
||||
fixed-by: mesh-catalog bbda88c, merged in #84 — every credential provider says whether it still holds a consumer
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -60,9 +60,3 @@ checks it after the first pass.
|
||||
instance and leaves the gap for the others.
|
||||
- Where does the record of what was applied live, if not in memory? ADR 0114, still
|
||||
proposed, puts rotation state with the vault. The same place may answer this.
|
||||
|
||||
## Closed
|
||||
|
||||
*2026-09-29, in a grooming pass rather than by whoever fixed it.* A provisioner asks the backend what is there rather than trusting what it remembers doing. Found by
|
||||
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
|
||||
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
|
||||
|
||||
@@ -1,9 +1,7 @@
|
||||
---
|
||||
status: resolved
|
||||
status: open
|
||||
opened: 2026-09-26
|
||||
located-in: [mesh-host internal/link (the apply report line), mesh-controller cmd/mesh-controller (status and its all-well condition)]
|
||||
fixed-by: mesh-host cbf5018, mesh-controller bfd983e
|
||||
amended-design:
|
||||
located-in: [mesh-host internal/apply, mesh-controller]
|
||||
---
|
||||
|
||||
# 125 — a hold is not a line in the apply report, and an operator flew blind into an outage
|
||||
|
||||
@@ -1,70 +0,0 @@
|
||||
# 125 — resolved: a hold is a line in the report, and it stops the mesh reading as well
|
||||
|
||||
*2026-09-30.*
|
||||
|
||||
## What the four surfaces say now
|
||||
|
||||
The report named four surfaces, none of which carried the one sentence that mattered. Two of them
|
||||
already did by the time this was picked up, and two did not.
|
||||
|
||||
**1. The apply report says what it held** — this was missing, and is the line the operator was reading
|
||||
when the count did not add up:
|
||||
|
||||
```
|
||||
applied 330 resource(s), 16 held until their module is taken (route-proxy: 13, mailu: 3)
|
||||
```
|
||||
|
||||
Grouped by module and ordered by name, because `take` acts on a module and that is the sentence an
|
||||
operator needs. An apply that held nothing says nothing extra — a line reporting `0 held` on every
|
||||
converged apply is one that stops being read.
|
||||
|
||||
**2. `status` counts holds, and a hold breaks "all well"** — this was missing. Status now says:
|
||||
|
||||
```
|
||||
15 resource(s) are held as found, because their module was assigned and never taken — so it is
|
||||
running none of what it declares:
|
||||
novox route-proxy (13), mailu (2)
|
||||
|
||||
`take <node> <module>` compares what runs against what it declares, and runs it
|
||||
```
|
||||
|
||||
**And the sentence that was the fault no longer prints.** "all doing what they were told, all heard
|
||||
from, running what the mesh would send them" was *true* for the whole outage, and acting on it stopped
|
||||
the predecessor's proxy. A hold now suppresses it; being adopted still does not, and the difference is
|
||||
deliberate — adopted is a mode somebody chose and can leave alone, a module assigned and never taken is
|
||||
a half-finished action with nothing left to finish it.
|
||||
|
||||
**3. `node show <node>` shows the node's own held list** — already true, and recorded here as checked
|
||||
rather than assumed. It reads `held` from what the machine last reported, with an `as of` beside it, and
|
||||
names each hold's kind, target, id, module, whether something other than the mesh has changed it, and
|
||||
where an original was kept.
|
||||
|
||||
**4. The push's count** is unchanged and now interpretable, which was the ask: `sent 346` against
|
||||
`applied 330, 16 held until their module is taken (…)` is a pair a reader can resolve without opening a
|
||||
file on the machine.
|
||||
|
||||
## Where the data came from
|
||||
|
||||
**The host already reported it.** `Held` has been on the wire since [ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md),
|
||||
and the controller already recorded it and showed it in `node show`. Nothing needed a new field, a new
|
||||
message or a migration — which is why this is additive, and why the report's framing (*"the semantics
|
||||
are consistent and right; the reporting is what let them be forgotten"*) was exactly right.
|
||||
|
||||
What was missing was that two surfaces never asked. Status read the mesh's take-time listing, so a
|
||||
module assigned after that listing showed nothing at all; the apply line counted what it applied and
|
||||
said nothing about the difference.
|
||||
|
||||
## One thing deliberately not done
|
||||
|
||||
**Status does not call a hold a fault.** It is correct behaviour, and a reader trained to see red for
|
||||
something the mesh did right will stop reading. It is reported as work outstanding, with the command
|
||||
that finishes it — and it withholds the all-well sentence, which is the part that carries the weight.
|
||||
|
||||
## How it is checked
|
||||
|
||||
- The apply line names the count and the module, and is empty when nothing is held (mesh-host).
|
||||
- Status finds a hold from what the machine reported, end to end through the store.
|
||||
- **A held module makes the all-well condition false**, asserted against the production condition
|
||||
rather than a copy of it — that condition is now one named function for this reason.
|
||||
- The JSON form carries a row per machine and module, and omits the field entirely when nothing is
|
||||
held.
|
||||
+3
-22
@@ -1,15 +1,10 @@
|
||||
---
|
||||
status: resolved
|
||||
status: located
|
||||
opened: 2026-09-27
|
||||
located-in: [mesh-controller cmd/mesh-controller/push.go, mesh-controller cmd/mesh-controller/sendable.go]
|
||||
fixed-by: mesh-controller sendable.go and push.go — an empty declaration is sent carrying `owns_nothing`, and the host refuses an empty body that does not carry it
|
||||
located-in: [mesh-controller cmd/mesh-controller/push.go]
|
||||
---
|
||||
|
||||
# 149 — a declaration that shrinks to empty is skipped, so the node keeps what it should drop
|
||||
|
||||
*Opened as 127 and renumbered on 2026-09-29: two records were given that number on the same day.*
|
||||
*The other kept it, because three documents and three source files cite it by number and nothing
|
||||
cited this one but a decision and a sibling issue, both corrected with this move.*
|
||||
# A declaration that shrinks to empty is skipped, so the node keeps what it should drop
|
||||
|
||||
## What was observed
|
||||
|
||||
@@ -42,17 +37,3 @@ mean "own nothing", which the host already applies correctly when it receives on
|
||||
|
||||
On ace, one command drops it permanently (the corrected controller never re-composes it):
|
||||
`sudo ufw delete allow 5671`. At ace's converge it would clear on its own.
|
||||
|
||||
## Closed
|
||||
|
||||
*2026-09-29, in a grooming pass.* Both halves are on `main` and both name this issue.
|
||||
|
||||
- The control plane **sends** it: a declaration that composes to no resources goes out with
|
||||
`owns_nothing`, and `push` says *sent, not skipped*.
|
||||
- The host **refuses an empty body that does not carry it**, so a truncated or mis-composed
|
||||
declaration can never be read as "own nothing" — which is the failure the fix had to avoid while
|
||||
making the empty case expressible.
|
||||
|
||||
Closed by reading the code rather than by watching a machine let go of a stray resource; the record
|
||||
says so rather than implying a run.
|
||||
|
||||
@@ -1,6 +1,5 @@
|
||||
---
|
||||
status: resolved
|
||||
fixed-by: mesh-host 1cb8953 and fdc768c — the mesh writes into a marked block of a text file instead of over it
|
||||
status: located
|
||||
opened: 2026-09-26
|
||||
located-in: [mesh-controller internal/catalogue/facts.go, mesh-host internal/apply]
|
||||
---
|
||||
@@ -72,9 +71,3 @@ private network loses that name too.
|
||||
- The host's file resource supports `into: "json"` only; anything else is a whole write.
|
||||
- `node show <node>` on the adopted workstation: `holds file /etc/hosts
|
||||
mesh-wireguard.fact-node-names`, original kept.
|
||||
|
||||
## Closed
|
||||
|
||||
*2026-09-29, in a grooming pass rather than by whoever fixed it.* A shared hosts file keeps every line that is not the mesh's. Found by
|
||||
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
|
||||
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
|
||||
|
||||
@@ -1,8 +1,7 @@
|
||||
---
|
||||
status: resolved
|
||||
status: located
|
||||
opened: 2026-09-26
|
||||
located-in: [mesh-catalog modules/ca-trust]
|
||||
fixed-by: the ca-trust module registered from the catalogue and assigned to a workstation; verified in both directions 2026-09-30 (02-resolution.md)
|
||||
located-in: [mesh-catalog ca-trust]
|
||||
---
|
||||
|
||||
# 129 — nothing makes a machine trust the mesh's own certificate authority
|
||||
|
||||
@@ -1,82 +1,32 @@
|
||||
# 129 — diagnosis
|
||||
# Diagnosis
|
||||
|
||||
*2026-09-30, from the workstation the issue was opened on.*
|
||||
*2026-09-29.*
|
||||
|
||||
## Still live, and reproduced exactly
|
||||
## What was ruled out
|
||||
|
||||
The certificate is genuine, the authority is the mesh's, and nothing on the machine trusts it:
|
||||
**That something already carries the root and it is only misplaced.** It does not. The authority
|
||||
serves its root at a path beside its ACME directory, and the one thing that fetches it — the route
|
||||
proxy — puts it in a directory of its own and hands it to one program. Nothing has ever written
|
||||
into a machine's trust store. Measured on three converged machines: the anchors present are the
|
||||
predecessor's authority and a developer tool's local root, and on the machines where the
|
||||
predecessor's was deliberately removed, every internal name fails verification.
|
||||
|
||||
```
|
||||
$ openssl s_client -connect keycloak.novox.internal:443 -servername keycloak.novox.internal
|
||||
subject=CN=keycloak.novox.internal
|
||||
issuer=O=Mesh Internal CA, CN=Mesh Internal CA Intermediate CA
|
||||
Verify return code: 20 (unable to get local issuer certificate)
|
||||
**That the private network could carry it, the way it carries the registry's trust.** That is what
|
||||
the report proposed, and it was rejected on consideration rather than on difficulty
|
||||
([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md), option 1): being on
|
||||
the network is what makes the registry *reachable* and is therefore the right trigger there, while
|
||||
trusting an authority is a separate fact from being able to reach it. The anchor's directory and
|
||||
the command that refreshes the extracted bundles are also one operating system's difference, which
|
||||
is the host's half of the mesh and not the controller's.
|
||||
|
||||
$ curl https://keycloak.novox.internal/
|
||||
curl: (60) SSL certificate OpenSSL verify result: unable to get local issuer certificate (20)
|
||||
```
|
||||
**That it needs a new host resource type.** It does not, today. A file and a service say the whole
|
||||
of it, which the packet filter already proves. The primitive becomes the right answer when a second
|
||||
operating system is in play, and not before.
|
||||
|
||||
`trust list` holds no entry for the mesh. The anchors present are two `mkcert` development roots and
|
||||
the predecessor's lab root — the report's account of the trust store is unchanged.
|
||||
## Where it belongs
|
||||
|
||||
The public name on the same proxy verifies cleanly (`CN=keycloak.novox.be`, Let's Encrypt, return code
|
||||
0), which places the fault exactly where the report puts it: not in the proxy, not in the authority,
|
||||
and not in the certificate.
|
||||
A module in the catalogue: it requires `internal-acme-ca`, fetches the root over the mesh's own
|
||||
network, installs it as a trust anchor, refreshes the machine's bundles, and — because being
|
||||
unassigned stops its unit, and stopping the unit is what undoes it — takes both away again.
|
||||
|
||||
**The name matters, and the report's "every HTTPS name the mesh serves internally" is too broad.** The
|
||||
served internal name is `<label>.<node>.internal`. The hosts file also carries
|
||||
`<label>.<public-domain>.internal`, which nothing serves and which fails differently — that is
|
||||
[issue 157](../157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md), found while
|
||||
reproducing this, and it cost the first several minutes of this diagnosis.
|
||||
|
||||
## The authority serves what the module needs
|
||||
|
||||
`step-ca` is up and healthy, and publishes `roots: /roots.pem` for both `acme-ca` and
|
||||
`internal-acme-ca`. That endpoint returns PEM:
|
||||
|
||||
```
|
||||
$ curl -sk https://127.0.0.1:9000/roots.pem
|
||||
-----BEGIN CERTIFICATE-----
|
||||
MIIBvzCCAWWgAwIBAgIQYa2CkdJk16JyG/dVy2qoEzAKBggqhkjOPQQDAjA+…
|
||||
```
|
||||
|
||||
So `${bound:internal-acme-ca:roots}` in the `ca-trust` module composes to a URL that returns a
|
||||
certificate, and the module's own check — refuse a body that is not one — is checking the right thing.
|
||||
|
||||
Worth recording because it was nearly filed as a defect: step-ca *also* serves `/roots`, which returns
|
||||
`{"crts":["-----BEGIN CERTIFICATE-----\n…"]}`. That body contains the literal text the module greps
|
||||
for, so had the module used `/roots` it would have installed JSON into the anchors directory and
|
||||
reported success. It does not use it. The guard is sound only because the published path is the PEM
|
||||
one, which is worth knowing before anybody changes either.
|
||||
|
||||
## What is actually in the way
|
||||
|
||||
**The module is not registered.** The report and the work plan both say it exists and is merged, which
|
||||
it does — `mesh-catalog modules/ca-trust`, on `main`. But the mesh has never been told about it:
|
||||
|
||||
```
|
||||
$ mesh-controller module list | grep -iE 'ca-trust|step-ca'
|
||||
step-ca 1 built 67f5f4cf on novox
|
||||
```
|
||||
|
||||
39 of the catalogue's 76 manifests are registered. `ca-trust` is one of the 37 that are not, so it
|
||||
cannot be assigned to anything — "assign it to one machine" has no module to name.
|
||||
|
||||
A dry run confirms it registers cleanly and needs no artifact built: it declares a directory, a script,
|
||||
a unit and a service, and no image.
|
||||
|
||||
```
|
||||
$ mesh-controller build <catalogue> --path modules/ca-trust --dry-run
|
||||
… the manifest, parsed and validated
|
||||
```
|
||||
|
||||
## So the remaining work is three steps, not one
|
||||
|
||||
1. **Register it** — build it from the catalogue, which pins nothing because it has no artifacts.
|
||||
2. **Assign it** to a machine. The workstation this was observed on is the honest first choice: it is
|
||||
where a person meets the fault, and it is where the check can be made with a plain client.
|
||||
3. **Verify** `curl https://<label>.<node>.internal/` with no flags, and `trust list` naming the mesh.
|
||||
|
||||
Then removal, which the module declares and nothing has exercised: unassigning must take the anchor
|
||||
away and refresh the bundles ([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md)),
|
||||
and that is the half most likely to be wrong, because it is the half nobody reaches by accident.
|
||||
The owner is therefore `mesh-catalog`, module `ca-trust`, and nothing in the control plane.
|
||||
|
||||
@@ -1,104 +0,0 @@
|
||||
# 129 — resolved: the workstation trusts the mesh, and stops when told to
|
||||
|
||||
*Done and measured on the machine, 2026-09-30.*
|
||||
|
||||
## What was done
|
||||
|
||||
Three steps, not the one the plan expected — the module was merged and had never been registered
|
||||
(see [the diagnosis](01-diagnosis.md)):
|
||||
|
||||
1. **Registered** `ca-trust` from the catalogue. No artifact to build: it declares a directory, a
|
||||
script, a unit and a service, and no image.
|
||||
2. **Assigned** it to the workstation the issue was opened on, and pushed.
|
||||
3. **Verified** with a plain client, then **unassigned and pushed again** to exercise removal, then
|
||||
assigned and pushed once more.
|
||||
|
||||
The host's own account of arriving:
|
||||
|
||||
```
|
||||
created ca-trust.state (/var/lib/ca-trust)
|
||||
created ca-trust.anchor (/var/lib/ca-trust/anchor)
|
||||
created ca-trust.unit (/etc/systemd/system/mesh-ca-trust.service)
|
||||
updated ca-trust.trust (mesh-ca-trust.service): boot disabled to enabled, stopped to running
|
||||
```
|
||||
|
||||
## It works, by the check the report asked for
|
||||
|
||||
The report's own reproduction, with no flags and nothing installed by hand:
|
||||
|
||||
```
|
||||
$ curl -sS -o /dev/null -w '%{http_code}' https://git.novox.internal/
|
||||
200
|
||||
$ openssl s_client -connect git.novox.internal:443 -servername git.novox.internal
|
||||
issuer=O=Mesh Internal CA, CN=Mesh Internal CA Intermediate CA
|
||||
Verify return code: 0 (ok)
|
||||
$ trust list | grep -A2 Mesh
|
||||
label: Mesh Internal CA Root CA
|
||||
trust: anchor
|
||||
category: authority
|
||||
```
|
||||
|
||||
Four internal names, all verifying: `git` 200, `umami` 200, `keycloak` 302, `drive` 302. Before this,
|
||||
every one of them was `curl: (60) … unable to get local issuer certificate (20)`.
|
||||
|
||||
**And the consequence the report named specifically**: git over HTTPS to the mesh's forge, which it said
|
||||
had forced the working clone URL to be ssh-only.
|
||||
|
||||
```
|
||||
$ git ls-remote https://git.novox.internal/novox/hq.git HEAD
|
||||
76fbe323ea3401fcdadbf500c61bd3fa5a0a8603 HEAD
|
||||
```
|
||||
|
||||
## Removal is symmetric, which nothing had ever shown
|
||||
|
||||
[ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md) says the module anchors the
|
||||
authority *and takes it away again*. That half had never run. Unassigning and pushing:
|
||||
|
||||
```
|
||||
removed ca-trust.trust (mesh-ca-trust.service)
|
||||
removed ca-trust.unit (/etc/systemd/system/mesh-ca-trust.service)
|
||||
removed ca-trust.anchor (/var/lib/ca-trust/anchor)
|
||||
removed ca-trust.state (/var/lib/ca-trust)
|
||||
```
|
||||
|
||||
Then: the anchor file gone, `trust list` naming no authority of the mesh's, and the plain client back to
|
||||
`unable to get local issuer certificate (20)`. A machine that leaves the mesh stops trusting it, as the
|
||||
record claims.
|
||||
|
||||
**The order is what makes it work, and is worth saying.** The service is removed *first*, so systemd
|
||||
runs the unit's `ExecStop` — which is what deletes the certificate and refreshes the bundles — while the
|
||||
script it calls still exists. Had the script or the state directory gone first, stopping the unit would
|
||||
have had nothing to run, and the anchor would have been left behind with nothing declaring it. Nothing
|
||||
in the module says this; it is the host's removal order that makes the module's symmetry real.
|
||||
|
||||
## What this leaves
|
||||
|
||||
- **Three machines of four** *(extended the same day, after the first was proven)*. `ca-trust` is now
|
||||
assigned to every converged machine, and each verifies with a plain client:
|
||||
|
||||
```
|
||||
novox mesh CA in trust store: 1 https://git.<node>.internal/ -> 200
|
||||
g14 mesh CA in trust store: 1 https://git.<node>.internal/ -> 200
|
||||
shanks mesh CA in trust store: 1 https://git.<node>.internal/ -> 200
|
||||
```
|
||||
|
||||
Before, each of the two that had not been assigned it answered
|
||||
`curl: (60) … unable to get local issuer certificate (20)` and held no entry for the mesh.
|
||||
|
||||
**`ace` is deliberately not among them.** It is the adopted machine, still carrying the
|
||||
predecessor's resolver and filter, and it is not converged until the network and module-assignment
|
||||
work is settled. Assigning a module to it would hold rather than run
|
||||
([ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)), which is
|
||||
correct and is not the same as trusting anything.
|
||||
|
||||
- **It arrives per assignment, which is a shape worth questioning.** The report's reasoning — that
|
||||
being on the private network is what makes a machine one that speaks to the mesh's names — argues
|
||||
for a fact carried to every machine on the network, the way the roster and the registry trust are.
|
||||
ADR 0147 chose a module assigned per machine, and this record does not reopen it; the cost is that
|
||||
a machine joining the mesh trusts nothing until somebody remembers a second command.
|
||||
- **`service-manager` reports `degraded` on this machine** and the module ran anyway. Worth knowing that
|
||||
the capability gate passes on a degraded service manager, since a module whose whole delivery is a
|
||||
unit is the kind that would be worst served by one.
|
||||
- **The predecessor's authority is still in the trust store**, beside the mesh's now rather than instead
|
||||
of it. The report notes that it cannot be retired while anything on the machine speaks TLS to a mesh
|
||||
name; that is no longer true here, and retiring it is its own piece of work.
|
||||
@@ -1,6 +1,5 @@
|
||||
---
|
||||
status: resolved
|
||||
fixed-by: mesh-host 3112c88 — undeclaring gives a unit back the state it was found in, and removes only a process the mesh made
|
||||
status: located
|
||||
opened: 2026-09-27
|
||||
located-in: [mesh-host internal/apply/apply.go (remove)]
|
||||
amended-design: 02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md
|
||||
@@ -16,7 +15,7 @@ found that the host's `remove` path stops every `service` resource that is no lo
|
||||
delete". `store.Orphans` matches by id alone. So any of these stops the unit:
|
||||
|
||||
- the module is unassigned — by mistake, or to switch it for another;
|
||||
- the node is sent a deliberately-empty declaration ([issue 149](../149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md));
|
||||
- the node is sent a deliberately-empty declaration ([issue 127](../127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md));
|
||||
- a later catalogue version renames the resource's `id`.
|
||||
|
||||
That is right for a service the mesh brought into being. It is wrong for a unit the mesh
|
||||
@@ -65,9 +64,3 @@ something to settle in passing.
|
||||
|
||||
The unassign preview is partly answered — the host's plan names each unit it will stop — and the
|
||||
controller's side is left open.
|
||||
|
||||
## Closed
|
||||
|
||||
*2026-09-29, in a grooming pass rather than by whoever fixed it.* Undeclaring no longer stops a unit the mesh only reloaded or only kept running. Found by
|
||||
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
|
||||
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
|
||||
|
||||
@@ -68,31 +68,3 @@ container runtime's shape, not a choice; the answer is to recreate, which is wha
|
||||
that would rather re-read a roster from a file can already ask for one as a fact
|
||||
([ADR 0120](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md)) and restart on
|
||||
it.
|
||||
|
||||
## The container was umami, and it had its own record (2026-09-30)
|
||||
|
||||
The container in the table above is umami, and its symptom had already been filed three days earlier as
|
||||
[issue 118](../118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md) — the store
|
||||
answering the dial and timing out the query, a restart loop, and a public name answering `502`. Neither
|
||||
record noticed the other: 118 was reasoning about path-MTU and conntrack on the network between the
|
||||
container and the store, which is what a stale name looks like from inside the container.
|
||||
|
||||
118 is resolved by this record's fix and says so. Noted here because a symptom filed twice, diagnosed
|
||||
once, and closed in one place is how a repository comes to disagree with itself.
|
||||
|
||||
## What replaced this fix (2026-09-30)
|
||||
|
||||
The fix here — putting the mesh's names into the digest the host compares, so a container whose names
|
||||
moved is recreated like one whose image moved — worked, and cost more than it was worth. It made the
|
||||
roster part of every container's identity, so one name moving replaced every container in the mesh: a
|
||||
module assigned on one machine restarted the store, the registry, the edge and mail on another
|
||||
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)).
|
||||
|
||||
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) goes at
|
||||
the cause this record only described: **a copy taken at creation is stale the moment the roster moves,
|
||||
and detecting that is not as good as not copying.** A container resolves the mesh's names through its
|
||||
machine's resolver, at the moment it asks, so the fault this record reports cannot occur rather than
|
||||
being noticed a restart later.
|
||||
|
||||
Recorded here because this is where somebody arrives to find out why the digest carries names, and the
|
||||
answer is that it did, for two days short of a month, and stopped.
|
||||
|
||||
@@ -1,9 +1,9 @@
|
||||
---
|
||||
status: resolved
|
||||
status: open
|
||||
opened: 2026-09-28
|
||||
located-in: [mesh-controller internal/catalogue/declaration.go (composeName took the consumer's own name as the internal domain)]
|
||||
fixed-by: mesh-controller PR 163 — the internal name composes under the node whose proxy serves the route (ADR 0151, 2026-09-30)
|
||||
amended-design: 03-DESIGN/01-to-be/08-connectivity.md
|
||||
located-in: [mesh-controller internal/catalogue]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 139 — An internal route name resolves to the consumer's node, not the one that serves it
|
||||
@@ -49,15 +49,3 @@ resolve whether or not anything answers.
|
||||
and if so, is `route` still one mesh-wide provision or a node-scoped seat with a mesh-wide fallback?
|
||||
- What certifies the name in either case? The certificate is obtained by whoever terminates TLS, and
|
||||
that is the question above in another form.
|
||||
|
||||
## Answered (2026-09-30)
|
||||
|
||||
[ADR 0151](../../02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md):
|
||||
the internal name is composed under the node that serves the route — the machine the request arrives
|
||||
at — because `<x>.<node>.internal` means *goes to that node* and nothing else. The public name stays
|
||||
the consumer node's, which is where the operator put it. The first question is answered that way; the
|
||||
second, a per-node route holder, is a decision about seats and is left where it is; the third is
|
||||
unchanged, since the proxy that terminates the name is given it and certifies it.
|
||||
|
||||
mesh-controller PR 163 carries it. On this mesh every route is served beside its module, so no name
|
||||
changed; the controller's tests hold the case where it would.
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
---
|
||||
status: resolved
|
||||
status: located
|
||||
opened: 2026-09-28
|
||||
located-in:
|
||||
- mesh-controller internal/catalogue/manifest.go
|
||||
@@ -7,7 +7,7 @@ located-in:
|
||||
- mesh-controller internal/catalogue/declaration.go
|
||||
- mesh-controller examples/route-proxy
|
||||
- mesh-catalog (every routed module manifest)
|
||||
fixed-by: mesh-controller bdf965d (a module names its endpoints) and c68d3a7 (an assignment configures an endpoint as one thing) — the filter, the proxy's names and both authorities now read one statement
|
||||
fixed-by:
|
||||
amended-design: 03-DESIGN/01-to-be/08-connectivity.md
|
||||
---
|
||||
|
||||
|
||||
@@ -1,47 +0,0 @@
|
||||
# Resolution
|
||||
|
||||
*2026-09-29.*
|
||||
|
||||
**Built, and this record did not say so.** The issue was written on 2026-09-28 and answered the same
|
||||
week by two commits in `mesh-controller`; nothing came back to close it, so the mesh's own account of
|
||||
itself said for a day that reach was declared nowhere while the code read it in three places.
|
||||
|
||||
- `bdf965d` — *a module names its endpoints, and a route names the one it serves*. `listens[].name`
|
||||
is the endpoint; a route contribution names the endpoint rather than repeating a port.
|
||||
- `c68d3a7` — *an assignment configures an endpoint as one thing*. The `endpoints` settings key, per
|
||||
node, by endpoint name: `{"endpoints": {"ssh": {"port": 20134, "reach": "public"}}}` — port, label
|
||||
and reach in one block, which is what [ADR 0138](../../02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md)
|
||||
asked for and what [ADR 0046](../../02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md)
|
||||
said configuration is.
|
||||
|
||||
## The three readers, which is what the issue was about
|
||||
|
||||
The complaint was that the per-node source override had exactly one caller. It now has three, and
|
||||
they are the three mechanisms reach was decided to settle at once:
|
||||
|
||||
| reader | what it does with it |
|
||||
|---|---|
|
||||
| the filter | `Reaches` turns each endpoint's reach into the rule for its machine port |
|
||||
| the proxy's names | `composeName` composes the public name, the internal name, or both — and a name nobody asked for is not composed |
|
||||
| the authorities | the proxy certifies only names it was actually given, each from its own authority, through two host policies rather than one |
|
||||
|
||||
**A routed endpoint keeps the manifest's port**, which is ADR 0138's own insight and older than it
|
||||
([ADR 0045](../../02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)): the
|
||||
proxy is how it is reached, so `public` there asks for a public *name*, not an open port.
|
||||
|
||||
**An endpoint that is not routed is reached and never named.** Git over ssh is that case — the one
|
||||
the issue said the model could not express — and it is now the ordinary one.
|
||||
|
||||
## How it is checked
|
||||
|
||||
`internal/catalogue/endpoints_setting_test.go`: a block says port, label and reach; a block may say
|
||||
only a reach; a name the module does not declare is refused; a reach outside the four values is
|
||||
refused; and saying the same thing twice — once in the block, once through the older per-port keys —
|
||||
is refused rather than resolved by whichever is read last. The proxy's half is `policy_test.go` and
|
||||
`authority_test.go`: a name the mesh did not send is not certified, by either authority.
|
||||
|
||||
## What is left, and it is not this
|
||||
|
||||
The older keys (`ports`, `expose`, and reach keyed by port) still work beside the block. They are
|
||||
what the block replaces, and retiring them is its own small change — not a gap in what reach can
|
||||
say.
|
||||
@@ -74,18 +74,3 @@ every future change of this shape, and it was paid today.
|
||||
argues for the former.
|
||||
- Does the same gap apply to the launcher and the units beside the binary, which are also files no
|
||||
declaration names?
|
||||
|
||||
## What now waits on this (2026-09-30)
|
||||
|
||||
[Issue 107](../107-a-declaration-carries-no-order/00-report.md) — a declaration carries no order, so a
|
||||
host cannot tell an older one from a newer. Its fix adds a field to the declaration, and a host refuses a
|
||||
declaration carrying a field it does not know, **whole**. So it is a flag day: every host upgraded, then
|
||||
the controller.
|
||||
|
||||
With nothing delivering a host, that means placing a binary by hand on every machine and unparking the
|
||||
adopted one, and a machine missed in the sequence is a machine the mesh cannot send anything to at all.
|
||||
107 is held for that reason rather than for anything in its own diagnosis.
|
||||
|
||||
**This is what makes a declaration field cost a rollout instead of an expedition**, which is a use for
|
||||
this record beyond keeping machines current: it is the thing standing between the mesh and its own
|
||||
protocol evolving.
|
||||
|
||||
@@ -1,185 +0,0 @@
|
||||
# 142 — the mesh can now build and publish its own host
|
||||
|
||||
*2026-09-30. Both of ADR 0141's named gaps are closed; the host is not yet placed by the mesh.*
|
||||
|
||||
## What ADR 0141's insight named, and what each cost
|
||||
|
||||
> The remaining work is a way to build the host and a way to name a version in a path, and until both
|
||||
> exist nothing delivers a version and every machine takes the fallback.
|
||||
|
||||
**1. Nothing could compile it.** The toolchain list was a closed set of typescript and python, and its
|
||||
own warning — every language is another implementation of the contracts modules share, so adding one
|
||||
commits to keeping N implementations in step — does not attach to Go. Go is how the host, the control
|
||||
plane and the builder are written, and none of them is a module in that sense: the host is what
|
||||
*applies* modules.
|
||||
|
||||
Closed by a Go toolchain naming a new base module, `mesh-tools-go`. Named and not pinned, so the mesh
|
||||
answers with the copy it holds and moving compiler is a build rather than an edit to the control
|
||||
plane's source ([ADR 0044](../../02-DECISIONS/0044-a-public-name-is-provisioned-like-any-capability.md),
|
||||
[0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md)).
|
||||
|
||||
**2. A version could not reach the path.** An archive named a fixed path and nothing interpolated the
|
||||
build into it, so nothing could ask for `…/versions/<version>/`.
|
||||
|
||||
Closed by `${version}` in any value of a resource that uses an archive or a bundle. **The version is
|
||||
the artifact's digest, short, and not the commit**: two builds of one commit are meant to be the same
|
||||
bytes — the toolchains are `-trimpath` for that — so a content-addressed version means an unchanged
|
||||
build resolves to the path it already had, where a commit-named path would move for an identical binary
|
||||
and recreate everything reading it.
|
||||
|
||||
## Three more things were in the way, and none was in the record
|
||||
|
||||
Found by doing it, in the order they appeared:
|
||||
|
||||
- **`sourcesFor` turned every entrypoint into a `.ts` file.** One language's file extension, written
|
||||
into the code that serves every language. The extension is the toolchain's now.
|
||||
- **The output directory was left to the compiler.** `tsc --outDir` makes one; `go build -o` writes into
|
||||
a directory and does not create it, failing with a message about a path rather than about a build.
|
||||
Made for every toolchain, because which compilers are forgiving is not something a reader should have
|
||||
to know.
|
||||
- **A bundle was refused if it named what it is built from.** The reason — a bundle is the module's own
|
||||
directory compiled whole — holds for an interpreted language and cannot hold for a compiled one: a Go
|
||||
repository carries several commands, the host and its bootstrap among them, and "the module's own
|
||||
directory" is then not a package at all. A compiled bundle may now say which package; the refusal
|
||||
stands for every interpreted one.
|
||||
|
||||
And a fourth, in the base module itself: its first Dockerfile ran `apt-get`, and the golang image the
|
||||
mesh holds is Alpine. That is
|
||||
[issue 136](../136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md) in an image
|
||||
rather than on a machine, and the build refused rather than a module failing later — which is the
|
||||
behaviour that issue wants.
|
||||
|
||||
## Measured: the mesh compiles its own host and publishes it to its own registry
|
||||
|
||||
```
|
||||
$ mesh-controller build <the host's repository>
|
||||
host-arch bundle artifact-store://mesh-host/host-arch/blobs/sha256:ad62528c…
|
||||
|
||||
mesh-host 1, built on novox from b5196e97
|
||||
```
|
||||
|
||||
Fetched from the mesh's registry and opened:
|
||||
|
||||
```
|
||||
mesh-host: ELF 64-bit LSB executable, x86-64, statically linked, stripped
|
||||
$ ./mesh-host help
|
||||
mesh-host — the node host
|
||||
```
|
||||
|
||||
Statically linked matters: what a machine holds is a file rather than a container, so a binary needing
|
||||
a libc it did not bring is a delivery that works until a machine differs.
|
||||
|
||||
**The host is a module now**, with one bundle for `arch` built from `cmd/mesh-host`. A system is
|
||||
required for a compiled artifact because a binary is pinned at link time so a host refuses to touch a
|
||||
machine it was not built for ([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)); `arch` is what all
|
||||
four of this mesh's machines report themselves to be, and another system is another artifact and
|
||||
another build.
|
||||
|
||||
## What is left
|
||||
|
||||
**The host module declares no resources**, so nothing places the built bundle on a machine yet. That is
|
||||
the next piece and it is the one with the interesting question in it: the resource is an archive
|
||||
unpacked to `…/versions/${version}/`, applied by the host that is running, and the bootstrap is not
|
||||
circular because the two are different versions ([ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md)).
|
||||
The host half of that mechanism — versions side by side, the newest runs, the running one stands aside
|
||||
between reconciles, rollback picks a directory — is built and tested and has never had a version to
|
||||
work on.
|
||||
|
||||
**One thing worth noticing while writing it.** The systems list exists because "the difference between
|
||||
two of them is a C library, not a kernel" — and a static Go binary has no C library. So one build would
|
||||
in fact run on all three. The pin is then a policy (a host refuses a machine it was not built for)
|
||||
rather than a necessity, which is a reasonable thing to keep and is worth knowing is a choice.
|
||||
|
||||
**And a sharper one, asked as a question and worth its own record.** The artifact's system is validated
|
||||
and then read by nothing: it does not reach the compiler, no machine is matched against it, and nothing
|
||||
chooses between two artifacts by it. The host built here is x86-64 because the build machine is, not
|
||||
because anything in the declaration said so — correct for this mesh by coincidence. That is
|
||||
[issue 159](../159-an-artifacts-system-is-checked-and-then-ignored/00-report.md).
|
||||
|
||||
## Delivered, started, and one fact short (2026-09-30, later)
|
||||
|
||||
The loop closed. The launcher is delivered as a **file** resource rather than inside the archive, and
|
||||
that difference is the safety: a file is written atomically — temp file, then rename — so the running
|
||||
launcher keeps the inode it was started from, where an archive writes in place with truncate and would
|
||||
cut the script a running shell is reading. The manifest carries a second copy of the launcher and a
|
||||
test refuses any difference from `packaging/nox-mesh-host-launch`.
|
||||
|
||||
On the workstation, in order: the version landed, the launcher was replaced, the running host saw a
|
||||
delivered version and stood aside, and after one restart of the unit the launcher started
|
||||
`/usr/lib/nox-mesh-host/versions/637f65559d16/nox-mesh-host`. **A host the mesh compiled, published,
|
||||
delivered and started.**
|
||||
|
||||
It would then have refused the first declaration it was asked to apply. The host's Makefile links in
|
||||
two facts the mesh's toolchain does not, and one of them — the system it was built for — is read before
|
||||
anything is applied. That is
|
||||
[issue 161](../161-a-delivered-host-carries-none-of-its-link-time-facts/00-report.md), and the machine
|
||||
is back on its hand-placed binary until it is answered.
|
||||
|
||||
**The fallback is what made that safe**, and it was not luck: the launcher runs the pinned version, or
|
||||
the newest delivered one, or the one placed by hand — so moving the delivered versions aside restored
|
||||
the machine in one step.
|
||||
|
||||
## Self-update works (2026-09-30, end of the day)
|
||||
|
||||
```
|
||||
running: /usr/lib/nox-mesh-host/versions/093231796eb0/nox-mesh-host
|
||||
mesh-controller node show shanks
|
||||
host 093231796eb0
|
||||
```
|
||||
|
||||
One machine runs a host the mesh compiled, published, delivered and started, applying declarations and
|
||||
reporting the version it was delivered as. The two facts a delivered binary was missing are
|
||||
[issue 161](../161-a-delivered-host-carries-none-of-its-link-time-facts/01-resolution.md) and resolved:
|
||||
the system comes from the artifact, the version from where the binary sits.
|
||||
|
||||
**Three machines still run a hand-placed host**, and rolling each forward is one assignment and one
|
||||
push. The control node is worth last.
|
||||
|
||||
One thing this found on the way out: an archive cannot be undeclared, and the attempt stops the machine
|
||||
applying anything at all — [issue 162](../162-an-archive-cannot-be-undeclared/00-report.md). It is how
|
||||
undoing the first delivery froze the workstation, and it is not specific to the host.
|
||||
|
||||
## Every machine self-updates (2026-09-30, evening)
|
||||
|
||||
```
|
||||
shanks 76f4566bef3d/nox-mesh-host active
|
||||
g14 76f4566bef3d/nox-mesh-host active
|
||||
novox 76f4566bef3d/nox-mesh-host active
|
||||
ace 76f4566bef3d/nox-mesh-host active
|
||||
|
||||
mesh-controller status: (no host split)
|
||||
```
|
||||
|
||||
The last delivery was unattended on all four: the fixed host was built, pushed, each machine stood
|
||||
aside exactly once for the genuinely newer version, and the delivered launcher started it — no
|
||||
restart by hand. A following push that delivered nothing new was applied and reported by every
|
||||
machine and stood nobody aside, which is the check
|
||||
[issue 163](../163-a-delivered-host-stood-aside-on-every-push-and-reported-nothing/00-report.md) asks
|
||||
for.
|
||||
|
||||
**Two more faults on the way, both mine, both found by reading the machine rather than the success
|
||||
line.** A delivered host compared the newest delivered version against its link-time stamp rather
|
||||
than the version it was running, so it stood aside on every push and — because standing aside cancels
|
||||
the report — never reported again (163). And the adopted machine kept its found launcher as the
|
||||
adoption rule says, so the delivery there needed a `take` before the launcher moved.
|
||||
|
||||
**The crossover needs one restart of the unit per machine, once.** The launcher process that was
|
||||
running on each machine was the old script, executing from its own inode; a new file beside it
|
||||
changes nothing until the unit restarts. Every subsequent delivery is unattended.
|
||||
|
||||
**Timing, measured:** on a machine, hearing a declaration to reporting it applied is about three
|
||||
seconds. A push as the operator sees it takes 17–20 seconds, and the difference is the control plane
|
||||
composing the declaration before it sends. A `--wait` shorter than that reads as "did not report" for
|
||||
a machine that did; the three-minute default read as slowness for a machine that never would. Neither
|
||||
number is a defect being chased here, and both are worth knowing before reading a push's answer.
|
||||
|
||||
## What this leaves
|
||||
|
||||
- [Issue 162](../162-an-archive-cannot-be-undeclared/00-report.md): an archive cannot be undeclared, so
|
||||
the host module — and any module with an archive — cannot be unassigned, and trying stops the machine
|
||||
applying anything.
|
||||
- [Issue 107](../107-a-declaration-carries-no-order/00-report.md) is unblocked: a declaration field is
|
||||
now a build and a push rather than an expedition.
|
||||
- Three stale version directories on the workstation from the first attempts, moved aside under
|
||||
`/var/lib/mesh-host/versions-held-back/`, and a backup of the adopted machine's hand-placed binary
|
||||
beside its state. Both are safe to delete and are not the mesh's to delete.
|
||||
+1
-1
@@ -4,7 +4,7 @@ opened: 2026-09-29
|
||||
located-in:
|
||||
- mesh-controller internal/catalogue/filtering.go (fixed for this instance)
|
||||
- mesh-controller (what status reports, and what it does not ask)
|
||||
fixed-by: partly — mesh-controller 1da96e8 makes the report state its own scope; nothing dials a provision yet, which is ADR 0146 and is not built
|
||||
fixed-by:
|
||||
amended-design: 03-DESIGN/01-to-be/10-delivery.md
|
||||
---
|
||||
|
||||
|
||||
-64
@@ -1,64 +0,0 @@
|
||||
# 145 — partly resolved: the report says what it is not a claim about
|
||||
|
||||
*2026-09-30.*
|
||||
|
||||
## What was done
|
||||
|
||||
The sentence that was true for eleven hours now states its own scope, immediately below itself:
|
||||
|
||||
```
|
||||
4 machine(s), all doing what they were told, all heard from, running what the mesh would send
|
||||
them, and every module current with its source
|
||||
|
||||
That is the mesh and the machines agreeing. Nothing here dials a provision:
|
||||
no grant the mesh composed has been tested, so a module unable to reach what it
|
||||
requires would not appear above (04-ISSUES/145)
|
||||
```
|
||||
|
||||
That is the whole of what this change does, and it is deliberately small. It does not check anything.
|
||||
It closes the distance between *the machines are as the mesh described them* and *it works* by naming
|
||||
it, and that distance is where the eleven hours went: the report was read as the second and only ever
|
||||
meant the first.
|
||||
|
||||
Two other reports in this group now carry real information they did not
|
||||
([issue 125](../125-a-hold-is-not-a-line-in-the-apply-report/00-report.md): a hold is a line in the
|
||||
apply report and breaks the all-well sentence;
|
||||
[issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md): the mesh knows which
|
||||
host runs a machine). Neither would have caught this fault, and both were the same shape of blindness.
|
||||
|
||||
`printStatus` is now separated from the asking, so these words can be read by a test with no store, bus
|
||||
or machine — they have been acted on and been misleading twice, which makes them worth holding still.
|
||||
|
||||
## What is NOT done, and why this record stays open
|
||||
|
||||
**Nothing dials a provision.** [ADR 0146](../../02-DECISIONS/0146-connectivity-is-checked-by-name-per-hosting-form.md)
|
||||
decides how it should be done — a module on every machine serving an endpoint of its own and dialling
|
||||
every other machine's, one name per hosting form, over TLS with the certificate verified. **Nothing is
|
||||
built**, the module that existed was deleted, and the work was deferred deliberately by the operator.
|
||||
It is not this record's to start.
|
||||
|
||||
So the measured fault of this issue — that the mesh can only report on itself — is unchanged. What
|
||||
changed is that the report no longer implies otherwise.
|
||||
|
||||
## The open questions, where they stand
|
||||
|
||||
- *Should a grant be checked, and from where?* Answered by ADR 0146 and not built: from the position
|
||||
the callers are in, by a module on every machine, per hosting form.
|
||||
- *What would it cost to be wrong in the other direction?* Unanswered and important. A check that
|
||||
reports a provision broken while it works trains a reader to ignore the report, which is the failure
|
||||
this whole issue is about arriving by the other door.
|
||||
- *What should `status` say about a machine whose modules cannot reach each other?* Answered in part:
|
||||
until something checks, it says that it has not checked. What it says when a check exists is ADR
|
||||
0146's to settle.
|
||||
- *Is there a cheaper signal than a probe?* Still open. The affected module logged the failure 6,154
|
||||
times; the mesh reads no module's logs and arguably should not, but something a module could *say*
|
||||
about its own provisions would have surfaced this in minutes.
|
||||
- *Does the same blindness apply to a provider that lost a consumer's grant?* Still open, and still
|
||||
nothing checks it.
|
||||
|
||||
## One thing worth carrying forward
|
||||
|
||||
**The certificate half of ADR 0146 is now possible where it was not.** Its check requires an internal
|
||||
name fetched over TLS with the certificate verified, and until 2026-09-30 no machine trusted the mesh's
|
||||
authority at all ([issue 129](../129-nothing-makes-a-machine-trust-the-meshs-authority/02-resolution.md)).
|
||||
Three of four do now. Whoever builds 0146 no longer has to solve that first.
|
||||
-12
@@ -71,15 +71,3 @@ against a suite nobody can execute.
|
||||
- The lab rewrites a bundle's registry-prefixed third-party references to upstream ones for a
|
||||
machine with an uplink (`test/integration/harness.ts`); the new bundle's bus reference needed
|
||||
that rule added, which is done and is not this issue.
|
||||
|
||||
## What one of its fixes then did to a running mesh (2026-09-30)
|
||||
|
||||
The change that stopped the doubling — putting the stream into a push consumer's delivery subject —
|
||||
is correct on a foundation being raised and fatal on a mesh that is already running: the server will
|
||||
not move that subject while a subscriber is bound, and a node is bound to its declaration consumer
|
||||
the whole time it is up. The control plane crash-looped on the first build that carried it.
|
||||
|
||||
Recorded and fixed as [issue 156](../156-moving-a-consumers-delivery-subject-stops-the-control-plane/00-report.md).
|
||||
Noted here because this record is where somebody will arrive when reading why the subject carries the
|
||||
stream at all, and the answer is incomplete without it: **the raise path was the only one exercised,
|
||||
and it is the one path on which nothing is bound.**
|
||||
|
||||
-88
@@ -82,100 +82,12 @@ managing, or genesis carries a user list that includes the first node's enrolmen
|
||||
takes over from there. Both are decisions, not patches, and both belong to the genesis step that was
|
||||
deliberately left until last.
|
||||
|
||||
## 5 — the composed user list has to be placed by hand at genesis *(fixed)*
|
||||
|
||||
The account a token is the password of is **not recorded at all**: the composer names an enrolment
|
||||
user for every machine with a live token, nothing minted a credential for it, and the composition
|
||||
left it out as a user with no password. The comment above the issuing code already claimed
|
||||
otherwise — *"the account is created before the token is handed over"* — which is how it went
|
||||
unnoticed. Issuing a token now records that account, with the token's own secret as its password,
|
||||
because that is the string the machine will present.
|
||||
|
||||
Placing it is the other half. The list reaches the machine running the bus in that machine's
|
||||
declaration, which a machine that has not enrolled does not get, so at genesis it cannot arrive
|
||||
that way. **The control plane composes and says what it composed** — `broker accounts`, to standard
|
||||
output — and whoever is raising the machine writes it beside the bus's configuration and makes the
|
||||
server re-read it. Twice, because two accounts come into existence at different moments: the
|
||||
enrolment when the token is issued, and the machine's own when it enrols. A control plane that
|
||||
wrote the file itself would have to know where the bus keeps its configuration and how to make it
|
||||
reload, which is the module's knowledge and is what the module takes over on the first push.
|
||||
|
||||
With that, **a first node enrols against the bus it just raised** — measured, from bare, in the
|
||||
lab.
|
||||
|
||||
## 6 — and is enrolled twice, keeping a credential the mesh has replaced *(fixed)*
|
||||
|
||||
```
|
||||
mesh-controller: enrolled anchor
|
||||
mesh-controller: enrolled anchor (the same second)
|
||||
```
|
||||
|
||||
One `enrol` on the machine, two enrolments in the control plane. Each mints the node a fresh bus
|
||||
password and returns it; the machine keeps the answer to the first, and the mesh keeps the hash of
|
||||
the second. The machine then reconnects for ever as a user whose password the mesh rotated out from
|
||||
under it — *authentication error - User "node.anchor"* on the bus, `Authorization Violation` in the
|
||||
host's log, and a node that never reports.
|
||||
|
||||
What is ruled out: the host asking twice — it asks again only when the mesh says *try again*, and
|
||||
a refused attempt is not logged as an enrolment. Redelivery by the consumer — there is one
|
||||
consumer, its acknowledgement window is thirty seconds, and the handler is quick.
|
||||
|
||||
What is left: the client re-publishing when an acknowledgement is slow, which is what its defaults
|
||||
do. That was addressed by giving the publish a message id derived from its own bytes, so the stream
|
||||
discards the copy — **and the duplicate survived it**, so either the id is not reaching the stream
|
||||
or the second copy is not a copy. This is where the trail stops.
|
||||
|
||||
Worth saying plainly: **the mint is the fragile part, not the delivery.** An enrolment answered
|
||||
twice is survivable if the answer is the same both times, and it cannot be — the mesh keeps only
|
||||
the hash, so a second answer is necessarily a different credential.
|
||||
|
||||
**Found, and it is not about enrolment at all** *(fixed)*. The bus's own counters settled it: one
|
||||
message published, one held in the stream, one delivery, nothing redelivered — and the controller
|
||||
enrolled the machine twice. So the handler ran twice on one delivery.
|
||||
|
||||
A push consumer delivers onto an ordinary subject, and **everything subscribed to that subject gets
|
||||
a copy**. The controller holds a consumer called `controller` on CONTROL and another called
|
||||
`controller` on EVENTS, and the delivery subject was derived from the consumer's name alone — so
|
||||
both were `_DELIVER.controller`, the one process held both subscriptions, and every message from
|
||||
either stream was acted on twice.
|
||||
|
||||
Enrolment is where it drew blood, because enrolling twice mints twice and the second credential
|
||||
replaces the first. But it applied to **every report and every event the controller follows**, and
|
||||
it is the kind of fault that leaves no trace: nothing is redelivered, no counter is wrong, the work
|
||||
simply happens twice. The comment in the receiving code about a merge that ran the whole catalogue
|
||||
five times over on 2026-09-28 is the same shape seen from the other end.
|
||||
|
||||
The stream is in the delivery subject now, because the pair is what identifies a consumer — the
|
||||
server scopes a durable's name to its stream, and this subject was the one place that scoping was
|
||||
dropped. A subscriber's permission gains the same shape, keeping the bare name so an existing
|
||||
consumer keeps working until the controller's next assertion moves it.
|
||||
|
||||
*How it is checked:* the consumers the mesh asks for are asserted to deliver onto distinct subjects,
|
||||
in the controller's own suite. Against a server it would be invisible, which is the point.
|
||||
|
||||
## Where it belongs
|
||||
|
||||
`mesh-host` (the bundle and the enrolment path) and `mesh-controller` (the certificate command, and
|
||||
the composition that cannot reach the bus at genesis). Three of the four are fixed on branches; the
|
||||
fourth is the genesis work.
|
||||
|
||||
## What made it slow, and what was changed so it is not
|
||||
|
||||
Six faults behind one another, each found by raising a machine and reading what it said. What cost
|
||||
the most was not the faults:
|
||||
|
||||
- **Every bed's own instructions named the bundle that cannot work**, so the first three attempts
|
||||
ended in a control plane crash-looping on a missing bus. They name the working one now.
|
||||
- **A host binary built without its system** refuses everything it is given with *this host was
|
||||
built for ""*, which reads like a broken bundle. The lab's README says so.
|
||||
- **`make image` in the control plane had been broken for as long as its base was pinned**: the
|
||||
Dockerfile's fallback is a Go older than the module asks for, and the pipeline never saw it
|
||||
because the pipeline passes the declared base in. It reads the base from the manifest now.
|
||||
- **Leaving the machine standing is what answers the question.** Every finding above came from
|
||||
shelling in afterwards — the host's log, the bus's log, the file the bus was actually handed —
|
||||
and none from the test's own output, which says only that nothing converged. The bed takes
|
||||
`MESH_LAB_KEEP`, and the README says to reach for it first.
|
||||
|
||||
## What it cost, for the next person
|
||||
|
||||
Every lab bed still names `foundation-first-node.lock` in its own instructions, and that bundle
|
||||
|
||||
@@ -1,75 +0,0 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-09-29
|
||||
located-in: [nothing in the mesh — the surface in use is the predecessor's, installed on the workstation]
|
||||
---
|
||||
|
||||
# 147 — the operator's tools still dial the bus that was removed
|
||||
|
||||
## What was observed
|
||||
|
||||
Every tool call an operator makes against the mesh fails, on every machine, with the same answer:
|
||||
|
||||
```
|
||||
AMQP not connected — cannot reach hal/mesh@novox
|
||||
AMQP not connected — cannot reach hal/mesh@shanks
|
||||
```
|
||||
|
||||
The mesh moved to one bus and the previous transport was deleted
|
||||
([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md), cut over
|
||||
2026-09-28). The surface an operator drives the mesh through — the tool bridge their assistant
|
||||
speaks to — still opens an AMQP connection, so it cannot reach anything. Not one node answers,
|
||||
including the machine the operator is sitting at.
|
||||
|
||||
**What that leaves.** The mesh itself is healthy: nodes are current, modules run, the bus carries
|
||||
the mesh's own traffic. What is gone is the way a person asks it anything. Every question — what a
|
||||
node runs, what is assigned, what a module's settings are — has to be asked by opening a shell on
|
||||
the machine and running the control plane's binary inside its container, which is precisely the
|
||||
path the tool surface exists to remove, and which nothing checks, records or permits.
|
||||
|
||||
It also silently changes how work gets done: an assistant told to use the mesh's tools finds them
|
||||
dead, falls back to `ssh` and `docker exec`, and the operator discovers the fallback rather than
|
||||
the fault.
|
||||
|
||||
## Why this is here and not a note in the knowledge base
|
||||
|
||||
The rule is that everything happens over the mesh's bus. A surface that cannot reach that bus is
|
||||
that rule enforced by nothing — and the mesh reports itself healthy throughout, because what broke
|
||||
is not a module, a node or a provision but the thing standing outside asking them questions.
|
||||
|
||||
The same cut-over removed the assistant's long-term memory (`recall`, `memorize`) for the same
|
||||
reason, which is why lessons from the last two days were written into this repository by hand.
|
||||
|
||||
## What would have prevented it
|
||||
|
||||
- **The tool surface as a consumer of the bus like any other**, so moving the bus moves it — rather
|
||||
than a separate bridge with its own connection settings that nothing resolves.
|
||||
- **A check that a tool call reaches a node**, run where the mesh's other checks run. Every check
|
||||
the mesh makes today is about the relationship between the mesh and a machine; none asks whether
|
||||
a person can ask it anything ([issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
|
||||
the same shape one level out).
|
||||
|
||||
## Diagnosed at once, because the answer was in the configuration
|
||||
|
||||
**The surface is not the mesh's.** The tool server the operator's assistant speaks to is the
|
||||
predecessor's brain, installed on the workstation and started as a local process, with the
|
||||
predecessor's broker URL — `amqp://…@<the control node>` — written into the assistant's own
|
||||
configuration. Nothing about it is a module: it has no manifest, no assignment, no seat, no account
|
||||
on the mesh's bus, and the mesh has never known it exists.
|
||||
|
||||
So nothing regressed. The mesh removed a transport that this program still dials, and the program
|
||||
was never part of the mesh to be moved. **The mesh has a tool model** — `tools` in a manifest, the
|
||||
calls a seat accepts, the subjects a module answers on, and a control-plane command that asks one —
|
||||
and **nothing publishes an operator-facing surface onto it.** What an operator uses is the thing
|
||||
that came before, kept alive by a URL in a file.
|
||||
|
||||
That is the issue, and it is larger than a broken connection: the way a person drives this mesh is
|
||||
outside the mesh.
|
||||
|
||||
## Evidence to carry into diagnosis
|
||||
|
||||
- The failure text names `hal/mesh@<node>`, so the bridge is resolving a node and then dialling the
|
||||
old transport.
|
||||
- It fails identically for the local machine, which rules out reachability and points at the
|
||||
transport alone.
|
||||
- The mesh's own traffic over the new bus is unaffected: nodes report, declarations apply.
|
||||
@@ -1,37 +0,0 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-29
|
||||
located-in: [mesh-controller cmd/mesh-controller]
|
||||
---
|
||||
|
||||
# 148 — a manifest outside this catalogue has no check
|
||||
|
||||
## What was observed
|
||||
|
||||
A module's manifest is validated by a **test** — `internal/catalogue`'s suite parses every manifest
|
||||
in the catalogue checkout beside it and fails on one it cannot resolve. That works, and it is how
|
||||
several real faults were caught before a machine saw them.
|
||||
|
||||
It is available to exactly one repository: this one. Somebody describing their own application in
|
||||
their own repository — the case
|
||||
[ADR 0037](../../02-DECISIONS/0037-where-a-module-lives.md) calls *the case that matters most* —
|
||||
has no check at all. They write a manifest, register it with a running mesh, and find out whether
|
||||
it is valid when the mesh refuses it, or later, when a machine applies something that resolved and
|
||||
should not have.
|
||||
|
||||
The same record asks for the answer: **a `module check` command on the control plane's binary**, so
|
||||
a manifest is checked by the tool rather than by a test that imports the tool's internals.
|
||||
|
||||
## What would have prevented it
|
||||
|
||||
Nothing prevents this; it was noticed and left. ADR 0037 named it on 2026-09-01 and the record sat
|
||||
`proposed` until 2026-09-29, so the missing half was never anybody's task.
|
||||
|
||||
## Evidence to carry into diagnosis
|
||||
|
||||
- `mesh-controller/internal/catalogue` — `ParseManifest` and `CatalogueProblems` are the check, and
|
||||
both are internal.
|
||||
- The catalogue-wide test is `TestEveryCatalogueManifestDeclaresWhatItMounts` and its siblings; they
|
||||
take a path from `MESH_CATALOG`, so the mechanism is already path-driven and not repository-bound.
|
||||
- `mesh-controller module add` refuses a bad manifest at registration, which is the same check far
|
||||
too late: by then it is in a running mesh's records.
|
||||
+15
-1
@@ -8,7 +8,7 @@ fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 153 — An adopted machine's data cannot be placed where it is
|
||||
# 149 — An adopted machine's data cannot be placed where it is
|
||||
|
||||
## What was observed
|
||||
|
||||
@@ -55,3 +55,17 @@ The two assignment halves 0112 decided: a setting that places a declared directo
|
||||
path on this node, and a setting that says where an access's data is — both validated like
|
||||
`endpoints` (unknown ids refused), and an access placed by the assignment still never created,
|
||||
chowned or removed.
|
||||
|
||||
## Addendum 2026-09-30 — whose the data is, not only where
|
||||
|
||||
The same modules need to run **as the owner of the data they access**: linuxserver images take
|
||||
`PUID`/`PGID`, and ace's library is `media:media` (`1001:2000`, group members `ace`, `n8n`, `media`).
|
||||
A manifest default is one value for every machine, and an assignment's settings do not reach a
|
||||
container's environment. Adding user and group management to the mesh would contradict ADR 0051 —
|
||||
the mesh owns nothing about the operator's data.
|
||||
|
||||
The data already says whose it is, and the host already looks at it when it confirms an access exists.
|
||||
Proposed: expose that as a fact the module can ask for, like `${port:…}` — e.g. `${access:<id>:uid}`
|
||||
and `${access:<id>:gid}`, read from the accessed path on the machine — so a media module declares
|
||||
`PUID=${access:media:uid}` and is right on every machine without the mesh creating, naming or
|
||||
chowning anyone.
|
||||
@@ -1,52 +0,0 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-29
|
||||
located-in:
|
||||
- mesh-controller internal/catalogue/declaration.go (contributions are emitted for an assigned module, taken or not)
|
||||
fixed-by:
|
||||
---
|
||||
|
||||
# 150 — A route is contributed before its module is taken
|
||||
|
||||
## What was observed
|
||||
|
||||
ace is adopted and runs `route-adapter` beside the predecessor's traefik (ADR 0104). The first web
|
||||
module migrated there was searxng. `assign ace searxng` + `push ace` — the step that is supposed to
|
||||
change nothing on an adopted machine (ADR 0100, hq 125: assign holds, take replaces) — took
|
||||
`searxng.zurag.be` down:
|
||||
|
||||
```
|
||||
https://searxng.zurag.be/ → 502
|
||||
```
|
||||
|
||||
for about five minutes, until the module was unassigned again.
|
||||
|
||||
## Why
|
||||
|
||||
Assign held everything it found on the machine — the predecessor's `searxng` container, its
|
||||
directories — exactly as designed. But the module's **route contribution** is not a resource on the
|
||||
machine, so nothing held it: it reached `route-adapter` at once, which wrote
|
||||
`mesh-searxng.zurag.be.yml` into traefik's file provider pointing at the mesh's assigned machine port
|
||||
(`http://ace.internal:20000`) — where nothing listened, because the container that would was held.
|
||||
traefik's file router for the name then won over the predecessor's docker-label router for the same
|
||||
name, and the name served a dead backend.
|
||||
|
||||
On this occasion the window was lengthened by an unrelated failure (the machine's docker address
|
||||
pools were exhausted, so the module's network could not be created), but the fault does not depend on
|
||||
it: **between assign and take, every routed module's public name points at a backend the mesh has
|
||||
deliberately not started.** The runbook's §6 step 2 ("check the preparation — HAL's service still
|
||||
serves") is false for every routed module on a node running route-adapter.
|
||||
|
||||
## What the operator did
|
||||
|
||||
Rolled back per runbook (unassign before take), then re-ran as assign → push → wait for the node to
|
||||
report its holds → take → push, keeping the window to the ~10 s between the two pushes. A workaround,
|
||||
not a fix: a module whose data must move between assign and take (runbook §6 step 3) cannot shrink
|
||||
its window this way.
|
||||
|
||||
## What would be right
|
||||
|
||||
A contribution from a module that is assigned but **not taken** on an adopted node should be held like
|
||||
the module's resources are — withheld from the provider until take — or the provider should be told
|
||||
the contributor is held and leave the predecessor's route alone. Either keeps "assign changes nothing"
|
||||
true for routed modules.
|
||||
@@ -1,106 +0,0 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-29
|
||||
located-in:
|
||||
- mesh-host internal/apply/apply.go (containerSpecReading hashes every `host` entry)
|
||||
- mesh-controller internal/catalogue/declaration.go (withMeshNames gives every container the mesh's names)
|
||||
fixed-by: mesh-controller PR 161 — the roster left every container's declaration and so its digest (ADR 0148 step 3, 2026-09-30)
|
||||
---
|
||||
|
||||
# 151 — A new name recreates every container in the mesh
|
||||
|
||||
## What was observed
|
||||
|
||||
Migrating one small module on ace (searxng) took four routine controller actions: `node public-domain
|
||||
ace zurag.be`, `assign ace searxng` + push, `unassign ace searxng` + push (a rollback, see
|
||||
[issue 150](../150-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take
|
||||
+ push again. Each push to ace also pushed novox ("this push left g14, novox, shanks behind … sending it
|
||||
too"). novox's host then **replaced every container it runs, twice**:
|
||||
|
||||
```
|
||||
22:46 … updated postgres.server (mesh-store): replaced; a container's configuration is fixed when it is created
|
||||
22:46 … updated distribution.store (mesh-registry): replaced; …
|
||||
22:46 … updated route-proxy.server (route-proxy): recreated …
|
||||
22:51 … updated postgres.server (mesh-store): replaced; …
|
||||
22:52 … updated gitea.server (gitea): replaced; …
|
||||
22:52–22:55 mesh-vault, builder, all of mailu, keycloak, minio, mongodb, invoicing, photos, umami, …
|
||||
```
|
||||
|
||||
Replacements per minute on novox, 22:46–22:55: 12, 8, 13, 5, 6, 8, 8, 8, 8, 4. The control plane was
|
||||
unreachable twice while its own store came back through crash recovery
|
||||
(`FATAL: the database system is starting up`), and dependants (umami) crash-looped until it did. None
|
||||
of the replaced containers belonged to the module being migrated, or to ace.
|
||||
|
||||
## Why (confirmed part)
|
||||
|
||||
Every container the mesh runs is given the mesh's names as `--add-host` entries
|
||||
(`withMeshNames`), and the host puts every entry into the container's spec digest — deliberately,
|
||||
since [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md): a container left alone when the roster moved kept a five-day-old
|
||||
address and restarted 2286 times. A running container cannot have its hosts changed, so a changed
|
||||
digest means a replace.
|
||||
|
||||
The consequence is that **the roster is part of every container everywhere**: anything that adds,
|
||||
removes or moves one name — a node's public domain, a routed name, an unassign — replaces every
|
||||
container on every machine that carries the list. On the hub that includes the control plane's store,
|
||||
the registry, the edge and mail.
|
||||
|
||||
## Not yet established
|
||||
|
||||
Which name moved in each pass. gitea's hosts after the fact list the `.internal` names and novox's
|
||||
internally routed `*.novox.be` names, and **not** `searxng.zurag.be` — so it is not simply "a routed
|
||||
name was added". Two passes suggest the set changed at assign and changed back at unassign; the
|
||||
declarations before and after would say, and nothing on the machine records the previous one.
|
||||
|
||||
## Why it matters now
|
||||
|
||||
The node-by-node migration adds names one module at a time. On ace alone that is ~25 web modules; if
|
||||
each assignment moves the roster, each is a full restart of every hub service, and a rollback is
|
||||
another. The migration is paused on this.
|
||||
|
||||
## What has since been ruled out as a cause (2026-09-30)
|
||||
|
||||
The roster was also moving for a reason that was not a name changing at all: a node whose plan would
|
||||
not compose had its routed names silently dropped from the roster handed to every machine, so a
|
||||
briefly unreachable store withdrew and restored a name on alternating passes. That is
|
||||
[issue 152](../152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md), and it is
|
||||
fixed — it is what made the churn on the control node repeat every six minutes for twenty-five
|
||||
minutes rather than once per operator action.
|
||||
|
||||
**It does not close this record.** A roster that changes for a real reason — a machine joining, a
|
||||
module assigned, a public domain set — still replaces every container on every machine that carries
|
||||
the list. 152 removed the false reasons; the question below is still open.
|
||||
|
||||
## What would be right (for diagnosis)
|
||||
|
||||
Keep 135's guarantee (no container runs with a stale address) without making the roster part of every
|
||||
container's identity — e.g. resolve mesh names through a resolver the container asks at lookup time
|
||||
rather than baked entries, or scope each container's entries to the names it actually binds.
|
||||
|
||||
## Answered (2026-09-30): the first of those two
|
||||
|
||||
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) takes
|
||||
the resolver, not the scoping. Scoping was the close call and was rejected for contradicting the
|
||||
standing intent that anything on the mesh can call anything on it: it would make a name reachable only
|
||||
where the mesh was told in advance it would be wanted, and it would leave the roster in the digest, so
|
||||
the churn returns whenever a widely-bound name moves.
|
||||
|
||||
Resolving at lookup time makes staleness impossible rather than detected — 109 and 135 stop being bugs
|
||||
that were fixed and become a shape that does not exist — and makes a name's blast radius nothing.
|
||||
|
||||
**This record stays open**, because the record answers it and the code does not. Nothing may stop
|
||||
copying names until a container can reach the resolver from any of the runtime's networks
|
||||
([issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)),
|
||||
which it cannot on two of four machines today. Removing the copy first reintroduces 109 and 135
|
||||
silently, on a live mesh, which is how both were found. The order is in the record.
|
||||
|
||||
## Resolved (2026-09-30)
|
||||
|
||||
Step 3 landed the day 110 did. mesh-controller PR 161 stops writing the roster into any container, so
|
||||
a container's digest no longer carries a name that is not its own. The controller's tests hold the
|
||||
record's check — a container's declaration does not move when the mesh's roster does, and does move
|
||||
when the module's own declared entries do.
|
||||
|
||||
The first push after the change recreated every container once, because every digest lost its host
|
||||
entries at the same moment. That was the last such event: from here a name added or moved on one
|
||||
machine changes no container anywhere, and the record's second check — add a routed name, watch every
|
||||
other machine's apply report say nothing changed — is what the next module assignment will show.
|
||||
@@ -1,124 +0,0 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-29
|
||||
located-in:
|
||||
- mesh-controller cmd/mesh-controller/plan.go (routeNamesInTheMesh skips a node whose plan will not compose)
|
||||
fixed-by: mesh-controller 6c5dfd0 (PR 147)
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 152 — A node whose plan will not compose silently removes its names from every machine
|
||||
|
||||
## What was observed
|
||||
|
||||
For twenty-five minutes, and for seventeen of them after the last operator action, the control-node's
|
||||
host applied all 327 of its resources every ~6.5 minutes without pause, replacing every container on
|
||||
the machine each time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail,
|
||||
and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine
|
||||
sat near 8. Nothing was converging: each pass ended and the next began four seconds later.
|
||||
|
||||
The two machines carrying no containers were not churning. They were only knocked off the bus each
|
||||
time the control node re-created it, reconnected, re-heard the same declaration and applied it again
|
||||
as a no-op.
|
||||
|
||||
## Why: the roster alternates between two values, and it is part of every container
|
||||
|
||||
Two consecutive declarations were compared by reading the `--add-host` entries of four containers
|
||||
the moment each pass created them:
|
||||
|
||||
```
|
||||
23:02 ace.internal drive.novox.be g14.internal keycloak.novox.be novox.internal
|
||||
office.novox.be portainer.novox.be shanks.internal umami.novox.be (9 names)
|
||||
23:07 … the same nine, and searxng.zurag.be (10 names)
|
||||
```
|
||||
|
||||
One routed name — belonging to a module on another machine entirely — leaves the roster and comes
|
||||
back. Because the roster is part of every container's spec digest ([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)),
|
||||
each flip is a different identity for every container on the machine, and a running container cannot
|
||||
have its hosts changed. So every flip replaces all of them.
|
||||
|
||||
**What makes it flip is a swallowed error.** `routeNamesInTheMesh` composes every node's plan to
|
||||
find the names it serves, and when one will not compose it moves on:
|
||||
|
||||
```
|
||||
plan, settings, err := planFor(ctx, open, n.Name)
|
||||
if err != nil {
|
||||
continue
|
||||
}
|
||||
```
|
||||
|
||||
A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does
|
||||
not know", but "the mesh states these names do not exist", to every machine at once.
|
||||
|
||||
## Why it sustains itself
|
||||
|
||||
The loop closes through the control plane's own database:
|
||||
|
||||
1. An apply replaces `mesh-store` — the store the control plane reads — by removing the container,
|
||||
so postgres comes back through crash recovery.
|
||||
2. While it recovers it refuses connections: `FATAL: the database system is not yet accepting
|
||||
connections / Consistent recovery state has not been yet reached` (observed, 21:08:39 UTC, every
|
||||
pass).
|
||||
3. `planFor` for the other machine fails against that store. `routeNamesInTheMesh` swallows it and
|
||||
drops its routed name.
|
||||
4. The roster changed, so all 327 resources differ, so all are replaced — including `mesh-store`,
|
||||
and including the bus, which is why the host also cannot report: `applied, and could not tell the
|
||||
mesh: reporting: nats: connection closed`.
|
||||
5. Back to 1.
|
||||
|
||||
Every pass destroys the evidence the next pass needs to decide it has nothing to do, and nothing
|
||||
outside the machine has to be wrong for it to continue.
|
||||
|
||||
**It is metastable, not permanent.** It ran from 22:46 to 23:11 — five full replacements of every
|
||||
container on the machine — and then stopped on its own, when one pass happened to read the store
|
||||
during a window it was up, composed the same roster twice running, and found nothing to do. Load fell
|
||||
from 7.8 to 1.7 and the machine returned to its five-minute idle tick.
|
||||
|
||||
That it ends by luck is the point, not a mitigation. The exit condition is a race the mesh does not
|
||||
control, the operator cannot see, and nothing reports; the same four actions on a slower machine, or
|
||||
a larger store, would not have found it. An outage that clears itself after twenty-five minutes and
|
||||
five restarts of the forge, the directory and mail is not a smaller fault than one that does not — it
|
||||
is the same fault, harder to catch.
|
||||
|
||||
## What it is not
|
||||
|
||||
- Not the operator's four actions on the other machine. Those explain the first passes
|
||||
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)); they were
|
||||
finished seventeen minutes and three full passes before these measurements.
|
||||
- Not a file that keeps changing. `/etc/hosts`, the bus's account list and the vault's export were
|
||||
hashed across passes and are byte-identical, and the directory reported `mode 755 to 700` every
|
||||
pass while already being `700`. Those resources are **misreported as changed** and are worth their
|
||||
own question, but they are not what moves a container's identity.
|
||||
- Not the lost report alone. A report that cannot be delivered explains a re-apply; it does not
|
||||
explain a re-apply that finds 327 differences.
|
||||
|
||||
## Why it matters beyond this outage
|
||||
|
||||
The same `continue` makes every routed name in the mesh conditional on every node's plan composing at
|
||||
the moment any machine is pushed to. One unreachable or half-migrated machine is enough to withdraw
|
||||
its names from everywhere — and the withdrawal is indistinguishable, on the receiving machine, from
|
||||
the operator having removed them.
|
||||
|
||||
The codebase already states the rule this breaks, forty lines away, about the same kind of lookup:
|
||||
|
||||
> A lookup failure is an error, never "not found": collapsing the two composed a declaration without
|
||||
> the trust whenever the inventory hiccuped, delivered by a push that reported success.
|
||||
|
||||
## How it was fixed, and how the fix is checked
|
||||
|
||||
`planFor` now marks the two failures that really are a statement about the node — its set not
|
||||
composing, and a setting that reaches nothing — and the three gatherers pass over those and only
|
||||
those. Every other failure is raised, naming the machine and the read.
|
||||
|
||||
Five tests hold it: a set that cannot compose is marked as the node's own; a store that cannot be
|
||||
read is *not*; one incoherent node still does not cost the rest their names; a roster is never
|
||||
returned beside an error; and the raised failure names what could not be read.
|
||||
|
||||
The two sibling gatherers were audited and fixed the same way — the grant composer, which would have
|
||||
withheld a consumer's credential, and the private-network membership, which would have taken a
|
||||
machine off the overlay. Three other `planFor` callers were audited and left alone: they refuse or
|
||||
report rather than silently withdraw, which is the safe direction.
|
||||
|
||||
**This does not close [issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md).**
|
||||
A roster that changes for a real reason still replaces every container in the mesh. This removes the
|
||||
false reasons; whether the roster belongs in a container's identity at all is that record's question.
|
||||
@@ -1,42 +0,0 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-29
|
||||
located-in:
|
||||
- hq 02-DECISIONS/0138 (reach: internal | public | both)
|
||||
- mesh-controller internal/catalogue/filtering.go (Reaches)
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 154 — A machine's own network is not a reach
|
||||
|
||||
## What was observed
|
||||
|
||||
Preparing ace's modules. ace sits on a home network (192.168.1.0/24) behind a router, and several of
|
||||
its services are reached **from that network by devices that will never be mesh machines**:
|
||||
|
||||
- mosquitto `1883` — an IoT light switch (`sonoff-office-light-switch`) and home-assistant;
|
||||
- unifi `8080`/`3478 udp`/`10001 udp` — the access points' inform, STUN and discovery;
|
||||
- plex `32400` — LAN streaming clients (three connected at survey time);
|
||||
- home-assistant `8123`, and the resolver on the LAN address.
|
||||
|
||||
ADR 0138 gives an endpoint's reach as `internal` (the private overlay), `public` (anywhere) or `both`.
|
||||
None of them says *this machine's own network*. The predecessor could: its unifi manifest opened
|
||||
inform/STUN/discovery `from: 192.168.0.0/16, 10.0.0.0/8, 172.16.0.0/12`.
|
||||
|
||||
## Consequence
|
||||
|
||||
The only reach that includes a LAN device is `public`. While ace is adopted that is harmless — its
|
||||
own firewall stays and admits the LAN — and behind NAT "anywhere" happens to mean the LAN. But:
|
||||
|
||||
- it states the wrong thing: an operator reading `reach: public` on an IoT broker believes it is on
|
||||
the internet, and a router port-forward added later for something else makes it so;
|
||||
- at `converge ace`, the mesh's filter is the sum of what it listens on (ADR 0045). An endpoint left
|
||||
`internal` cuts every LAN device off at the flip; one set `public` opens it to the internet on any
|
||||
machine with a public address.
|
||||
|
||||
## What would be right (for diagnosis)
|
||||
|
||||
A reach — or a source — that means the networks the machine is directly attached to (its uplink's
|
||||
subnets, as the machine reports them), so a LAN-only service is declared as exactly that and the
|
||||
filter can admit it without admitting the internet.
|
||||
@@ -1,58 +0,0 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-29
|
||||
located-in: [hq 00-META/checks/cycle.py]
|
||||
fixed-by: hq f89aef9 (PR 192)
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 155 — Two records may share a number, and every check passes
|
||||
|
||||
## What was observed
|
||||
|
||||
On 2026-09-29 two machines opened issues against this repository within the same hour. Both read
|
||||
`main` correctly and both took "the next free number", and they collided twice:
|
||||
|
||||
| | one machine opened | the other had already used |
|
||||
|---|---|---|
|
||||
| first | 147, 148 | 147, 148 on an unmerged branch |
|
||||
| second | 149, 150 | 149 on an unmerged branch, 150 from renumbering the first collision |
|
||||
|
||||
The first collision was reconciled by hand before merging. The second was **merged into `main`**, and
|
||||
`records.py`, `cycle.py` and `index.py` all reported success over a tree holding
|
||||
`149-a-declaration-that-shrinks-to-empty` beside `149-an-adopted-machines-data-cannot-be-placed-where-it-is`,
|
||||
and two folders numbered 150.
|
||||
|
||||
## Why
|
||||
|
||||
The number is allocated as `max(main) + 1`, and `main` lags every open pull request — seven of them
|
||||
that evening. Two readers of the same `main` therefore compute the same next number, and neither is
|
||||
doing anything wrong. The existing reconciliation precedent (a second record numbered 127 became 149)
|
||||
assumed a single writer, which stopped being true when a second machine began filing its own findings.
|
||||
|
||||
## Why it matters
|
||||
|
||||
An issue number is how every other record cites this one — `fixed-by:`, `located-in:`, a decision
|
||||
record's consequence, a commit message. Two records answering to one number is a citation that
|
||||
resolves to whichever folder the reader happened to open, and the failure is silent on both sides:
|
||||
the citer is not wrong, and the cited record exists.
|
||||
|
||||
It is also exactly the class this repository says it does not permit — a rule (`00-META/process/03-issues.md`:
|
||||
"take the next free number") enforced by nothing.
|
||||
|
||||
## How it was fixed, and how the fix is checked
|
||||
|
||||
`cycle.py` now refuses a tree in which two issue folders share a leading number, and names both.
|
||||
Proven by adding a duplicate and watching it fail, then removing it and watching it pass.
|
||||
|
||||
The colliding records were renumbered 153 and 154, in the branch that landed last — renumbering a
|
||||
branch whose author is still pushing only moves the race.
|
||||
|
||||
**The check catches the collision; it does not prevent it.** Allocating a number still needs the open
|
||||
pull requests read as well as `main`. That is a habit the check now backstops rather than one it
|
||||
replaces, and playbook [03](../../00-META/process/03-issues.md) now says so at the step where the
|
||||
number is taken.
|
||||
|
||||
[ADR 0080](../../02-DECISIONS/0080-the-development-cycle-is-checked.md) enumerates what `cycle.py`
|
||||
enforces and named four things; it carries a progressive insight naming the fifth. The decision
|
||||
stands — this is one more thing frontmatter and file names can carry, found by its absence.
|
||||
-98
@@ -1,98 +0,0 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-29
|
||||
located-in:
|
||||
- mesh-controller internal/broker/jetstream.go (EnsureConsumer)
|
||||
fixed-by: mesh-controller e7da39d (PR 148)
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 156 — Moving a consumer's delivery subject stops the control plane, and only on a mesh that is running
|
||||
|
||||
## What was observed
|
||||
|
||||
The control plane crash-looped, every restart ending the same way:
|
||||
|
||||
```
|
||||
mesh-controller: asserting how novox hears its declaration:
|
||||
bringing consumer novox on NODES to match: nats: consumer name already in use
|
||||
```
|
||||
|
||||
It came up on the first build of the controller in eight hours. The machines kept running what they
|
||||
already held — this stops the mesh being *changed*, not the services it placed — and nothing could be
|
||||
pushed, no report was consumed and no enrolment answered, for as long as it lasted.
|
||||
|
||||
## Why
|
||||
|
||||
[Issue 146](../146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md) put the
|
||||
stream into a push consumer's delivery subject, because one process holding two consumers of the same
|
||||
name on two streams was given one subject and acted on every message twice.
|
||||
|
||||
**The server will not move a push consumer's delivery subject while a subscriber is bound to it.** It
|
||||
refuses with `consumer name already in use` — a message about the name, for a conflict about the
|
||||
subject, which is why the trail starts in the wrong place.
|
||||
|
||||
A node is bound to its declaration consumer the whole time it is up. That *is* a node listening for
|
||||
what it should be. So every node consumer in a mesh that is running is one the assertion cannot bring
|
||||
to match — and the assertion happens before the controller serves, so it never serves.
|
||||
|
||||
The controller's own two consumers moved without trouble, and are on the new subject in the live mesh.
|
||||
It asserts them before it subscribes, so nothing was bound.
|
||||
|
||||
## Why nothing caught it
|
||||
|
||||
The change was exercised on a mesh being raised, where every consumer is created rather than updated
|
||||
and nothing is bound to any of them. On that path the code is correct. The test that would have caught
|
||||
it needs a mesh that is already running: an existing consumer, a subscriber still attached, and then
|
||||
the assertion.
|
||||
|
||||
Reproduced exactly that way before the fix — same server version, same stream shape, same consumer —
|
||||
and it fails with the same words as the machine did. An earlier version of the same test unsubscribed
|
||||
first and passed against the code that was crash-looping on the control node.
|
||||
|
||||
## What it is not
|
||||
|
||||
- Not the change it shipped beside. The merge that triggered this build carried
|
||||
[issue 152](../152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md)'s fix and three
|
||||
other commits that had never been deployed; this one is 146's.
|
||||
- Not a version difference. The test server and the mesh's broker are both nats-server v2.10.29.
|
||||
|
||||
## How it was fixed
|
||||
|
||||
The consumer that works is kept, and the assertion says so instead of failing.
|
||||
|
||||
**Not deleted and re-made.** Re-making moves the subject, and a holder may not yet be allowed to
|
||||
subscribe to the new one: the wider grant travels in the bus's user list, which the control plane
|
||||
composes and a machine applies minutes later. On the live mesh the nodes are granted `_DELIVER.<node>`
|
||||
and not `_DELIVER.<node>.>` — re-making would have silenced every machine in the mesh, which is worse
|
||||
than the collision it was fixing and far harder to undo. That was the first fix written here, and the
|
||||
permission is the reason it was not shipped.
|
||||
|
||||
**Not fatal**, which is what 146's change intended and did not do: the bare subject still delivers, and
|
||||
collides only where one holder has two consumers of one name. A node has one.
|
||||
|
||||
## How the fix is checked
|
||||
|
||||
Two tests against a real server: a consumer with a subscriber bound keeps its subject, is reported,
|
||||
and still delivers to that subscriber; a consumer with nothing bound moves, so 146's fix still applies
|
||||
where the collision actually was.
|
||||
|
||||
## What is left
|
||||
|
||||
The node consumers stay on the bare subject, which is correct and not tidy. Nothing is wrong while
|
||||
they do not move: one consumer per name per stream cannot collide with itself.
|
||||
|
||||
The controller reports each one it kept, and did, on the start that fixed this — four node consumers
|
||||
and the build machine's worker, which is bound the same way and was not anticipated here:
|
||||
|
||||
```
|
||||
consumer novox on NODES still delivers to "_DELIVER.novox" and not "_DELIVER.novox.NODES":
|
||||
nats: consumer name already in use. It keeps working; the subject moves on an assertion
|
||||
made while nothing is bound to it
|
||||
```
|
||||
|
||||
**The wider grant has since landed** (2026-09-30, measured on the mesh's own broker config): every
|
||||
node is now allowed `_DELIVER.<node>.>` as well as the bare subject. That was the thing missing when
|
||||
this was diagnosed, and it is why re-making the consumers then would have silenced every machine.
|
||||
What remains is only the second half — an assertion made while each node is detached from its
|
||||
consumer — and that is its own piece of work, not a side effect of a restart.
|
||||
@@ -1,70 +0,0 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-30
|
||||
located-in:
|
||||
- mesh-controller internal/catalogue/roster.go (the roster's entries for a routed name)
|
||||
fixed-by: mesh-controller PR 163 — a routed name is published as itself, once, with no suffixed alias (ADR 0151, 2026-09-30)
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 157 — A routed name is published with an `.internal` alias that nothing serves
|
||||
|
||||
## What was observed
|
||||
|
||||
Every machine's hosts file carries two entries for each routed name — the name, and the name with
|
||||
`.internal` appended:
|
||||
|
||||
```
|
||||
10.10.0.1 keycloak.novox.be.internal keycloak.novox.be
|
||||
10.10.0.1 drive.novox.be.internal drive.novox.be
|
||||
10.10.0.1 umami.novox.be.internal umami.novox.be
|
||||
```
|
||||
|
||||
The suffixed one resolves and is served by nothing. The proxy refuses it during the handshake, and
|
||||
says so exactly:
|
||||
|
||||
```
|
||||
http: TLS handshake error from 10.10.0.3:33480:
|
||||
no public route for "keycloak.novox.be.internal" in this mesh, so no certificate is asked for
|
||||
```
|
||||
|
||||
A client sees `curl: (35) TLS connect error ... tlsv1 alert internal error` and no peer certificate —
|
||||
a server-side refusal, with nothing in it to say the name was never real.
|
||||
|
||||
The name the proxy does serve is `<label>.<node>.internal` — `keycloak.novox.internal` — which is what
|
||||
[issue 139](../139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md) describes. So
|
||||
there are two internal shapes for one service, one of them composed by appending the suffix to a name
|
||||
that already has a domain.
|
||||
|
||||
## Why it matters
|
||||
|
||||
**It sends a reader to the wrong diagnosis.** Reproducing
|
||||
[issue 129](../129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md) on 2026-09-30, the
|
||||
first three names tried came from the hosts file, all failed with a TLS alert rather than the
|
||||
verification error 129 reports, and the evidence pointed at the proxy having lost its internal
|
||||
certificates — a regression that had not happened. The correct name reproduces 129 exactly. Several
|
||||
minutes went into a fault that did not exist, and the only thing that distinguished the two was reading
|
||||
the proxy's own log.
|
||||
|
||||
It is also a name in every machine's hosts file, and in every container's, that cannot be reached: the
|
||||
shape the mesh is otherwise careful about — writing a name that resolves to something that does not
|
||||
answer is worse than not writing it, because a connection to an address that does not answer hangs
|
||||
where a name that does not resolve fails at once ([ADR 0007](../../02-DECISIONS/0007-connectivity.md),
|
||||
and the same reasoning in `namesInTheMesh`).
|
||||
|
||||
## Where to look
|
||||
|
||||
The roster template renders one entry per name as `{{.FQDN}} {{.Name}}`, and `FQDN` is composed by
|
||||
appending the mesh suffix to the bare name. For a node that is right — `novox` becomes
|
||||
`novox.internal`. For a routed name the bare name is already a fully qualified public name, so the
|
||||
composition produces `keycloak.novox.be.internal`, which is not a name anything was told to serve.
|
||||
|
||||
The fix is a judgement about what a routed name's internal form is, and 139 is the record that asks it;
|
||||
this one is the evidence that the current answer publishes a third thing that is neither.
|
||||
|
||||
## Resolved (2026-09-30)
|
||||
|
||||
The judgement 139 asked for is [ADR 0151](../../02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md):
|
||||
a routed name has no mesh form. The roster now publishes it as itself, once, at the serving node's
|
||||
address; the `<domain>.internal` line is gone from every machine's hosts file, and a controller test
|
||||
refuses it coming back.
|
||||
-49
@@ -1,49 +0,0 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-30
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 158 — The proxy re-reads and re-logs every route it serves, every two seconds
|
||||
|
||||
## What was observed
|
||||
|
||||
The route proxy on the control node logs the whole of what it serves about every two seconds —
|
||||
measured 2026-09-30, **31 times in sixty seconds**, each line naming all 52 routes:
|
||||
|
||||
```
|
||||
23:27:19 serving 52 route(s): autoconfig.novox.be, autoconfig.novox.internal, … www.praktijkdespiegel.be
|
||||
23:27:21 serving 52 route(s): autoconfig.novox.be, autoconfig.novox.internal, … www.praktijkdespiegel.be
|
||||
23:27:23 serving 52 route(s): autoconfig.novox.be, autoconfig.novox.internal, … www.praktijkdespiegel.be
|
||||
```
|
||||
|
||||
Nothing is changing. The route set is identical every time.
|
||||
|
||||
## Why it matters
|
||||
|
||||
**It buries the only line that matters.** Between two of those entries sits the one error that explained
|
||||
a failing name:
|
||||
|
||||
```
|
||||
23:27:23 http: TLS handshake error from 10.10.0.3:33480:
|
||||
no public route for "keycloak.novox.be.internal" in this mesh, so no certificate is asked for
|
||||
```
|
||||
|
||||
One line of signal to roughly 5,000 characters of repetition, and the diagnosis it belonged to
|
||||
([issue 157](../157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md)) was found by
|
||||
grepping past it. A log that says the same true thing every two seconds is a log nobody reads, which is
|
||||
the same failure as one that says nothing — and it is the mesh's own accuracy rule pointed the other
|
||||
way: a report that is loud about the unchanging is not reporting.
|
||||
|
||||
Whether the *re-read* is also wasteful is secondary and unmeasured — it may be a cheap file stat. The
|
||||
logging is not in question.
|
||||
|
||||
## Where to look
|
||||
|
||||
Not localised. The proxy is `mesh-controller examples/route-proxy`; whether it re-reads on a timer or on
|
||||
a file watch, and whether it logs unconditionally or only on change, is the first thing to read. Saying
|
||||
what changed — or saying nothing — is the behaviour wanted, and the mesh already has the rule written
|
||||
down for its own reports: a log that is quiet on success and loud on failure reads as broken when it is
|
||||
working, and one that is loud always reads as nothing.
|
||||
@@ -1,100 +0,0 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-09-30
|
||||
located-in:
|
||||
- mesh-controller internal/catalogue/build.go (System is validated and read by nothing else)
|
||||
- mesh-controller internal/builder (the compile invocation names no target)
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 159 — An artifact's system is checked, and then nothing uses it
|
||||
|
||||
## What was observed
|
||||
|
||||
Asked whether the host is built for more than one architecture, 2026-09-30, having just built it
|
||||
through the new Go toolchain.
|
||||
|
||||
[ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md) says an
|
||||
artifact declares what it targets, and that one artifact per target is one build each. The manifest
|
||||
layer enforces the first half strictly: a bundle in a language that compiles to a binary **must** name
|
||||
a system, must name one of `alpine`, `android`, `arch`, and must not name one at all if its language is
|
||||
interpreted. A manifest that gets any of that wrong is refused with a reason.
|
||||
|
||||
**The field is then read by nothing.** Every use of it in the control plane is in the function that
|
||||
validates it. It does not reach the compiler, no machine is matched against it, and nothing chooses
|
||||
between two artifacts by it.
|
||||
|
||||
So the compile runs with no target named and produces a binary for whatever the build machine happens
|
||||
to be. The host, declared `system: arch` and built on this mesh's only build machine:
|
||||
|
||||
```
|
||||
mesh-host: ELF 64-bit LSB executable, x86-64, statically linked, stripped
|
||||
```
|
||||
|
||||
Correct for every machine in this mesh, which are all x86-64 Arch — and correct by coincidence rather
|
||||
than by anything the declaration did.
|
||||
|
||||
## Why it matters
|
||||
|
||||
**A module declaring two systems would get two identical binaries.** Both would be published, both
|
||||
pinned, both delivered, and the one sent to the machine it was not built for would fail at exec with a
|
||||
message about a format — which is the shape ADR 0005's link-time pin exists to prevent, arriving
|
||||
because the pin was never applied.
|
||||
|
||||
`android` in the list is the sharp end: it is not an x86-64 platform, and an artifact declared for it
|
||||
today would be an x86-64 binary wearing the label. Nothing would say so until a machine tried to run
|
||||
it.
|
||||
|
||||
**And the field reads as implemented.** It is required, validated against a closed list, and refused
|
||||
with a careful message — every signal a manifest author gets says the mesh is acting on it. A field
|
||||
that is checked and ignored is worse than one that does not exist, because the check is what persuades
|
||||
you it works.
|
||||
|
||||
## Two things this is not
|
||||
|
||||
- **Not the same axis as the distribution.** `alpine`, `android`, `arch` are what a machine reports
|
||||
itself to be, and the comment on the list says why: "the difference between two of them is a C
|
||||
library, not a kernel". The processor is a second dimension and the manifest has no word for it at
|
||||
all — so even a correct implementation of the current field would not answer the question that
|
||||
started this.
|
||||
- **Not urgent for this mesh.** Four machines, all x86-64 Arch, one build machine. Nothing is broken
|
||||
today and nothing will be until a machine differs — which is exactly how long a field like this stays
|
||||
invisible.
|
||||
|
||||
## The machine already says what it is
|
||||
|
||||
Found while filing [issue 160](../160-a-machine-says-little-about-itself-and-only-when-asked/00-report.md):
|
||||
every machine reports its architecture and kernel in the same profile that carries its capabilities, and
|
||||
the mesh keeps them.
|
||||
|
||||
```
|
||||
ace | amd64 | linux
|
||||
g14 | amd64 | linux
|
||||
novox | amd64 | linux
|
||||
shanks | amd64 | linux
|
||||
```
|
||||
|
||||
Nothing in the control plane reads either, and `node show` prints the capabilities beside them without
|
||||
printing them. **So a fix does not need a new fact from the machine** — matching an artifact's declared
|
||||
system against what a machine reported is possible today, and the missing piece is only the comparison
|
||||
and a compiler told what to target.
|
||||
|
||||
## Where to look
|
||||
|
||||
The compile invocation is assembled in `mesh-controller internal/builder`, and for Go it would need
|
||||
`GOOS`/`GOARCH` set from the artifact rather than inherited from the build machine. That needs the
|
||||
manifest to carry a processor as well as a system, or the systems list to mean both — which is a
|
||||
decision, not a fix, and belongs with whoever answers whether one static binary should serve several
|
||||
distributions at all.
|
||||
|
||||
**That last question is live.** The Go toolchain builds statically, so a single binary has no C library
|
||||
to differ about and would in fact run on Alpine and Arch alike. The per-system pin is then a policy — a
|
||||
host refuses a machine it was not built for — rather than a technical necessity, and worth knowing is a
|
||||
choice.
|
||||
|
||||
## How a fix is checked
|
||||
|
||||
An artifact declared for a system the build machine is not produces a binary for that system, shown by
|
||||
reading the file rather than by the build reporting success; and two artifacts declared for two systems
|
||||
do not have the same digest.
|
||||
@@ -1,92 +0,0 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-30
|
||||
located-in:
|
||||
- mesh-host internal/profile (what a machine collects about itself)
|
||||
- mesh-controller internal/inventory (what the mesh keeps of it)
|
||||
- mesh-controller cmd/mesh-controller (node show, which displays a part of it)
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 160 — A machine says little about itself, and only when the mesh asks it something
|
||||
|
||||
*Filed as a to-do rather than a fault: nothing is broken by it today. `open` is the status this
|
||||
repository has for "written down, nobody has started" — there is no `todo`.*
|
||||
|
||||
## What is wanted
|
||||
|
||||
**When a machine joins, the mesh should collect as much as it reasonably can about it**, and refresh
|
||||
that daily or thereabouts. It already asks what the machine *can do*; what it is made of is the same
|
||||
question one level down, and the mesh has no habit of asking it.
|
||||
|
||||
## What is already collected, which is more than it looks
|
||||
|
||||
A machine reports, and the mesh keeps:
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| eight capabilities | container-runtime, firewall, graphical-session, overlay, package-manager, privileged, seat, service-manager — each with a version or a reason it is absent |
|
||||
| **its architecture and kernel** | in the same profile, beside the capabilities |
|
||||
| which links face outside it | read from its own routing table on every apply |
|
||||
| the version of the host running on it | added 2026-09-30 |
|
||||
| what it found and is holding | on an adopted machine |
|
||||
| what is reachable on it | every listening socket and published port, on an adopted machine |
|
||||
| the firewall it was found with, and the tunnel it carried | on an adopted machine |
|
||||
|
||||
**The architecture and the kernel are already there and nothing reads them.** Measured on the mesh:
|
||||
|
||||
```
|
||||
ace | amd64 | linux
|
||||
g14 | amd64 | linux
|
||||
novox | amd64 | linux
|
||||
shanks | amd64 | linux
|
||||
```
|
||||
|
||||
`node show` prints the capabilities and not these two. No code in the control plane reads either.
|
||||
|
||||
## What is missing
|
||||
|
||||
**The facts.** Nothing is collected about memory, disk, the processor beyond its architecture, the
|
||||
distribution and its version, whether the machine is virtual or physical, its uptime, its timezone, how
|
||||
many cores it has. Several of those are what somebody actually wants when deciding where a module
|
||||
should go, and the comment on module placement already imagines them: *"`seat: card1-DP-1`, an
|
||||
architecture, an amount of memory"*.
|
||||
|
||||
**The refresh.** A machine publishes what it knows about itself **after an apply**, and the five-minute
|
||||
reconcile publishes nothing. So the mesh's picture of a machine is as old as the last time it sent that
|
||||
machine something — found the same day while checking host versions, where three machines read as
|
||||
having said nothing until each was pushed
|
||||
([issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/01-resolution.md)). A daily refresh is
|
||||
the same missing mechanism: a machine saying something on its own schedule rather than only when spoken
|
||||
to.
|
||||
|
||||
**Somewhere to read it.** Whatever is collected has to be visible, or it joins the architecture in
|
||||
being true and unread.
|
||||
|
||||
## The pattern this is the third instance of
|
||||
|
||||
Three times in one day the mesh was found to be collecting something and reading it nowhere:
|
||||
|
||||
- what an adopted machine holds — reported since the adoption work, and no surface counted it until
|
||||
[issue 125](../125-a-hold-is-not-a-line-in-the-apply-report/01-resolution.md);
|
||||
- the host version — reported for a week, and the control plane's own copy of the report did not have
|
||||
the field ([issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md));
|
||||
- the architecture and kernel — collected, stored, read by nothing, found today by being asked whether
|
||||
the host is built for more than one processor
|
||||
([issue 159](../159-an-artifacts-system-is-checked-and-then-ignored/00-report.md)).
|
||||
|
||||
**Collecting is the easy half and it is the half that gets done.** Whatever this issue adds should say,
|
||||
in the same breath, which surface shows it — or it will be the fourth.
|
||||
|
||||
## Why it is not urgent
|
||||
|
||||
Every machine in this mesh is `amd64` and every one reports so. One architecture is enough for now, and
|
||||
that is a decision rather than an oversight: a second processor is a second build of every component,
|
||||
and nothing needs one.
|
||||
|
||||
## How a fix is checked
|
||||
|
||||
A machine's own account of itself is visible in one place; it names at least what is listed as missing
|
||||
above; the newest of it is no older than a day on a machine nobody has pushed to; and a machine that
|
||||
cannot determine one of them says so rather than reporting a zero.
|
||||
@@ -1,95 +0,0 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-30
|
||||
located-in:
|
||||
- mesh-host cmd/mesh-host/main.go (version and builtFor, both set at link time)
|
||||
- mesh-controller internal/builder (the toolchain, which deliberately takes nothing from the module)
|
||||
fixed-by: mesh-controller (the system stamp, and one linker flag rather than two), mesh-host (the version read from the path) — verified on a machine 2026-09-30, 01-resolution.md
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 161 — A host the mesh built carries none of the facts its Makefile stamps in
|
||||
|
||||
## What was observed
|
||||
|
||||
*2026-09-30, on the workstation, having just made the host self-updating.*
|
||||
|
||||
The mesh compiled the host, published it, delivered it and the launcher started it. It ran, read the
|
||||
machine correctly, and **would have refused the first declaration it was asked to apply.**
|
||||
|
||||
The host's own Makefile links in two facts:
|
||||
|
||||
```
|
||||
LDFLAGS := -s -w -X main.builtFor=$(SYSTEM) -X main.version=$(VERSION)
|
||||
```
|
||||
|
||||
The mesh's Go toolchain links in neither, on purpose: a toolchain accepts nothing from the module,
|
||||
because anything a module could override there it would be writing a Dockerfile to override
|
||||
([ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md)). So a
|
||||
delivered host has `builtFor = ""` and `version = "development build"`.
|
||||
|
||||
**`builtFor` empty is the one that bites.** Before applying anything, the host asks which system it
|
||||
was built for:
|
||||
|
||||
```go
|
||||
sys, err := system.For(builtFor)
|
||||
```
|
||||
|
||||
and that answers, for an empty name:
|
||||
|
||||
```
|
||||
this host was built for "", which is not a system it knows. Built hosts are: …
|
||||
```
|
||||
|
||||
It is called before any resource is applied, so the failure is in the safe direction — the machine is
|
||||
not half-configured. It is still a host that cannot do its job, and nothing about it looks wrong: the
|
||||
unit is active, the link to the bus is up, and the log says it is hearing what the node should be.
|
||||
|
||||
Measured: after the crossover the machine logged nothing further, where the previous host had written
|
||||
a reconcile line every five minutes.
|
||||
|
||||
## Why this was found rather than reported
|
||||
|
||||
Nothing reports it. The host does not check its own stamps at start, the mesh does not ask, and the
|
||||
declaration that would fail is the same declaration that would deliver a fix — so **a machine in this
|
||||
state cannot be repaired by the mesh.** It was restored by moving the delivered versions aside and
|
||||
letting the launcher fall back to the hand-placed binary, which is the fallback working exactly as
|
||||
designed.
|
||||
|
||||
## What the records already say about half of it
|
||||
|
||||
[ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md) settles the
|
||||
version and its answer is not implemented:
|
||||
|
||||
> A component's version comes from where it sits, not from its linker. It is unpacked into a directory
|
||||
> named for its version, so it can read its own version from its path. The stamp goes, and with it the
|
||||
> need for a build to know what it will be called.
|
||||
|
||||
That is exactly right and would also fix what the mesh reports: a delivered host would say
|
||||
`637f65559d16` rather than `development build`, and
|
||||
[issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md)'s host comparison would
|
||||
mean something for delivered hosts.
|
||||
|
||||
**The system pin has no answer yet**, and it needs one before any mesh-built host can apply anything.
|
||||
The tension is real: the target is a property of the artifact and 0142 says so, but a toolchain that
|
||||
passed it would be linking a value into a variable whose name belongs to the module — which is the
|
||||
coupling the toolchain exists to avoid. Candidates, none decided:
|
||||
|
||||
- the path carries it as well as the version, so the host reads both from where it sits, as 0142 does
|
||||
for the version;
|
||||
- the bundle carries a small file beside the binary saying what it was built for, written by the
|
||||
builder from the artifact's declaration;
|
||||
- the host stops being pinned at link time and refuses on a fact it reads from the machine instead —
|
||||
which changes what [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) decided and is the biggest of
|
||||
the three.
|
||||
|
||||
## What is true in the meantime
|
||||
|
||||
Self-update works end to end and is one fact short of usable: the mesh builds the host, publishes it,
|
||||
delivers it to a machine, the running host stands aside, and the launcher starts the delivered one. The
|
||||
machine is left on its hand-placed binary until this is answered, which is one command to undo.
|
||||
|
||||
## How a fix is checked
|
||||
|
||||
A host the mesh built and delivered applies a declaration on a machine, shown by the machine's own
|
||||
reconcile line; and it reports a version that names the build it came from rather than a placeholder.
|
||||
@@ -1,59 +0,0 @@
|
||||
# 161 — resolved: a host the mesh built runs a machine
|
||||
|
||||
*2026-09-30. Measured on the workstation.*
|
||||
|
||||
```
|
||||
running: /usr/lib/nox-mesh-host/versions/093231796eb0/nox-mesh-host
|
||||
agent: active
|
||||
reconciles in the last six minutes: 1
|
||||
|
||||
mesh-controller node show shanks
|
||||
host 093231796eb0
|
||||
```
|
||||
|
||||
A binary the mesh compiled, published to its own registry, delivered over the bus, started by the
|
||||
launcher, applying declarations, and reporting a version that names the build it came from.
|
||||
|
||||
## The two facts, and where each now comes from
|
||||
|
||||
**The system it was built for comes from the artifact.** ADR 0142 already made the target a property
|
||||
of the artifact rather than of the recipe, so the toolchain names the variable it fills and the
|
||||
artifact supplies the value. It is the one thing a toolchain takes from a module, and it is stated
|
||||
rather than inferred.
|
||||
|
||||
**The version comes from where the binary sits**, which is what
|
||||
[ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md) decided and
|
||||
nothing had implemented: a delivered host reads the directory it was unpacked into. A host placed by
|
||||
hand keeps its link-time stamp, which is the honest answer for one the mesh did not deliver — and is
|
||||
every other machine today.
|
||||
|
||||
## Two mistakes on the way, both found by reading the output
|
||||
|
||||
**A repeated flag is not a merged one.** The stamp was appended as a second `-ldflags`, and the Go
|
||||
command takes the last and drops the first. The binary gained its system and lost `-s -w`: 12.2MB
|
||||
against 8.5MB, with its debug info. The comment I had written said the linker "accepts and merges"
|
||||
them. It does not. Linker flags are the toolchain's own list now, composed into one flag, and a test
|
||||
refuses a compile line that carries `-ldflags` itself.
|
||||
|
||||
**The delivered binary was named after its package.** `cmd/mesh-host` builds `mesh-host`; every
|
||||
machine runs `nox-mesh-host`, which is what the launcher looks for inside a version. The first
|
||||
delivery landed, reported `created … 1 file(s)`, and was invisible. An artifact says what its
|
||||
executable is called now.
|
||||
|
||||
Both were caught by listing the directory and reading the binary rather than believing the line that
|
||||
said it worked.
|
||||
|
||||
## What this cost while it was wrong, and what saved it
|
||||
|
||||
A delivered host that cannot apply is a machine the mesh cannot repair, because the declaration that
|
||||
would fix it is the declaration it cannot apply. The workstation was restored by moving the delivered
|
||||
versions aside so the launcher fell back to the hand-placed binary — **the fallback in the launcher,
|
||||
working exactly as designed**, and the reason this was an inconvenience rather than an expedition.
|
||||
|
||||
It also loops if you are not careful: the working binary applies, delivers a version, stands aside,
|
||||
and the broken one starts. Stopping the unit while the fix was built was the way through.
|
||||
|
||||
## What is left
|
||||
|
||||
**Three machines still run a hand-placed host.** Rolling them forward is one assignment and one push
|
||||
each, and the control node is worth doing last and watching.
|
||||
@@ -1,56 +0,0 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-09-30
|
||||
located-in: [mesh-host internal/apply (no removal for an archive)]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 162 — An archive cannot be undeclared, and trying stops the machine applying anything
|
||||
|
||||
## What was observed
|
||||
|
||||
*2026-09-30, unassigning the host module from the workstation to undo a delivery.*
|
||||
|
||||
```
|
||||
holding this machine: 0 applied, and map[apply:applying "mesh-host.next": no way to remove a "archive"
|
||||
0 resource(s) were applied and remain; everything was attempted, so what is not listed as failed was done.
|
||||
```
|
||||
|
||||
**Nothing was applied at all** — not the archive, not the other forty resources that had nothing to do
|
||||
with it. The machine stopped reconciling and stayed that way until the module was assigned again.
|
||||
|
||||
## Why it matters
|
||||
|
||||
Every other resource kind can be taken away. A file is removed and what was found under it is put
|
||||
back; a container is stopped and removed; a unit is given back the state it was found in
|
||||
([ADR 0118](../../02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md)). An
|
||||
archive has no removal at all, so:
|
||||
|
||||
- **a module with an archive can never be unassigned** — the attempt fails for ever;
|
||||
- **the failure takes the whole apply with it**, so the machine applies nothing else either, and one
|
||||
unassignable resource is a machine frozen against every other change;
|
||||
- it is silent in the mesh's terms: the push reported sent, and only the machine's own journal says
|
||||
what happened.
|
||||
|
||||
The host module is the obvious case and not the only one. An archive is for what inlining cannot
|
||||
serve — a theme, an icon set, a tree of configuration — and any module using one is in the same
|
||||
position.
|
||||
|
||||
## What the right answer probably is, and the question in it
|
||||
|
||||
The other kinds answer this by remembering what they found. An archive unpacks many files into a
|
||||
directory the mesh did not necessarily create, so removal has a real question in it: **remove what the
|
||||
archive put there, or remove the directory?** The first needs the applier to have recorded the file
|
||||
list; the second would delete whatever else lives there — and for the host's own versions directory,
|
||||
that is every other delivered version.
|
||||
|
||||
Recording what was unpacked is the answer that matches how the rest of the host behaves, and it is
|
||||
what [issue 126](../126-a-volume-path-is-not-in-the-spec-comparison/00-report.md) and ADR 0118 already
|
||||
argue for elsewhere: the mesh gives back what it found.
|
||||
|
||||
## How a fix is checked
|
||||
|
||||
A module with an archive is assigned, pushed, unassigned and pushed again; what the archive put on the
|
||||
machine is gone, anything that was in the directory beforehand is still there, and the apply that
|
||||
removed it applied everything else in the same declaration.
|
||||
-56
@@ -1,56 +0,0 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-30
|
||||
located-in: [mesh-host cmd/mesh-host/main.go (the successor check after an apply)]
|
||||
fixed-by: mesh-host PR 58 — the check asks with the running version, read from the binary's path, not the link-time stamp
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 163 — A delivered host stood aside on every push, and reported nothing
|
||||
|
||||
## What was observed
|
||||
|
||||
*2026-09-30, rolling the mesh-built host onto the last two machines.*
|
||||
|
||||
Every push to a machine running a delivered host produced, in order:
|
||||
|
||||
```
|
||||
host 093231796eb0 is delivered; standing aside so the launcher runs it
|
||||
applied 333 resource(s)
|
||||
applied, and could not tell the mesh: reporting: context canceled
|
||||
nox-mesh-host-launch: the host exited cleanly; starting it again
|
||||
nox-mesh-host-launch: running /usr/lib/nox-mesh-host/versions/093231796eb0/nox-mesh-host
|
||||
```
|
||||
|
||||
— for the version it was **already running**. It restarted itself on every push, for ever, and the mesh
|
||||
never received a single report from it: `node show` kept the version from before the crossover, and
|
||||
the operator's push waited its full three minutes for an answer that was never coming.
|
||||
|
||||
Read as healthy throughout: unit active, bus link up, "hearing what this node should be".
|
||||
|
||||
## Why
|
||||
|
||||
After an apply the host asks whether a newer host has been delivered than the one running, and the
|
||||
question was asked with the **link-time version stamp**. Since
|
||||
[issue 161](../161-a-delivered-host-carries-none-of-its-link-time-facts/01-resolution.md) a delivered
|
||||
host's version comes from where it sits and its stamp is `development build` — so the comparison never
|
||||
matched the newest delivered version, and "a newer host is waiting" was always true.
|
||||
|
||||
Standing aside cancels the context the report is published with, so the report was lost on every one
|
||||
of those applies. Two faults from one wrong argument.
|
||||
|
||||
The change that moved the version to the path was applied to the report and to the known-good record,
|
||||
and not here. Half a change, and the half left behind was the one that decides whether to exit.
|
||||
|
||||
## Why the three-minute wait made it invisible
|
||||
|
||||
The push's `--wait` timing out read as *slow*. It was not slow: **the report was never going to arrive.**
|
||||
The operator put it exactly: *if you don't get a response in five seconds, something is wrong.* A wait
|
||||
long enough to absorb a machine's whole apply is a wait long enough to hide that the machine never
|
||||
answered.
|
||||
|
||||
## How it is checked
|
||||
|
||||
A machine running a delivered host is pushed a declaration that delivers nothing new; it applies,
|
||||
reports, and does not stand aside. A machine running a delivered host is pushed a genuinely newer
|
||||
version; it stands aside once, and the next push it does not.
|
||||
@@ -1,78 +0,0 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-30
|
||||
located-in:
|
||||
- mesh-catalog modules/mailu (eight containers name Mailu's own resolver, and one of them binds a mesh name)
|
||||
- mesh-catalog modules/dnsmasq (dropped the DNSSEC bit its upstreams set)
|
||||
fixed-by: mesh-catalog PR 178 (mailu-admin uses the machine's resolver) and PR 179 (the machine's resolver passes the DNSSEC bit down) — 2026-09-30, the same afternoon
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 171 — A module that names its own resolver knows no mesh name
|
||||
|
||||
## What was observed
|
||||
|
||||
The afternoon [ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
|
||||
landed — no container is given the mesh's names any more; it asks the machine's resolver — the mail
|
||||
system's admin container began logging, 523 times in three minutes:
|
||||
|
||||
```
|
||||
psycopg2.OperationalError: could not translate host name "novox.internal" to address: Name does not resolve
|
||||
```
|
||||
|
||||
Mail was accepted on every port and the web front answered; the admin and the spam filter beside it
|
||||
were unhealthy, and anything that needed the database — a mailbox change through the API, the spam
|
||||
filter's domain list — failed. Found by the operator asking whether mail was back, forty minutes in.
|
||||
|
||||
Mailu ships its own resolver, an unbound in a container, and every other Mailu container is told to
|
||||
use it — the module carries `dns: [192.168.203.254]` on eight containers. That resolver recurses from the
|
||||
root and knows nothing under `.internal`. Until that afternoon the admin container had the database's
|
||||
name anyway, because the mesh wrote every name into every container at creation; the copy was the only
|
||||
reason a container behind its own resolver could reach anything by a mesh name, and nobody knew it was
|
||||
load-bearing.
|
||||
|
||||
**Removing the override was not enough.** Given the machine's resolver instead, the admin refused to
|
||||
start: `Your DNS resolver at 127.0.0.11 isn't doing DNSSEC validation`. Mailu checks, at start, that
|
||||
its resolver returns the Authenticated Data bit for a signed name. The mesh's resolver forwards to two
|
||||
upstreams that validate and set the bit, and dropped it on the way down — dnsmasq does unless told
|
||||
otherwise. Mailu's own unbound has no hook to forward a zone elsewhere, so it could not be taught the
|
||||
mesh's names either.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
**A container with a resolver of its own has opted out of the machine's, and nothing says so.** 0148
|
||||
made the machine's resolver load-bearing for every container; a `dns` on a container is a quiet
|
||||
exception to that, and the exception used to be papered over by the copy the record removed. The
|
||||
manifest field reads like a preference and is a decision about whether mesh names exist inside the
|
||||
container.
|
||||
|
||||
**A resolver that forwards to validating upstreams and hides the fact is less useful than it could
|
||||
be**, and the first program to check found out.
|
||||
|
||||
**The mesh reported nothing.** Every container ran; the failing one accepted connections; the report
|
||||
was about bytes. It is [issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
|
||||
again, and the check that would have caught it is the same unbuilt one.
|
||||
|
||||
## What was done
|
||||
|
||||
- The one Mailu container that binds a mesh name — the admin, through the database it is granted —
|
||||
no longer names Mailu's resolver and uses the machine's, like every container without a `dns` of
|
||||
its own (mesh-catalog PR 178). The other seven keep unbound: the spam filter needs a validating
|
||||
resolver for its blocklist lookups, and none of them asks for a mesh name.
|
||||
- The machine's resolver passes the DNSSEC bit down from its upstreams, `proxy-dnssec` (PR 179). It
|
||||
does not validate itself; the trust is the upstream's and the path to it, as a forwarding resolver's
|
||||
always was, and the configuration says so.
|
||||
|
||||
## What checks it
|
||||
|
||||
The admin container's own start-up check, which is what failed, and the mesh's status once it reads
|
||||
healthy. A container-level check that a mesh name resolves from inside every declared container is
|
||||
the one 110 and 145 both ask for and is not built.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should a container's `dns` be refused, or made to say what it gives up? A module that names its
|
||||
own resolver and binds a mesh name is a contradiction the controller can see at composition — the
|
||||
grant hands it a name its resolver will not answer.
|
||||
- Should the machine's resolver validate rather than proxy? It would cost a trust anchor on every
|
||||
machine and make the resolver slower to start; proxying was enough for the one program that asked.
|
||||
@@ -1,55 +0,0 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-09-30
|
||||
located-in:
|
||||
- the predecessor's terminal module (still generating the operator's ssh client blocks on every workstation)
|
||||
- mesh-controller internal/catalogue (the ssh-client roster, tested and not yet a catalogue module)
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 172 — The ssh client block for a machine matches one spelling of its name, and the other gets the wrong user
|
||||
|
||||
## What was observed
|
||||
|
||||
On a workstation, 2026-09-30, reported by the operator. `ssh home-server` logs in; `ssh
|
||||
home-server.internal` is refused with `Permission denied (publickey)`. The operator expected the
|
||||
opposite, if either: the full mesh name is the one the resolver serves.
|
||||
|
||||
The name is not the fault. Both spellings resolve to the machine's private address — the mesh's roster
|
||||
region in the hosts file carries `<node>.internal <node>` on one line, and the resolver answers
|
||||
anything under the node's name. What differs is the login: the generated client configuration has a
|
||||
`Host home-server` block naming the account to log in as, and `home-server.internal` matches no block,
|
||||
so ssh falls back to the operator's local username, which has no account on that machine. Spelled
|
||||
`account@home-server.internal` it works.
|
||||
|
||||
The file is the predecessor's. `~/.ssh/config.d/mesh` says in its own header that it is generated by
|
||||
the predecessor's terminal module, which only ever wrote the bare name. The mesh's own ssh-client
|
||||
roster — every other machine's Host block, written as a marked region of the operator's `~/.ssh/config`
|
||||
with the account the mesh knows for that machine (to-be 29) — already matches both spellings, and a
|
||||
controller test holds `Host marge marge.internal`. It is composed and tested in the controller and is
|
||||
not a module in the catalogue, so no machine receives it; every workstation still runs the
|
||||
predecessor's generator.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
**A name the mesh serves and a name a person can use are not the same set**, and the difference is
|
||||
silent. The resolver, the hosts file and the certificate authority all treat `<node>.internal` as the
|
||||
machine's name; the one file that decides who you log in as does not know it. A person who learns the
|
||||
mesh's name from `status` or from a certificate and types it is refused with an error that says
|
||||
nothing about a missing Host block.
|
||||
|
||||
**It is the migration story for the operator's own tooling, arriving as a symptom.** The mesh has the
|
||||
right file and does not ship it. Until the ssh-client roster is a module and is assigned to the
|
||||
workstations, the predecessor's generator keeps writing a file the mesh has already superseded, and
|
||||
every such file is one the mesh cannot correct.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should the ssh-client roster become a catalogue module now, assigned to every workstation, and take
|
||||
the predecessor's `config.d/mesh` out of the operator's `Include`? Its content is settled; what is
|
||||
not is the takeover of a file in a person's home that another generator still writes.
|
||||
- Should the block match a third spelling — the machine's public name, where it has one — or is that
|
||||
a different key and a different account?
|
||||
- What checks it? A controller test holds the two spellings; nothing checks that the file a workstation
|
||||
actually has is the mesh's rather than the predecessor's.
|
||||
Reference in New Issue
Block a user