Merge remote-tracking branch 'origin/main' into decision/docker-module

This commit is contained in:
2026-10-02 00:02:57 +02:00
20 changed files with 512 additions and 25 deletions
@@ -1,7 +1,7 @@
---
status: open
status: located
opened: 2026-09-22
located-in: []
located-in: [mesh-controller cmd/mesh-controller/adoption.go (take previews nothing), mesh-host internal/apply (the comparison and the record)]
fixed-by:
amended-design:
---
@@ -35,3 +35,8 @@ it changes before it changes it, and for taking a module this one does not.
included, and ask for the same kind of confirmation as the flip?
- Or should taking refuse while a port of the module is reachable more widely than the module
declares, until the operator either changes the module's exposure or confirms the narrowing?
## Decided, 2026-10-01
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rules 1 and 2: the preview names a narrowing. Building follows,
host first, then the controller's `take`.
@@ -1,7 +1,7 @@
---
status: open
status: located
opened: 2026-09-22
located-in: []
located-in: [mesh-controller cmd/mesh-controller/adoption.go (take previews nothing), mesh-host internal/apply (the comparison and the record)]
fixed-by:
amended-design:
---
@@ -48,3 +48,8 @@ network, or it is not a takeover.
directory — so the module adopts it by the rule that already exists?
- Should something refuse to call a module the successor of a bootstrap service it cannot adopt?
- Is the forge's own address better resolved than set, which is [issue 088](../088-the-forges-own-address-names-a-port-it-may-not-have/00-report.md)?
## Decided, 2026-10-01
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rule 7: genesis raises as the module declares. Building follows,
host first, then the controller's `take`.
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-09-22
located-in: [mesh-catalog, mesh-controller internal/catalogue]
fixed-by:
fixed-by: ADR 0104 — the route adapter module (mesh-catalog modules/route-adapter) writes each migrated route into the predecessor's proxy; it runs on the home server's migration
amended-design:
---
@@ -71,3 +71,10 @@ answered by an **adapter** that writes into the predecessor's own configuration.
the predecessor's proxy keeps serving every name and keeps its certificates, while each migrated
module's name is pointed at the mesh's container. The proxy is the last cutover again, and by then
every route is one the mesh contributed.
## Resolved, 2026-10-01
The adapter ADR 0104 decided exists and runs: `route-adapter` provides `route` on an adopted
machine by writing each migrated module's route where the predecessor's proxy reads it, and the
proxy itself is the last cutover. [ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md)
records the rest of what a take compares.
@@ -1,7 +1,7 @@
---
status: open
status: located
opened: 2026-09-23
located-in: []
located-in: [mesh-controller cmd/mesh-controller/adoption.go (take previews nothing), mesh-host internal/apply (the comparison and the record)]
fixed-by:
amended-design:
---
@@ -58,3 +58,8 @@ knowing the code.
an operator to undo it without reading the source?
- Is there anything a node must never be pushed without, such that sending a partial declaration is
worse than sending none?
## Decided, 2026-10-01
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rule 6: judged where stored; an impossible statement costs a module. Building follows,
host first, then the controller's `take`.
@@ -1,7 +1,7 @@
---
status: open
status: located
opened: 2026-09-23
located-in: []
located-in: [mesh-controller cmd/mesh-controller/adoption.go (take previews nothing), mesh-host internal/apply (the comparison and the record)]
fixed-by:
amended-design:
---
@@ -75,3 +75,8 @@ found, and so would be kept for ever on purpose.
module unassigned between the two declarations?
- What reports this? Nothing on the machine currently answers "what is running here that the mesh
did not ask for", which is the question that would have found this in seconds.
## Decided, 2026-10-01
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rule 5: former targets are removed and strays reported. Building follows,
host first, then the controller's `take`.
@@ -1,7 +1,7 @@
---
status: open
status: located
opened: 2026-09-23
located-in: []
located-in: [mesh-controller cmd/mesh-controller/adoption.go (take previews nothing), mesh-host internal/apply (the comparison and the record)]
fixed-by:
amended-design:
---
@@ -63,3 +63,8 @@ written.
substitutes settings into content today.
- Is the kept original enough of an answer, given nothing restores it and nothing points at it
when the service starts behaving differently?
## Decided, 2026-10-01
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rules 1 and 2: the difference is shown and a differing file refuses. Building follows,
host first, then the controller's `take`.
@@ -1,7 +1,7 @@
---
status: open
status: located
opened: 2026-09-23
located-in: []
located-in: [mesh-controller cmd/mesh-controller/adoption.go (take previews nothing), mesh-host internal/apply (the comparison and the record)]
fixed-by:
amended-design:
---
@@ -60,3 +60,8 @@ expected rate.
nothing answers the first.
- Is a digest pin the right thing for a module that takes over an existing service at all, or
should a cutover be able to say *keep what is running* and record what that was?
## Decided, 2026-10-01
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rules 1 and 2: the images are compared by age and a downgrade refuses. Building follows,
host first, then the controller's `take`.
@@ -1,7 +1,7 @@
---
status: open
status: located
opened: 2026-09-23
located-in: []
located-in: [mesh-controller cmd/mesh-controller/adoption.go (take previews nothing), mesh-host internal/apply (the comparison and the record)]
fixed-by:
amended-design:
---
@@ -65,3 +65,8 @@ the module can only be installed fresh.
Should it, so the dangerous case can be refused rather than discovered?
- What is the reverse path: the mesh has minted one, the service ignored it, and the working value
is still on the machine. Nothing reconciles those.
## Decided, 2026-10-01
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rules 2 and 3: a minted secret for found data refuses; secret accept reaches required secrets. Building follows,
host first, then the controller's `take`.
@@ -1,7 +1,7 @@
---
status: open
status: located
opened: 2026-09-23
located-in: []
located-in: [mesh-controller cmd/mesh-controller/adoption.go (take previews nothing), mesh-host internal/apply (the comparison and the record)]
fixed-by:
amended-design:
---
@@ -61,3 +61,8 @@ exercise.
learn to take a group atomically? Nothing takes more than one module at a time today.
- Does the same hole exist for anything else the predecessor's runtime resolves and the mesh's does
not — a network alias, a `depends_on`, a name in a shared `/etc/hosts`?
## Decided, 2026-10-01
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rules 1 and 4: the neighbours are named; a found network may be kept by a setting. Building follows,
host first, then the controller's `take`.
@@ -1,5 +1,5 @@
---
status: open
status: located
opened: 2026-09-26
located-in: [mesh-host internal/apply]
---
@@ -45,3 +45,8 @@ Instant renames both ways broke the circular dependency (forge needed for builds
builds needed for the push, push needed for the forge): data back to the old path,
old-spec forge started, artifacts rebuilt, data renamed forward, push. Nothing lost;
the install-page junk was discarded twice.
## Decided, 2026-10-01
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rules 5 and 7: every field compared; build says the policy. Building follows,
host first, then the controller's `take`.
@@ -1,8 +1,8 @@
---
status: open
status: resolved
opened: 2026-10-01
located-in: [mesh-controller internal/link/serve.go (act handles one message at a time; sourceMoved waits for every build the merge asks)]
fixed-by:
located-in: [mesh-controller cmd/mesh-controller/upgrades.go (the merge handler builds inside the receive loop)]
fixed-by: mesh-controller PR 197 (ADR 0162: the merge handler writes a plan, asks the first tier and returns; outcomes and a ticker advance it) and PR 199
amended-design: []
---
@@ -50,3 +50,20 @@ busy should say so where `status` is read.
and a build outcome for another module arrive together, and the outcome is recorded before the
build finishes; live, the controller's log during the next catalogue merge shows registrations
interleaved with the merge's own.
## Decided, 2026-10-01
[ADR 0162](../../02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md): a merge
produces a tiered plan the store keeps; the handler asks the first tier and returns; outcomes advance
the plan; a controller replaced mid-plan resumes it. The loop is never held by a build again.
## Resolved, 2026-10-01 evening
Since the controller holding ADR 0162's plan rolled, a merge announcement is handled in
milliseconds: the plan is written, the first tier asked, the loop free. The builds the merge
implies are asked from the store's record, tier by tier, so a controller replaced mid-plan resumes
it rather than losing it. The first live merge under it (a controller change) is the proof the
decision's table asks for; its tiers are read with `plans`.
*How it is checked:* the plan tests in mesh-controller; live, `plans` after a merge and the loop's
log taking reports in while the plan builds.
@@ -65,3 +65,10 @@ this report asks for.
*How this would be checked:* a builder restarted between an ask and its build still builds it; a
merge of a dependent repository before its prerequisite is held and named; `builds` lists asked,
running and built.
## Decided, 2026-10-01
The third fault is answered by [ADR 0162](../../02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md):
a release across repositories is a plan whose dependency edges cross repositories, sorted into
tiers and deployed tier by tier, read in `status`. The order a person kept is the order the tiers
give.
@@ -0,0 +1,63 @@
---
status: resolved
opened: 2026-10-01
located-in: [mesh-controller cmd/mesh-controller/plan.go (theRestOfTheMesh resolves every other machine without its pins and skips one that refuses, saying nothing), mesh-controller cmd/mesh-controller/network.go (onTheNetwork, the same)]
fixed-by: mesh-controller PR 199 (each machine resolved with its own pins; a dropped machine said in both passes), rolled 2026-10-01 evening; PR 200 keeps the per-machine view quiet
amended-design: []
---
# 188 — A refusal inside "who is on the network" drops a machine silently, and every symptom points elsewhere
## What was observed
At 15:28Z on 2026-10-01 the controller rolled to a build that refuses a machine with two modules
answering one provision and no pin naming which (mesh-controller 195). The control node had two
issuers of `acme-ca`. From that moment every plan of the control node failed with *step-ca has a
content that says `${machine:at}`, and this machine says mesh-range or name*; `seats` listed every
seat as unheld; the build machine refused the builder's and the proxy's builds with *no clone base
for that seat — nothing holds it*; and the roll-out of the next controller was refused with the
`${machine:at}` words. Not one of those names the cause. It was found by running the previous image
as a one-shot beside the current one and reading the difference, forty minutes later.
## Why this is here
`onTheNetwork` decides which machines have an address by resolving each one, unchecked, and
*skipping* any whose resolution errs. A machine skipped there has no `at`, so its own plan fails on
the first placeholder that needs one, in another module's words; everything held on it reads as
unheld; everything built from it cannot be built. The design lets one refusal become four unrelated
symptoms and no sentence about the refusal itself. It is the same shape as
[issue 187](../187-the-mesh-tells-nobody-when-it-stops-working/00-report.md): a fault that is swallowed
where it happens and discovered where it hurts.
## What a fix needs
- A machine whose resolution refuses is said, by `onTheNetwork`'s caller or in `status`: *the
control node does not resolve: more than one module provides acme-ca; pin one* — the resolver's
own words, which exist and were dropped.
- A refusal that a release introduces for a machine already converged — a new rule the stored
state does not meet — must not be silent at the roll either; the controller's prepare or first
resolution after a roll should name every machine it now refuses.
*How this would be checked:* a controller test where one machine's unchecked resolution refuses:
`status` names the machine and the refusal, and the other machines keep their addresses.
## Resolved in the live mesh, 2026-10-01
Two faults, one on top of the other. The refusal was mesh-controller 195's new rule — two modules
answering one provision on one machine need a pin — which its author hotfixed for the first pass
(196). The second pass of "the rest of the mesh" resolves every machine *without its pins*, so the
control node, pinned or not, was refused there and vanished: every seat it holds read as unheld,
the builder's and the proxy's builds were refused for want of the git seat's clone base, the
roll-out of the next controller was refused, and the first tiered plan failed at its first tier.
Found by a diagnostic build counting what each machine yielded. Fixed by mesh-controller PR
`fix/a-machine-not-on-the-network-is-said`: each machine is resolved with its own pins, and a machine
left out is named with the resolver's words in both places. The pin itself (`step-ca`, the issuer
the proxy already had) was made by hand and stands.
## Resolved, 2026-10-01 evening
The controller holding the fix was rolled onto the control node by the operator and a colleague
(the running one could not roll itself); after it every seat read as held again, a build asked
through the git seat worked, all four machines resolved and pushed. The line naming a dropped
machine spoke once too often — in the per-machine view, where the others are resolved without the
planned machine's offers and may fail by design — and is quiet there since mesh-controller PR 200.
@@ -0,0 +1,43 @@
---
status: located
opened: 2026-10-01
located-in: [mesh-controller internal/inventory/catalogue.go (RegisterModule records the source commit; a moved event follows a commit that changed, not an artifact that did), mesh-controller cmd/mesh-controller/upgrades.go (the roll-out follows the moved event)]
fixed-by:
amended-design: []
---
# 189 — A rebuild from the same commit is not a move, so a packaging module's new image never rolls out
## What was observed
The build machine's definition packages the controller's source. A controller merge rebuilds it, and
the rebuilt image carries the new controller; its own source commit in the catalogue is unchanged.
Registering that build therefore moves nothing the catalogue announces: no *moved* event, no
roll-out, although the module's upgrade policy says roll out and the artifact is new. On 2026-10-01
at 19:45Z the first tiered plan under [ADR 0162](../../02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md)
built the build machine and waited for the machine running it to apply the new build — a wait
that nothing would end, because nothing had sent it. A push by hand opened the gate.
## Why this is here
A version is what a module *runs*, and that is the artifact. The source commit is how the mesh
knows the artifact's provenance, not what makes it new: a build that reads another repository, or
pulls a base image, produces a different artifact from the same commit. The roll-out followed the
commit, so every module that packages another's source, and every dependent rebuilt because its
base moved, is rebuilt and then left behind on every machine until somebody pushes. The plan now
sends what it waits for (mesh-controller PR `fix/a-plan-sends-what-it-waits-for`), which covers the
gate; the general rule — a new artifact for a module with a roll-out policy is sent, moved commit or
not — is the decision this report asks for, and the catalogue's *moved* event should say what moved:
the artifact.
*How this would be checked:* a controller test registering a build of an unchanged commit with a
new artifact digest for a module whose policy rolls out: the machines are sent; `status` shows no
machine behind afterwards.
## Built in part, 2026-10-01
mesh-controller 203 and 204: a plan sends the machines of every module in a built tier whose policy
rolls out, once, moved commit or not, and waits for the ones a later tier is built by. After a merge
nothing is left for a hand to push, except what a *record* policy leaves by design. What remains for
a decision is the catalogue's own word: *moved* should follow the artifact, so a module rebuilt
outside any plan — by `build` by hand — rolls out the same way.