Issues 224, 248-252, 254, 255 resolved, 256 opened and resolved, 253 where it stands

Each resolved with the pull request that fixed it and what was seen live;
253 waits for real collection, the operator's word to give.
This commit is contained in:
jochen
2026-10-05 18:36:19 +02:00
parent 91aa481269
commit 0fdcb15200
12 changed files with 92 additions and 17 deletions
@@ -1,8 +1,8 @@
---
status: open
status: resolved
opened: 2026-10-04
located-in: [mesh-host internal/apply/apply.go, mesh-host internal/apply/schedule.go]
fixed-by:
fixed-by: mesh-host PR #23
amended-design:
---
@@ -62,3 +62,9 @@ is a hole in something new rather than something that broke.
at the cost of a second source for "is this container meant to be running".
- Either way: should the *report* say a window is open, so a machine that looks half-stopped at
03:31 reads as working rather than broken?
## Resolved — 2026-10-05
An apply that arrives during a window now leaves the containers the window holds alone, and reports them `held-still`, naming the step. The first apply after the window converges them. The window is recorded under the host's state directory, the one place every applier on a machine shares: the daemon, a hand-run apply and the installer. It is released on every path, and lapses after six hours or when its process is gone. The machine's report lists open windows. Live on all four machines the same evening. Of the issue's two options, the narrower was taken: nothing blocks, so a push is never held for the length of a window.
Accepted and said in the change: an apply that inspected a container as running in the instant a window opens can still recreate it.
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-10-05
located-in: [mesh-controller internal/broker]
fixed-by:
fixed-by: mesh-controller PR #51
amended-design:
---
@@ -48,3 +48,7 @@ reached all four machines.
**Not explained by this issue:** the record's merge was still logged twice, with the consumer fresh and
nothing replayed. The duplication has a cause of its own, still to be found. It may be that the forge
announces a merge on two paths, or that one event is handled twice.
## Resolved — 2026-10-05
The controller's event consumer is made from now when it is made, and `broker consumer-reset` re-makes a stuck one from now. Live: after the restart, merges reached the controller within seconds. Two tests main then failed, both skipped without a store, were fixed in mesh-controller PR #52. The merge heard twice was a separate cause: [issue 250](../250-a-merge-made-through-the-forges-tool-is-announced-twice/00-report.md).
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-10-05
located-in: [mesh-controller cmd/mesh-controller]
fixed-by:
fixed-by: mesh-controller PR #54
amended-design: 03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md
---
@@ -18,3 +18,13 @@ before it takes effect). That decision leaves an add-only change alone on purpos
**Decided** in [ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md):
grants before code, and one machine first unless a module's policy says *together*.
## Resolved — 2026-10-05
[ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md), live the same evening:
- every push now issues memberships before it sends declarations;
- the bus's machine goes first when its list of users moved;
- a plan sent the build agent to one machine first and to the rest once that machine reported.
The first live rollout exposed a fault in the tier gate. The first machine's report, made between the two sends, was read as stale ([issue 256](../256-a-first-machines-report-read-as-stale-between-the-two-sends/00-report.md)).
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-10-05
located-in: [mesh-catalog modules/gitea]
fixed-by:
fixed-by: mesh-catalog PR #63, PR #66
amended-design:
---
@@ -43,3 +43,7 @@ it would get a second plan. The record module synced twice for each merge.
The poll is the only emitter: it sees every path and carries the clone address. The tool merges and
answers the merge commit. A merge made through the tool is heard up to thirty seconds later, which the
module already accepts ("an event a minute late is still an event").
## Resolved — 2026-10-05
The forge module's poll is the only announcer of a merge. It asks only the repositories that moved since its last look, and one pass at a time. Live: three merges were each heard once, and a later merge was planned within a minute.
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-10-05
located-in: [mesh-catalog modules/records]
fixed-by:
fixed-by: mesh-catalog PR #64
amended-design:
---
@@ -30,3 +30,7 @@ The module, ported to Go as part of the fix, clones into a directory of its own
a directory it makes, and so owns. What the old layout left behind is removed where it is the module's.
Where it is not, the module names it in its status, with the one command that deletes it. The
container-runtime capability is dropped. A sync that fails is still said in the status, as before.
## Resolved — 2026-10-05
The record module, ported to Go, clones into a directory it makes and owns. Live: it synced to the newest commit, with 692 documents. It names seven leftovers of the old checkout that it cannot remove, with the command that removes them; that is the operator's act.
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-10-05
located-in: [mesh-catalog modules/gitea, mesh-controller cmd/mesh-controller]
fixed-by:
fixed-by: mesh-catalog PR #63, mesh-controller PR #54
amended-design: 03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md
---
@@ -32,3 +32,7 @@ rebuild not hidden it, a module whose own files changed would have stayed unbuil
shared, when that module's definition is among the changed paths. A sibling without a definition, such
as a shared library, still means everything, which is the safe direction. Root files still mean
everything.
## Resolved — 2026-10-05
The forge's module reads every page of a pull request's files. The controller counts a new module's own directory as that module's when its definition is among the changed paths. Live: catalogue merges planned one module each.
@@ -2,7 +2,7 @@
status: located
opened: 2026-10-05
located-in: [mesh-controller internal/builder, mesh-controller internal/artifacts, mesh-catalog modules/distribution]
fixed-by:
fixed-by: mesh-controller PR #53, mesh-host PR #23, mesh-catalog PR #62
amended-design: 03-DESIGN/01-to-be/18-building-a-module.md
---
@@ -49,3 +49,7 @@ forever.
- letting an archive go removes its manifest first;
- a forgotten module keeps nothing.
- Real collection returns once the controller reports no kept archive unheld.
## Where it stands — 2026-10-05
Every kept archive is held: the controller's collection command reports 134 of 134 held, none missing. The window the collector needs is no longer reopened by an apply ([issue 224](../224-an-apply-reopens-a-maintenance-window-by-recreating-what-it-held-still/00-report.md)). The collector still runs as a dry run. Turning it to real collection deletes the layers nothing keeps, which is the operator's word to give; this issue resolves when that change lands.
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-10-05
located-in: [mesh-controller cmd/mesh-controller, mesh-controller internal/inventory]
fixed-by:
fixed-by: mesh-controller PR #54
amended-design: 03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md
---
@@ -30,3 +30,7 @@ build's output wins.
[ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md)
§3. A newer merge's plan supersedes the older open plans for the same repository and branch, and takes in
the modules they had not built. A person can close a plan by its id.
## Resolved — 2026-10-05
A newer merge's plan supersedes the older open plans of its repository and branch, and a person can close a plan by its id. Live: the plan left open since the afternoon was closed by hand, and its note said what it had built and never sent. A merge made while the controller's own plan was open took it over: the older plan reads "superseded at tier 0 by" the newer, which planned its three modules again.
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-10-05
located-in: [mesh-catalog modules/systemd]
fixed-by:
fixed-by: mesh-catalog PR #65
amended-design:
---
@@ -30,3 +30,7 @@ the sudo module's check confirms.
A read of the system journal escalates like an act does. The account's own journal, in the user scope,
does not. The module was ported to Go as part of the fix, its tests with it.
## Resolved — 2026-10-05
The journal verb reads a system unit's journal escalated, in the module ported to Go. Live: the controller's and the host's journals read through the verb on the control node.
@@ -0,0 +1,31 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-controller cmd/mesh-controller]
fixed-by: mesh-controller PR #55
amended-design:
---
# 256 — A first machine's report read as stale between the two sends
## What was observed
The first plan to roll out under [ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md)
sent the build agent to one machine first, and to the other two once that machine reported it applied, 27
seconds later. All three applied it within seconds. The plan then said, for seven minutes, that it was
"waiting for build-agent" on the first machine to be applied. It went on only when that machine happened to
report again for another reason.
## Diagnosis
The tier gate asked every machine for a report made after the module's last send. With one machine first,
a module is sent twice, and the second send is the later one. The first machine's report came between the
two sends, so it read as older than the build. The plan waited for that machine's next report, one report
cycle, and never knew why.
## Fix
The gate judges each machine from its own send: the first machine from the first send, the rest from the
second. A plan's wait is also printed to the second rather than the minute, where it had read "0s", and the
first send is printed in the machine's own time rather than in UTC. Checked by the controller's test: a first
machine's report between the two sends opens the gate, and a report from before its send does not.