Reviews 0110/0121 after a session where renaming seats cost three freezes, a
builder deadlock, and hand-resolved manifests. The seat rules were right; the
set being a compiled Go slice referenced by name-string everywhere was the
mistake. Seats become a table keyed by a stable id; claims/held/production code
reference the id; a rename is one UPDATE, no rebuild, no re-registration, no
freeze. The build machine reads the set from the mesh instead of embedding it,
removing the controller/builder seat coupling. Closed set and scope naming
unchanged; only storage and reference change. Outstanding renames (registry
seats, private-network scope) wait for this — as data each is a write.
Records the manual update process (module moved -> build -> reconcile; and the
breaking-change freeze/re-register recovery), and the two things that make
self-update more than a webhook: the build-on-push trigger is currently HAL's
(hal-gitea-tools on :9877), a retirement gap the mesh must replace with its own
forge-webhook trigger wired to every repo including mesh-controller; and the
builder validates manifests too, so a breaking change couples controller +
builder + manifests + hosts, and renaming the builder's own seat deadlocks its
rebuild. Names the transition discipline (accept old+new for one release) that
self-update needs so a push does not auto-freeze.
Records the reversal: distribution stays as the mesh's OCI registry (it serves
every artifact-store:// image); only verdaccio, a redundant second npm registry,
is removed. The 'consolidate onto gitea / retire distribution' direction was
dropped. Also records that the node-* rename was executed as one controlled
migration with a brief compose freeze, and why the delivering registry seats
are deferred rather than folded in.
The control plane's seats grew a second naming style (the-*) beside mesh-*,
and the closed set was the only place any seat could be defined. This settles
both: system seats are mesh-* (one, mesh-wide) or node-* (one per node), named
for scope; a module may define its own seat outside the closed set. Folds in
the seat review: mesh-build-machine (scope fix), mesh-private-network (one
server + client modules, dropping per-node VPN choice), showcase becomes the
first module-defined seat, node-uplink, and the node-* renames — plus the
registry consolidation onto gitea, which reshapes the registry seats and gates
retiring distribution/verdaccio. Records why the renames are a coordinated
migration and why distribution cannot be removed until gitea serves images.
A roster fact may be shared — written into a marked region of the machine's
file (into: block, hq 128) rather than as the whole file. The template
renders the content; shared decides how the host lays it down. Composes with
hq 128: the region mechanism is the host's, the format is the module's.
The facts mechanism formatted the roster in Go in the control plane — one
formatter per fact, in the consumer's own configuration language. ADR 0120
makes a fact a path and a template: the mesh owns the data, the module owns
the format, and the control plane holds no format at all.
to-be 29 (operator accounts + what lives under a home) is rewritten to ride
it: the ssh files become roster templates, the whole ~/.ssh is owned with a
found/owned boundary that cannot lock the operator out, keys are mesh-owned
through an SSH CA (existing keys adopted not regenerated, the operator's
personal key signed not minted), and the ssh-agent is a user-scoped service.
The mesh models machines but not the people on them — a node record
holds no username, and no module places anything under a home. So who
you are on each node (jochens/ace/jochen) is unknown to the mesh, and
nothing owns ~/.ssh, dotfiles or ~/.config. HAL knew it; the nox mesh
dropped it. Proposes the account as a node fact and a home-scoped
resource class (the ~/ mirror of ADR 0112's /var/lib placement), with
the login key staying the operator's (ADR 0051). Not urgent — HAL's
generators still run — load-bearing at node-by-node retirement. Found
generating ~/.ssh/config from HAL's registry, which nox has no
equivalent for.
Both repository checks fail on main. 0115 is cited by no design, and it is
marked accepted while resting on 0112, which is proposed.
Design 27 already states the rule the record decides — "a module is assigned
at most once to a node, and that pair is the assignment's identity" — so it
is the home, and now says so. And 0112, 0113, 0114 and design 27 are all
proposed: the batch is under review, so the record is too. Promoting 0112
instead would be marking a record accepted to satisfy a check, which
check_rests_on names as a failure this repository already made once.
PR #133 landed a different 0115 while this branch was open. The bus record
is now 0116, with every citation in designs 19, 25, 28 and the index
following it.
Note: cycle.py and records.py both fail on main as merged, on that record —
nothing cites it, and it rests on 0112, which is still proposed. Both
pre-date this branch and are left for their own change.
A record can assert a fact that goes stale while the decision it supports
stays right. Superseding for that buries a sound record under a second one
and makes every reader work out which is live. So a correction of fact is
now made in place, marked and dated, with the old wording quoted — bounded
by three conditions and checked by records.py, which fires on an unmarked,
undated or back-dated note. Judgements still supersede.
Applied to 0115: no conformance suite exists to recapture, and the full
genesis bed cannot run until the links exist. Designs 25 and 28 follow.
Counting the surface first changed the plan twice: the genesis bed cannot run
until the links exist, so it belongs to step 4, and there is no conformance
suite to recapture — step 3 builds one against the current bus before moving
it. Both corrections are recorded in the breakdown rather than edited into
ADR 0115. Also indexes design 25, which was never listed.
The NATS change was recorded as one undivided item, which hid three gaps:
a mesh already running had no adoption path, the protocol specification did
not know its transport was being replaced, and nothing was runnable until
everything was. Dividing it is what surfaced them.
125: a hold is not a line in the apply report — sixteen resources held
for an untaken module while four surfaces reported success, and the
operator stopped the edge's predecessor on their word (the route-proxy
flip outage). 126: a changed volume path neither recreates a running
container nor warns, and a roll-out upgrade policy makes a build a
deployment — together they turned a data-path migration into a forge
outage (the /var/lib move). Filed as 119/121 in the novox session
before syncing; renumbered past the other session's 119-124.
Extends ADR 0075. Surfaced fixing builder's hand-faked package-registry
binding tonight: gitea's manifest declares the provision once with a single
npm-path, conflating what should be independently assignable per ecosystem
(npm/cargo/docker/...) the same way artifact-store and package-registry
were themselves split. Cited in 22-the-work-ahead.md's Phase 2, where the
target state this decision points at was already described a week ago.
Numbered 109, not 108: route-proxy's policy feature (mesh-controller PR
still-unwritten decision record — reserved but never committed. Renumbered
around it rather than colliding.
Fixing builder's hand-faked package-registry binding tonight (requires:
package-registry, a real mesh grant instead of a hardcoded JSON fragment)
broke three tests describing a deliberate carried-binding fallback for
exactly this: gitea's own image is built by builder, so builder cannot
yet hold a real grant from gitea the first time either has to exist.
Invisible on novox (already bootstrapped, gitea already live) — real on
any genesis from scratch. Fix left in place, tests left failing rather
than reverted or hacked, so the gap stays visible.
Filed after a session where every mesh-controller interaction went through
docker exec — its manifest runs it as a container with network: host, using
none of the isolation that resource type usually buys, while ADR 0006 makes
it the mesh's single point of coordination. Open question, not a claimed
defect: does type: container get the controller anything type: process
(supervised the way the host supervises its own unit, per ADR 0005) would not.
Comparing route-proxy against what HAL's actual Traefik config does today,
not Traefik's general feature set, per the standing rule that the nox mesh
must do at minimum what the HAL mesh it replaces already does. Everything
else checked out even or better; these two are real, confirmed gaps —
RedisInsight has no login of its own and depends entirely on Traefik's
basicauth middleware, and the gitea-internal route depends on an IP-scoped
deny rule. Neither has any equivalent in route-proxy's single-lookup
request path.
Records the rule the operator gave directly, mid-session, after checking
that HAL's own postgres and lavinmq both used a directory bind and the
mesh's adoption of them three weeks ago switched to a named volume without
a reason recorded anywhere.
Already built and rolled out on novox (mesh-catalog PR #54) before this
record -- urgent enough to fix first and write down after. Includes the
incident: the new host directories needed the container's own UID, which a
named volume gets for free and a directory bind does not; mesh-store
crash-looped on Permission denied until ownership was matched to what the
original volume already had.
Closes issue 115. Checks pass.
The decision (the mesh assigns a container's machine-side port; a module
says only what it needs) was proposed 2026-09-01, and the machinery
already implements it in full -- internal/inventory/ports.go's PortFor,
declaration.go's publishedOn. What was missing was the catalogue actually
complying: 14 of 46 modules baked a machine-side number into their own
manifest anyway. mesh-catalog PR fixes 11 of them (the two defensible
kinds -- foundation, protocol-fixed -- are left alone, per the issue's own
categories). Accepting the decision now that it's actually enforced, and
closing the issue it was blocking.
Checks pass.
Playbook 03 step 2: move status to diagnosing, then located once the
owner is known. located-in is filled with four confirmed packages —
the owner is known.
The carried-peer record (CarriedPeer/TunnelPeer) lives in mesh-controller
internal/inventory, not internal/catalogue. internal/catalogue is the
right package for the zone-generation side of the fix (facts.go's
nodeZones), but a different concern from where the name field itself
would go. Split the two so a decision record doesn't get pointed at the
wrong package.
Fixes, each named where it was wrong:
- reload-on is a service field; a container only has restart-on, which
recreates. Cited precedent (registry-trust-reload) is a service resource,
not a container. Fix: nats-server's own SIGHUP reload, triggered by an
in-image entrypoint watching a directory-mounted config file (issue 103's
recreate-on-change applies to a directly-mounted file, not a directory's
contents) -- asks nothing new of the host.
- A JetStream-delivered message's Reply field is already claimed by the
consumer's own ack address, so a responder using it answers nobody. Fix:
every CONTROL message needing a reply carries its reply subject in its own
payload; the controller publishes there explicitly, never via Respond().
Enrolment is the case this design actually depends on, so it's fixed there
too, not just noted.
- The listed permissions never granted publish on a durable consumer's own
ack-reply subject -- a module could receive but never ack, so every
message redelivers forever. Fixed with a scoped grant per module's own
consumer.
- One account (a deliberate choice, kept) means inbox privacy is the
permission list or nothing. The design granted 'its reply inbox' without
scoping it, which read as any user reaching any inbox. Fixed: each user's
inbox prefix is derived from its own identity and its permissions name
only that prefix.
New open question from this revision, not closed: whether the in-image
watch-and-SIGHUP shape belongs in mesh-sdk if a second module ever needs it.
Checks pass (records.py, cycle.py, index.py).
Checked why ADR 0104's forward-to-predecessor shape doesn't transfer to the
resolver the way it did the proxy: DNS is one process on one port, and
assigning the mesh's dnsmasq module replaces it in place, so there is no
predecessor process left standing to forward to.
But /etc/dnsmasq.d/hal-dns.conf's static address= lines for ace/shanks/g14
match the mesh's own carried-peer addresses from overlay show exactly. The
name a carried peer needs isn't a guess the operator has to make under
pressure — it's a transcription of a record the predecessor already has and
has been correctly serving for six days. Located in mesh-controller (no way
to attach a name to a carried peer today) and the dnsmasq module (doesn't
emit a wildcard for a named-but-uncarried peer). Not implemented.
hq cannot hold the migration's operational record — it names machines, addresses and
paths, and this repository is public — but it can say where that record is, which is what
this map is for. Asked for by the operator, who had to be told.
111, resolved: the map the control plane hands a resolution holds the machines and the
names the mesh merely serves, and the resolver's zones were given both — inventing names
under a suffix it answers authoritatively for. 112, open: adopting a tunnel gives the mesh
the peers' addresses and none of their names, so taking the resolver before they enrol
stops three machines resolving at all.
Both found by reading the plan before pushing it.
Found reviewing the resolver's conversion. Nothing fails while the node is adopted; it
fails at the flip, and it is the same split that decided which container survived the
hub's address change.
Found when adopting the tunnel moved the hub's address: the declaration followed, the
running container did not, and the forge lost its database. Issue 102's rule broken one
level down, and issue 103's fix stopping one input short.
Implemented in mesh-controller #49 and mesh-host #24. One proposal was rejected on the
record's own terms: converging the hub is not made to wait on other machines' migrations.
The addresses follow: the control plane holds the ports the node gave, a recorded build
holds no address at all, and the two forwarders that had been holding the control plane
together are removed. 097's stranded container turned out to be listening on every
interface and connected to the mesh's store — by its own old database, which is the only
reason nothing was at risk.
Both from the review of the issue-104 fix: a hand-applied file on an enrolled node is
recorded as carried and would remove the foundation, and nothing on the wire orders one
declaration against another.