Both repository checks fail on main. 0115 is cited by no design, and it is
marked accepted while resting on 0112, which is proposed.
Design 27 already states the rule the record decides — "a module is assigned
at most once to a node, and that pair is the assignment's identity" — so it
is the home, and now says so. And 0112, 0113, 0114 and design 27 are all
proposed: the batch is under review, so the record is too. Promoting 0112
instead would be marking a record accepted to satisfy a check, which
check_rests_on names as a failure this repository already made once.
PR #133 landed a different 0115 while this branch was open. The bus record
is now 0116, with every citation in designs 19, 25, 28 and the index
following it.
Note: cycle.py and records.py both fail on main as merged, on that record —
nothing cites it, and it rests on 0112, which is still proposed. Both
pre-date this branch and are left for their own change.
A record can assert a fact that goes stale while the decision it supports
stays right. Superseding for that buries a sound record under a second one
and makes every reader work out which is live. So a correction of fact is
now made in place, marked and dated, with the old wording quoted — bounded
by three conditions and checked by records.py, which fires on an unmarked,
undated or back-dated note. Judgements still supersede.
Applied to 0115: no conformance suite exists to recapture, and the full
genesis bed cannot run until the links exist. Designs 25 and 28 follow.
Counting the surface first changed the plan twice: the genesis bed cannot run
until the links exist, so it belongs to step 4, and there is no conformance
suite to recapture — step 3 builds one against the current bus before moving
it. Both corrections are recorded in the breakdown rather than edited into
ADR 0115. Also indexes design 25, which was never listed.
The NATS change was recorded as one undivided item, which hid three gaps:
a mesh already running had no adoption path, the protocol specification did
not know its transport was being replaced, and nothing was runnable until
everything was. Dividing it is what surfaced them.
125: a hold is not a line in the apply report — sixteen resources held
for an untaken module while four surfaces reported success, and the
operator stopped the edge's predecessor on their word (the route-proxy
flip outage). 126: a changed volume path neither recreates a running
container nor warns, and a roll-out upgrade policy makes a build a
deployment — together they turned a data-path migration into a forge
outage (the /var/lib move). Filed as 119/121 in the novox session
before syncing; renumbered past the other session's 119-124.
Extends ADR 0075. Surfaced fixing builder's hand-faked package-registry
binding tonight: gitea's manifest declares the provision once with a single
npm-path, conflating what should be independently assignable per ecosystem
(npm/cargo/docker/...) the same way artifact-store and package-registry
were themselves split. Cited in 22-the-work-ahead.md's Phase 2, where the
target state this decision points at was already described a week ago.
Numbered 109, not 108: route-proxy's policy feature (mesh-controller PR
still-unwritten decision record — reserved but never committed. Renumbered
around it rather than colliding.
Fixing builder's hand-faked package-registry binding tonight (requires:
package-registry, a real mesh grant instead of a hardcoded JSON fragment)
broke three tests describing a deliberate carried-binding fallback for
exactly this: gitea's own image is built by builder, so builder cannot
yet hold a real grant from gitea the first time either has to exist.
Invisible on novox (already bootstrapped, gitea already live) — real on
any genesis from scratch. Fix left in place, tests left failing rather
than reverted or hacked, so the gap stays visible.
Filed after a session where every mesh-controller interaction went through
docker exec — its manifest runs it as a container with network: host, using
none of the isolation that resource type usually buys, while ADR 0006 makes
it the mesh's single point of coordination. Open question, not a claimed
defect: does type: container get the controller anything type: process
(supervised the way the host supervises its own unit, per ADR 0005) would not.
Comparing route-proxy against what HAL's actual Traefik config does today,
not Traefik's general feature set, per the standing rule that the nox mesh
must do at minimum what the HAL mesh it replaces already does. Everything
else checked out even or better; these two are real, confirmed gaps —
RedisInsight has no login of its own and depends entirely on Traefik's
basicauth middleware, and the gitea-internal route depends on an IP-scoped
deny rule. Neither has any equivalent in route-proxy's single-lookup
request path.
Records the rule the operator gave directly, mid-session, after checking
that HAL's own postgres and lavinmq both used a directory bind and the
mesh's adoption of them three weeks ago switched to a named volume without
a reason recorded anywhere.
Already built and rolled out on novox (mesh-catalog PR #54) before this
record -- urgent enough to fix first and write down after. Includes the
incident: the new host directories needed the container's own UID, which a
named volume gets for free and a directory bind does not; mesh-store
crash-looped on Permission denied until ownership was matched to what the
original volume already had.
Closes issue 115. Checks pass.
The decision (the mesh assigns a container's machine-side port; a module
says only what it needs) was proposed 2026-09-01, and the machinery
already implements it in full -- internal/inventory/ports.go's PortFor,
declaration.go's publishedOn. What was missing was the catalogue actually
complying: 14 of 46 modules baked a machine-side number into their own
manifest anyway. mesh-catalog PR fixes 11 of them (the two defensible
kinds -- foundation, protocol-fixed -- are left alone, per the issue's own
categories). Accepting the decision now that it's actually enforced, and
closing the issue it was blocking.
Checks pass.
Playbook 03 step 2: move status to diagnosing, then located once the
owner is known. located-in is filled with four confirmed packages —
the owner is known.
The carried-peer record (CarriedPeer/TunnelPeer) lives in mesh-controller
internal/inventory, not internal/catalogue. internal/catalogue is the
right package for the zone-generation side of the fix (facts.go's
nodeZones), but a different concern from where the name field itself
would go. Split the two so a decision record doesn't get pointed at the
wrong package.
Fixes, each named where it was wrong:
- reload-on is a service field; a container only has restart-on, which
recreates. Cited precedent (registry-trust-reload) is a service resource,
not a container. Fix: nats-server's own SIGHUP reload, triggered by an
in-image entrypoint watching a directory-mounted config file (issue 103's
recreate-on-change applies to a directly-mounted file, not a directory's
contents) -- asks nothing new of the host.
- A JetStream-delivered message's Reply field is already claimed by the
consumer's own ack address, so a responder using it answers nobody. Fix:
every CONTROL message needing a reply carries its reply subject in its own
payload; the controller publishes there explicitly, never via Respond().
Enrolment is the case this design actually depends on, so it's fixed there
too, not just noted.
- The listed permissions never granted publish on a durable consumer's own
ack-reply subject -- a module could receive but never ack, so every
message redelivers forever. Fixed with a scoped grant per module's own
consumer.
- One account (a deliberate choice, kept) means inbox privacy is the
permission list or nothing. The design granted 'its reply inbox' without
scoping it, which read as any user reaching any inbox. Fixed: each user's
inbox prefix is derived from its own identity and its permissions name
only that prefix.
New open question from this revision, not closed: whether the in-image
watch-and-SIGHUP shape belongs in mesh-sdk if a second module ever needs it.
Checks pass (records.py, cycle.py, index.py).
Checked why ADR 0104's forward-to-predecessor shape doesn't transfer to the
resolver the way it did the proxy: DNS is one process on one port, and
assigning the mesh's dnsmasq module replaces it in place, so there is no
predecessor process left standing to forward to.
But /etc/dnsmasq.d/hal-dns.conf's static address= lines for ace/shanks/g14
match the mesh's own carried-peer addresses from overlay show exactly. The
name a carried peer needs isn't a guess the operator has to make under
pressure — it's a transcription of a record the predecessor already has and
has been correctly serving for six days. Located in mesh-controller (no way
to attach a name to a carried peer today) and the dnsmasq module (doesn't
emit a wildcard for a named-but-uncarried peer). Not implemented.
hq cannot hold the migration's operational record — it names machines, addresses and
paths, and this repository is public — but it can say where that record is, which is what
this map is for. Asked for by the operator, who had to be told.
111, resolved: the map the control plane hands a resolution holds the machines and the
names the mesh merely serves, and the resolver's zones were given both — inventing names
under a suffix it answers authoritatively for. 112, open: adopting a tunnel gives the mesh
the peers' addresses and none of their names, so taking the resolver before they enrol
stops three machines resolving at all.
Both found by reading the plan before pushing it.
Found reviewing the resolver's conversion. Nothing fails while the node is adopted; it
fails at the flip, and it is the same split that decided which container survived the
hub's address change.
Found when adopting the tunnel moved the hub's address: the declaration followed, the
running container did not, and the forge lost its database. Issue 102's rule broken one
level down, and issue 103's fix stopping one input short.
Implemented in mesh-controller #49 and mesh-host #24. One proposal was rejected on the
record's own terms: converging the hub is not made to wait on other machines' migrations.
The addresses follow: the control plane holds the ports the node gave, a recorded build
holds no address at all, and the two forwarders that had been holding the control plane
together are removed. 097's stranded container turned out to be listening on every
interface and connected to the mesh's store — by its own old database, which is the only
reason nothing was at risk.
Both from the review of the issue-104 fix: a hand-applied file on an enrolled node is
recorded as carried and would remove the foundation, and nothing on the wire orders one
declaration against another.
Measured: AMQP is spoken in three places of the mesh's own code and in none of the
sdk or the modules; the predecessor's world is AMQP and retiring. NATS answers every
guarantee the bus relies on, durability via JetStream. Recommended: not underneath
the migration, not after it either — in parallel, one rehearsed rollout.
Two birth-address outages and a registry that would have been the third; a
container that keeps a stale environment after its file changes; a host command
that applied a converged declaration to an adopted node; the hub and the vault
without seats. And the decision the operator made under it all: the hub adopts
the predecessor's tunnel in place, key and peers and range and port.
Found checking the third cutover rather than running it. The first two were safe by
accident — both are reached through a host port, which survives a change of owner.
This is the first constraint found that decides the order of the migration.
Found at the second cutover. Carrying a value in works only for a module's own
secrets; a secret answered by the provision is minted, and six catalogue modules
take one that way.
Three modules in a row on one machine; the first was found by taking it and cost a
three-minute outage. The runbook's answer is a rule a person must remember, which is
the shape this repository says not to settle for.
The first pass answered under both ends, which review showed is the same fault seen
from the other side where two mappings share a number. Verified on the machine: the
forge is back on the port its own configuration has always advertised.
Found reading the second module's cutover rather than running it: the catalogue's
config drops a rule the machine's has, and no step puts the two side by side.
094's cause is one blind spot read from two ends, written up in its diagnosis; the fix
answers the first open question and not the other two, which become 096. 097 was found
looking at what the forge's cutover left running.