Asked what the "sidecar" is and whether a supervised process would do instead. The repository
answers both ways. ADR 0047 (accepted, unsuperseded) says a module with tools or events runs a
container carrying its compiled code. To-be 18 and 20 (both proposed) define a `process` resource
type — the module's own code, a unit the machine's supervisor keeps up — and the worked guide says
plainly "it is why these are `process` rather than four containers." Neither design doc names 0047,
and no decision record mentions a `process` shape at all.
Diagnosed rather than left open, because the ground truth settles what the report could not.
The shape is real: mesh-host defines TypeProcess, applies it, and tests it, and the host's
vocabulary is twelve shapes rather than the nine ADR 0029 counted. So the alternative the report
offered — that two proposed documents describe a type that does not exist — is disproven.
ADR 0029's mechanism is intact and was not enough. The vocabulary-count test names the decision
behind each addition: network 0029, access 0051, opening 0100. The eleventh names a *proposed
design document*, and TypeProcess is the only shape in the vocabulary whose doc comment cites no
ADR. Requiring every addition to name something does not require it to name a decision.
The argument this issue asked for already exists — as a Go test comment. "It is a full-host shape
rather than a portable one: it needs a process supervisor to install into. It does NOT need a
container runtime, which is the point — only software that genuinely needs isolation asks for a
container." That is a decision's context and consequences, in another repository.
What the catalogue does is a third thing: 115 container declarations against 3 process, all three
in showcase — the module to-be 20 documents. There the tools resource is a container running
`sleep infinity` on a bare upstream base with the broker credential mounted, and the tools and
provisioner entrypoints are run by nothing. That is the condition 0047 was written to end, back
in a new shape.
Where the isolation argument leaks is narrower than expected and worth having precisely: the
serving key and the credential shape both conform. But serveTools serves every registered module
over one broker connection, the runtime takes its modules from a comma-separated list, and
x-source is stamped from the single credential — so two modules in one runtime means the second's
events are attributed to the first. Nothing refuses it and no test asserts against it.
Located on hq rather than on a code repository: the implementation and the design layer agree,
and the missing thing is the record. Which shape is right is left open, deliberately — this
establishes that the question was answered in practice and never written down, not which answer
is correct.
One correction kept in the trail: the first search here was for len(Vocabulary()), found nothing,
and was two steps from being written up as "the mechanism ADR 0029 relied on is gone." The test
binds the slice to a local first. A negative search result read as a fact about the world is the
same error issue 113 recorded.
The gap is closed in the proxy: policy applies, the four capabilities exist, the table is keyed
by host and path with a total ordering, and the two failure modes that rot quietly are held by
tests — a declaration carrying a credential refused rather than served, an unreadable secret
failing closed.
Resolved rather than left open because the issue reports a gap in the proxy and that gap is
gone. But the record says plainly what it does not yet allow: an operator still cannot move the
affected routes, because that needs the mesh side — a manifest able to declare these values and
the controller minting the secret auth names. Until both exist the capability is reachable only
by writing the routes file by hand. That is the ordinary build-out of a contract this issue's
decision created, and it belongs to to-be 08 rather than here.
The open questions are marked answered and kept rather than deleted, pointing at ADR 0108 —
what was rejected and why is the half worth having, and a section still saying "the fix should
not be written before these are answered" after the fix was written reads as though nobody
looked.
One finding kept in the record: priority was read with the reader for ports, which caps at
65535, and the one real rule this reproduces is declared at 100000. It parsed to zero, so
refusal and path scoping would have shipped looking complete and doing nothing on the only case
that motivated them. A validator borrowed from a neighbouring field is a silent default.
Issue 116 found the mesh's proxy applies nothing to a request — host lookup, forward. Against
what the replaced ingress actually relies on, four capabilities are missing: authentication
(three dependents, each gating an admin surface with no login of its own), refusal scoped to a
path (one, a live incident mitigation), path-scoped routing with priority, and redirect.
Policy goes on the route rather than beside it. A proxy-side settings layer keyed by route name
would keep the grant literally clean, but then "what protects this route" is answered from two
files nothing keeps in step — and a route's protection is part of what a route is.
The set is closed at those four, so a fifth is an amendment and each addition is earned by a
dependent that exists. An open middleware surface was rejected: it recreates what is being
replaced, and narrowing one later is far harder than widening a closed one.
Where policy needs a credential the declaration names a secret and never carries the value,
which keeps the existing secret machinery the only thing holding credentials. Inlining a hash
was rejected as the first credential in a declaration — a precedent easier to set than withdraw.
This re-keys the routing table by host and path with priority, which follows from the decision
rather than being a separate one: two of the four need one host routed more than one way. Equal
priorities must resolve identically every time or the proxy stops being reproducible.
The record says how it is checked, including the negative case that rots quietly — a
declaration carrying a credential value rather than a reference must be refused, so the
rejected option cannot return by accident.
08-connectivity §3 names the record and gains the subsection; issue 116 gains amended-design.
Two things the report got wrong, and one it could not have found the way it looked.
Disclosure first: it carried a real hostname and an absolute node path, in a public
repository. Both are gone; the ingress, the modules and the routes are named by role, as the
rest of 04-ISSUES does.
The count was low. Basic authentication has three dependents in the catalogue, not one — the
key-value store's browser UI, a database web UI, and the ingress's own dashboard. All three
are credential-less admin surfaces whose only gate is a middleware the mesh's proxy lacks.
The earlier version read only the node's dynamic configuration directory, which cannot see
what modules declare as container labels; counting needs both sources, and the report now
says so.
Two gaps were missing entirely. Redirect rules: two live routes canonicalise a www name onto
its apex, they exist only on the node and not in the catalogue, and they fail silently rather
than erroring. And path-scoped routing with priority, which is the one that reorders the
issue: the table maps host to exactly one target, so a host cannot be routed two ways, and
the refusal rule matches a path on a host already routed elsewhere. Authentication and a
source filter would not make it expressible. Path scoping is a prerequisite, not a sibling.
Also corrected: the refusal rule was described as an address-scoped deny. It is an allow-list
holding a single documentation-range address — deny-everyone — so reading it as address-scoped
points at the wrong fix. And its severity was understated: its own header records it as
incident response closing an abused write primitive, which is not "a real exposure" but a live
mitigation.
The open questions now say plainly that they are design questions and the fix should not be
written before they are answered, and one is added: whether a declaration may carry a
credential at all.
Comparing route-proxy against what HAL's actual Traefik config does today,
not Traefik's general feature set, per the standing rule that the nox mesh
must do at minimum what the HAL mesh it replaces already does. Everything
else checked out even or better; these two are real, confirmed gaps —
RedisInsight has no login of its own and depends entirely on Traefik's
basicauth middleware, and the gitea-internal route depends on an IP-scoped
deny rule. Neither has any equivalent in route-proxy's single-lookup
request path.
The previous version ranked candidates on whether they preserved single sign-on to the
object store's console. That is not a requirement: a "user" of the store is normally an
application, so the requirement is per-application keys scoped to buckets — which the mesh
already mints. And the console login it ranked on never worked; the module's own hook comment
records "policy claim missing", a failing login written up as progress.
It also omitted the incumbent's own maintained fork, which changes the question from "which
product replaces it" into two decisions: repoint, or migrate — and if migrating, to which.
Repointing costs an image reference; migrating costs a data copy, two handler rewrites and a
maintenance window. Repointing does not foreclose migrating, which is the argument for taking
it first.
On the corrected requirement Garage ranks first — its per-key-per-bucket model is the
requirement verbatim, its admin API matches how the mesh provisions, and the highest-risk
consumer is first-party documented against it. Its remaining gap (no versioning, no
server-side encryption, partial lifecycle) is unmeasured against the buckets and is the one
thing that could still disqualify it.
Measured and folded in: 230 GiB logical, 82,496 objects, 468 GiB raw at 2.03x, eight drive
directories on one filesystem on one machine. That last fact decides more than any feature —
the erasure coding is not buying independent-drive redundancy, so the redundancy model is
close to irrelevant and only storage overhead remains, which at this volume is a rounding
error against the headroom.
Both errors are recorded at the end of 01 rather than quietly fixed. A configured feature is
not an observed one; and when a dependency dies, "who took it over" precedes "what replaces
it" — searching for alternatives by construction returns things that are not the incumbent.
The diagnosis carried a table headed "claims that could not be substantiated", denying a
module.json, a digest pin, and an all-zeros runtime digest. All three exist. The table is
withdrawn in full and replaced with what is actually true, plus the two claims that remain
genuinely unverified rather than disproven.
The cause: one repository was searched and absence in it was written up as absence. The
catalogue of the mesh being built is a separate repository, not checked out where the search
ran, and all four claims were about that repository. Compounding it, the predecessor's
object-store module and the one being cut over to were treated as one thing — they are
different files in different repositories, one pinning a tag with no sidecar, the other a
digest with two container resources.
Also corrected in the report: located-in named the wrong repository; the "pins a tag" passage
described the predecessor; the open question about pinning by digest is struck, because this
module already does and it made no difference — a deleted digest resolves to nothing either
way. The section on why nothing broke is now scoped explicitly to the predecessor's
machinery.
The lesson kept in the record: "zero occurrences anywhere in the tree" is only as strong as
the tree searched, and a diagnosis must say which tree. A confident rebuttal of a correct
report is worse than no diagnosis — it sends the next person to the wrong place with a
written record behind them.
A thought raised mid-task needed remembering but not working on, and there was no
mechanism for that — so it was recorded by hand. This is that, made repeatable.
Records to Claude's persistent memory rather than the repository, deliberately. A parked
thought has no number, owner or status: giving it one asserts triage that deferring says
has not happened. A shared "deferred" document would be a central status file, which
AGENTS.md forbids. And a repository write means a branch, a commit and an MR — the drift
the skill exists to prevent.
Wraps no playbook, because deferring precedes the development cycle rather than being part
of it. It does say which playbook a thought would need if it graduates, and that recording
"undetermined" is the honest answer when the evidence does not say.
The stop condition is the substance: at most two lines, then return to what was in
progress. No plan, no triage question, nothing opened.
Records the rule the operator gave directly, mid-session, after checking
that HAL's own postgres and lavinmq both used a directory bind and the
mesh's adoption of them three weeks ago switched to a named volume without
a reason recorded anywhere.
Already built and rolled out on novox (mesh-catalog PR #54) before this
record -- urgent enough to fix first and write down after. Includes the
incident: the new host directories needed the container's own UID, which a
named volume gets for free and a directory bind does not; mesh-store
crash-looped on Permission denied until ownership was matched to what the
original volume already had.
Closes issue 115. Checks pass.
SeaweedFS was scoped as primary because it looked like the only candidate preserving
OIDC console login. Measured: its admin UI is Apache-2.0 but its identity-provider
integration is not — console SSO sits behind the per-TB commercial licence, alongside
point-in-time recovery and automatic EC repair. The free build gives OIDC on the S3 API
via STS and a console authenticated by local username and password.
So the answer to the gating question is that no candidate preserves the current feature
set for free, which this effort had written down as a possible outcome. Reopened across
three candidates with the requirement-by-requirement evidence in 01.
Two corrections to what the overview recorded. RustFS is not a binary-level drop-in
retaining existing data: API and on-disk compatibility are separate paths and the on-disk
one is preview-scoped. And it carries an open defect in the credential path the bucket
provision depends on, which gates it specifically.
Nothing graduates before two measurements named in 01: whether an authenticating proxy
is an acceptable answer to console SSO, and which S3 endpoints consumers actually call —
the latter because Garage does not implement the full span and cannot be ranked until
that is counted.
The decision (the mesh assigns a container's machine-side port; a module
says only what it needs) was proposed 2026-09-01, and the machinery
already implements it in full -- internal/inventory/ports.go's PortFor,
declaration.go's publishedOn. What was missing was the catalogue actually
complying: 14 of 46 modules baked a machine-side number into their own
manifest anyway. mesh-catalog PR fixes 11 of them (the two defensible
kinds -- foundation, protocol-fixed -- are left alone, per the issue's own
categories). Accepting the decision now that it's actually enforced, and
closing the issue it was blocking.
Checks pass.
Review of my own text found an unattributed claim — "the mesh's claim that a node
can be rebuilt from its declarations" — which is not a stated principle anywhere.
Replaced with the design position that genuinely covers it: to-be 07 chooses
references over payload because "reproducibility comes from pinning the identity of
a thing rather than carrying its bytes". This incident is that choice's failure mode
when the identity stops resolving, which is a sharper point than the one I made.
Scope stated honestly: the passage is about the foundation bundle and this module is
not in it, but pin-identity-fetch-bytes is how every module gets third-party images.
Also names the tension the mirroring question actually carries — mirroring is a move
away from references-over-payload, so it is a decision, not a fix.
The symptom arrived diagnosed as "the registry disabled anonymous pulls for the
whole vendor namespace". It did not hold: sibling repositories in that namespace
pull normally, the "$disabled" token field appears on every repository including
working ones and describes signing rather than access, and "actions": [] with a
401 is byte-identical to what an invented repository name returns. Both registries'
own APIs establish deletion instead.
Recorded because the correction is the expensive part to rediscover, and because
the instance was harmless while the standing condition is not: no node that does
not already hold the images can ever provision the module again, and nothing
detects that until one tries.
Research 015 scopes the replacement. It is not a redesign — the foundation design
already commits to S3 the protocol rather than the product, and the object store
is an ordinary module, so this instantiates an existing principle. The live OIDC
wiring is the requirement that gates the choice, and it is checked first.
Playbook 03 step 2: move status to diagnosing, then located once the
owner is known. located-in is filled with four confirmed packages —
the owner is known.
The carried-peer record (CarriedPeer/TunnelPeer) lives in mesh-controller
internal/inventory, not internal/catalogue. internal/catalogue is the
right package for the zone-generation side of the fix (facts.go's
nodeZones), but a different concern from where the name field itself
would go. Split the two so a decision record doesn't get pointed at the
wrong package.
Checked why ADR 0104's forward-to-predecessor shape doesn't transfer to the
resolver the way it did the proxy: DNS is one process on one port, and
assigning the mesh's dnsmasq module replaces it in place, so there is no
predecessor process left standing to forward to.
But /etc/dnsmasq.d/hal-dns.conf's static address= lines for ace/shanks/g14
match the mesh's own carried-peer addresses from overlay show exactly. The
name a carried peer needs isn't a guess the operator has to make under
pressure — it's a transcription of a record the predecessor already has and
has been correctly serving for six days. Located in mesh-controller (no way
to attach a name to a carried peer today) and the dnsmasq module (doesn't
emit a wildcard for a named-but-uncarried peer). Not implemented.
hq cannot hold the migration's operational record — it names machines, addresses and
paths, and this repository is public — but it can say where that record is, which is what
this map is for. Asked for by the operator, who had to be told.
111, resolved: the map the control plane hands a resolution holds the machines and the
names the mesh merely serves, and the resolver's zones were given both — inventing names
under a suffix it answers authoritatively for. 112, open: adopting a tunnel gives the mesh
the peers' addresses and none of their names, so taking the resolver before they enrol
stops three machines resolving at all.
Both found by reading the plan before pushing it.
Found reviewing the resolver's conversion. Nothing fails while the node is adopted; it
fails at the flip, and it is the same split that decided which container survived the
hub's address change.
Found when adopting the tunnel moved the hub's address: the declaration followed, the
running container did not, and the forge lost its database. Issue 102's rule broken one
level down, and issue 103's fix stopping one input short.
Implemented in mesh-controller #49 and mesh-host #24. One proposal was rejected on the
record's own terms: converging the hub is not made to wait on other machines' migrations.
The addresses follow: the control plane holds the ports the node gave, a recorded build
holds no address at all, and the two forwarders that had been holding the control plane
together are removed. 097's stranded container turned out to be listening on every
interface and connected to the mesh's store — by its own old database, which is the only
reason nothing was at risk.
Both from the review of the issue-104 fix: a hand-applied file on an enrolled node is
recorded as carried and would remove the foundation, and nothing on the wire orders one
declaration against another.
Measured: AMQP is spoken in three places of the mesh's own code and in none of the
sdk or the modules; the predecessor's world is AMQP and retiring. NATS answers every
guarantee the bus relies on, durability via JetStream. Recommended: not underneath
the migration, not after it either — in parallel, one rehearsed rollout.
Two birth-address outages and a registry that would have been the third; a
container that keeps a stale environment after its file changes; a host command
that applied a converged declaration to an adopted node; the hub and the vault
without seats. And the decision the operator made under it all: the hub adopts
the predecessor's tunnel in place, key and peers and range and port.