Commit Graph
570 Commits
Author SHA1 Message Date
jochen a11da86591 ADR 0108: a route carries the policy applied to a request
Issue 116 found the mesh's proxy applies nothing to a request — host lookup, forward. Against
what the replaced ingress actually relies on, four capabilities are missing: authentication
(three dependents, each gating an admin surface with no login of its own), refusal scoped to a
path (one, a live incident mitigation), path-scoped routing with priority, and redirect.

Policy goes on the route rather than beside it. A proxy-side settings layer keyed by route name
would keep the grant literally clean, but then "what protects this route" is answered from two
files nothing keeps in step — and a route's protection is part of what a route is.

The set is closed at those four, so a fifth is an amendment and each addition is earned by a
dependent that exists. An open middleware surface was rejected: it recreates what is being
replaced, and narrowing one later is far harder than widening a closed one.

Where policy needs a credential the declaration names a secret and never carries the value,
which keeps the existing secret machinery the only thing holding credentials. Inlining a hash
was rejected as the first credential in a declaration — a precedent easier to set than withdraw.

This re-keys the routing table by host and path with priority, which follows from the decision
rather than being a separate one: two of the four need one host routed more than one way. Equal
priorities must resolve identically every time or the proxy stops being reproducible.

The record says how it is checked, including the negative case that rots quietly — a
declaration carrying a credential value rather than a reference must be refused, so the
rejected option cannot return by accident.

08-connectivity §3 names the record and gains the subsection; issue 116 gains amended-design.
2026-09-25 13:48:18 +02:00
jochen c839d9ac26 Issue 116: scrub the disclosure, and correct the count and the shape of the gap
Two things the report got wrong, and one it could not have found the way it looked.

Disclosure first: it carried a real hostname and an absolute node path, in a public
repository. Both are gone; the ingress, the modules and the routes are named by role, as the
rest of 04-ISSUES does.

The count was low. Basic authentication has three dependents in the catalogue, not one — the
key-value store's browser UI, a database web UI, and the ingress's own dashboard. All three
are credential-less admin surfaces whose only gate is a middleware the mesh's proxy lacks.
The earlier version read only the node's dynamic configuration directory, which cannot see
what modules declare as container labels; counting needs both sources, and the report now
says so.

Two gaps were missing entirely. Redirect rules: two live routes canonicalise a www name onto
its apex, they exist only on the node and not in the catalogue, and they fail silently rather
than erroring. And path-scoped routing with priority, which is the one that reorders the
issue: the table maps host to exactly one target, so a host cannot be routed two ways, and
the refusal rule matches a path on a host already routed elsewhere. Authentication and a
source filter would not make it expressible. Path scoping is a prerequisite, not a sibling.

Also corrected: the refusal rule was described as an address-scoped deny. It is an allow-list
holding a single documentation-range address — deny-everyone — so reading it as address-scoped
points at the wrong fix. And its severity was understated: its own header records it as
incident response closing an abused write primitive, which is not "a real exposure" but a live
mitigation.

The open questions now say plainly that they are design questions and the fix should not be
written before they are answered, and one is added: whether a declaration may carry a
credential at all.
2026-09-25 13:24:34 +02:00
jschoubben 226d556743 Issue 116: route-proxy has no authentication or IP-restriction mechanism
Comparing route-proxy against what HAL's actual Traefik config does today,
not Traefik's general feature set, per the standing rule that the nox mesh
must do at minimum what the HAL mesh it replaces already does. Everything
else checked out even or better; these two are real, confirmed gaps —
RedisInsight has no login of its own and depends entirely on Traefik's
basicauth middleware, and the gitea-internal route depends on an IP-scoped
deny rule. Neither has any equivalent in route-proxy's single-lookup
request path.
2026-09-25 11:51:56 +02:00
jschoubben cdcd4da27e Merge pull request 'Research 015: reopen the comparison — the premise for narrowing to one candidate was false' (#105) from storage/015-reopen-the-candidate-comparison into main 2026-09-24 16:47:18 +00:00
jochen 1d524a1fa8 Research 015: rewrite the comparison — wrong axis, and a missing candidate
The previous version ranked candidates on whether they preserved single sign-on to the
object store's console. That is not a requirement: a "user" of the store is normally an
application, so the requirement is per-application keys scoped to buckets — which the mesh
already mints. And the console login it ranked on never worked; the module's own hook comment
records "policy claim missing", a failing login written up as progress.

It also omitted the incumbent's own maintained fork, which changes the question from "which
product replaces it" into two decisions: repoint, or migrate — and if migrating, to which.
Repointing costs an image reference; migrating costs a data copy, two handler rewrites and a
maintenance window. Repointing does not foreclose migrating, which is the argument for taking
it first.

On the corrected requirement Garage ranks first — its per-key-per-bucket model is the
requirement verbatim, its admin API matches how the mesh provisions, and the highest-risk
consumer is first-party documented against it. Its remaining gap (no versioning, no
server-side encryption, partial lifecycle) is unmeasured against the buckets and is the one
thing that could still disqualify it.

Measured and folded in: 230 GiB logical, 82,496 objects, 468 GiB raw at 2.03x, eight drive
directories on one filesystem on one machine. That last fact decides more than any feature —
the erasure coding is not buying independent-drive redundancy, so the redundancy model is
close to irrelevant and only storage overhead remains, which at this volume is a rounding
error against the headroom.

Both errors are recorded at the end of 01 rather than quietly fixed. A configured feature is
not an observed one; and when a dependency dies, "who took it over" precedes "what replaces
it" — searching for alternatives by construction returns things that are not the incumbent.
2026-09-24 18:46:47 +02:00
jochen 999636e2a8 Issue 113: retract the diagnosis table — the original report was right
The diagnosis carried a table headed "claims that could not be substantiated", denying a
module.json, a digest pin, and an all-zeros runtime digest. All three exist. The table is
withdrawn in full and replaced with what is actually true, plus the two claims that remain
genuinely unverified rather than disproven.

The cause: one repository was searched and absence in it was written up as absence. The
catalogue of the mesh being built is a separate repository, not checked out where the search
ran, and all four claims were about that repository. Compounding it, the predecessor's
object-store module and the one being cut over to were treated as one thing — they are
different files in different repositories, one pinning a tag with no sidecar, the other a
digest with two container resources.

Also corrected in the report: located-in named the wrong repository; the "pins a tag" passage
described the predecessor; the open question about pinning by digest is struck, because this
module already does and it made no difference — a deleted digest resolves to nothing either
way. The section on why nothing broke is now scoped explicitly to the predecessor's
machinery.

The lesson kept in the record: "zero occurrences anywhere in the tree" is only as strong as
the tree searched, and a diagnosis must say which tree. A confident rebuttal of a correct
report is worse than no diagnosis — it sends the next person to the wrong place with a
written record behind them.
2026-09-24 18:46:47 +02:00
jochen 48abc36b5b Merge remote-tracking branch 'origin/main' into work/object-store-records 2026-09-24 18:41:53 +02:00
jschoubben 2c5f805467 Merge pull request 'Add hq-defer: park a thought without moving the work off course' (#107) from meta/hq-defer-skill into main 2026-09-24 16:39:22 +00:00
jochen db2c950ba6 hq-defer: make the MEMORY.md pointer an explicit placeholder
Review caught it reading as a real relative link, so a link checker flags
.claude/skills/hq-defer/file.md forever. Angle brackets say placeholder.
2026-09-24 18:38:43 +02:00
jochen 971d0839f5 Add hq-defer: park a thought without moving the work off course
A thought raised mid-task needed remembering but not working on, and there was no
mechanism for that — so it was recorded by hand. This is that, made repeatable.

Records to Claude's persistent memory rather than the repository, deliberately. A parked
thought has no number, owner or status: giving it one asserts triage that deferring says
has not happened. A shared "deferred" document would be a central status file, which
AGENTS.md forbids. And a repository write means a branch, a commit and an MR — the drift
the skill exists to prevent.

Wraps no playbook, because deferring precedes the development cycle rather than being part
of it. It does say which playbook a thought would need if it graduates, and that recording
"undetermined" is the honest answer when the evidence does not say.

The stop condition is the substance: at most two lines, then return to what was in
progress. No plan, no triage question, nothing opened.
2026-09-24 16:40:24 +02:00
jschoubben 5677e97508 Merge pull request 'ADR 0107: persistent data is a directory bind, never a named volume' (#106) from decide/0107-persistent-data-is-a-directory-bind into main 2026-09-24 14:24:17 +00:00
jschoubben a93743708c ADR 0107: persistent data is a directory bind, never a named volume
Records the rule the operator gave directly, mid-session, after checking
that HAL's own postgres and lavinmq both used a directory bind and the
mesh's adoption of them three weeks ago switched to a named volume without
a reason recorded anywhere.

Already built and rolled out on novox (mesh-catalog PR #54) before this
record -- urgent enough to fix first and write down after. Includes the
incident: the new host directories needed the container's own UID, which a
named volume gets for free and a directory bind does not; mesh-store
crash-looped on Permission denied until ownership was matched to what the
original volume already had.

Closes issue 115. Checks pass.
2026-09-24 16:21:51 +02:00
jochen c8cbbcfb8a Research 015: reopen the comparison — the premise for narrowing to one candidate was false
SeaweedFS was scoped as primary because it looked like the only candidate preserving
OIDC console login. Measured: its admin UI is Apache-2.0 but its identity-provider
integration is not — console SSO sits behind the per-TB commercial licence, alongside
point-in-time recovery and automatic EC repair. The free build gives OIDC on the S3 API
via STS and a console authenticated by local username and password.

So the answer to the gating question is that no candidate preserves the current feature
set for free, which this effort had written down as a possible outcome. Reopened across
three candidates with the requirement-by-requirement evidence in 01.

Two corrections to what the overview recorded. RustFS is not a binary-level drop-in
retaining existing data: API and on-disk compatibility are separate paths and the on-disk
one is preview-scoped. And it carries an open defect in the credential path the bucket
provision depends on, which gates it specifically.

Nothing graduates before two measurements named in 01: whether an authenticating proxy
is an acceptable answer to console SSO, and which S3 endpoints consumers actually call —
the latter because Garage does not implement the full span and cannot be ranked until
that is counted.
2026-09-24 16:10:33 +02:00
jschoubben a458f751c6 Merge pull request 'Accept ADR 0038; close issue 091' (#104) from issue/091-ports-are-mesh-assigned-not-manifest-fixed into main 2026-09-24 13:53:36 +00:00
jschoubben 402b798ab6 Accept ADR 0038; close issue 091
The decision (the mesh assigns a container's machine-side port; a module
says only what it needs) was proposed 2026-09-01, and the machinery
already implements it in full -- internal/inventory/ports.go's PortFor,
declaration.go's publishedOn. What was missing was the catalogue actually
complying: 14 of 46 modules baked a machine-side number into their own
manifest anyway. mesh-catalog PR fixes 11 of them (the two defensible
kinds -- foundation, protocol-fixed -- are left alone, per the issue's own
categories). Accepting the decision now that it's actually enforced, and
closing the issue it was blocking.

Checks pass.
2026-09-24 15:52:45 +02:00
jschoubben f10dce4f9e Merge pull request 'Issue 113 and research 015: the object store's images are gone upstream, not access-restricted' (#103) from storage/113-the-object-store-lost-its-upstream into main 2026-09-24 13:45:34 +00:00
jochen 4beb6629db Issue 113: ground the rebuildability point in what the design actually says
Review of my own text found an unattributed claim — "the mesh's claim that a node
can be rebuilt from its declarations" — which is not a stated principle anywhere.
Replaced with the design position that genuinely covers it: to-be 07 chooses
references over payload because "reproducibility comes from pinning the identity of
a thing rather than carrying its bytes". This incident is that choice's failure mode
when the identity stops resolving, which is a sharper point than the one I made.

Scope stated honestly: the passage is about the foundation bundle and this module is
not in it, but pin-identity-fetch-bytes is how every module gets third-party images.

Also names the tension the mirroring question actually carries — mirroring is a move
away from references-over-payload, so it is a decision, not a fix.
2026-09-24 15:44:44 +02:00
jochen d497b37e43 Issue 113 and research 015: the object store's images are gone upstream, not access-restricted
The symptom arrived diagnosed as "the registry disabled anonymous pulls for the
whole vendor namespace". It did not hold: sibling repositories in that namespace
pull normally, the "$disabled" token field appears on every repository including
working ones and describes signing rather than access, and "actions": [] with a
401 is byte-identical to what an invented repository name returns. Both registries'
own APIs establish deletion instead.

Recorded because the correction is the expensive part to rediscover, and because
the instance was harmless while the standing condition is not: no node that does
not already hold the images can ever provision the module again, and nothing
detects that until one tries.

Research 015 scopes the replacement. It is not a redesign — the foundation design
already commits to S3 the protocol rather than the product, and the object store
is an ordinary module, so this instantiates an existing principle. The live OIDC
wiring is the requirement that gates the choice, and it is checked first.
2026-09-24 15:27:03 +02:00
jschoubben 183b22997c Merge pull request 'Issue 112: diagnose — the predecessor's own DNS config already names the carried peers' (#101) from issue/112-diagnosis into main 2026-09-24 12:49:50 +00:00
jschoubben 8d67cf63c5 Issue 112: status located, not diagnosing
Playbook 03 step 2: move status to diagnosing, then located once the
owner is known. located-in is filled with four confirmed packages —
the owner is known.
2026-09-24 14:27:25 +02:00
jschoubben 75c104c355 Issue 112 diagnosis: correct located-in attribution
The carried-peer record (CarriedPeer/TunnelPeer) lives in mesh-controller
internal/inventory, not internal/catalogue. internal/catalogue is the
right package for the zone-generation side of the fix (facts.go's
nodeZones), but a different concern from where the name field itself
would go. Split the two so a decision record doesn't get pointed at the
wrong package.
2026-09-24 13:54:20 +02:00
jschoubben 6a56738d7b Issue 112: diagnose — the predecessor's own DNS config already names the carried peers
Checked why ADR 0104's forward-to-predecessor shape doesn't transfer to the
resolver the way it did the proxy: DNS is one process on one port, and
assigning the mesh's dnsmasq module replaces it in place, so there is no
predecessor process left standing to forward to.

But /etc/dnsmasq.d/hal-dns.conf's static address= lines for ace/shanks/g14
match the mesh's own carried-peer addresses from overlay show exactly. The
name a carried peer needs isn't a guess the operator has to make under
pressure — it's a transcription of a record the predecessor already has and
has been correctly serving for six days. Located in mesh-controller (no way
to attach a name to a carried peer today) and the dnsmasq module (doesn't
emit a wildcard for a named-but-uncarried peer). Not implemented.
2026-09-24 13:25:44 +02:00
jschoubben 91a5c63d65 Merge pull request 'Name the migration repository in the map, so nobody has to be told it exists' (#99) from meta/name-the-migration-repository into main 2026-09-24 00:02:07 +00:00
jschoubben 545d038198 Name the migration repository in the map, so nobody has to be told it exists
hq cannot hold the migration's operational record — it names machines, addresses and
paths, and this repository is public — but it can say where that record is, which is what
this map is for. Asked for by the operator, who had to be told.
2026-09-24 02:01:41 +02:00
jschoubben 7f438d049f Merge pull request 'Issue 092: genesis publishes to a registry the container runtime does not yet trust' (#79) from issue/092-genesis-registry-trust into main 2026-09-23 23:38:46 +00:00
jschoubben 60e43f9446 Merge pull request 'Issue 091: a module definition carries a machine port' (#78) from issue/091-machine-ports-in-manifests into main 2026-09-23 23:38:40 +00:00
jschoubben 36aa722c2a Merge pull request 'Issues 111 and 112: the resolver was told the wrong set of names, twice over' (#98) from issues/111-112-the-resolver-was-told-the-wrong-names into main 2026-09-23 23:32:42 +00:00
jschoubben f704e2ca64 Issues 111 and 112: the resolver was told the wrong set of names, twice over
111, resolved: the map the control plane hands a resolution holds the machines and the
names the mesh merely serves, and the resolver's zones were given both — inventing names
under a suffix it answers authoritatively for. 112, open: adopting a tunnel gives the mesh
the peers' addresses and none of their names, so taking the resolver before they enrol
stops three machines resolving at all.

Both found by reading the plan before pushing it.
2026-09-24 01:32:04 +02:00
jschoubben e7a90be3ee Merge pull request 'Issue 110: on a converged node a container on the runtime's own network cannot reach the resolver' (#97) from issues/110-the-resolver-and-the-default-network into main 2026-09-23 23:13:48 +00:00
jschoubben 3a8515273d Issue 110: on a converged node a container on the runtime's own network cannot reach the resolver
Found reviewing the resolver's conversion. Nothing fails while the node is adopted; it
fails at the flip, and it is the same split that decided which container survived the
hub's address change.
2026-09-24 01:13:30 +02:00
jschoubben 0e70ca0808 Merge pull request 'Issue 109: a container keeps the address it was made with' (#96) from issues/109-a-container-keeps-the-address-it-was-made-with into main 2026-09-23 23:02:48 +00:00
jschoubben 6ccd138729 Issue 109: a container keeps the address it was made with
Found when adopting the tunnel moved the hub's address: the declaration followed, the
running container did not, and the forge lost its database. Issue 102's rule broken one
level down, and issue 103's fix stopping one input short.
2026-09-24 01:02:16 +02:00
jschoubben 07ee79199a Merge pull request 'ADR 0105: what review settled — the tunnel adoption is implemented' (#95) from decide/0105-implemented into main 2026-09-23 22:39:08 +00:00
jschoubben f660620637 ADR 0105: what review settled — carried peers, the flip, the refusals, and keeping the hub's identity
Implemented in mesh-controller #49 and mesh-host #24. One proposal was rejected on the
record's own terms: converging the hub is not made to wait on other machines' migrations.
2026-09-24 00:38:54 +02:00
jschoubben f61f047a5d Merge pull request 'Issue 102 resolved and verified on the machine; issue 097's orphan was on the host network' (#94) from issues/102-resolved-and-097-worse into main 2026-09-23 22:15:49 +00:00
jschoubben 133e10a738 Issue 102 resolved, verified on the machine with both forwarders gone; 097's orphan was on the host network
The addresses follow: the control plane holds the ports the node gave, a recorded build
holds no address at all, and the two forwarders that had been holding the control plane
together are removed. 097's stranded container turned out to be listening on every
interface and connected to the mesh's store — by its own old database, which is the only
reason nothing was at risk.
2026-09-24 00:02:13 +02:00
jschoubben c209a575e1 Merge pull request 'ADR 0106: the bus is NATS; issue 104 resolved' (#92) from decide/0106-the-bus-is-nats into main 2026-09-23 21:40:03 +00:00
jschoubben e022798858 ADR 0106: the bus is NATS — native, built beside the migration, cut over after its core; issue 104 resolved 2026-09-23 23:39:17 +02:00
jschoubben 6f173a7912 Merge pull request 'Issue 108: the registry has no garbage collection, and two doors make it harder to add' (#91) from issues/108-registry-gc into main 2026-09-23 21:32:52 +00:00
jschoubben 73091dcb4c Issue 108: the registry has no garbage collection, and two doors make it harder to add 2026-09-23 23:32:29 +02:00
jschoubben f6ec64ee4e Merge pull request 'Issue 107: a declaration carries no order; rescue on an enrolled node is reconcile, not apply FILE' (#90) from issues/107-declarations-carry-no-order into main 2026-09-23 21:27:24 +00:00
jschoubben 671c2f3881 Issue 107: a declaration carries no order; rescue on an enrolled node is reconcile, not apply FILE
Both from the review of the issue-104 fix: a hand-applied file on an enrolled node is
recorded as carried and would remove the foundation, and nothing on the wire orders one
declaration against another.
2026-09-23 23:27:08 +02:00
jschoubben 2445d80565 Merge pull request 'Research 014: the bus on NATS — decide now, build in the lab, cut over once after the core' (#89) from research/014-nats into main 2026-09-23 21:15:22 +00:00
jschoubben a548b34f5d Research 014: fix the reference to ADR 0039 2026-09-23 23:14:57 +02:00
jschoubben b59907ee08 Research 014: the bus on NATS — decide now, build in the lab, cut over once after the core
Measured: AMQP is spoken in three places of the mesh's own code and in none of the
sdk or the modules; the predecessor's world is AMQP and retiring. NATS answers every
guarantee the bus relies on, durability via JetStream. Recommended: not underneath
the migration, not after it either — in parallel, one rehearsed rollout.
2026-09-23 23:14:39 +02:00
jschoubben 8638ba3a4f Merge pull request 'Issues 102–106 and ADR 0105: what the core migration found, and the hub adopting the predecessor's tunnel' (#88) from core/issues-102-106-and-tunnel-adr into main 2026-09-23 20:54:53 +00:00
jschoubben cb2117f1c4 Issues 102–106 and ADR 0105 from the core migration
Two birth-address outages and a registry that would have been the third; a
container that keeps a stale environment after its file changes; a host command
that applied a converged declaration to an adopted node; the hub and the vault
without seats. And the decision the operator made under it all: the hub adopts
the predecessor's tunnel in place, key and peers and range and port.
2026-09-23 22:50:10 +02:00
jschoubben ee2bdf220c Merge pull request 'Issue 101: taking a service its neighbours reach by container name cuts them off' (#87) from issues/101-a-service-reached-by-name-loses-its-network into main 2026-09-23 18:14:08 +00:00
jschoubben 61e4e971a4 Issue 101: taking a service reached by container name cuts its neighbours off
Found checking the third cutover rather than running it. The first two were safe by
accident — both are reached through a host port, which survives a change of owner.
This is the first constraint found that decides the order of the migration.
2026-09-23 20:13:53 +02:00
jschoubben f4f58e7c30 Merge pull request 'Issue 100: a secret the mesh mints cannot be the one the service it takes over already uses' (#86) from issues/100-a-minted-secret-cannot-be-the-one-the-service-already-uses into main 2026-09-23 00:54:21 +00:00