Commit Graph
236 Commits
Author SHA1 Message Date
jochen a1d8b478ec Issue 121: retract step 4 — the forge's service is not built
Step 4 claimed the forge cannot exist as a container before the builder has built its image. The
forge's server is an upstream public image pinned by digest; only its runtime sidecar is built. The
store and the broker have the same shape, and ADR 0078 raises both at genesis as plumbing and adopts
them in place — so the forge can be raised the same way and serve git, packages and OCI before
anything is built.

What survives: a grant is minted by the provider's runtime sidecar, which is built, so the question
is whether raise-service, grant, build-sidecar simply works. Sequencing inside the mesh, not images.

The wrong version came from taking a record's bootstrap argument at face value instead of comparing it
to how the store and broker are raised — one command away in the manifests.
2026-09-26 15:36:19 +02:00
jochen c343fc68c5 Issue 108: the second door was attached to the wrong thing
Two things were treated as one. The mesh's own artifact store holds the store seat and is internal by
design — reached by name over the overlay, no accounts, ADR 0082. Serving a registry publicly is a
service the mesh can host: a module with its own name, accounts and storage, like anything else it
runs for somebody. The conversion this report was written beside gave the seat holder a second public
door over the same filesystem, which is neither.

Keeps the original issue whole — no garbage collection, and the settings a routine needs are not
enabled — drops the two-door complication, and sharpens one thing: deletion on the only door is
deletion on a door with no accounts, which the predecessor kept behind its authenticated one.
2026-09-26 15:22:09 +02:00
jochen e399a2c148 Issue 122: the pattern is already on main, in five modules
Asked whether merging the three reviewed changes would set a precedent. It would not: five modules
already carry a name belonging to this one mesh — a workflow module stating its host, protocol and
absolute webhook URL, and two carrying a full clone URL for a repository on the mesh's own forge.

That changes what the issue is for. There is no version of this catalogue today that does not name
the mesh it was written in, so refusing three changes buys nothing and a mechanism is the only thing
that removes any of them. The three were merged on that reading, each PR saying so.
2026-09-26 15:07:22 +02:00
jochen d7078061ea Issue 122: a module cannot ask for its own public name
Three open module changes independently wrote one mesh's names into the catalogue — two literal
public URLs, because the software generates absolute URLs behind a proxy, and one bucket renamed to
match what exists here. None was careless: the mesh composes <label>.<public-domain> for the proxy
and never hands it back to the module that asked for the route, and no interpolation yields a public
name, so writing the answer down is the only expressible option.

Files it rather than blocking the three, because the fix is a mechanism and the instances are live
needs. The cost is stated: a second mesh installing the identity provider gets the first mesh's
hostname, and nothing distinguishes a literal domain from a version number.
2026-09-26 14:58:18 +02:00
jochen 9766a3afce Issue 121 diagnosed: the seats work renamed the requirement, not the order
Asked first whether the seats change fixed this in passing, since it landed the same day and
touches both manifests the report names. It did not: the seat's holder is consulted only where
several nodes provide the thing, and with none providing it resolution refuses outright. Genesis
has no exemption — the unchecked first pass exists to learn what each node offers, and a
declaration is never built from it.

Records the part that did change: the three tests left failing on purpose were deleted by the
controller's seats PR and replaced with passing seat-based ones, so the gap is invisible again.
Adds the resolver to located-in, since that is where the refusal is.
2026-09-26 14:47:45 +02:00
jschoubben 99ffa6474d Merge pull request 'Issue 121: builder's real package-registry grant deadlocks a genesis bootstrap' (#117) from issue/117-builder-package-registry-deadlocks-genesis into main 2026-09-26 12:20:53 +00:00
jochen 09e502058b Issue 121: builder's real package-registry grant deadlocks a genesis bootstrap
Renumbered from 117, which is taken on main by 'a module's own code is a container in one record
and a process in another' — two reports claimed the same number and git would not have said so.

Scrubbed the node's name and a real registry path; this repository is public.
2026-09-26 14:20:05 +02:00
jschoubben 74b88d8efc Merge main 2026-09-26 14:19:58 +02:00
jochen b083790b21 Issue 118: the analytics store answers the dial and times out the query
Keeps 118: the other claimant to this number is on main as issue 119, where ADR 0112 points.

Scrubbed the service's public name — this repository is public — and completed the report's
frontmatter with the fixed-by and amended-design keys every other report carries.
2026-09-26 14:19:40 +02:00
jschoubben 0740de9d14 Merge main 2026-09-26 14:19:33 +02:00
jochen d2044fb7b4 Merge main: the bus design, issues 113/114/117/120 and research 017 landed
# Conflicts:
#	04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md
2026-09-26 14:16:05 +02:00
jschoubben 01c6b89cc5 Merge pull request 'Design 25: the bus on NATS — proposed architecture for review; issue 103 resolved' (#93) from design/25-the-bus-on-nats into main 2026-09-26 12:15:19 +00:00
jschoubben 93502a05dc Merge pull request 'Issue 117: a module's own code is a container in one record and a process in another' (#110) from issue/117-a-modules-own-code-is-a-container-and-a-process into main 2026-09-26 12:14:38 +00:00
jschoubben de0c9b6cfc Merge pull request 'Issue 120: a provisioner remembers what it did, not what is there' (#114) from issue/120-a-provisioner-remembers-what-it-did-not-what-is into main 2026-09-26 12:14:17 +00:00
jschoubben 77934991c4 Merge pull request 'Issue 113 resolved by the repin, and what issue 064 did not cover' (#108) from issue/113-record-the-repin-and-fold-114 into main 2026-09-26 12:13:52 +00:00
jochen b3bd50c588 Issue 120: a provisioner remembers what it did, not what is there
The harness compares against its own memory, so a backend that loses
what was provisioned (the cache's ACL users on a server restart) is
never provisioned again, silently.
2026-09-26 00:44:10 +02:00
jochen e387c4bd0e Apply review: two credentials, staged admin rotation, a ninth provider
The fact-check found mailu, whose user is its mailbox, so 0114 rotates
over two credentials rather than two logins, the adapter choosing what a
credential is. Also: minio keeps non-empty buckets; five backends take
their admin credential only at first init, so single-party rotation is
staged; postgres ownership moves to a non-login role; the harness keys by
consumer; rotation state lives with the vault. Consistency fixes across
0110-0113, 26 and 27; issue 103 resolved by mesh-host PR #22.
2026-09-26 00:38:06 +02:00
jochen 6e3373c879 Design pass: address the review
0113 — the plaintext claim was false under its own mechanism: handing a provider's answer to the
controller puts every secret on the broker and in the controller in the clear. The provider now seals
each secret field itself, to the consumer node's public key the mesh hands it, and the controller
carries sealed fields it cannot open. That is stricter than today, where the controller holds every
minted credential in the clear. Option 3 (plaintext to the controller) is recorded and rejected. The
foundation exception now covers root-secret rotation (0085) and forms like the broker admin's hash, so
no phase claims to remove the broker's bootstrap step. To-be 24 and 13 are named among what it amends.

27 — resolution is consistent with 0110: co-location and the only provider apply only where no seat
delivers the provision, so an unheld seat is refused even with one provider. The secret-field rule now
matches 0086 exactly (a declared env-file, never a container environment value). The seat placeholder
is the controller's, and the one module reading it moves to a host port. Contracts are held by the
controller and written down in phase 1, so they can be checked; every rule has a check. An operator's
secret is still the operator's, with the vault as custodian. Which seats a module holds is listed as
not settled.

0110 — the unheld-seat-with-one-provider case and the one-answer-for-everyone rule have checks; the
claim about moved manifests is corrected. 26 — the table governs and the code catches up, not the
reverse; scope and capacity agree with the glossary; moving a seat is described as it really is today.
0112 — aligned with 27, and lists 0049 and 26 among what it changes.

Issue 118 is renumbered 119: another branch took 118 first. 'Control-plane' is gone from 0110 and 0111.
2026-09-25 22:46:10 +02:00
jochen 7668190154 ADR 0112 and issue 118: address the review
- Secrets follow ADR 0085 as amended: a module's own secret is a provision the controller mints and
  the vault records. The previous commit had that backwards. Whether the vault should generate
  instead is recorded as an open question, not decided.
- A directory's contract is owner and mode only. The persistence flag was the keep flag ADR 0030
  refused; a directory is kept while it holds anything, and disposable data is a named volume (0107).
- An operator's shared data stays an access (ADR 0051), which rejected an operator-owned directory.
  Only where its path is written moves to the assignment.
- The records it changes on acceptance are named: 0051, 0091, 0046 (settings keyed by instance),
  0084 (a provider is a node and an instance), and the glossary, which gains its new words only
  when the record is accepted.
- How it is checked covers every stated rule. Container-side paths are no longer flagged by the
  host-path rule, and code fallbacks are covered.
- Provisions are what other modules provide. A seat's occupant is not listed as one, and the vault
  is not described as selectable per assignment.
- 'Control plane' becomes 'controller'. The provider count is ten of eleven, not eleven of twelve.
2026-09-25 22:29:38 +02:00
jochen ca235e775f Issue 118 and ADR 0112 (proposed): a module definition names no node, no mesh and no path
Issue 118 records what a review of where module code reads its files found: 789 host-path strings
in 70 of the catalogue's 71 definitions, every one a decision the definition makes about a machine.
Mounts are checked (ADR 0091); the same paths retyped as values are not. It records what that has
already allowed — a DNS provider that would provision nobody silently, a contributions file that
names credentials by host path and so forces every provider to mount at the identical path, an SDK
loop that treats an unwritten contributions file as empty without a word, defaults in code that
disagree with their own manifests — and that no module can be assigned to one node twice, because
every identity is keyed by the module's name.

ADR 0112, proposed for review, answers it the way ADR 0038 answered ports: a definition names
variables, and installing it resolves every one or refuses, from three sources — the assignment's
own configuration, provisions the mesh resolves against a contract, and what the mesh generates or
knows. A directory becomes a provision: the module requires one by name with its owner, mode and
persistence, and where it lands is the assignment's. The mesh's own files stop carrying host paths.
An assignment gets an identity of its own, so a module may run twice on one node.

Checking copies for agreement was rejected as checking something that should not exist; rewriting
paths per assignment was rejected as inferring which strings are paths by their shape. Syntax, a
node's default layout, and when a second instance becomes possible are left to the design.
2026-09-25 22:29:38 +02:00
jschoubben 74ae0609cb issue 118: umami's store answers the dial and times out the query 2026-09-25 22:01:28 +02:00
jschoubben 2ca63ae54e 117: builder's real package-registry grant deadlocks a genesis bootstrap
Fixing builder's hand-faked package-registry binding tonight (requires:
package-registry, a real mesh grant instead of a hardcoded JSON fragment)
broke three tests describing a deliberate carried-binding fallback for
exactly this: gitea's own image is built by builder, so builder cannot
yet hold a real grant from gitea the first time either has to exist.
Invisible on novox (already bootstrapped, gitea already live) — real on
any genesis from scratch. Fix left in place, tests left failing rather
than reverted or hacked, so the gap stays visible.
2026-09-25 16:59:04 +02:00
jochen 82a6badc7c Issue 114: land the controller's container-or-process question, renumbered
Filed 2026-09-24 on a branch of its own and never merged, numbered 113, which is taken. 114 is
free because a sibling branch folded it, so it takes that number and keeps its commit.

Kept separate from issue 117 rather than folded into it. 117 asks the same question of every
module and locates the missing decision; this asks it of the controller, where `network: host`
means container network isolation — the property that resource type usually buys — is not in use.
That observation is this report's own and is nowhere in 117, and folding would lose it.

Its first open question is answered by 117's diagnosis and now says so: the host's `process` shape
is built, applied and tested, restart and run-to-completion semantics included, so deciding this
does not wait on host-side work.
2026-09-25 16:16:21 +02:00
jschoubben 10a2b706c6 Issue 113: should the controller be a container or a process the host supervises
Filed after a session where every mesh-controller interaction went through
docker exec — its manifest runs it as a container with network: host, using
none of the isolation that resource type usually buys, while ADR 0006 makes
it the mesh's single point of coordination. Open question, not a claimed
defect: does type: container get the controller anything type: process
(supervised the way the host supervises its own unit, per ADR 0005) would not.
2026-09-25 16:15:27 +02:00
jochen 35db2aaa41 Issue 117: a module's own code is a container in one record and a process in another
Asked what the "sidecar" is and whether a supervised process would do instead. The repository
answers both ways. ADR 0047 (accepted, unsuperseded) says a module with tools or events runs a
container carrying its compiled code. To-be 18 and 20 (both proposed) define a `process` resource
type — the module's own code, a unit the machine's supervisor keeps up — and the worked guide says
plainly "it is why these are `process` rather than four containers." Neither design doc names 0047,
and no decision record mentions a `process` shape at all.

Diagnosed rather than left open, because the ground truth settles what the report could not.
The shape is real: mesh-host defines TypeProcess, applies it, and tests it, and the host's
vocabulary is twelve shapes rather than the nine ADR 0029 counted. So the alternative the report
offered — that two proposed documents describe a type that does not exist — is disproven.

ADR 0029's mechanism is intact and was not enough. The vocabulary-count test names the decision
behind each addition: network 0029, access 0051, opening 0100. The eleventh names a *proposed
design document*, and TypeProcess is the only shape in the vocabulary whose doc comment cites no
ADR. Requiring every addition to name something does not require it to name a decision.

The argument this issue asked for already exists — as a Go test comment. "It is a full-host shape
rather than a portable one: it needs a process supervisor to install into. It does NOT need a
container runtime, which is the point — only software that genuinely needs isolation asks for a
container." That is a decision's context and consequences, in another repository.

What the catalogue does is a third thing: 115 container declarations against 3 process, all three
in showcase — the module to-be 20 documents. There the tools resource is a container running
`sleep infinity` on a bare upstream base with the broker credential mounted, and the tools and
provisioner entrypoints are run by nothing. That is the condition 0047 was written to end, back
in a new shape.

Where the isolation argument leaks is narrower than expected and worth having precisely: the
serving key and the credential shape both conform. But serveTools serves every registered module
over one broker connection, the runtime takes its modules from a comma-separated list, and
x-source is stamped from the single credential — so two modules in one runtime means the second's
events are attributed to the first. Nothing refuses it and no test asserts against it.

Located on hq rather than on a code repository: the implementation and the design layer agree,
and the missing thing is the record. Which shape is right is left open, deliberately — this
establishes that the question was answered in practice and never written down, not which answer
is correct.

One correction kept in the trail: the first search here was for len(Vocabulary()), found nothing,
and was two steps from being written up as "the mechanism ADR 0029 relied on is gone." The test
binds the slice to a local first. A negative search result read as a fact about the world is the
same error issue 113 recorded.
2026-09-25 16:15:00 +02:00
jochen 367df38e6d Issue 116: resolved by mesh-controller PR #58
The gap is closed in the proxy: policy applies, the four capabilities exist, the table is keyed
by host and path with a total ordering, and the two failure modes that rot quietly are held by
tests — a declaration carrying a credential refused rather than served, an unreadable secret
failing closed.

Resolved rather than left open because the issue reports a gap in the proxy and that gap is
gone. But the record says plainly what it does not yet allow: an operator still cannot move the
affected routes, because that needs the mesh side — a manifest able to declare these values and
the controller minting the secret auth names. Until both exist the capability is reachable only
by writing the routes file by hand. That is the ordinary build-out of a contract this issue's
decision created, and it belongs to to-be 08 rather than here.

The open questions are marked answered and kept rather than deleted, pointing at ADR 0108 —
what was rejected and why is the half worth having, and a section still saying "the fix should
not be written before these are answered" after the fix was written reads as though nobody
looked.

One finding kept in the record: priority was read with the reader for ports, which caps at
65535, and the one real rule this reproduces is declared at 100000. It parsed to zero, so
refusal and path scoping would have shipped looking complete and doing nothing on the only case
that motivated them. A validator borrowed from a neighbouring field is a silent default.
2026-09-25 14:25:24 +02:00
jochen a11da86591 ADR 0108: a route carries the policy applied to a request
Issue 116 found the mesh's proxy applies nothing to a request — host lookup, forward. Against
what the replaced ingress actually relies on, four capabilities are missing: authentication
(three dependents, each gating an admin surface with no login of its own), refusal scoped to a
path (one, a live incident mitigation), path-scoped routing with priority, and redirect.

Policy goes on the route rather than beside it. A proxy-side settings layer keyed by route name
would keep the grant literally clean, but then "what protects this route" is answered from two
files nothing keeps in step — and a route's protection is part of what a route is.

The set is closed at those four, so a fifth is an amendment and each addition is earned by a
dependent that exists. An open middleware surface was rejected: it recreates what is being
replaced, and narrowing one later is far harder than widening a closed one.

Where policy needs a credential the declaration names a secret and never carries the value,
which keeps the existing secret machinery the only thing holding credentials. Inlining a hash
was rejected as the first credential in a declaration — a precedent easier to set than withdraw.

This re-keys the routing table by host and path with priority, which follows from the decision
rather than being a separate one: two of the four need one host routed more than one way. Equal
priorities must resolve identically every time or the proxy stops being reproducible.

The record says how it is checked, including the negative case that rots quietly — a
declaration carrying a credential value rather than a reference must be refused, so the
rejected option cannot return by accident.

08-connectivity §3 names the record and gains the subsection; issue 116 gains amended-design.
2026-09-25 13:48:18 +02:00
jochen c839d9ac26 Issue 116: scrub the disclosure, and correct the count and the shape of the gap
Two things the report got wrong, and one it could not have found the way it looked.

Disclosure first: it carried a real hostname and an absolute node path, in a public
repository. Both are gone; the ingress, the modules and the routes are named by role, as the
rest of 04-ISSUES does.

The count was low. Basic authentication has three dependents in the catalogue, not one — the
key-value store's browser UI, a database web UI, and the ingress's own dashboard. All three
are credential-less admin surfaces whose only gate is a middleware the mesh's proxy lacks.
The earlier version read only the node's dynamic configuration directory, which cannot see
what modules declare as container labels; counting needs both sources, and the report now
says so.

Two gaps were missing entirely. Redirect rules: two live routes canonicalise a www name onto
its apex, they exist only on the node and not in the catalogue, and they fail silently rather
than erroring. And path-scoped routing with priority, which is the one that reorders the
issue: the table maps host to exactly one target, so a host cannot be routed two ways, and
the refusal rule matches a path on a host already routed elsewhere. Authentication and a
source filter would not make it expressible. Path scoping is a prerequisite, not a sibling.

Also corrected: the refusal rule was described as an address-scoped deny. It is an allow-list
holding a single documentation-range address — deny-everyone — so reading it as address-scoped
points at the wrong fix. And its severity was understated: its own header records it as
incident response closing an abused write primitive, which is not "a real exposure" but a live
mitigation.

The open questions now say plainly that they are design questions and the fix should not be
written before they are answered, and one is added: whether a declaration may carry a
credential at all.
2026-09-25 13:24:34 +02:00
jschoubben 226d556743 Issue 116: route-proxy has no authentication or IP-restriction mechanism
Comparing route-proxy against what HAL's actual Traefik config does today,
not Traefik's general feature set, per the standing rule that the nox mesh
must do at minimum what the HAL mesh it replaces already does. Everything
else checked out even or better; these two are real, confirmed gaps —
RedisInsight has no login of its own and depends entirely on Traefik's
basicauth middleware, and the gitea-internal route depends on an IP-scoped
deny rule. Neither has any equivalent in route-proxy's single-lookup
request path.
2026-09-25 11:51:56 +02:00
jochen ab7d216fc6 Issue 113 resolved by the repin, and what issue 064 did not cover
The object-store module was repinned to a maintained fork of the withdrawn server image, its
runtime sidecar built rather than pulled, and its data moved off the predecessor's live
directory. The instance is closed; the three general points the report makes are not, and What
was done says so rather than letting a resolved status imply otherwise.

Folds in the one thing a duplicate report of this symptom had that this one did not: issue 064
asked whether the build environment can reach a declared vendor image and assumed that, once
declared, it stays fetchable. Withdrawal is the case that assumption does not cover. The
duplicate is not merged — it carried the reading this report's diagnosis retracts.
2026-09-25 00:55:13 +02:00
jochen 999636e2a8 Issue 113: retract the diagnosis table — the original report was right
The diagnosis carried a table headed "claims that could not be substantiated", denying a
module.json, a digest pin, and an all-zeros runtime digest. All three exist. The table is
withdrawn in full and replaced with what is actually true, plus the two claims that remain
genuinely unverified rather than disproven.

The cause: one repository was searched and absence in it was written up as absence. The
catalogue of the mesh being built is a separate repository, not checked out where the search
ran, and all four claims were about that repository. Compounding it, the predecessor's
object-store module and the one being cut over to were treated as one thing — they are
different files in different repositories, one pinning a tag with no sidecar, the other a
digest with two container resources.

Also corrected in the report: located-in named the wrong repository; the "pins a tag" passage
described the predecessor; the open question about pinning by digest is struck, because this
module already does and it made no difference — a deleted digest resolves to nothing either
way. The section on why nothing broke is now scoped explicitly to the predecessor's
machinery.

The lesson kept in the record: "zero occurrences anywhere in the tree" is only as strong as
the tree searched, and a diagnosis must say which tree. A confident rebuttal of a correct
report is worse than no diagnosis — it sends the next person to the wrong place with a
written record behind them.
2026-09-24 18:46:47 +02:00
jschoubben a93743708c ADR 0107: persistent data is a directory bind, never a named volume
Records the rule the operator gave directly, mid-session, after checking
that HAL's own postgres and lavinmq both used a directory bind and the
mesh's adoption of them three weeks ago switched to a named volume without
a reason recorded anywhere.

Already built and rolled out on novox (mesh-catalog PR #54) before this
record -- urgent enough to fix first and write down after. Includes the
incident: the new host directories needed the container's own UID, which a
named volume gets for free and a directory bind does not; mesh-store
crash-looped on Permission denied until ownership was matched to what the
original volume already had.

Closes issue 115. Checks pass.
2026-09-24 16:21:51 +02:00
jschoubben 402b798ab6 Accept ADR 0038; close issue 091
The decision (the mesh assigns a container's machine-side port; a module
says only what it needs) was proposed 2026-09-01, and the machinery
already implements it in full -- internal/inventory/ports.go's PortFor,
declaration.go's publishedOn. What was missing was the catalogue actually
complying: 14 of 46 modules baked a machine-side number into their own
manifest anyway. mesh-catalog PR fixes 11 of them (the two defensible
kinds -- foundation, protocol-fixed -- are left alone, per the issue's own
categories). Accepting the decision now that it's actually enforced, and
closing the issue it was blocking.

Checks pass.
2026-09-24 15:52:45 +02:00
jochen 4beb6629db Issue 113: ground the rebuildability point in what the design actually says
Review of my own text found an unattributed claim — "the mesh's claim that a node
can be rebuilt from its declarations" — which is not a stated principle anywhere.
Replaced with the design position that genuinely covers it: to-be 07 chooses
references over payload because "reproducibility comes from pinning the identity of
a thing rather than carrying its bytes". This incident is that choice's failure mode
when the identity stops resolving, which is a sharper point than the one I made.

Scope stated honestly: the passage is about the foundation bundle and this module is
not in it, but pin-identity-fetch-bytes is how every module gets third-party images.

Also names the tension the mirroring question actually carries — mirroring is a move
away from references-over-payload, so it is a decision, not a fix.
2026-09-24 15:44:44 +02:00
jochen d497b37e43 Issue 113 and research 015: the object store's images are gone upstream, not access-restricted
The symptom arrived diagnosed as "the registry disabled anonymous pulls for the
whole vendor namespace". It did not hold: sibling repositories in that namespace
pull normally, the "$disabled" token field appears on every repository including
working ones and describes signing rather than access, and "actions": [] with a
401 is byte-identical to what an invented repository name returns. Both registries'
own APIs establish deletion instead.

Recorded because the correction is the expensive part to rediscover, and because
the instance was harmless while the standing condition is not: no node that does
not already hold the images can ever provision the module again, and nothing
detects that until one tries.

Research 015 scopes the replacement. It is not a redesign — the foundation design
already commits to S3 the protocol rather than the product, and the object store
is an ordinary module, so this instantiates an existing principle. The live OIDC
wiring is the requirement that gates the choice, and it is checked first.
2026-09-24 15:27:03 +02:00
jschoubben 8d67cf63c5 Issue 112: status located, not diagnosing
Playbook 03 step 2: move status to diagnosing, then located once the
owner is known. located-in is filled with four confirmed packages —
the owner is known.
2026-09-24 14:27:25 +02:00
jschoubben 75c104c355 Issue 112 diagnosis: correct located-in attribution
The carried-peer record (CarriedPeer/TunnelPeer) lives in mesh-controller
internal/inventory, not internal/catalogue. internal/catalogue is the
right package for the zone-generation side of the fix (facts.go's
nodeZones), but a different concern from where the name field itself
would go. Split the two so a decision record doesn't get pointed at the
wrong package.
2026-09-24 13:54:20 +02:00
jschoubben 6a56738d7b Issue 112: diagnose — the predecessor's own DNS config already names the carried peers
Checked why ADR 0104's forward-to-predecessor shape doesn't transfer to the
resolver the way it did the proxy: DNS is one process on one port, and
assigning the mesh's dnsmasq module replaces it in place, so there is no
predecessor process left standing to forward to.

But /etc/dnsmasq.d/hal-dns.conf's static address= lines for ace/shanks/g14
match the mesh's own carried-peer addresses from overlay show exactly. The
name a carried peer needs isn't a guess the operator has to make under
pressure — it's a transcription of a record the predecessor already has and
has been correctly serving for six days. Located in mesh-controller (no way
to attach a name to a carried peer today) and the dnsmasq module (doesn't
emit a wildcard for a named-but-uncarried peer). Not implemented.
2026-09-24 13:25:44 +02:00
jschoubben 7f438d049f Merge pull request 'Issue 092: genesis publishes to a registry the container runtime does not yet trust' (#79) from issue/092-genesis-registry-trust into main 2026-09-23 23:38:46 +00:00
jschoubben 60e43f9446 Merge pull request 'Issue 091: a module definition carries a machine port' (#78) from issue/091-machine-ports-in-manifests into main 2026-09-23 23:38:40 +00:00
jschoubben f704e2ca64 Issues 111 and 112: the resolver was told the wrong set of names, twice over
111, resolved: the map the control plane hands a resolution holds the machines and the
names the mesh merely serves, and the resolver's zones were given both — inventing names
under a suffix it answers authoritatively for. 112, open: adopting a tunnel gives the mesh
the peers' addresses and none of their names, so taking the resolver before they enrol
stops three machines resolving at all.

Both found by reading the plan before pushing it.
2026-09-24 01:32:04 +02:00
jschoubben 3a8515273d Issue 110: on a converged node a container on the runtime's own network cannot reach the resolver
Found reviewing the resolver's conversion. Nothing fails while the node is adopted; it
fails at the flip, and it is the same split that decided which container survived the
hub's address change.
2026-09-24 01:13:30 +02:00
jschoubben 6ccd138729 Issue 109: a container keeps the address it was made with
Found when adopting the tunnel moved the hub's address: the declaration followed, the
running container did not, and the forge lost its database. Issue 102's rule broken one
level down, and issue 103's fix stopping one input short.
2026-09-24 01:02:16 +02:00
jschoubben 133e10a738 Issue 102 resolved, verified on the machine with both forwarders gone; 097's orphan was on the host network
The addresses follow: the control plane holds the ports the node gave, a recorded build
holds no address at all, and the two forwarders that had been holding the control plane
together are removed. 097's stranded container turned out to be listening on every
interface and connected to the mesh's store — by its own old database, which is the only
reason nothing was at risk.
2026-09-24 00:02:13 +02:00
jschoubben f3ad60b98c Issue 103 resolved by mesh-host #22 2026-09-23 23:44:31 +02:00
jschoubben e022798858 ADR 0106: the bus is NATS — native, built beside the migration, cut over after its core; issue 104 resolved 2026-09-23 23:39:17 +02:00
jschoubben 73091dcb4c Issue 108: the registry has no garbage collection, and two doors make it harder to add 2026-09-23 23:32:29 +02:00
jschoubben 671c2f3881 Issue 107: a declaration carries no order; rescue on an enrolled node is reconcile, not apply FILE
Both from the review of the issue-104 fix: a hand-applied file on an enrolled node is
recorded as carried and would remove the foundation, and nothing on the wire orders one
declaration against another.
2026-09-23 23:27:08 +02:00
jschoubben cb2117f1c4 Issues 102–106 and ADR 0105 from the core migration
Two birth-address outages and a registry that would have been the third; a
container that keeps a stale environment after its file changes; a host command
that applied a converged declaration to an adopted node; the hub and the vault
without seats. And the decision the operator made under it all: the hub adopts
the predecessor's tunnel in place, key and peers and range and port.
2026-09-23 22:50:10 +02:00
jschoubben 61e4e971a4 Issue 101: taking a service reached by container name cuts its neighbours off
Found checking the third cutover rather than running it. The first two were safe by
accident — both are reached through a host port, which survives a change of owner.
This is the first constraint found that decides the order of the migration.
2026-09-23 20:13:53 +02:00