The controller kept one report per machine, replaced, so a resource nothing can
ever apply looked like a failure that had just happened, every few minutes, for
ever. It now counts identical reports and status says stuck after three.
The playbook, the README and the status skill knew five statuses; the cycle check
knew a sixth, 'fixed', and not 'wontfix'. Eleven issues sat in the sixth for weeks
with their fixes shipped, one step short of closed. They are resolved; the check
refuses the word from now on and accepts the one the playbook allows.
ADR 0069 had already placed the controller's manifest in its own repository, and the
raising design called the catalogue copy a thing to remove. The installer's step 3 was
handed that manifest by the build and step 9 read a second copy anyway.
The end-to-end design held 'the run rebuilds what it tests' for binaries and images
and not for manifests. The beds' inline copies fell into three kinds; only the first
is a stale copy. The other two are named: a mesh test wearing a catalogue module's
name (074) and a stocked runtime image the run never rebuilds (075).
A seeded file is created once (0087); the foundation filters before anything
listens (0088); a machine becomes the last thing it was told (design 05).
Each says how it is checked.
Closes issue 041 by decision and by code on the same branch: the catalogue
engine refuses a secret in a container's env, and a secret-carrying env-file
unless the container declares its reason; the controller reads all six of
its credentials from files; design 13 states the rule and how it is checked.
An audit of the six code repositories found eleven open issues fixed on main
with commits and beds to show (025, 027, 033, 036, 037, 040, 045, 047, 050,
052, 053), three partly (007, 026, 035), nine not (020, 031, 041, 046, 049,
054, 064, 065, 066) and one whose fix would live outside those repos (006).
Resolved ones name their evidence; partly ones say what remains; 041 records
that the exposure has widened since it was reported.
The installer half of the amended ADR 0085 is built and proven by the
genesis bed's root-secrets step; the foundation design closes its open item
and the installation design says what the installer does and what it still
cannot check.
A bed run from .work/<slug>/mesh-lab derives mesh-tools and mesh-sdk by
sibling path and fails at once when the directory holds only the touched
repos. Detached worktrees on main, never symlinks.
Recorded on the record, dated, before anything shipped against the sentences
that change. The vault is installed at genesis like the store and broker, one
per mesh, and holds every secret a module has for itself sealed a second time
to an operator key whose private half never enters the mesh — the break-glass
path the first version left open, without a key one place holds.
Design 24 says how; 07 and 21 say what genesis does not yet do; issue 071
names the fixed credentials the foundation is raised with today.
Design 24 flips to in-progress with mesh-catalog as its owner (playbook 04).
Starting the build surfaced two gaps the decision did not settle: a module
requiring `secret` receives exactly one value (069), and no command can
accept an operator's value into a consumer↔vault pair (070). Both opened as
issues rather than improvised around.
Also fills fixed-by on 067 and 068, which the cycle check refused as resolved
with no reference.
ADR 0084 (extends 0027) — a provision is served by a node-scoped provider the
consumer selects, defaulting to co-location; a module may instead carry a private
embedded instance that is not a provision. Design: 01-to-be/23-choosing-a-provider.
ADR 0085 (extends 0031) — a secret is a provision and the vault is the module that
provides it; a module's own local secret becomes an ordinary pair credential that
rotates through the existing machinery, while the controller's provisioning-credential
mint (0048) is unchanged. Design: 01-to-be/24-the-secrets-vault; doc 13 amended to
cross-link the non-pair secret.
Issues 067/068 marked resolved with amended-design set. ADR index regenerated;
records and index checks pass.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Store and broker seats were made ordinary modules; secret-minting is still a
privileged property of the controller that no module owns. Propose the vault
become a module that provides a secret provision (generate/hold/rotate/backup/
audit), node-scoped like every other provider (issue 067), subsuming the three
secret paths and giving local-secret rotation a home.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Naming which provider is only half of how a module gets a database. The other
half: a module may carry its own version/fork-pinned instance, module-network
only, no published port, not a provision — and the model has no word for it.
Add it as a second axis with its own open questions.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The mesh models provisions as mesh-scoped (one provider of a kind, a single
mesh-store). But node-specific services delivered to the mesh was the plan from
the start: both nodes already run their own postgres, SQL server, redis and
object store, and identity — currently single — already serves apps on a second
node. The model cannot express which provider serves a consumer, so it collapses
a deliberately per-node fleet to one. Provider scoping is a whole-mesh decision
across postgres/s3-bucket/oidc, not an SSO patch.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Surfaced while reviewing the host apply loop for the controller/host
split discussion. 065: a permanently-failing resource retries for ever
with no escalation — reported, but never raises its hand as stuck.
066: a partly-applied declaration leaves a mixed state with no rollback,
which is harmless for independent resources and unexamined for pairs
that are only correct together. Both are questions HQ must answer, not
incidents — status open, no fix proposed.
The P1-sweep PR left both at 'located'; the fixes have since merged
(mesh-controller #31, mesh-tools #10) and the bed proves them. Flip to
resolved with their fixed-by.
060: 44 of 70 modules now mesh-buildable (was 8) — the mechanical
majority, proven by direct builds and the bed. The structural
remainders: external-dependency fetch (new issue 064) and route-proxy's
cross-repo build context. 064 records the build environment's isolation
from public npm and Docker Hub.
The broker's port was accepted on input but not forwarded, so a joined
node reached a DNAT'd broker only until it restarted. Fixed in
mesh-controller (25e42b3), proven by the built-store-cross-node bed
adopting the foundation's broker (which restarts it) and the joined
node still receiving declarations.
One push leaves the mesh consistent (the 057 decision, proposed for
acceptance); the shared runtime waits for its broker (058). Fixes on
mesh-control fix/one-push-is-enough and mesh-tools
fix/the-runtime-waits-for-its-broker; the built-store-cross-node bed
enforces both.
Fixed by mesh-controller PR 30 (79a9e17), rebased onto the merged
registry-reach train and proven by a green fresh run of the no-fake
two-node bed built from that commit.
One green fresh run of the no-fake two-node bed is the proof: node2's
consumers open the store and broker the mesh built and adopted, over
the overlay. Fixed by mesh-controller PR 29, mesh-catalog PR 26, gated
by mesh-lab PR 34.
061: the broker module's provisioner never ran; its runtime container
named no command and the image default is the tool host. 062: a failed
artifact-store lookup composed the network without the registry trust,
turning a transient error into permanent silent state. Both located,
fixes on the 042/048 train branches.
Only 8 of 70 modules carry a build section; the other 62 exist through an out-of-band
build script plus placeholder rewriting, so the mesh's own build-and-deliver path has
never run for them. Found by the no-fake bed; blocks it and the migration's delivery
assumption.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Decisions were the one link the cycle checks skipped, and measuring found 19 of 70 records
orphaned — the credential flow and the module-runtime cluster among them, which is how a
stale premise about a settled decision survived in working memory. cycle.py now refuses an
accepted record nothing cites; the 19 got true homes (design frontmatter, the playbook that
implements 0021, META for the process records). The overview names the practice: spec-driven
development with provenance.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Findings of an adversarial review of the 055 fix: hub!=broker conflated, membership tested
as has-address, silent staleness, portless silent fallback, two hubs unrefused. One review
claim recorded as disputed against live wire measurements.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The flow the process overview draws — idea/symptom -> decision -> to-be design -> code ->
as-is — was enforced by nothing. cycle.py now refuses a to-be design naming no decision, an
in-progress/implemented design naming no owning code, a located/fixed issue with no owner,
a fixed/resolved issue with no fix, and a graduated research overview that does not say what
it became. AGENTS.md carries the cycle and a where-to-look table so a fresh session (or a
cleared context) finds the chain in frontmatter instead of assuming it. Grounding the check
surfaced two real gaps, fixed here: the work-ahead design named no owning code, and research
003 listed one became target twice.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The glossary's authority page still named the controller's seat the-controller in two
entries, contradicting its own seat section after ADR 0079; issue 058's heading kept the
pre-renumber 059; 055's fixed-by named branches that stop existing after merge (now merge
commits/PRs) and its located-in listed file paths where the convention wants repos; 056's
located-in named mesh-host, which received no fix, instead of mesh-catalog; and the design
layer never said the one-store/one-broker property is enforced — 07-the-foundation and the
installation table now state the seats.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Observed in the two-node bed: a cross-node runtime exits on 'timed out fetching the
broker's certificate' until the tunnel forms, then settles. The retry belongs in-process.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The store and broker were "one per mesh" by convention only. Each foundation module now
claims a mesh-scoped seat named after the server it guards — postgres/mesh-store,
lavinmq/mesh-broker — and the controller's seat is renamed the-controller -> mesh-controller
so all three follow one rule. The resolver refuses a second holder, closing 056. Glossary,
the foundation and installation docs, and the decisions index follow the new name.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
055 is proven end-to-end by the two-node bed: a consumer on a joined node reaches the
adopted broker over the overlay, its binding names the control-node, and its vhost is
minted. Three fixes — the broker's amqps port in the firewall, the broker credential
naming the overlay not the public address, and pushing the provider node after the remote
consumer arrives. The last is an operator ordering, not a code fix, and its silent-failure
edge (a cross-node consumer that never provisions and only crash-loops) is opened as 057.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A two-node bed pins it: the firewall gap (broker's 5671 not in listens) is fixed
in mesh-catalog; the real bug is modules.go:294/build.go:231 building the broker
URL with the genesis public MESH_BROKER_ADDRESS, which a joined node cannot reach.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
"One postgres, one lavinmq" (ADR 0078) holds only by genesis assigning the
adopted modules to the control-node alone; nothing enforces it. A second assign
raises a divergent second server, silently. Open questions: a mesh-scoped
exclusive seat, or explicit adoption the controller refuses to place elsewhere.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Settles the design repository now that the self-upgrade build is on main:
- Records the two decisions that shipped without a record — ADR 0077 (the
controller/foundation/node vocabulary) and ADR 0078 (the store and broker are
ordinary modules); accepts ADR 0075 and 0076, which shipped work rests on.
- Fills issue 051's amended-design and wires ADR 0078 into 07-the-foundation.
- Sweeps the repo rename (mesh-control -> mesh-controller) into the mutable docs
now that the forge repo is renamed; updates the glossary note and repos.md.
- Fixes the six broken links from the design-doc renames, indexes the glossary,
regenerates the decisions reading order.
Both checks (records.py, index.py) are green. Statuses stay honest: the build is
on main and lab-proven but not deployed as the production mesh, so the to-be docs
remain in-progress and the as-is layer (the hal mesh) is unchanged — graduation
to implemented + as-is belongs to deployment, not merge.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Marks WBS 3.2/3.3/3.4 done and resolves issue 051: the foundation's store and
broker are adopted in place as the postgres and lavinmq modules, upgradeable
through their stated windows, source-tracked by status. A bare-metal mesh runs
one postgres and one lavinmq, proven 22/22 in the one-node lab. The two follow-up
gaps are tracked as issues 054 and 055.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
054: the adopted servers bind 0.0.0.0 from genesis but the packet filter is
installed later, so there is a window where they are open with only bootstrap
credentials. 055: the servers bind on the control-node and the one-node bed
cannot prove a consumer on another machine can reach them over the overlay.
Both are follow-ups to issue 051's adoption (WBS Phase 3), tracked rather than
rushed.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
"control plane" -> controller and "substrate" -> foundation throughout
03-DESIGN, 00-META and the README, with 06-the-control-plane.md and
07-the-substrate.md renamed to 06-the-controller.md and 07-the-foundation.md.
The immutable 02-DECISIONS records keep their original wording (and links to
them are unchanged) — a term retired here may still appear there, which the
glossary explains how to read.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Locks the vocabulary that kept drifting in conversation — controller (not
"control plane"), foundation (not "substrate"), node and control-node, seat /
bench / claim, package vs artifact. AGENTS.md points at it as the authority.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Records the decision the package-registry work turns on — the SDK is built on a
public base and published before the toolchain that consumes it, so nothing is
circular; mesh-tools stays the thin toolchain base but resolves the SDK by
version. Reconciles docs 12/17/22 and indexes the record.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The protocol spec claimed the envelope and grant drifted across implementations.
Inspection showed the wire agrees — envelope required headers match, the two
optional ones are legitimately optional, and the grant wire (the contributions
file) is identical on both sides. The disagreement was in dead types, now removed.
So phase 1's 'make them agree' work is done by deletion and correction. What
remains is a conformance fixture as prevention — pinning the envelope and the
contributions file so a future change that breaks agreement fails a test — and a
full per-capability suite is deferred until a third language actually needs it.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The record claimed the two implementations already disagreed. Inspection showed
the live wire agrees: the disagreeing grant types were dead (removed), and the
envelope's two extra headers are optional and set when relevant, not missing.
The danger was dead types contradicting the wire, not live disagreement — which
is a sharper reason for specifying the wire and checking against it, not a weaker
one. The model stands; the conformance suite's job is prevention rather than
repairing a present break.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The correction the operator pushed: stop running a 20-minute lab against a mesh
mid-transformation, debugging paths the next phase deletes. The base-build hang
is almost certainly the SDK resolving from a git URL inside a docker build (issue
053), which Phase 2 removes — so debugging it on the current shape is debugging
deprecated code.
Phase 0 folds in: the installer's own regressions are fixed and committed;
whether it runs green is the final acceptance test, after the phases that change
its build path are in. Faults that can be reasoned out of the code path are, by
reading rather than running.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Everything decided this cycle and not yet built. Phase 0 gets the installer
green, because nothing else is testable end to end without it. Phase 1 makes the
protocol one thing and fixes the Go/TS drift the installer's own provisioning
exercises. Phase 2 stands up the private package registry ADR 0014 assumes and
publishes the SDK into it. Phase 3 adopts the substrate so one postgres and one
lavinmq serve everything, which is the hardest and needs all three above.
Order is dependency, not preference. Each phase ends at a run rather than a
paragraph, because a phase that ends at a claim is how things went missing this
cycle without anything complaining.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The lavinmq module raises its own server container and provisions vhosts on it,
separate from the substrate's mesh-broker. A mesh with the module assigned runs
two LavinMQ servers where one belongs — the exact AMQP twin of the two postgres
containers.
The module's own provisioner already assumes one server: it creates a vhost per
consumer, named for the login, isolated by the vhost boundary — the analog of
postgres's database-per-login. So the mesh's own control traffic is the / vhost
and every consumer's broker is a vhost beside it, all on one server.
Adoption therefore means the module does not run a server of its own: its server
is the substrate broker, adopted, and the module contributes the provisioner,
tools, events and bootstrap against it. The isolation model is already built;
what remains to design is only raising the one broker at genesis and then holding
it as a module, over the broker it is.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The substrate's store and the postgres module are the same thing. They were two
rows only while the substrate was a different KIND of thing — a store raised from
a bundle cannot provide postgres-database, so anything wanting a database needed
a second server. It is visible on any mesh built today: mesh-store and postgres,
two containers, the same image.
Which name survives is settled by the naming rule. Where a consumer speaks a
protocol the interface is the protocol, and "database" is not a capability. The
control plane's own queries use distinct on and on conflict, so the coupling is
to postgres and a "store" module would advertise a swap that fails the first time
anyone tries it.
The broker collapses the same way with a different outcome: amqp IS a protocol
that several implementations speak, so amqp is a legitimate provision and lavinmq
is one provider of it.
So adopting the substrate is not only an upgrade path — it is two rows of a
mesh's module list becoming one, twice. And it leaves nothing that is a specialty
after the pivot, which is the claim the whole design rests on and is not true
today for exactly those two.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Reported as 'the host does not survive a reboot', which reads as a mesh that
cannot come back. mesh-host/packaging/ ships nox-mesh-host.service and two
companions. The installer declines to place them because a unit file is a
packaging decision, and the lab starts the host with --host-in-background, which
says in its own help that it does not survive a reboot.
So the lab run failing this was the lab being honest, and the gap is the step
that puts a shipped unit on a machine — narrower and more fixable than what I
wrote.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The first draft implied gitea and the registry could not share a machine, because
registry claims the-artifact-store at node scope and I carried that across to
gitea without asking what the claim is for.
A machine running gitea for git and packages alongside a registry serving
artifacts is an ordinary arrangement. They are different ports doing different
jobs, and nothing about one being the mesh's artifact store requires the other
not to exist.
The exclusivity that matters is mesh-wide and already expressed: provides at mesh
scope means two providers are two answers, and the resolver refuses until one is
assigned. Forbidding co-residence adds nothing and forbids something reasonable.
Whether registry should still hold that claim is left open rather than answered
from outside its manifest — it may be protecting something about its port or its
data directory that nobody wrote down.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Two questions circling turned out to be one asked twice: should gitea be the
mesh's registry, and where does the SDK come from. The framing that dissolves
both is that artifact-store is already a provision and the registry already
provides it — so this was never about replacing a component. It is a second
provider of an existing provision, which this mesh has a mechanism for and uses
for certificate authorities already.
So: two provisions, because they are two jobs. artifact-store is content
addressed, pinned by digest, no versions and no ranges — what the mesh delivers
to machines. package-registry is an ecosystem's own, addressed by name and
version — what code resolves when compiled. Conflating them is how a mesh that
pins everything ends up rebuilding one commit into two different things.
The small registry stays the provider genesis installs, not because it is better
but because of what it is: a directory and one container, installable where there
is no database and no control plane. Gitea needs both, and the pivot needs
somewhere to publish before either exists.
Gitea also provides artifact-store, so a mesh may choose it — and choosing it
answers issues 042 and 048 by adopting something that already has accounts and
TLS, rather than reimplementing them in a registry that has neither.
Moving between providers is a designed act with a verification step that is easy
to skip and is the only thing between it and a mesh that cannot restart its own
control plane.
And the bootstrap still has no package registry when the first build needs one.
Named rather than solved, so the next person does not discover it.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Every step from a bare machine to a mesh that maintains itself, in three phases,
with each step named as the installer prints it.
The point of writing it out is the shape it exposes. The installer owns twelve
steps and ends at a mesh that RUNS. Seven more turn that into a mesh that WORKS —
the shared base, a store that is a provider rather than the control plane's own
memory, the catalogue, the replay of what was built before the catalogue existed,
the control plane rebuilt through the module path, the private network with the
node actually placed on it, and the packet filter. None of those seven is the
installer's. They are things somebody types, which is why a test had to be
written to discover they were missing.
Machines arrive last, in phase three, because a machine joining a mesh that
cannot build anything proves enrolment works and nothing else.
And five things that are not yet true are named rather than implied: phase two is
manual, a second machine cannot pull what the mesh built, ADR 0014 assumes a
private package registry that genesis has not installed when the first build
needs it, the host agent does not survive a reboot, and nothing can contradict a
claim that a machine was installed this way.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
ADR 0014 is accepted and unambiguous: each module consumes its dependencies from
the private registry, the mesh's own shared library included, and a cross-package
change is publish then consume.
So the git dependency at a pinned commit is not a mechanism under consideration.
It is the shared library being consumed a way the record rules out, and the lock
naming a sibling directory is what that looks like when nobody publishes. Issue
053 is reclassified from a question about mechanism to a violation with a
direction.
What stays open is narrower and real: which software serves the private registry,
and that a fresh mesh has none when the first SDK is built — ADR 0014 assumes one
exists, and at genesis nothing has installed it.
Recorded after arguing at length for a bespoke content-addressed alternative,
against a decision that was already made and that I had not read.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx