Commit Graph
232 Commits
Author SHA1 Message Date
jschoubben 03fc13b84b 051: one lavinmq, not two — the duplication is running, not latent
The lavinmq module raises its own server container and provisions vhosts on it,
separate from the substrate's mesh-broker. A mesh with the module assigned runs
two LavinMQ servers where one belongs — the exact AMQP twin of the two postgres
containers.

The module's own provisioner already assumes one server: it creates a vhost per
consumer, named for the login, isolated by the vhost boundary — the analog of
postgres's database-per-login. So the mesh's own control traffic is the / vhost
and every consumer's broker is a vhost beside it, all on one server.

Adoption therefore means the module does not run a server of its own: its server
is the substrate broker, adopted, and the module contributes the provisioner,
tools, events and bootstrap against it. The isolation model is already built;
what remains to design is only raising the one broker at genesis and then holding
it as a module, over the broker it is.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 22:31:17 +02:00
jschoubben 87f464cc5f A finished mesh holds twelve, and two rows were one module each
The substrate's store and the postgres module are the same thing. They were two
rows only while the substrate was a different KIND of thing — a store raised from
a bundle cannot provide postgres-database, so anything wanting a database needed
a second server. It is visible on any mesh built today: mesh-store and postgres,
two containers, the same image.

Which name survives is settled by the naming rule. Where a consumer speaks a
protocol the interface is the protocol, and "database" is not a capability. The
control plane's own queries use distinct on and on conflict, so the coupling is
to postgres and a "store" module would advertise a swap that fails the first time
anyone tries it.

The broker collapses the same way with a different outcome: amqp IS a protocol
that several implementations speak, so amqp is a legitimate provision and lavinmq
is one provider of it.

So adopting the substrate is not only an upgrade path — it is two rows of a
mesh's module list becoming one, twice. And it leaves nothing that is a specialty
after the pivot, which is the claim the whole design rests on and is not true
today for exactly those two.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 20:26:42 +02:00
jschoubben 27b2161379 The SDK's delivery is decided; the git URL violates it
ADR 0014 is accepted and unambiguous: each module consumes its dependencies from
the private registry, the mesh's own shared library included, and a cross-package
change is publish then consume.

So the git dependency at a pinned commit is not a mechanism under consideration.
It is the shared library being consumed a way the record rules out, and the lock
naming a sibling directory is what that looks like when nobody publishes. Issue
053 is reclassified from a question about mechanism to a violation with a
direction.

What stays open is narrower and real: which software serves the private registry,
and that a fresh mesh has none when the first SDK is built — ADR 0014 assumes one
exists, and at genesis nothing has installed it.

Recorded after arguing at length for a bespoke content-addressed alternative,
against a decision that was already made and that I had not read.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 16:50:40 +02:00
jschoubben 033a5d384c Issue 053 — the SDK is pinned twice and the two disagree, and ADR 0074 accepted
The tool runtime's manifest names the SDK as a git dependency at a pinned commit;
its lock file names a sibling directory that exists on one workstation. It builds
only because the recipe runs npm install, which tolerates a lock disagreeing with
its manifest and re-resolves from the manifest — the one command that hides this.

It matters because a lock exists to make a build reproducible and this one
describes one machine, and because it is the first thing a fresh mesh builds: the
toolchain carries the SDK and everything with code of its own compiles inside it,
so a dependency resolved differently on the build machine than on a workstation
is a difference in every module the mesh will ever build.

And it is about to be copied. Each language's toolchain will carry that
language's SDK the same way, so the shape is worth settling before there are four
of them.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 13:44:02 +02:00
jschoubben b163ed1fcc Issue 052 — the firewall closes the port the mesh runs on
The packet filter generates its rules from what modules declare they listen on.
The broker is not a module, so it declares nothing, so its port is not opened.
Every machine dials that port to enrol and to receive every declaration it is
ever sent.

Invisible where it is assembled and fatal on the next machine: a mesh of one
never dials its own broker across the network, so the ruleset looks right. The
first machine to join is refused at the packet filter during enrolment, several
steps from anything that reports it — and assigning the firewall before joining
machines is both the natural order and the one that breaks.

This is 051 in a second place. That issue says the substrate cannot be updated
because the mesh holds no record of it; the same absence means the firewall
cannot know it exists. Ssh is already a floor for the same reason — a machine
nobody can reach is a machine nobody can repair — and the broker may be the same
class of fact.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:59:33 +02:00
jschoubben f52895a646 Issue 051 — the mesh can update everything except what it depends on
The store and broker come from a bundle the installer writes once, with images
pinned in it, and nothing can change them afterwards: no build, no version to be
behind, no roll-out, and no way to report being out of date, because the mesh
holds no record of them as modules at all.

That is backwards. They are what everything else depends on, so their updates
matter most, and they are the only things with no mechanism to deliver one. A
mesh with a year-old broker reports itself entirely current.

The fix probably already exists: the control plane is carried, raised and then
adopted as an ordinary module pinned to what is running. Nothing in that pattern
is specific to the control plane. It would also remove a duplication visible on
any one-node mesh — the same postgres image running twice, because a store that
cannot be a module cannot provide a database to anything.

Recorded with the two hard parts stated rather than waved at: upgrading a store
the control plane is reading from, and upgrading a broker over the broker.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:44:10 +02:00
jschoubben 5d13f9c83d Issue 050 — the catalogue knows nothing built before it started
A mesh raised from bare metal built six modules and its catalogue reported three:
exactly those built after it began running. Missing were the shared base, the
store, and the catalogue itself.

The hole is never random. On a fresh mesh the modules built before the catalogue
are by necessity the ones it needed in order to exist, so the foundation is
always what is absent, on every mesh, at the moment the graph is first populated.

It breaks the question the catalogue is for: build edges hang off the base, so a
catalogue with no record of it answers 'what must be rebuilt' confidently and
wrongly. And nothing reports the gap, because a catalogue cannot know what it was
never told — this was found by comparing its answer against what the mesh had
just been watched doing.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:16:23 +02:00
jschoubben 0aa8f62e0c Issue 049 — a module serves tools and nothing may call them
A module's broker account is scoped to what it declares it emits and consumes. A
tool call needs a reply queue, which that scope does not cover and should not. So
the account is right, the request is reasonable, and no account exists that can
make it — asking a module its own question, from its own container, with its own
credential, is refused.

It matters because a module's tools are its operator-facing surface: the
catalogue serves the five questions it exists to answer and nothing can reach
them. It is also why a running module keeps being mistaken for a working one — a
test that cannot ask anything checks a container is up, and that substitution has
hidden two faults this week.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:08:40 +02:00
jschoubben ee70d5b451 Issue 048 — a stated rule about the registry is enforced by nothing
The registry says every machine pulls from it and opens its port to the mesh for
that reason. A machine that tries is refused by its own container runtime: the
registry serves plain HTTP and anything but loopback is treated as HTTPS.

It has never failed, and that is the finding. Every proof that a machine can
fetch a mesh-built artifact was a proof about the machine that built it, where
the reference was loopback. The bed that uses a routable address gets away with
it because the harness writes the runtime's configuration before the mesh exists.

Kept separate from 042, which they are easy to confuse: that one is about not
being allowed to pull, this one about not being able to whatever the credential
says. Fixing 042 alone leaves a machine with a valid account for a registry its
runtime will not talk to.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 21:01:32 +02:00
jschoubben 261064ec63 Issue 047 — and the half that runs the other way
The report covered ports the firewall cannot close. The matching fault is ports
it does close and should not: rules come from what modules declare, a migrating
mesh knows about almost nothing, and anything listening on the host is dropped.

No module declares an ssh port and the generator has no allowance for one. The
session that loads the rules survives on conntrack until it drops, and then the
machine is reached from a rescue console.
2026-09-14 15:09:20 +02:00
jschoubben 027c41d125 Issue 047 — the firewall leaves published container ports open
Rehearsed on three lab machines before doing it on an anchor: loading the mesh's
rules refuses an ordinary host port and leaves a published container port
reachable, same prober, same second.

It is deliberate — no forward chain, because dropping there would stop every
container — but traffic to a published port never reaches the input chain, so
the firewall is silent about it. The anchor publishes 38 such ports today, all
filtered by the system being replaced, so the cutover would open every one of
them while reporting a firewall that is up.
2026-09-14 14:59:11 +02:00
jschoubben 24d52a8380 Issue 046 — an upstream image cannot be mirrored into the mesh's registry
Found while migrating the first module. Five approaches were tried and observed
to fail, including resolving the index and naming the platform at both ends; a
speculative fix was written and reverted rather than shipped, because it did not
make the mirror work.

One module uses this today, which is why it went unnoticed — and it is the shape
most of a migration wants, because the services being moved are third-party.
2026-09-14 13:27:42 +02:00
jschoubben 2eb65a17a3 A module names its base, so the mesh can act on what it notices
The mesh could already say which modules a base change invalidated, and could
not do anything about it: each recipe named one particular copy of the base by
fingerprint, and rebuilding produced the old one. Worse, the copy each named
existed only inside a throwaway lab, so those three modules could not be built
anywhere at all — and the line each replaced had the same fault.
2026-09-13 23:58:06 +02:00
jschoubben 3983188b8a Issue 044 resolved — the mesh builds its own floor
Cloned alone, built by the mesh, every module moved onto it, and the base then
changed for real: all three went stale naming what moved. A comment-only change
correctly makes nothing stale, which found a second bug — staleness compared
commits where it should compare artifacts.
2026-09-13 02:51:08 +02:00
jschoubben 15ac7fd8dc The base has no circularity — the builder does not stand on it
Written as an open question; it has an answer, and leaving it open would have
made the fix look harder than it is.
2026-09-13 02:25:07 +02:00
jschoubben 3a9ed9d3fd Two findings from building the catalogue on a live mesh
The runtime every module compiles against cannot be built by the mesh, so the
one rule that would catch it moving can never fire. And a container keeps the
values it was created with, so two good applies can leave it running on neither.
2026-09-13 01:49:02 +02:00
jschoubben 837b5df2f7 Withdraw 043: the capability existed and the wrong verb was used
A build machine was refused the build queue, and this was raised as a gap in what
a manifest can express. It is not: `builder issue` creates exactly that account,
three lines from the code being read at the time.

Kept rather than deleted, for the one real thing in it — the wrong verb succeeds
and reports success, producing an account that authenticates and can do nothing,
so the failure surfaces a layer away as a permissions error that reads like a
missing feature.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 16:51:37 +02:00
jschoubben 26bec28c60 Issue 042 — nothing gives a node an account for a registry
Found by deleting the lab's registry and giving the machines a real path out:
public images fetched, the operator's own could not be fetched at all. There is
no provision for a registry credential, no manifest field, and no step in
enrolment that establishes one.

It applies to the mesh's own store too, which today asks for nothing — a
decision that has never been written down as one, and so reads as an absence.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 01:36:00 +02:00
jschoubben c29accd1be Issue 039 resolved — the operator's images are pinned, and why that is a stopgap
Nine references now name the digest their tag resolved to. What closed is the
immediate fault; the open questions stand, because a digest in a repository is
wrong the moment anybody rebuilds — which is the reason the design wants the
repository to name artifacts and the mesh to hold digests.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 01:06:35 +02:00
jschoubben 423fde8534 Place the two new records in the reading order, and fix a link the merge moved
Both carried a topic outside the six the index knows, so neither had a place to
be read in — 0067 had none at all. Both are 'the tiers', beside 0036 (bootstrap
ends at a usable mesh) and 0007 (connectivity), which is what they extend.

0067 cited 0041 for tier 0's property; on this trunk 0041 is events, and the
record it meant is 0005. A citation that resolves to the wrong record reads as
corroboration, which is worse than a dead link.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:48:20 +02:00
jschoubben 84c95c53af Merge branch 'issue/021-provider-port-published-on-loopback' into design/bootstrap-is-a-pivot
# Conflicts:
#	03-DESIGN/01-to-be/04-lab-installation.md
2026-09-11 00:45:59 +02:00
jschoubben f9e129f734 Issue 040 — the only description of how a mesh is stood up is a test
There is no installer. The complete account of standing up a mesh is an
integration test in the lab, and a fixture may invent what it needs — this one
raised a registry no production has and rewrote every image reference through
it, concealing both 039 and the fact that a first node outside the lab had no
bootstrap path at all. The bed being green said nothing about whether a mesh
could be installed.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:34:59 +02:00
jschoubben e228355a52 Issue 039 — the lab's registry was pinning what the catalogue left unpinned
Nine images across seven modules name a tag, not a digest. ADR 0006 forbids it
and the host refuses it by name — and the refusal has never fired in a bed,
because the lab pushed every image into its own registry and rewrote every
reference to the digest it had just assigned. The harness was supplying the
property under test. Found by deleting the harness.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:20:10 +02:00
jschoubben a4440c9acf Issue 038 resolved — corrected root cause (same-node served-port mismatch, not loopback) + fix reference
The real defect is same-node consumers announced the declared port instead of
the assigned/published one; the loopback observation was a stale pre-0038 build.
Fixed in mesh-control fix/same-node-provider-announced-port (c147a26) with a
regression test.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 23:06:58 +02:00
jschoubben 9ffe9e97c9 Renumber to issue 038 — 021 is already taken (same-node credentials, closed) and issues run to 037 across branches
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 22:46:11 +02:00
jschoubben d16b5294f6 Issue 021: narrow to mesh-assigned ports (bare decl binds loopback, explicit host mapping binds 0.0.0.0)
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 22:35:06 +02:00
jschoubben 315bd7118a Issue 021 — a provider is announced at a name its port is not bound to
A from:mesh provider is announced (per 018's fix) at the node's private-network
name, but its port is published bound to loopback, so consumers dialing the
announced <node>.internal:port reach nothing. Diagnosed from the lab: DB
consumers that require the database at startup crash-loop; the provider is
healthy on loopback only.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 22:29:04 +02:00
jschoubben 2405d72fb0 ADR 0052 (proposed) — an init step is a container run once to completion
A module can declare state but not a step that runs. mosquitto must seed its
dynsec admin into dynamic-security.json before the broker starts, or the plugin
aborts; the database providers need the same for first-boot migrations and
health gates (04-ISSUES/037). The old event-hook engine that did this was
powerful and flaky; this is the narrowest sound mechanism instead.

A run-once step is an ordinary container marked `run-once: true`: the host runs
it to completion, requires exit 0, and gates the apply on it — so what the
declaration places after it (the broker) starts only once it has finished.
Gating is by declaration order, not a resolved dependency (ADR 0005); the
completion marker is the recorded declaration digest (ADR 0018), so a re-apply
does not re-run it unless the declaration changed. No new host shape and no
arbitrary host command: strictly less powerful than an `action`.

Points 04-ISSUES/037 fixed-by/amended-design at the record; index regenerated;
records.py and index.py pass.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 23:52:26 +02:00
jschoubben 220e3c72f5 ADR 0051 (proposed) — shared data is the operator's
Resolves 04-ISSUES/036: eight media modules each declared the shared
library and download directories as their own resources, and the
resolver's duplicate-owner refusal — right in general — would refuse the
stack's only sensible assignment the first time two landed on one node.

The decision, from the operator: shared, pre-existing data is
operator-owned and external. The mesh does not create, chown, reconcile
or remove it. A module declares it needs access to such a path (read or
read-write); the host mounts it and owns nothing. Several modules
accessing one path is normal — the duplicate-path refusal is about
ownership, not use. An accessed path absent at apply is refused clearly,
not created. Extends ADR 0030: the third case the host had no word for,
what it neither made nor configured and must not touch.

Point 036's fixed-by/amended-design at the record; mark it located in
mesh-control, mesh-catalog and mesh-host. Regenerate the decision index.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 22:22:30 +02:00
jschoubben 495db6dc89 Restore initialization's issue 003 — the status flip was based on a wrong premise
Init's 003 was already resolved with its own consolidated attribution
(03-DESIGN/01-to-be/08-connectivity.md); the re-homing overwrote it with the
session's ADR-0045 firewall attribution. Init is canonical and the firewall
decision is recorded in the ported ADR 0045 regardless, so 003 is restored
untouched. No initialization record is modified by this reconciliation.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 12:26:23 +02:00
jschoubben e269f9a185 Re-home this session's new ADRs (0039-0049) and issues (032-037) onto the consolidated scheme; flip issue 003; port repos.md sdk line + feature-branches playbook (07); regenerate index
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 12:24:07 +02:00
jschoubben 9638aa02c5 031 — a machine becomes each thing it was told, in turn
Declarations queue and the host applies all of them, oldest first, so a
machine pushed five things in a minute spends five applies becoming the
last one. Correct every step — each declaration is the whole machine —
and wasted in all but the final step.

Only visible since a report names its declaration: the reports arriving
were about ever-older ones, while timestamp comparisons used to happen
to pass. The fix is consumption order, not the queue; the open question
is what a superseded declaration's report should say, because silence
reads as disobedience and "applied" would be a lie.
2026-09-02 01:06:12 +02:00
jschoubben 6bdd3f2bfa 030 — asking what a machine should be re-signed its certificate
Composing a declaration signed the machine's certificate anew each
time, and a signing carries a fresh random serial — so what the mesh
would send differed from what it had sent by one byte, for ever, and
every machine carrying a certificate stood eternally waiting.

Third find of the same rule: issued once and kept. The port had it, the
secret had it, the certificate composed fresh on every asking — and the
keeping column had existed since the serving-key migration, written by
nothing, the same shape ReleasePorts was found in.

Found by keeping the scenario standing and diffing two plans seconds
apart: one line, where four theories had none.
2026-09-01 22:28:26 +02:00
jschoubben 6385d52bc0 029 fixed — the store is added, not built, and the cycle is refused
The registry module names its image by digest, the way the bundle names
the three a first node starts from, and goes in as a manifest. The lab
now walks that path and passes.

A manifest that provides the store and also builds artifacts is refused
at parse. The lab had been passing only because its Docker Hub stand-in
quietly received the push — a prop covering for the thing under test.
2026-09-01 21:43:42 +02:00
jschoubben 013dbce4a9 029 — the artifact store cannot be delivered by the artifact store
Installing the module that provides `artifact-store` requires something
that provides `artifact-store`: its image is mirrored in, mirroring
publishes to the store, and the builder refuses to run without one.

Never seen, because the lab always has a registry standing before the
mesh asks for one, and so does any mesh built on a machine that already
had one.

What it blocks is larger than a registry. The store holds images and
packed archives both — it is the module catalogue in artefact form — so
until it exists a mesh can run only what its bundle already carries.

The substrate record already answers it. It asks of each candidate
whether it can grant itself the thing it provides: the store cannot
create its own database, the broker cannot create its own virtual host,
and the registry cannot grant itself a repository. So the registry
module names its image and is never built.

With a limit worth saying out loud rather than discovering: a module
providing the store may not build artifacts of its own, a UI or a tool
server included, because there is nowhere to put them until it runs.
Such a registry is two modules.
2026-09-01 21:15:05 +02:00
jschoubben e75683c01f 026 reopened — the rule is right about data, wrong about facilities
Enforcing "declare what you mount" refuses the builder, which mounts the
container runtime's socket. That socket is not the builder's data: it
exists already, the machine owns it, and declaring it as one of the
module's directories would be a lie the host would act on.

Two kinds of mount are spelled identically today — the directory my data
lives in, and a machine facility I was granted. Until a manifest can say
which, the fourteen declared mounts are right by coincidence, which is
what this issue was opened about.

Recorded rather than decided: separating them is new vocabulary, and
inventing it to turn a check green is how a mechanism nobody chose ends
up load-bearing.
2026-09-01 19:45:55 +02:00
jschoubben bd8f09d647 026 fixed — and the coincidence turned into a rule
The fourteen undeclared mounts are declared. More to the point, a
manifest that does not declare one is now refused: they were right by
coincidence, and a checklist nothing enforces is a checklist that is
true until the next commit.

Refused in the control plane, because the machine cannot tell the
difference — asked to mount a path that does not exist, it makes the
directory, which is a thing it is perfectly able to do.
2026-09-01 19:36:15 +02:00
jschoubben 0bcfeb4e80 028 fixed — the mesh assigns the port, and knows what it cannot move
A module now says its port once, in `listens`, and the container's
mapping, the rule set and what a consumer is told are all derived from
one assignment. The three hand-written copies that agreed only because
one person wrote them are gone.

The half that made this an issue rather than an inconvenience was that
the substrate is not a module: nothing in the mesh had heard of its own
store, so it handed a database module the port the store already had.
The machine now says what it carries, and the mesh assigns around it.

Ports the protocol fixes became claims, which needed no new mechanism —
the mesh already had one for what is singular on a machine.

Left open, and unchanged by any of this: whether a module should publish
to the machine at all. Assignment makes publishing safe without making
it necessary.
2026-09-01 18:32:58 +02:00
jschoubben c2a37ab7f8 38 — the mesh assigns the port, and a module does not care
Jochen's call, and the right one: a module cannot choose a port well,
because it is written once and assigned anywhere. Any number it picks is
a guess about a machine it has never seen, and two modules guessing the
same number is not a mistake either of them made.

Writing it up turned up something the issue had missed. The same number
appears three times in every module — the rule set, what a consumer is
told, and what the runtime publishes — and nothing checks that they
agree. They agree today because one person wrote all three. A module
whose `serves` said one thing and whose container published another
would resolve, compose, apply, and hand every consumer a port that
answers nothing.

So the decision is one source with the other two derived, and an
assignment made once and kept, as a credential is.

The part that needed thought is ports that cannot move — mail on 25,
submission on 587. Those become claims, which is what the mesh already
has for what is singular on a machine. Two modules wanting 25 is the
same shape as two wanting the seat, and gets refused by name at
assignment rather than by a container runtime at apply. That makes this
mostly a matter of pointing an existing mechanism at ports.

Left open: whether a module should publish to the machine at all.
Assignment makes publishing safe without making it necessary.
2026-09-01 17:42:27 +02:00
jschoubben 6330abce5f 028 — two things want one port, and nothing says so until the machine
Found by fixing 027 and pushing again. The declaration is now accepted
and the database container still cannot start: the mesh's own store
holds 5432 on that machine, and the module publishes 5432.

Nothing catches it because the substrate is not a module. It arrives
from the bundle before there is a mesh to ask, so the control plane has
never heard of the store and does not know it holds a port. Resolution
can compare modules with each other and cannot compare one against what
the mesh is built on.

Nor does it compare modules with each other. A port is exclusive on a
machine in exactly the way a claim is, and the mesh has a mechanism for
that which ports do not use.

It has been met before: the end-to-end test that exercises a real
database publishes 5433 rather than 5432, inline, with nothing saying
why. That is how a constraint becomes folklore.

The open question is bigger than the bug. Whether a module should
publish to the machine at all decides how a consumer reaches it, and
changes what `serves` means.
2026-09-01 17:21:31 +02:00
jschoubben 03266fd4a2 027 — a container cannot follow a file, and a rotated credential is the case
Found by the forge failing to start. I had put `restart-on` on nine
containers so they would pick up a rotated credential; it belongs to a
service, and the host refused the whole declaration.

Removing it fixes the modules and leaves the reason I reached for it.
The mechanism is written against exactly this, in the host's own words:
a running service does not re-read its configuration, so replace the
file, find it already running, do nothing, and the machine keeps
behaving as before while every check passes. Every word of that applies
to a container, and nearly everything the mesh runs is one.

The cost is concrete. Rotation replaces the file and tells the provider
to accept the new credential. A provider reconciles, so it takes it. A
consumer is usually a container, so it does not — and the two ends hold
different passwords, which is the fault ADR 0001 records costing two
days. The test that proves rotation works uses a consumer that reads the
file on each attempt, so it does not meet this.

Two things a fix has to keep: it stays declared state rather than a
command, because the link may not carry an action; and where an env-file
changed, the honest verb is recreate rather than restart, because a
container's environment is fixed at creation.
2026-09-01 17:19:43 +02:00
jschoubben 3f00d3b413 026 filed; 025 corrected; the image store is a module
Three corrections, two of them to things I wrote today.

026 is the serious one. Four modules mounted fourteen host paths nothing
declared — the mail spool, the databases, the object store's data. The
runtime creates those as root, so owner and mode go unapplied, and the
rule that keeps a directory holding data the mesh did not put there is
written in terms of declared directories. It reached the configuration
and missed the data. The cause was carrying compose files across: a
container shape that can express one gets filled in like one.

025 claimed nothing turns a tag into a digest. That is false, and the
answer was designed and built before I wrote it. A module names an
artifact, not an image, and `kind: upstream` mirrors somebody else's
image into the mesh's own registry, pinned by the digest it lands with.
The two-document split the issue described as the shape of a fix is the
design. Pinning twelve images by hand was treating the symptom, and left
them pointing at a public registry rather than the mesh's.

And the image store was written up as something the mesh does. It is an
ordinary module — considered for the substrate and removed, because the
test is whether the control plane needs it before its first instruction,
not whether it can grant itself one. So somebody's own registry is the
same module as the mesh's.
2026-09-01 16:10:13 +02:00
jschoubben 2bbdc52440 025 — half done: nothing unrunnable reaches a machine now
The refusal landed, and the examples pin images that exist. What is
still missing is the part that makes it unnecessary: nothing in the mesh
turns a tag into a digest, so it was done by hand — which is precisely
what the issue says a person should not be asked to do.

Recorded because the resolution mechanism turned out to be trivial:
asking a registry what a tag points at takes about a second and pulls
nothing. That removes the main argument for leaving this open.

Also records the two faults that fell out of pinning for real. The mail
system named seven repositories that do not exist, because it publishes
to a different registry than the manifest assumed, and one of the seven
had been renamed upstream. Nothing checking only the shape of a
reference could have found either.
2026-09-01 15:13:52 +02:00
jschoubben fc7717f6b1 021 closed — the record was left open after the fix landed
The code and its tests went in hours before the record was touched, so
an issue that read `located` had been fixed all along. That is the exact
failure the frontmatter exists to prevent: status is meant to be
answerable from the record rather than by reading the code.

Closed with the commit that did it, and cross-referenced to 022 and 023,
which came out of the same mistaken instinct — treating the machine as
a boundary, then as an identity, then finding a consumer had a password
and no name to present with it.
2026-09-01 15:04:36 +02:00
jschoubben f6b1834ea9 024 fixed — the registry was addressed by hand and nothing else was
The cause was one line. Machines get a systemd-networkd unit with a
static address, so networkd finishes and reports the link configured.
The registry ran `ip addr add` inline, which leaves networkd waiting to
configure something it was never told about — and
systemd-networkd-wait-online has an infinite timeout.

So network-online.target was never reached and everything ordered after
it never started. On these machines that is Docker, so `docker load`
blocked on a socket whose daemon was queued behind a target that would
never come, and three bounded timeouts stacked to thirty-five minutes.
These machines have no DHCP by design, so that wait was never going to
end.

The hypothesis in this record was wrong and the record now says so.
Stocking had just been changed, so stocking looked guilty; stocking
takes 34 seconds and always did, timed directly before changing
anything.

Fixed with two things that made it cost hours instead of minutes: an
image placement now waits for the runtime and refuses after 120s naming
what systemd is waiting on, and the end-to-end test passes onProgress —
the raise reported every step and the test discarded it, which is why
thirty-five minutes and four minutes of silence looked the same.

The suite then ran to completion, 23 of 24, the one failure a check of
its own flagging a path as a credential because `/` is in the base64
alphabet.

Also recorded: a redirected log lags, because Node block-buffers stdout
to a file. Read as a stall twice, the second time right after the real
fix — where a buffering artifact argues the fix did not work.
2026-09-01 09:53:08 +02:00
jschoubben 407416e6d0 024 — a run stalls before the host is placed, and says nothing while it does
Seen twice today. Once mid-run: thirteen passes, then the process ended
with no summary, no failure and no receipt. Once from the start: the
first test ran 35 minutes against a measured 4.5 and was still running
when it was stopped.

Ruled out rather than assumed: not memory (84 GiB free, no OOM), not the
daemon (the stalled machine answered `incus exec` immediately), and not
the changes under test — the anchor VM had no host log and no
containers, so the run never reached placing the host.

What changed just before is that the rebuild went from two artifacts to
six, and every one of them is pushed into the scenario's registry, which
is the step the second stall sat in. Recorded as what changed, not as
the diagnosis.

The reason this is an issue and not a slow test: the suite prints
nothing between starting a scenario and finishing its first test, so
four minutes and thirty-five look identical from outside, and the only
recourse is to guess. That is how a workstation was left unbootable in
August. And a run that ends silently after thirteen passes is a run
somebody may believe.
2026-09-01 03:20:40 +02:00
jschoubben f583502fc9 023 fixed — the mesh says who a consumer is, and what it is bound to
Both halves had one cause: the mesh knew something and did not say it.

Who a consumer is now comes from one derivation, sent to the provider in
its grant and to the consumer in its binding, so the two agree by
construction. The provisioners use the name they are given and refuse to
invent one, because a name of their own would create a login the
consumer could never guess while everything reported success.

Bound values reach the file that needs them through the symmetric twin
of the sealed placeholder — simpler, because they are not secret, so the
control plane fills them in and the host gains nothing.

The lab run meant to prove this failed in a way that looked like the fix
being wrong: rotation could not authenticate against a real database.
The cause was the suite rebuilding the control plane's image and not the
provisioner's, so an image built that minute ran against a provisioner
built the day before. That is 005's family and is recorded with the
issue, because the misleading part is worth more than the fix.
2026-09-01 03:05:43 +02:00
jschoubben 80f18caf03 022 fixed; 023 filed — a password is not a connection
022 turned out to have a silent half worth recording: the provider
refuses loudly and names the modules, which reads as a decision, while
the consuming node does not refuse at all. Three modules wanting one
database produce one need, so two of them get no credential file and
each starts and fails to authenticate with nothing saying why.

023 is what remained after fixing it. A consumer now gets its own
password, in whatever shape its configuration wants, and still cannot
connect: the user name is invented by the provisioner and recorded
nowhere in the mesh, and the host and port sit in a JSON binding that an
application reading KEY=value cannot use.

The asymmetry is backwards and the coverage document now says so. The
secret is the hard case, because the mesh must not be able to read it,
and the secret is the part that arrives. The host and port are ordinary
facts the mesh holds in the clear, and they are the ones stuck.

Keycloak, Gitea, Mailu and MinIO all parse and resolve and none of them
can start. This is what stands between the module set and a running one.
2026-09-01 02:40:40 +02:00
jschoubben c86adbe3cc 022 — a credential belongs to a node, so a second consumer refuses
Found while checking whether the module vocabulary covers real use
cases. A node running three modules that all want a database cannot be
planned at all:

  anchor has 3 modules asking for "postgres-database" and they would
  share one credential: gitea, keycloak, umami

The refusal is right about what it says and wrong about what it implies.
They would share one credential, and sharing is worse than refusing —
but the arrangement being refused is the ordinary one, and the node this
mesh exists to take over runs eight modules against one database server.

The cause is the key: a credential is keyed by provision, consumer node
and provider node, so `consumer` is a machine. The provisioner inherits
it and names the role `mesh_<node>`. The refusal is not a check that
caught something; it is the only honest thing that function can do with
a key that cannot tell two consumers apart.

It is the same mistake as 021 with a different face. There the machine
was treated as a trust boundary; here it is treated as an identity, as
though "who is asking" is answered by naming a host. Two modules on one
node are as separate as two on different nodes.

Worth stating plainly: without the refusal, gitea's login would have
opened keycloak's database, and nothing would have said so — from the
provisioner's side it created exactly what it was asked to create.

Not a local fix. It crosses the control plane, the grant file naming and
every provisioner that names something after a consumer.
2026-09-01 02:28:39 +02:00
jschoubben 36d342f176 What a module must be able to say, measured against 127 that exist
Every manifest in the system being replaced was read and every key
counted, then set against what the new one can express. Three findings
worth more than the table.

**The most-used key was already covered and I expected a gap.**
Depending on another module — 65 manifests, the commonest thing any of
them says — is a requirement naming a module, which already means that
module rather than anything providing the name.

**The largest real gap is tool servers: 56 modules, over half.** A
module can already run one; what is missing is anything saying it offers
tools. That is plausibly a provision rather than new vocabulary, which
would need nothing added — not yet decided, and recorded as undecided.

**The gap most worth closing is health, at seven modules.** The mesh
knows a container is running, which is not whether it answers, and this
project has paid for that distinction twice. An action with a verify is
exactly the right shape and may not arrive over the link, so a module
cannot declare one.

Two things are missing deliberately and say so: stage hooks, because the
link may not carry an action and a module needing setup ships a program;
and flavours, retired in favour of claims.

Config merging is missing and should stay missing. A mechanism that
understands TOML gets asked for YAML, then INI, which is how the thing
being replaced became unholdable.

Also records what the survey found that is not about coverage: manifests
that had stopped matching what was actually brokered, one fact derived
in two places giving two answers, and a live listing returning
credentials in plaintext.
2026-09-01 02:20:34 +02:00