Commit Graph
33 Commits
Author SHA1 Message Date
jschoubben 76fcbda9a7 route-proxy: an ACME account belongs to the authority that issued it
autocert keeps its account key at one fixed name, `acme_account+key`, in whatever
directory it is given, and reuses it for ever. That is right while the authority
stays the same and silently wrong the moment it does not. Re-initialising the
internal CA makes a new authority with a new root: it has never heard of the
account in the cache, rejects every use of it, and autocert has no path back from
that. Nothing is re-registered, no order ever reaches the CA, and issuance stops
with nothing saying why — until somebody guesses that deleting the cache directory
by hand is the answer.

Name the directory after the authority instead of sharing one between all of them:
a digest of the ACME directory URL and the root this proxy was told to verify it
with. A re-initialised CA has a new root, the mesh delivers it as a changed bundle,
the proxy restarts on that file and lands in a directory with no account in it, so
autocert registers afresh and orders again. The healing is that "is this account
still valid" never has to be asked — an account is only ever found where it is
still valid, which needs no error codes, no probe at startup and no network call
that can itself fail.

It closes a latent one of the same shape: pointing ACME_DIRECTORY at production
after testing against staging reused the staging account, because the cache had no
idea the two were different.

Trailing whitespace around the delivered root is not a new authority — the mesh
writes that file, and a newline coming or going must not throw away an account.
Old directories are left on disk, unused: they hold the only copy of certificates
that may still be valid, and this program is not the thing that should decide a
certificate is finished with.

A proxy already holding certificates orders them once more on the first start after
this, because its account moves. Free against the lab's own CA and against staging;
one issuance per name against a public authority.

novox/hq ADR 0056

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 21:02:00 +02:00
jschoubben 98f5d34610 route-proxy: an empty ACME_CA_BUNDLE means the system trust store
Piece A of ADR 0056 (selectable issuer). A provider that serves an empty root —
public-acme, whose root already ships in the OS trust store — leaves route-proxy's
CA bundle file existing but empty, because the mesh writes it unconditionally from
${bound:acme-ca:root}. Read that as "trust the system roots", the same as an unset
bundle, instead of failing with "holds no certificate this can trust". A bundle
that holds bytes but no parseable certificate is still refused.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 00:22:06 +02:00
jschoubben 68a9235792 The mount gate knows a facility from a directory
Portainer mounts the container runtime's socket, and the catalogue's
mount gate refused it — rightly by its own lights, since nothing in the
manifest distinguishes a machine facility from the module's data. That
distinction is 04-ISSUES/026's open question, so the gate now carries
the one facility the catalogue mounts as a named exception beside the
citation, one line per facility, never a pattern.

Also the confession: the previous commit landed with this gate red,
because a pipeline's tail swallowed go test's exit code. The gate was
right and the process around it briefly was not.
2026-09-02 02:09:10 +02:00
jschoubben d278edabe0 Three more: nodered, icecast, portainer
Node-RED and Icecast are the plain shapes — a data directory owned by
the number inside, generated passwords in a host-written env file, one
declared port each.

Portainer mounts the container runtime's socket, which is the mount
04-ISSUES/026 is reopened about: a machine facility, not the module's
data, spelled today exactly like a data directory. It is converted as
it runs now rather than held hostage to that vocabulary — the same
mount the builder already carries — and it will be the second citation
when 026 gets its answer.

Letta stays unconverted for now: the arrangement being replaced pins a
year-old image of a fast-moving project, and converting a pin nobody
would keep is not fidelity. It wants a fresh look at what version to
run, which is a decision and not a translation.

All pinned by real digests, resolved on this workstation today.
2026-09-02 02:08:33 +02:00
jschoubben 1c4e10e0d8 The cache keeps no ACL file, and the watch keeps the cache
The lab's diagnostics said it in one line: AUTH called without any
password configured for the default user. With an aclfile configured,
redis takes the default user from the file and quietly ignores
requirepass — so the empty seed this module shipped left the store
without any password at all, politely refusing the credential the mesh
had sealed for it.

And the file could never have worked here anyway: it was host-declared
content, which the host reconciles, so every re-apply would have wiped
what ACL SAVE wrote — a fight between two reconcilers with the tenants
as the ball.

So no file. requirepass alone does what it says, and durability moves
to the watch, which now checks the store and not only its inputs: a
restarted store comes back empty and is re-granted within a tick,
because reconciling is against reality, not against a diff of
instructions. A run that failed leaves last empty, so the next tick
retries instead of believing the inputs were handled.
2026-09-02 01:23:38 +02:00
jschoubben 649ce9bc3b Directories belong to the number that runs inside
The lab named both failures in one run: the store restarted forever on
a conf file it could not read, and the forge could not traverse into
the directory that held its files. Both are the same fault — a file the
mesh declares root-owned, consumed by a container process that dropped
to a uid the machine has never heard of.

The forge's data now belongs to 1000, the user its container runs as.
The store's conf, ACL file and data belong to 999, which is what redis
becomes after its entrypoint drops privileges. The package registry's
conf directory belongs to 10001, which writes htpasswd into it.

Made expressible by the host in the commit beside this one: an owner
may be numeric, because a container's user has no name on the machine.
2026-09-01 23:45:17 +02:00
jschoubben d85be55e3c The media tail: bazarr, jackett, nzbget, tautulli, ombi
The same shape as the four before them — a config directory of their
own, the shared library paths they actually touch, one owner and mode
across every sharer — which is the point of doing them together: five
manifests that differ only in name, port and which shelves they read
prove the shape is a shape and not a coincidence.

All pinned by real digests, resolved on this workstation today. The
catalogue stands at 27.
2026-09-01 23:26:29 +02:00
jschoubben 1dce1fc5e7 The media four: plex, sonarr, radarr, qbittorrent
The shape they share is the interesting part: a media library is a fact
about the machine that several modules mount, so each declares exactly
the paths it touches, with one owner and mode across all of them. The
host reconciles a directory rather than creating it, so the second
module ensuring a path the first already ensured is maintenance, not a
fight — and ADR 0030 keeps a shared path alive when one of its sharers
is unassigned, because it holds what the mesh did not put there.

Plex runs on the machine's own network like home-assistant, and for the
same reason: discovery is the point. The other three publish, so the
mesh chooses their machine ports.

All pinned by real digests, resolved on this workstation today. The
catalogue stands at 22, of which the arrangement being replaced ran 95
— but its number counts workstation ricing and one-machine tooling
beside services, and only the services convert to manifests like these.
2026-09-01 23:23:43 +02:00
jschoubben 0ba52d03a3 Two more: nextcloud and home-assistant
Nextcloud is the first consumer of two provisions at once: a database
from the mesh's postgres and primary storage in a bucket from the
mesh's own object store — which is how the arrangement being replaced
ran it, minus the bundled MariaDB it no longer needs. Every credential
in one env file the host writes; the manifest holds placeholders and
the mesh holds nothing readable.

Home automation runs on the machine's own network, because discovering
devices is the point and a bridge would hide them — so nothing is
published, the declared port is the bound port, and the rule set opens
exactly it. The case MachineSide was corrected for, in the catalogue.

Both pinned by real digests, resolved on this workstation today.
2026-09-01 23:09:40 +02:00
jschoubben ef7750f816 Three more: searxng, influxdb, verdaccio
Converted from the arrangement being replaced. The search engine brings
its own valkey on its own network — a sidecar is just a second
container resource. The time-series database initialises itself from
two generated secrets, and its data and config directories are declared
with the owner the image runs as. The package registry's configuration
is a declared file rather than a merged one, which is the position
16-module-coverage takes on config merging: the module knows its own
format because it wrote the rest of the file.

Two were read and deliberately not converted, which is worth recording
where the next person will look:

n8n builds a custom image, so it is a module with a repository rather
than a manifest in this catalogue — where a module's own code lives is
ADR 0037's question, and pretending otherwise here would prejudge it.

mosquitto authenticates from a hashed password file that only
mosquitto_passwd can write, and the mesh delivers plaintext sealed
files — so an honest conversion needs a small provisioner, the same
shape as the cache's. Without one, the manifest would compose a broker
nobody can log in to, which is exactly the kind of module that parses,
resolves, and stops on the machine.

All images pinned by real digests, resolved on this workstation today.
2026-09-01 22:49:18 +02:00
jschoubben 8c4a478b09 Four more of the catalogue: registry, redis, umami, grafana
Converted from the arrangement being replaced, in its shapes rather
than theirs.

The registry is the manifest the lab already proved, promoted: names
its image by digest and is never built (04-ISSUES/029), provides the
artifact store, claims it once per machine.

Redis is the third provision after a database and a bucket, and the
first whose tenancy is a pattern in a shared keyspace rather than a
namespace something else enforces. Its provisioner mirrors the postgres
one's contract line for line — the manifest, the sealed per-consumer
files, the mark, the withdrawal of orphans — and speaks RESP directly:
five commands are needed, and a client library large enough to hide
them would be most of the program's size. A prefix Redis would read as
a pattern is refused, because the grant must mean what the manifest
said; grants are persisted with ACL SAVE, or said loudly, because a
cache that forgets its tenants on restart reports success until then.

Umami asks the mesh for its database and a generated app secret, and
carries no state of its own — the arrangement being replaced ran a
bundled second postgres beside it. Grafana keeps its dashboards in a
declared directory with the image's own owner. Both listen on 3000, as
does the forge — which is the mesh's port assignment earning its keep.

Traefik is deliberately not converted: the mesh's route provider is
mesh-route-proxy, which speaks route grants natively, and a traefik
that consumed them would be an adapter nobody has written pretending
to be a conversion.

All images pinned by real digests, resolved on this workstation today.
2026-09-01 22:44:53 +02:00
jschoubben 4d6ec5b10c One object store, not two
object-store.json and minio.json described the same thing: same image,
same provision at the same scope, same provisioner. Not two
implementations a person could choose between — one module written
twice. Assigning both to a node would have collided on `s3-bucket`.

It exists because it was written first, to pair with photos.json for the
README's worked edge, and minio.json was the fuller version of the same
module written later. Nobody removed the first.

The pair test keeps its point and now reads the surviving one. Checked
across the rest: this was the only duplicate.
2026-09-01 21:06:17 +02:00
jschoubben 1f5b70a995 The mesh assigns the port, and a module says it once
novox/hq ADR 0038. A module cannot choose a port: it is written once and
assigned anywhere, so any number it picks is a guess about a machine it
has never seen. A database module met the mesh's own store on 5432 and
was told, by a container runtime three layers down, that the port was
already allocated.

The number used to appear three times in every module — the rule set,
what a consumer is told, and what the runtime publishes — agreeing only
because one person wrote all three. Now it appears once, in `listens`,
and the other two are derived: the container publishes `20000:5432`, the
consumer is told 20000, and the rule set opens 20000.

An assignment is made once and kept, as a credential is. A port that
moved on every declaration would restart both ends each time and hand a
consumer a number that was true when it was read.

Ports the protocol fixes — mail on 25, submission on 587, DNS on 53 —
say so, and are then claims: one holder per machine, and the second is
refused by name at assignment. That is the mechanism the mesh already
has for what is singular on a machine, pointed at ports.

A mapping written the long way is left exactly as it is. Some things
must be pinned by hand, and quietly overruling somebody who wrote both
halves would be worse than not offering the short form.

Still open, and known: the substrate is not a module, so the mesh has
never heard of its own store and cannot yet assign around it. That is
what 028 will still be about after this.
2026-09-01 17:52:53 +02:00
jschoubben a5d85266d0 A container does not take restart-on, and nine of them did
This is what stopped the forge. The host refused the whole declaration:

  resource "postgres.server": a container does not use "restart-on",
  and it is set. Refused rather than ignored

`restart-on` belongs to a service. I put it on containers this morning
so one would pick up a rotated credential — nine times across seven
modules — and nothing between the manifest and the machine said a word.
The control plane composed it happily; the parser accepted it; the
manifest tests passed. The only thing that knew was the host, five steps
downstream, and hearing from it cost a seventeen-minute run.

The host was right twice over. It refused, and it refused *everything*,
because applying the parts it understood would leave a machine that
looks configured and is not. One misplaced key therefore stops a module
dead, which is the correct severity and an argument for catching it
where it is written.

So the shapes and their keys are now written down here and checked. They
are duplicated from another repository deliberately — this is its wire
format, like the shape of a grant file — and a contract with two copies
and no check is a contract until somebody edits one.

What this does not fix is why I reached for it: a container cannot
follow a file. Filed separately.
2026-09-01 17:19:04 +02:00
jschoubben f5b03e1474 Declare the directories that hold the data
novox/hq 04-ISSUES/026. Four modules mounted fourteen host paths that no
resource declared — the mail spool, the databases, the object store's
data. Each would be created by the container runtime as root, with a
mode nobody chose, so `owner` and `mode` went unapplied on exactly the
directories that matter.

The worse half: a directory the mesh declared and no longer wants is
kept rather than removed when it holds anything the mesh did not put
there. That rule is the answer to what happens to data when a module
goes away, and it is written in terms of declared directories. An
undeclared one is not covered. So the one rule guarding against data
loss reached the configuration directories, which are cheap to lose, and
missed the data directories, which are why the rule exists.

The cause is worth naming. These manifests were written by reading the
arrangement being replaced and carrying its compose files across —
service, image, ports, volumes, environment. The container shape can
express all of that, which is what made the transliteration feel like
progress. A shape that can express a compose file gets filled in like
one, and a volume line borrowed from compose declares no owner, no mode
and no intent.

Declared parent-first, because the host applies in the order written and
does not sort. The check is mechanical now, because a person comparing
volumes against directories by hand is the process that produced this.

Still open, and bigger: whether these paths are where a module's data
should live at all. They were inherited whole, and they decide what a
person backs up.
2026-09-01 16:09:23 +02:00
jschoubben ee3cc1b6f4 Pin the example modules to images that exist
novox/hq 04-ISSUES/025. Every image reference in every example module
was sixty-four zeros — eighteen of them across five modules. Each
parsed, resolved, and composed into a declaration a host accepts, and
none could ever have started: the machine reaches `docker pull` and
stops. That is why those modules were written and not running, and no
check saw it because every check passed.

The host validates the shape of a reference and nothing more, which is
correct: verifying a digest exists means reaching a registry, and that
is the one thing a host must never have to do. So the last place that
could catch this is the wrong place to try.

The guard therefore sits where a declaration is composed, not where a
manifest is parsed. A file in a repository is allowed to await a pin —
the design already says the manifest in a repository names artifacts
while the manifest the mesh holds names digests, and the bundle works
exactly that way. What must never happen is a placeholder reaching a
machine, and composing is the last moment before one does.

Twelve third-party images resolved to real digests without pulling
anything, which is also the mechanism the open issue needs. Two
discoveries came free: mailu publishes to ghcr rather than Docker Hub,
so seven references named repositories that do not exist at all; and it
renamed roundcube to webmail, so that one would have failed even with
the right registry.

What stays a placeholder is the mesh's own provisioner images, which
genuinely have no digest until built and pushed — the bundle's problem,
legitimately unresolved here. The stand-in consumer now stands in with
a real image rather than an invented one.
2026-09-01 15:13:33 +02:00
jschoubben a4090014f3 An example may not name an image nothing builds
Found by reading the manifests rather than by running them. Two of the
provisioner images the examples name had no way to be produced: the
object store's had a Dockerfile and no target, and Keycloak's did not
exist at all — no image, no Dockerfile, no program.

A module naming an image nothing produces resolves, plans, pushes and
stops on the machine at `docker pull`, which is the fault arriving as
far from its cause as it can get.

The object store's target is added. Keycloak's provisioner is removed
from its manifest, because writing a manifest for a program that does
not exist is the same mistake as the .env files: it parses, it resolves,
and it could never work.

That makes keycloak's manifest true about today — a server the mesh
runs, with its database and its admin credential — and it makes the gap
loud. Keycloak no longer claims to provide oidc-client, so a consumer
asking for one is refused at plan time by name, rather than resolving
cleanly and never having a client created.

The check covers only images beginning `mesh-`. Postgres and the rest
come from a registry and are somebody else's to build; what this bounds
is the set this repository is responsible for and might forget.
2026-09-01 03:12:49 +02:00
jschoubben e5243cd753 Ask every consumer for usable configuration, not just the one in hand
The keycloak check was written while keycloak was the module being
worked on, which is how a check ends up proving one thing about one
file. It now runs over every example that requires something, and asks
the two questions that matter for all of them: that no ${bound:...}
reached the machine as a value, and that anything named PASSWORD is
still a hole only the host can fill.

The first is the one worth having. A placeholder written through is read
as a value by whatever parses the file — a connection to a host called
"${bound:postgres-database:at}" — and the failure names neither the
module nor the mesh.

Modules whose requirements nothing in the examples answers are logged
and passed over, because that is a fact about the example set rather
than about them.
2026-09-01 03:10:43 +02:00
jschoubben 122680b554 A consumer can write its own connection string
novox/hq 04-ISSUES/023. A consumer was given its password, the address,
the port and where its credential lives, and still could not connect —
the user name was invented by the provisioner and recorded nowhere, and
the rest sat in a JSON binding that a program reading KEY=value cannot
use.

Both halves have the same cause: the mesh knew something and did not say
it.

**Who a consumer is, said once.** The provisioner used to derive
mesh_<node>_<module> and that string existed nowhere else — not in the
control plane, not in the binding, and above all not at the consumer,
which has to present it. Now the mesh derives it once and sends it to
both ends, so they agree by construction rather than by two conventions
that were the same on the day they were written. The provisioners refuse
to invent one if the mesh says nothing, because falling back to a name
of their own would create a role the consumer would never guess and
everything would report success.

**Bound values reach the file that needs them.** ${bound:provision:key}
is the symmetric twin of the sealed placeholder, and simpler: these
values are not secret, so the control plane fills them in before sending
and the host gains no field and learns no format. It stays
name-agnostic — at, as and from are true of any provision, and every
other key comes from what the provider said it serves.

The asymmetry it removes was backwards. The secret is the hard case,
because the mesh must not be able to read it, and the secret was the
part that already arrived.

Keycloak and Gitea now produce complete connections, asserted from the
manifests on disk rather than from fixtures: every part filled, no
placeholder surviving as a value, and the password still a hole only the
host can close. Three faults injected, each caught.
2026-09-01 03:03:07 +02:00
jschoubben be2dca27ab The example modules put their credentials where the programs read them
Every one of these declared `own-secrets` pointing at a path called
`.env` and then mounted it as `env-file`. The file's whole content is
the password. Docker reads that as a malformed line and the container
starts with no password set — which is not a failure to start, it is a
service running with the wrong credential.

They parsed, they resolved, and none of them could ever have worked.
That is what a manifest checked only by the parser buys.

Each now keeps the sealed file as what it is — a password, alone — and
declares a file beside it whose content says ${secret:name}. The host
fills the hole on the machine, which is the only place both halves
exist. The provisioners mount the bare file, because they read a
password file and always did.

Two tests, both driven from the manifests on disk rather than from
fixtures: every ${secret:x} must name something the module declared, and
nothing may read a bare password file as an env file. Injecting the
shipped bug reproduces it word for word.

Keycloak, Gitea and Mailu still cannot connect to their databases, for
the reason in 04-ISSUES/023 — the user name is the provisioner's
invention and the bound values cannot reach a config file. Their own
credentials are right now; that half was independent and is done.
2026-09-01 02:52:00 +02:00
jschoubben 0af3ea1acf A consumer is a module on a machine, not a machine
novox/hq 04-ISSUES/022. A credential was keyed by provision, consumer
node and provider node, so "who is asking" was answered by naming a
host. The node this mesh exists to take over runs eight modules against
one database server.

The symptom had two halves and only one was loud. The provider refused,
naming the modules and explaining they would share one credential, which
reads as a decision rather than a limit. The consumer did not refuse: it
resolved cleanly, wrote one module's credential file and left the others
absent — a service that starts and cannot authenticate, with nothing
saying why. That is 021 again on a different axis.

Three modules wanting one database produced one need, carrying whichever
module mentioned it first, because the resolution walk is a work-list
over names. The fan-out now happens in one place, after the walk. The
record path already did this correctly and said why: a consumer here is
a module on a machine. It is the same rule.

Downstream: the secret's key gains the consuming module, the grant file
is named after both halves, needs are matched by provision and module
rather than provision alone, and the provisioners name the role and the
access key after the module. The refusal in ContributionsTo is gone
because there is nothing left to refuse.

Worth stating plainly: without that refusal, gitea's login would have
opened keycloak's database. From the provisioner's side it created
exactly what it was asked to create.

Existing secrets are discarded rather than backfilled. They cannot say
which module they were for, and a secret is remade and delivered to both
ends on the next push — so this costs one rotation and invents nothing.

Also guards the role name against PostgreSQL's 63-byte truncation, which
is a notice rather than an error and would reintroduce exactly this
collision at a length nobody tests.

Three faults injected — the fan-out removed, needs matched by name
alone, the grant file named after the machine — each caught.
2026-09-01 02:40:09 +02:00
jschoubben d0511ee3fe The command API, which refuses everything until it knows who is asking
novox/hq ADR 0035: one implementation, several surfaces, and a surface
holds no decisions. The act of assigning — including that an assignment
which does not resolve is kept and still refused — moved into acts.go,
and the command line now calls it too. Two surfaces, one refusal, in the
same words.

It will not run without --issuer, and refuses at start rather than per
request so it is found by whoever ran it rather than by whoever finds
it. There is no flag that removes the check.

The authenticator is honest about what it is: no token can be verified
until an identity provider exists, because that is a module and none is
running, so every request is refused and told that the command line
still works. A surface that functioned without authentication would be
one somebody left running — and the board this stands behind is
published on a public name.

Four refusals, four tests. The last one first asserted "not 200", which
passed because a request with no database fails at the store anyway — it
proved nothing about whether the input was checked. It now asserts the
specific refusal, and bites when the check is removed.
2026-09-01 01:55:56 +02:00
jschoubben f04d00b411 The proxy can obtain a public certificate, and asks staging by default
Work breakdown 1.4. The mesh's own authority certifies internal names
and always did; a name reachable from outside needs one the world
already trusts, and there was no ACME anywhere in this repository.

Uses acme/autocert from x/crypto, which was already a dependency — one
indirect addition (x/net, for idna) and no new direct one.

Three things worth more than the feature:

**Staging is the default** (novox/hq 04-ISSUES/004). Production issuance
is rate-limited per domain and per account and does not replenish
quickly. Defaulting to production would leave the safe path depending on
remembering to opt out, on exactly the work most likely to iterate. A
staging certificate is trusted by no browser, so the mistake announces
itself on the first request rather than a fortnight later.

**A certificate is only asked for on a name the mesh routes here.**
Without that policy, anything that can reach the port and send a name
triggers an order for it — a scan becomes a stream of failed orders
against the account's rate limit, and the proxy looks healthy
throughout. What it may certify is what it was told to serve.

**A private issuer is trusted by naming a file, never by skipping
verification.** Skip would still apply on the day this points at a
public issuer, and nothing would say so.

TLS is opt-in: without TLS_LISTEN the proxy serves plain HTTP exactly as
before, which is what an internal-only mesh wants. With it and no cache,
it refuses rather than defaulting — every restart would otherwise order
new certificates, silently, until the rate limit says it does not.
2026-08-31 19:35:34 +02:00
jschoubben 981cd4139b Reference the record rather than link across repositories
A relative link out of this repository resolves nowhere on a forge. Every
other reference here names the record in prose.
2026-08-31 17:52:25 +02:00
jschoubben 950ebb9e28 The two halves of an object-store edge, as a readable pair
Manifests for a provider and a consumer, so the contract can be read
rather than only exercised through a lab fixture that stages the grants
by hand.

Checked as a pair rather than separately, because two manifests that
only ever parse alone are two manifests nobody has held against each
other. The test asserts the names match, that each side says where it
wants to be told, and that the consumer contributes the key the
provisioner actually reads.

That last one is the trap worth having a test for: a consumer
contributing "name" — which is exactly what a database consumer
contributes — resolves cleanly, deploys, and then fails on the machine
with "asked for a bucket and did not name it". Nothing in that message
points back at the manifest that caused it. Both mistakes were made
while writing these two files.
2026-08-31 17:52:14 +02:00
jschoubben ed9a30f22d A module can be given a bucket: the provisioner that makes a secret true
Phase 1.1 of the work breakdown. The finding that shaped it came before
any code: **the control plane special-cases nothing.** provides,
requires, contributes and grants are entirely name-agnostic, so asking
for a bucket needed no change to the mesh at all — only a provider that
answers. What was missing was the last step, where something on the
machine turns a delivered secret into a key that works.

Named `s3-bucket` by ADR 0027's test: a consumer's code is written
against the S3 API, and swapping one store for another does not break
it, so the coupling is to the protocol rather than the product — which
is what the substrate design already said about AMQP, S3 and OCI.

Proven on a real store, 7 assertions: a generated secret becomes a
working key; rotation makes the new one work and the old one stop; a
consumer that goes away loses its key; a key nobody here made is left
alone; a manifest naming a credential that was never written is refused;
an unusable bucket name is refused naming the consumer that asked.

**And the one a database does not need.** One PostgreSQL server holds
separate databases and the product enforces the boundary; one object
store holds every bucket behind one endpoint, so a consumer being unable
to reach another's is a policy somebody wrote. A policy granting
arn:aws:s3:::* would pass every other test in the file, so the unit
tests assert what the policy does NOT say.

It drives the vendor's command line rather than an SDK: the admin API
encrypts its request bodies, which is why a separate admin library
exists, and pulling that in would add a system-metrics dependency tree
to a repository with none in order to create a user.
2026-08-31 17:51:12 +02:00
jschoubben ab1dd34d12 The resolver does not ask itself for upstreams
dnsmasq read /etc/resolv.conf to find where to forward. Whatever points a
machine at the mesh writes its own address into that file — so dnsmasq's
upstream was dnsmasq, and every query it could not answer locally looped. Its
receive queue filled with 15KB of them and every lookup on the machine hung,
which is why this arrived as a thirty-second timeout rather than a wrong
answer.

It needs no upstream at all: the asking module routes only the mesh's suffix
here and leaves everything else where the machine already sent it. And it names
none, because choosing one would send every query this machine makes somewhere
nobody agreed to.

Also corrected: the comment claiming it takes only 127.0.0.55. Listening on a
loopback address makes dnsmasq take the rest of loopback with it, 127.0.0.1
included — which is what claiming `the-dns-port` already says, and which the
comment was quietly denying. That is the same comfortable claim as ".54 is
free", in the same file, made twice.
2026-08-31 14:09:50 +02:00
jschoubben e3a2790acd The resolver answers on an address systemd does not hold
`127.0.0.54` is systemd-resolved's DNS *proxy* stub. The module asserted it was
free, in a comment that read as reasoned — "not .53, that is
systemd-resolved's" — and it was simply wrong: resolved holds both. dnsmasq
could not create the socket and never started.

Nothing in a unit test could have caught it. They checked the module names an
address and that the asking modules point at the same one, and all of that
passed while the daemon could not start. Only a machine knows which addresses
are spare, which is the argument for proving a module that asserts facts about
machines on a machine, before believing the assertions.

So it moves to .55, and says what that is: a convention, not a reservation. If
a future systemd takes it, this line changes and nothing else does.

The tests now derive the address from the serving module and check the two
asking modules agree with it, rather than naming it a fourth time — that fourth
place is the one nobody would think to change.

And the lab assigns `resolved-split-dns` rather than `resolv-conf`: those
machines run systemd-resolved, which owns the file. The two claim the same
thing precisely so the wrong choice is a refusal rather than a fight, and
picking the wrong one was testing the fight.
2026-08-31 13:39:58 +02:00
jschoubben bff893f8af Resolver modules: one that serves, and two ways of deciding what a machine asks
Three manifests and the rule that keeps them apart. Serving and asking are
genuinely different roles, and systemd-resolved can only do the second — it
cannot answer a wildcard, it routes the mesh's suffix to something that can. A
module that treated them as one role could not work, which is the mistake worth
naming rather than discovering.

So `the-dns-port` and `the-resolver-configuration` are two claims. A machine
gets one of each, and two of either is refused by the mesh rather than fought
over on the machine — which is what ADR 0009's table meant by listing resolvers
beside the seat and pid 1. That table names the resource `/etc/resolv.conf`,
which is what it is; a claim is a name in the catalogue's own form, and the
catalogue refuses the path as one.

Neither module knows anything about the machine it is on, which is what lets
them be static manifests: they name `mesh0` and `127.0.0.54`, both chosen by
the mesh, rather than an address only that machine has. Not 127.0.0.1 and not
127.0.0.53 — taking either would be a module claiming something it did not say
it claims.

A service can now reflect a file another module put on the machine, written
`<module>.<id>`. The resolver has to restart when the mesh rewrites the names;
without it, it would serve the names it started with for ever, with every
machine that joined afterwards unreachable and every check passing.
2026-08-31 12:21:58 +02:00
jschoubben d0c0ee8dab A route is a grant, and a provider is told where its consumer is
novox/hq 08-connectivity §3, built. The mirror of a database grant: there the
consumer supplies a name and receives credentials; here it supplies a target
and receives a name. Nothing new in the vocabulary — a route is a provision
like any other.

One field was missing and it is the one that matters for anything reaching
back: a contribution now carries where the mesh says that machine is. A reverse
proxy is told to send traffic to a consumer and has to open a connection, so
without it every provider implementing a provision would have to know how the
mesh names machines — a convention leaking into every module.

The proxy itself is an example, not part of the control plane: the contract is
the file, not this program. It replaces its table whole rather than merging,
because the file is the whole truth about who has a route and merging would
keep serving a name whose module was unassigned — the stale-route fault
08-connectivity lists as open, reintroduced one level down. A name it does not
serve is refused by saying which it does: a route withdrawn and a name that
never existed are different things.
2026-08-31 02:43:19 +02:00
jschoubben ebcfd37b92 Rotate a credential and move both ends together
The invariant novox/hq ADR 0001 records as unowned, and it was measurably
false in HAL: a provision documented as never rotating minted a new password on
every adoption and updated only the provider's row. Consumers on three nodes
held dead credentials for two days while the mesh reported success. Nothing
enumerated who held the old one.

Three things make that impossible here. The holders are a set the mesh can name
— each pair has its own credential, so rotating one consumer touches one role
and the affected list is a query rather than an assumption. Both ends are
pushed by this command rather than a later one, because leaving the sending to
whoever remembered is the fault exactly. And it is all-or-nothing: if any
affected machine cannot be resolved, nothing is sent and the old credential
keeps working, which is a mesh that has not rotated rather than one that has
half-rotated.

The window is stated rather than hidden: a role's password changes on the
provider and the file changes on the consumer, and they cannot be simultaneous.

The provisioner now takes its superuser password from the file the mesh wrote,
which is how the mesh delivers one. Passing it through the environment needed a
person in the middle of the one path that exists so there is not one — and put
a superuser password where `docker inspect` prints it.
2026-08-31 02:38:52 +02:00
jschoubben c37d368f65 A module may need a secret of its own, and the provisioner watches
Two things, both found by trying to write a real postgres module and
discovering it could not be said.

A database has a superuser password, a broker an administrator, a
registry an account. None of them is *for* anybody — they are not the
credential a consumer is given, and the mechanism that hands those out
has a consumer in the middle of it. So a module may declare what it needs
and where to put it, and the mesh generates one per node, seals it, and
reads it no more than it reads any other.

Per node, deliberately: a module running on three machines has three
passwords. One in the manifest instead would put the same secret on every
machine that ever runs it, in a file anybody can read, for ever. Made
once and kept, or a running database would be handed a password it was
not started with; remade when the machine's sealing key changes, like
everything else sealed here.

A need declared and not made is refused rather than skipped, because a
module whose own credential is silently absent starts, fails to
authenticate, and the reason is three layers from the machine reporting
it.

And the provisioner can watch. That is what lets it be a module rather
than a binary somebody places: run once, it needs invoking after every
declaration by a timer or a unit wired to a file; watching, it is an
ordinary long-running service the host already supervises. It polls
rather than watching the filesystem, because the host writes atomically —
the file is replaced, so a watch on the path stops seeing anything after
the first replacement, and a watcher that silently stops working is worse
than a poll. Credentials are compared by digest and never held: this runs
for as long as the machine is up.
2026-08-30 18:22:05 +02:00
jschoubben c3046dcf56 A provider is told who its consumers are, and a reference provisioner
Contributions were node-local, so a mesh-scoped provider — the one case
that most needs them — never heard from its consumers. A database was
given a password and no idea what to create it for.

Cross-node consumers now reach the provider's `receives` file, merged in
with the ones on its own machine: from the provider's side they are the
same thing, and a provider that had to read two lists would read one of
them. Each names the file its credential is in rather than carrying it,
because the mesh discarded the value and could not put it there. The
readable half therefore stays readable.

And examples/postgres-provisioner, which is the last step: it reads what
the host wrote and makes PostgreSQL accept it. Explicitly not part of the
control plane — the control plane decides and never touches a machine.
This runs on the machine and touches it, and a real one ships with the
module that ships PostgreSQL. It lives here because this is where the
contract is defined, written as something that runs so it can be read.

It reconciles rather than applying a change, because it is never told
what changed. Three things that follow, and each is a fault somebody has
shipped:

- the password is set every time, not only on creation, or a rotation
  reports success and changes nothing
- what it made and nobody asks for any more is revoked, or a departed
  consumer keeps a working login for ever
- what it did not make is left alone, or it cannot be run on a database
  that predates it

Proven in the lab against a real PostgreSQL, each assertion confirmed to
fail with the behaviour removed. The suite is in mesh-lab, which also
records the two ways the test itself was wrong first.
2026-08-30 01:31:25 +02:00