Commit Graph
150 Commits
Author SHA1 Message Date
jschoubben 0d975a051d Asking what the mesh would send must not change it
`status` hung. It composes a declaration for every node to answer *is
this machine running what I would send it*, and composing one assigns
each module a machine port — so the question wrote to the database, and
wrote to the same rows as the machine it was asking about.

`port_assignment` is unique on (node, machine). Two transactions
inserting the same port do not race, they queue: the second waits on the
index until the first commits. A status polled every two seconds while a
node applies is two writers on those rows, and the poll stopped
returning rather than returning something wrong — which is the better
failure of the two, and still a failure.

The latent version of this was there before anything polled: two
compositions running at once could both allocate.

So allocation belongs to the send path alone. The mesh chooses a port
when it commits to sending one; every other caller reads what was
chosen. A module with nothing assigned has never been sent, which is
precisely what "waiting" means — the read needs no number to be right
about that, and inventing one would make the answer worse.

Named rather than passed as a bare bool: at three call sites, `true` and
`false` say nothing about which of these two things is meant.

Checked by the lab, which now polls status throughout an apply.
2026-09-01 20:05:08 +02:00
jschoubben 83c6a2f244 Withdraw the mount check: it refuses the builder
The rule was right about data and wrong about everything else. The
builder mounts the container runtime's socket, which is not its data,
does not belong to it, and must not be declared as one of its
directories — and the check refused the builder's own manifest.

Caught by the lab, though not honestly: the run was already going when
this went in, so the builder binary was rebuilt mid-run with the check
compiled into it and the failure was mine, not the mesh's. Confirmed
against the manifest directly rather than inferred from the log.

What it was protecting is real and stands — the fourteen mounts are all
declared. But enforcing it needs a way to tell "the directory my data
lives in" from "a machine facility I was granted", and the mesh has no
vocabulary for the second. `capabilities` is the closest thing and does
not name paths. That is a design decision, so it goes back to
04-ISSUES/026 rather than being invented here to make a check pass.

The check that every real manifest still parses is kept. It costs
nothing and it is how the next attempt at this finds out sooner.
2026-09-01 19:45:36 +02:00
jschoubben 53eb000a84 A container may not mount a path the module never declared
Closes the half of 04-ISSUES/026 that would otherwise come back. The
fourteen mounts across the forge, the mail system, the store and the
object store are all declared now — but nothing said they had to be, so
they were right by coincidence and the next volume added would not be.

A bind mount whose source does not exist is created by the container
runtime, as root, with a mode it picks. So `owner` and `mode` — which
exist precisely so a module can say who its data belongs to — were
silently not applied to the only directories holding data.

And the rule written for exactly this case did not reach them. A
directory the mesh declared and no longer wants is kept, not removed,
when it holds anything the mesh did not put there (ADR 0030). That is
the answer to *what happens to my data when a module goes away*, and it
is written in terms of declared directories: an undeclared one sits
outside it, because the mesh does not know it is there.

Refused where it is written rather than on the machine, which cannot
tell the difference — by the time the host sees the mount it is being
asked to make a directory, which it is perfectly able to do. The fault
is in the manifest, so it is named at the manifest. Same argument as the
action refusal directly above it.

A path under a declared directory counts as declared, as do the files a
module already names: its own secrets, its grants, what it receives.

Every real manifest is checked to still parse, and the refusal bites.
2026-09-01 19:36:01 +02:00
jschoubben c67f836185 The mesh may only move a port it actually publishes
The lab caught this: a module declaring a port and running no container
had its rule set opened on 20000 while its service sat on 9101. The
firewall reported success and blocked the thing it was told to admit,
which is the precise failure the filtering comment warns about, arrived
at from the other side.

Assignment was applied to every declared port. But a container's mapping
is the thing that translates, and where there is none the software binds
what it binds — the mesh choosing a number does not move the service, it
only makes the mesh wrong about where it is.

The declaration side already knew this: publishedOn rewrites container
ports and nothing else. Filtering did not, so the two disagreed about
the same fact. MachineSide is now the one derivation both follow.

It also fixes a second case nobody had hit yet: a mapping the manifest
wrote itself, like the mail system's 7080:80. That is passed through
untouched when composing, so assigning it a machine port would have
opened a rule on a port the container does not publish. Either side of
such a mapping now names it, and the host side is the answer — a module
may read `listens` as what its software binds or as what the machine
exposes, and both readings want the same number.

Recorded either way, assigned or not: the map means where this module's
port is on this machine, and every reader needs that answer regardless
of who chose it.

Tests bite — making it always assignable reproduces the lab failure.
2026-09-01 19:31:05 +02:00
jschoubben 41f7c51032 Assign around what a machine already holds
The other half of ADR 0038, and what 04-ISSUES/028 was actually about.
A module can now avoid colliding with another module; until this it
could not avoid colliding with the mesh itself.

The substrate is not a module. A node raises it from the bundle it
carries before any mesh exists, so the control plane had never heard of
the store, the broker, or its own container — and handed a database
module 5432, which the store already had.

So the machine says. The host records what each resource binds,
distinguishing what it carried from what the mesh sent — a distinction
that already existed so the two never remove each other — and reports
the carried ones. The node states and this context writes, which is the
shape of every message between them.

What the declaration binds, not what is open. A machine's open ports are
a moving target, and assigning around them would mean a port that was
free when it was asked for and taken when it was used.

Replaced whole each time rather than merged: a machine that gave a port
back must be believed about that too, and a set that only grows keeps a
port reserved for something no longer there.

Tested against a real database, and the tests bite — removing the check
hands the module 20000, which the machine had said it holds.
2026-09-01 18:32:38 +02:00
jschoubben 1f5b70a995 The mesh assigns the port, and a module says it once
novox/hq ADR 0038. A module cannot choose a port: it is written once and
assigned anywhere, so any number it picks is a guess about a machine it
has never seen. A database module met the mesh's own store on 5432 and
was told, by a container runtime three layers down, that the port was
already allocated.

The number used to appear three times in every module — the rule set,
what a consumer is told, and what the runtime publishes — agreeing only
because one person wrote all three. Now it appears once, in `listens`,
and the other two are derived: the container publishes `20000:5432`, the
consumer is told 20000, and the rule set opens 20000.

An assignment is made once and kept, as a credential is. A port that
moved on every declaration would restart both ends each time and hand a
consumer a number that was true when it was read.

Ports the protocol fixes — mail on 25, submission on 587, DNS on 53 —
say so, and are then claims: one holder per machine, and the second is
refused by name at assignment. That is the mechanism the mesh already
has for what is singular on a machine, pointed at ports.

A mapping written the long way is left exactly as it is. Some things
must be pinned by hand, and quietly overruling somebody who wrote both
halves would be worse than not offering the short form.

Still open, and known: the substrate is not a module, so the mesh has
never heard of its own store and cannot yet assign around it. That is
what 028 will still be about after this.
2026-09-01 17:52:53 +02:00
jschoubben a5d85266d0 A container does not take restart-on, and nine of them did
This is what stopped the forge. The host refused the whole declaration:

  resource "postgres.server": a container does not use "restart-on",
  and it is set. Refused rather than ignored

`restart-on` belongs to a service. I put it on containers this morning
so one would pick up a rotated credential — nine times across seven
modules — and nothing between the manifest and the machine said a word.
The control plane composed it happily; the parser accepted it; the
manifest tests passed. The only thing that knew was the host, five steps
downstream, and hearing from it cost a seventeen-minute run.

The host was right twice over. It refused, and it refused *everything*,
because applying the parts it understood would leave a machine that
looks configured and is not. One misplaced key therefore stops a module
dead, which is the correct severity and an argument for catching it
where it is written.

So the shapes and their keys are now written down here and checked. They
are duplicated from another repository deliberately — this is its wire
format, like the shape of a grant file — and a contract with two copies
and no check is a contract until somebody edits one.

What this does not fix is why I reached for it: a container cannot
follow a file. Filed separately.
2026-09-01 17:19:04 +02:00
jschoubben f5b03e1474 Declare the directories that hold the data
novox/hq 04-ISSUES/026. Four modules mounted fourteen host paths that no
resource declared — the mail spool, the databases, the object store's
data. Each would be created by the container runtime as root, with a
mode nobody chose, so `owner` and `mode` went unapplied on exactly the
directories that matter.

The worse half: a directory the mesh declared and no longer wants is
kept rather than removed when it holds anything the mesh did not put
there. That rule is the answer to what happens to data when a module
goes away, and it is written in terms of declared directories. An
undeclared one is not covered. So the one rule guarding against data
loss reached the configuration directories, which are cheap to lose, and
missed the data directories, which are why the rule exists.

The cause is worth naming. These manifests were written by reading the
arrangement being replaced and carrying its compose files across —
service, image, ports, volumes, environment. The container shape can
express all of that, which is what made the transliteration feel like
progress. A shape that can express a compose file gets filled in like
one, and a volume line borrowed from compose declares no owner, no mode
and no intent.

Declared parent-first, because the host applies in the order written and
does not sort. The check is mechanical now, because a person comparing
volumes against directories by hand is the process that produced this.

Still open, and bigger: whether these paths are where a module's data
should live at all. They were inherited whole, and they decide what a
person backs up.
2026-09-01 16:09:23 +02:00
jschoubben ee3cc1b6f4 Pin the example modules to images that exist
novox/hq 04-ISSUES/025. Every image reference in every example module
was sixty-four zeros — eighteen of them across five modules. Each
parsed, resolved, and composed into a declaration a host accepts, and
none could ever have started: the machine reaches `docker pull` and
stops. That is why those modules were written and not running, and no
check saw it because every check passed.

The host validates the shape of a reference and nothing more, which is
correct: verifying a digest exists means reaching a registry, and that
is the one thing a host must never have to do. So the last place that
could catch this is the wrong place to try.

The guard therefore sits where a declaration is composed, not where a
manifest is parsed. A file in a repository is allowed to await a pin —
the design already says the manifest in a repository names artifacts
while the manifest the mesh holds names digests, and the bundle works
exactly that way. What must never happen is a placeholder reaching a
machine, and composing is the last moment before one does.

Twelve third-party images resolved to real digests without pulling
anything, which is also the mechanism the open issue needs. Two
discoveries came free: mailu publishes to ghcr rather than Docker Hub,
so seven references named repositories that do not exist at all; and it
renamed roundcube to webmail, so that one would have failed even with
the right registry.

What stays a placeholder is the mesh's own provisioner images, which
genuinely have no digest until built and pushed — the bundle's problem,
legitimately unresolved here. The stand-in consumer now stands in with
a real image rather than an invented one.
2026-09-01 15:13:33 +02:00
jschoubben 2835f41a64 Stop committing a 12 MB binary I added by accident today
The lab is pointed at a path for the builder it should write, and I
pointed it at the repository root instead of build/, which .gitignore
already covers. Two of today's commits carry the compiled binary as a
result.

Untracked and ignored by name, so the same slip does not land it again.
2026-09-01 03:21:45 +02:00
jschoubben e49586646b A comment claimed a test that does not exist
I wrote that a test asserts the control plane's placeholder expression
and the host's still agree. None does, and none in this repository could
— a unit test here can only assert what this repository already
believes.

That is precisely the thing this project refuses to tolerate: a stated
rule with no way to check it, which costs more than no rule because
people believe it. Written by me, today, in the same file that closes a
gap of the same kind.

What actually proves it is the lab, and the comment now says so.
2026-09-01 03:15:22 +02:00
jschoubben a4090014f3 An example may not name an image nothing builds
Found by reading the manifests rather than by running them. Two of the
provisioner images the examples name had no way to be produced: the
object store's had a Dockerfile and no target, and Keycloak's did not
exist at all — no image, no Dockerfile, no program.

A module naming an image nothing produces resolves, plans, pushes and
stops on the machine at `docker pull`, which is the fault arriving as
far from its cause as it can get.

The object store's target is added. Keycloak's provisioner is removed
from its manifest, because writing a manifest for a program that does
not exist is the same mistake as the .env files: it parses, it resolves,
and it could never work.

That makes keycloak's manifest true about today — a server the mesh
runs, with its database and its admin credential — and it makes the gap
loud. Keycloak no longer claims to provide oidc-client, so a consumer
asking for one is refused at plan time by name, rather than resolving
cleanly and never having a client created.

The check covers only images beginning `mesh-`. Postgres and the rest
come from a registry and are somebody else's to build; what this bounds
is the set this repository is responsible for and might forget.
2026-09-01 03:12:49 +02:00
jschoubben e5243cd753 Ask every consumer for usable configuration, not just the one in hand
The keycloak check was written while keycloak was the module being
worked on, which is how a check ends up proving one thing about one
file. It now runs over every example that requires something, and asks
the two questions that matter for all of them: that no ${bound:...}
reached the machine as a value, and that anything named PASSWORD is
still a hole only the host can fill.

The first is the one worth having. A placeholder written through is read
as a value by whatever parses the file — a connection to a host called
"${bound:postgres-database:at}" — and the failure names neither the
module nor the mesh.

Modules whose requirements nothing in the examples answers are logged
and passed over, because that is a fact about the example set rather
than about them.
2026-09-01 03:10:43 +02:00
jschoubben 71f77617e3 Omit a consumer's identity where there is none
A contribution that is not a credential grant — a module offering
something to another on its own machine — has nobody to be identified
to, and was carrying an empty `as`. A field that is always present and
usually empty teaches a reader to ignore it, including when it is not.
2026-09-01 03:09:18 +02:00
jschoubben 122680b554 A consumer can write its own connection string
novox/hq 04-ISSUES/023. A consumer was given its password, the address,
the port and where its credential lives, and still could not connect —
the user name was invented by the provisioner and recorded nowhere, and
the rest sat in a JSON binding that a program reading KEY=value cannot
use.

Both halves have the same cause: the mesh knew something and did not say
it.

**Who a consumer is, said once.** The provisioner used to derive
mesh_<node>_<module> and that string existed nowhere else — not in the
control plane, not in the binding, and above all not at the consumer,
which has to present it. Now the mesh derives it once and sends it to
both ends, so they agree by construction rather than by two conventions
that were the same on the day they were written. The provisioners refuse
to invent one if the mesh says nothing, because falling back to a name
of their own would create a role the consumer would never guess and
everything would report success.

**Bound values reach the file that needs them.** ${bound:provision:key}
is the symmetric twin of the sealed placeholder, and simpler: these
values are not secret, so the control plane fills them in before sending
and the host gains no field and learns no format. It stays
name-agnostic — at, as and from are true of any provision, and every
other key comes from what the provider said it serves.

The asymmetry it removes was backwards. The secret is the hard case,
because the mesh must not be able to read it, and the secret was the
part that already arrived.

Keycloak and Gitea now produce complete connections, asserted from the
manifests on disk rather than from fixtures: every part filled, no
placeholder surviving as a value, and the password still a hole only the
host can close. Three faults injected, each caught.
2026-09-01 03:03:07 +02:00
jschoubben 96f90ab986 Refuse an action where it was written, not on the machine
A module may not declare an action: the link may not carry a command to
run, and that bound is what limits a compromised control plane to shapes
it cannot turn into arbitrary code (novox/hq ADR 0005). The host
enforces it, correctly and in the right place.

But a module's resources reach a machine over the link, so a manifest
carrying an action was accepted here, stored, resolved, planned and
pushed — and refused on the machine, in the host's log, with nothing
connecting it back to the manifest that caused it.

The rule held. It was just unusable, which is the same shape as the
network shape earlier today: the refusal was right, arrived far from its
cause, and nobody was reading the log.

The refusal names the rule and what to do instead, because "you may not"
with no alternative is where a module author stops.

Found while checking a claim I had written in the coverage document —
that a module cannot declare one. It could; it just could not deliver
it. The document is corrected.
2026-09-01 02:55:56 +02:00
jschoubben be2dca27ab The example modules put their credentials where the programs read them
Every one of these declared `own-secrets` pointing at a path called
`.env` and then mounted it as `env-file`. The file's whole content is
the password. Docker reads that as a malformed line and the container
starts with no password set — which is not a failure to start, it is a
service running with the wrong credential.

They parsed, they resolved, and none of them could ever have worked.
That is what a manifest checked only by the parser buys.

Each now keeps the sealed file as what it is — a password, alone — and
declares a file beside it whose content says ${secret:name}. The host
fills the hole on the machine, which is the only place both halves
exist. The provisioners mount the bare file, because they read a
password file and always did.

Two tests, both driven from the manifests on disk rather than from
fixtures: every ${secret:x} must name something the module declared, and
nothing may read a bare password file as an env file. Injecting the
shipped bug reproduces it word for word.

Keycloak, Gitea and Mailu still cannot connect to their databases, for
the reason in 04-ISSUES/023 — the user name is the provisioner's
invention and the bound values cannot reach a config file. Their own
credentials are right now; that half was independent and is done.
2026-09-01 02:52:00 +02:00
jschoubben 0af3ea1acf A consumer is a module on a machine, not a machine
novox/hq 04-ISSUES/022. A credential was keyed by provision, consumer
node and provider node, so "who is asking" was answered by naming a
host. The node this mesh exists to take over runs eight modules against
one database server.

The symptom had two halves and only one was loud. The provider refused,
naming the modules and explaining they would share one credential, which
reads as a decision rather than a limit. The consumer did not refuse: it
resolved cleanly, wrote one module's credential file and left the others
absent — a service that starts and cannot authenticate, with nothing
saying why. That is 021 again on a different axis.

Three modules wanting one database produced one need, carrying whichever
module mentioned it first, because the resolution walk is a work-list
over names. The fan-out now happens in one place, after the walk. The
record path already did this correctly and said why: a consumer here is
a module on a machine. It is the same rule.

Downstream: the secret's key gains the consuming module, the grant file
is named after both halves, needs are matched by provision and module
rather than provision alone, and the provisioners name the role and the
access key after the module. The refusal in ContributionsTo is gone
because there is nothing left to refuse.

Worth stating plainly: without that refusal, gitea's login would have
opened keycloak's database. From the provisioner's side it created
exactly what it was asked to create.

Existing secrets are discarded rather than backfilled. They cannot say
which module they were for, and a secret is remade and delivered to both
ends on the next push — so this costs one rotation and invents nothing.

Also guards the role name against PostgreSQL's 63-byte truncation, which
is a notice rather than an error and would reintroduce exactly this
collision at a length nobody tests.

Three faults injected — the fan-out removed, needs matched by name
alone, the grant file named after the machine — each caught.
2026-09-01 02:40:09 +02:00
jschoubben 1314be5282 A file may hold a credential where its content says one belongs
The gap that stopped keycloak and gitea from starting. A granted
credential arrives as a file whose entire content is the password, which
is what a program reading a password file wants — and most programs do
not read one. They read KEY=value, or a JSON document with the token at
an attribute inside it. A module in that position could be handed the
bare value or nothing, and both are useless.

The host has been able to do this all along: content with ${secret:name}
in it, sealed values beside it, substitution on the machine, which is
the only place both halves exist. Nothing filled the values in, so the
hole could be written and never closed and the host refused the file.
That refusal was correct and the feature was unreachable.

A module reaches its own secrets and the credentials it was granted —
both things it wrote in its own manifest — and nothing else. Naming
another module's is refused: two modules on one machine are as separate
as two on different machines, and letting one read the other's
credential by guessing a name would end that to save writing a file.

Filling runs after settings, which is the whole reason it sits where it
does. A setting is how a placeholder gets into a JSON document in the
first place — the desktop client that reads its token from an attribute,
not an environment variable. Before the merge that file's content is
"{}" and asks for nothing.

Tested through Declaration rather than through the helper. Three times
in this repository a test asserted on a helper while the code calling it
was wrong, and each time the injected fault stayed silent. Three faults
injected here — the call removed, the call moved before settings, and
the module boundary widened — each caught by the test meant for it.
2026-09-01 02:28:50 +02:00
jschoubben df62bb57e5 A requirement answered on this machine is still a requirement
novox/hq 04-ISSUES/021. Two modules where one provided what the other
required, on one node, resolved cleanly with zero needs: no credential
was made, the consumer's secret file was never written, and whatever
read it would fail somewhere else entirely. Nothing was refused and
nothing was reported.

The world a node resolves against is every OTHER node, so a provider on
the same machine never became a Needed, and the credential loop walks
Needs. Every step reasonable, the sum a silent gap.

It survived because everything proven until now was cross-machine —
the interesting case for a mesh and the rare one in practice. The first
module to want a database on its own machine was the first real one.

The assumption underneath was that a local consumer needs no credential,
which holds for a process reaching a unix socket where the system can
vouch for the caller. It does not hold for containers, which is how
nearly everything here runs: the consumer reaches the provider over TCP
from its own container and the database asks for a password exactly as
it would from another machine. **The machine stops being a trust
boundary once both ends are containers.**

A brokered provision answered here is now a need naming this node, and
carries what the provider serves — which a local provider never
contributes through the world. A name nothing grants is unchanged: a
shell answered here is answered, and nothing more is owed. Both
directions tested, both injections bite.
2026-09-01 02:15:26 +02:00
jschoubben d0511ee3fe The command API, which refuses everything until it knows who is asking
novox/hq ADR 0035: one implementation, several surfaces, and a surface
holds no decisions. The act of assigning — including that an assignment
which does not resolve is kept and still refused — moved into acts.go,
and the command line now calls it too. Two surfaces, one refusal, in the
same words.

It will not run without --issuer, and refuses at start rather than per
request so it is found by whoever ran it rather than by whoever finds
it. There is no flag that removes the check.

The authenticator is honest about what it is: no token can be verified
until an identity provider exists, because that is a module and none is
running, so every request is refused and told that the command line
still works. A surface that functioned without authentication would be
one somebody left running — and the board this stands behind is
published on a public name.

Four refusals, four tests. The last one first asserted "not 200", which
passed because a request with no database fails at the store anyway — it
proved nothing about whether the input was checked. It now asserts the
specific refusal, and bites when the check is removed.
2026-09-01 01:55:56 +02:00
jschoubben 8cf1ecf6a5 secret accept: see the flag that comes after the arguments
The command refused every real invocation. Go's flag package stops
parsing at the first non-flag argument, so with the positionals first —
the order that reads correctly — `--from -` stayed among them and the
count check rejected it.

The host's own parser carries a note about this exact fault, and the
version it describes is worse: there a flag somebody passed was silently
ignored and the command succeeded anyway. This one at least refused.

The tests did not catch it because every case in them was a rejection.
The command was broken in the only way that matters — it refused what it
is for — and the suite was green. The lab found it at the first call.

Two tests now: the helper, and the command itself with a --from naming a
file that is not there, so the complaint must be about the file rather
than about usage. The second exists because injecting against the first
stayed silent: testing the helper alone left the command free to ignore
it entirely.
2026-08-31 23:05:10 +02:00
jschoubben 7e9c28fd9e secret accept — carry a value the mesh did not make
The entry point for adopting something already running, and the half
that was missing. The store has carried the distinction since the
beginning — a module secret records whether it was `made` or `accepted`,
and refuses to invent a replacement for the second — and
AcceptSecretForModule existed, with exactly one caller: the broker
account issued to a build machine. Nothing else could write one.

Without it every module secret is generated, which against a database
that already exists puts 32 random bytes where a working credential was.
The machine applies it, reports success, and whatever reads it fails to
authenticate somewhere else entirely, with the mesh insisting the secret
was delivered — which it was.

The value is read from a file or from standard input, never from an
argument: a value on the command line is in the shell's history and in
the process list. Same path a model-access key already takes, and no new
dependency — the first version reached for x/term and the existing one
needed nothing.

Sealed on the way in, plaintext discarded, and not printed back. The
only difference from a generated secret is where the value came from.

Two rules with a test each, and the second is the one that would have
been got wrong: only the line ending is removed, never surrounding
space. Trimming both ends is the obvious thing and would deliver a
password chosen with a leading space as a different password, silently.

Both were briefly untested for different reasons — the trimming lived
where no test could reach it, and then a -run filter matched neither
test. Extracted, and injected against the whole suite.
2026-08-31 21:57:57 +02:00
jschoubben f04d00b411 The proxy can obtain a public certificate, and asks staging by default
Work breakdown 1.4. The mesh's own authority certifies internal names
and always did; a name reachable from outside needs one the world
already trusts, and there was no ACME anywhere in this repository.

Uses acme/autocert from x/crypto, which was already a dependency — one
indirect addition (x/net, for idna) and no new direct one.

Three things worth more than the feature:

**Staging is the default** (novox/hq 04-ISSUES/004). Production issuance
is rate-limited per domain and per account and does not replenish
quickly. Defaulting to production would leave the safe path depending on
remembering to opt out, on exactly the work most likely to iterate. A
staging certificate is trusted by no browser, so the mistake announces
itself on the first request rather than a fortnight later.

**A certificate is only asked for on a name the mesh routes here.**
Without that policy, anything that can reach the port and send a name
triggers an order for it — a scan becomes a stream of failed orders
against the account's rate limit, and the proxy looks healthy
throughout. What it may certify is what it was told to serve.

**A private issuer is trusted by naming a file, never by skipping
verification.** Skip would still apply on the day this points at a
public issuer, and nothing would say so.

TLS is opt-in: without TLS_LISTEN the proxy serves plain HTTP exactly as
before, which is what an internal-only mesh wants. With it and no cache,
it refuses rather than defaulting — every restart would otherwise order
new certificates, silently, until the rate limit says it does not.
2026-08-31 19:35:34 +02:00
jschoubben 16faadfe52 A session is a consumer of a licence, and (node, module) already names one
Work breakdown 1.2. Two sessions run on the control-plane node — the
node's own and the mesh's (novox/hq ADR 0026) — so a machine stopped
being a usable answer to "whose licence is this".

14-model-access.md called per-module-per-machine "a step toward it and
not it", and that is true of a worker: many run on one machine from one
module, so the pair cannot name them apart. It is not true of a session.
The two sessions are two modules — the same mechanism started in
different context roots, and a context root is what a module delivers —
so (node, module) tells them apart and nothing needed adding.

Checked rather than argued: different licences on one machine, each with
its own key, and a session on no licence is not handed the other's.

The third test exists because a fault injection stayed silent. The first
two put the sessions on different licences, so the licence alone
disambiguates and the module argument is never load-bearing — removing
it from the query changed nothing and everything still passed. Two
sessions on the SAME licence is the case that needs the pair to be the
identity: releasing one must leave the other, and a machine-shaped
answer takes both.
2026-08-31 18:41:19 +02:00
jschoubben 981cd4139b Reference the record rather than link across repositories
A relative link out of this repository resolves nowhere on a forge. Every
other reference here names the record in prose.
2026-08-31 17:52:25 +02:00
jschoubben 950ebb9e28 The two halves of an object-store edge, as a readable pair
Manifests for a provider and a consumer, so the contract can be read
rather than only exercised through a lab fixture that stages the grants
by hand.

Checked as a pair rather than separately, because two manifests that
only ever parse alone are two manifests nobody has held against each
other. The test asserts the names match, that each side says where it
wants to be told, and that the consumer contributes the key the
provisioner actually reads.

That last one is the trap worth having a test for: a consumer
contributing "name" — which is exactly what a database consumer
contributes — resolves cleanly, deploys, and then fails on the machine
with "asked for a bucket and did not name it". Nothing in that message
points back at the manifest that caused it. Both mistakes were made
while writing these two files.
2026-08-31 17:52:14 +02:00
jschoubben ed9a30f22d A module can be given a bucket: the provisioner that makes a secret true
Phase 1.1 of the work breakdown. The finding that shaped it came before
any code: **the control plane special-cases nothing.** provides,
requires, contributes and grants are entirely name-agnostic, so asking
for a bucket needed no change to the mesh at all — only a provider that
answers. What was missing was the last step, where something on the
machine turns a delivered secret into a key that works.

Named `s3-bucket` by ADR 0027's test: a consumer's code is written
against the S3 API, and swapping one store for another does not break
it, so the coupling is to the protocol rather than the product — which
is what the substrate design already said about AMQP, S3 and OCI.

Proven on a real store, 7 assertions: a generated secret becomes a
working key; rotation makes the new one work and the old one stop; a
consumer that goes away loses its key; a key nobody here made is left
alone; a manifest naming a credential that was never written is refused;
an unusable bucket name is refused naming the consumer that asked.

**And the one a database does not need.** One PostgreSQL server holds
separate databases and the product enforces the boundary; one object
store holds every bucket behind one endpoint, so a consumer being unable
to reach another's is a policy somebody wrote. A policy granting
arn:aws:s3:::* would pass every other test in the file, so the unit
tests assert what the policy does NOT say.

It drives the vendor's command line rather than an SDK: the admin API
encrypts its request bodies, which is why a separate admin library
exists, and pulling that in would add a system-metrics dependency tree
to a repository with none in order to create a user.
2026-08-31 17:51:12 +02:00
jschoubben 9f5d7a1a83 Name ADR 0013 in the test that defends it
A migration refused when the record and the files disagree is the
property that makes a schema trustworthy months later. The test asserted
it without naming the decision, so an audit of which decisions are
defended could not see it. novox/hq ADR 0017.
2026-08-31 17:33:34 +02:00
jschoubben ee84b624b1 A provision names the engine, because a consumer is coupled to one
Provisions were named after roles: provides "database", requires
"database". Nothing distinguished engines, so a module written against
PostgreSQL could be matched to a provider of SQL Server, resolve as
satisfied, deploy, and fail on its first query — with nothing
connecting that error back to a match made elsewhere by something that
believed it had done its job.

The failure is in the direction that hides. Refusing on ambiguity
exists precisely so this does not happen, and the generic name walked
around it: with one provider of each name nothing is ambiguous, so
nothing is asked.

How it got in: every resolver test had exactly one provider per name,
so no mismatch was expressible and none was caught. The fixtures agreed
with the design — the same fault as the imagined test output in
04-ISSUES/005, at the level of a name.

Refused rather than documented, because the old naming *was* the
documented convention. Providing database/db/sql/sql-database is now a
parse error naming what to write instead.

The rule is about coupling, not specificity everywhere: route and
resolver stay role-named, because a consumer genuinely cannot tell
which proxy answered. novox/hq ADR 0027.
2026-08-31 17:12:46 +02:00
jschoubben ab1dd34d12 The resolver does not ask itself for upstreams
dnsmasq read /etc/resolv.conf to find where to forward. Whatever points a
machine at the mesh writes its own address into that file — so dnsmasq's
upstream was dnsmasq, and every query it could not answer locally looped. Its
receive queue filled with 15KB of them and every lookup on the machine hung,
which is why this arrived as a thirty-second timeout rather than a wrong
answer.

It needs no upstream at all: the asking module routes only the mesh's suffix
here and leaves everything else where the machine already sent it. And it names
none, because choosing one would send every query this machine makes somewhere
nobody agreed to.

Also corrected: the comment claiming it takes only 127.0.0.55. Listening on a
loopback address makes dnsmasq take the rest of loopback with it, 127.0.0.1
included — which is what claiming `the-dns-port` already says, and which the
comment was quietly denying. That is the same comfortable claim as ".54 is
free", in the same file, made twice.
2026-08-31 14:09:50 +02:00
jschoubben a18c3b9d13 needs is now own-secrets, named for whose it is
It sat beside `secrets` — where a *provision's* credential lands on a consumer.
Both were name-to-path, both held something secret, and the names
distinguished them not at all. Reaching for the wrong one parsed cleanly and
failed somewhere else entirely, which is the shape of fault this whole design
exists to prevent, sitting in the manifest format.

The axis that separates them is not how secret they are — both are — but
whose. `secrets` is keyed by the provision it is for and belongs to a
relationship with another machine. `own-secrets` is keyed by a name the module
chose and belongs to nobody else.

A manifest using the old name is told the new one rather than refused with
"unknown field": whoever wrote it knew what they meant, and the mesh knows what
it is called now. An invented key is still refused as one rather than guessed
at.

Found by auditing the 19 manifest fields for whether any could be mistaken for
another. This was the only pair that could — and while checking it, a second
instance of the same collision turned up one layer down: `Manifest.Needs` and
`Resolution.Needs` were different concepts sharing a name in Go. The rename
separates those too.
2026-08-31 13:47:21 +02:00
jschoubben e3a2790acd The resolver answers on an address systemd does not hold
`127.0.0.54` is systemd-resolved's DNS *proxy* stub. The module asserted it was
free, in a comment that read as reasoned — "not .53, that is
systemd-resolved's" — and it was simply wrong: resolved holds both. dnsmasq
could not create the socket and never started.

Nothing in a unit test could have caught it. They checked the module names an
address and that the asking modules point at the same one, and all of that
passed while the daemon could not start. Only a machine knows which addresses
are spare, which is the argument for proving a module that asserts facts about
machines on a machine, before believing the assertions.

So it moves to .55, and says what that is: a convention, not a reservation. If
a future systemd takes it, this line changes and nothing else does.

The tests now derive the address from the serving module and check the two
asking modules agree with it, rather than naming it a fourth time — that fourth
place is the one nobody would think to change.

And the lab assigns `resolved-split-dns` rather than `resolv-conf`: those
machines run systemd-resolved, which owns the file. The two claim the same
thing precisely so the wrong choice is a refusal rather than a fight, and
picking the wrong one was testing the fight.
2026-08-31 13:39:58 +02:00
jschoubben 6eeb10f066 A command opens each store once, not once per machine
Working out what a machine should be reaches the identity context for its
certificate and the licence context for its model access. Both were opened —
and waited on — inside functions called for every node in a push. Two machines
hid it. Fifty would be fifty connect-and-wait cycles for data that does not
change while the push runs.

So a command holds what it has open, and passes it. Each context is opened on
first use rather than up front, because most commands need one and paying to
reach three would be the same waste from the other side.

The contexts stay separate, which is the point: this is one struct holding
three connections to three databases, not one connection to a shared one. No
context reaches another's store, and each still holds only its own credential
(novox/hq ADR 0008).

A pure move again — the gate is green before and after, and no test changed.
2026-08-31 13:35:39 +02:00
jschoubben 8623613704 Split main.go along the seams it already had
2,769 lines and 59 functions, holding command parsing, store opening,
resolution, the board, rotation, licences, builds and status rendering.
Nothing in it was wrong. It grew because appending was always the cheapest next
step, and no single edit was the one that should have been a new file.

That is exactly how novox/hq ADR 0001 records `hal/sdk` reaching 155 files and
34,636 lines — "containing code from every context", with each addition
avoiding a cycle and none of them the mistake. This is the same shape at 8% of
the size, which is why it is worth doing now rather than noting.

Eight files, along boundaries that already existed: what a machine is; the
private network; the catalogue; working out what one machine should be; sending
it; builds; the three questions; and reaching each context's store. main.go
keeps what a main is for — parsing arguments and dispatching.

A pure move. No behaviour changed, no test changed, and the gate is green
before and after — which is the only thing that makes a refactor this size
safe to do in one commit.
2026-08-31 13:30:54 +02:00
jschoubben 68b0d0c10a What answers a requirement is applied before what asked for it
The host does not sort, so the order written here is the order a machine
applies. Selection walks outward from what was assigned, which puts a consumer
before the thing it pulled in — and a service that reads a file another module
writes then starts before the file exists.

It fails, and the next reconcile fixes it. That is the worst shape a fault can
take: what gets remembered is that it works, and nobody looks again. It is
04-ISSUES/013 one level up from where that was found — there, the mesh's own
computed files came after a module's resources; here, a whole module comes
after the one that needed it.

Nothing had hit it because no module until now both required something with
resources of its own and had a resource depending on it. Writing the resolver
module was what made it reachable, and it would have shown up as dnsmasq
failing once on every fresh machine and working ever after.

Unrelated modules keep the order selection gave them — assigned first, then
what they pulled in. That order is meaningful, and reshuffling it would make
every declaration's diff unreadable for no gain.

Two modules requiring each other are both applied rather than refused: a cycle
is not a machine that cannot work, and refusing would make a cooperating pair
impossible to assign.
2026-08-31 12:54:27 +02:00
jschoubben bff893f8af Resolver modules: one that serves, and two ways of deciding what a machine asks
Three manifests and the rule that keeps them apart. Serving and asking are
genuinely different roles, and systemd-resolved can only do the second — it
cannot answer a wildcard, it routes the mesh's suffix to something that can. A
module that treated them as one role could not work, which is the mistake worth
naming rather than discovering.

So `the-dns-port` and `the-resolver-configuration` are two claims. A machine
gets one of each, and two of either is refused by the mesh rather than fought
over on the machine — which is what ADR 0009's table meant by listing resolvers
beside the seat and pid 1. That table names the resource `/etc/resolv.conf`,
which is what it is; a claim is a name in the catalogue's own form, and the
catalogue refuses the path as one.

Neither module knows anything about the machine it is on, which is what lets
them be static manifests: they name `mesh0` and `127.0.0.54`, both chosen by
the mesh, rather than an address only that machine has. Not 127.0.0.1 and not
127.0.0.53 — taking either would be a module claiming something it did not say
it claims.

A service can now reflect a file another module put on the machine, written
`<module>.<id>`. The resolver has to restart when the mesh rewrites the names;
without it, it would serve the names it started with for ever, with every
machine that joined afterwards unreachable and every check passing.
2026-08-31 12:21:58 +02:00
jschoubben 3954157555 The mesh computes every name under a machine, for a resolver to answer
Services are named under the machine they run on — postgres.novox.internal,
plex.ace.internal. The first label is the service and the rest is the node, so
what has to resolve is anything under a node's name. What routes it once it
arrives is a proxy's concern and stays separate.

A hosts file cannot do that. It answers exact names, and a wildcard there would
mean writing down every service in advance — which is the enumeration the
arrangement exists to avoid. novox/hq 08-connectivity named this exact case as
the trigger for needing a resolver rather than a file, and it is the first
thing to meet it.

The mesh writes the data and runs no daemon. A resolver is third-party
software, and third-party software runs on the mesh rather than being of it
(ADR 0001): the mesh has no business shipping one, choosing which one, or
knowing its configuration language. What only the mesh can know is which
machines exist and where they are. A module that runs a resolver requires what
this provides and reads one file, so swapping the daemon changes that module
and nothing here.

Separate from names rather than part of them: a machine with no container
runtime can still have a hosts file, and folding them together would take exact
names away from a machine that cannot run a daemon in order to give it a
wildcard it cannot use either.

A machine with no address is left out. A wildcard pointing at nothing is worse
than no wildcard — every name under it resolves and then hangs, where an
unresolvable name fails at once and says which name it was.
2026-08-31 11:39:21 +02:00
jschoubben 4d67c48342 Every container is given the mesh's names
Internal names are written to the machine's hosts file, which serves the
machine and not what the machine runs: a container gets its own hosts file
holding only its own hostname. So every name the mesh wrote was invisible to
the majority of things that need one — and on the machine it always worked,
which is exactly what made it easy to miss.

It was hit for real in the lab, and worked around by resolving the address on
the machine and passing it in. That workaround is now removed, and its absence
is the assertion.

A file rather than a resolver, which is the decision the mesh already made
about names and this extends rather than overturns: it works on every runtime,
needs no package and has no failure mode of its own. The stated trigger for a
resolver — names that are not one-per-node, service names, wildcards — is
still not met.

Given by the mesh, not chosen by a module: a module that listed the machines
would go stale the day one joins, and one that did not would be a module whose
containers cannot reach anything by name. A container that named its own keeps
them and gets the mesh's beside them.

Only containers, and not the ones on the machine's own network: a runtime
refuses to write a hosts file for those, and a file or a service given the
field is a declaration the host refuses outright — so getting it wrong breaks
the whole machine for something that was never about names.
2026-08-31 11:13:34 +02:00
jschoubben 092109debc A computed module says what its machine opens, so a hub can be filtered
The machine that most needed a firewall was the one that could not have one. A
hub is dialled by every node at other sites and needs its port open; a machine
that is not a hub dials out and needs nothing open. They are the same module,
and `listens` in a manifest is one answer for every machine that runs it — so
the machine a static answer gets wrong is the one facing the public internet.

A generator can now say what it opens, in a second interface rather than a
method on every generator: most have nothing to say here, and requiring an
empty method of each would be a cost paid everywhere for one caller.

The port is the one in the endpoint, which is where the interface takes its
ListenPort from. One source, so a rule set cannot open a port the interface is
not on. Open to everywhere and deliberately: a node at another site is not on
the private network until this port lets it on, so restricting it to the mesh
would be a rule that can never be satisfied by the thing it exists for.

And a generator that cannot say is refused rather than read as silence. Closing
a port on the evidence of a failure to look is how a machine is severed by a
fault somewhere else — and the machine it would sever is the hub, whose only
route to being fixed is the network it just closed.
2026-08-31 10:07:01 +02:00
jschoubben d1c256c2b1 The board says what is waiting too, or it disagrees with the command
Adding "not running what the mesh would send it" to `status` and not to the
board would have left two answers to one question with a person in front of
each — which is the single thing this page's design forbids, introduced by the
change that was supposed to make the question answerable.

The published JSON carries it as well, so the page, the command and anything
built against either say the same thing from the same read. Additive, because
that shape is hard to change once anything is built against it.

Never told stays separate from out of date on the page as it is everywhere
else: same remedy, and nobody has ever asked that machine to be anything.
2026-08-31 06:00:04 +02:00
jschoubben 5a28434ba8 "Behind" means not running what the mesh would send
It meant "failed or refused". So a machine that applied cleanly and whose
declaration has since changed was not behind — and novox/hq ADR 0010's
question, did my change go out?, was answerable exactly for the machines that
broke. For every machine that worked, the answer was silence whether the change
had gone out or not, which is the thing replacing a pipeline was supposed not
to cost.

The mesh now records a digest of what it last sent each machine. A digest
rather than the declaration: it can compute what a machine should be at any
moment, and keeping a copy would be a second account of it able to disagree
with the first. What cannot be recomputed is what was actually sent.

Recorded after the send, not before — a digest kept for something that failed
to send would make the machine look current for a declaration it never
received.

Never told stays separate from out of date. The remedy is the same push and the
situations are not alike: nobody has ever asked that machine to be anything.
And a machine the mesh could not work out is not reported as waiting, because
saying so would invent a comparison — that is `plan`'s answer to give.

`status` says it and `push --behind` sends it, or the flag would know something
the person reading the status does not.
2026-08-31 05:17:21 +02:00
jschoubben dfca21fa55 Keep what a machine said about itself, not just the yes
novox/hq ADR 0009: a capability's presence gates an assignment and its detail
carries a value — seat: card1-DP-1, an architecture, an amount of memory. So
'can this run here' and 'what should it be configured as' are one fact read two
ways, and the mesh was keeping the first read and discarding the second.

The reason an absent capability is absent went the same way, which is the case
a person most needs: 'this machine has no container runtime' is the answer and
'docker is not installed' is why, and only the machine knows why.

`node show` says it back. Never reported and reported nothing stay different
things there — one machine has not run the host, the other ran it and can do
nothing, and those send a person to different places.
2026-08-31 04:53:45 +02:00
jschoubben 92133c340b A board that reads through the same interfaces and holds nothing
novox/hq 03-DESIGN/01-to-be/11-a-board.md, built. The board being replaced is
one service reading every context's database directly — ADR 0008 violated by
the one component with a reason to violate it. The cost is not hypothetical: a
boundary nothing may cross can move, and one thing crossing it is enough to
freeze it. A board that reads the provisioning tables breaks when provisioning
changes them, and the change then gets weighed against the board.

So the three questions are read once, by one function, for all three ways of
saying them — a person's status, its JSON, and this page. Three
implementations of "which machine is not doing what it was told" would be three
chances to disagree.

Refused and failed stay distinct all the way to the page: refused means the
machine is exactly as it was and what is wrong is in what was sent; failed
means it is in a state nobody declared. Different places to fix, so one word
for both would send half the readers to the wrong one.

It stores nothing, changes nothing, and every action it might offer already
exists as a command. A board that cannot reach the mesh says so rather than
rendering an empty page — an empty page says "nothing is wrong" in the one
situation where nobody can know that.

One test earns its place twice: a machine's own words are the whole reason the
page is useful and the one thing on it nobody in this repository wrote, so they
are shown and are not markup.
2026-08-31 04:49:21 +02:00
jschoubben 29b336bb8d Say the remedy beside the problem in status
A status that names what is wrong and not what to do about it makes somebody go
and find the command — and the command is the whole point of having noticed.
The hint existed on one of the two paths that print this.
2026-08-31 03:22:24 +02:00
jschoubben faf5ecd70f One machine's unanswerable requirement does not remove it from the mesh
The pass that answers *what does this node offer* takes a failed resolution to
mean it learned nothing about that node. So refusing an unanswerable
requirement there made the machine disappear — and every other machine was then
told, wrongly, that the two of them shared no private network.

A wrong answer about a machine nobody asked about, caused by a fault on a
third. The lab found it: one module needing a licence that had not been added
yet made two unrelated machines look disconnected.

The second pass still refuses it, where the question is actually being asked.
2026-08-31 03:06:33 +02:00
jschoubben 005b928d36 Test the licence store, and name an unknown licence rather than a constraint
Its own store, its own test database, the same shape every other context has.
Five properties: a key with nobody to seal it to is refused rather than kept
readably; a key is sealed once per holder and the blobs differ because they are
sealed to different machines; a holder recorded afterwards has none and the
existing ones keep theirs; releasing a consumer takes its key; and a licence
nobody recorded is refused by name.

The last was the only one whose message mattered and whose message was not
checked — the database's own foreign-key error is true and mentions a
constraint, which sends somebody to read a schema instead of typing the name
they meant.

Partial sealing now says how far it got. The person holding the key is the only
one who can finish, and running it again knowing what it will do is different
from running it hoping.
2026-08-31 02:59:39 +02:00
jschoubben fbae2f1f9f Build everything behind its source, in one command
novox/hq ADR 0010 replaced a pipeline with a comparison, and named the risk:
losing the question "did my change go out?". The mesh could already answer
which modules are behind their source — and then a person read that list and
retyped each repository, which is a person being the loop, and the loop is the
thing the pipeline was doing before it was taken away.

The mirror of `push --behind`, with the same argument and the same refusal to
combine the two forms: naming a repository and asking which need building are
different requests.

One failing does not stop the others, for the same reason one broken module no
longer blocks a machine's whole declaration: a mesh where one bad repository
holds back nine good ones is a mesh where nobody dares add the tenth.

Each is built from its own recorded ref rather than the commit the mesh
happened to notice — pinning to that would quietly turn a tracked branch into
a pin.
2026-08-31 02:55:10 +02:00
jschoubben 87c6a56b81 Model access is a provision answered by a record, not a machine
novox/hq ADR 0024, gaps 1 and 2. The user's stated requirement, and the first
thing here that no machine can answer: a hosted model is on nobody's node and
is reached over the public internet, so the rule that refuses two ends sharing
no private network must not apply to it.

A licence is a named thing and the name is the operator's — *the personal
account*, *the organisation's* — because the whole point is saying which one a
given consumer uses, and an anonymous credential hanging off a provider cannot
be said. Many to many, so deliberately not a claim: two machines sharing an
account is ordinary rather than a collision.

Gap 2 is the missing verb, *accept*: take a value somebody supplied, seal it to
each holder, discard the plaintext. With the consequence stated rather than
hidden — a holder recorded after the key was supplied has no key and the mesh
cannot make one, so it is refused by name with the remedy, not silently handed
an empty file.

Refusal is felt, as the record warns: a mesh holding three ways to reach a
model refuses every consumer that has not chosen. So the refusal names the
candidates and the exact command. Being right is not the same as being usable.

Gaps 3 and 4 — a consumer that is not a machine, and switching as a reaction
rather than a declaration — remain gaps. Half-building them would put a
conditional in the declaration language, which is what ADR 0024 says plainly to
avoid.

Its own context, with its own store and its own credential: a licence is a
different aggregate from anything inventory owns, and it refers to nodes by
name because that is what crossing a context boundary may carry.
2026-08-31 02:50:38 +02:00
jschoubben d0c0ee8dab A route is a grant, and a provider is told where its consumer is
novox/hq 08-connectivity §3, built. The mirror of a database grant: there the
consumer supplies a name and receives credentials; here it supplies a target
and receives a name. Nothing new in the vocabulary — a route is a provision
like any other.

One field was missing and it is the one that matters for anything reaching
back: a contribution now carries where the mesh says that machine is. A reverse
proxy is told to send traffic to a consumer and has to open a connection, so
without it every provider implementing a provision would have to know how the
mesh names machines — a convention leaking into every module.

The proxy itself is an example, not part of the control plane: the contract is
the file, not this program. It replaces its table whole rather than merging,
because the file is the whole truth about who has a route and merging would
keep serving a name whose module was unassigned — the stale-route fault
08-connectivity lists as open, reintroduced one level down. A name it does not
serve is refused by saying which it does: a route withdrawn and a name that
never existed are different things.
2026-08-31 02:43:19 +02:00