Commit Graph
29 Commits
Author SHA1 Message Date
jschoubben 1314be5282 A file may hold a credential where its content says one belongs
The gap that stopped keycloak and gitea from starting. A granted
credential arrives as a file whose entire content is the password, which
is what a program reading a password file wants — and most programs do
not read one. They read KEY=value, or a JSON document with the token at
an attribute inside it. A module in that position could be handed the
bare value or nothing, and both are useless.

The host has been able to do this all along: content with ${secret:name}
in it, sealed values beside it, substitution on the machine, which is
the only place both halves exist. Nothing filled the values in, so the
hole could be written and never closed and the host refused the file.
That refusal was correct and the feature was unreachable.

A module reaches its own secrets and the credentials it was granted —
both things it wrote in its own manifest — and nothing else. Naming
another module's is refused: two modules on one machine are as separate
as two on different machines, and letting one read the other's
credential by guessing a name would end that to save writing a file.

Filling runs after settings, which is the whole reason it sits where it
does. A setting is how a placeholder gets into a JSON document in the
first place — the desktop client that reads its token from an attribute,
not an environment variable. Before the merge that file's content is
"{}" and asks for nothing.

Tested through Declaration rather than through the helper. Three times
in this repository a test asserted on a helper while the code calling it
was wrong, and each time the injected fault stayed silent. Three faults
injected here — the call removed, the call moved before settings, and
the module boundary widened — each caught by the test meant for it.
2026-09-01 02:28:50 +02:00
jschoubben df62bb57e5 A requirement answered on this machine is still a requirement
novox/hq 04-ISSUES/021. Two modules where one provided what the other
required, on one node, resolved cleanly with zero needs: no credential
was made, the consumer's secret file was never written, and whatever
read it would fail somewhere else entirely. Nothing was refused and
nothing was reported.

The world a node resolves against is every OTHER node, so a provider on
the same machine never became a Needed, and the credential loop walks
Needs. Every step reasonable, the sum a silent gap.

It survived because everything proven until now was cross-machine —
the interesting case for a mesh and the rare one in practice. The first
module to want a database on its own machine was the first real one.

The assumption underneath was that a local consumer needs no credential,
which holds for a process reaching a unix socket where the system can
vouch for the caller. It does not hold for containers, which is how
nearly everything here runs: the consumer reaches the provider over TCP
from its own container and the database asks for a password exactly as
it would from another machine. **The machine stops being a trust
boundary once both ends are containers.**

A brokered provision answered here is now a need naming this node, and
carries what the provider serves — which a local provider never
contributes through the world. A name nothing grants is unchanged: a
shell answered here is answered, and nothing more is owed. Both
directions tested, both injections bite.
2026-09-01 02:15:26 +02:00
jschoubben ee84b624b1 A provision names the engine, because a consumer is coupled to one
Provisions were named after roles: provides "database", requires
"database". Nothing distinguished engines, so a module written against
PostgreSQL could be matched to a provider of SQL Server, resolve as
satisfied, deploy, and fail on its first query — with nothing
connecting that error back to a match made elsewhere by something that
believed it had done its job.

The failure is in the direction that hides. Refusing on ambiguity
exists precisely so this does not happen, and the generic name walked
around it: with one provider of each name nothing is ambiguous, so
nothing is asked.

How it got in: every resolver test had exactly one provider per name,
so no mismatch was expressible and none was caught. The fixtures agreed
with the design — the same fault as the imagined test output in
04-ISSUES/005, at the level of a name.

Refused rather than documented, because the old naming *was* the
documented convention. Providing database/db/sql/sql-database is now a
parse error naming what to write instead.

The rule is about coupling, not specificity everywhere: route and
resolver stay role-named, because a consumer genuinely cannot tell
which proxy answered. novox/hq ADR 0027.
2026-08-31 17:12:46 +02:00
jschoubben a18c3b9d13 needs is now own-secrets, named for whose it is
It sat beside `secrets` — where a *provision's* credential lands on a consumer.
Both were name-to-path, both held something secret, and the names
distinguished them not at all. Reaching for the wrong one parsed cleanly and
failed somewhere else entirely, which is the shape of fault this whole design
exists to prevent, sitting in the manifest format.

The axis that separates them is not how secret they are — both are — but
whose. `secrets` is keyed by the provision it is for and belongs to a
relationship with another machine. `own-secrets` is keyed by a name the module
chose and belongs to nobody else.

A manifest using the old name is told the new one rather than refused with
"unknown field": whoever wrote it knew what they meant, and the mesh knows what
it is called now. An invented key is still refused as one rather than guessed
at.

Found by auditing the 19 manifest fields for whether any could be mistaken for
another. This was the only pair that could — and while checking it, a second
instance of the same collision turned up one layer down: `Manifest.Needs` and
`Resolution.Needs` were different concepts sharing a name in Go. The rename
separates those too.
2026-08-31 13:47:21 +02:00
jschoubben 68b0d0c10a What answers a requirement is applied before what asked for it
The host does not sort, so the order written here is the order a machine
applies. Selection walks outward from what was assigned, which puts a consumer
before the thing it pulled in — and a service that reads a file another module
writes then starts before the file exists.

It fails, and the next reconcile fixes it. That is the worst shape a fault can
take: what gets remembered is that it works, and nobody looks again. It is
04-ISSUES/013 one level up from where that was found — there, the mesh's own
computed files came after a module's resources; here, a whole module comes
after the one that needed it.

Nothing had hit it because no module until now both required something with
resources of its own and had a resource depending on it. Writing the resolver
module was what made it reachable, and it would have shown up as dnsmasq
failing once on every fresh machine and working ever after.

Unrelated modules keep the order selection gave them — assigned first, then
what they pulled in. That order is meaningful, and reshuffling it would make
every declaration's diff unreadable for no gain.

Two modules requiring each other are both applied rather than refused: a cycle
is not a machine that cannot work, and refusing would make a cooperating pair
impossible to assign.
2026-08-31 12:54:27 +02:00
jschoubben bff893f8af Resolver modules: one that serves, and two ways of deciding what a machine asks
Three manifests and the rule that keeps them apart. Serving and asking are
genuinely different roles, and systemd-resolved can only do the second — it
cannot answer a wildcard, it routes the mesh's suffix to something that can. A
module that treated them as one role could not work, which is the mistake worth
naming rather than discovering.

So `the-dns-port` and `the-resolver-configuration` are two claims. A machine
gets one of each, and two of either is refused by the mesh rather than fought
over on the machine — which is what ADR 0009's table meant by listing resolvers
beside the seat and pid 1. That table names the resource `/etc/resolv.conf`,
which is what it is; a claim is a name in the catalogue's own form, and the
catalogue refuses the path as one.

Neither module knows anything about the machine it is on, which is what lets
them be static manifests: they name `mesh0` and `127.0.0.54`, both chosen by
the mesh, rather than an address only that machine has. Not 127.0.0.1 and not
127.0.0.53 — taking either would be a module claiming something it did not say
it claims.

A service can now reflect a file another module put on the machine, written
`<module>.<id>`. The resolver has to restart when the mesh rewrites the names;
without it, it would serve the names it started with for ever, with every
machine that joined afterwards unreachable and every check passing.
2026-08-31 12:21:58 +02:00
jschoubben 4d67c48342 Every container is given the mesh's names
Internal names are written to the machine's hosts file, which serves the
machine and not what the machine runs: a container gets its own hosts file
holding only its own hostname. So every name the mesh wrote was invisible to
the majority of things that need one — and on the machine it always worked,
which is exactly what made it easy to miss.

It was hit for real in the lab, and worked around by resolving the address on
the machine and passing it in. That workaround is now removed, and its absence
is the assertion.

A file rather than a resolver, which is the decision the mesh already made
about names and this extends rather than overturns: it works on every runtime,
needs no package and has no failure mode of its own. The stated trigger for a
resolver — names that are not one-per-node, service names, wildcards — is
still not met.

Given by the mesh, not chosen by a module: a module that listed the machines
would go stale the day one joins, and one that did not would be a module whose
containers cannot reach anything by name. A container that named its own keeps
them and gets the mesh's beside them.

Only containers, and not the ones on the machine's own network: a runtime
refuses to write a hosts file for those, and a file or a service given the
field is a declaration the host refuses outright — so getting it wrong breaks
the whole machine for something that was never about names.
2026-08-31 11:13:34 +02:00
jschoubben 092109debc A computed module says what its machine opens, so a hub can be filtered
The machine that most needed a firewall was the one that could not have one. A
hub is dialled by every node at other sites and needs its port open; a machine
that is not a hub dials out and needs nothing open. They are the same module,
and `listens` in a manifest is one answer for every machine that runs it — so
the machine a static answer gets wrong is the one facing the public internet.

A generator can now say what it opens, in a second interface rather than a
method on every generator: most have nothing to say here, and requiring an
empty method of each would be a cost paid everywhere for one caller.

The port is the one in the endpoint, which is where the interface takes its
ListenPort from. One source, so a rule set cannot open a port the interface is
not on. Open to everywhere and deliberately: a node at another site is not on
the private network until this port lets it on, so restricting it to the mesh
would be a rule that can never be satisfied by the thing it exists for.

And a generator that cannot say is refused rather than read as silence. Closing
a port on the evidence of a failure to look is how a machine is severed by a
fault somewhere else — and the machine it would sever is the hub, whose only
route to being fixed is the network it just closed.
2026-08-31 10:07:01 +02:00
jschoubben faf5ecd70f One machine's unanswerable requirement does not remove it from the mesh
The pass that answers *what does this node offer* takes a failed resolution to
mean it learned nothing about that node. So refusing an unanswerable
requirement there made the machine disappear — and every other machine was then
told, wrongly, that the two of them shared no private network.

A wrong answer about a machine nobody asked about, caused by a fault on a
third. The lab found it: one module needing a licence that had not been added
yet made two unrelated machines look disconnected.

The second pass still refuses it, where the question is actually being asked.
2026-08-31 03:06:33 +02:00
jschoubben 87c6a56b81 Model access is a provision answered by a record, not a machine
novox/hq ADR 0024, gaps 1 and 2. The user's stated requirement, and the first
thing here that no machine can answer: a hosted model is on nobody's node and
is reached over the public internet, so the rule that refuses two ends sharing
no private network must not apply to it.

A licence is a named thing and the name is the operator's — *the personal
account*, *the organisation's* — because the whole point is saying which one a
given consumer uses, and an anonymous credential hanging off a provider cannot
be said. Many to many, so deliberately not a claim: two machines sharing an
account is ordinary rather than a collision.

Gap 2 is the missing verb, *accept*: take a value somebody supplied, seal it to
each holder, discard the plaintext. With the consequence stated rather than
hidden — a holder recorded after the key was supplied has no key and the mesh
cannot make one, so it is refused by name with the remedy, not silently handed
an empty file.

Refusal is felt, as the record warns: a mesh holding three ways to reach a
model refuses every consumer that has not chosen. So the refusal names the
candidates and the exact command. Being right is not the same as being usable.

Gaps 3 and 4 — a consumer that is not a machine, and switching as a reaction
rather than a declaration — remain gaps. Half-building them would put a
conditional in the declaration language, which is what ADR 0024 says plainly to
avoid.

Its own context, with its own store and its own credential: a licence is a
different aggregate from anything inventory owns, and it refers to nodes by
name because that is what crossing a context boundary may carry.
2026-08-31 02:50:38 +02:00
jschoubben d0c0ee8dab A route is a grant, and a provider is told where its consumer is
novox/hq 08-connectivity §3, built. The mirror of a database grant: there the
consumer supplies a name and receives credentials; here it supplies a target
and receives a name. Nothing new in the vocabulary — a route is a provision
like any other.

One field was missing and it is the one that matters for anything reaching
back: a contribution now carries where the mesh says that machine is. A reverse
proxy is told to send traffic to a consumer and has to open a connection, so
without it every provider implementing a provision would have to know how the
mesh names machines — a convention leaking into every module.

The proxy itself is an example, not part of the control plane: the contract is
the file, not this program. It replaces its table whole rather than merging,
because the file is the whole truth about who has a route and merging would
keep serving a name whose module was unassigned — the stale-route fault
08-connectivity lists as open, reintroduced one level down. A name it does not
serve is refused by saying which it does: a route withdrawn and a name that
never existed are different things.
2026-08-31 02:43:19 +02:00
jschoubben 8b2d1bce9d Something answered on this machine is still bound
A binding was skipped when the provider turned out to be on the same node,
reasoning that a file saying "it is on this node" is a fact nobody needs. That
is right about the location and wrong about everything beside it: a binding
also carries what the provider said a consumer must know, which is the port,
and a consumer cannot invent that.

A build machine sharing a node with the registry it pushes to sat in a loop
saying it could not read its own binding. Nothing was wrong with the machine,
the module, the credential or the provision — the file was never written, and
the absence looked exactly like a mistake in the module.

The original intent is kept where it was right: a provision whose provider said
nothing a consumer must know is still not written. A shell is answered here and
there is nothing to say about it. A registry is answered here and the port is
still unguessable.

The address is this machine's name on the private network, or loopback when it
has none — a machine off the network still reaches itself, and a name nothing
resolves is worse than an address that always works.
2026-08-31 01:35:23 +02:00
jschoubben 1eb1b69cae What the mesh computes is applied before what the module declared
The host does not sort — order is stated (novox/hq ADR 0005) — so the order the
mesh writes down is the order a machine applies. Certificates, credentials,
bound files and the rule set were appended after a module's own resources, so a
service or container that depends on one was applied before it existed.

It failed and the next reconcile fixed it, which is why nothing caught it. A
fault that repairs itself on the second attempt is worse than one that does
not: what gets remembered is that it works.

Nothing the mesh computes depends on a module's resources, so putting all of it
first is unconditionally right. Merged after the computed-resources branch,
which replaces a module's resources wholesale and would otherwise discard them.
2026-08-31 00:37:01 +02:00
jschoubben d9bee18d44 A mesh on both address families renders both
nftables matches ip and ip6 separately and one set holding both is a syntax
error, so the file would not load: the service reports a configuration fault
and the machine filters nothing. Also a make target for the builder image,
which the lab now stocks.
2026-08-31 00:34:20 +02:00
jschoubben 48735171eb A machine's filtering is computed from what it was assigned
A rule nobody derives is a rule somebody keeps in step by hand, and five HAL
manifests carry a `scope:` key that reads as a restriction and restricts
nothing. Both halves are closed here.

Manifests are parsed strictly. An unknown key is refused, which is the
discipline the host's declaration parser has always had; `scope:` survived
because nothing rejected it.

A module says what it listens on and who may reach it, and saying from where is
required — a rule with no source is open, and must say so rather than appear to
restrict something. The mesh gathers every assigned module's ports, widens
where two overlap, names every module that wanted each one, and renders one
nftables file per node. What no module declared is closed.

Three things it deliberately does not do: it writes no forward policy, because
what a machine routes is the container runtime's business and dropping there
stops every container on the node; it never flushes the whole ruleset, only
its own table; and it carries no command to load itself, because the link may
not carry an action. A service declares `restart-on` the file instead, which is
the shape that rule leaves.

Also fixes a fault the lab found: certificateFor asked where every node is
without the catalogue, so nothing resolved, every machine looked like it was on
no private network, and every certificate the mesh was asked for was refused
with a reason that was not true. Asking that question without the catalogue is
now refused rather than answered wrongly.
2026-08-31 00:25:12 +02:00
jschoubben 646609c1b2 The mesh certifies names inside it
08-connectivity keeps two authorities apart on purpose: a public one for
names the outside world reaches, and the mesh's own for names only the
mesh knows. Nothing implemented the second, so anything between machines
was plaintext or trust-on-first-use — which the design refuses everywhere
else.

A node now generates a fourth key at enrolment and reports the public
half. A fourth, because a key used for two purposes is one rotation away
from breaking the other: the identity key signs messages to the mesh and
would do for TLS, and reusing it would mean rotating a node's identity
every time its certificate is replaced.

**Nothing secret travels and nothing is sealed.** A certificate authority
says "this name belongs to the holder of this key", so the mesh signs a
public half it cannot use, and the certificate it issues is public. A
module asks for one and is given the certificate and, if it wants,
the mesh's own — the private key is a path to a file the machine already
has, the same arrangement the private network's key uses.

Asserted by verifying rather than inspecting, because a certificate that
parses and does not chain fails at the moment something connects:

- what the mesh issues verifies against the mesh, for the name asked for
- the name is in the subject alternative names, since a certificate
  carrying it only in the common name is refused by every modern client
- it certifies the key the node generated and no other
- another mesh's certificate does not verify, which is the whole point of
  two authorities being separate
- the authority cannot sign another authority — one that could is one
  that can be delegated without anybody deciding to
- two control planes starting together agree on one authority, or a mesh
  has certificates half its machines refuse

Certificates last ten years, which is a choice: a short life needs
something to renew it, and a renewal that fails silently is a mesh that
stops trusting itself on a date nobody wrote down. What makes one
replaceable is that the mesh reissues on demand, not that it expires.
2026-08-31 00:09:13 +02:00
jschoubben 45c3853f4e A consumer that stopped asking is withdrawn
Found by testing removal, which is the half nobody tests.

A grant was emitted for every secret the mesh held, whether or not the
machine still asked for it. So a consumer that was unassigned kept
appearing in its provider's manifest — and the provisioner's rule about
removing what nobody asks for can only fire if the mesh stops asking. The
login would have stayed live for ever, and nothing would have said so.

Skipped where the declaration is built rather than where grants are
gathered, so the rule holds whoever gathers them. No credential file is
written for a withdrawn consumer either, or the provisioner would find a
file its manifest does not mention and have to guess what that means.

The secret itself is deliberately kept. It is sealed and unusable to the
mesh, and a machine that comes back gets what it had — what withdraws the
login is the manifest, which is the thing that reconciles.
2026-08-30 19:17:13 +02:00
jschoubben 15fd70e3ce A module may mirror an image it did not write
A module usually runs software somebody else built: a database module
ships configuration and a provisioner and does not build a database. It
could name the upstream reference directly, and then every machine needs
a route to a public registry and the reference is a tag somebody else can
move — which is what pinning exists to prevent.

So an artifact may be `upstream`: pulled by the reference the module
names, pushed into the mesh's own registry, and pinned by the digest that
registry assigns. This is what the bootstrap already does by hand; it is
now something a module can say.

Refused: an upstream reference with no tag or digest, because what gets
mirrored would be whatever `latest` means today and a module pinned to
that is not pinned. And the rule that a build reads only its own
repository does not apply to it — applying it anyway refused every
reference with a registry host in it, which the test caught.

Written by trying to write a real postgres module and finding it could
not be said. It can now: two directories, two containers pinned by
digest, a superuser password sealed to the machine, and the grants
manifest — six resources from one assignment, all accepted by the host's
own parser.

That exercise also found my manifest wrong rather than the host: a
container declared `restart-on`, which is a service field, and the host
refused it by name. It is right to. A container whose own definition
changes is recreated, and a file it mounts is read by the process inside,
which is that image's business.
2026-08-30 18:28:05 +02:00
jschoubben c37d368f65 A module may need a secret of its own, and the provisioner watches
Two things, both found by trying to write a real postgres module and
discovering it could not be said.

A database has a superuser password, a broker an administrator, a
registry an account. None of them is *for* anybody — they are not the
credential a consumer is given, and the mechanism that hands those out
has a consumer in the middle of it. So a module may declare what it needs
and where to put it, and the mesh generates one per node, seals it, and
reads it no more than it reads any other.

Per node, deliberately: a module running on three machines has three
passwords. One in the manifest instead would put the same secret on every
machine that ever runs it, in a file anybody can read, for ever. Made
once and kept, or a running database would be handed a password it was
not started with; remade when the machine's sealing key changes, like
everything else sealed here.

A need declared and not made is refused rather than skipped, because a
module whose own credential is silently absent starts, fails to
authenticate, and the reason is three layers from the machine reporting
it.

And the provisioner can watch. That is what lets it be a module rather
than a binary somebody places: run once, it needs invoking after every
declaration by a timer or a unit wired to a file; watching, it is an
ordinary long-running service the host already supervises. It polls
rather than watching the filesystem, because the host writes atomically —
the file is replaced, so a watch on the path stops seeing anything after
the first replacement, and a watcher that silently stops working is worse
than a poll. Credentials are compared by digest and never held: this runs
for as long as the machine is up.
2026-08-30 18:22:05 +02:00
jschoubben 44ba100595 A module says what it builds, and the built manifest is a different
document

The manifest in a repository names artifacts; the manifest the mesh holds
names digests. Keeping them the same file would mean a repository
carrying a digest — wrong the moment anybody edits anything, and pinning
a value nobody could have checked.

So a resource says `"artifact": "server"`, and resolving a build rewrites
it to the image reference or the archive's source and digest, removing
the build-time word entirely. The host has never heard of an artifact and
its strict decoder would refuse one, at the worst moment.

A module that builds nothing is ordinary and needs no build section —
most of what a person installs is configuration, and a field that exists
to be left blank is a field nobody fills in correctly.

Refusals worth having:

- an artifact declared and not produced blames THE BUILD, not the
  resource. Both are failures and the remedies are in different places;
  telling somebody to fix the wrong one costs an afternoon. Found by
  injection: the first version's message could not be told apart from
  the resource-level one, so the check was not actually tested.
- a build reads its own repository and nothing else. An input path
  leaving it makes what gets built depend on whatever happens to be on
  the machine building it.
- two artifacts with one name, because a resource naming it could mean
  either.
2026-08-30 03:29:17 +02:00
jschoubben d978712d7f Split resolving from rendering a declaration
resolve.go had grown to 796 lines doing four jobs: working out what a
machine should run, applying settings, collecting contributions, and
placing credentials. They answer different questions — the first is
"what", the rest are "what does that look like as resources" — and one
file doing both is how a thing starts becoming the kernel everything
imports.

Prompted by looking at why HAL's shared library became unmaintainable.
Measured while here, and the shape is the inverse of that one: the large
packages import nothing internal, and only inventory and link compose. A
change to module resolution cannot reach connectivity, because
connectivity does not import it.
2026-08-30 02:54:51 +02:00
jschoubben c3046dcf56 A provider is told who its consumers are, and a reference provisioner
Contributions were node-local, so a mesh-scoped provider — the one case
that most needs them — never heard from its consumers. A database was
given a password and no idea what to create it for.

Cross-node consumers now reach the provider's `receives` file, merged in
with the ones on its own machine: from the provider's side they are the
same thing, and a provider that had to read two lists would read one of
them. Each names the file its credential is in rather than carrying it,
because the mesh discarded the value and could not put it there. The
readable half therefore stays readable.

And examples/postgres-provisioner, which is the last step: it reads what
the host wrote and makes PostgreSQL accept it. Explicitly not part of the
control plane — the control plane decides and never touches a machine.
This runs on the machine and touches it, and a real one ships with the
module that ships PostgreSQL. It lives here because this is where the
contract is defined, written as something that runs so it can be read.

It reconciles rather than applying a change, because it is never told
what changed. Three things that follow, and each is a fault somebody has
shipped:

- the password is set every time, not only on creation, or a rotation
  reports success and changes nothing
- what it made and nobody asks for any more is revoked, or a departed
  consumer keeps a working login for ever
- what it did not make is left alone, or it cannot be run on a database
  that predates it

Proven in the lab against a real PostgreSQL, each assertion confirmed to
fail with the behaviour removed. The suite is in mesh-lab, which also
records the two ways the test itself was wrong first.
2026-08-30 01:31:25 +02:00
jschoubben 20f78cd5f1 Credentials the mesh delivers and cannot read
HAL keeps env vars in the registry, encrypted at rest. Its own tooling
records what that bought and what it did not. `secret_locate` matches by
value rather than by name — because the same password sits in
mesh_provisions, in module_env, in each node's .env in plain text, and
inside every connection string composed from it, and its documentation
says those URL copies "are often the only copies actually in use". And a
query against the encrypted column returns zero rows and proves nothing,
so auditing moved to the decrypted copies on the nodes.

Two faults there, and encryption at rest addresses neither: the control
plane can read what it stores, so a copy of the database is a copy of
every credential; and one secret has many homes with nothing tracking
them.

So here the mesh generates a password, seals it to each end with keys
those nodes generated, stores both blobs, and discards the plaintext. It
cannot read what it holds. Neither can the broker relaying it. And
nothing is composed centrally — a connection string is assembled on the
machine that needs one — so no copy is ever minted in a shape nothing
tracks. `Compromise of a node is compromise of that node` (ADR 0004) is
now true of secrets, not only of identity.

Two files rather than one, because the mesh cannot compose a document
containing a value it discarded: `binds` carries the readable facts,
`secrets` carries the credential alone. The readable half stays readable
in the declaration; the secret half changes only when the secret does,
which makes restart-on precise. The provider gets a directory, one file
per consumer, for the same reason.

It is made once and kept — regenerating per declaration would restart
both ends on every push, and the password a provider was told to create
would never be the one its consumer was given. It is remade when either
end's sealing key changes, and both ends learn the new one in the same
push, so there is no window where half the mesh holds a dead credential.

Two tests found passing for the wrong reason, both caught because their
injection came back clean:

- the provider's copy was asserted non-empty, which reads the same
  whichever column is selected. It now opens the blob with the
  provider's own key.
- RotateSecret deleted and re-created; the re-create was dead, because
  the next read makes one anyway. Removed, and a second path to the same
  act is how two ends come to disagree.

And one real fault: three places built a declaration, and the one behind
`--json` predated credentials, so it silently produced a declaration
missing them — a difference between what `plan` showed and what anything
reading `--json` got. There is one path now.
2026-08-30 00:21:18 +02:00
jschoubben c4782ae2fd An app is told where its database is
Knowing that a machine needs the anchor's database is useless to the
program that needs it unless the program is told. It knew; nothing was
written anywhere it could read.

Two fields, mirroring contributes/receives in the other direction:

  serves: {database: {port: 5432, driver: postgres}}   on the provider
  binds:  {database: /etc/app/database.json}           on the consumer

The provider says what a consumer needs to know; the mesh adds the half
only it has — which machine, and what that machine is called on the
private network. The file says, in itself, that it carries no credential
and why. A missing field looks like a bug; a stated absence looks like a
boundary.

Binding something answered on this machine writes nothing. A file saying
"it is on this node" is a fact nobody needs and one more thing to keep
true.

And two machines that share no private network are refused rather than
wired together. An app here and a database there with no path between
them is a mesh that reports itself configured and does not work — the
failure surfaces as a connection timing out, which is the slowest place
to find it. This is checkable now only because the network became
something a machine is given rather than something it has by having an
address.

One fault, found by running it: working out who is on the private network
resolved the mesh, and resolving the mesh asks who is on the private
network. It hung for two minutes. The comment above the function said not
to do that and the function did it anyway; it now resolves each node
locally, which is the right answer to the question regardless — whether a
machine is on the network depends on what it was assigned, not on what it
takes from others.
2026-08-30 00:02:18 +02:00
jschoubben d4064122d6 Where the answer to a requirement is allowed to live
Two different things were both written `requires`. A shell, a display
server and a private network have to be on the machine that needs them.
A database does not — it runs somewhere and is reached over the network.
Both were answered the same way, so requiring a database installed
PostgreSQL on every machine that ran a web application.

What a module provides now carries a scope, the same idea claims already
use, written short in the ordinary case:

  "provides": ["shell"]
  "provides": [{"name": "database", "scope": "mesh"}]

A mesh-scoped requirement is answered by finding the node already running
it — never by installing it here. Choosing a machine to put a database on
is a decision with consequences, and nothing resolving a web application
should make it silently. With nothing anywhere it refuses and says which
module to assign; with two it refuses and says how to choose.

Choosing is `pin <node> <provision> <from>`, kept per node because that
is the granularity the choice has. A pin at a machine that does not
provide it refuses rather than falling back — a fallback would quietly
move somebody's data. One provider does not overrule a pin either.

Resolving a node now needs to know what the others offer, and working
that out needs them resolved, so it is two passes: the first answers only
what each node offers, the second answers everything. Nothing is ever
declared from the first.

A node's plan says what it takes from elsewhere. It is the only part of a
set that stops working when a different machine goes away, and nothing
else in that output would have said so. It is also where a credential
will hang once there is a mechanism for handing one back.

One test found passing for the wrong reason: it read pins through a join
on the provider, which hides a dangling row whether or not it was cleaned
up. It counts rows now, and bites when the cascade is removed.
2026-08-29 23:51:50 +02:00
jschoubben 5a3a87e8c3 A module can tell its provider what it needs
`requires` said a thing must be there. It never said what to do with it,
so a web application requiring a reverse proxy had nowhere to put "this
name, this port". The two modules that needed it most went round the
outside and opened a connection to the control plane's database, which is
why every node holds a credential to it permanently.

Two fields close it:

  contributes: {reverse-proxy: {host: board, port: 8080}}
  receives:    {reverse-proxy: /etc/traefik/dynamic/mesh.json}

The control plane collects every contribution on a node and writes them
to the path the provider named, ordered by module so the file does not
churn. Contributing to something is requiring it — asking to be published
means a publisher must exist, and a module that had to say both would
eventually say one.

The control plane does not know what a reverse proxy is and does not
write one's configuration. It delivers facts; the module turns them into
whatever it runs. That is why swapping the proxy touches nothing that
publishes through it, and why the host needs no new vocabulary — a
received file is a file.

Settings reach a contribution the same way they reach a file, because a
hostname is exactly what differs between one mesh and the next.

Two things found by running it:

- the file had a `//` header, so it said "do not edit" to a person and
  failed to parse for the program meant to read it. The note is inside
  the document now.
- a provider with no consumers gets an empty file rather than none. It
  cannot otherwise tell "nothing asked for me" from "the mesh never
  wrote it", and those want different responses.

Also `plan <node> --json`, which is how the declaration gets handed to
the host's own parser.
2026-08-29 23:35:43 +02:00
jschoubben 44d134ba25 Networking is a module, and a domain module is how you avoid choosing
Connectivity was code beside the module system doing the module system's
job: every machine with an address was on the private network and there
was no way to keep one off.

A manifest can now say its resources are computed by the control plane,
which is what a peer list needs — it is derived from every machine at
once, so nothing could be written in advance. The network is a module
from there on: assigned, resolved, settled, and absent from a machine
nobody gave it to.

Three modules rather than one, because WireGuard is one VPN of several:

  mesh-wireguard   provides private-network, mesh-addressing
                   claims the-private-network, one per node
  mesh-names       provides name-resolution, requires mesh-addressing
  networking       requires both, and ships no files of its own

The last is the point. Most people want the network up and do not want
to choose a VPN, so `assign networking` takes the only answer to each
requirement silently. The day the catalogue holds a second one there are
two answers, the resolver refuses and names them, and choosing is
assigning the one you want. No flavor field, nothing to configure.

Names left the WireGuard declaration for their own module. They would be
identical over a different private network, and bundling them made one
module out of two things.

Three faults the walk found:

- choosing tailscale still installed WireGuard, dragged back in by the
  names needing the mesh's own addresses. Caught now by a claim: running
  two VPNs is fine, being *the* mesh network is singular.
- a requirement wanted by two modules was reported twice, identically.
- "this mesh has no hub" was reported when the real cause was that a
  node could not be resolved at all. It now names the node and the why.

And a test that asserts the manifests actually shipped, after the claim
went missing from the real one while every test stayed green.
2026-08-29 23:19:32 +02:00
jschoubben 65ade756f2 Settings: changing a module's config without editing its file
Managed files are generated and never edited, so somebody's intention about one
has to live where the generator can see it. It does now: the module ships
defaults, settings go over the top by key, and the file is produced from both.
Upstream can rewrite its half freely and the keys somebody chose survive.

Two layers, both from the start. The mesh's settings for a module, then one
machine's over those. A node that differs is expressed by differing, rather
than by restating everything the rest already say -- which would pin all of it
against future changes for no reason.

An override beats a default and there is nothing to resolve. A setting is a
statement about that key made deliberately; the default was only ever what to
do in the absence of one. So when upstream changes a key somebody has set,
there is no conflict, no merge markers, and nothing to ask.

Nested blocks merge and lists are replaced whole. Setting one field of a block
must not delete its siblings, or every setting would restate the whole block
and pin all of it. A list that merged element-wise could neither be shortened
nor reordered, and there is no correct guess about which element is "the same
one".

A module can keep specific keys for itself -- a socket path its own code
depends on -- and setting one is REFUSED rather than ignored. A setting quietly
dropped is somebody believing they changed something.

Settings that reach nothing are named at the moment they would be used, not
discovered later by the machine not behaving differently.

`plan --files` prints what a machine would be given before it is sent, because
"1 resource" does not tell you whether the merge landed.

One test kept with a note that it does not defend this code: output stability
comes from Go's encoder sorting map keys, so it passes with the merging
removed. Worth having as the thing that would catch a change of encoder, but it
is not evidence about anything written here, and it was checked.
2026-08-29 23:01:09 +02:00
jschoubben 409cd16a09 The mesh decides what a node runs
The gap that has been named at the end of every report for a week. Until now a
declaration came from a person handing over a file; now it comes from what was
assigned, resolved against the catalogue, and the control plane is deciding
rather than relaying.

Everything from the module conversation, built and run on real machines:

  assign laptop i3      -> accepted, brings xorg, because nothing else provides
                           it and there was no choice to make
  assign laptop sway    -> refused: xorg and wayland both claim the-seat
  assign laptop editor  -> refused: three modules provide a shell -- bash,
                           fish, zsh -- choose one
  assign laptop zsh     -> accepted, and the editor's requirement is answered
  bash, fish beside it  -> fine, nothing is claimed

Claims rather than pairwise exclusion, so a third display server would say what
it claims and need no edit to xorg or wayland. Scoped to node, site or mesh:
two DHCP servers at one site collide and at two sites do not, and the mesh-wide
one is the hub said as a claim instead of hard-coded.

Some conflicts cost no manifest field at all. The refusal above names the seat
AND the two files, because the mesh already holds every resource of every
module -- neither i3 nor sway knows the other exists.

Resource identities carry their module, so two modules may both call something
"config" without the second silently replacing the first. What a service
reflects is qualified the same way, or it would name a resource that no longer
exists and stop being restarted when its own configuration changes.

Nothing is sent until every node resolves. A push that configured three and
refused on the fourth would leave the mesh in a state nobody asked for, and the
fourth is exactly where a claim collision appears.

One real flaw found by using it rather than by testing it: assigning zsh did
not satisfy a requirement for a shell. Requirements were counted against the
catalogue without first asking what the set already offers, so "choose one and
assign it" named three modules and then ignored the one you chose. The remedy
was useless and every test passed.
2026-08-29 22:00:06 +02:00