Commit Graph
7 Commits
Author SHA1 Message Date
jschoubben 20f78cd5f1 Credentials the mesh delivers and cannot read
HAL keeps env vars in the registry, encrypted at rest. Its own tooling
records what that bought and what it did not. `secret_locate` matches by
value rather than by name — because the same password sits in
mesh_provisions, in module_env, in each node's .env in plain text, and
inside every connection string composed from it, and its documentation
says those URL copies "are often the only copies actually in use". And a
query against the encrypted column returns zero rows and proves nothing,
so auditing moved to the decrypted copies on the nodes.

Two faults there, and encryption at rest addresses neither: the control
plane can read what it stores, so a copy of the database is a copy of
every credential; and one secret has many homes with nothing tracking
them.

So here the mesh generates a password, seals it to each end with keys
those nodes generated, stores both blobs, and discards the plaintext. It
cannot read what it holds. Neither can the broker relaying it. And
nothing is composed centrally — a connection string is assembled on the
machine that needs one — so no copy is ever minted in a shape nothing
tracks. `Compromise of a node is compromise of that node` (ADR 0004) is
now true of secrets, not only of identity.

Two files rather than one, because the mesh cannot compose a document
containing a value it discarded: `binds` carries the readable facts,
`secrets` carries the credential alone. The readable half stays readable
in the declaration; the secret half changes only when the secret does,
which makes restart-on precise. The provider gets a directory, one file
per consumer, for the same reason.

It is made once and kept — regenerating per declaration would restart
both ends on every push, and the password a provider was told to create
would never be the one its consumer was given. It is remade when either
end's sealing key changes, and both ends learn the new one in the same
push, so there is no window where half the mesh holds a dead credential.

Two tests found passing for the wrong reason, both caught because their
injection came back clean:

- the provider's copy was asserted non-empty, which reads the same
  whichever column is selected. It now opens the blob with the
  provider's own key.
- RotateSecret deleted and re-created; the re-create was dead, because
  the next read makes one anyway. Removed, and a second path to the same
  act is how two ends come to disagree.

And one real fault: three places built a declaration, and the one behind
`--json` predated credentials, so it silently produced a declaration
missing them — a difference between what `plan` showed and what anything
reading `--json` got. There is one path now.
2026-08-30 00:21:18 +02:00
jschoubben c4782ae2fd An app is told where its database is
Knowing that a machine needs the anchor's database is useless to the
program that needs it unless the program is told. It knew; nothing was
written anywhere it could read.

Two fields, mirroring contributes/receives in the other direction:

  serves: {database: {port: 5432, driver: postgres}}   on the provider
  binds:  {database: /etc/app/database.json}           on the consumer

The provider says what a consumer needs to know; the mesh adds the half
only it has — which machine, and what that machine is called on the
private network. The file says, in itself, that it carries no credential
and why. A missing field looks like a bug; a stated absence looks like a
boundary.

Binding something answered on this machine writes nothing. A file saying
"it is on this node" is a fact nobody needs and one more thing to keep
true.

And two machines that share no private network are refused rather than
wired together. An app here and a database there with no path between
them is a mesh that reports itself configured and does not work — the
failure surfaces as a connection timing out, which is the slowest place
to find it. This is checkable now only because the network became
something a machine is given rather than something it has by having an
address.

One fault, found by running it: working out who is on the private network
resolved the mesh, and resolving the mesh asks who is on the private
network. It hung for two minutes. The comment above the function said not
to do that and the function did it anyway; it now resolves each node
locally, which is the right answer to the question regardless — whether a
machine is on the network depends on what it was assigned, not on what it
takes from others.
2026-08-30 00:02:18 +02:00
jschoubben d4064122d6 Where the answer to a requirement is allowed to live
Two different things were both written `requires`. A shell, a display
server and a private network have to be on the machine that needs them.
A database does not — it runs somewhere and is reached over the network.
Both were answered the same way, so requiring a database installed
PostgreSQL on every machine that ran a web application.

What a module provides now carries a scope, the same idea claims already
use, written short in the ordinary case:

  "provides": ["shell"]
  "provides": [{"name": "database", "scope": "mesh"}]

A mesh-scoped requirement is answered by finding the node already running
it — never by installing it here. Choosing a machine to put a database on
is a decision with consequences, and nothing resolving a web application
should make it silently. With nothing anywhere it refuses and says which
module to assign; with two it refuses and says how to choose.

Choosing is `pin <node> <provision> <from>`, kept per node because that
is the granularity the choice has. A pin at a machine that does not
provide it refuses rather than falling back — a fallback would quietly
move somebody's data. One provider does not overrule a pin either.

Resolving a node now needs to know what the others offer, and working
that out needs them resolved, so it is two passes: the first answers only
what each node offers, the second answers everything. Nothing is ever
declared from the first.

A node's plan says what it takes from elsewhere. It is the only part of a
set that stops working when a different machine goes away, and nothing
else in that output would have said so. It is also where a credential
will hang once there is a mechanism for handing one back.

One test found passing for the wrong reason: it read pins through a join
on the provider, which hides a dangling row whether or not it was cleaned
up. It counts rows now, and bites when the cascade is removed.
2026-08-29 23:51:50 +02:00
jschoubben 5a3a87e8c3 A module can tell its provider what it needs
`requires` said a thing must be there. It never said what to do with it,
so a web application requiring a reverse proxy had nowhere to put "this
name, this port". The two modules that needed it most went round the
outside and opened a connection to the control plane's database, which is
why every node holds a credential to it permanently.

Two fields close it:

  contributes: {reverse-proxy: {host: board, port: 8080}}
  receives:    {reverse-proxy: /etc/traefik/dynamic/mesh.json}

The control plane collects every contribution on a node and writes them
to the path the provider named, ordered by module so the file does not
churn. Contributing to something is requiring it — asking to be published
means a publisher must exist, and a module that had to say both would
eventually say one.

The control plane does not know what a reverse proxy is and does not
write one's configuration. It delivers facts; the module turns them into
whatever it runs. That is why swapping the proxy touches nothing that
publishes through it, and why the host needs no new vocabulary — a
received file is a file.

Settings reach a contribution the same way they reach a file, because a
hostname is exactly what differs between one mesh and the next.

Two things found by running it:

- the file had a `//` header, so it said "do not edit" to a person and
  failed to parse for the program meant to read it. The note is inside
  the document now.
- a provider with no consumers gets an empty file rather than none. It
  cannot otherwise tell "nothing asked for me" from "the mesh never
  wrote it", and those want different responses.

Also `plan <node> --json`, which is how the declaration gets handed to
the host's own parser.
2026-08-29 23:35:43 +02:00
jschoubben 44d134ba25 Networking is a module, and a domain module is how you avoid choosing
Connectivity was code beside the module system doing the module system's
job: every machine with an address was on the private network and there
was no way to keep one off.

A manifest can now say its resources are computed by the control plane,
which is what a peer list needs — it is derived from every machine at
once, so nothing could be written in advance. The network is a module
from there on: assigned, resolved, settled, and absent from a machine
nobody gave it to.

Three modules rather than one, because WireGuard is one VPN of several:

  mesh-wireguard   provides private-network, mesh-addressing
                   claims the-private-network, one per node
  mesh-names       provides name-resolution, requires mesh-addressing
  networking       requires both, and ships no files of its own

The last is the point. Most people want the network up and do not want
to choose a VPN, so `assign networking` takes the only answer to each
requirement silently. The day the catalogue holds a second one there are
two answers, the resolver refuses and names them, and choosing is
assigning the one you want. No flavor field, nothing to configure.

Names left the WireGuard declaration for their own module. They would be
identical over a different private network, and bundling them made one
module out of two things.

Three faults the walk found:

- choosing tailscale still installed WireGuard, dragged back in by the
  names needing the mesh's own addresses. Caught now by a claim: running
  two VPNs is fine, being *the* mesh network is singular.
- a requirement wanted by two modules was reported twice, identically.
- "this mesh has no hub" was reported when the real cause was that a
  node could not be resolved at all. It now names the node and the why.

And a test that asserts the manifests actually shipped, after the claim
went missing from the real one while every test stayed green.
2026-08-29 23:19:32 +02:00
jschoubben 65ade756f2 Settings: changing a module's config without editing its file
Managed files are generated and never edited, so somebody's intention about one
has to live where the generator can see it. It does now: the module ships
defaults, settings go over the top by key, and the file is produced from both.
Upstream can rewrite its half freely and the keys somebody chose survive.

Two layers, both from the start. The mesh's settings for a module, then one
machine's over those. A node that differs is expressed by differing, rather
than by restating everything the rest already say -- which would pin all of it
against future changes for no reason.

An override beats a default and there is nothing to resolve. A setting is a
statement about that key made deliberately; the default was only ever what to
do in the absence of one. So when upstream changes a key somebody has set,
there is no conflict, no merge markers, and nothing to ask.

Nested blocks merge and lists are replaced whole. Setting one field of a block
must not delete its siblings, or every setting would restate the whole block
and pin all of it. A list that merged element-wise could neither be shortened
nor reordered, and there is no correct guess about which element is "the same
one".

A module can keep specific keys for itself -- a socket path its own code
depends on -- and setting one is REFUSED rather than ignored. A setting quietly
dropped is somebody believing they changed something.

Settings that reach nothing are named at the moment they would be used, not
discovered later by the machine not behaving differently.

`plan --files` prints what a machine would be given before it is sent, because
"1 resource" does not tell you whether the merge landed.

One test kept with a note that it does not defend this code: output stability
comes from Go's encoder sorting map keys, so it passes with the merging
removed. Worth having as the thing that would catch a change of encoder, but it
is not evidence about anything written here, and it was checked.
2026-08-29 23:01:09 +02:00
jschoubben 409cd16a09 The mesh decides what a node runs
The gap that has been named at the end of every report for a week. Until now a
declaration came from a person handing over a file; now it comes from what was
assigned, resolved against the catalogue, and the control plane is deciding
rather than relaying.

Everything from the module conversation, built and run on real machines:

  assign laptop i3      -> accepted, brings xorg, because nothing else provides
                           it and there was no choice to make
  assign laptop sway    -> refused: xorg and wayland both claim the-seat
  assign laptop editor  -> refused: three modules provide a shell -- bash,
                           fish, zsh -- choose one
  assign laptop zsh     -> accepted, and the editor's requirement is answered
  bash, fish beside it  -> fine, nothing is claimed

Claims rather than pairwise exclusion, so a third display server would say what
it claims and need no edit to xorg or wayland. Scoped to node, site or mesh:
two DHCP servers at one site collide and at two sites do not, and the mesh-wide
one is the hub said as a claim instead of hard-coded.

Some conflicts cost no manifest field at all. The refusal above names the seat
AND the two files, because the mesh already holds every resource of every
module -- neither i3 nor sway knows the other exists.

Resource identities carry their module, so two modules may both call something
"config" without the second silently replacing the first. What a service
reflects is qualified the same way, or it would name a resource that no longer
exists and stop being restarted when its own configuration changes.

Nothing is sent until every node resolves. A push that configured three and
refused on the fourth would leave the mesh in a state nobody asked for, and the
fourth is exactly where a claim collision appears.

One real flaw found by using it rather than by testing it: assigning zsh did
not satisfy a requirement for a shell. Requirements were counted against the
catalogue without first asking what the set already offers, so "choose one and
assign it" named three modules and then ignored the one you chose. The remedy
was useless and every test passed.
2026-08-29 22:00:06 +02:00