b70f0d626aea69cc5193b6bb2076ad73b4324017
20
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b70f0d626a |
What review found in the port machinery, fixed
Three faults, one file split. All from reading, all verified to bite. Unassign now releases the module's ports. ReleasePorts existed, said "for when it is unassigned" in its own comment, and was called by nothing — so a fixed port stayed claimed in the name of a module that was gone, and the next module needing it was refused by a ghost. Kept-once-chosen is a promise about a module that is still here. MachineSide reads addressed mappings. "127.0.0.1:8080:80" was split at the first colon, "127.0.0.1" failed to parse as a port, and the mapping was silently skipped — putting the filter back on the declared port, the exact fault the function was written to end. The machine side is the second-from-last part, which is the reading the host already applies, and the substrate bundle writes that shape today. An allocation race answers in the mesh's words. Two concurrent picks of the same port used to surface as a Postgres constraint violation, verbatim. The table has two keys, so the collision is one of two facts: the racer was this same assignment — then its answer is the answer, kept-once-chosen does not care who chose — or another module took the machine port, and an unfixed pick is simply made again against the moved free list. A fixed port that lost the race is refused by name. Told apart by re-reading the row, not by the constraint's name, so this does not couple to the migration's spelling. And the artifact-store cycle tests moved to bootstrap_cycle_test.go; machineside_test.go had quietly become three subjects. |
||
|
|
be62f49eab |
What provides the artifact store cannot be delivered through it
Refused where it is written: a module that provides `artifact-store` and also builds artifacts asks the mesh to put an artifact into the thing that artifact is needed to create. Building publishes to the store, and the builder will not start without one — "a built artifact nobody can fetch is not built". This is the question the substrate record asks of every candidate: can it grant itself the thing it provides? The store cannot create its own database, the broker cannot create its own virtual host, and a registry cannot grant itself a repository. The first two are why they are in the bundle. This is the same sentence, unenforced. So such a module names its image, exactly as the bundle names the three a first node starts from. One that wants an interface or a tool server beside it is a second module, mirrored the ordinary way once the first is running — a real limit, and better said here than discovered on a mesh new enough that nobody is watching it. Refused at the manifest because the alternative is a build that never returns. The provision name is now a constant. A string compared in one place is a convention; a string a rule turns on is a fact. |
||
|
|
83c6a2f244 |
Withdraw the mount check: it refuses the builder
The rule was right about data and wrong about everything else. The builder mounts the container runtime's socket, which is not its data, does not belong to it, and must not be declared as one of its directories — and the check refused the builder's own manifest. Caught by the lab, though not honestly: the run was already going when this went in, so the builder binary was rebuilt mid-run with the check compiled into it and the failure was mine, not the mesh's. Confirmed against the manifest directly rather than inferred from the log. What it was protecting is real and stands — the fourteen mounts are all declared. But enforcing it needs a way to tell "the directory my data lives in" from "a machine facility I was granted", and the mesh has no vocabulary for the second. `capabilities` is the closest thing and does not name paths. That is a design decision, so it goes back to 04-ISSUES/026 rather than being invented here to make a check pass. The check that every real manifest still parses is kept. It costs nothing and it is how the next attempt at this finds out sooner. |
||
|
|
53eb000a84 |
A container may not mount a path the module never declared
Closes the half of 04-ISSUES/026 that would otherwise come back. The fourteen mounts across the forge, the mail system, the store and the object store are all declared now — but nothing said they had to be, so they were right by coincidence and the next volume added would not be. A bind mount whose source does not exist is created by the container runtime, as root, with a mode it picks. So `owner` and `mode` — which exist precisely so a module can say who its data belongs to — were silently not applied to the only directories holding data. And the rule written for exactly this case did not reach them. A directory the mesh declared and no longer wants is kept, not removed, when it holds anything the mesh did not put there (ADR 0030). That is the answer to *what happens to my data when a module goes away*, and it is written in terms of declared directories: an undeclared one sits outside it, because the mesh does not know it is there. Refused where it is written rather than on the machine, which cannot tell the difference — by the time the host sees the mount it is being asked to make a directory, which it is perfectly able to do. The fault is in the manifest, so it is named at the manifest. Same argument as the action refusal directly above it. A path under a declared directory counts as declared, as do the files a module already names: its own secrets, its grants, what it receives. Every real manifest is checked to still parse, and the refusal bites. |
||
|
|
c67f836185 |
The mesh may only move a port it actually publishes
The lab caught this: a module declaring a port and running no container had its rule set opened on 20000 while its service sat on 9101. The firewall reported success and blocked the thing it was told to admit, which is the precise failure the filtering comment warns about, arrived at from the other side. Assignment was applied to every declared port. But a container's mapping is the thing that translates, and where there is none the software binds what it binds — the mesh choosing a number does not move the service, it only makes the mesh wrong about where it is. The declaration side already knew this: publishedOn rewrites container ports and nothing else. Filtering did not, so the two disagreed about the same fact. MachineSide is now the one derivation both follow. It also fixes a second case nobody had hit yet: a mapping the manifest wrote itself, like the mail system's 7080:80. That is passed through untouched when composing, so assigning it a machine port would have opened a rule on a port the container does not publish. Either side of such a mapping now names it, and the host side is the answer — a module may read `listens` as what its software binds or as what the machine exposes, and both readings want the same number. Recorded either way, assigned or not: the map means where this module's port is on this machine, and every reader needs that answer regardless of who chose it. Tests bite — making it always assignable reproduces the lab failure. |
||
|
|
1f5b70a995 |
The mesh assigns the port, and a module says it once
novox/hq ADR 0038. A module cannot choose a port: it is written once and assigned anywhere, so any number it picks is a guess about a machine it has never seen. A database module met the mesh's own store on 5432 and was told, by a container runtime three layers down, that the port was already allocated. The number used to appear three times in every module — the rule set, what a consumer is told, and what the runtime publishes — agreeing only because one person wrote all three. Now it appears once, in `listens`, and the other two are derived: the container publishes `20000:5432`, the consumer is told 20000, and the rule set opens 20000. An assignment is made once and kept, as a credential is. A port that moved on every declaration would restart both ends each time and hand a consumer a number that was true when it was read. Ports the protocol fixes — mail on 25, submission on 587, DNS on 53 — say so, and are then claims: one holder per machine, and the second is refused by name at assignment. That is the mechanism the mesh already has for what is singular on a machine, pointed at ports. A mapping written the long way is left exactly as it is. Some things must be pinned by hand, and quietly overruling somebody who wrote both halves would be worse than not offering the short form. Still open, and known: the substrate is not a module, so the mesh has never heard of its own store and cannot yet assign around it. That is what 028 will still be about after this. |
||
|
|
96f90ab986 |
Refuse an action where it was written, not on the machine
A module may not declare an action: the link may not carry a command to run, and that bound is what limits a compromised control plane to shapes it cannot turn into arbitrary code (novox/hq ADR 0005). The host enforces it, correctly and in the right place. But a module's resources reach a machine over the link, so a manifest carrying an action was accepted here, stored, resolved, planned and pushed — and refused on the machine, in the host's log, with nothing connecting it back to the manifest that caused it. The rule held. It was just unusable, which is the same shape as the network shape earlier today: the refusal was right, arrived far from its cause, and nobody was reading the log. The refusal names the rule and what to do instead, because "you may not" with no alternative is where a module author stops. Found while checking a claim I had written in the coverage document — that a module cannot declare one. It could; it just could not deliver it. The document is corrected. |
||
|
|
ee84b624b1 |
A provision names the engine, because a consumer is coupled to one
Provisions were named after roles: provides "database", requires "database". Nothing distinguished engines, so a module written against PostgreSQL could be matched to a provider of SQL Server, resolve as satisfied, deploy, and fail on its first query — with nothing connecting that error back to a match made elsewhere by something that believed it had done its job. The failure is in the direction that hides. Refusing on ambiguity exists precisely so this does not happen, and the generic name walked around it: with one provider of each name nothing is ambiguous, so nothing is asked. How it got in: every resolver test had exactly one provider per name, so no mismatch was expressible and none was caught. The fixtures agreed with the design — the same fault as the imagined test output in 04-ISSUES/005, at the level of a name. Refused rather than documented, because the old naming *was* the documented convention. Providing database/db/sql/sql-database is now a parse error naming what to write instead. The rule is about coupling, not specificity everywhere: route and resolver stay role-named, because a consumer genuinely cannot tell which proxy answered. novox/hq ADR 0027. |
||
|
|
a18c3b9d13 |
needs is now own-secrets, named for whose it is
It sat beside `secrets` — where a *provision's* credential lands on a consumer. Both were name-to-path, both held something secret, and the names distinguished them not at all. Reaching for the wrong one parsed cleanly and failed somewhere else entirely, which is the shape of fault this whole design exists to prevent, sitting in the manifest format. The axis that separates them is not how secret they are — both are — but whose. `secrets` is keyed by the provision it is for and belongs to a relationship with another machine. `own-secrets` is keyed by a name the module chose and belongs to nobody else. A manifest using the old name is told the new one rather than refused with "unknown field": whoever wrote it knew what they meant, and the mesh knows what it is called now. An invented key is still refused as one rather than guessed at. Found by auditing the 19 manifest fields for whether any could be mistaken for another. This was the only pair that could — and while checking it, a second instance of the same collision turned up one layer down: `Manifest.Needs` and `Resolution.Needs` were different concepts sharing a name in Go. The rename separates those too. |
||
|
|
48735171eb |
A machine's filtering is computed from what it was assigned
A rule nobody derives is a rule somebody keeps in step by hand, and five HAL manifests carry a `scope:` key that reads as a restriction and restricts nothing. Both halves are closed here. Manifests are parsed strictly. An unknown key is refused, which is the discipline the host's declaration parser has always had; `scope:` survived because nothing rejected it. A module says what it listens on and who may reach it, and saying from where is required — a rule with no source is open, and must say so rather than appear to restrict something. The mesh gathers every assigned module's ports, widens where two overlap, names every module that wanted each one, and renders one nftables file per node. What no module declared is closed. Three things it deliberately does not do: it writes no forward policy, because what a machine routes is the container runtime's business and dropping there stops every container on the node; it never flushes the whole ruleset, only its own table; and it carries no command to load itself, because the link may not carry an action. A service declares `restart-on` the file instead, which is the shape that rule leaves. Also fixes a fault the lab found: certificateFor asked where every node is without the catalogue, so nothing resolved, every machine looked like it was on no private network, and every certificate the mesh was asked for was refused with a reason that was not true. Asking that question without the catalogue is now refused rather than answered wrongly. |
||
|
|
646609c1b2 |
The mesh certifies names inside it
08-connectivity keeps two authorities apart on purpose: a public one for names the outside world reaches, and the mesh's own for names only the mesh knows. Nothing implemented the second, so anything between machines was plaintext or trust-on-first-use — which the design refuses everywhere else. A node now generates a fourth key at enrolment and reports the public half. A fourth, because a key used for two purposes is one rotation away from breaking the other: the identity key signs messages to the mesh and would do for TLS, and reusing it would mean rotating a node's identity every time its certificate is replaced. **Nothing secret travels and nothing is sealed.** A certificate authority says "this name belongs to the holder of this key", so the mesh signs a public half it cannot use, and the certificate it issues is public. A module asks for one and is given the certificate and, if it wants, the mesh's own — the private key is a path to a file the machine already has, the same arrangement the private network's key uses. Asserted by verifying rather than inspecting, because a certificate that parses and does not chain fails at the moment something connects: - what the mesh issues verifies against the mesh, for the name asked for - the name is in the subject alternative names, since a certificate carrying it only in the common name is refused by every modern client - it certifies the key the node generated and no other - another mesh's certificate does not verify, which is the whole point of two authorities being separate - the authority cannot sign another authority — one that could is one that can be delegated without anybody deciding to - two control planes starting together agree on one authority, or a mesh has certificates half its machines refuse Certificates last ten years, which is a choice: a short life needs something to renew it, and a renewal that fails silently is a mesh that stops trusting itself on a date nobody wrote down. What makes one replaceable is that the mesh reissues on demand, not that it expires. |
||
|
|
15fd70e3ce |
A module may mirror an image it did not write
A module usually runs software somebody else built: a database module ships configuration and a provisioner and does not build a database. It could name the upstream reference directly, and then every machine needs a route to a public registry and the reference is a tag somebody else can move — which is what pinning exists to prevent. So an artifact may be `upstream`: pulled by the reference the module names, pushed into the mesh's own registry, and pinned by the digest that registry assigns. This is what the bootstrap already does by hand; it is now something a module can say. Refused: an upstream reference with no tag or digest, because what gets mirrored would be whatever `latest` means today and a module pinned to that is not pinned. And the rule that a build reads only its own repository does not apply to it — applying it anyway refused every reference with a registry host in it, which the test caught. Written by trying to write a real postgres module and finding it could not be said. It can now: two directories, two containers pinned by digest, a superuser password sealed to the machine, and the grants manifest — six resources from one assignment, all accepted by the host's own parser. That exercise also found my manifest wrong rather than the host: a container declared `restart-on`, which is a service field, and the host refused it by name. It is right to. A container whose own definition changes is recreated, and a file it mounts is read by the process inside, which is that image's business. |
||
|
|
c37d368f65 |
A module may need a secret of its own, and the provisioner watches
Two things, both found by trying to write a real postgres module and discovering it could not be said. A database has a superuser password, a broker an administrator, a registry an account. None of them is *for* anybody — they are not the credential a consumer is given, and the mechanism that hands those out has a consumer in the middle of it. So a module may declare what it needs and where to put it, and the mesh generates one per node, seals it, and reads it no more than it reads any other. Per node, deliberately: a module running on three machines has three passwords. One in the manifest instead would put the same secret on every machine that ever runs it, in a file anybody can read, for ever. Made once and kept, or a running database would be handed a password it was not started with; remade when the machine's sealing key changes, like everything else sealed here. A need declared and not made is refused rather than skipped, because a module whose own credential is silently absent starts, fails to authenticate, and the reason is three layers from the machine reporting it. And the provisioner can watch. That is what lets it be a module rather than a binary somebody places: run once, it needs invoking after every declaration by a timer or a unit wired to a file; watching, it is an ordinary long-running service the host already supervises. It polls rather than watching the filesystem, because the host writes atomically — the file is replaced, so a watch on the path stops seeing anything after the first replacement, and a watcher that silently stops working is worse than a poll. Credentials are compared by digest and never held: this runs for as long as the machine is up. |
||
|
|
44ba100595 |
A module says what it builds, and the built manifest is a different
document The manifest in a repository names artifacts; the manifest the mesh holds names digests. Keeping them the same file would mean a repository carrying a digest — wrong the moment anybody edits anything, and pinning a value nobody could have checked. So a resource says `"artifact": "server"`, and resolving a build rewrites it to the image reference or the archive's source and digest, removing the build-time word entirely. The host has never heard of an artifact and its strict decoder would refuse one, at the worst moment. A module that builds nothing is ordinary and needs no build section — most of what a person installs is configuration, and a field that exists to be left blank is a field nobody fills in correctly. Refusals worth having: - an artifact declared and not produced blames THE BUILD, not the resource. Both are failures and the remedies are in different places; telling somebody to fix the wrong one costs an afternoon. Found by injection: the first version's message could not be told apart from the resource-level one, so the check was not actually tested. - a build reads its own repository and nothing else. An input path leaving it makes what gets built depend on whatever happens to be on the machine building it. - two artifacts with one name, because a resource naming it could mean either. |
||
|
|
20f78cd5f1 |
Credentials the mesh delivers and cannot read
HAL keeps env vars in the registry, encrypted at rest. Its own tooling records what that bought and what it did not. `secret_locate` matches by value rather than by name — because the same password sits in mesh_provisions, in module_env, in each node's .env in plain text, and inside every connection string composed from it, and its documentation says those URL copies "are often the only copies actually in use". And a query against the encrypted column returns zero rows and proves nothing, so auditing moved to the decrypted copies on the nodes. Two faults there, and encryption at rest addresses neither: the control plane can read what it stores, so a copy of the database is a copy of every credential; and one secret has many homes with nothing tracking them. So here the mesh generates a password, seals it to each end with keys those nodes generated, stores both blobs, and discards the plaintext. It cannot read what it holds. Neither can the broker relaying it. And nothing is composed centrally — a connection string is assembled on the machine that needs one — so no copy is ever minted in a shape nothing tracks. `Compromise of a node is compromise of that node` (ADR 0004) is now true of secrets, not only of identity. Two files rather than one, because the mesh cannot compose a document containing a value it discarded: `binds` carries the readable facts, `secrets` carries the credential alone. The readable half stays readable in the declaration; the secret half changes only when the secret does, which makes restart-on precise. The provider gets a directory, one file per consumer, for the same reason. It is made once and kept — regenerating per declaration would restart both ends on every push, and the password a provider was told to create would never be the one its consumer was given. It is remade when either end's sealing key changes, and both ends learn the new one in the same push, so there is no window where half the mesh holds a dead credential. Two tests found passing for the wrong reason, both caught because their injection came back clean: - the provider's copy was asserted non-empty, which reads the same whichever column is selected. It now opens the blob with the provider's own key. - RotateSecret deleted and re-created; the re-create was dead, because the next read makes one anyway. Removed, and a second path to the same act is how two ends come to disagree. And one real fault: three places built a declaration, and the one behind `--json` predated credentials, so it silently produced a declaration missing them — a difference between what `plan` showed and what anything reading `--json` got. There is one path now. |
||
|
|
c4782ae2fd |
An app is told where its database is
Knowing that a machine needs the anchor's database is useless to the
program that needs it unless the program is told. It knew; nothing was
written anywhere it could read.
Two fields, mirroring contributes/receives in the other direction:
serves: {database: {port: 5432, driver: postgres}} on the provider
binds: {database: /etc/app/database.json} on the consumer
The provider says what a consumer needs to know; the mesh adds the half
only it has — which machine, and what that machine is called on the
private network. The file says, in itself, that it carries no credential
and why. A missing field looks like a bug; a stated absence looks like a
boundary.
Binding something answered on this machine writes nothing. A file saying
"it is on this node" is a fact nobody needs and one more thing to keep
true.
And two machines that share no private network are refused rather than
wired together. An app here and a database there with no path between
them is a mesh that reports itself configured and does not work — the
failure surfaces as a connection timing out, which is the slowest place
to find it. This is checkable now only because the network became
something a machine is given rather than something it has by having an
address.
One fault, found by running it: working out who is on the private network
resolved the mesh, and resolving the mesh asks who is on the private
network. It hung for two minutes. The comment above the function said not
to do that and the function did it anyway; it now resolves each node
locally, which is the right answer to the question regardless — whether a
machine is on the network depends on what it was assigned, not on what it
takes from others.
|
||
|
|
d4064122d6 |
Where the answer to a requirement is allowed to live
Two different things were both written `requires`. A shell, a display
server and a private network have to be on the machine that needs them.
A database does not — it runs somewhere and is reached over the network.
Both were answered the same way, so requiring a database installed
PostgreSQL on every machine that ran a web application.
What a module provides now carries a scope, the same idea claims already
use, written short in the ordinary case:
"provides": ["shell"]
"provides": [{"name": "database", "scope": "mesh"}]
A mesh-scoped requirement is answered by finding the node already running
it — never by installing it here. Choosing a machine to put a database on
is a decision with consequences, and nothing resolving a web application
should make it silently. With nothing anywhere it refuses and says which
module to assign; with two it refuses and says how to choose.
Choosing is `pin <node> <provision> <from>`, kept per node because that
is the granularity the choice has. A pin at a machine that does not
provide it refuses rather than falling back — a fallback would quietly
move somebody's data. One provider does not overrule a pin either.
Resolving a node now needs to know what the others offer, and working
that out needs them resolved, so it is two passes: the first answers only
what each node offers, the second answers everything. Nothing is ever
declared from the first.
A node's plan says what it takes from elsewhere. It is the only part of a
set that stops working when a different machine goes away, and nothing
else in that output would have said so. It is also where a credential
will hang once there is a mechanism for handing one back.
One test found passing for the wrong reason: it read pins through a join
on the provider, which hides a dangling row whether or not it was cleaned
up. It counts rows now, and bites when the cascade is removed.
|
||
|
|
5a3a87e8c3 |
A module can tell its provider what it needs
`requires` said a thing must be there. It never said what to do with it,
so a web application requiring a reverse proxy had nowhere to put "this
name, this port". The two modules that needed it most went round the
outside and opened a connection to the control plane's database, which is
why every node holds a credential to it permanently.
Two fields close it:
contributes: {reverse-proxy: {host: board, port: 8080}}
receives: {reverse-proxy: /etc/traefik/dynamic/mesh.json}
The control plane collects every contribution on a node and writes them
to the path the provider named, ordered by module so the file does not
churn. Contributing to something is requiring it — asking to be published
means a publisher must exist, and a module that had to say both would
eventually say one.
The control plane does not know what a reverse proxy is and does not
write one's configuration. It delivers facts; the module turns them into
whatever it runs. That is why swapping the proxy touches nothing that
publishes through it, and why the host needs no new vocabulary — a
received file is a file.
Settings reach a contribution the same way they reach a file, because a
hostname is exactly what differs between one mesh and the next.
Two things found by running it:
- the file had a `//` header, so it said "do not edit" to a person and
failed to parse for the program meant to read it. The note is inside
the document now.
- a provider with no consumers gets an empty file rather than none. It
cannot otherwise tell "nothing asked for me" from "the mesh never
wrote it", and those want different responses.
Also `plan <node> --json`, which is how the declaration gets handed to
the host's own parser.
|
||
|
|
44d134ba25 |
Networking is a module, and a domain module is how you avoid choosing
Connectivity was code beside the module system doing the module system's
job: every machine with an address was on the private network and there
was no way to keep one off.
A manifest can now say its resources are computed by the control plane,
which is what a peer list needs — it is derived from every machine at
once, so nothing could be written in advance. The network is a module
from there on: assigned, resolved, settled, and absent from a machine
nobody gave it to.
Three modules rather than one, because WireGuard is one VPN of several:
mesh-wireguard provides private-network, mesh-addressing
claims the-private-network, one per node
mesh-names provides name-resolution, requires mesh-addressing
networking requires both, and ships no files of its own
The last is the point. Most people want the network up and do not want
to choose a VPN, so `assign networking` takes the only answer to each
requirement silently. The day the catalogue holds a second one there are
two answers, the resolver refuses and names them, and choosing is
assigning the one you want. No flavor field, nothing to configure.
Names left the WireGuard declaration for their own module. They would be
identical over a different private network, and bundling them made one
module out of two things.
Three faults the walk found:
- choosing tailscale still installed WireGuard, dragged back in by the
names needing the mesh's own addresses. Caught now by a claim: running
two VPNs is fine, being *the* mesh network is singular.
- a requirement wanted by two modules was reported twice, identically.
- "this mesh has no hub" was reported when the real cause was that a
node could not be resolved at all. It now names the node and the why.
And a test that asserts the manifests actually shipped, after the claim
went missing from the real one while every test stayed green.
|
||
|
|
409cd16a09 |
The mesh decides what a node runs
The gap that has been named at the end of every report for a week. Until now a
declaration came from a person handing over a file; now it comes from what was
assigned, resolved against the catalogue, and the control plane is deciding
rather than relaying.
Everything from the module conversation, built and run on real machines:
assign laptop i3 -> accepted, brings xorg, because nothing else provides
it and there was no choice to make
assign laptop sway -> refused: xorg and wayland both claim the-seat
assign laptop editor -> refused: three modules provide a shell -- bash,
fish, zsh -- choose one
assign laptop zsh -> accepted, and the editor's requirement is answered
bash, fish beside it -> fine, nothing is claimed
Claims rather than pairwise exclusion, so a third display server would say what
it claims and need no edit to xorg or wayland. Scoped to node, site or mesh:
two DHCP servers at one site collide and at two sites do not, and the mesh-wide
one is the hub said as a claim instead of hard-coded.
Some conflicts cost no manifest field at all. The refusal above names the seat
AND the two files, because the mesh already holds every resource of every
module -- neither i3 nor sway knows the other exists.
Resource identities carry their module, so two modules may both call something
"config" without the second silently replacing the first. What a service
reflects is qualified the same way, or it would name a resource that no longer
exists and stop being restarted when its own configuration changes.
Nothing is sent until every node resolves. A push that configured three and
refused on the fourth would leave the mesh in a state nobody asked for, and the
fourth is exactly where a claim collision appears.
One real flaw found by using it rather than by testing it: assigning zsh did
not satisfy a requirement for a shell. Requirements were counted against the
catalogue without first asking what the set already offers, so "choose one and
assign it" named three modules and then ignored the one you chose. The remedy
was useless and every test passed.
|