cc86f8b433bbab5feff3e2a4e94aae6a0bd9fb7f
29
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
97448194ac |
Seats are a closed set, a seat's holder answers for what it delivers, and a build source may live on the git seat
Implements novox/hq ADR 0110 and 0111. The seat set lives in internal/catalogue/seats.go: fourteen seats, each with a scope, what occupying it delivers, and the record that made it one. A test asserts the count and a decision per entry, so changing the set means finding the argument, as the host's vocabulary test does. The first set is every seat already claimed — including the-private-network, which the network module claims from a manifest composed in this repository's code, not from any module.json — plus npm-package-registry (ADR 0109) and git (ADR 0111). A test parses every catalogue manifest and this repository's own and fails on any refused claim, so closing the set refuses nothing in use. ParseManifest now refuses a claim on a seat the mesh does not define, a seat claimed at another scope, and a delivering seat claimed by a module that does not provide what it delivers. A malformed claim is refused once, for being malformed. Resolution: among several providers of a mesh provision, a pin still wins; then the holder of the seat that delivers it; then the only provider; otherwise refused as before. ADR 0009's "never guessed" holds — the seat is the choice made once, mesh-wide, rather than a pin per consumer node. A provider now carries the module it came from, because a provider is a (node, module) pair and the pair is what tells a holder from a neighbour on the same machine. The planner's second pass is now given the first pass's holdings. Without them, a node consuming a seat-delivered provision was refused there, and a refused node's own claims dropped out of what the mesh holds — letting a second holder of one of its seats pass unrefused. `seats [--json]` lists every seat, what it delivers, and each holder, derived from assignments every time and never stored. Unheld seats are listed. A stored claim outside the set — possible for a manifest registered before the set closed, since stored manifests are not re-validated — is shown rather than hidden. `build --self <owner>/<repo>` builds from a repository on the git seat's holder. The clone URL is composed at build time from the holder's node and what it serves for git; the recorded source is the path and the seat (migration 0032), never an address, so a moved forge changes nothing recorded. Nobody holding the seat refuses self-hosted builds and says so; external URLs are unchanged. An address passed with --self is refused rather than recorded as a path. Replaces three foundation tests that defended the builder's carried package binding. The catalogue removed that binding when the builder began requiring the registry through a real grant, so the tests were already failing on main; they now assert the builder requires what the npm seat delivers and carries no copy of its own, and that the forge holds the npm and git seats. Verified: go vet clean; the whole suite passes against a throwaway Postgres (make postgres), the new inventory tests included; gofmt clean apart from cmd/mesh-builder/stdout_test.go, which fails on main too. |
||
|
|
9f3790dcda |
Review: one local name is still a local name; a local name is unique; recovery knows it; recipes read as instructions; a tag before a digest; ask fails at once when nothing serves
A secrets object with one local name delivered no file. Two requirements could share a local name. secret recover and the export could not tell two locals apart. The recipe check missed continued lines and read heredoc bodies as bases. repo:tag@digest kept the tag in the repository. ask now publishes mandatory, so a tool nothing serves is said at once rather than after the wait. |
||
|
|
3e5c010c5e |
A need is kept once per provision, consumer, local name and provider
Two consumers of one same-node provision produced two raw needs and, fanned out per consumer, four — the same credential twice for each. Harmless, since a pair is one row however often it is asked for, and wrong all the same. |
||
|
|
76773dfbf5 |
Several secrets expand where the consumer is known, not on the first module to mention the provision
The lab's two-secrets consumer was given one credential and no file: the expansion ran on the resolver's walk over names, on whichever module mentioned the provision first, and the per-consumer pass copied that. It expands in that pass now, and a test has two consumers of one provision, one keeping one file and one keeping two. |
||
|
|
6ae4ae1dba |
A module may hold several secrets from one provider, each a pair of its own
secrets: maps a requirement to several files under local names. Each local name is its own need, its own pair credential (the pair is keyed on it: migration 0027), its own file on the consumer, its own holder at the provider (the identity with the local name after it) and rotates apart from the others. The plain shape is unchanged and every existing row is the credential it was (novox/hq 04-ISSUES/069, ADR 0094). |
||
|
|
5062c36fc9 |
A binding answered on this very node still carries an address
Two sibling branches resolve a provision answered by the consumer's own machine. The one for a node-scoped provider falls back to loopback when the machine is on no private network, with a comment saying why and a test holding it. The one for a mesh-scoped provider passed node.At straight through, and nothing noticed because nothing had yet composed a host out of it. The mesh's own artifact store is mesh-scoped and sits on the same machine as the builder that pushes to it. Give the builder the address from its binding and it gets MESH_REGISTRY=:5000 — a name with no host, written into its environment without complaint. It surfaces much later as ":5000/mesh-tools/build" is not a valid repository/tag which is a message about a tag for a fault in how a binding was resolved, on a machine several steps from the decision. A machine off the private network still reaches itself, which is what the neighbouring branch already said. The test fails without the fix, showing the empty address rather than only the symptom. Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx |
||
|
|
c4030947b0 |
The routing record is 0066, not 0056
0056 is 'the authority is the control plane, not a database'. A citation pointing at the wrong decision is worse than none: it reads as corroboration. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF |
||
|
|
b824c65ab0 |
catalogue: a co-located provider's served values, and a co-located contribution's port
Two more of one fault, and the fault is the same as 04-ISSUES/038: the same-node path diverging from the cross-node one. The mesh works out what a provider on ANOTHER machine serves by walking that node — reading its manifest with that machine's port assignments, then settling the result with that node's settings layers — before offering it to a consumer. A provider on the consumer's OWN machine never passes through that walk, so every step of it had to be repeated in resolve.go's servedHere and declaration.go's here(). 038 repeated the port. Nothing repeated the settling. So a served value the operator supplied reached a co-located consumer as the manifest's empty default. On the ADR 0056 anchor that value is an internal CA's root: step-ca and route-proxy on one node, route-proxy's binding carrying root: "", an empty CA bundle written, a silent fall back to the system trust store, and issuance stopping with nothing saying why. The same step-ca on another node would have worked. The second is the mirror direction. gitea declares a bare container port 3000 and the machine publishes it as 20000:3000, but gitea's route CONTRIBUTION still said 3000 — so the proxy beside it dialled a port nothing listens on and answered 502. 038 fixed what a consumer is TOLD about a provider; this is what a workload TELLS a provider about itself. The redirect uses the CONTRIBUTING module's assignment, because the port is the workload's, not the proxy's; a contribution carried here from another machine is left exactly as it is, its port being that machine's to assign. Both are settled in Declaration, which is the first moment the machine's ports and the provider's settings both exist. That also removes an order dependence: the resolver built its same-node needs mid-walk, from whichever modules had been chosen by the time the requirement came up and in whatever order a map iterated, so what a co-located binding carried depended on the order somebody happened to assign things in. Re-deriving from the finished closure does not. servedHere keeps its job — deciding whether a same-node provider serves anything at all, which is what makes the need exist — and now says that its values are provisional. novox/hq ADR 0056 Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF |
||
|
|
232862315c |
catalogue: compose a route's name from a label and its node's domain, and resolve it in-mesh
A public route used to carry its whole hostname as a literal in the module manifest, so running the same catalogue against a different domain meant overriding that literal on every routed module, per node. The mesh was, in effect, holding a map of names to services: the one thing it should never hold, because the subdomain is the operator's choice and the domain is the node's. Compose instead. A route contribution carries a `label` (the subdomain); a node carries its `public_domain` as node-level configuration; the mesh joins `<label>.<public-domain>` and grants exactly that, interpreting neither half. Held as a node property beside the node's other node-level facts (endpoint, site, overlay address), not in a module's settings — the ADR calls it node-level, and the settings table is keyed per module. Additive, so an unmigrated catalogue keeps working: a contribution that still carries a full `name` and no `label` passes through unchanged, and the catalogue can migrate module by module. A labelled contribution on a node with no public domain composes nothing, reading downstream as a route that named no host. And propagate: each granted route name is published into internal resolution mesh-wide, mapped to the node that serves it, alongside the `<node>.internal` names every container already gets. So a container — and an internal ACME validator, which cannot complete a challenge for a name it cannot reach — resolves a routed name to the proxy that serves it. Name-agnostic throughout: the mesh propagates whatever names it was told to serve and knows nothing about what they mean. novox/hq 02-DECISIONS/0056 Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF |
||
|
|
a92c11be12 |
Resolve: one un-hostable assignment no longer refuses the whole node
A module a person assigns to a machine that cannot host it — its declared capability has no detector there, as fail2ban does on a host with no firewall — made Resolve refuse the entire node, so a whole-node push refused to send the healthy modules beside it too. One module on the wrong machine took down every other module on that node. Assign already keeps such an assignment on purpose (it is what a person meant, and acts.go says so), so the fix is on the resolve/push side: a directly-assigned module the machine cannot host is left out of the closure and reported as un-applied on the Resolution, rather than refusing the set. The healthy modules still resolve, declare, and converge. A module that is *required* by something running here and cannot be hosted still refuses — that set is genuinely incoherent — so the distinction is who wanted it. assign, plan and push now name the un-applied module and the missing capability, via a shared WrongMachine message, so it is neither silently dropped nor fatal. Reconciled two tests that encoded the old whole-node refusal for directly-assigned un-hostable modules; added coverage for the healthy-modules-still-converge case and the required-un-hostable-still-refuses distinction. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF |
||
|
|
b78a911e34 |
catalogue: deliver a keyless same-node provider's served facts
A node-scope provider that answers a requirement on the same machine and
serves connection facts (a port) but mints no credential delivered
nothing to a co-located consumer. resolve.go only built the delivering
Needed when brokered[want] was set — true only for mesh-scope providers;
a node-scope keyless provider set local[want] instead and fell through,
so knownFor saw no binding and boundInto refused the consumer's
${bound:model-access:port} file.
Deliver the served facts as a need whenever the same-node answer serves a
non-empty set, with a loopback fallback for the address when the node is
off the private network — the reachability rule does not apply to two ends
on one machine. The brokered (credentialed, mesh-scope) path is untouched.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
|
||
|
|
33fd28ffa6 |
licences: deliver the refresh token by the ordinary sealed path, not a bespoke envelope
The refreshable-grant refresh token no longer rides a custom at-rest envelope that a
module opens with a node private key. A module is never given a node's private sealing
key, so that path could not exist -- the gap Phase C hit.
Instead the refresh token is a credential sealed to the MANAGER holder with the same
anonymous box (secrets.Seal / crypto_box_seal) every credential uses, stored as one
sealed blob, and delivered by the existing host-unseal-and-mount: the host opens it with
the node's real key and mounts the cleartext at the manager module's bound path, exactly
as a consumer's db password is delivered.
- refresh_grant now stores { sealed, manager_key }, dropping the AtRest token/wrapped_key
columns; internal/secrets/atrest.go is retired (nothing else used it).
- the licence records its manager as (node, module); KeyFor delivers the refresh token to
the manager holder and the access token to consumers, disambiguated by module so the two
can co-locate. Accept and the reseal skip the manager holder.
- the manager holder is delivered the node's PUBLIC sealing key in its bound facts, so the
module can re-seal a rotated refresh token with no private key of its own; the
declaration tolerates its empty pre-adoption secret rather than refusing.
- SubmitRefresh / set-grant take a sealed blob, never a refresh token in the clear.
The invariant holds unchanged: the control plane never reads the refresh token, and no node
but the manager holds it. A committed cross-language test proves the TypeScript module seal
opens under Go box.OpenAnonymous (the host's Unseal) -- both are NaCl crypto_box_seal.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
|
||
|
|
aeb65a3e1d |
catalogue: a module accesses operator-owned data, and does not own it
04-ISSUES/036: the media stack is several modules that must share the library and download directories on one machine, but the manifest could only say "a directory I own". Six modules each declared the same paths as their own resources, and the resolver's duplicate-owner refusal — right in general — would refuse the stack's only sensible assignment the first time two of them landed on one node. Add an `accesses` field: a pre-existing, operator-owned path a module is granted use of but does not own (novox/hq ADR 0051). Distinct from a `directory` resource on every axis the host acts on — the mesh creates, chowns and reconciles a directory; it mounts an access and owns nothing. An access is not a resource, so it never enters the duplicate-owner map and several modules may name one path with no conflict. What is refused is the contradiction: a path one module owns and another accesses. Rendered into the declaration as an `access` resource, before the container that mounts it, so the host can find it present or refuse clearly. Unit tests cover co-resolution (the exact 036 case), the unchanged owner-vs-owner refusal, the owner-vs-accessor refusal, and access validation. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF |
||
|
|
1f5b70a995 |
The mesh assigns the port, and a module says it once
novox/hq ADR 0038. A module cannot choose a port: it is written once and assigned anywhere, so any number it picks is a guess about a machine it has never seen. A database module met the mesh's own store on 5432 and was told, by a container runtime three layers down, that the port was already allocated. The number used to appear three times in every module — the rule set, what a consumer is told, and what the runtime publishes — agreeing only because one person wrote all three. Now it appears once, in `listens`, and the other two are derived: the container publishes `20000:5432`, the consumer is told 20000, and the rule set opens 20000. An assignment is made once and kept, as a credential is. A port that moved on every declaration would restart both ends each time and hand a consumer a number that was true when it was read. Ports the protocol fixes — mail on 25, submission on 587, DNS on 53 — say so, and are then claims: one holder per machine, and the second is refused by name at assignment. That is the mechanism the mesh already has for what is singular on a machine, pointed at ports. A mapping written the long way is left exactly as it is. Some things must be pinned by hand, and quietly overruling somebody who wrote both halves would be worse than not offering the short form. Still open, and known: the substrate is not a module, so the mesh has never heard of its own store and cannot yet assign around it. That is what 028 will still be about after this. |
||
|
|
0af3ea1acf |
A consumer is a module on a machine, not a machine
novox/hq 04-ISSUES/022. A credential was keyed by provision, consumer node and provider node, so "who is asking" was answered by naming a host. The node this mesh exists to take over runs eight modules against one database server. The symptom had two halves and only one was loud. The provider refused, naming the modules and explaining they would share one credential, which reads as a decision rather than a limit. The consumer did not refuse: it resolved cleanly, wrote one module's credential file and left the others absent — a service that starts and cannot authenticate, with nothing saying why. That is 021 again on a different axis. Three modules wanting one database produced one need, carrying whichever module mentioned it first, because the resolution walk is a work-list over names. The fan-out now happens in one place, after the walk. The record path already did this correctly and said why: a consumer here is a module on a machine. It is the same rule. Downstream: the secret's key gains the consuming module, the grant file is named after both halves, needs are matched by provision and module rather than provision alone, and the provisioners name the role and the access key after the module. The refusal in ContributionsTo is gone because there is nothing left to refuse. Worth stating plainly: without that refusal, gitea's login would have opened keycloak's database. From the provisioner's side it created exactly what it was asked to create. Existing secrets are discarded rather than backfilled. They cannot say which module they were for, and a secret is remade and delivered to both ends on the next push — so this costs one rotation and invents nothing. Also guards the role name against PostgreSQL's 63-byte truncation, which is a notice rather than an error and would reintroduce exactly this collision at a length nobody tests. Three faults injected — the fan-out removed, needs matched by name alone, the grant file named after the machine — each caught. |
||
|
|
df62bb57e5 |
A requirement answered on this machine is still a requirement
novox/hq 04-ISSUES/021. Two modules where one provided what the other required, on one node, resolved cleanly with zero needs: no credential was made, the consumer's secret file was never written, and whatever read it would fail somewhere else entirely. Nothing was refused and nothing was reported. The world a node resolves against is every OTHER node, so a provider on the same machine never became a Needed, and the credential loop walks Needs. Every step reasonable, the sum a silent gap. It survived because everything proven until now was cross-machine — the interesting case for a mesh and the rare one in practice. The first module to want a database on its own machine was the first real one. The assumption underneath was that a local consumer needs no credential, which holds for a process reaching a unix socket where the system can vouch for the caller. It does not hold for containers, which is how nearly everything here runs: the consumer reaches the provider over TCP from its own container and the database asks for a password exactly as it would from another machine. **The machine stops being a trust boundary once both ends are containers.** A brokered provision answered here is now a need naming this node, and carries what the provider serves — which a local provider never contributes through the world. A name nothing grants is unchanged: a shell answered here is answered, and nothing more is owed. Both directions tested, both injections bite. |
||
|
|
68b0d0c10a |
What answers a requirement is applied before what asked for it
The host does not sort, so the order written here is the order a machine applies. Selection walks outward from what was assigned, which puts a consumer before the thing it pulled in — and a service that reads a file another module writes then starts before the file exists. It fails, and the next reconcile fixes it. That is the worst shape a fault can take: what gets remembered is that it works, and nobody looks again. It is 04-ISSUES/013 one level up from where that was found — there, the mesh's own computed files came after a module's resources; here, a whole module comes after the one that needed it. Nothing had hit it because no module until now both required something with resources of its own and had a resource depending on it. Writing the resolver module was what made it reachable, and it would have shown up as dnsmasq failing once on every fresh machine and working ever after. Unrelated modules keep the order selection gave them — assigned first, then what they pulled in. That order is meaningful, and reshuffling it would make every declaration's diff unreadable for no gain. Two modules requiring each other are both applied rather than refused: a cycle is not a machine that cannot work, and refusing would make a cooperating pair impossible to assign. |
||
|
|
faf5ecd70f |
One machine's unanswerable requirement does not remove it from the mesh
The pass that answers *what does this node offer* takes a failed resolution to mean it learned nothing about that node. So refusing an unanswerable requirement there made the machine disappear — and every other machine was then told, wrongly, that the two of them shared no private network. A wrong answer about a machine nobody asked about, caused by a fault on a third. The lab found it: one module needing a licence that had not been added yet made two unrelated machines look disconnected. The second pass still refuses it, where the question is actually being asked. |
||
|
|
87c6a56b81 |
Model access is a provision answered by a record, not a machine
novox/hq ADR 0024, gaps 1 and 2. The user's stated requirement, and the first thing here that no machine can answer: a hosted model is on nobody's node and is reached over the public internet, so the rule that refuses two ends sharing no private network must not apply to it. A licence is a named thing and the name is the operator's — *the personal account*, *the organisation's* — because the whole point is saying which one a given consumer uses, and an anonymous credential hanging off a provider cannot be said. Many to many, so deliberately not a claim: two machines sharing an account is ordinary rather than a collision. Gap 2 is the missing verb, *accept*: take a value somebody supplied, seal it to each holder, discard the plaintext. With the consequence stated rather than hidden — a holder recorded after the key was supplied has no key and the mesh cannot make one, so it is refused by name with the remedy, not silently handed an empty file. Refusal is felt, as the record warns: a mesh holding three ways to reach a model refuses every consumer that has not chosen. So the refusal names the candidates and the exact command. Being right is not the same as being usable. Gaps 3 and 4 — a consumer that is not a machine, and switching as a reaction rather than a declaration — remain gaps. Half-building them would put a conditional in the declaration language, which is what ADR 0024 says plainly to avoid. Its own context, with its own store and its own credential: a licence is a different aggregate from anything inventory owns, and it refers to nodes by name because that is what crossing a context boundary may carry. |
||
|
|
8b2d1bce9d |
Something answered on this machine is still bound
A binding was skipped when the provider turned out to be on the same node, reasoning that a file saying "it is on this node" is a fact nobody needs. That is right about the location and wrong about everything beside it: a binding also carries what the provider said a consumer must know, which is the port, and a consumer cannot invent that. A build machine sharing a node with the registry it pushes to sat in a loop saying it could not read its own binding. Nothing was wrong with the machine, the module, the credential or the provision — the file was never written, and the absence looked exactly like a mistake in the module. The original intent is kept where it was right: a provision whose provider said nothing a consumer must know is still not written. A shell is answered here and there is nothing to say about it. A registry is answered here and the port is still unguessable. The address is this machine's name on the private network, or loopback when it has none — a machine off the network still reaches itself, and a name nothing resolves is worse than an address that always works. |
||
|
|
d978712d7f |
Split resolving from rendering a declaration
resolve.go had grown to 796 lines doing four jobs: working out what a machine should run, applying settings, collecting contributions, and placing credentials. They answer different questions — the first is "what", the rest are "what does that look like as resources" — and one file doing both is how a thing starts becoming the kernel everything imports. Prompted by looking at why HAL's shared library became unmaintainable. Measured while here, and the shape is the inverse of that one: the large packages import nothing internal, and only inventory and link compose. A change to module resolution cannot reach connectivity, because connectivity does not import it. |
||
|
|
c3046dcf56 |
A provider is told who its consumers are, and a reference provisioner
Contributions were node-local, so a mesh-scoped provider — the one case that most needs them — never heard from its consumers. A database was given a password and no idea what to create it for. Cross-node consumers now reach the provider's `receives` file, merged in with the ones on its own machine: from the provider's side they are the same thing, and a provider that had to read two lists would read one of them. Each names the file its credential is in rather than carrying it, because the mesh discarded the value and could not put it there. The readable half therefore stays readable. And examples/postgres-provisioner, which is the last step: it reads what the host wrote and makes PostgreSQL accept it. Explicitly not part of the control plane — the control plane decides and never touches a machine. This runs on the machine and touches it, and a real one ships with the module that ships PostgreSQL. It lives here because this is where the contract is defined, written as something that runs so it can be read. It reconciles rather than applying a change, because it is never told what changed. Three things that follow, and each is a fault somebody has shipped: - the password is set every time, not only on creation, or a rotation reports success and changes nothing - what it made and nobody asks for any more is revoked, or a departed consumer keeps a working login for ever - what it did not make is left alone, or it cannot be run on a database that predates it Proven in the lab against a real PostgreSQL, each assertion confirmed to fail with the behaviour removed. The suite is in mesh-lab, which also records the two ways the test itself was wrong first. |
||
|
|
20f78cd5f1 |
Credentials the mesh delivers and cannot read
HAL keeps env vars in the registry, encrypted at rest. Its own tooling records what that bought and what it did not. `secret_locate` matches by value rather than by name — because the same password sits in mesh_provisions, in module_env, in each node's .env in plain text, and inside every connection string composed from it, and its documentation says those URL copies "are often the only copies actually in use". And a query against the encrypted column returns zero rows and proves nothing, so auditing moved to the decrypted copies on the nodes. Two faults there, and encryption at rest addresses neither: the control plane can read what it stores, so a copy of the database is a copy of every credential; and one secret has many homes with nothing tracking them. So here the mesh generates a password, seals it to each end with keys those nodes generated, stores both blobs, and discards the plaintext. It cannot read what it holds. Neither can the broker relaying it. And nothing is composed centrally — a connection string is assembled on the machine that needs one — so no copy is ever minted in a shape nothing tracks. `Compromise of a node is compromise of that node` (ADR 0004) is now true of secrets, not only of identity. Two files rather than one, because the mesh cannot compose a document containing a value it discarded: `binds` carries the readable facts, `secrets` carries the credential alone. The readable half stays readable in the declaration; the secret half changes only when the secret does, which makes restart-on precise. The provider gets a directory, one file per consumer, for the same reason. It is made once and kept — regenerating per declaration would restart both ends on every push, and the password a provider was told to create would never be the one its consumer was given. It is remade when either end's sealing key changes, and both ends learn the new one in the same push, so there is no window where half the mesh holds a dead credential. Two tests found passing for the wrong reason, both caught because their injection came back clean: - the provider's copy was asserted non-empty, which reads the same whichever column is selected. It now opens the blob with the provider's own key. - RotateSecret deleted and re-created; the re-create was dead, because the next read makes one anyway. Removed, and a second path to the same act is how two ends come to disagree. And one real fault: three places built a declaration, and the one behind `--json` predated credentials, so it silently produced a declaration missing them — a difference between what `plan` showed and what anything reading `--json` got. There is one path now. |
||
|
|
c4782ae2fd |
An app is told where its database is
Knowing that a machine needs the anchor's database is useless to the
program that needs it unless the program is told. It knew; nothing was
written anywhere it could read.
Two fields, mirroring contributes/receives in the other direction:
serves: {database: {port: 5432, driver: postgres}} on the provider
binds: {database: /etc/app/database.json} on the consumer
The provider says what a consumer needs to know; the mesh adds the half
only it has — which machine, and what that machine is called on the
private network. The file says, in itself, that it carries no credential
and why. A missing field looks like a bug; a stated absence looks like a
boundary.
Binding something answered on this machine writes nothing. A file saying
"it is on this node" is a fact nobody needs and one more thing to keep
true.
And two machines that share no private network are refused rather than
wired together. An app here and a database there with no path between
them is a mesh that reports itself configured and does not work — the
failure surfaces as a connection timing out, which is the slowest place
to find it. This is checkable now only because the network became
something a machine is given rather than something it has by having an
address.
One fault, found by running it: working out who is on the private network
resolved the mesh, and resolving the mesh asks who is on the private
network. It hung for two minutes. The comment above the function said not
to do that and the function did it anyway; it now resolves each node
locally, which is the right answer to the question regardless — whether a
machine is on the network depends on what it was assigned, not on what it
takes from others.
|
||
|
|
d4064122d6 |
Where the answer to a requirement is allowed to live
Two different things were both written `requires`. A shell, a display
server and a private network have to be on the machine that needs them.
A database does not — it runs somewhere and is reached over the network.
Both were answered the same way, so requiring a database installed
PostgreSQL on every machine that ran a web application.
What a module provides now carries a scope, the same idea claims already
use, written short in the ordinary case:
"provides": ["shell"]
"provides": [{"name": "database", "scope": "mesh"}]
A mesh-scoped requirement is answered by finding the node already running
it — never by installing it here. Choosing a machine to put a database on
is a decision with consequences, and nothing resolving a web application
should make it silently. With nothing anywhere it refuses and says which
module to assign; with two it refuses and says how to choose.
Choosing is `pin <node> <provision> <from>`, kept per node because that
is the granularity the choice has. A pin at a machine that does not
provide it refuses rather than falling back — a fallback would quietly
move somebody's data. One provider does not overrule a pin either.
Resolving a node now needs to know what the others offer, and working
that out needs them resolved, so it is two passes: the first answers only
what each node offers, the second answers everything. Nothing is ever
declared from the first.
A node's plan says what it takes from elsewhere. It is the only part of a
set that stops working when a different machine goes away, and nothing
else in that output would have said so. It is also where a credential
will hang once there is a mechanism for handing one back.
One test found passing for the wrong reason: it read pins through a join
on the provider, which hides a dangling row whether or not it was cleaned
up. It counts rows now, and bites when the cascade is removed.
|
||
|
|
5a3a87e8c3 |
A module can tell its provider what it needs
`requires` said a thing must be there. It never said what to do with it,
so a web application requiring a reverse proxy had nowhere to put "this
name, this port". The two modules that needed it most went round the
outside and opened a connection to the control plane's database, which is
why every node holds a credential to it permanently.
Two fields close it:
contributes: {reverse-proxy: {host: board, port: 8080}}
receives: {reverse-proxy: /etc/traefik/dynamic/mesh.json}
The control plane collects every contribution on a node and writes them
to the path the provider named, ordered by module so the file does not
churn. Contributing to something is requiring it — asking to be published
means a publisher must exist, and a module that had to say both would
eventually say one.
The control plane does not know what a reverse proxy is and does not
write one's configuration. It delivers facts; the module turns them into
whatever it runs. That is why swapping the proxy touches nothing that
publishes through it, and why the host needs no new vocabulary — a
received file is a file.
Settings reach a contribution the same way they reach a file, because a
hostname is exactly what differs between one mesh and the next.
Two things found by running it:
- the file had a `//` header, so it said "do not edit" to a person and
failed to parse for the program meant to read it. The note is inside
the document now.
- a provider with no consumers gets an empty file rather than none. It
cannot otherwise tell "nothing asked for me" from "the mesh never
wrote it", and those want different responses.
Also `plan <node> --json`, which is how the declaration gets handed to
the host's own parser.
|
||
|
|
44d134ba25 |
Networking is a module, and a domain module is how you avoid choosing
Connectivity was code beside the module system doing the module system's
job: every machine with an address was on the private network and there
was no way to keep one off.
A manifest can now say its resources are computed by the control plane,
which is what a peer list needs — it is derived from every machine at
once, so nothing could be written in advance. The network is a module
from there on: assigned, resolved, settled, and absent from a machine
nobody gave it to.
Three modules rather than one, because WireGuard is one VPN of several:
mesh-wireguard provides private-network, mesh-addressing
claims the-private-network, one per node
mesh-names provides name-resolution, requires mesh-addressing
networking requires both, and ships no files of its own
The last is the point. Most people want the network up and do not want
to choose a VPN, so `assign networking` takes the only answer to each
requirement silently. The day the catalogue holds a second one there are
two answers, the resolver refuses and names them, and choosing is
assigning the one you want. No flavor field, nothing to configure.
Names left the WireGuard declaration for their own module. They would be
identical over a different private network, and bundling them made one
module out of two things.
Three faults the walk found:
- choosing tailscale still installed WireGuard, dragged back in by the
names needing the mesh's own addresses. Caught now by a claim: running
two VPNs is fine, being *the* mesh network is singular.
- a requirement wanted by two modules was reported twice, identically.
- "this mesh has no hub" was reported when the real cause was that a
node could not be resolved at all. It now names the node and the why.
And a test that asserts the manifests actually shipped, after the claim
went missing from the real one while every test stayed green.
|
||
|
|
65ade756f2 |
Settings: changing a module's config without editing its file
Managed files are generated and never edited, so somebody's intention about one has to live where the generator can see it. It does now: the module ships defaults, settings go over the top by key, and the file is produced from both. Upstream can rewrite its half freely and the keys somebody chose survive. Two layers, both from the start. The mesh's settings for a module, then one machine's over those. A node that differs is expressed by differing, rather than by restating everything the rest already say -- which would pin all of it against future changes for no reason. An override beats a default and there is nothing to resolve. A setting is a statement about that key made deliberately; the default was only ever what to do in the absence of one. So when upstream changes a key somebody has set, there is no conflict, no merge markers, and nothing to ask. Nested blocks merge and lists are replaced whole. Setting one field of a block must not delete its siblings, or every setting would restate the whole block and pin all of it. A list that merged element-wise could neither be shortened nor reordered, and there is no correct guess about which element is "the same one". A module can keep specific keys for itself -- a socket path its own code depends on -- and setting one is REFUSED rather than ignored. A setting quietly dropped is somebody believing they changed something. Settings that reach nothing are named at the moment they would be used, not discovered later by the machine not behaving differently. `plan --files` prints what a machine would be given before it is sent, because "1 resource" does not tell you whether the merge landed. One test kept with a note that it does not defend this code: output stability comes from Go's encoder sorting map keys, so it passes with the merging removed. Worth having as the thing that would catch a change of encoder, but it is not evidence about anything written here, and it was checked. |
||
|
|
409cd16a09 |
The mesh decides what a node runs
The gap that has been named at the end of every report for a week. Until now a
declaration came from a person handing over a file; now it comes from what was
assigned, resolved against the catalogue, and the control plane is deciding
rather than relaying.
Everything from the module conversation, built and run on real machines:
assign laptop i3 -> accepted, brings xorg, because nothing else provides
it and there was no choice to make
assign laptop sway -> refused: xorg and wayland both claim the-seat
assign laptop editor -> refused: three modules provide a shell -- bash,
fish, zsh -- choose one
assign laptop zsh -> accepted, and the editor's requirement is answered
bash, fish beside it -> fine, nothing is claimed
Claims rather than pairwise exclusion, so a third display server would say what
it claims and need no edit to xorg or wayland. Scoped to node, site or mesh:
two DHCP servers at one site collide and at two sites do not, and the mesh-wide
one is the hub said as a claim instead of hard-coded.
Some conflicts cost no manifest field at all. The refusal above names the seat
AND the two files, because the mesh already holds every resource of every
module -- neither i3 nor sway knows the other exists.
Resource identities carry their module, so two modules may both call something
"config" without the second silently replacing the first. What a service
reflects is qualified the same way, or it would name a resource that no longer
exists and stop being restarted when its own configuration changes.
Nothing is sent until every node resolves. A push that configured three and
refused on the fourth would leave the mesh in a state nobody asked for, and the
fourth is exactly where a claim collision appears.
One real flaw found by using it rather than by testing it: assigning zsh did
not satisfy a requirement for a shell. Requirements were counted against the
catalogue without first asking what the set already offers, so "choose one and
assign it" named three modules and then ignored the one you chose. The remedy
was useless and every test passed.
|