eab870fff9e9112e64b81d4793f14a0ab8a451ac
100
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
eab870fff9 |
A board, and the one constraint that is not a feature
Read from the board that exists. Eight sections; four are about work and workers and are held back with that domain. The other four are the mesh itself, and everything behind the main one already exists here — it is a reader, not a second source of truth. The constraint is the point of writing this down now. The existing board is one service that reads every context's database, because that is the shortest path to a page showing all of them at once. That is ADR 0008 violated by the one component with a reason to violate it, and the cost is the same one the shared library has: a boundary nothing may cross is a boundary that can move, and one thing crossing it is enough to freeze it. So a board reads through interfaces and stores nothing. If a question is slow, the answer belongs in the context that owns it, where everything else asking gets it too. |
||
|
|
5253742773 |
ADR 0024 — model access is a provision, and a licence has a name
A new requirement, and it is mostly a shape the mesh already has. A module that needs to think requires model-access; several vendors and a locally-run model are several modules providing it; choosing is assigning the one you want. A model the mesh runs itself needs nothing new at all — it is a mesh-scoped provision on the node with the hardware, credential included. A licence is a named thing because the whole point is saying which one a given consumer uses, and the names are the operator's. Many to many, so not a claim: two machines sharing an account is ordinary, not a collision. Four gaps, written as gaps rather than design: - a provider that is on no node, reached over the public internet, which the reachability rule must not refuse - a secret the mesh is GIVEN rather than mints. Every credential it handles today it generated and discarded; an API key arrives from a person, and accepting one must still discard the plaintext - a consumer that is not a machine. Which licence a worker uses is a binding to an agent, and the provisions model has no consumer identity other than a node - switching on exhaustion is a reaction to something observed, not a declaration. It belongs with observability, changing a binding — saying so is what stops the declaration language growing a conditional The existing auto-refresh and switching is not being replaced because it was wrong. It is being rebuilt because it lives somewhere that cannot express the rest. |
||
|
|
a73014dcd5 |
A bare machine became a mesh, and something joined it
First end-to-end raise. A machine with a container runtime applied the bundle its host carries and ended with a store, databases, schemas, a broker holding a certificate it generated itself, and the control plane serving. Then it took a token, checked the broker against the pinned fingerprint, generated three keypairs and enrolled — the first node being a node whose mesh is not up yet, observed rather than argued. And a credential crossed. Declared the provider of a database for a second node and pushed to over the broker, the machine ended with the password in one file at mode 0600, and that password appears nowhere in the declaration that crossed the broker, nowhere in the control plane's database, and nowhere in what the node reported back. That is the whole secrets argument, measured. One fault, in the joining: the token did not say what the mesh calls the machine, so enrolment needed a flag its own help said it did not, and failed at the broker with an empty username. It is the fifth thing a token carries now — the node cannot work its own name out, because the broker account it authenticates as is named after it and exists before the mesh has told it anything. |
||
|
|
0f7e4ab597 |
The provisioner, which is where the mesh stops
A password nothing was told to create authenticates nowhere. The mesh generates one, seals it to both ends and cannot read it — so it cannot tell the software to accept it either. Something on the providing machine reads what arrived and makes it true. That something belongs to the module, not to the mesh. The control plane decides and never touches a machine; a provisioner runs on the machine and touches it. What the mesh owns is the contract: a manifest of who asked and where each credential is, and one file per consumer holding it. It reconciles and is never told what changed, which forces three things that are each a fault somebody has shipped: set the password every time or a rotation changes nothing; remove what nobody asks for or a departed consumer keeps a login for ever; leave alone what it did not make or it cannot be run on anything that predates it. Saying where the mesh stops is the point. It decides, delivers, and can prove what it delivered; the last inch belongs to whoever knows what `create role` means. |
||
|
|
82a5b9548a |
The secret is delivered without ever being held
Written after looking at how the existing mesh does it, so this is a reaction to a measurement rather than a preference. There, credentials sit in a column encrypted at rest. Its own tooling records what that bought: the tool for finding a secret matches by value rather than by name, because the same password is in three tables, in each node's environment file in plain text, and inside every connection string composed from it — copies its documentation calls the ones usually in use. And a query against the encrypted column returns zero rows and proves nothing, so auditing moved to the decrypted copies. Encryption at rest addresses neither fault. The control plane can read what it stores, so a copy of its database is a copy of everything. And composition is what mints the untracked copies. So the value is sealed to the node that will use it before it is stored, with a key that node generated. Nothing central is composed. What it costs is auditing by value, which was never real anyway; what stays answerable is which node holds what, which is what rotation asks. What remains is a provisioner. The mesh generates the secret and tells both ends; nothing yet acts on the telling. |
||
|
|
cc872a58ce |
Binding is built except for the secret
Which turned out to be the useful way to cut it. A provider says what a consumer needs in order to use it; a consumer says where it wants to be told; the mesh adds which machine and what that machine is called on the private network. So an app on one node reaches its database on another, by a name the mesh also created. The file says it carries no credential and why, because a missing field looks like a bug and a stated absence looks like a boundary. What remains is the secret itself, and the shape it will arrive in now exists. Also: two machines wired together across no private network is refused, and that only became checkable when the network stopped being something a machine has by virtue of holding an address. |
||
|
|
80b74d32d6 |
Where the answer to a requirement is allowed to live
0009 distinguishes presence from instantiation — what the edge hands over. It never distinguished where the thing on the other end is, and that turned out to be the half doing the damage: a shell and a database were both written `requires`, so requiring a database installed one on every machine that used one. A provided name now carries a scope, as a claim already does. Scope belongs to the name rather than to each provider, or one requirement means two things depending on which module answers it. A requirement answered from the mesh is never satisfied locally. Nothing provides it, and it says which module to assign somewhere; two do, and it says how to choose. Choosing is recorded per node, because two machines may reasonably use two different databases. And knowing which node answers is the first half of handing a credential back — you cannot be given a database's password before it is settled whose database it is. |
||
|
|
90ecfe6a01 |
An edge has two directions, and only one of them is built
0009 already said a consumer supplies a target and receives a name. What it did not say is that those are two separate mechanisms. Contribution — publish me at this name, on this port — now exists. Binding — and hand me back a credential — does not, and is the larger half: a secret has to exist, be stored, reach one node and not the others, and rotate with every holder informed. That is the invariant set found violated three ways at once, so it is not something to add in passing. The absence had a measured cost. Exactly two modules opened a direct connection to the control plane's database, and they are the reason every node permanently holds a credential to it. Both were doing by hand what this edge is for. Neither needed a new kind of thing. |
||
|
|
7fe2c31bdf |
Networking is a module, and what a domain module actually is
Two records, from building it. 0009 has a section titled "there are no domain modules", and `networking` now exists. It is not a contradiction and it reads as one, so the difference is written down: what was refused contains WireGuard and a proxy and is assigned where half of it is unwanted. What exists contains nothing — requirements and a name — so there is no half. Every artifact it leads to is still an ordinary module assigned on its own terms. With the cost stated, because it is real: adding a second implementation turns a settled question into an open one for everyone using the bundle, not only for whoever wanted the alternative. That is the refusing rule applied consistently, and the alternative is a default, which is the flavor field returning under a better name. 08-connectivity gains why the network stopped being code beside the module system: a machine was on the private network because it had an address, and there was no way to keep one off. A manifest can now say its resources are computed, which is what a peer list needs. And three modules rather than one, because WireGuard is one VPN of several. Naming a module after the job and putting one implementation inside it is flavor wearing a generic name — the second VPN has nowhere to go. |
||
|
|
554f6bd7a4 |
A capability may carry a value, and adding one is not free
Recorded while building the seat detector. A capability is a named fact about a machine: its presence gates an assignment and its detail can carry a value, so "can this run here" and "what should it be configured as" are the same fact read two ways. A verdict has always had a detail beside its yes or no, so panel: oled needs no new concept. Two things that keep the set honest, both worth writing down before anyone adds the fiftieth capability. It must be detected and the detector must say how it knows -- so nobody can add one they cannot check, which is the whole of issue 007. And detectors ship inside the host, which is one static binary, so adding a capability means shipping a new host everywhere. That argues for a small general vocabulary rather than a specific one. |
||
|
|
f140303257 |
A module claims; it does not list its rivals. And flavor is retired.
Three decisions, all Jochen's, and the first is the one that unlocked it. Exclusivity is not a property of a module. It is a property of a singular resource the module takes over. Two shells compete for nothing and any number may be installed; two display servers both want the seat. So a module declares what it CLAIMS, and two modules claiming the same thing cannot both be assigned within that claim's scope. Not "xorg conflicts with wayland". Pairwise exclusion has a property that only shows up later: adding a third display server means editing xorg and wayland to know about it. Every new module requires changing modules nobody who wrote it owns, and the edits grow as the square of the count. With a claim the third one says what it claims and nothing else changes anywhere. Claims have a scope -- node, site, mesh -- which is not new. The mesh already enforces exactly one hub with a unique index. Scope is that idea said once rather than hard-coded per case. And some conflicts need no claim at all: two modules declaring the same file or binding the same port are visible from what they declare. A claim is only written for the abstract ones. A requirement with several answers is refused, never guessed. One candidate is assigned silently because there was no choice to make; none is refused naming what is missing; several is refused naming them. That is what makes a solver unnecessary -- counting candidates has no surprising behaviour, and a solver can be added later without changing a single manifest. Flavor is retired. It was carrying three unrelated meanings: variants of a thing, a subset of a module a node installs, and whatever the current system does, which earned two knowledge-base entries about going wrong. A word with three meanings cannot be reasoned about. What it reached for is two ordinary things -- different modules providing the same thing, and one module with a setting. |
||
|
|
974985b3d1 |
Four things the lab found about the private network
All on the first three machines to actually run it, and all invisible from the mesh's own state: the graph was right, the files were right, the services were up, every node reported success, and the network did not work. A running interface does not re-read its configuration, so a node joining left every existing node carrying a network that no longer existed. A hub sharing a site with a spoke was emitted twice, which WireGuard refuses. Two nodes at one site that neither can be dialled were peered directly, so nobody opened the path and the more specific route blackholed -- this document's own warning arriving in its implementation. And Docker sets the FORWARD policy to DROP, so a hub with forwarding enabled still carried nothing between its spokes. The last one is the sharpest: the substrate at tier 1 silently breaks the network at tier 2, and nothing in either tier's state says so. None of these is reachable by reasoning, and each was found within minutes of a real machine trying it. That is the argument for the lab in one line. |
||
|
|
6bcf0e4f9f |
Issue 010 fixed: origins keep the bundle and the mesh apart
The store records where each resource came from and each origin removes only its own. Verified on the scenario that caused it -- eleven resources raised, enrolled, sent the same two-resource declaration, and the store, broker and control plane were all still running. A later declaration dropping a resource still removed it, so removal by omission survived the fix. Two more faults found while fixing it, both the same shape. A report published to a routing key nobody bound vanishes: the broker accepts it, finds no queue, drops it, and tells the publisher nothing -- so nodes announced what they had applied into a void. And publishReport was discarding its error, so a node that could not tell the mesh looked exactly like one that had. Reports are mandatory now, so an unroutable one comes back and is said out loud, and the binding covers every key a node may publish. |
||
|
|
594ea10b07 |
Issue 010: the first declaration destroys the substrate
Found in the lab, doing the ordinary thing: raise a first node, enrol it, send it a declaration. Both declared resources applied correctly and every container on the machine was removed -- the store, the broker, and the control plane that had sent the message. The link died mid-sentence because the broker carrying it had just been torn down by what it carried. Nothing is behaving incorrectly. Apply removes what the store holds and the declaration does not name, which is what reconciliation means. The fault is that the carried bundle and mesh declarations share one store, so the host cannot tell what this machine raised for itself before there was a mesh from what the mesh told it to have. It is invisible until those two meet, which happens exactly once per mesh: on the first node, after enrolment, the moment the control plane first speaks. The report says what is not the answer, including the tempting one -- having the control plane send the substrate back. It cannot: it was never told what the bundle contained, and the bundle exists precisely because there was no control plane to ask. |
||
|
|
02afb7516b |
What connecting to the mesh is, and what a node presents
Two things this record never said, both asked directly. Connecting to the mesh is one outbound AMQP connection from the node to the broker, held open. There is no second connection and nothing is ever dialled at a node. Being in the mesh means that connection is up. Two different things ride on it and conflating them is what made this murky. An AMQP account, which the mesh issues per node at enrolment, answers whether the connection is accepted at all -- per node rather than shared, because a shared one lets any node consume another's queue, which is the shared-credential fault this record exists to remove reappearing at the transport. The node's own keypair answers which node is speaking, on every message. It is not made redundant by the account: with only an account the control plane knows who is speaking because the broker says so, and that is the same transitive authority this record already refuses in the other direction. A compromised broker could attribute reports to whichever node it liked. So a node holds two things after enrolment -- a credential the mesh issued for reaching the broker, and a key it generated that the mesh only sees the public half of. Both are its own, neither reaches anything else. |
||
|
|
004057d85c |
A node's identity is a keypair it generates. This was never open.
I have been treating "what a node presents to prove it is that node" as an undecided design question for weeks, and blocking on it. It was decided. 08-connectivity says of the overlay keys: each node generates its own keypair, the private key never leaves the machine, the public key is published to the mesh -- and says explicitly that this IS ADR 0004's "a node holds its own identity", applied. Nobody had applied it to the thing 0004 is actually about. What caused it was a word. The lifecycle said a joining node receives its own durable identity, which reads as the mesh issuing something, and then the question is what. The mesh issues nothing. A node arrives holding its identity; what it receives is being known. That line now says what happens: it presents the one-time secret and its own public key, which the mesh records. The rule above it then holds literally rather than aspirationally. The mesh stores a public key, so a copy of the mesh's database grants nothing, and compromise of a node really is compromise of only that node. Also recorded, since it was asked directly: same principle as SSH, own key, not the machine's SSH host key. Host keys are regenerated by reinstalls and image clones, which would silently un-enrol a node; their lifecycle belongs to sshd rather than the mesh; and a partial host has no SSH daemon at all, so an identity scheme resting on one excludes a supported kind of node. The good half of that idea is kept: the mesh knows every node, so it can distribute host keys the way it distributes authorised keys, and node-to-node SSH stops depending on trust-on-first-use. |
||
|
|
5fd522b8da |
A node is a machine; the session is a feature of it
Correcting an overstatement from the previous commit, where I had written that a node IS a conversation. It is not. A node is a machine inside the mesh, and the session is one of the things running on it -- like the host, like any workload. That also dissolves the conflict I flagged as unresolved rather than needing anyone to decide it. 0001 says a node does not authenticate to a model provider, agents do. Still true: the session authenticates, and the session is not the machine. The node does not think, something on the node does. I had manufactured the contradiction by promoting a feature into an identity. 0001's summary row is corrected the same way, and says explicitly that neither the node's session nor a hired worker makes the node itself a thinking thing -- both run on a machine, which is what leaves that line untouched. |
||
|
|
066f14b5f8 |
A node is a conversation, and that is not the employee model
Moving this out of 0003 and out of its vocabulary. I had spent three attempts fitting the node's own session into the agent-as-employee record, each time bending hired, draining, reassigned and retired to cover something none of them describe. 0003 is back to its original text. It belongs in 0004, under what a node is, because that is what it is -- not a program installed on a node but part of the node. It holds one session permanently, anything in the mesh can message it, and it remembers across callers and across weeks. Its system prompt is the engram, which is recorded here for the first time despite running on every node. Also recorded: it has its own narrower tool list, so it can go and look rather than only report about itself; there is no authorisation between nodes, because every node is the operator's own; and how a node passes a question on is its own business rather than a protocol field. Switched off it still answers, and that is the point of having an off state rather than an absent one. A node with nothing there is a silence somebody has to diagnose. A node that says it is switched off is not. Same rule the host follows about a service that does not exist. 0001's summary is corrected too: it had one row for "agents", which is the conflation being complained about. Two rows now. A node's own session and a hired worker are built from the same parts and run on entirely different terms. Left standing and NOT resolved here: 0001 says a node does not authenticate to a model provider, agents do. A node that holds a session does. That is a real conflict between what is recorded and what runs, and it needs deciding rather than a fourth reconciliation from me. |
||
|
|
079c488d5e |
Provisioned and immutable beats exempt
Replacing the framing I wrote an hour ago. I had the node's own agent sitting outside the lifecycle as an exemption, which is a rule somebody has to remember. Provisioned the ordinary way and constrained is a rule the system enforces, and it is one row like any other rather than a category every query listing agents has to special-case. It also reads the original sentence more carefully. "Exempt from the hiring lifecycle" is exempt from hiring, not from having a lifecycle. Its lifecycle is the node's -- provisioned at enrolment, retired when the node is retired. Same states, a different thing driving them, and no exemption needed. The constraints are now the four nonsense states written as things that cannot happen rather than as an argument: not retirable, reassignable or deletable while its node exists; exactly one per node. And a distinction that was missing -- its existence is immutable, its engram is not. Freezing the personality would remove the way a node is configured. Disabling is the better half of this. A node with no agent is a silence somebody has to diagnose; a node whose agent is disabled answers saying so, immediately, with no model invoked -- the queue is still consumed and the state is the reply. That is the host's own rule about a service that does not exist, applied one tier up: absence must never be indistinguishable from a failure to answer. |
||
|
|
fd7f7557bd |
The node's own session, and why it is not hired
Answering a question that was asked three times and that I kept not answering: should the node's session just be an agent per node, since otherwise the functionality exists at two levels? Same mechanism, different lifecycle. A persistent session, accumulating memory, a system prompt, a scoped tool list, addressable by message -- identical, and building that twice is the duplication the question was worried about. What must not be shared is the lifecycle, because if a node's own voice were an ordinary hired agent it could be retired, leaving a node nothing can talk to; reassigned, moving one machine's mind onto another; hired twice, with no answer to which one replies; or never hired, leaving a node mute. The exemption in this record exists to make those four unreachable. I had this backwards earlier today and said so out loud: I called "a node itself is an agent of a kind exempt from the hiring lifecycle" a fossil of the old model and recommended striking it. It is the design. And it does not conflict with 0001 -- "the two agent rows per node merge" means one per node, not zero. I read merge as delete and invented a contradiction between two records that agree. Engrams are recorded for the first time. They are in use on every node and appear in no record, which is how a decided thing comes to look accidental. The engram is the node's system prompt, and it is what makes one node's answers recognisably its own rather than generic. Also recorded: there is no authorisation between nodes, because every node is the operator's own and a prompt from one is a prompt from them. The consequence is stated once rather than left to be discovered -- the mesh boundary is the security boundary, which is what puts the whole perimeter on the token and the overlay. And how a node passes a question on is the node's choice, not a protocol field. A node may say who is asking or may simply ask, the way a person relaying a question decides how to phrase it. That follows from the engram. The cost is that there is no machine-readable chain of who ultimately asked; each node still holds what it was asked and by whom. |
||
|
|
88ba81e9c1 |
Agents reaching nodes is the capability, not a hole in it
Correcting what I wrote an hour ago. I had recorded node-to-node SSH as "not a mesh function" and "a second control path through the back door", reasoning from ADR 0004's rule that the host has no inbound control surface. That conflated two different things and got the product backwards. There is no node-to-node SSH to forbid. The actor is always an agent; a node is only where it happens to be running -- ADR 0001 already says a node is a place where an agent can run and that is the entire relationship. An agent hired onto one node reaching another to do work is the capability the whole arrangement exists to provide. The credential is the agent's, in its own credential directory, which ADR 0001 already established. So a node's authorized_keys lists agents and never nodes, and three things follow: no node holds a key reaching another node, so 0004's "a node holds its own identity and nothing else" stays literally true; a compromised node costs the credentials of the agents that were on it rather than a way into everything; and who may reach what stays a mesh-wide fact, which is why it is identity's. The rule I misapplied is about how a node's declared state changes -- over the broker, never by being dialled. An agent with a shell is not the mesh reconfiguring a machine, it is what a person with a terminal has always been, and this design already depends on that working: the overlay is the way back in when a declaration breaks something. What such a session leaves behind is drift, and drift is what reconciliation is for. 0001 also stops underselling the fourth layer. It read as "the layer the other three exist to carry", which is true and flat. The value is that an agent can work across a set of machines as though they were one -- centrally configurable machines are ordinary; that is not. |
||
|
|
918dc04916 |
What this actually is, and three things that were assumed
Four things settled by talking them through, all of which had been true in somebody's head and written nowhere. It is not a mesh in the peer-to-peer sense and will not become one. 0001 now says what it is instead: machines linked by a private network, one node holding knowledge of all of them, modules as the way anything is built and delivered, and agents hired onto nodes to do the work. The word describes what machines can reach, not how they are governed. "Master" overstates it the other way -- nothing needs that node to keep running, only to change. 0006 gains the option that would make it a real mesh, recorded as considered rather than rejected by silence: every node holding the whole inventory, a replication process, an elected master with promotion on failure. What settles it is not the complexity but that it still would not deliver the name, because application databases are not replicated -- so a genuine peer-to-peer mesh means becoming a replicated database system for every consumer's data too. That is a larger product than the thing it would support. Also in 0006: three central roles, not one. Losing the control plane costs change, losing the broker costs being told anything, and losing the hub costs nodes in different places reaching each other at all -- which is operation, not administration. Whether they are one node is not decided. And SSH access is identity's. It appeared three times as something that uses the overlay and never as something the mesh provides, which reads as settled when nothing decided it. Nobody else could: the mesh is the only thing that knows which humans and agents exist and which nodes they may reach. Node to node SSH stays out -- the host has no inbound control surface by decision, and nodes reaching each other that way is a second control path through the back door. 0007 gains the requirement underneath all of it. Reachability was recorded as a fact to track and never as a thing some node must have. The broker's node and the hub must be dialable by every node at a stable address, or nothing can join and a disconnected node cannot return. A mesh entirely behind NAT cannot be raised. That is a precondition and it belongs with the others. The link staying on the underlay is also argued now rather than asserted. At join time it is forced; afterwards it is a choice, and the reason is that a repair channel carried over the thing being repaired is not one. Moving it onto the overlay, with fallback, is recorded as open with what it would have to get right -- a WireGuard interface has no link state to test, and a silent fallback is this repository's recurring fault in a new place. 0010 says in one line what was the intention throughout: the module system is the CI/CD. Not a pipeline beside the mesh. Build, test, publish and deploy are one reconciliation seen at four points, which is why a thing that cannot be a module cannot be delivered. |
||
|
|
5218b06c02 |
Fold the control plane's build decisions into 0006 and 0008
Back to 23 records. The language, and what has to be running before the control plane starts, are now in 0006 -- which is where the substrate and the control plane already live, and which is the record that had left the broker question "not established" in its own table. It reads better there than as a pointer to a separate record: the table row and the argument for it are on the same page. The store mechanics went into 0008. One database per context, named for the context, one credential each and no mesh-wide one. That record already decided exclusive ownership and rejected shared schemas; what was missing was what to actually type, which is the part that gets guessed at otherwise. Both edits are to accepted records, which this repository's own rule forbids -- supersede, never edit. Recorded here so it is visible rather than silent. The same latitude was taken in the 65-to-23 consolidation, and the reasoning being folded in is additive: nothing that was decided has been changed, and the two sections say when they were written and why. |
||
|
|
82a3065f82 |
Tier 2 exists, and the token was missing a quarter of itself
mesh-control is built as far as it can honestly go: one context of seven, inventory, with its schema and the command that applies it. The repos map and the control plane design say so, and point at ADR 0024 for what it took. Separately, and more importantly: this repository described the enrolment token as carrying three things when ADR 0004 says four. The missing one is the control plane's signing identity -- the reason a node does not have to trust the broker it dials. Without it the control plane's authority is transitive through the broker, and 0004 spells out what that costs: a compromised broker could forge declarations, and since the host applies whatever the link delivers, that is the whole machine. The record has the argument in full; the design doc had dropped the conclusion. Found by reading the two together while deciding what the control plane must store, which is roughly the only way it would have been found -- both documents are internally consistent and only disagree with each other. |
||
|
|
84f4425fd6 |
The broker precedes the control plane, and it is written in Go
Two things found by trying to build tier 2. The substrate design asked whether the message broker has to be running before the control plane, and framed it as depending on whether the control plane's own parts talk to each other over it. They do not -- it is one process -- so under that framing the broker stays out of the bundle. The framing cannot answer the question. What decides it is how the control plane reaches a node, and the answer was already decided: only ever over the link, and the link is the broker. So provisioning the broker would require the broker. The first node does not escape this by being local, because it enrols the ordinary way, by dialling the broker at the address in its token -- which was deliberate, and worth keeping. The bundle is two images now. The record says what that costs, including a certificate the broker needs at a moment when there is no mesh to issue one. The language had never been decided for tier 2. Go, for the same reason the host is: the bundle pins this image by digest and runs it where nothing can check it, so the image should hold the program and nothing else. Also corrects something already built: the bootstrap created one database and called it 'mesh'. ADR 0008 grants a context only what it exclusively owns and ADR 0006 says the mesh database names a thing that will not exist. One database per context, so one today, called inventory. |
||
|
|
6a2b107fb8 |
Restore a consequence the consolidation dropped
I said nothing was lost when 65 records became 23. That was too strong, and here is a counterexample: ADR 0046's consequence that the lab needs a way to place images did not survive into the merged substrate record. The compression kept the decision and dropped one of the things it implied. It was not lost from the repository -- 04-ISSUES/009 had already picked it up, which is why it was found at all. But the record no longer carried it, and the record is where somebody would look. Restored, now as a resolved fact rather than an open consequence: the lab raises a registry inside the scenario, which is the real path since that is what every node after the first pulls from. The digests it serves are its own, and that satisfies the pinning rule -- what is required is a reference that is exact and cannot move. Worth recording the wrong assumption too, because it is what made this look impossible for two days: I took "pinned by digest" to mean the UPSTREAM digest had to be preserved. It does not. Any digest that is exact and immutable satisfies the rule, and a registry assigns one. |
||
|
|
087a8f4144 |
Close 009: a sealed machine now pulls by digest
The resolution was the one the issue predicted -- a registry inside the scenario -- and it is the real path rather than a stand-in, since that is what every node after the first pulls from. The digests are the lab registry's own, which satisfies the pinning rule: what is required is a reference that is exact and cannot move, and one this registry assigned is both. That was the insight that unblocked it; I had assumed the upstream digest had to be preserved, which is what made it look impossible. The fault worth keeping is recorded in the issue: the read-back checked that the catalog endpoint answered by matching the substring 'repositories', which an empty catalog also contains. It passed on a registry holding nothing. This repository's own subject, arriving in the tooling built to catch it. |
||
|
|
b4607dfc03 |
Numbers are identity; the reading order is a generated, checked index
Decided after measuring what renumbering actually costs: 96 references in code comments across two repositories, none of which would have failed to compile. They would have pointed at the wrong reasoning, which is worse than a broken link because nothing reports it. So a number identifies a record and never changes. It cannot also be a position -- a position moves when the set changes, and an identity that moves is not one. The reading order moves into an index generated from each record's `topic:`. Six topics, in the order somebody learns the system. The index is WRITTEN rather than only generated on demand, which reverses what this repository previously said. The reason it said otherwise is that a hand-written index drifts -- but a reader looking at the folder on a forge sees the folder, not a command, and the drift objection is answered by checking rather than by refusing to write one. That is §5's own rule: a rule states how it is checked. Two checks, both confirmed to bite. index.py fails when the written order no longer matches the records. records.py fails when a record has no topic or one nobody defined -- the quiet failure being a record that vanishes from the order rather than appearing in the wrong place. |
||
|
|
333356cff3 |
Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when things happened to be decided, which after consolidation is fictional anyway since record 5 alone folds decisions taken across a week. Concretely wrong before: the domain statement sat at 8, after five engineering rules; the constitution was scattered across 5, 12 and 17; the tiers landed at 15, 16, 21 and 22 with process records in between. Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what runs on them and how it gets there (9-10), how it is built (11-16), how it is checked (17-18), how we work (19-23). Two things made this safe rather than free. It is a permutation, not a compaction, so the renames go through temporary names -- otherwise two files want one slot and one is lost. And the reference rewrite is a single simultaneous pass, because almost every number moved into a slot another number was vacating; replacing one at a time would have cascaded and pointed things at the wrong record while still resolving. Verified: 284 [ADR NNNN](path) links across the repository, all with matching text and target. The ordering principle is now stated in 19 rather than left implicit -- the repository already said "the numbering is the flow" about its folders, and there was no reason for the records to be the exception. |
||
|
|
e1febe8e0f |
Renumber the records 1 to 23
The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18, 19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only the archaeology of what used to be there. Renumbered contiguously. Renames run in ascending order, so every target number is already free and no two files ever collide. The reference rewrite is one simultaneous pass rather than a sequence of replacements. Numbers moved into slots other numbers were vacating -- the node host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time would have cascaded and silently pointed things at the wrong record. Seven plain-text references survived the merges as prose rather than links, naming records that no longer existed: the enrolment token, the link boundary, what a declaration is, reachability, the repository structure. Each mapped to the consolidated record that now holds it. Verified rather than assumed: every [ADR NNNN](path) link now has matching text and target, checked across the whole repository, and the checker passes. Frontmatter `consolidates:` lists dropped -- they named records that are gone, and each consolidated record already says in prose what it absorbed. |
||
|
|
77f3a4cea7 |
Consolidate: 65 decision records to 23
Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.
the node host 8 -> 1 applies not decides, depends on nothing,
per operating system, root service, the
launcher, episodic, what a declaration is,
actions from the bundle only
a node and how it joins 4 -> 1 what a node is, joining, the link as
security boundary, the enrolment token
modules and the graph 7 -> 1 everything is a module, no domain modules,
three edges, provisioning, the core library
substrate and control 6 -> 1 the test, seven contexts, one control plane,
plane the authority is not a database, the named
products, the pinned bundle
connectivity 3 -> 1 a route is a grant, reachability declared,
filter rules
delivery 5 -> 1 reconciliation not a pipeline, artifacts,
the three silos, a failed step, the verdict
the lab 5 -> 1 (earlier)
how this repository 10 -> 1 (earlier)
works
Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.
The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.
The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
|
||
|
|
5e83ac2c22 |
Consolidate: 65 decision records to 52
Jochen: a normal application has 3-5 ADRs, maybe 10 for a large one, and we are at 65. Fair, and the cause is mine -- I recorded every FINDING as a decision rather than every fork in the road. Two merges, both cases where one decision had been split across many records because it was taken over several days rather than at once. 0019 absorbs ten records about how this repository works: what it is and that it is public, the folder flow, the two design layers, the issue front door, status in frontmatter, playbooks, the naming rule, the product name. Those were never ten decisions -- they were one, seen from ten angles as the repository took shape. 0016 absorbs the five about the lab: a node is a virtual machine, a router is scenery, a scenario declares the underlay, a scenario is a closed address space, and the two scenario classes. Same pattern -- one design, split by the order it was worked out in. The consolidated 0019 also raises the bar for what earns a record, since that is what produced 65: a record is warranted when there is a genuine fork -- a direction reversed, an alternative that will be proposed again, something contested. A finding is not a decision, and a bug is certainly not. Everything else belongs in the design document where the reasoning is actually read. The checker earned its place here. Deleting nine records left 13 dangling links across the repository and it named every one, including in AGENTS.md. Nothing was found by reading. Remaining clusters worth the same treatment: the host (8 records), delivery (5), modules (6), connectivity (4), substrate and control plane (4). That would be 52 down to roughly 30. |
||
|
|
10365f2eae |
Consolidate the design layer: one place per topic
Jochen: a jungle of specs that slightly contradict or patch each other, and what matters is a working state rather than history. Both are fair and both are mine. Measured rather than assumed. 05-the-node-host and 09-the-node-lifecycle both covered enrolment, the install commands, the unit file, the launcher and reconcile -- I wrote 09 without taking anything out of 05, so the same things were said twice and could drift apart. Split by what each document IS. 05 is the component: what the host is, its parts, the declaration vocabulary, the build order, how it is verified. 09 is what happens to it: install, enrol, run, upgrade, retire. The whole "The process" section left 05, and the unit file moved to 09 where installing is described. 05 goes from 338 lines to 245 and now points at 09 rather than restating it. 09 also carried a 105-line "Resolved" section -- six mechanisms framed as "these were open and here is the answer". The content is needed; the framing is history, and history is what makes a document read as a changelog rather than a description. Renamed to what it actually is and the was-open phrasing removed. Also added 10-delivery.md, which did not exist: four accepted decisions -- 0054, 0063, 0064, 0065 -- had no design document at all, which is the specific reason the delivery picture felt scattered. It is now one document covering modules, the three edges, the core library, and how a change becomes a running thing, with a table of what each property is designed against and what must exist before it can be built. |
||
|
|
9d091c81e0 |
A build edge, a core library that is a domain, and 0063 corrected
Three things from walking a real dev cycle through 0063, all of which Jochen caught by pushing on where I had glossed. 0064 -- a build edge is a third kind. Research 011 established presence and instantiation, and both are RUNTIME edges: they answer what a module needs in order to run. Delivery needs a different question -- what has to be rebuilt when this changes -- and that relationship is fixed inside an artifact rather than negotiated when it runs. So the graph as designed could not drive delivery, which is the real reason 0063 was not approvable. It is derived rather than declared, read from what a module actually imports, because a declared list and the imports it describes drift and the imports are the true ones. The runtime edges stay declared, and that asymmetry is not an inconsistency: a runtime edge is an intention somebody has, a build edge is a fact about code that exists. It also makes design quality measurable. A module with many inbound build edges is one whose every change is expensive, and the current shared library is exactly that -- nobody could see it because nothing drew the edges. 0065 -- the core library is the mesh's domain. Jochen disagreed with 0030's "types, not behaviour" and was right: that guard is aimed at the wrong thing. A library everything depends on is a hub whether it holds types or code, and the fan-in is what makes a change expensive. So types ship with the module that owns them -- trading one wide edge for several narrow ones -- and the core library holds what is true of the mesh regardless of context, which research 011 already found: a module, a node, an assignment. The test is "would this still mean the same thing in a context that had never heard of the one it came from". A node does; a pipeline stage does not. Domain-driven is the point rather than the label: "who else might want this" always answers yes, which is how the current one grew. And it changes the check for the better. "The build output contains no runtime code" would have enforced a rule now withdrawn. Inbound build edges is a measurement rather than a prohibition, and it is visible while a hub is forming rather than after. 0063 revised on both counts, plus a third: I had written "the lab judges it" as though that were a step. A lab run takes tens of seconds, occupies a VM, and fails for environmental reasons -- and a shared-library change produces dozens. One expensive non-deterministic gate fails both ways, and neither failure looks like itself. Verdicts are now tiered, and a run that failed environmentally is explicitly not a verdict. 0063 also now carries what must exist before it can be implemented, rather than leaving that to be discovered. |
||
|
|
4ab8a0507f |
Delivery is reconciliation, not a pipeline; research 008 closes
Jochen: don't rebuild the current coordinator, use it as a pitfall list. That reframed the last open question rather than answering it. 0058 stopped deploy being a stage that pushes to nodes, and said plainly what it did not fix: detection. A merge that created no pipeline, and nothing said so. That is not a defect in the detector -- it is what happens when correctness depends on an event ARRIVING. 0063 applies 0058's move one level up. The control plane holds what source exists and what has been built from it, and builds the difference. A change becomes a build because source is ahead of artifacts, which is a comparison answerable at any moment. An event makes it fast; nothing makes it necessary, so a missed webhook costs latency and cannot cost correctness. The mesh becomes one idea at two layers: the control plane reconciles artifacts against source, the host reconciles machine state against declarations. The pipeline as a state machine disappears, and with it the stage list that a verify step was once omitted from. That reframing answered the three questions still open in 008, so it graduates with all six closed. A deployed state is two comparisons rather than an event. A verdict is about an ARTIFACT and gates whether it may be declared -- sharper than the question expected. And "before self-hosting" mostly dissolves, because a reconciler needs source and artifacts as bindings where a pipeline's stages name their targets. Four costs recorded, and one is a real risk rather than a trade: a reconciler that cannot reach its target retries forever, and without something noticing, the failure is silence -- the exact fault this removes, reintroduced elsewhere. Also named: the run identity people actually use is lost, and "did my change go out?" needs a replacement or this will be worse to live with than what it replaces, whatever its properties. |
||
|
|
f728c3fd98 |
File 009: a digest-pinned image cannot be placed in the lab
Two accepted decisions collide, and testing found it rather than review. 0046 pins images by digest and has the host refuse anything unpinned. The lab places images by exporting them from the workstation, because a sealed scenario cannot reach a registry -- and that loses the digest, since a repo digest only exists for an image a registry served. Measured: the load says 'Loaded image ID:' rather than 'Loaded image:', and the image lands dangling. So a tag is refused by the host and a digest is unusable in the lab. There is currently no declaration the lab can raise that exercises the container shape, which matters because the container shape IS the substrate -- every bootstrap step past the runtime is one. The resolution is a registry inside the scenario, and that is not a workaround: 0048 already names an OCI registry as substrate and every node after the first pulls from the mesh's own. It also removes the lab's export-and-push mechanism rather than repairing it. 0046 now carries a pointer, since its own consequence is where the collision was predicted -- half of it is closed and the other half turned out to be harder than 'not solved here' suggested. |
||
|
|
9dc57b4712 |
Graduate 005; record what 0058 answered in 008
Continuing the sweep. Both were answered by records that did not cite them, which is the same pattern 003 showed -- an effort stays active because the decision that resolved it was reached from another direction. 005 graduates. Three of its four questions are answered: provider modules do not group (0044), the ~50 modules that co-change with nothing stay as they are, and 'group or leave' was never the right pair -- 0054 reframes it as authority versus package. Worth noting the debt runs the other way too: this effort's measurement, that reachability is the ONLY place modules genuinely co-change, is what 0054 rests on and why connectivity is a context while nothing else needed one. Its fourth question moves rather than closes. Whether applications leave the monorepo before or after they group is a sequencing question, so it belongs to 009-migration. 008 stays active, with its central question marked answered: the coordinator converges nodes on a declaration rather than dispatching stages (0058). The three-silo split survives with the third redefined. What 0058 explicitly does NOT answer is how a change becomes a pipeline reliably -- detection is upstream of everything it changed and remains the fragile input. |
||
|
|
0a37d751e2 |
Graduate research 003; file the rescue that does not exist
A sweep of the nine active research efforts. 003 was answered five days ago and nobody closed it -- the decision it asked for was taken without citing it, which is how an effort stays `active` after being resolved. Its recommendation is what the mesh adopted, and the match is exact rather than approximate. "Run the daemons as containers, making Docker the supervisor for everything" is ADR 0057. Its warning that a mesh-native supervisor inherits fate-sharing "unless it sits outside the mesh's own process tree" is where ADR 0061 put the launcher. And its insistence that it cannot be all-or-nothing is why the host itself is the one thing an init starts. Its incidental finding does not graduate with it, so it is now issue 008: the automatic node rescue the documentation describes does not exist. No unit declares OnFailure=, nothing calls the rescue script on a timer. That is worse than having no rescue. A rescue nobody wrote is a gap somebody can see; a documented one that is absent is a gap nobody looks for, and the documentation is read exactly when a node has failed and somebody is deciding whether to intervene. The issue names two honest resolutions -- implement it, or delete the documentation and say a failed node needs a person -- and says the choice is scheduling rather than technical, since the new host's recovery is built and tested. It also says what would make the finding certain: it came from reading the repository, and confirming it on a running node is the difference between "no unit declares this" and "no unit in the source declares this". |
||
|
|
ba0d01788e |
0062: a host may be episodic; 0060's Android gap closed
0060 named the gap and did not close it: everywhere else an init runs the launcher at boot, and Android grants neither an init to register with nor anything worth supervising, because a supervisor would be killed alongside what it supervises. Closed by narrowing what is required rather than building something. A host is resident or episodic, and both are hosts. Being killed by the platform is disconnection, which 0036 already made ordinary -- and every mechanism an episodic host needs already exists because it was built for laptops that close. A partial host can join a mesh and cannot be the first node, since every bootstrap step is a shape it refuses. Its bundle says so. Two consequences that are easy to miss: last-heard-from means much less on an episodic host, so a healthy phone reads as a dead server unless the reader knows which kind it is; and a declaration may take a long time to land, which makes 0058's outstanding-versus-failed distinction load-bearing. Still open, and in that order: what an Android node is FOR, and only then how it is started. |
||
|
|
f1b1cd9aa0 |
Review: three ADRs no longer said what we had concluded
A sweep for claims overtaken by the last few days. Annotated rather than rewritten, following the pattern already in 0049 -- what changed and why is the useful part, and an accepted record should not quietly become something else. 0057's init section was wrong on all three of its claims. It said the host needs FOUR things from an init; 0061 reduced that to one. It said every machine the mesh targets already has systemd; Alpine does not, and it is the intended first node. It said there is no second init to abstract over; there is now, and the answer is still not an abstraction -- it is a four-line file per system. What survives is the part that was always right: an init is not a dependency in 0041's sense, because it is not installed, it is what the machine already is. 0048 named Docker as the container runtime. It is now docker or podman, detected rather than chosen -- because adoption keeps what a machine already has, so naming one contradicted a rule already decided. That row is the only one of the five that names two, and the record now says why. 0060 claimed the bundle is portable across operating systems. Its mechanism is; its contents are not -- package names, unit names, service names all differ, so an Arch host embeds an Arch bundle. That was my error, and it is the exact confusion behind the question that found it. The design layer had the same drift: 07 and 09 said "Docker" where they meant a container runtime, 09 said systemd restarts the host after an upgrade when the launcher does, and both install snippets assumed Arch. They now show Alpine and Arch side by side, which makes the point better than prose did -- step 1 differs per system, step 2 never does. Checked and NOT changed: 0047's "the vocabulary grows by one shape" is a claim about the rate, not the count, and is still true. 0037 lists docker among tools the host manages, which it does. 0041 says nothing about either. |
||
|
|
66df0eb53e |
0061: the launcher supervises; init is asked for one thing
The record said an init is asked for two things -- start at boot and restart on exit -- which was half a change. It moved the give-up logic out of unit files and left the restart in one, so init still decided when the host came back. The launcher no longer execs the host. It supervises it, so restarting is ours too, and init is asked only to run it at boot. There is an OpenRC script beside the systemd unit now. Records the cost honestly: not exec'ing means the launcher must trap the shutdown signal and pass it down, because a supervisor that exits while its child runs leaves the host to be killed rather than to stop. And records what the implementation found: the counter counts consecutive FAILURES, not starts. Counting starts meant a host that upgraded itself three times rolled itself back, having worked perfectly every time -- because a clean exit IS the upgrade path. That is now the second time a clean exit has been mishandled, so it is called out as the thing to check. |
||
|
|
c557f99cba |
Record what testing podman actually showed
0060 said the container runtime was a separate decision. It is now made, and the reasoning is worth keeping because it is the opposite answer to the same question one paragraph earlier. Abstracting service managers is lossy -- systemd and OpenRC are different models and LoadState has no equivalent. Container runtimes converged on one CLI deliberately, so almost nothing is lost: checked against podman 6.1.0, run, rm -f and docker's own template syntax for state and labels all work unchanged. Only the probe differs. So: a two-entry lookup, not an interface. The difference that is NOT in the CLI is the one that would have shipped silently. Podman accepts --restart unless-stopped, records it, and has no daemon to act on it -- containers do not return after a reboot unless podman-restart.service is enabled, which by default it is not. Every command reports success and the effect does not happen. That belongs in the declaration rather than the host: a node using podman is told to enable the unit. Which is what made the service shape's missing 'boot' field visible, and it is now built. |
||
|
|
e1ad39b500 |
Per-OS hosts, and an init asked for only start and restart
0060 -- the host is built per operating system. systemd and pacman are the Arch host's implementation, not abstractions the mesh has to grow. They are not independent choices: a machine has pacman because it is Arch, and the package manager, service manager and packaging format arrive together as one decision somebody made at install time. Rejected abstracting them, and the reason is correctness rather than effort. The service applier reads LoadState to tell "not installed" apart from "stopped", which is what stops it reporting absence as success. An interface spanning systemd and OpenRC degrades to what both express, and the lowest common denominator is exactly where that fault lives. Almost all of it is shared -- the vocabulary, store, apply loop, read-back discipline, refusal model, bundle and link are portable. Two appliers differ. And delivery was already per-OS, since a .pkg.tar.zst is an Arch artifact, so this is the seam that already existed. Android is the interesting case rather than Debian: no service manager, no package installation, usually no root. Such a host implements file, directory and action and refuses the rest -- the same refusal a host already gives an unknown type, with a different reason. Those three are the portable floor. The container runtime is deliberately left open: it is not an OS split, since Arch runs docker or podman. 0061 -- the init is asked for start-at-boot and restart-on-exit, and nothing else. Both are expressible in OpenRC, runit, s6 and an Android init.rc. Counting failed starts and rolling back moves into a launcher, because that is the one piece which must work when the host does not, and a script with a counter can be tested where OnFailure= can only be hoped for. Supersedes 0059, keeping its reasoning in full. The checker found all six places citing 0059 and refused the commit until they named the replacement. |
||
|
|
dcc4b8339c |
Say who consumes the broker and who writes the registry
Left implicit by the previous commit, which said the owning context writes without saying what does the consuming. The control plane is the consumer, and there is one of it. Seven contexts but one deployable, so it is one process dispatching internally rather than seven consumers racing -- which matters because the as-is records two consumers accidentally sharing a queue and silently splitting the traffic, each getting half of what it expected. With one consumer that cannot arise. The broker is also the buffer while the control plane is down: nodes keep publishing, messages queue, the control plane drains them on return. That is what makes a single control plane tolerable -- an outage delays the mesh's knowledge rather than losing it. One consequence named because it will otherwise be discovered: an unbounded queue grows until the broker's disk is full, and the broker is the component every node depends on. The bound is per queue and undecided -- dropping the oldest health report is obviously right, dropping the oldest declaration acknowledgement is not. |
||
|
|
19997d56c3 |
Approve 0057-0059, with four corrections from review
Not approved as drafted -- four things came out of checking them against each other, and one was a bug that would have broken every upgrade. The bug: 0059 specified Restart=on-failure while 0057 has the host restart onto a new binary by exiting CLEANLY. on-failure does not restart a process that exited zero, so every upgraded node would have been left stopped, having successfully upgraded. Found by reading the two records against each other rather than by either alone. Now Restart=always in all three places that mention it. The host cannot run in a container, and the reason is decisive rather than stylistic: step 0 of the substrate bootstrap installs the container runtime, so a host inside a container would need the thing it exists to install. It would also break 0041 -- copy it onto a machine and run it stops being true when the machine must already have a runtime. Everything above tier 0 is a container; the host is not. That split is the tier boundary, not an inconsistency. systemd is named rather than abstracted. An init is not a dependency in 0041's sense: 0041 is about what must be installed before the host works, and an init is not installed, it is what the machine already is. The unit file is the only systemd-specific artefact and it belongs to the package, so a machine with a different supervisor ships a different package. The mesh is a watchdog, and my first draft was half an answer. Recovery must be local -- nothing dials a node, and a host that cannot start cannot report. But detection is the mesh's, and a local supervisor structurally cannot do it: it sees one process failing and cannot tell a broken machine from a broken release. Only something watching every node can, and that distinction decides whether the response is "fix this machine" or "stop shipping this version". So a host rollout is staged -- a few nodes, wait for heartbeats, continue or stop on silence. Local rollback still needed, because the canary nodes break and because a node offline during the rollout gets the declaration later with no batch around it. The first declaration is the overlay and nothing else. Forced, because a node's address and peers are assigned rather than chosen. But also the way back in: a node reachable over the overlay can be fixed by hand if a later declaration breaks it, and a large first declaration risks a node that is broken and unreachable at once. Also stated plainly, because it reads as a contradiction: nodes reach each other over the overlay and every node consumes from the broker; what 0039 forbids is an inbound CONTROL surface, not reachability. And in 06: no node holds a credential to any control-plane store, for reads or writes. Four ADRs already say this separately and none of them said it in one place. Nodes state over the broker; the owning context writes. With a note that most high-frequency writes are observability's, not the registry's -- routing logs into the registry would be the shared-schema mistake arriving through a door marked performance. |
||
|
|
605c9fd441 |
Changes are pushed, not polled; and a stuck host rolls itself back
Two corrections and one new decision, all from Jochen catching things. Pushed, not polled. I described updates as landing "on the next reconcile", which reads as polling and is not the design. A declaration arrives as a message on a link that is already open; the host applies it then. Polling over an existing connection would be slower to land AND constant traffic to learn nothing. The timer is for drift and nothing else, and it cannot be replaced by an event for a definitional reason: drift is change the mesh did not make -- somebody edited a managed file, a distribution upgrade replaced a config -- so nothing will ever publish a message about it. Only looking finds it. Separated the heartbeat from the reconcile timer, which I had been conflating. They point in opposite directions and answer different questions: the timer looks at the machine and asks whether it still matches; the heartbeat reports upward and is what makes silence mean something. A node with nothing to do sends nothing, and without a heartbeat that is indistinguishable from a node that stopped. 0059 -- a host that cannot start is rolled back by the service manager. I had left this open on the grounds that recovery meant the host judging its own health. That objection does not survive being asked properly: a keepalive is something else judging the host. The watchdog must be local, because nothing dials a node and a host that cannot start cannot report -- so it is the service manager, which is already there. The failure it prevents is sharper than "the node is down": a host that will not start looks exactly like a machine somebody switched off, which is the one condition this design has deliberately decided not to alarm on. So a bad release reaches every node, each goes quiet, and the mesh reports a fleet of sleeping laptops. Confirmed means started and completed one reconcile -- deliberately not "the link is up", or a laptop on a train would roll itself back. The rollback is a script shipped by the package, not a host subcommand, because a binary that will not start cannot be its own recovery. It rolls back once: a second failure means the machine is the problem, not the binary. Also refined the records checker, which produced a false positive: a proposed record may extend another proposed one, because decisions are drafted in chains and the alternative is marking things accepted to satisfy a check. An accepted document resting on a proposed record still fails, and that was verified. 0057, 0058 and 0059 are all proposed. |
||
|
|
aeea2a9f9a |
Resolve the host lifecycle's open items, and say how the host is delivered
The upgrade question turned out to be a delivery question, so 0058 answers both. Today's third silo runs once per node and sends each one a command to install and start. That is where the as-is records a package install that 404ed from every mirror while the job went green, an image pull failure that did not fail the deploy, and a verify stage that was built and never scheduled because it was missing from a list. The shape underneath all of those is that the thing reporting success was not the thing doing the work. Meanwhile ADR 0037 has given every node a component that applies state, reads back and reports -- so two mechanisms now change a node and only one checks its work. 0058: a pipeline ends when the declaration is updated. Deploy stops sending commands to nodes and becomes one write. The host applies it on its next reconcile, and the host cannot report success it did not verify. The verify stage disappears as a stage, which is the point -- verification stops being a step that can be left off a list. A pipeline result now means "the declaration is updated, and here is which nodes have applied it". It does not wait for every node, because a node may be legitimately switched off for a week. Outstanding is reported separately from failed, since conflating them is how the old system produced a stall with no error anywhere. The host is delivered by exactly this path and needs no new resource type: a `file` writes the package manager's config pointing at the mesh's repository, a `package` names the version. Added a step I had missed -- before exiting for a restart, the host runs the new binary once. A package can install something that does not execute here, and that turns "the node never came back" into "the apply failed and said why". Six open items resolved: re-enrolment is decided when the token is issued and revokes the previous identity; the mesh keeps a recovery copy of what each node reports it owns, which un-strands the orphans; last-contact is reported with no threshold, because a laptop off for three weeks is doing nothing wrong; adoption always completes but a failed line makes a node ineligible for assignment; a briefing is a structured document whose outcome is computed from its lines; and the token is printed once and carried by hand, which is the property that makes it worth anything. Still open and named: automatic rollback of a host version that will not start. 0057 and 0058 are both proposed. |
||
|
|
2204b01909 |
Design the node lifecycle end to end
The host was described as a component and never as something that runs for years on a machine somebody else also uses. 09 covers every state a machine can be in and every transition between them. Four states: unmanaged, hosted, enrolled, disconnected. Only the last two are nodes, and they are the same node in two situations. `hosted` -- the host installed but never told which mesh it belongs to -- had no name before and is where a machine sits between the two adoption commands. Things that were unclear and now are not: The first node walks the same path in an unusual order: reconcile from the bundle, the control plane it just raised issues a token, enrol against it. Its specialness lasts two commands. A side effect worth having -- enrolment is exercised on node one, rather than being written and first used on node two. Enrolment reports profile and inventory BEFORE the control plane decides anything. The profile is the input to that decision, not a diagnostic; the control plane cannot decide what a machine should run without knowing what it can run. Rebooting mid-apply is safe by construction. The store records each resource after it worked, so a host that dies half way through comes back and applies the rest. The rule that stops the host lying about what it did also makes it crash-safe. Retiring splits in two. Graceful is a final empty declaration. A node that is gone will reconcile its last declaration forever -- the honest consequence of making disconnection ordinary. The answer is not to make the host expire but that the node holds nothing that outlives revocation: every grant is a per-node credential revoked at the provider. A lost node keeps running and stops being able to reach anything. Said plainly rather than implying the mesh can switch a machine off, which it cannot and should not. Losing the store is quiet and permanent, so it gets its own section. The host re-enrols and re-applies fine; what does not come back is removal, because resources it no longer has a record of become unowned and sit there indefinitely. Also corrects 0057, which said the mesh must not upgrade the host at all. That conflated two acts. Replacing the binary is safe -- Unix keeps the running inode. Stopping the unit is not. So the host may apply a package naming itself, and restarts by finishing its apply and exiting cleanly, letting the supervisor start it on the new binary. It never asks the service manager to restart it. That makes a fleet-wide host upgrade an ordinary declaration, which the first draft gave up on. 0057 remains proposed. |
||
|
|
3ab11c96ef |
Say what the host process is: a root service, installed as a package
The design described what the host does and never what it is at runtime. The words daemon, long-running, interval, poll and heartbeat appeared nowhere in it or in the relevant decisions. What exists is a command that runs and exits; what the design needs is a process holding a link. Nobody had written down that those differ, so several questions had no answer. 0057 settles them. It runs on every node -- the host is what makes a machine managed, so a machine without one is not a node. Root, because no useful subset of the job is unprivileged. A systemd unit, because something must survive a reboot to hold the link. It never manages its own unit. The temptation is obvious and it ends with a host stopping itself half way through an apply, leaving a machine with nothing running to fix it. The installation owns the host; the host owns everything else. Installed as a package, with a tarball as the floor. The package carries the unit file, the state directory and an upgrade path, which a bare binary does not. But the mesh's package repository is hosted on the mesh, so any route that needs the mesh to install the thing that joins the mesh is a circle -- the tarball is the path that must never acquire a dependency. Reconciles on start, on a declaration, on a timer and on reconnect. The timer is the one easy to leave out, and without it `owned` reports what the host applied rather than what is there -- ADR 0035 violated by omission. The records checker caught this commit on its first attempt: 05 listed 0057 in its frontmatter while 0057 is still proposed, and a to-be document may not rest on an unaccepted record. The section now says so in the body instead. |
||
|
|
2330d74c1b |
The host's vocabulary is complete; 05 and 07 said otherwise
All six shapes are built. 07 still said the last three did not exist, and 05 still described stage 2 as having built three of six. Records what the lab still cannot do, because that is now the only thing between here and an end-to-end substrate bootstrap: a sealed scenario cannot fetch an image and its machines carry no container runtime, so package, container and action were verified against a real machine instead. |
||
|
|
03874f3fe2 |
Add a structural check over HQ's own records
Nothing in this repository was verified by anything but reading, which is how a superseded decision stayed live in the constitution and in the to-be README at the same time. Both were found by a person looking, and nothing stopped a third. Five checks: links resolve; `decisions:`/`extends:` name records that exist and are accepted; a governing document citing a superseded record must name its replacement in the same paragraph; supersession is symmetric; filename number matches heading number. Each was made to fail before it was made to pass. The live-citation check was verified against a reconstruction of the actual incident -- the to-be README citing ADR 0017 as live guidance -- and reports it with file and line. It found one thing nobody had noticed: ADR 0018 never declared that it superseded 0011, though 0011 has named 0018 as its superseder since August. Fixed. Deliberately not checked, and said so in the README: 02-DECISIONS and 01-RESEARCH may cite superseded records freely, because a decision record discusses history and research records what was observed. 00-as-is may rest on one, per 0056. Flagging those would put noise on correct documents, and a check that cries wolf gets suppressed -- which costs more than not having it. Two bugs found by running it: the frontmatter reader iterated an inline list as characters, and the as-is exemption was missing entirely. |
||
|
|
e1f4c7d9e0 |
Approve 0054-0056, apply them, and fix the two smaller findings
0003 is now superseded by 0056. Nothing is left proposed. Applied: - 06 corrected from ten contexts to seven plus the api, each row now stating why it passes the more-than-one-node test. work, knowledge and stream are named as mesh-hosted rather than dropped; `ai` folds into config; `record` is deferred explicitly rather than listed. Its frontmatter now cites 0055. - how-we-build §4 amended per 0054, and the derived page republished by playbook 05. The sync found the drift the playbook exists to catch: the published §4 and the source did not say the same thing. The source said "four accidents, not four boundaries"; the published page said "one intent expressed four times", and only the published page carried the scope caveat. Same rule, two texts, already diverging. Verified the republish by reading back -- the new rule is present and the old section's body returns nothing -- rather than trusting the success message. The two smaller findings: - 0051 separated the transport identity from the declaring authority. It said the token carries "an address" and "the identity to expect" without saying what the node dials. It dials the broker, so pinning only that would make the control plane's authority transitive and let a compromised broker forge declarations -- which, since the host applies whatever the link delivers, is the whole machine. The token now carries four things, and declarations are signed and verified per declaration. Cost recorded: rotating the signing identity is fleet-wide. - 0026 no longer restates 0022's rule about generated views. 0022's own words are "prose does not restate status; one place, and two is one too many", which is what 0026 was doing to it. |
||
|
|
f49d177a31 |
Draft three records for the contradictions the review found
0054 -- things that change together share an authority, not a package. The constitution instructs agents to group "how a node is reachable" into one module, citing superseded ADR 0017; ADR 0044 says there is no networking thing to install. Since the constitution is injected where work is decided, the superseded rule is the one actually steering work. The observation behind it was right -- research 005 measured that reachability is the only place modules genuinely change together -- but the conclusion was wrong: tight coupling means a shared authority, not one artifact. wireguard and traefik deploy to different node sets, so the merged module would be assigned where half is unwanted. Requires amending how-we-build and republishing the derived page. 0055 -- the control plane is the node-coordinating contexts. Three context lists were in circulation (0015 says nine, 06 says ten, the README said eight) and none was decided. Research 006 said explicitly that the change "belongs in a new record -- not written here", and the design used the list anyway. Reconciling them shows `stream` and `ai` were dropped with no reasoning at all. Applying 06's own test -- needs to know about more than one node -- gives seven contexts plus the api, with work, knowledge and stream as hosted applications and `ai` folded into config as an ordinary grant. The record defers rather than lists. The cost is stated rather than reassured away: a board composing across the boundary reads more than one interface. That was raised before as "only moves the problem up a layer", and the answer is that 0045 already requires surfaces to read interfaces rather than stores -- what changes is the count, not the kind of work. 0056 -- the authority is the control plane, not a database. Every clause of 0003 has been decided against in four separate records and it is still accepted and cited as live. The error underneath is the same category error 0054 corrects: "source of truth" named a storage location when it meant an authority, and once the store is the answer, shared schemas follow. The half that was right -- the repository defines what exists, the mesh defines what runs where -- survives untouched. Best consequence: the cache mode disappears, so a node that has not heard from the mesh is no longer indistinguishable from one that has. 0003 is left accepted until 0056 is. |
||
|
|
ef5dd0751b |
Approve 0049-0053; drop a to-be item superseded by ADR 0044
The 'domain grouping' item cited ADR 0017 as live guidance. 0044 superseded it -- there is no domain module to group into, so there is no domain list to settle. |
||
|
|
ccbbfa9c8a |
One node runs the control plane, and nothing takes over
Closes the two open questions in 06 and 08, which turned out to be one question: how many control planes run, and what happens when the hub is down. Both were drifting toward redundancy by default -- a standby plane, a second hub, an election to pick between them. That is not one feature but a property every layer must then honour, and each layer gets it wrong independently. Not wanted, and not needed. A handful of machines with one node hosting the registry is not a distributed system. The argument for why this is sound rather than merely cheap is that the design already tolerates it by construction. ADR 0036 makes reachability state rather than class; the host reconciles from its own store (0043) and never needed to ask anybody to hold the state it was last given. So the control plane being down is not a new failure mode -- it is every node in the ordinary disconnected situation at once. What is lost is change, not operation. No node holds a contended role: the control plane is assigned like any other module, and the overlay hub is declared (0050). No promotion, no quorum, no fencing, no split brain, no replicated store, and no "which node is authoritative" recurring at every layer. Two consequences stated plainly rather than buried. The control-plane node is a single point of failure -- deliberate, and said out loud so it stays deliberate. And recovery is restore rather than failover, which makes backup the availability story rather than hygiene. The sharpest one is the clock: the control plane owns certificate issuance (0049), so an outage outlasting a renewal window expires every public name. That bounds how long recovery may take, and nothing measures it today. |
||
|
|
4e80820e2f |
Design connectivity in full: overlay, resolution, exposure, filtering, certificates
Written as one document because the five are one design. They share inputs, they must agree, and every one of them today is computed in a different place by a different module from a different copy of the same facts. The through-line is that none of the five can be answered by a machine alone, so all five are decided centrally and delivered as `file` resources. That costs no new host vocabulary and removes both remaining direct database connections from nodes -- wireguard and traefik are the only two, and both are connectivity. Three decisions fall out, all proposed: 0050 -- reachability is declared, not inferred from an address. The RFC1918 regex is wrong for carrier-grade NAT (100.64/10 tests as public, so an endpoint is written to an address nothing can reach), wrong for IPv6, and wrong for a routable address behind a closed firewall. The lab needing TEST-NET-3 to satisfy the regex is the same bug from the other side. Also kills hub election by address prefix, which fails silently and makes renumbering an outage. 0051 -- the enrolment token carries where the mesh is and how to recognise it. Closes two circles with one mechanism: verifying the mesh needed the CA, and obtaining the CA meant trusting whoever handed it over; and a node had to reach the mesh before it could resolve any mesh name. An address plus a fingerprint, carried out of band, resolves both -- and closes the CA question 0049 deferred. 0052 -- a filter rule names its source. `scope:` is declared in five manifests, is part of no rule type, and is referenced by no code, so those manifests appear to restrict ports and restrict nothing. Removed rather than implemented; the general fix is refusing unknown keys, which the host already does and manifests do not. Also corrects two claims in 0049 asserting wireguard was already handled. Research 006 says both modules still reach upward; neither is. |
||
|
|
8d9282d86b |
Resolve the ingress gap: a route is a grant
ADR 0048 named ingress as an unclosed hole -- nothing said what terminates TLS, how a public name reaches a container, or which tier owned it. Resolving it needed no new concepts, which is why it survived: nobody had applied the rules already written to it. Ingress is not substrate. The control plane does not need a route to start, and no node needs one to reach it -- the node dials out and has no listening control surface. It grants itself a route afterwards, like a bucket. A route is an instantiation edge under ADR 0044. The direction mirrors a database -- the consumer supplies a target and receives a name rather than credentials -- but it is the same edge. The substantive finding is that exposure is three facts at two scopes: name resolution and certificate issuance need to know which node is publicly reachable, and only the proxy mapping is a single machine's business. That is why it belongs to the connectivity context, and why Traefik doing all three on the node is wrong. Which matters beyond tidiness: research 006 counted traefik as one of two modules opening a direct Postgres connection, reading nodes and mesh_ca. That violates 0037, 0045 and 0039 at once, and is why every node permanently holds a credential to the control plane's database. Deriving the config centrally and delivering it as `file` resources removes it, costs zero new host vocabulary, and closes the set 0039 identified -- wireguard was the other. Left open deliberately: the mesh's internal CA is the other thing traefik reads, and it belongs to the link's mutual authority, not to exposure. Conflating the two is what made the gap hard to see. Also fixes an inconsistency from the previous commit: 06 still claimed the virtual host was raised from the bundle. Proposed, not accepted -- for review. |
||
|
|
4d19e93900 |
Name the substrate's actual products
The design layer described every service by role and never once by name: Postgres appeared in zero design documents. That was over-application of the research rule "never identify the mesh it observed", which is about node names and domains, not software. Two things were actually broken by it. substrate.lock pins images by digest and a digest belongs to a named image, so the bundle could not be written from the design. And a reader could not tell a settled choice from an unexamined one -- "a relational store" reads identically either way. ADR 0048 names them: PostgreSQL, LavinMQ, MinIO, an OCI registry, Docker. The argument for each is continuity, which is a real argument -- replacing a substrate service migrates the mesh's own state. Role and product are now both written, because the design depends on the protocol while the installer needs the product. Also separates two questions the substrate doc had merged: being substrate and being in the bundle. Only Postgres must precede the control plane; the rest are substrate by role and ordinary by delivery. Whether the bus joins it is left open, because it turns on the control plane's internal shape. Names the forge as Gitea, and records ingress/Traefik as an unclosed gap rather than a naming one -- nothing says what terminates TLS or which tier owns it. Fixes a miscount: the host's bootstrap vocabulary is six shapes, not five. |
||
|
|
c631cbd07c |
The bootstrap starts a step earlier than recorded
Asked whether postgres has to be installed, and the answer exposed a missing step. The store is a container, so something must run containers before anything else happens — and a container runtime is a PACKAGE, not a container. Step 0 is where several threads meet. It is what the host's capability detection already reports, and the first use of that report by something other than a person. It is adopted rather than installed when the machine already has a runtime with configuration somebody chose. And it is a package, needing the machine's own package manager and a network, both of which ADR 0046 permits. So the host's bootstrap vocabulary is six shapes: package, container, file, directory, service, action. Stage 2 built three of them. The node host design now names which three remain and why the lab cannot yet exercise them — a sealed scenario fetches nothing and its machines carry no container runtime, which is lab-installation work rather than a constraint on the design, because production machines have a network. |
||
|
|
93470f6162 |
ADR 0047 — the bundle may carry actions the link may not
The bootstrap's sharpest open question, and the framing was wrong. "State on this machine" was being read as the filesystem and the service manager. A service running on this machine IS part of this machine — writing a file and creating a database in a local store differ in mechanism, not in scope. The real question was underneath: must the host learn what a database is? It must not. Giving it a `database` resource type means tier 0 knows Postgres, then a bucket, then a virtual host — the host acquiring the substrate's vocabulary one service at a time, which is what ADR 0037 exists to stop. So the bundle declares an ACTION and the host runs it and verifies it. What a database means stays with the module that provides one; the host knows only how to run a declared action against something local and check the result. Its vocabulary grows by one shape rather than by one resource type per service. Actions are permitted in the bundle and forbidden over the link, and the asymmetry is deliberate. A bundle arrives WITH the binary: anyone able to put a hostile action in it could equally have put it in the host itself, so refusing actions there buys nothing and costs the bootstrap. The link is a separate party, reachable separately, and an action there is the unbounded blast radius ADR 0039 refuses. That decision stands unchanged. And ongoing provisioning is not the host's at all — the control plane does it once a mesh exists — so the asymmetry costs nothing. Which dissolves the earlier worry about one mechanism with a tier boundary inside it: there are two mechanisms, with different actors, scopes and trust models, and that is the answer rather than a compromise. Named rather than hidden: this is the escape hatch research 011 warned about, arbitrary code in the place hardest to remove later. It is bounded by being bundle-only and by every action having to declare how it verifies itself, and that boundary is the whole defence. |
||
|
|
5b3d0ebd4f |
ADR 0046 — the installer fetches what it pins
The blocking question was where a container image comes from, and the version that blocked assumed the machine might have no network. That assumption came from the LAB: a scenario is a closed address space by design, which is what lets two scenarios hold the same addresses without meeting. Production is not sealed — a machine being adopted has a network, and one that does not is a machine where very little works anyway. So substrate.lock carries references, not payload: an image name and a digest, fetched at apply time. A first node pulls from upstream because no mesh registry exists yet; every node after that pulls from the mesh's own. The lab is the exception and places images itself, the way it already places the host binary — a property of a test environment, and letting it dictate the production design would be the tail wagging the dog. Pinned by DIGEST rather than tag. Reproducibility comes from pinning the identity of a thing, not from carrying its bytes, which is what makes fetching acceptable rather than a compromise. ADR 0041 survives untouched, which was the point. "Copy it onto a machine and run it" stays literally true — one binary, a few megabytes, which then fetches what it was told to. Carrying images would have quietly redefined the property that decision rests on. Costs accepted and named: an apply can now fail because something is unreachable, which a self-contained artifact could not, so it must fail legibly — naming what it could not fetch and from where. And the lab needs a way to place images into a machine that also has no container runtime, both of which are lab-installation concerns and neither solved here. Research 012's build-time-versus-apply-time reframing narrows accordingly: it still holds for what a tailored installer contains, and no longer has to hold for images. |
||
|
|
0531d6fc38 |
ADRs 0044 and 0045 — the module design, closed; 011 graduates
011 opened asking what a graph deletes and found the graph already existed. The work became design, worked through twenty cases and one provider in full. Two decisions close it. 0044 — a module declares presence, instantiation and exclusion. Two kinds of edge because a game wanting a database is not a game wanting postgres to exist: one creates something per consumer, carries credentials back, can be revoked, and leaves the provider holding state. Names are concrete unless providers are genuinely substitutable — `terminal` passes, `database` fails, and the adapter is what creates an interface. Where there is no contract there is a tag, which describes and does not bind. Exclusion is a third relation and is not derivable. A node provides names too, which makes capability checking stop being a separate mechanism and makes the host's detection an input to resolution. Constraints, never placement. Scope decides which provider and the binding is written down and sticky, in a place that follows the scope. And there are three entities, not two — the assignment carries what belongs to neither end, which is what node-agnostic modules ran out of. 0045 — a context owns its store, exclusively. No shared writes and no read roles on another context's store, because reading couples you to its layout just as firmly and invisibly. The unit is the CONTEXT, not the process: a board showing the mesh's own data is the mesh showing its own data. Asking or subscribing is derived from ADR 0036 rather than chosen. And it is the first clear list of what the design removes: grant kinds, table ownership, cross-context migration ordering, and a class of permission modelling. 0017 is superseded rather than narrowed — its text unchanged, its status changed. Folders assert relationships where edges record them, and the domain module goes with it. Left explicitly undecided in both: what a resolver delegates rather than reimplements, how many instances a module should have, and what a provider hands back. |
||
|
|
60aea14935 |
Define the substrate, and answer 006's four-or-five conditionally
Same gap as the control plane: load-bearing and unpinned. The substrate is what the control plane CONSUMES AND CANNOT GRANT ITSELF. Every module needing a database asks provisioning for one; the control plane needs one too and cannot ask itself, because it is not running yet. That circularity is not an awkwardness to work around — it is the definition, and anything on the wrong side of it must be raised by the bundle the host carries. Which answers 006's open question in the honest form rather than with a number. The identity provider is substrate only if the control plane DELEGATES authentication — then it cannot serve anybody before the provider exists and cannot grant itself a client. If it authenticates natively, the provider is an ordinary hosted service. So the count follows from a decision not yet taken, and asserting four was asserting that decision. The test also rules out the tempting wrong answer: an identity provider, a mail server and an analytics service are all infrastructure by any ordinary reading, and none are substrate, because the control plane starts and runs without them. Important is not the test. Records why the bundle is pinned by hand — it is applied when no mesh exists, so nothing can resolve a version or ask a registry — and why it must be self-contained, which makes it an artifact built on a machine with a network for a machine that may have none. |
||
|
|
148395ca54 |
Define the control plane, which was used 79 times and defined nowhere
Nineteen files, seventy-nine mentions, no definition. That is how-we-build §5 failing on this repository's own vocabulary — ubiquitous language is checked, not assumed. The definition, and it is not arbitrary: the control plane is everything that needs to know about MORE THAN ONE NODE. It follows from ADR 0037, which has the host applying rather than deciding precisely because deciding needs knowledge the machine does not have. So the line falls exactly there — writing a file is the host's, choosing which nodes run the store is the control plane's, and anything a single machine could answer alone does not belong here at all. That last consequence is worth having: putting a single-machine concern in tier 2 is a mistake the tier rule will NOT catch, because the dependency direction stays correct. Also states what it is not — not the thing that changes machines, not a surface, not the substrate, and not privileged on a node beyond what the declaration vocabulary allows. And the property that makes tier 2 unlike the others: it is itself a consumer, with the same requirements as any module, which is the circularity the bundle exists to resolve rather than hide. Scoped deliberately: this defines the term and does not design the contexts inside it. Ten is the skeleton's claim rather than a settled list, and research 006 still asks whether the record belongs here or in the substrate. |
||
|
|
a4ab3e15c2 |
011: one interface, many contexts — and the constraint that hides in it
The objection is right: if every context runs its own service with its own interface, the board is coupled to N interfaces instead of N schemas, something has to compose them, and composition is logic — which tier 3 says a surface does not hold. That moves the problem up a layer rather than solving it. The skeleton already answers it, and the previous entry talked past it. `work` and `knowledge` are not separate services; they are contexts INSIDE the control plane, alongside the record, inventory and delivery — and `api` is listed there as the one interface every surface speaks to. So the board speaks to one interface. Behind it the contexts stay separate, integrating through the record, but they are one tier, one repository, one deployable — and coupling within a tier is not what the tier rule forbids. The problem does move up a layer, and the layer it moves to already exists and has this as its job. The caveat is load-bearing and now recorded as an open question: this holds only while the contexts are not separate deployables. The moment one becomes its own service with its own interface, the board is back to N clients, something must compose them, and the composition has nowhere to live that tier 3 permits. That is a real constraint on how far the control plane may be split, and it is worth knowing before splitting rather than after. |
||
|
|
a4a25ca7e3 |
011: one surface over several contexts is normal
The board visualises the mesh, the work engine, the knowledge base and more, and the alternative — a web application per context — is worse for everyone using it. Composing several sources into one view is what a surface IS, so this is not a compromise with the ownership rule. What changes is only where it reads from: each context's interface rather than each context's store. Most of that already exists — 56 of 126 modules carry a tool surface, more than carry a service. And the unified board is what keeps those interfaces honest. A view that cannot be built from a context's interface proves the interface inadequate, discovered where it is cheap to notice rather than the first time something else needs the same data and quietly reaches for the store instead. If composing many calls proves too slow, the answer is a projection the board owns and keeps current from events, not access to somebody else's tables. |
||
|
|
fa62c7f0e4 |
011: correct the rule — contexts, not processes
An earlier version argued a dashboard reading a dozen stores was caught by exclusive ownership, because a dashboard is a surface and surfaces speak to an interface. Wrong, and it drew the line in the wrong place. The mesh's own board showing nodes, modules and deployments is not a separate context reaching across a boundary — it is the mesh showing its own data. Requiring it to go through an interface to reach facts its own context owns is ceremony. The rule is that a CONTEXT is granted what it exclusively owns. Everything inside it — service, surface, tools — reads that store freely. What is forbidden is a different context reading it. Which is what the consumer count already showed: the problem was never surfaces, it was three other contexts keeping their tables in the mesh's database. |
||
|
|
e71d532c2e |
011: request or subscription is derived, not chosen
Asked what the distinction actually is, and the SQL half needed correcting first: under exclusive ownership SQL runs against your own database and nothing else, whatever transport a query might travel over. Both options are the mesh's own channel and both ride the broker, so the transport is not the distinction. The distinction is where the answer lives when you need it. A request asks at the moment and waits — always current, costs a round trip, cannot answer when the other side is down. A subscription keeps a local copy — instant, works offline, as current as the last event received, and you must handle what you missed. What decides is not taste. ADR 0036 makes disconnection an ordinary situation rather than an exception, so anything that must keep working while disconnected CANNOT use a request: there is nobody to ask. And the converse — anything where a stale answer is worse than no answer cannot use a subscription. A display can lag; a decision about whether a grant is still valid cannot. So an apparently open question turns out to be derived from a decision already taken. What stays open is narrower: what a consumer does about the events it missed while disconnected — replay from a point, ask once for a full picture and resume, or rebuild. The question every projection has. Also recorded: separate databases are required in the new design, and the shared registry is a leftover rather than a pattern. |
||
|
|
7e83723b7b |
011: rewrite the question table, which had gone stale silently
Several edits to the overview matched nothing and returned success, so the question table still carried answers superseded two or three exchanges ago — "when two modules provide one name, who chooses" was still open in the table while answered in the file it pointed at, and nothing recorded the instantiation edge, instance counts, grants, bootstrap provisioning, the tool audience, or the registry consumer check. That is the fault this repository catalogues, committed by the thing cataloguing it: a string replacement that found no match, reported nothing, and left the document claiming a state it did not have. Rewritten from what the documents actually say rather than patched again. Nine questions settled, fourteen live, and the split is now visible instead of implied. |
||
|
|
afcc355744 |
011: checked the registry's real consumers, and the question was the wrong shape
The exclusive-ownership rule turned on whether every reader of the mesh registry could be served another way. Eighteen consumers open a direct connection. Four groups, and only one is work. The owner and its machinery keep reading, because they own it. The node appliers are already resolved — ADR 0037 stops the host querying the mesh database, decided for tier reasons with nothing to do with this. The bulk are FOREIGN TENANTS. The work engine holds ten of its own tables in the registry's database, the knowledge base two, pipeline logs one. Thirteen foreign tables across three contexts, which is how-we-build §4's shared schema counted. So the question was the wrong shape: the problem is not readers needing a new route to data, it is tenants needing to move out. Tasks, agents and teams have nothing to do with nodes and modules and are co-located by history. Give that context its own database and its dependency on the registry shrinks to one table. A handful of genuine cross-context reads remain, small enough to enumerate rather than estimate. The rule holds. Left open: whether those reads want an interface or events. Asking which nodes exist at the moment you need to know is a request; reacting when a node appears is a subscription, and some consumers want both. |
||
|
|
6b1aab6a1e |
011: the dashboard case, and why exclusive ownership is the tier rule
Raised as the hardest test of the rule: a board showing nodes, modules, pipelines, agents and tasks wants to read a dozen stores, and under exclusive ownership it can read none of them. It survives, and not by luck. The board is a SURFACE, and surfaces already may not do this — the skeleton puts `api/` in the control plane as the one interface every surface speaks to, and tier 3 as thin, no logic. A board reading stores directly is a surface reaching past the context that owns the data, which the tier rule forbids for reasons that have nothing to do with databases. So it is not a counter-example; it is an instance the rule catches. And the two rules turn out to be one rule seen from two sides: exclusive ownership is the tier rule expressed in terms of storage. The general shape for anything needing to see across many things: consume the record and own your own view. A reporting context builds a projection from events and reads its own store, never anybody else's. The cost said plainly rather than buried: a projection is more work than a join, and it lags. A board queries the mesh's own database directly today — ordinary, working — and this rule makes that a migration rather than a preference. The reason to pay it is §4's already-measured cost, not elegance. |
||
|
|
aa767d17a8 |
011: a module is granted only what it exclusively owns
Reconsidered by the operator — maybe shared databases should not be allowed at all — and the stricter version is better and goes further than the schemas it replaces. No shared writes, and no read-only role on another module's database either. Reading another context's tables couples you to its layout exactly as firmly as writing them does, and the coupling is harder to see because nothing breaks until the owner changes a column. That is how-we-build §4 taken at its word rather than at its letter. The permissive version — a per-consumer schema, revocable, with cross-context joins possible but deliberate — kept the letter and left the temptation. A boundary that is merely inconvenient to cross is a boundary that gets crossed. The cost is cross-module reporting, and it is the point rather than a regrettable side effect: anything wanting to know what several modules hold consumes their events or calls their interface. That is §4's whole argument, and the mesh already has both mechanisms. What gets harder is precisely the thing that was making work belonging to one context keep having to be implemented in another. And it is the first clear instance of what this effort has been hunting — what the design DELETES rather than adds. Grant kinds collapse to one: an exclusive resource. With them go the question of who owns which table, the guessing at revocation time, cross-module migration ordering, and a class of permission modelling a shared store would otherwise need. One thing it does not answer, recorded because it could make the rule unworkable: the mesh's own registry is read directly by many things today, and under this rule they consume events or call tools instead. Achievable in principle. Whether EVERY current consumer can be served that way is unchecked, and should be before this becomes a decision. |
||
|
|
4c8515507a |
011: where a binding lives follows the scope, and a grant is not always a whole resource
Two questions asked directly, and the second collides with a rule in force. One module on two nodes sharing a database corrects something stated flatly: the binding is not "recorded on the assignment". Where it is written down FOLLOWS THE SCOPE. A shared grant belongs to the module and every assignment references the same one — which is the answer for two nodes wanting one database between them. A per-instance grant belongs to the assignment. Same relation, two homes, and which home is what makes two instances share something or not. Several modules adding their own tables to one database is three needs wearing one sentence, and a provider offers KINDS of grant rather than one: a database for a consumer whose tables are nobody else's business, a read-only role for one that needs to see what another holds, and a SCHEMA within a shared database for the case actually asked about. Loose tables in a shared database is what how-we-build §4 warns against in as many words — several domains sharing one forty-five-table schema, which is why work belonging to one context keeps having to be implemented in another. Not a style objection; the observed cost, already paid. A per-consumer schema keeps what the request wants and drops what §4 objects to. Same database, same connection, same backup, and a cross-schema read remains physically possible when genuinely needed. What it adds is ownership: migrations touch one namespace, two modules cannot collide over a table name, and revoking drops the schema rather than guessing which tables belonged to whom. So the fault §4 names is still possible and no longer accidental — a cross-context join becomes something somebody deliberately writes rather than the path of least resistance. And revocation becomes answerable, which the whole-database version never was. |
||
|
|
fa9889536c |
011: tools have a different audience, migrations cross the edge, provisioning is early
Three additions, and the third kills an assumption. Tools are the most common content in the catalogue — 56 of 126 modules, more than carry a service — and they survive the split without fitting either half. A tool is not an artifact and not node state; it is a contract the mesh publishes on a module's behalf, and what consumes it is an AGENT rather than another module. That is a second audience the design has not described. Whether it is one relation with two audiences or two relations is cheap to decide now and expensive later. A migration belongs to the CONSUMER and runs on the PROVIDER. A game's migrations run against the database the store granted it: owned by the consumer, hosted inside something it does not control, ordered after the provisioning edge because there is nothing to migrate until the grant exists, and scoped to that grant. Ownership crosses the edge, which nothing in provides and requires expresses — and it gives a consumer's own install an internal order, provisioned then migrated then started, that depends on an edge rather than on its contents. And provisioning is EARLY, not late. The assumption worth killing is that it is something the control plane does for consumers once a mesh is running. The mesh's own registry database is provisioned before there is a mesh, and so is its virtual host on the broker: the store runs from the carried bundle, a database is created in it, the mesh's own schema is applied, and only then does a control plane exist. Steps two and three happen before there is a mesh to do them, so provisioning is part of the bootstrap and part of what the bundle has to express. Which strains ADR 0043. The host applies declared state ON THIS MACHINE, and a database inside a running store is not a file or a unit. At bootstrap it is at least local — the store is on the same machine. Afterwards a consumer on one node provisioned from a store on another is the ordinary case and reaching it is not the host's job. The same operation is local at bootstrap and remote later, which is either two mechanisms or one with a tier boundary crossing inside it. Currently the sharpest unresolved thing in the effort. |
||
|
|
a3c7e7e1f1 |
011: providing is a facet, and the assignment is a third thing
Any hosted service can be a factory — an identity provider grants clients, an analytics service grants a tracking identity, a mail server grants mailboxes, an application platform grants a project that is several of those at once. Providing is a FACET a module may have, not a kind of module it is, which is the same conclusion this effort reached about services and applications arriving from the other direction. So `provider` stops being a category too. Two relational stores from different vendors both grant "a database" and are the sharpest possible test of the substitutability rule. They fail it completely — different protocol, dialect, driver, client library compiled into the consumer — so `database` stays a tag, now with two real providers rather than a thought experiment. The assignment is a third entity, recorded because the operator tried the alternative: modules were once node-agnostic and it did not survive. Several of a provider's properties belong to neither end — where its state lives, how it is reached, tuning derived from the machine's hardware, which instance serves a given consumer. Not the catalogue, because they differ per node; not the node, because they are about this module. A design with only modules and nodes has nowhere to put them, which is what node-agnostic ran out of. The current system already stores environment values per module AND per node, arriving the same way. Which answers the question asked directly: two nodes both run a store, so which serves a consumer? Neither obvious answer. Not the consumer naming a node — that is placement in the consumer's manifest, a game edited because a database moved. Not the consumer not caring — for presence it genuinely does not, for instantiation it cares permanently. What the consumer knows is the SCOPE of its own need: one instance shared across every instance of itself, or one each. That decides, and needs no node named. Then the mesh binds, and the binding is recorded on the assignment and is sticky — a resolver that re-derives which store serves a consumer will one day derive a different answer and relocate a database. |
||
|
|
13c6068874 |
011: the provider shape generalises, and two things differ inside it
The broker has all nine properties the store has. So do the object store and the image registry. A substrate service is a SERVICE PLUS A FACTORY, there are four of them, and the pattern generalises past the substrate: anything granting something per consumer has this shape. Two differences matter more than the similarity. The broker cannot be managed over the broker. ADR 0001 makes it the channel every node takes work from and ADR 0039 makes it the security boundary, so the module providing it is also the way modules are managed — a declaration cannot be delivered to it over itself. Nothing else has that property; the store is consumed by the control plane but is not how the control plane REACHES anything. This is what the carried bundle exists for: the broker is raised from what the host carries because there is no other way to raise it. A constraint on one module, not a general rule, and a schema with no way to say so hides it. And two modules of identical shape want opposite instance counts. The broker is one per mesh by decision. The store cannot be, because a node that must keep working while disconnected cannot depend on a database elsewhere. Which settles what cases.md left open: how many instances is NOT derivable from what a module is. It is a per-module decision, it has to be declared, and nothing in provides, requires or excludes says it. Revocation differs in consequence too. Dropping a database leaves data until something removes it — a leak, recoverable. Dropping a virtual host loses whatever was undelivered — silent, and not. Same relation, different blast radius, which argues for the provider deciding what revocation means rather than the mesh applying one rule. File renamed: it was never really about postgres. |
||
|
|
f160b28a71 |
011: postgres worked through, and "one kind of edge" was wrong
The tidy version said a module provides names and requires names and that is the only edge. Working postgres through completely disproves it. A small game wanting to store data does not require postgres to EXIST. It requires postgres to MAKE IT A DATABASE and hand back credentials. Those are different relations in every way that matters: one creates something per consumer, carries a payload back, can be revoked, and leaves the provider holding state about who was granted what. The other creates nothing. So: two kinds of edge, one graph. Instantiation implies presence; presence does not imply instantiation. The current system already had exactly this split — `dependencies` for presence, `requires: provision:` for instantiation, with the resolver deriving one from the other. analysis.md called that derivation a convenience. It is not: it is the correct relationship between two genuinely different relations, and the design had collapsed them. Postgres also turns out to be nine things, not one. A container. Persistent state where moving nodes is a migration rather than a reschedule. Configuration partly derived from the machine's hardware. A tool surface. A provisioner. Its own bookkeeping about what it granted, which is not the data it stores. An exposure decision per node it runs on. Credentials it generates, which means a provisioning edge carries a secret. And health that is not "the container is up". Four questions the worked example makes concrete rather than abstract. WHICH postgres, when there are two — a consumer of `terminal` does not care and a consumer of a database cares permanently. How many instances a module should have, which cannot be a global rule because one-per-mesh is wrong for a store a disconnected node needs and one-per-node is wrong for the mesh's own registry. What happens to a grant when its consumer is removed, where dropping is data loss and keeping is a leak. And whether a declaration is composed PER NODE from what that node reported — because tuning follows hardware the control plane cannot know, and the alternative is the host deciding, which ADR 0037 forbids. |
||
|
|
c9c2dfe686 |
011: what a feature is, and what it splits into
The operator wants features gone, and 006 left it open. Measured, and the answer is that nothing replaces them because they were never one concept. A feature is a kind of content a module carries, detected from its directory: twenty-one of them, each with a handler owning six stages — build, publish, install, configure, start, verify. The structural finding: EVERY handler implements EVERY stage. `configs` writes files onto a node, has nothing to build, and has a build stage. `npm` publishes to a registry, has nothing to start, and has a start stage. One interface spans build-time and apply-time, so every kind of content must implement both halves and most do nothing in one — and a stage that does nothing looks exactly like a stage that failed to do anything. They split four ways, across three tiers. Artifacts built once per version and published, where no node is involved — delivery. Resources that are desired state on a machine, which is what ADR 0043 already describes and the host already does — tier 0. Actions run once against something that is not this machine, like a migration against a database on another node — delivery, and seeds go entirely. And checks: the prerequisites are REQUIREMENTS IN DISGUISE, a module saying what must be true before it can be installed, which is what an edge in the graph says; the verifiers are the read-back the host already performs. So `feature` is one word for four things spanning three tiers, which is why the pipeline is hard to reason about. One property must survive the split, and it is the thing the current design got right: content is DETECTED, relationships are DECLARED. A module that says it has migrations and has none is a fault nobody sees until it matters — but what it requires and provides is not visible in a directory and has to be said. |
||
|
|
ae099482a9 |
011: twenty cases, and two axes nothing covers
Before settling a schema, what a module can actually be. Twenty kinds of thing, with the hard ones at the end because they are the point. The ordinary nine are unsurprising: a supervised service, a system package with configuration, an application a person launches, a command-line tool, a library that never runs, a one-shot task, a scheduled one, an adapter, and a standalone application whose only difference is where its source lives. The eleven that break a naive schema are where the work is. Something that is a service AND an application — a git forge is consumed as a remote and operated through a web interface, and neither reading is wrong. Something that provides and consumes, because provider and consumer are ends of edges rather than kinds of module. Something the mesh installs that then becomes a node CAPABILITY, which means a node's provides-list is partly derived from what is installed on it and not only detected. Something that must be adopted rather than installed. Something that is a set rather than a thing. Something with exactly one instance for the whole mesh, where assigning it twice is not redundancy but two meshes. Something that is not software at all — a firewall policy, a DNS record, pure desired state, which fits the host's declaration model exactly and an installable package model not at all. An agent. The host itself, which is not a module and needs a schema that can say so. And the things the mesh depends on and does not control, which are why a node can be perfectly configured and still not work. Nine axes come out of it. Two are covered by nothing anyone has proposed: HOW MANY INSTANCES a thing may have, and WHETHER TWO CAN COEXIST — `excludes` covers part of the second and nothing covers the first. And one question the cases sharpen: is "runs" a property or a kind? The axes say property — one schema with a field saying how it runs, `never` included. The alternative is several kinds of module with different schemas, which is the taxonomy this effort already rejected once for services and applications. |
||
|
|
f55ecc1a47 |
011: an abstract name needs providers that are actually substitutable
Two corrections from the operator, and the first improves the design rather than narrowing it. `database` is not an edge. The test it fails, and the test the proposal was missing: can a consumer be switched from one provider to another WITHOUT CHANGING? A module speaking Postgres does not speak MongoDB or SQL Server — different wire protocol, dialect, driver — so a consumer declaring `requires: database` and handed any of them breaks. The name promises what no provider can deliver, and the resolver would report a requirement satisfied that is not. `terminal` passes: anything that runs a command in a terminal works and the consumer never learns which it got. So the ADAPTER is what creates an interface. `ai-assistant` is legitimate exactly because adapters normalise what is behind it. Without one there is no interface, there is a category — and a category is a TAG. Tags describe, edges bind, and keeping them apart is what stops the catalogue acquiring a second kind of relationship that looks like a dependency and is not, which is what a folder named after a domain already was. And the domain module goes. A `networking` module gathering a firewall, a resolver and a proxy under one name came from an older shape and does not fit — there is no such thing to install. There is core infrastructure: concrete modules named individually, not flavourable, with no grouping module standing in front of them. Fixed three places where the revision left the old rule standing, including an example manifest still requiring `database` — the kind of contradiction that would have been read as the design rather than as a leftover. |
||
|
|
e20a09ae80 |
011: one kind of edge
The design, rather than an account of what exists. A module provides names and requires names, and that single relation absorbs three things this effort had listed separately: requiring another module is requiring a concrete name, requiring a resource is requiring an abstract one, and an interface is simply a name with more than one provider. Nothing has to declare that it is an interface — it either has one provider or several. The move that does the most work: a NODE provides names too. Its profile is a set of them — display-server, container-runtime, an architecture — so a module requiring a display server is satisfied by the node exactly as one requiring a database is satisfied by another module. One resolution instead of two, and a graphical application cannot land on a node without a display server for the same reason, through the same code, that it cannot land without its libraries. Which makes the host's capability detection an input to resolution rather than something a person reads. It was built to be read; it turns out to be a provides-list. `excludes` is the one genuinely new relation, because it is not derivable: two modules that both provide message-bus look interchangeable when installing both would break the machine. Constraints are not placement. They say what must be true of a node, never which node — which is the mistake the measurement found in the current catalogue, where a module pins its database to a named node so a second node cannot provide it without editing the consumer. What it deletes, for the design: the module/resource distinction, the interface as a kind of thing, capability checking as a separate mechanism, domain grouping — folders assert relationships where edges record them, so a domain becomes a query over the graph rather than a directory somebody keeps true — and possibly tiers, if a tier is just a computed level. What it does not delete, stated so it is not discovered later: a resolver still has to exist, with version constraints and conflicts, and the design owes an answer on what it delegates rather than reimplements. |
||
|
|
c3a2984b3e |
011: measured, and the premise was wrong — the graph is not missing
The effort was opened to ask whether the catalogue's missing structure is a graph. It is not missing. 126 manifests, 103 edges, no cycles, nothing dangling, deepest chain of five — and a resolver in the SDK that topologically sorts them, already called by the tool loader at startup, the installer when syncing modules onto a node, and the delivery coordinator when expanding what a change affects. It already does something this effort assumed would need designing: a requirement on another module's provision is treated as an implicit edge to the module that provides it. So "ordering by the graph", which ADR 0043 makes the control plane's job, is a thing to call rather than a thing to build. The one place the graph is wrong, it is wrong about the substrate. A module needing a database declares `provider: postgres` inside `provisions:` — which is what a module OFFERS — so the resolver, which reads `dependencies:` and `requires:`, never sees it. Three edges are invisible this way, and they are the mesh's own database, the mesh's own broker, and the work engine's database. The consequence is measurable: computing what a working mesh needs from the declared graph gives registry -> sdk -> mesh -> meshware. Four modules, four levels, no database. Arithmetically correct and obviously wrong, for exactly one reason — a field that means "depends on" is not read as one. That is 04-ISSUES/003 in a new form: not a key nothing reads, but a key read as something other than what it means. Two latent defects, both contrary to ADR 0008 and both in the component ADR 0043 makes responsible for ordering a host will apply without question: a cycle warns and falls back to input order, and a dependency that does not exist warns and continues. Neither has fired, because the catalogue currently has no cycles and nothing dangling, which is why nobody has noticed. And placement is decided in the catalogue: a provision pins itself to a named node in the manifest. Which node runs what is an inventory decision — tier 2 by the skeleton's own test — so a second node cannot provide the mesh's database without editing the module that consumes it. What the graph would DELETE is currently nothing. What it would add is three declarations that no manifest uses today: excludes, a required node capability, and an interface with adapters. Whether they would be used is not measured, and zero usage is equally consistent with nobody needing them and nobody being able to express them. |
||
|
|
278f7427ed |
012: the briefing carries an outcome, derived from its lines
Proposed by the operator: state plainly whether adoption succeeded, partly succeeded or failed, with a severity per line. Taken with one change — the overall is DERIVED as the worst mark present, never written alongside. Two fields maintained independently drift, and a briefing reading "full success" while carrying a failed line is exactly the fault this record keeps cataloguing. An outcome computed from its lines cannot disagree with them. Four marks: ok, kept, unknown, failed. "unknown" is not a shade of success — adoption will meet configuration it cannot parse and state it cannot read, and folding those into "fine" is the same move as reporting an installed package as a capability. And adding severity reopens something the earlier rule did not cover. "Flags inform, they do not block" was decided about CONFLICTS, where the mesh chose deliberately and the machine still works. A failure is not "we chose" but "we could not". Treating both the same makes a node where something the mesh needed never happened indistinguishable from one where a log level differed. |
||
|
|
0106318bcb |
012: on conflict, keep the machine's configuration
Reversed by the operator, and both directions are recorded because the reasoning for each is the useful part. What is already on the machine stays, the conflict is flagged, adoption completes. This buys non-destructiveness by construction: the class that made the opposite rule dangerous — a storage driver against the filesystem it is actually on, a data directory pointing at a mount that exists — cannot arise, because nothing tied to the machine's physical reality is overwritten. It exposes the mirror. The mesh's configuration is not only preference; some of it is what a module needs to function. Keeping the machine's version there produces a module that is installed and does not work, which is 04-ISSUES/007 arriving from a direction that issue did not anticipate. And a fleet where every node kept its own settings is one where a module works on one node and fails on another with nothing able to say why. So neither direction is right as a blanket, and the question is not whose configuration wins. It is whether the module REQUIRES the setting or merely PREFERS it — required contradictions cannot be kept without breaking the module, preferences should always yield to what is there. That is a property of the module's declaration rather than of the adoption algorithm, which makes it one more thing the graph would carry. Until modules can say which of their settings are load-bearing, adoption is defaulting in the dark, and the default chosen is the one that does not break the machine it is adopting. |
||
|
|
bcb18c7329 |
012: on conflict, install the mesh's version
Decided by the operator. Where the existing configuration and the mesh's disagree, the mesh's version is installed, the conflict is flagged, and it is reconciled afterwards — the mesh's configuration is known to work, the machine's is not, and a half-adopted machine is a state nobody understands. So adoption always completes and flags inform rather than block, which also settles what 'adopted with open questions' prevents: nothing. The node is a node. The original is kept, so nothing is unrecoverable. One class left open rather than folded in, because it is the one place the oldest rule in this record argues the other way. 'Known to work' is true of the mesh's configuration in isolation, not on this machine. Most disagreements are preference and overwriting them is right. A few are tied to what is physically present — a storage driver against the filesystem it is actually on, a data directory pointing at a mount that exists — and installing ours there does not discard a preference, it can make existing data unreadable. Restoring the configuration file afterwards does not undo that. The default is settled. The exception is not 'there is a conflict' but 'applying ours would destroy something a configuration backup cannot restore', and identifying that class is open. |
||
|
|
60736199a4 |
012: keep the original, and flag what cannot be decided
Two additions from the operator, and the second answers a question this effort had open with two bad answers. Nothing is taken over without keeping what was there. Adoption happens on machines somebody is already using, and the configuration being taken over is configuration somebody chose. This is a never rule rather than a courtesy, and it earns that by the same incident the mesh's strongest rule carries: the worst loss in this record came from a tool acting on a path it did not own. Adoption is that act made deliberate, which makes the safeguard obligatory. And adoption produces a briefing, not just a result. It meets things a script cannot decide — a runtime configured one way against a mesh wanting another, a package pinned for a reason, local settings the mesh has no opinion about. Silently winning is wrong in both directions and refusing outright makes a machine in use unadoptable. So conflicts are FLAGGED: what it found, what it took over, what it could not resolve, written to be read by a person or an agent as the first thing a session on that node has to work with. That is the declaration parser's principle at a larger scale — name every problem at once, to somebody who can act on it. The question it turns on is recorded rather than assumed away: are flags advisory or blocking? A briefing nobody opens is worse than a failure, because the machine is in service and the record says it went well — 04-ISSUES/003 again. Working position: the node is usable and the mesh KNOWS it has unresolved adoption questions, as a state something can ask about rather than a document in a log directory. What that state prevents is undecided. |
||
|
|
ddb8091f68 |
Research 012 — the minimum viable node, and adopting what is already there
Building tier 0 reached a wall that looked like a packaging problem and is not. The host can be told to run a container or install a package; both need a file, and asking where the host gets it produced a bad trilemma — carry everything, download at apply time, or push the files in first. Downloading fails on the first node, which cannot fetch the image registry from the image registry it is trying to start. The reframing came from the operator: the machine is not offline, and what matters is WHEN the fetching happens. Move it from apply time to build time — build the installer on a machine with a network, tailored to the target, apply it on a target that then needs nothing. The same move the lab already made for its router image. Which makes the question not where artifacts come from but what is missing from THIS machine, and that needs two things answered: the closure for a one-node mesh, and how a machine already in use becomes one. Adoption is the second half, and it is sharper than it sounds. Having a package installed is not owning it: a container runtime found already present carries settings somebody chose, and noticing the binary exists discovers none of them. It was also the original path — 00-as-is/05 records adoption of a pre-existing machine's configuration as the original mechanism, since made legacy and explicitly out of scope for the lab. It returns for a different reason than it was dropped for. Two collisions recorded rather than discovered later. ADR 0004 has managed files generated and never edited, and adoption needs a one-time import before that rule starts applying — three states, and the middle one is new. And ADR 0043 says the host never touches what it did not create, which is exactly what adoption does; that rule needs a companion rather than an exception. Eight open questions, including whether 'tier' is just a coarse view of a graph level, whether owning a package means owning its version, and what cannot be precomputed at all — because tailoring moves the cost of building from source rather than removing it. |
||
|
|
64913ed0d3 | Merge pull request 'The approval is the checkpoint, and what a declaration is' (#10) from design/approval-is-the-checkpoint into main | ||
|
|
00d5ba8376 |
0042 and 0043, as approved
0042 becomes the operator's own rule and stops there: every merge into the main branch is notified and approved. Notified means proposed and said out loud, not performed and mentioned; approved means a person says yes to THAT merge. Who performs it is not the thing worth constraining, which is what makes an agent merging its own work unremarkable — the checkpoint already happened. 0043 answers the question it did not: where the ordered list comes from. By hand today, in substrate.lock, because the first node has no control plane to derive anything from. Afterwards the control plane derives it from module assignments, resolved configuration, and what each module declares it needs — ordered by the dependency graph, which is research 011. So the record is complete on the consumer side and deliberately silent on the producer side, and that is a legitimate order to settle them in: the host must refuse what it does not understand whoever wrote it. One consequence that only appeared when the question was asked: if the graph turns out not to determine a total order, that is 011's problem and not the host's. The host is still handed a list and still applies it as given. Recorded because it is the seam where a future difficulty would otherwise try to migrate into tier 0. Both were marked accepted before they had been read. Approved now, so the field is true — which it was not when it was written. |
||
|
|
bea052753e |
ADR 0043 — what a declaration is
Stage 2 could not start without it. Three constraints already bound the shape and between them they decide most of it. JSON, because the standard library carries it and carries no YAML, and a YAML declaration would put a third-party parser inside the one binary whose whole argument is that it needs nothing — to gain authoring comfort in a document generated by a machine and read by a machine. An ordered list, because ordering is a DECISION. A host deriving order from declared dependencies would be deciding the thing most likely to differ between what the control plane intended and what the machine does. The control plane knows what depends on what; it says so by saying when. Unknown is refused, never skipped — an unknown version, type or field refuses the whole declaration. A host that skipped what it did not understand would apply most of a declaration and report success, which is 04-ISSUES/003 with the declaration on the other side of the wire. Complete for what the host OWNS, and only that. It removes what it previously applied and is no longer declared, which it knows from the store rather than by inference, and never removes what it did not create — a converger that treats 'not declared' as 'must not exist' deletes what the mesh never put there. Two consequences arriving earlier than the build order suggested: the store is load-bearing at stage 2, because nothing can be removed without knowing what was applied. And a closed address space bounds the first vocabulary to what needs no network, because a scenario has no route to a package repository. |
||
|
|
23232e019a |
ADR 0042 — the approval is the checkpoint, not the second pair of hands
§2 said "never merge your own", written for people. Applied to an agent it produced a contradiction that surfaced immediately: an agent asked to merge cannot merge, because it authored what it is being asked to merge. So every merge here was either performed by the thing that wrote it, or not performed. Rejected the literal reading, because an operator clicking merge dozens of times without reading is not a checkpoint — it is the SHAPE of one, which is worse, since the record then claims a review that did not happen. Rejected dropping the rule, because the failure it prevents is not one an agent is less prone to. So the rule names what the checkpoint actually is: a person deciding, not a person clicking. Work may be merged by whoever wrote it once a human has explicitly approved that merge. What "explicit" excludes is the half that can rot, so it is enumerated: a standing permission cited forever, an instruction to do the work read as approval to merge it, silence, and the author's own judgement that it is ready. This narrows rather than relaxes. The obligation moves from who performs the merge to whether a person decided — a higher bar in the case the old wording permits, where a reviewer merges someone else's work without reading it. Synced, and verified by reading the rule back out of the live page rather than by trusting the publish. |
||
|
|
7ef1c561e0 | Merge pull request 'Tier 0: the questions answered, the decisions taken, and the design' (#9) from design/close-the-record into main | ||
|
|
92e8c74ce4 |
ADR 0041 and the build handoff for the node host
Building tier 0 forced the question "the one binary installed by hand" had been carrying unexamined. A TypeScript host needs a runtime present before it runs, so the thing installed by hand becomes two — and the second must be installed by the means the host exists to replace. So the host is a statically linked binary that requires nothing present, written in Go. Rejected: a runtime installed first, which breaks the property the tier rests on; and bundling the runtime into the executable, which carries ninety megabytes to preserve a language choice and puts a young feature at the bottom of the stack. The argument that decided it is architectural rather than about taste. 0037 means the host never queries the mesh database and 0039 means it only receives declarations, so the host shares NO code with any other tier — not a client, not a schema, not the SDK. The language boundary falls exactly on a boundary that already exists, and a second language usually costs duplicated logic where here there is none to duplicate. §8 gains a scope: it said "TypeScript throughout" when everything was a service or a surface, and is now scoped to those with tier 0 named. Another sync owed. Playbook 04 steps 2 and 4: repos.md records mesh-host as existing, the design takes code: [mesh-host] and status: in-progress. |
||
|
|
b9facf9375 |
Design the node host
Playbook 02 step 3, on four recorded decisions. Tier 0 has one job — apply declared state on this machine — and the six absorbed concerns are instances of it, not additions to it. Specifies the six parts and what each owns, and the two properties that make apply trustworthy rather than merely present: every applier reads back, because setting a value is not evidence the value took; and what was applied is recorded after it works, never before, because a failed apply leaves the machine wherever it reached and nothing must claim otherwise. Build order is staged so each stage is verifiable in the lab before the next exists. Stage 1 is profile and inventory — no control plane, no declarations, no network — and it is deliberately the smallest useful thing, because `place:` has nothing to place and the lab therefore raises empty machines. Stage 1 ends that, and every later stage is tested by a lab that already works. Stage 2 is the one that could invalidate the tier boundary: whether one host can raise the substrate alone is Move 1's assumption and has never been proved. Every decision the design rests on is given the test that asserts it, per 0034 — including the dependency-direction lint, which is what makes "the host never queries the mesh database" a rule rather than an intention. Six things left open and named, including the one that host-size.md could not measure: zero dependencies, but still six vocabularies. |
||
|
|
6e4fc5d69b |
ADR 0040 and the constitution sync — absorb, then publish
The sync came due for three accepted rules. Reading the target before overwriting it found the source and the enforced copy had diverged unrecorded, and that a literal republish would have DELETED rules the mesh enforces: the live page's §4 carried SOLID, layering, TypeScript strictness and DRY/YAGNI, which appear in no decision record anywhere and have been checked against for six weeks. Removal was not the safe alternative either. The orchestrator reads "when absent, no constitution is injected (backward-compatible)" — so deleting the page would not fail, it would silently inject nothing, and every design meeting would run unchecked. Three unenforced rules would have become all of them. So: absorb first. The code-quality rules land as §8 rather than §4, because appending renumbers nothing and every existing citation stays valid. They are marked as inherited — every other rule states the incident behind it, these state nothing because nothing was written down, and importing them silently would have claimed a provenance the document does not have. The review bar is resolved to a person who is not the proposer. The live page required two node operators; there is one, so the rule was never met and could not be — a rule that cannot be satisfied is not a high standard, it is one everything silently violates. Then published, and the read-back earned its place in the playbook: the FIRST publish reported success and changed nothing. New revision, new title, body unapplied — a malformed argument dropped silently. §5 demonstrating itself during its own publication. |
||
|
|
34e7a4780a |
Accept 0035 — a report is read from the system
The principle is obvious; the obligations it imposes are not, and those are what was actually being decided. The applier records what it did, including facts it never reads itself, purely so something else can read them back. It records them AFTER the thing works, so a half-finished run ends up with less metadata rather than optimistic metadata — rejecting the simpler alternative of tagging at creation with a status field, because a failed raise deliberately leaves wreckage standing and wreckage tagged at creation claims things that never happened. And it binds anything that reports on the mesh, not only the drawing. §5 carries the rule unqualified now. The constitution sync is owed for this and for 0034 and 0018, and has not been run. |
||
|
|
66a89cefbe |
Accept 0039 — a node owns no password, only an identity
Settled in the operator's own words: nodes should not own passwords, only an identity when communicating to the broker. Recorded that way at the top of the record, because it is the whole decision in one line and the rest is why. What this obliges, in order of newness: enrolment is the one mechanism that does not exist. Per-node broker users, virtual hosts and per-queue permissions are broker configuration. Mutual authority is certificates on a connection already open. And the boundary must fail legibly, which is the requirement the debugging objection earned. |
||
|
|
9ee0c63c7a |
0039: the link already exists, and most of the cost is already paid
Written first as though the link were a thing to build. It is not. ADR 0001 already has it — every node connects outbound to a single broker, nothing ever connects to a node, each node declaring an exchange named for itself and consuming from its own queue. Already outbound-only, already per-node addressed, already the one channel everything arrives through. So this record is not proposing a channel. It proposes that the channel carry per-node identity instead of one shared credential. The same as-is records the fault it fixes, for the broker rather than the database: the broker's credential is mesh-wide, rotating it is a mesh-wide operation, and doing it wrong has taken the broker down. That reframes the overhead objection, which was fair against what the record said and not against what it means. Of six properties, four are already true, one is broker configuration — users, vhosts and per-queue permissions the broker already implements — and exactly one is new machinery: enrolment. Meanwhile 0037 subtracts, since a node under it holds no database credential at all. Three hand-carried shared secrets become one identity that grants only identity. Adds the option that was actually being weighed and was missing: accept the exposure as the cost of simplicity. Rejected because the simplicity IS the unfixability — the credential cannot be rotated precisely because everything holds the same one. And adds a requirement from the debugging objection, which was the strongest part of it: it must fail legibly. A boundary that refuses a node without saying why is worse than the password it replaced, because a wrong password at least announces itself. That is §5 applied to a security mechanism. |
||
|
|
902739acb6 |
Research 011 — the module graph
The proposal to split modules into provisioning services and applications was worked through and abandoned, for a reason worth keeping: it cannot be filed consistently. A git forge is consumed as a service and operated through a web interface; an analytics service grants tracking identity and is a dashboard. The operator's correction is the sharper form — what runs on the machine is a supervised container, not something a user started. That is a fact about HOW a thing runs, not about what kind of thing it is. So it is a facet, and 0002 survives: everything is a module. What the catalogue is missing is not a taxonomy but a graph. Grouping asserts relationships; a graph records them. Five declarations, of which two exist: requires/provides a resource (yes), requires/excludes another module (no), requires a node capability (no). Plus interface modules that carry no implementation, with adapters providing them. Recorded because it matters: this is a package manager's model, and pacman already has all of it — depends, conflicts, and provides as virtual packages, which is exactly the interface/adapter idea. Arriving there independently is evidence for the shape. It is also a warning about what not to reimplement. Working position on capabilities, to be tested: intrinsic ones (hardware, architecture, network position) are detected and never installed, and a module requiring one it lacks is impossible rather than unresolved. Provided ones (a display server, a container runtime) are not a separate kind of thing — they are modules that provide a capability, so "may the mesh install a capability" is not policy, it is dependency resolution. Issue 007 then bears directly: an installed package is not a capability. Also captured: the operator's assessment that the machinery around a module — scheduled tasks, hooks, migrations, config and env — is worth keeping, seeds are not, and the integration is wrong enough to need a major refactor. Research 005 found supporting evidence from another direction, that the densest apparent coupling in the catalogue is manifest boilerplate churn. The first open question is the one that decides whether this is progress: what does the graph DELETE? If modules gain declarations and lose nothing, it is motion. |
||
|
|
f5073b9b00 |
Accept 0034 and 0018; 0011 is superseded
0034: a test defends a decision. §5 carried it marked "proposed, pending review"; the marker is removed and the rule now stands unqualified. The lab was already built to it, which is the inversion §5 exists to catch — closed now rather than left standing. 0018: the mesh creates no symlinks. §3 said the installer owns the links today and the INTENT was that the mesh creates none. It is no longer an intent, so the wording says so, and ADR 0011 becomes superseded rather than edited — its reasoning is why the rule exists at all, and the incident behind it is the reason anyone believes either record. The links the installer still reconciles are a migration, not a permission. The constitution sync (§6 step 4, playbook 05) is NOT done. An unsynced rule is a rule the mesh does not enforce, whatever this document says — and publishing it changes what every design meeting is checked against, so it wants saying out loud rather than doing quietly. |