2baf22ac43047916f3d36ee0d116b980595c9560
87
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a1a10e9ed2 |
Bootstrap ends at a usable mesh, and the first credential comes from a person
Bootstrap stopped when the control plane started — a mesh that runs and cannot be used by anybody not standing at the machine, since the networked surfaces need an identity provider and no module has been assigned yet. It now runs through the provider and the first login. The obstacle was not incidental. The mesh has never held a readable secret: Make generates and seals, keeping no readable copy. An initial administrator's credential is the first value a person must read. Generating it and printing it once was the convenient option and is refused. It would give the control plane a plaintext secret for the first time — briefly, and to one terminal, but the capability would then exist, and an exception made for one case does not stay one. The next awkward credential gets printed too, and "a copy of the database is a copy of nothing" stops being checkable by reading the code. So the operator supplies it, on standard input, not echoed — the path that already exists for a model-access key. What is created is an account in the identity provider, not a user of the mesh; there is still no user model. Unattended bootstrap remains possible and the value still comes from outside: automation supplying it is the operator supplying it. What is refused is the mesh inventing one, so an unattended bootstrap with nothing provided yields a mesh with no administrator — correct rather than broken. |
||
|
|
0ef6d3f574 |
One implementation, several surfaces, and what that costs
The mesh is operated from a command line and must be operable from a browser and from a model's tools, without becoming three systems. The pattern is already in the code and was unnamed: `board` serves HTTP by calling the same functions the CLI calls, holding nothing. Takes the decision 0034 said had to be taken deliberately rather than arrive with a feature: the HTTP surface is not read-only, so a browser login now carries authority over the mesh. Names the dependency by protocol — an OAuth2 identity provider — as the mesh does for AMQP, S3 and OCI. Keycloak is what fills the role; what the control plane knows is that it validates a token, and replacing the provider is a migration rather than a redesign. Says what this must not become, because it is the failure the project was started over: a kernel every module imports, 155 files of code from every context. Shared surfaces are not a shared library. Three adapters calling the same functions is not the same as logic leaving the context that owns it. And records the loop it creates. The networked surfaces depend on a module the control plane assigns, so when identity is down nobody can authenticate — including whoever is trying to fix it. The way out is the command line, which authenticates through nothing and is available to the account that owns the machine. Hence the rule: no capability exists only behind an authenticated surface, because that is a capability which disappears exactly when identity does. |
||
|
|
fccac61e58 |
The board is a web application, not a category
Supersedes 0032, which decided the right thing and described it wrongly. The decision is unchanged: the account that installed the host owns the mesh, and there is no user model. What was wrong was inventing "a surface that delegates authentication" for the board. It is a web application with a login, in the way every web application has a login. That is a fact about an application, not a property of the mesh. The cost was not cosmetic. It made the identity module look like part of the mesh's authority — something the mesh depends on to know who anybody is — when the mesh knows nothing about people at all and one of the applications running on it happens to have a login. Keeps the line that is worth writing down, and states it more plainly: signing in to an application must not become authority over the mesh. Today it cannot, because the board reads and does not act. The moment it can assign a module, whoever it lets in has mesh authority — and it would arrive as a feature rather than as a decision. So a surface that can change the mesh is a change to who owns the mesh, and is taken as one. Not forbidden; just not something that turns up in a pull request titled "add assign button". |
||
|
|
10fa7c76d7 |
The substrate is a store and a broker
Third correction to one table today, found the same way as the other two: by asking whether both halves of the test were answered, or only the easy one. 0006 admits the registry because "it cannot grant itself a repository" — true, and the second half. Nothing established that the control plane needs one in order to run. Counted rather than argued: the bundle raises twelve resources and no registry is among them. The registry arrives afterwards as an ordinary module, which is exactly what the lab asserts. 0006 half-said this already, calling it "substrate by role and ordinary by delivery, provisioned once there is a control plane to do it". A member provisioned by the thing it supposedly precedes is not a member; that phrase was carrying a contradiction rather than resolving one. The registry is a closer call than the object store and the difference is worth keeping: the control plane never touches an object store at all, but it genuinely uses the registry. So the registry is a real dependency of the mesh operating and not of the control plane starting — and it is the second that the word means. The substrate is now exactly what the bundle raises, which is the strongest form the list can take: checkable by counting rather than by reading an argument, and the two cannot drift. The finding is not about substrates. A test with two conditions is a test only when both are asked. |
||
|
|
a028337490 |
The local account owns the mesh; a surface delegates to a module
Answers what 0031 left open, and a question it did not ask — who owns the mesh at all. There was no answer, and the absence was invisible because every operation so far has been run by the person sitting at the machine, so nothing had to say whether that was the design or the circumstance. The account that installed the host owns the mesh on that node. No user model, no roles, nothing to administer. It follows from 0004 rather than adding to it: there is no authorisation between nodes because every node is the operator's own, so a user model inside that boundary would guard nothing — anyone it could stop could read the node's key off the disk. The board is different, and the difference is the network. A surface reachable by a browser has to know who is asking, because those people are not by construction people with a shell on the machine. So it delegates to an OAuth provider, which is a module. That does not make identity substrate. A surface delegating authentication is not the control plane delegating it: the control plane runs, applies declarations and reaches nodes with no identity provider in existence. Only the board needs one. Records the cost plainly: anybody with a shell on a node has full authority there, and there is no way to give somebody authority over one node without giving them a login on it. |
||
|
|
e3934e4449 |
The control plane authenticates nobody, so identity is a module
Closes the last open question about what the substrate contains. 0006 left an identity provider conditional — substrate only if the control plane delegated authentication — and said the decision had not been taken. It is now: it delegates to nothing. The conditional was never about machines. A node proves itself with a keypair it generated over a broker account issued at enrolment, and declarations are verified by signature; none of that involves an identity provider. It was only ever about whether a person signing in to a mesh surface would be authenticated by something else. So the substrate is three — a relational store, a message bus, an image registry — and with 0028 having removed the object store, no member is conditional and every one is there for the same reason. It does not settle how a person signs in to a surface, deliberately. What is settled is that whatever answers that is not something which must exist before the mesh does, so it can be decided late or replaced — which being substrate would have prevented. |
||
|
|
d1ab2dc0b4 |
Data outlives the mesh that declared it, and the conversion starts where it lives
0030, found by asking what the conversion actually needs rather than by reviewing anything. The host deleted a directory and everything under it when it stopped being declared — which happens when a module is unassigned, or when a manifest is edited to move a data folder, which is the exact operation this plan needs. A database's files, a mail spool. The report said "removed". A directory still holding something is now kept and said so. No flag and nothing to remember: emptiness is the test, and it works because the removal order was already right — the mesh's own contents are gone by the time the directory is reached, so what remains is by definition something nobody declared. The plan now says data outranks its own ordering: copy, read back through the service that owns it, and only then point anything at the new location. Never move and then check. And it records where this starts — the node holding all the production data — with what that costs stated rather than argued with. Everything proven so far was proven on machines that could be destroyed and raised again. A scenario proves the mechanism, not the state on that machine. |
||
|
|
e4327a3a5e |
Phase 1.3 done: ordering was already there, the network was not
Ordering needed no change for the third time running — resources apply in the order declared and nothing sorts them — and is now asserted, because sorting them for any sensible reason would have passed every other test. Separates ordering from readiness, which the task had run together: a container started is not a container ready. Nothing waits, and what needs something usable retries. That is deliberate and more robust than start ordering, since a dependency can restart long after apply. The network was the first thing in Phase 1 that genuinely needed building, and the first that needed a decision: 0029 records why a shape rather than an action, and the vocabulary is nine. |
||
|
|
cb1954e7b5 |
Accept 0024, and rewrite the work breakdown around what is actually being done
**0024 accepted.** Model access was decided, built, and proven in the lab, and two design documents rest on it; only the status had never moved. The gate is green again. **The work breakdown rewritten.** It planned a decomposition of the existing system in place — extract contexts, declared features, shrink the shared library. That is not the work. A replacement is being built beside it, and only the old Phase 0 survived contact with reality, so the one document meant to say what happens next was describing a system being retired. Now ordered by what "modules move across one at a time until the old registry is off" actually requires: - Phase 0 is marked done against the twenty-two lab assertions, **and carries its own limitation**: every module exercised was written to test the mechanism, so the vocabulary was shaped by its own fixtures. - Phase 1 is the vocabulary gaps found by asking what real modules need — an object-store provision, a session as a licence consumer, a network shape with ordering, public certificate issuance. - Phase 2 is one module, then a week of running it, because the point of going first is to find what Phase 1 missed. - Phase 3 picks modules that each prove something the first did not; the mail system is last because it is the one that may send work back into the declaration language. - Phase 4 is switching the registry off, named as a phase so it is not mistaken for the goal. Keeps the rules of engagement unchanged — they were about how work is done, not what it is — with one addition: stop and ask before anything that touches a machine outside the lab. Adds a section on keeping the list true, since the document it replaces was wrong for weeks and nothing said so. A claim here is counted, not reasoned, and a phase is done when the lab says so. |
||
|
|
cbcbba8099 |
A provision names its engine; the substrate supplies only the control plane
**0027 — provisions.** A module written against PostgreSQL could be matched to a provider of SQL Server, resolve as satisfied, and fail on its first query. The name said the role, so nothing distinguished engines. Refusing on ambiguity could not help: with one provider of each name nothing is ambiguous. Enforced at parse rather than documented, because the old naming was the documentation. **0028 — the substrate.** 0006 admits an object store on the grounds that it cannot grant itself a bucket. That answers the second half of the test and assumes the first: the control plane does not need one. Verified — no S3 client in mesh-control, and internal/builder/registry.go records the deliberate choice to put artifacts in the OCI registry as content-addressed blobs. The row was inherited from the system being replaced, where an object store distributed module tarballs, and was never re-tested against the definition above it. So an object store is an ordinary module, and a mesh with nothing needing one runs none. Migrating it is module work, not substrate work. 0028 also states what 0006 left unsaid: a substrate service and a module of the same product are different instances. The substrate is raised from the bundle before any mesh exists, so it is not in the module graph — a workload depending on it would depend on something the graph cannot see, cannot rotate a credential for, and cannot move, and would put workload data in the store the control plane keeps its own state in. Both records were found by reading code against design rather than design against itself, which is the review that should have happened sooner. |
||
|
|
3c6c16abdf |
The mesh has a session of its own, and it is the node session's mechanism
A session for the mesh itself, addressed as the mesh, differing from a node's in exactly three things: the context it starts in, its engram, and its licence binding. Not a new kind of agent — the same mechanism pointed at a different root. Two implementations of one mechanism drift, and the vocabulary collision 0001 exists to undo began exactly that way. It runs on the control-plane node, and the reasoning is easy to get backwards: not "the important agent on the important machine", but that this node is already the one place excepted from "compromise of a node is compromise of that node". Placed anywhere else it would create a second such place. It is an addition to per-node messaging and never a replacement. 0001 holds that losing the control plane costs change, not operation — and a mesh whose only conversational surface lived there would lose the ability to ask anything while every machine kept running perfectly. Writing it up exposed that the node session's setup was never designed at all. 0004 gives behaviour and stops: nothing said how a session starts, where its context lives, or how a broker message becomes a prompt. That gap was invisible until something had to be built *like* a node session. 15-the-agent-session.md covers both as one mechanism. It also makes "a consumer that is not a machine" undeferrable. The control-plane node now hosts two sessions that must hold different licences, and a per-machine binding cannot express that at all. Noted in 14-model-access.md against the gap it was already recorded as. Also completes the to-be index, which stopped at 10 and omitted four documents. Pre-existing broken ADR references in the older rows are left alone rather than guessed at. |
||
|
|
e823cc1cc5 |
The design record is read where it is written, never copied to be found
Decides the question 006 narrowed to. An agent reads this repository directly and the search consults it, so these documents surface beside ordinary results instead of only when somebody already suspects they exist. A scheduled sync into the mesh's memory was the option that works with what exists today, and lost on the ground this repository can least afford: it makes a second copy, and the copy that is searched quietly stops matching the copy that is edited. A design record that has silently diverged from the reasoning it claims to carry is worse than one that cannot be found — the first misleads, the second merely fails. Amends what 0019 promised rather than satisfying it: these documents will not be indexed, they will be read. The commitment that survives is the one that mattered — that a searcher finds them without already suspecting they exist. Gated on an agent that does not exist yet, so 006 stays open on the build with a decided shape. What closes it is a check that fails today by design: search the mesh's memory for a phrase that appears only in a design document here, and require it back. |
||
|
|
5253742773 |
ADR 0024 — model access is a provision, and a licence has a name
A new requirement, and it is mostly a shape the mesh already has. A module that needs to think requires model-access; several vendors and a locally-run model are several modules providing it; choosing is assigning the one you want. A model the mesh runs itself needs nothing new at all — it is a mesh-scoped provision on the node with the hardware, credential included. A licence is a named thing because the whole point is saying which one a given consumer uses, and the names are the operator's. Many to many, so not a claim: two machines sharing an account is ordinary, not a collision. Four gaps, written as gaps rather than design: - a provider that is on no node, reached over the public internet, which the reachability rule must not refuse - a secret the mesh is GIVEN rather than mints. Every credential it handles today it generated and discarded; an API key arrives from a person, and accepting one must still discard the plaintext - a consumer that is not a machine. Which licence a worker uses is a binding to an agent, and the provisions model has no consumer identity other than a node - switching on exhaustion is a reaction to something observed, not a declaration. It belongs with observability, changing a binding — saying so is what stops the declaration language growing a conditional The existing auto-refresh and switching is not being replaced because it was wrong. It is being rebuilt because it lives somewhere that cannot express the rest. |
||
|
|
a73014dcd5 |
A bare machine became a mesh, and something joined it
First end-to-end raise. A machine with a container runtime applied the bundle its host carries and ended with a store, databases, schemas, a broker holding a certificate it generated itself, and the control plane serving. Then it took a token, checked the broker against the pinned fingerprint, generated three keypairs and enrolled — the first node being a node whose mesh is not up yet, observed rather than argued. And a credential crossed. Declared the provider of a database for a second node and pushed to over the broker, the machine ended with the password in one file at mode 0600, and that password appears nowhere in the declaration that crossed the broker, nowhere in the control plane's database, and nowhere in what the node reported back. That is the whole secrets argument, measured. One fault, in the joining: the token did not say what the mesh calls the machine, so enrolment needed a flag its own help said it did not, and failed at the broker with an empty username. It is the fifth thing a token carries now — the node cannot work its own name out, because the broker account it authenticates as is named after it and exists before the mesh has told it anything. |
||
|
|
0f7e4ab597 |
The provisioner, which is where the mesh stops
A password nothing was told to create authenticates nowhere. The mesh generates one, seals it to both ends and cannot read it — so it cannot tell the software to accept it either. Something on the providing machine reads what arrived and makes it true. That something belongs to the module, not to the mesh. The control plane decides and never touches a machine; a provisioner runs on the machine and touches it. What the mesh owns is the contract: a manifest of who asked and where each credential is, and one file per consumer holding it. It reconciles and is never told what changed, which forces three things that are each a fault somebody has shipped: set the password every time or a rotation changes nothing; remove what nobody asks for or a departed consumer keeps a login for ever; leave alone what it did not make or it cannot be run on anything that predates it. Saying where the mesh stops is the point. It decides, delivers, and can prove what it delivered; the last inch belongs to whoever knows what `create role` means. |
||
|
|
82a5b9548a |
The secret is delivered without ever being held
Written after looking at how the existing mesh does it, so this is a reaction to a measurement rather than a preference. There, credentials sit in a column encrypted at rest. Its own tooling records what that bought: the tool for finding a secret matches by value rather than by name, because the same password is in three tables, in each node's environment file in plain text, and inside every connection string composed from it — copies its documentation calls the ones usually in use. And a query against the encrypted column returns zero rows and proves nothing, so auditing moved to the decrypted copies. Encryption at rest addresses neither fault. The control plane can read what it stores, so a copy of its database is a copy of everything. And composition is what mints the untracked copies. So the value is sealed to the node that will use it before it is stored, with a key that node generated. Nothing central is composed. What it costs is auditing by value, which was never real anyway; what stays answerable is which node holds what, which is what rotation asks. What remains is a provisioner. The mesh generates the secret and tells both ends; nothing yet acts on the telling. |
||
|
|
cc872a58ce |
Binding is built except for the secret
Which turned out to be the useful way to cut it. A provider says what a consumer needs in order to use it; a consumer says where it wants to be told; the mesh adds which machine and what that machine is called on the private network. So an app on one node reaches its database on another, by a name the mesh also created. The file says it carries no credential and why, because a missing field looks like a bug and a stated absence looks like a boundary. What remains is the secret itself, and the shape it will arrive in now exists. Also: two machines wired together across no private network is refused, and that only became checkable when the network stopped being something a machine has by virtue of holding an address. |
||
|
|
80b74d32d6 |
Where the answer to a requirement is allowed to live
0009 distinguishes presence from instantiation — what the edge hands over. It never distinguished where the thing on the other end is, and that turned out to be the half doing the damage: a shell and a database were both written `requires`, so requiring a database installed one on every machine that used one. A provided name now carries a scope, as a claim already does. Scope belongs to the name rather than to each provider, or one requirement means two things depending on which module answers it. A requirement answered from the mesh is never satisfied locally. Nothing provides it, and it says which module to assign somewhere; two do, and it says how to choose. Choosing is recorded per node, because two machines may reasonably use two different databases. And knowing which node answers is the first half of handing a credential back — you cannot be given a database's password before it is settled whose database it is. |
||
|
|
90ecfe6a01 |
An edge has two directions, and only one of them is built
0009 already said a consumer supplies a target and receives a name. What it did not say is that those are two separate mechanisms. Contribution — publish me at this name, on this port — now exists. Binding — and hand me back a credential — does not, and is the larger half: a secret has to exist, be stored, reach one node and not the others, and rotate with every holder informed. That is the invariant set found violated three ways at once, so it is not something to add in passing. The absence had a measured cost. Exactly two modules opened a direct connection to the control plane's database, and they are the reason every node permanently holds a credential to it. Both were doing by hand what this edge is for. Neither needed a new kind of thing. |
||
|
|
7fe2c31bdf |
Networking is a module, and what a domain module actually is
Two records, from building it. 0009 has a section titled "there are no domain modules", and `networking` now exists. It is not a contradiction and it reads as one, so the difference is written down: what was refused contains WireGuard and a proxy and is assigned where half of it is unwanted. What exists contains nothing — requirements and a name — so there is no half. Every artifact it leads to is still an ordinary module assigned on its own terms. With the cost stated, because it is real: adding a second implementation turns a settled question into an open one for everyone using the bundle, not only for whoever wanted the alternative. That is the refusing rule applied consistently, and the alternative is a default, which is the flavor field returning under a better name. 08-connectivity gains why the network stopped being code beside the module system: a machine was on the private network because it had an address, and there was no way to keep one off. A manifest can now say its resources are computed, which is what a peer list needs. And three modules rather than one, because WireGuard is one VPN of several. Naming a module after the job and putting one implementation inside it is flavor wearing a generic name — the second VPN has nowhere to go. |
||
|
|
554f6bd7a4 |
A capability may carry a value, and adding one is not free
Recorded while building the seat detector. A capability is a named fact about a machine: its presence gates an assignment and its detail can carry a value, so "can this run here" and "what should it be configured as" are the same fact read two ways. A verdict has always had a detail beside its yes or no, so panel: oled needs no new concept. Two things that keep the set honest, both worth writing down before anyone adds the fiftieth capability. It must be detected and the detector must say how it knows -- so nobody can add one they cannot check, which is the whole of issue 007. And detectors ship inside the host, which is one static binary, so adding a capability means shipping a new host everywhere. That argues for a small general vocabulary rather than a specific one. |
||
|
|
f140303257 |
A module claims; it does not list its rivals. And flavor is retired.
Three decisions, all Jochen's, and the first is the one that unlocked it. Exclusivity is not a property of a module. It is a property of a singular resource the module takes over. Two shells compete for nothing and any number may be installed; two display servers both want the seat. So a module declares what it CLAIMS, and two modules claiming the same thing cannot both be assigned within that claim's scope. Not "xorg conflicts with wayland". Pairwise exclusion has a property that only shows up later: adding a third display server means editing xorg and wayland to know about it. Every new module requires changing modules nobody who wrote it owns, and the edits grow as the square of the count. With a claim the third one says what it claims and nothing else changes anywhere. Claims have a scope -- node, site, mesh -- which is not new. The mesh already enforces exactly one hub with a unique index. Scope is that idea said once rather than hard-coded per case. And some conflicts need no claim at all: two modules declaring the same file or binding the same port are visible from what they declare. A claim is only written for the abstract ones. A requirement with several answers is refused, never guessed. One candidate is assigned silently because there was no choice to make; none is refused naming what is missing; several is refused naming them. That is what makes a solver unnecessary -- counting candidates has no surprising behaviour, and a solver can be added later without changing a single manifest. Flavor is retired. It was carrying three unrelated meanings: variants of a thing, a subset of a module a node installs, and whatever the current system does, which earned two knowledge-base entries about going wrong. A word with three meanings cannot be reasoned about. What it reached for is two ordinary things -- different modules providing the same thing, and one module with a setting. |
||
|
|
02afb7516b |
What connecting to the mesh is, and what a node presents
Two things this record never said, both asked directly. Connecting to the mesh is one outbound AMQP connection from the node to the broker, held open. There is no second connection and nothing is ever dialled at a node. Being in the mesh means that connection is up. Two different things ride on it and conflating them is what made this murky. An AMQP account, which the mesh issues per node at enrolment, answers whether the connection is accepted at all -- per node rather than shared, because a shared one lets any node consume another's queue, which is the shared-credential fault this record exists to remove reappearing at the transport. The node's own keypair answers which node is speaking, on every message. It is not made redundant by the account: with only an account the control plane knows who is speaking because the broker says so, and that is the same transitive authority this record already refuses in the other direction. A compromised broker could attribute reports to whichever node it liked. So a node holds two things after enrolment -- a credential the mesh issued for reaching the broker, and a key it generated that the mesh only sees the public half of. Both are its own, neither reaches anything else. |
||
|
|
004057d85c |
A node's identity is a keypair it generates. This was never open.
I have been treating "what a node presents to prove it is that node" as an undecided design question for weeks, and blocking on it. It was decided. 08-connectivity says of the overlay keys: each node generates its own keypair, the private key never leaves the machine, the public key is published to the mesh -- and says explicitly that this IS ADR 0004's "a node holds its own identity", applied. Nobody had applied it to the thing 0004 is actually about. What caused it was a word. The lifecycle said a joining node receives its own durable identity, which reads as the mesh issuing something, and then the question is what. The mesh issues nothing. A node arrives holding its identity; what it receives is being known. That line now says what happens: it presents the one-time secret and its own public key, which the mesh records. The rule above it then holds literally rather than aspirationally. The mesh stores a public key, so a copy of the mesh's database grants nothing, and compromise of a node really is compromise of only that node. Also recorded, since it was asked directly: same principle as SSH, own key, not the machine's SSH host key. Host keys are regenerated by reinstalls and image clones, which would silently un-enrol a node; their lifecycle belongs to sshd rather than the mesh; and a partial host has no SSH daemon at all, so an identity scheme resting on one excludes a supported kind of node. The good half of that idea is kept: the mesh knows every node, so it can distribute host keys the way it distributes authorised keys, and node-to-node SSH stops depending on trust-on-first-use. |
||
|
|
5fd522b8da |
A node is a machine; the session is a feature of it
Correcting an overstatement from the previous commit, where I had written that a node IS a conversation. It is not. A node is a machine inside the mesh, and the session is one of the things running on it -- like the host, like any workload. That also dissolves the conflict I flagged as unresolved rather than needing anyone to decide it. 0001 says a node does not authenticate to a model provider, agents do. Still true: the session authenticates, and the session is not the machine. The node does not think, something on the node does. I had manufactured the contradiction by promoting a feature into an identity. 0001's summary row is corrected the same way, and says explicitly that neither the node's session nor a hired worker makes the node itself a thinking thing -- both run on a machine, which is what leaves that line untouched. |
||
|
|
066f14b5f8 |
A node is a conversation, and that is not the employee model
Moving this out of 0003 and out of its vocabulary. I had spent three attempts fitting the node's own session into the agent-as-employee record, each time bending hired, draining, reassigned and retired to cover something none of them describe. 0003 is back to its original text. It belongs in 0004, under what a node is, because that is what it is -- not a program installed on a node but part of the node. It holds one session permanently, anything in the mesh can message it, and it remembers across callers and across weeks. Its system prompt is the engram, which is recorded here for the first time despite running on every node. Also recorded: it has its own narrower tool list, so it can go and look rather than only report about itself; there is no authorisation between nodes, because every node is the operator's own; and how a node passes a question on is its own business rather than a protocol field. Switched off it still answers, and that is the point of having an off state rather than an absent one. A node with nothing there is a silence somebody has to diagnose. A node that says it is switched off is not. Same rule the host follows about a service that does not exist. 0001's summary is corrected too: it had one row for "agents", which is the conflation being complained about. Two rows now. A node's own session and a hired worker are built from the same parts and run on entirely different terms. Left standing and NOT resolved here: 0001 says a node does not authenticate to a model provider, agents do. A node that holds a session does. That is a real conflict between what is recorded and what runs, and it needs deciding rather than a fourth reconciliation from me. |
||
|
|
079c488d5e |
Provisioned and immutable beats exempt
Replacing the framing I wrote an hour ago. I had the node's own agent sitting outside the lifecycle as an exemption, which is a rule somebody has to remember. Provisioned the ordinary way and constrained is a rule the system enforces, and it is one row like any other rather than a category every query listing agents has to special-case. It also reads the original sentence more carefully. "Exempt from the hiring lifecycle" is exempt from hiring, not from having a lifecycle. Its lifecycle is the node's -- provisioned at enrolment, retired when the node is retired. Same states, a different thing driving them, and no exemption needed. The constraints are now the four nonsense states written as things that cannot happen rather than as an argument: not retirable, reassignable or deletable while its node exists; exactly one per node. And a distinction that was missing -- its existence is immutable, its engram is not. Freezing the personality would remove the way a node is configured. Disabling is the better half of this. A node with no agent is a silence somebody has to diagnose; a node whose agent is disabled answers saying so, immediately, with no model invoked -- the queue is still consumed and the state is the reply. That is the host's own rule about a service that does not exist, applied one tier up: absence must never be indistinguishable from a failure to answer. |
||
|
|
fd7f7557bd |
The node's own session, and why it is not hired
Answering a question that was asked three times and that I kept not answering: should the node's session just be an agent per node, since otherwise the functionality exists at two levels? Same mechanism, different lifecycle. A persistent session, accumulating memory, a system prompt, a scoped tool list, addressable by message -- identical, and building that twice is the duplication the question was worried about. What must not be shared is the lifecycle, because if a node's own voice were an ordinary hired agent it could be retired, leaving a node nothing can talk to; reassigned, moving one machine's mind onto another; hired twice, with no answer to which one replies; or never hired, leaving a node mute. The exemption in this record exists to make those four unreachable. I had this backwards earlier today and said so out loud: I called "a node itself is an agent of a kind exempt from the hiring lifecycle" a fossil of the old model and recommended striking it. It is the design. And it does not conflict with 0001 -- "the two agent rows per node merge" means one per node, not zero. I read merge as delete and invented a contradiction between two records that agree. Engrams are recorded for the first time. They are in use on every node and appear in no record, which is how a decided thing comes to look accidental. The engram is the node's system prompt, and it is what makes one node's answers recognisably its own rather than generic. Also recorded: there is no authorisation between nodes, because every node is the operator's own and a prompt from one is a prompt from them. The consequence is stated once rather than left to be discovered -- the mesh boundary is the security boundary, which is what puts the whole perimeter on the token and the overlay. And how a node passes a question on is the node's choice, not a protocol field. A node may say who is asking or may simply ask, the way a person relaying a question decides how to phrase it. That follows from the engram. The cost is that there is no machine-readable chain of who ultimately asked; each node still holds what it was asked and by whom. |
||
|
|
88ba81e9c1 |
Agents reaching nodes is the capability, not a hole in it
Correcting what I wrote an hour ago. I had recorded node-to-node SSH as "not a mesh function" and "a second control path through the back door", reasoning from ADR 0004's rule that the host has no inbound control surface. That conflated two different things and got the product backwards. There is no node-to-node SSH to forbid. The actor is always an agent; a node is only where it happens to be running -- ADR 0001 already says a node is a place where an agent can run and that is the entire relationship. An agent hired onto one node reaching another to do work is the capability the whole arrangement exists to provide. The credential is the agent's, in its own credential directory, which ADR 0001 already established. So a node's authorized_keys lists agents and never nodes, and three things follow: no node holds a key reaching another node, so 0004's "a node holds its own identity and nothing else" stays literally true; a compromised node costs the credentials of the agents that were on it rather than a way into everything; and who may reach what stays a mesh-wide fact, which is why it is identity's. The rule I misapplied is about how a node's declared state changes -- over the broker, never by being dialled. An agent with a shell is not the mesh reconfiguring a machine, it is what a person with a terminal has always been, and this design already depends on that working: the overlay is the way back in when a declaration breaks something. What such a session leaves behind is drift, and drift is what reconciliation is for. 0001 also stops underselling the fourth layer. It read as "the layer the other three exist to carry", which is true and flat. The value is that an agent can work across a set of machines as though they were one -- centrally configurable machines are ordinary; that is not. |
||
|
|
918dc04916 |
What this actually is, and three things that were assumed
Four things settled by talking them through, all of which had been true in somebody's head and written nowhere. It is not a mesh in the peer-to-peer sense and will not become one. 0001 now says what it is instead: machines linked by a private network, one node holding knowledge of all of them, modules as the way anything is built and delivered, and agents hired onto nodes to do the work. The word describes what machines can reach, not how they are governed. "Master" overstates it the other way -- nothing needs that node to keep running, only to change. 0006 gains the option that would make it a real mesh, recorded as considered rather than rejected by silence: every node holding the whole inventory, a replication process, an elected master with promotion on failure. What settles it is not the complexity but that it still would not deliver the name, because application databases are not replicated -- so a genuine peer-to-peer mesh means becoming a replicated database system for every consumer's data too. That is a larger product than the thing it would support. Also in 0006: three central roles, not one. Losing the control plane costs change, losing the broker costs being told anything, and losing the hub costs nodes in different places reaching each other at all -- which is operation, not administration. Whether they are one node is not decided. And SSH access is identity's. It appeared three times as something that uses the overlay and never as something the mesh provides, which reads as settled when nothing decided it. Nobody else could: the mesh is the only thing that knows which humans and agents exist and which nodes they may reach. Node to node SSH stays out -- the host has no inbound control surface by decision, and nodes reaching each other that way is a second control path through the back door. 0007 gains the requirement underneath all of it. Reachability was recorded as a fact to track and never as a thing some node must have. The broker's node and the hub must be dialable by every node at a stable address, or nothing can join and a disconnected node cannot return. A mesh entirely behind NAT cannot be raised. That is a precondition and it belongs with the others. The link staying on the underlay is also argued now rather than asserted. At join time it is forced; afterwards it is a choice, and the reason is that a repair channel carried over the thing being repaired is not one. Moving it onto the overlay, with fallback, is recorded as open with what it would have to get right -- a WireGuard interface has no link state to test, and a silent fallback is this repository's recurring fault in a new place. 0010 says in one line what was the intention throughout: the module system is the CI/CD. Not a pipeline beside the mesh. Build, test, publish and deploy are one reconciliation seen at four points, which is why a thing that cannot be a module cannot be delivered. |
||
|
|
5218b06c02 |
Fold the control plane's build decisions into 0006 and 0008
Back to 23 records. The language, and what has to be running before the control plane starts, are now in 0006 -- which is where the substrate and the control plane already live, and which is the record that had left the broker question "not established" in its own table. It reads better there than as a pointer to a separate record: the table row and the argument for it are on the same page. The store mechanics went into 0008. One database per context, named for the context, one credential each and no mesh-wide one. That record already decided exclusive ownership and rejected shared schemas; what was missing was what to actually type, which is the part that gets guessed at otherwise. Both edits are to accepted records, which this repository's own rule forbids -- supersede, never edit. Recorded here so it is visible rather than silent. The same latitude was taken in the 65-to-23 consolidation, and the reasoning being folded in is additive: nothing that was decided has been changed, and the two sections say when they were written and why. |
||
|
|
84f4425fd6 |
The broker precedes the control plane, and it is written in Go
Two things found by trying to build tier 2. The substrate design asked whether the message broker has to be running before the control plane, and framed it as depending on whether the control plane's own parts talk to each other over it. They do not -- it is one process -- so under that framing the broker stays out of the bundle. The framing cannot answer the question. What decides it is how the control plane reaches a node, and the answer was already decided: only ever over the link, and the link is the broker. So provisioning the broker would require the broker. The first node does not escape this by being local, because it enrols the ordinary way, by dialling the broker at the address in its token -- which was deliberate, and worth keeping. The bundle is two images now. The record says what that costs, including a certificate the broker needs at a moment when there is no mesh to issue one. The language had never been decided for tier 2. Go, for the same reason the host is: the bundle pins this image by digest and runs it where nothing can check it, so the image should hold the program and nothing else. Also corrects something already built: the bootstrap created one database and called it 'mesh'. ADR 0008 grants a context only what it exclusively owns and ADR 0006 says the mesh database names a thing that will not exist. One database per context, so one today, called inventory. |
||
|
|
6a2b107fb8 |
Restore a consequence the consolidation dropped
I said nothing was lost when 65 records became 23. That was too strong, and here is a counterexample: ADR 0046's consequence that the lab needs a way to place images did not survive into the merged substrate record. The compression kept the decision and dropped one of the things it implied. It was not lost from the repository -- 04-ISSUES/009 had already picked it up, which is why it was found at all. But the record no longer carried it, and the record is where somebody would look. Restored, now as a resolved fact rather than an open consequence: the lab raises a registry inside the scenario, which is the real path since that is what every node after the first pulls from. The digests it serves are its own, and that satisfies the pinning rule -- what is required is a reference that is exact and cannot move. Worth recording the wrong assumption too, because it is what made this look impossible for two days: I took "pinned by digest" to mean the UPSTREAM digest had to be preserved. It does not. Any digest that is exact and immutable satisfies the rule, and a registry assigns one. |
||
|
|
b4607dfc03 |
Numbers are identity; the reading order is a generated, checked index
Decided after measuring what renumbering actually costs: 96 references in code comments across two repositories, none of which would have failed to compile. They would have pointed at the wrong reasoning, which is worse than a broken link because nothing reports it. So a number identifies a record and never changes. It cannot also be a position -- a position moves when the set changes, and an identity that moves is not one. The reading order moves into an index generated from each record's `topic:`. Six topics, in the order somebody learns the system. The index is WRITTEN rather than only generated on demand, which reverses what this repository previously said. The reason it said otherwise is that a hand-written index drifts -- but a reader looking at the folder on a forge sees the folder, not a command, and the drift objection is answered by checking rather than by refusing to write one. That is §5's own rule: a rule states how it is checked. Two checks, both confirmed to bite. index.py fails when the written order no longer matches the records. records.py fails when a record has no topic or one nobody defined -- the quiet failure being a record that vanishes from the order rather than appearing in the wrong place. |
||
|
|
333356cff3 |
Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when things happened to be decided, which after consolidation is fictional anyway since record 5 alone folds decisions taken across a week. Concretely wrong before: the domain statement sat at 8, after five engineering rules; the constitution was scattered across 5, 12 and 17; the tiers landed at 15, 16, 21 and 22 with process records in between. Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what runs on them and how it gets there (9-10), how it is built (11-16), how it is checked (17-18), how we work (19-23). Two things made this safe rather than free. It is a permutation, not a compaction, so the renames go through temporary names -- otherwise two files want one slot and one is lost. And the reference rewrite is a single simultaneous pass, because almost every number moved into a slot another number was vacating; replacing one at a time would have cascaded and pointed things at the wrong record while still resolving. Verified: 284 [ADR NNNN](path) links across the repository, all with matching text and target. The ordering principle is now stated in 19 rather than left implicit -- the repository already said "the numbering is the flow" about its folders, and there was no reason for the records to be the exception. |
||
|
|
e1febe8e0f |
Renumber the records 1 to 23
The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18, 19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only the archaeology of what used to be there. Renumbered contiguously. Renames run in ascending order, so every target number is already free and no two files ever collide. The reference rewrite is one simultaneous pass rather than a sequence of replacements. Numbers moved into slots other numbers were vacating -- the node host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time would have cascaded and silently pointed things at the wrong record. Seven plain-text references survived the merges as prose rather than links, naming records that no longer existed: the enrolment token, the link boundary, what a declaration is, reachability, the repository structure. Each mapped to the consolidated record that now holds it. Verified rather than assumed: every [ADR NNNN](path) link now has matching text and target, checked across the whole repository, and the checker passes. Frontmatter `consolidates:` lists dropped -- they named records that are gone, and each consolidated record already says in prose what it absorbed. |
||
|
|
77f3a4cea7 |
Consolidate: 65 decision records to 23
Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.
the node host 8 -> 1 applies not decides, depends on nothing,
per operating system, root service, the
launcher, episodic, what a declaration is,
actions from the bundle only
a node and how it joins 4 -> 1 what a node is, joining, the link as
security boundary, the enrolment token
modules and the graph 7 -> 1 everything is a module, no domain modules,
three edges, provisioning, the core library
substrate and control 6 -> 1 the test, seven contexts, one control plane,
plane the authority is not a database, the named
products, the pinned bundle
connectivity 3 -> 1 a route is a grant, reachability declared,
filter rules
delivery 5 -> 1 reconciliation not a pipeline, artifacts,
the three silos, a failed step, the verdict
the lab 5 -> 1 (earlier)
how this repository 10 -> 1 (earlier)
works
Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.
The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.
The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
|
||
|
|
5e83ac2c22 |
Consolidate: 65 decision records to 52
Jochen: a normal application has 3-5 ADRs, maybe 10 for a large one, and we are at 65. Fair, and the cause is mine -- I recorded every FINDING as a decision rather than every fork in the road. Two merges, both cases where one decision had been split across many records because it was taken over several days rather than at once. 0019 absorbs ten records about how this repository works: what it is and that it is public, the folder flow, the two design layers, the issue front door, status in frontmatter, playbooks, the naming rule, the product name. Those were never ten decisions -- they were one, seen from ten angles as the repository took shape. 0016 absorbs the five about the lab: a node is a virtual machine, a router is scenery, a scenario declares the underlay, a scenario is a closed address space, and the two scenario classes. Same pattern -- one design, split by the order it was worked out in. The consolidated 0019 also raises the bar for what earns a record, since that is what produced 65: a record is warranted when there is a genuine fork -- a direction reversed, an alternative that will be proposed again, something contested. A finding is not a decision, and a bug is certainly not. Everything else belongs in the design document where the reasoning is actually read. The checker earned its place here. Deleting nine records left 13 dangling links across the repository and it named every one, including in AGENTS.md. Nothing was found by reading. Remaining clusters worth the same treatment: the host (8 records), delivery (5), modules (6), connectivity (4), substrate and control plane (4). That would be 52 down to roughly 30. |
||
|
|
10365f2eae |
Consolidate the design layer: one place per topic
Jochen: a jungle of specs that slightly contradict or patch each other, and what matters is a working state rather than history. Both are fair and both are mine. Measured rather than assumed. 05-the-node-host and 09-the-node-lifecycle both covered enrolment, the install commands, the unit file, the launcher and reconcile -- I wrote 09 without taking anything out of 05, so the same things were said twice and could drift apart. Split by what each document IS. 05 is the component: what the host is, its parts, the declaration vocabulary, the build order, how it is verified. 09 is what happens to it: install, enrol, run, upgrade, retire. The whole "The process" section left 05, and the unit file moved to 09 where installing is described. 05 goes from 338 lines to 245 and now points at 09 rather than restating it. 09 also carried a 105-line "Resolved" section -- six mechanisms framed as "these were open and here is the answer". The content is needed; the framing is history, and history is what makes a document read as a changelog rather than a description. Renamed to what it actually is and the was-open phrasing removed. Also added 10-delivery.md, which did not exist: four accepted decisions -- 0054, 0063, 0064, 0065 -- had no design document at all, which is the specific reason the delivery picture felt scattered. It is now one document covering modules, the three edges, the core library, and how a change becomes a running thing, with a table of what each property is designed against and what must exist before it can be built. |
||
|
|
9d091c81e0 |
A build edge, a core library that is a domain, and 0063 corrected
Three things from walking a real dev cycle through 0063, all of which Jochen caught by pushing on where I had glossed. 0064 -- a build edge is a third kind. Research 011 established presence and instantiation, and both are RUNTIME edges: they answer what a module needs in order to run. Delivery needs a different question -- what has to be rebuilt when this changes -- and that relationship is fixed inside an artifact rather than negotiated when it runs. So the graph as designed could not drive delivery, which is the real reason 0063 was not approvable. It is derived rather than declared, read from what a module actually imports, because a declared list and the imports it describes drift and the imports are the true ones. The runtime edges stay declared, and that asymmetry is not an inconsistency: a runtime edge is an intention somebody has, a build edge is a fact about code that exists. It also makes design quality measurable. A module with many inbound build edges is one whose every change is expensive, and the current shared library is exactly that -- nobody could see it because nothing drew the edges. 0065 -- the core library is the mesh's domain. Jochen disagreed with 0030's "types, not behaviour" and was right: that guard is aimed at the wrong thing. A library everything depends on is a hub whether it holds types or code, and the fan-in is what makes a change expensive. So types ship with the module that owns them -- trading one wide edge for several narrow ones -- and the core library holds what is true of the mesh regardless of context, which research 011 already found: a module, a node, an assignment. The test is "would this still mean the same thing in a context that had never heard of the one it came from". A node does; a pipeline stage does not. Domain-driven is the point rather than the label: "who else might want this" always answers yes, which is how the current one grew. And it changes the check for the better. "The build output contains no runtime code" would have enforced a rule now withdrawn. Inbound build edges is a measurement rather than a prohibition, and it is visible while a hub is forming rather than after. 0063 revised on both counts, plus a third: I had written "the lab judges it" as though that were a step. A lab run takes tens of seconds, occupies a VM, and fails for environmental reasons -- and a shared-library change produces dozens. One expensive non-deterministic gate fails both ways, and neither failure looks like itself. Verdicts are now tiered, and a run that failed environmentally is explicitly not a verdict. 0063 also now carries what must exist before it can be implemented, rather than leaving that to be discovered. |
||
|
|
4ab8a0507f |
Delivery is reconciliation, not a pipeline; research 008 closes
Jochen: don't rebuild the current coordinator, use it as a pitfall list. That reframed the last open question rather than answering it. 0058 stopped deploy being a stage that pushes to nodes, and said plainly what it did not fix: detection. A merge that created no pipeline, and nothing said so. That is not a defect in the detector -- it is what happens when correctness depends on an event ARRIVING. 0063 applies 0058's move one level up. The control plane holds what source exists and what has been built from it, and builds the difference. A change becomes a build because source is ahead of artifacts, which is a comparison answerable at any moment. An event makes it fast; nothing makes it necessary, so a missed webhook costs latency and cannot cost correctness. The mesh becomes one idea at two layers: the control plane reconciles artifacts against source, the host reconciles machine state against declarations. The pipeline as a state machine disappears, and with it the stage list that a verify step was once omitted from. That reframing answered the three questions still open in 008, so it graduates with all six closed. A deployed state is two comparisons rather than an event. A verdict is about an ARTIFACT and gates whether it may be declared -- sharper than the question expected. And "before self-hosting" mostly dissolves, because a reconciler needs source and artifacts as bindings where a pipeline's stages name their targets. Four costs recorded, and one is a real risk rather than a trade: a reconciler that cannot reach its target retries forever, and without something noticing, the failure is silence -- the exact fault this removes, reintroduced elsewhere. Also named: the run identity people actually use is lost, and "did my change go out?" needs a replacement or this will be worse to live with than what it replaces, whatever its properties. |
||
|
|
f728c3fd98 |
File 009: a digest-pinned image cannot be placed in the lab
Two accepted decisions collide, and testing found it rather than review. 0046 pins images by digest and has the host refuse anything unpinned. The lab places images by exporting them from the workstation, because a sealed scenario cannot reach a registry -- and that loses the digest, since a repo digest only exists for an image a registry served. Measured: the load says 'Loaded image ID:' rather than 'Loaded image:', and the image lands dangling. So a tag is refused by the host and a digest is unusable in the lab. There is currently no declaration the lab can raise that exercises the container shape, which matters because the container shape IS the substrate -- every bootstrap step past the runtime is one. The resolution is a registry inside the scenario, and that is not a workaround: 0048 already names an OCI registry as substrate and every node after the first pulls from the mesh's own. It also removes the lab's export-and-push mechanism rather than repairing it. 0046 now carries a pointer, since its own consequence is where the collision was predicted -- half of it is closed and the other half turned out to be harder than 'not solved here' suggested. |
||
|
|
ba0d01788e |
0062: a host may be episodic; 0060's Android gap closed
0060 named the gap and did not close it: everywhere else an init runs the launcher at boot, and Android grants neither an init to register with nor anything worth supervising, because a supervisor would be killed alongside what it supervises. Closed by narrowing what is required rather than building something. A host is resident or episodic, and both are hosts. Being killed by the platform is disconnection, which 0036 already made ordinary -- and every mechanism an episodic host needs already exists because it was built for laptops that close. A partial host can join a mesh and cannot be the first node, since every bootstrap step is a shape it refuses. Its bundle says so. Two consequences that are easy to miss: last-heard-from means much less on an episodic host, so a healthy phone reads as a dead server unless the reader knows which kind it is; and a declaration may take a long time to land, which makes 0058's outstanding-versus-failed distinction load-bearing. Still open, and in that order: what an Android node is FOR, and only then how it is started. |
||
|
|
f1b1cd9aa0 |
Review: three ADRs no longer said what we had concluded
A sweep for claims overtaken by the last few days. Annotated rather than rewritten, following the pattern already in 0049 -- what changed and why is the useful part, and an accepted record should not quietly become something else. 0057's init section was wrong on all three of its claims. It said the host needs FOUR things from an init; 0061 reduced that to one. It said every machine the mesh targets already has systemd; Alpine does not, and it is the intended first node. It said there is no second init to abstract over; there is now, and the answer is still not an abstraction -- it is a four-line file per system. What survives is the part that was always right: an init is not a dependency in 0041's sense, because it is not installed, it is what the machine already is. 0048 named Docker as the container runtime. It is now docker or podman, detected rather than chosen -- because adoption keeps what a machine already has, so naming one contradicted a rule already decided. That row is the only one of the five that names two, and the record now says why. 0060 claimed the bundle is portable across operating systems. Its mechanism is; its contents are not -- package names, unit names, service names all differ, so an Arch host embeds an Arch bundle. That was my error, and it is the exact confusion behind the question that found it. The design layer had the same drift: 07 and 09 said "Docker" where they meant a container runtime, 09 said systemd restarts the host after an upgrade when the launcher does, and both install snippets assumed Arch. They now show Alpine and Arch side by side, which makes the point better than prose did -- step 1 differs per system, step 2 never does. Checked and NOT changed: 0047's "the vocabulary grows by one shape" is a claim about the rate, not the count, and is still true. 0037 lists docker among tools the host manages, which it does. 0041 says nothing about either. |
||
|
|
66df0eb53e |
0061: the launcher supervises; init is asked for one thing
The record said an init is asked for two things -- start at boot and restart on exit -- which was half a change. It moved the give-up logic out of unit files and left the restart in one, so init still decided when the host came back. The launcher no longer execs the host. It supervises it, so restarting is ours too, and init is asked only to run it at boot. There is an OpenRC script beside the systemd unit now. Records the cost honestly: not exec'ing means the launcher must trap the shutdown signal and pass it down, because a supervisor that exits while its child runs leaves the host to be killed rather than to stop. And records what the implementation found: the counter counts consecutive FAILURES, not starts. Counting starts meant a host that upgraded itself three times rolled itself back, having worked perfectly every time -- because a clean exit IS the upgrade path. That is now the second time a clean exit has been mishandled, so it is called out as the thing to check. |
||
|
|
c557f99cba |
Record what testing podman actually showed
0060 said the container runtime was a separate decision. It is now made, and the reasoning is worth keeping because it is the opposite answer to the same question one paragraph earlier. Abstracting service managers is lossy -- systemd and OpenRC are different models and LoadState has no equivalent. Container runtimes converged on one CLI deliberately, so almost nothing is lost: checked against podman 6.1.0, run, rm -f and docker's own template syntax for state and labels all work unchanged. Only the probe differs. So: a two-entry lookup, not an interface. The difference that is NOT in the CLI is the one that would have shipped silently. Podman accepts --restart unless-stopped, records it, and has no daemon to act on it -- containers do not return after a reboot unless podman-restart.service is enabled, which by default it is not. Every command reports success and the effect does not happen. That belongs in the declaration rather than the host: a node using podman is told to enable the unit. Which is what made the service shape's missing 'boot' field visible, and it is now built. |
||
|
|
e1ad39b500 |
Per-OS hosts, and an init asked for only start and restart
0060 -- the host is built per operating system. systemd and pacman are the Arch host's implementation, not abstractions the mesh has to grow. They are not independent choices: a machine has pacman because it is Arch, and the package manager, service manager and packaging format arrive together as one decision somebody made at install time. Rejected abstracting them, and the reason is correctness rather than effort. The service applier reads LoadState to tell "not installed" apart from "stopped", which is what stops it reporting absence as success. An interface spanning systemd and OpenRC degrades to what both express, and the lowest common denominator is exactly where that fault lives. Almost all of it is shared -- the vocabulary, store, apply loop, read-back discipline, refusal model, bundle and link are portable. Two appliers differ. And delivery was already per-OS, since a .pkg.tar.zst is an Arch artifact, so this is the seam that already existed. Android is the interesting case rather than Debian: no service manager, no package installation, usually no root. Such a host implements file, directory and action and refuses the rest -- the same refusal a host already gives an unknown type, with a different reason. Those three are the portable floor. The container runtime is deliberately left open: it is not an OS split, since Arch runs docker or podman. 0061 -- the init is asked for start-at-boot and restart-on-exit, and nothing else. Both are expressible in OpenRC, runit, s6 and an Android init.rc. Counting failed starts and rolling back moves into a launcher, because that is the one piece which must work when the host does not, and a script with a counter can be tested where OnFailure= can only be hoped for. Supersedes 0059, keeping its reasoning in full. The checker found all six places citing 0059 and refused the commit until they named the replacement. |
||
|
|
19997d56c3 |
Approve 0057-0059, with four corrections from review
Not approved as drafted -- four things came out of checking them against each other, and one was a bug that would have broken every upgrade. The bug: 0059 specified Restart=on-failure while 0057 has the host restart onto a new binary by exiting CLEANLY. on-failure does not restart a process that exited zero, so every upgraded node would have been left stopped, having successfully upgraded. Found by reading the two records against each other rather than by either alone. Now Restart=always in all three places that mention it. The host cannot run in a container, and the reason is decisive rather than stylistic: step 0 of the substrate bootstrap installs the container runtime, so a host inside a container would need the thing it exists to install. It would also break 0041 -- copy it onto a machine and run it stops being true when the machine must already have a runtime. Everything above tier 0 is a container; the host is not. That split is the tier boundary, not an inconsistency. systemd is named rather than abstracted. An init is not a dependency in 0041's sense: 0041 is about what must be installed before the host works, and an init is not installed, it is what the machine already is. The unit file is the only systemd-specific artefact and it belongs to the package, so a machine with a different supervisor ships a different package. The mesh is a watchdog, and my first draft was half an answer. Recovery must be local -- nothing dials a node, and a host that cannot start cannot report. But detection is the mesh's, and a local supervisor structurally cannot do it: it sees one process failing and cannot tell a broken machine from a broken release. Only something watching every node can, and that distinction decides whether the response is "fix this machine" or "stop shipping this version". So a host rollout is staged -- a few nodes, wait for heartbeats, continue or stop on silence. Local rollback still needed, because the canary nodes break and because a node offline during the rollout gets the declaration later with no batch around it. The first declaration is the overlay and nothing else. Forced, because a node's address and peers are assigned rather than chosen. But also the way back in: a node reachable over the overlay can be fixed by hand if a later declaration breaks it, and a large first declaration risks a node that is broken and unreachable at once. Also stated plainly, because it reads as a contradiction: nodes reach each other over the overlay and every node consumes from the broker; what 0039 forbids is an inbound CONTROL surface, not reachability. And in 06: no node holds a credential to any control-plane store, for reads or writes. Four ADRs already say this separately and none of them said it in one place. Nodes state over the broker; the owning context writes. With a note that most high-frequency writes are observability's, not the registry's -- routing logs into the registry would be the shared-schema mistake arriving through a door marked performance. |
||
|
|
605c9fd441 |
Changes are pushed, not polled; and a stuck host rolls itself back
Two corrections and one new decision, all from Jochen catching things. Pushed, not polled. I described updates as landing "on the next reconcile", which reads as polling and is not the design. A declaration arrives as a message on a link that is already open; the host applies it then. Polling over an existing connection would be slower to land AND constant traffic to learn nothing. The timer is for drift and nothing else, and it cannot be replaced by an event for a definitional reason: drift is change the mesh did not make -- somebody edited a managed file, a distribution upgrade replaced a config -- so nothing will ever publish a message about it. Only looking finds it. Separated the heartbeat from the reconcile timer, which I had been conflating. They point in opposite directions and answer different questions: the timer looks at the machine and asks whether it still matches; the heartbeat reports upward and is what makes silence mean something. A node with nothing to do sends nothing, and without a heartbeat that is indistinguishable from a node that stopped. 0059 -- a host that cannot start is rolled back by the service manager. I had left this open on the grounds that recovery meant the host judging its own health. That objection does not survive being asked properly: a keepalive is something else judging the host. The watchdog must be local, because nothing dials a node and a host that cannot start cannot report -- so it is the service manager, which is already there. The failure it prevents is sharper than "the node is down": a host that will not start looks exactly like a machine somebody switched off, which is the one condition this design has deliberately decided not to alarm on. So a bad release reaches every node, each goes quiet, and the mesh reports a fleet of sleeping laptops. Confirmed means started and completed one reconcile -- deliberately not "the link is up", or a laptop on a train would roll itself back. The rollback is a script shipped by the package, not a host subcommand, because a binary that will not start cannot be its own recovery. It rolls back once: a second failure means the machine is the problem, not the binary. Also refined the records checker, which produced a false positive: a proposed record may extend another proposed one, because decisions are drafted in chains and the alternative is marking things accepted to satisfy a check. An accepted document resting on a proposed record still fails, and that was verified. 0057, 0058 and 0059 are all proposed. |
||
|
|
aeea2a9f9a |
Resolve the host lifecycle's open items, and say how the host is delivered
The upgrade question turned out to be a delivery question, so 0058 answers both. Today's third silo runs once per node and sends each one a command to install and start. That is where the as-is records a package install that 404ed from every mirror while the job went green, an image pull failure that did not fail the deploy, and a verify stage that was built and never scheduled because it was missing from a list. The shape underneath all of those is that the thing reporting success was not the thing doing the work. Meanwhile ADR 0037 has given every node a component that applies state, reads back and reports -- so two mechanisms now change a node and only one checks its work. 0058: a pipeline ends when the declaration is updated. Deploy stops sending commands to nodes and becomes one write. The host applies it on its next reconcile, and the host cannot report success it did not verify. The verify stage disappears as a stage, which is the point -- verification stops being a step that can be left off a list. A pipeline result now means "the declaration is updated, and here is which nodes have applied it". It does not wait for every node, because a node may be legitimately switched off for a week. Outstanding is reported separately from failed, since conflating them is how the old system produced a stall with no error anywhere. The host is delivered by exactly this path and needs no new resource type: a `file` writes the package manager's config pointing at the mesh's repository, a `package` names the version. Added a step I had missed -- before exiting for a restart, the host runs the new binary once. A package can install something that does not execute here, and that turns "the node never came back" into "the apply failed and said why". Six open items resolved: re-enrolment is decided when the token is issued and revokes the previous identity; the mesh keeps a recovery copy of what each node reports it owns, which un-strands the orphans; last-contact is reported with no threshold, because a laptop off for three weeks is doing nothing wrong; adoption always completes but a failed line makes a node ineligible for assignment; a briefing is a structured document whose outcome is computed from its lines; and the token is printed once and carried by hand, which is the property that makes it worth anything. Still open and named: automatic rollback of a host version that will not start. 0057 and 0058 are both proposed. |