77a1493df4c569fc5ef8b1132deb03ae64e6c028
193
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
66df0eb53e |
0061: the launcher supervises; init is asked for one thing
The record said an init is asked for two things -- start at boot and restart on exit -- which was half a change. It moved the give-up logic out of unit files and left the restart in one, so init still decided when the host came back. The launcher no longer execs the host. It supervises it, so restarting is ours too, and init is asked only to run it at boot. There is an OpenRC script beside the systemd unit now. Records the cost honestly: not exec'ing means the launcher must trap the shutdown signal and pass it down, because a supervisor that exits while its child runs leaves the host to be killed rather than to stop. And records what the implementation found: the counter counts consecutive FAILURES, not starts. Counting starts meant a host that upgraded itself three times rolled itself back, having worked perfectly every time -- because a clean exit IS the upgrade path. That is now the second time a clean exit has been mishandled, so it is called out as the thing to check. |
||
|
|
c557f99cba |
Record what testing podman actually showed
0060 said the container runtime was a separate decision. It is now made, and the reasoning is worth keeping because it is the opposite answer to the same question one paragraph earlier. Abstracting service managers is lossy -- systemd and OpenRC are different models and LoadState has no equivalent. Container runtimes converged on one CLI deliberately, so almost nothing is lost: checked against podman 6.1.0, run, rm -f and docker's own template syntax for state and labels all work unchanged. Only the probe differs. So: a two-entry lookup, not an interface. The difference that is NOT in the CLI is the one that would have shipped silently. Podman accepts --restart unless-stopped, records it, and has no daemon to act on it -- containers do not return after a reboot unless podman-restart.service is enabled, which by default it is not. Every command reports success and the effect does not happen. That belongs in the declaration rather than the host: a node using podman is told to enable the unit. Which is what made the service shape's missing 'boot' field visible, and it is now built. |
||
|
|
e1ad39b500 |
Per-OS hosts, and an init asked for only start and restart
0060 -- the host is built per operating system. systemd and pacman are the Arch host's implementation, not abstractions the mesh has to grow. They are not independent choices: a machine has pacman because it is Arch, and the package manager, service manager and packaging format arrive together as one decision somebody made at install time. Rejected abstracting them, and the reason is correctness rather than effort. The service applier reads LoadState to tell "not installed" apart from "stopped", which is what stops it reporting absence as success. An interface spanning systemd and OpenRC degrades to what both express, and the lowest common denominator is exactly where that fault lives. Almost all of it is shared -- the vocabulary, store, apply loop, read-back discipline, refusal model, bundle and link are portable. Two appliers differ. And delivery was already per-OS, since a .pkg.tar.zst is an Arch artifact, so this is the seam that already existed. Android is the interesting case rather than Debian: no service manager, no package installation, usually no root. Such a host implements file, directory and action and refuses the rest -- the same refusal a host already gives an unknown type, with a different reason. Those three are the portable floor. The container runtime is deliberately left open: it is not an OS split, since Arch runs docker or podman. 0061 -- the init is asked for start-at-boot and restart-on-exit, and nothing else. Both are expressible in OpenRC, runit, s6 and an Android init.rc. Counting failed starts and rolling back moves into a launcher, because that is the one piece which must work when the host does not, and a script with a counter can be tested where OnFailure= can only be hoped for. Supersedes 0059, keeping its reasoning in full. The checker found all six places citing 0059 and refused the commit until they named the replacement. |
||
|
|
19997d56c3 |
Approve 0057-0059, with four corrections from review
Not approved as drafted -- four things came out of checking them against each other, and one was a bug that would have broken every upgrade. The bug: 0059 specified Restart=on-failure while 0057 has the host restart onto a new binary by exiting CLEANLY. on-failure does not restart a process that exited zero, so every upgraded node would have been left stopped, having successfully upgraded. Found by reading the two records against each other rather than by either alone. Now Restart=always in all three places that mention it. The host cannot run in a container, and the reason is decisive rather than stylistic: step 0 of the substrate bootstrap installs the container runtime, so a host inside a container would need the thing it exists to install. It would also break 0041 -- copy it onto a machine and run it stops being true when the machine must already have a runtime. Everything above tier 0 is a container; the host is not. That split is the tier boundary, not an inconsistency. systemd is named rather than abstracted. An init is not a dependency in 0041's sense: 0041 is about what must be installed before the host works, and an init is not installed, it is what the machine already is. The unit file is the only systemd-specific artefact and it belongs to the package, so a machine with a different supervisor ships a different package. The mesh is a watchdog, and my first draft was half an answer. Recovery must be local -- nothing dials a node, and a host that cannot start cannot report. But detection is the mesh's, and a local supervisor structurally cannot do it: it sees one process failing and cannot tell a broken machine from a broken release. Only something watching every node can, and that distinction decides whether the response is "fix this machine" or "stop shipping this version". So a host rollout is staged -- a few nodes, wait for heartbeats, continue or stop on silence. Local rollback still needed, because the canary nodes break and because a node offline during the rollout gets the declaration later with no batch around it. The first declaration is the overlay and nothing else. Forced, because a node's address and peers are assigned rather than chosen. But also the way back in: a node reachable over the overlay can be fixed by hand if a later declaration breaks it, and a large first declaration risks a node that is broken and unreachable at once. Also stated plainly, because it reads as a contradiction: nodes reach each other over the overlay and every node consumes from the broker; what 0039 forbids is an inbound CONTROL surface, not reachability. And in 06: no node holds a credential to any control-plane store, for reads or writes. Four ADRs already say this separately and none of them said it in one place. Nodes state over the broker; the owning context writes. With a note that most high-frequency writes are observability's, not the registry's -- routing logs into the registry would be the shared-schema mistake arriving through a door marked performance. |
||
|
|
605c9fd441 |
Changes are pushed, not polled; and a stuck host rolls itself back
Two corrections and one new decision, all from Jochen catching things. Pushed, not polled. I described updates as landing "on the next reconcile", which reads as polling and is not the design. A declaration arrives as a message on a link that is already open; the host applies it then. Polling over an existing connection would be slower to land AND constant traffic to learn nothing. The timer is for drift and nothing else, and it cannot be replaced by an event for a definitional reason: drift is change the mesh did not make -- somebody edited a managed file, a distribution upgrade replaced a config -- so nothing will ever publish a message about it. Only looking finds it. Separated the heartbeat from the reconcile timer, which I had been conflating. They point in opposite directions and answer different questions: the timer looks at the machine and asks whether it still matches; the heartbeat reports upward and is what makes silence mean something. A node with nothing to do sends nothing, and without a heartbeat that is indistinguishable from a node that stopped. 0059 -- a host that cannot start is rolled back by the service manager. I had left this open on the grounds that recovery meant the host judging its own health. That objection does not survive being asked properly: a keepalive is something else judging the host. The watchdog must be local, because nothing dials a node and a host that cannot start cannot report -- so it is the service manager, which is already there. The failure it prevents is sharper than "the node is down": a host that will not start looks exactly like a machine somebody switched off, which is the one condition this design has deliberately decided not to alarm on. So a bad release reaches every node, each goes quiet, and the mesh reports a fleet of sleeping laptops. Confirmed means started and completed one reconcile -- deliberately not "the link is up", or a laptop on a train would roll itself back. The rollback is a script shipped by the package, not a host subcommand, because a binary that will not start cannot be its own recovery. It rolls back once: a second failure means the machine is the problem, not the binary. Also refined the records checker, which produced a false positive: a proposed record may extend another proposed one, because decisions are drafted in chains and the alternative is marking things accepted to satisfy a check. An accepted document resting on a proposed record still fails, and that was verified. 0057, 0058 and 0059 are all proposed. |
||
|
|
aeea2a9f9a |
Resolve the host lifecycle's open items, and say how the host is delivered
The upgrade question turned out to be a delivery question, so 0058 answers both. Today's third silo runs once per node and sends each one a command to install and start. That is where the as-is records a package install that 404ed from every mirror while the job went green, an image pull failure that did not fail the deploy, and a verify stage that was built and never scheduled because it was missing from a list. The shape underneath all of those is that the thing reporting success was not the thing doing the work. Meanwhile ADR 0037 has given every node a component that applies state, reads back and reports -- so two mechanisms now change a node and only one checks its work. 0058: a pipeline ends when the declaration is updated. Deploy stops sending commands to nodes and becomes one write. The host applies it on its next reconcile, and the host cannot report success it did not verify. The verify stage disappears as a stage, which is the point -- verification stops being a step that can be left off a list. A pipeline result now means "the declaration is updated, and here is which nodes have applied it". It does not wait for every node, because a node may be legitimately switched off for a week. Outstanding is reported separately from failed, since conflating them is how the old system produced a stall with no error anywhere. The host is delivered by exactly this path and needs no new resource type: a `file` writes the package manager's config pointing at the mesh's repository, a `package` names the version. Added a step I had missed -- before exiting for a restart, the host runs the new binary once. A package can install something that does not execute here, and that turns "the node never came back" into "the apply failed and said why". Six open items resolved: re-enrolment is decided when the token is issued and revokes the previous identity; the mesh keeps a recovery copy of what each node reports it owns, which un-strands the orphans; last-contact is reported with no threshold, because a laptop off for three weeks is doing nothing wrong; adoption always completes but a failed line makes a node ineligible for assignment; a briefing is a structured document whose outcome is computed from its lines; and the token is printed once and carried by hand, which is the property that makes it worth anything. Still open and named: automatic rollback of a host version that will not start. 0057 and 0058 are both proposed. |
||
|
|
2204b01909 |
Design the node lifecycle end to end
The host was described as a component and never as something that runs for years on a machine somebody else also uses. 09 covers every state a machine can be in and every transition between them. Four states: unmanaged, hosted, enrolled, disconnected. Only the last two are nodes, and they are the same node in two situations. `hosted` -- the host installed but never told which mesh it belongs to -- had no name before and is where a machine sits between the two adoption commands. Things that were unclear and now are not: The first node walks the same path in an unusual order: reconcile from the bundle, the control plane it just raised issues a token, enrol against it. Its specialness lasts two commands. A side effect worth having -- enrolment is exercised on node one, rather than being written and first used on node two. Enrolment reports profile and inventory BEFORE the control plane decides anything. The profile is the input to that decision, not a diagnostic; the control plane cannot decide what a machine should run without knowing what it can run. Rebooting mid-apply is safe by construction. The store records each resource after it worked, so a host that dies half way through comes back and applies the rest. The rule that stops the host lying about what it did also makes it crash-safe. Retiring splits in two. Graceful is a final empty declaration. A node that is gone will reconcile its last declaration forever -- the honest consequence of making disconnection ordinary. The answer is not to make the host expire but that the node holds nothing that outlives revocation: every grant is a per-node credential revoked at the provider. A lost node keeps running and stops being able to reach anything. Said plainly rather than implying the mesh can switch a machine off, which it cannot and should not. Losing the store is quiet and permanent, so it gets its own section. The host re-enrols and re-applies fine; what does not come back is removal, because resources it no longer has a record of become unowned and sit there indefinitely. Also corrects 0057, which said the mesh must not upgrade the host at all. That conflated two acts. Replacing the binary is safe -- Unix keeps the running inode. Stopping the unit is not. So the host may apply a package naming itself, and restarts by finishing its apply and exiting cleanly, letting the supervisor start it on the new binary. It never asks the service manager to restart it. That makes a fleet-wide host upgrade an ordinary declaration, which the first draft gave up on. 0057 remains proposed. |
||
|
|
3ab11c96ef |
Say what the host process is: a root service, installed as a package
The design described what the host does and never what it is at runtime. The words daemon, long-running, interval, poll and heartbeat appeared nowhere in it or in the relevant decisions. What exists is a command that runs and exits; what the design needs is a process holding a link. Nobody had written down that those differ, so several questions had no answer. 0057 settles them. It runs on every node -- the host is what makes a machine managed, so a machine without one is not a node. Root, because no useful subset of the job is unprivileged. A systemd unit, because something must survive a reboot to hold the link. It never manages its own unit. The temptation is obvious and it ends with a host stopping itself half way through an apply, leaving a machine with nothing running to fix it. The installation owns the host; the host owns everything else. Installed as a package, with a tarball as the floor. The package carries the unit file, the state directory and an upgrade path, which a bare binary does not. But the mesh's package repository is hosted on the mesh, so any route that needs the mesh to install the thing that joins the mesh is a circle -- the tarball is the path that must never acquire a dependency. Reconciles on start, on a declaration, on a timer and on reconnect. The timer is the one easy to leave out, and without it `owned` reports what the host applied rather than what is there -- ADR 0035 violated by omission. The records checker caught this commit on its first attempt: 05 listed 0057 in its frontmatter while 0057 is still proposed, and a to-be document may not rest on an unaccepted record. The section now says so in the body instead. |
||
|
|
03874f3fe2 |
Add a structural check over HQ's own records
Nothing in this repository was verified by anything but reading, which is how a superseded decision stayed live in the constitution and in the to-be README at the same time. Both were found by a person looking, and nothing stopped a third. Five checks: links resolve; `decisions:`/`extends:` name records that exist and are accepted; a governing document citing a superseded record must name its replacement in the same paragraph; supersession is symmetric; filename number matches heading number. Each was made to fail before it was made to pass. The live-citation check was verified against a reconstruction of the actual incident -- the to-be README citing ADR 0017 as live guidance -- and reports it with file and line. It found one thing nobody had noticed: ADR 0018 never declared that it superseded 0011, though 0011 has named 0018 as its superseder since August. Fixed. Deliberately not checked, and said so in the README: 02-DECISIONS and 01-RESEARCH may cite superseded records freely, because a decision record discusses history and research records what was observed. 00-as-is may rest on one, per 0056. Flagging those would put noise on correct documents, and a check that cries wolf gets suppressed -- which costs more than not having it. Two bugs found by running it: the frontmatter reader iterated an inline list as characters, and the as-is exemption was missing entirely. |
||
|
|
e1f4c7d9e0 |
Approve 0054-0056, apply them, and fix the two smaller findings
0003 is now superseded by 0056. Nothing is left proposed. Applied: - 06 corrected from ten contexts to seven plus the api, each row now stating why it passes the more-than-one-node test. work, knowledge and stream are named as mesh-hosted rather than dropped; `ai` folds into config; `record` is deferred explicitly rather than listed. Its frontmatter now cites 0055. - how-we-build §4 amended per 0054, and the derived page republished by playbook 05. The sync found the drift the playbook exists to catch: the published §4 and the source did not say the same thing. The source said "four accidents, not four boundaries"; the published page said "one intent expressed four times", and only the published page carried the scope caveat. Same rule, two texts, already diverging. Verified the republish by reading back -- the new rule is present and the old section's body returns nothing -- rather than trusting the success message. The two smaller findings: - 0051 separated the transport identity from the declaring authority. It said the token carries "an address" and "the identity to expect" without saying what the node dials. It dials the broker, so pinning only that would make the control plane's authority transitive and let a compromised broker forge declarations -- which, since the host applies whatever the link delivers, is the whole machine. The token now carries four things, and declarations are signed and verified per declaration. Cost recorded: rotating the signing identity is fleet-wide. - 0026 no longer restates 0022's rule about generated views. 0022's own words are "prose does not restate status; one place, and two is one too many", which is what 0026 was doing to it. |
||
|
|
f49d177a31 |
Draft three records for the contradictions the review found
0054 -- things that change together share an authority, not a package. The constitution instructs agents to group "how a node is reachable" into one module, citing superseded ADR 0017; ADR 0044 says there is no networking thing to install. Since the constitution is injected where work is decided, the superseded rule is the one actually steering work. The observation behind it was right -- research 005 measured that reachability is the only place modules genuinely change together -- but the conclusion was wrong: tight coupling means a shared authority, not one artifact. wireguard and traefik deploy to different node sets, so the merged module would be assigned where half is unwanted. Requires amending how-we-build and republishing the derived page. 0055 -- the control plane is the node-coordinating contexts. Three context lists were in circulation (0015 says nine, 06 says ten, the README said eight) and none was decided. Research 006 said explicitly that the change "belongs in a new record -- not written here", and the design used the list anyway. Reconciling them shows `stream` and `ai` were dropped with no reasoning at all. Applying 06's own test -- needs to know about more than one node -- gives seven contexts plus the api, with work, knowledge and stream as hosted applications and `ai` folded into config as an ordinary grant. The record defers rather than lists. The cost is stated rather than reassured away: a board composing across the boundary reads more than one interface. That was raised before as "only moves the problem up a layer", and the answer is that 0045 already requires surfaces to read interfaces rather than stores -- what changes is the count, not the kind of work. 0056 -- the authority is the control plane, not a database. Every clause of 0003 has been decided against in four separate records and it is still accepted and cited as live. The error underneath is the same category error 0054 corrects: "source of truth" named a storage location when it meant an authority, and once the store is the answer, shared schemas follow. The half that was right -- the repository defines what exists, the mesh defines what runs where -- survives untouched. Best consequence: the cache mode disappears, so a node that has not heard from the mesh is no longer indistinguishable from one that has. 0003 is left accepted until 0056 is. |
||
|
|
ef5dd0751b |
Approve 0049-0053; drop a to-be item superseded by ADR 0044
The 'domain grouping' item cited ADR 0017 as live guidance. 0044 superseded it -- there is no domain module to group into, so there is no domain list to settle. |
||
|
|
ccbbfa9c8a |
One node runs the control plane, and nothing takes over
Closes the two open questions in 06 and 08, which turned out to be one question: how many control planes run, and what happens when the hub is down. Both were drifting toward redundancy by default -- a standby plane, a second hub, an election to pick between them. That is not one feature but a property every layer must then honour, and each layer gets it wrong independently. Not wanted, and not needed. A handful of machines with one node hosting the registry is not a distributed system. The argument for why this is sound rather than merely cheap is that the design already tolerates it by construction. ADR 0036 makes reachability state rather than class; the host reconciles from its own store (0043) and never needed to ask anybody to hold the state it was last given. So the control plane being down is not a new failure mode -- it is every node in the ordinary disconnected situation at once. What is lost is change, not operation. No node holds a contended role: the control plane is assigned like any other module, and the overlay hub is declared (0050). No promotion, no quorum, no fencing, no split brain, no replicated store, and no "which node is authoritative" recurring at every layer. Two consequences stated plainly rather than buried. The control-plane node is a single point of failure -- deliberate, and said out loud so it stays deliberate. And recovery is restore rather than failover, which makes backup the availability story rather than hygiene. The sharpest one is the clock: the control plane owns certificate issuance (0049), so an outage outlasting a renewal window expires every public name. That bounds how long recovery may take, and nothing measures it today. |
||
|
|
4e80820e2f |
Design connectivity in full: overlay, resolution, exposure, filtering, certificates
Written as one document because the five are one design. They share inputs, they must agree, and every one of them today is computed in a different place by a different module from a different copy of the same facts. The through-line is that none of the five can be answered by a machine alone, so all five are decided centrally and delivered as `file` resources. That costs no new host vocabulary and removes both remaining direct database connections from nodes -- wireguard and traefik are the only two, and both are connectivity. Three decisions fall out, all proposed: 0050 -- reachability is declared, not inferred from an address. The RFC1918 regex is wrong for carrier-grade NAT (100.64/10 tests as public, so an endpoint is written to an address nothing can reach), wrong for IPv6, and wrong for a routable address behind a closed firewall. The lab needing TEST-NET-3 to satisfy the regex is the same bug from the other side. Also kills hub election by address prefix, which fails silently and makes renumbering an outage. 0051 -- the enrolment token carries where the mesh is and how to recognise it. Closes two circles with one mechanism: verifying the mesh needed the CA, and obtaining the CA meant trusting whoever handed it over; and a node had to reach the mesh before it could resolve any mesh name. An address plus a fingerprint, carried out of band, resolves both -- and closes the CA question 0049 deferred. 0052 -- a filter rule names its source. `scope:` is declared in five manifests, is part of no rule type, and is referenced by no code, so those manifests appear to restrict ports and restrict nothing. Removed rather than implemented; the general fix is refusing unknown keys, which the host already does and manifests do not. Also corrects two claims in 0049 asserting wireguard was already handled. Research 006 says both modules still reach upward; neither is. |
||
|
|
8d9282d86b |
Resolve the ingress gap: a route is a grant
ADR 0048 named ingress as an unclosed hole -- nothing said what terminates TLS, how a public name reaches a container, or which tier owned it. Resolving it needed no new concepts, which is why it survived: nobody had applied the rules already written to it. Ingress is not substrate. The control plane does not need a route to start, and no node needs one to reach it -- the node dials out and has no listening control surface. It grants itself a route afterwards, like a bucket. A route is an instantiation edge under ADR 0044. The direction mirrors a database -- the consumer supplies a target and receives a name rather than credentials -- but it is the same edge. The substantive finding is that exposure is three facts at two scopes: name resolution and certificate issuance need to know which node is publicly reachable, and only the proxy mapping is a single machine's business. That is why it belongs to the connectivity context, and why Traefik doing all three on the node is wrong. Which matters beyond tidiness: research 006 counted traefik as one of two modules opening a direct Postgres connection, reading nodes and mesh_ca. That violates 0037, 0045 and 0039 at once, and is why every node permanently holds a credential to the control plane's database. Deriving the config centrally and delivering it as `file` resources removes it, costs zero new host vocabulary, and closes the set 0039 identified -- wireguard was the other. Left open deliberately: the mesh's internal CA is the other thing traefik reads, and it belongs to the link's mutual authority, not to exposure. Conflating the two is what made the gap hard to see. Also fixes an inconsistency from the previous commit: 06 still claimed the virtual host was raised from the bundle. Proposed, not accepted -- for review. |
||
|
|
4d19e93900 |
Name the substrate's actual products
The design layer described every service by role and never once by name: Postgres appeared in zero design documents. That was over-application of the research rule "never identify the mesh it observed", which is about node names and domains, not software. Two things were actually broken by it. substrate.lock pins images by digest and a digest belongs to a named image, so the bundle could not be written from the design. And a reader could not tell a settled choice from an unexamined one -- "a relational store" reads identically either way. ADR 0048 names them: PostgreSQL, LavinMQ, MinIO, an OCI registry, Docker. The argument for each is continuity, which is a real argument -- replacing a substrate service migrates the mesh's own state. Role and product are now both written, because the design depends on the protocol while the installer needs the product. Also separates two questions the substrate doc had merged: being substrate and being in the bundle. Only Postgres must precede the control plane; the rest are substrate by role and ordinary by delivery. Whether the bus joins it is left open, because it turns on the control plane's internal shape. Names the forge as Gitea, and records ingress/Traefik as an unclosed gap rather than a naming one -- nothing says what terminates TLS or which tier owns it. Fixes a miscount: the host's bootstrap vocabulary is six shapes, not five. |
||
|
|
93470f6162 |
ADR 0047 — the bundle may carry actions the link may not
The bootstrap's sharpest open question, and the framing was wrong. "State on this machine" was being read as the filesystem and the service manager. A service running on this machine IS part of this machine — writing a file and creating a database in a local store differ in mechanism, not in scope. The real question was underneath: must the host learn what a database is? It must not. Giving it a `database` resource type means tier 0 knows Postgres, then a bucket, then a virtual host — the host acquiring the substrate's vocabulary one service at a time, which is what ADR 0037 exists to stop. So the bundle declares an ACTION and the host runs it and verifies it. What a database means stays with the module that provides one; the host knows only how to run a declared action against something local and check the result. Its vocabulary grows by one shape rather than by one resource type per service. Actions are permitted in the bundle and forbidden over the link, and the asymmetry is deliberate. A bundle arrives WITH the binary: anyone able to put a hostile action in it could equally have put it in the host itself, so refusing actions there buys nothing and costs the bootstrap. The link is a separate party, reachable separately, and an action there is the unbounded blast radius ADR 0039 refuses. That decision stands unchanged. And ongoing provisioning is not the host's at all — the control plane does it once a mesh exists — so the asymmetry costs nothing. Which dissolves the earlier worry about one mechanism with a tier boundary inside it: there are two mechanisms, with different actors, scopes and trust models, and that is the answer rather than a compromise. Named rather than hidden: this is the escape hatch research 011 warned about, arbitrary code in the place hardest to remove later. It is bounded by being bundle-only and by every action having to declare how it verifies itself, and that boundary is the whole defence. |
||
|
|
5b3d0ebd4f |
ADR 0046 — the installer fetches what it pins
The blocking question was where a container image comes from, and the version that blocked assumed the machine might have no network. That assumption came from the LAB: a scenario is a closed address space by design, which is what lets two scenarios hold the same addresses without meeting. Production is not sealed — a machine being adopted has a network, and one that does not is a machine where very little works anyway. So substrate.lock carries references, not payload: an image name and a digest, fetched at apply time. A first node pulls from upstream because no mesh registry exists yet; every node after that pulls from the mesh's own. The lab is the exception and places images itself, the way it already places the host binary — a property of a test environment, and letting it dictate the production design would be the tail wagging the dog. Pinned by DIGEST rather than tag. Reproducibility comes from pinning the identity of a thing, not from carrying its bytes, which is what makes fetching acceptable rather than a compromise. ADR 0041 survives untouched, which was the point. "Copy it onto a machine and run it" stays literally true — one binary, a few megabytes, which then fetches what it was told to. Carrying images would have quietly redefined the property that decision rests on. Costs accepted and named: an apply can now fail because something is unreachable, which a self-contained artifact could not, so it must fail legibly — naming what it could not fetch and from where. And the lab needs a way to place images into a machine that also has no container runtime, both of which are lab-installation concerns and neither solved here. Research 012's build-time-versus-apply-time reframing narrows accordingly: it still holds for what a tailored installer contains, and no longer has to hold for images. |
||
|
|
0531d6fc38 |
ADRs 0044 and 0045 — the module design, closed; 011 graduates
011 opened asking what a graph deletes and found the graph already existed. The work became design, worked through twenty cases and one provider in full. Two decisions close it. 0044 — a module declares presence, instantiation and exclusion. Two kinds of edge because a game wanting a database is not a game wanting postgres to exist: one creates something per consumer, carries credentials back, can be revoked, and leaves the provider holding state. Names are concrete unless providers are genuinely substitutable — `terminal` passes, `database` fails, and the adapter is what creates an interface. Where there is no contract there is a tag, which describes and does not bind. Exclusion is a third relation and is not derivable. A node provides names too, which makes capability checking stop being a separate mechanism and makes the host's detection an input to resolution. Constraints, never placement. Scope decides which provider and the binding is written down and sticky, in a place that follows the scope. And there are three entities, not two — the assignment carries what belongs to neither end, which is what node-agnostic modules ran out of. 0045 — a context owns its store, exclusively. No shared writes and no read roles on another context's store, because reading couples you to its layout just as firmly and invisibly. The unit is the CONTEXT, not the process: a board showing the mesh's own data is the mesh showing its own data. Asking or subscribing is derived from ADR 0036 rather than chosen. And it is the first clear list of what the design removes: grant kinds, table ownership, cross-context migration ordering, and a class of permission modelling. 0017 is superseded rather than narrowed — its text unchanged, its status changed. Folders assert relationships where edges record them, and the domain module goes with it. Left explicitly undecided in both: what a resolver delegates rather than reimplements, how many instances a module should have, and what a provider hands back. |
||
|
|
00d5ba8376 |
0042 and 0043, as approved
0042 becomes the operator's own rule and stops there: every merge into the main branch is notified and approved. Notified means proposed and said out loud, not performed and mentioned; approved means a person says yes to THAT merge. Who performs it is not the thing worth constraining, which is what makes an agent merging its own work unremarkable — the checkpoint already happened. 0043 answers the question it did not: where the ordered list comes from. By hand today, in substrate.lock, because the first node has no control plane to derive anything from. Afterwards the control plane derives it from module assignments, resolved configuration, and what each module declares it needs — ordered by the dependency graph, which is research 011. So the record is complete on the consumer side and deliberately silent on the producer side, and that is a legitimate order to settle them in: the host must refuse what it does not understand whoever wrote it. One consequence that only appeared when the question was asked: if the graph turns out not to determine a total order, that is 011's problem and not the host's. The host is still handed a list and still applies it as given. Recorded because it is the seam where a future difficulty would otherwise try to migrate into tier 0. Both were marked accepted before they had been read. Approved now, so the field is true — which it was not when it was written. |
||
|
|
bea052753e |
ADR 0043 — what a declaration is
Stage 2 could not start without it. Three constraints already bound the shape and between them they decide most of it. JSON, because the standard library carries it and carries no YAML, and a YAML declaration would put a third-party parser inside the one binary whose whole argument is that it needs nothing — to gain authoring comfort in a document generated by a machine and read by a machine. An ordered list, because ordering is a DECISION. A host deriving order from declared dependencies would be deciding the thing most likely to differ between what the control plane intended and what the machine does. The control plane knows what depends on what; it says so by saying when. Unknown is refused, never skipped — an unknown version, type or field refuses the whole declaration. A host that skipped what it did not understand would apply most of a declaration and report success, which is 04-ISSUES/003 with the declaration on the other side of the wire. Complete for what the host OWNS, and only that. It removes what it previously applied and is no longer declared, which it knows from the store rather than by inference, and never removes what it did not create — a converger that treats 'not declared' as 'must not exist' deletes what the mesh never put there. Two consequences arriving earlier than the build order suggested: the store is load-bearing at stage 2, because nothing can be removed without knowing what was applied. And a closed address space bounds the first vocabulary to what needs no network, because a scenario has no route to a package repository. |
||
|
|
23232e019a |
ADR 0042 — the approval is the checkpoint, not the second pair of hands
§2 said "never merge your own", written for people. Applied to an agent it produced a contradiction that surfaced immediately: an agent asked to merge cannot merge, because it authored what it is being asked to merge. So every merge here was either performed by the thing that wrote it, or not performed. Rejected the literal reading, because an operator clicking merge dozens of times without reading is not a checkpoint — it is the SHAPE of one, which is worse, since the record then claims a review that did not happen. Rejected dropping the rule, because the failure it prevents is not one an agent is less prone to. So the rule names what the checkpoint actually is: a person deciding, not a person clicking. Work may be merged by whoever wrote it once a human has explicitly approved that merge. What "explicit" excludes is the half that can rot, so it is enumerated: a standing permission cited forever, an instruction to do the work read as approval to merge it, silence, and the author's own judgement that it is ready. This narrows rather than relaxes. The obligation moves from who performs the merge to whether a person decided — a higher bar in the case the old wording permits, where a reviewer merges someone else's work without reading it. Synced, and verified by reading the rule back out of the live page rather than by trusting the publish. |
||
|
|
92e8c74ce4 |
ADR 0041 and the build handoff for the node host
Building tier 0 forced the question "the one binary installed by hand" had been carrying unexamined. A TypeScript host needs a runtime present before it runs, so the thing installed by hand becomes two — and the second must be installed by the means the host exists to replace. So the host is a statically linked binary that requires nothing present, written in Go. Rejected: a runtime installed first, which breaks the property the tier rests on; and bundling the runtime into the executable, which carries ninety megabytes to preserve a language choice and puts a young feature at the bottom of the stack. The argument that decided it is architectural rather than about taste. 0037 means the host never queries the mesh database and 0039 means it only receives declarations, so the host shares NO code with any other tier — not a client, not a schema, not the SDK. The language boundary falls exactly on a boundary that already exists, and a second language usually costs duplicated logic where here there is none to duplicate. §8 gains a scope: it said "TypeScript throughout" when everything was a service or a surface, and is now scoped to those with tier 0 named. Another sync owed. Playbook 04 steps 2 and 4: repos.md records mesh-host as existing, the design takes code: [mesh-host] and status: in-progress. |
||
|
|
6e4fc5d69b |
ADR 0040 and the constitution sync — absorb, then publish
The sync came due for three accepted rules. Reading the target before overwriting it found the source and the enforced copy had diverged unrecorded, and that a literal republish would have DELETED rules the mesh enforces: the live page's §4 carried SOLID, layering, TypeScript strictness and DRY/YAGNI, which appear in no decision record anywhere and have been checked against for six weeks. Removal was not the safe alternative either. The orchestrator reads "when absent, no constitution is injected (backward-compatible)" — so deleting the page would not fail, it would silently inject nothing, and every design meeting would run unchecked. Three unenforced rules would have become all of them. So: absorb first. The code-quality rules land as §8 rather than §4, because appending renumbers nothing and every existing citation stays valid. They are marked as inherited — every other rule states the incident behind it, these state nothing because nothing was written down, and importing them silently would have claimed a provenance the document does not have. The review bar is resolved to a person who is not the proposer. The live page required two node operators; there is one, so the rule was never met and could not be — a rule that cannot be satisfied is not a high standard, it is one everything silently violates. Then published, and the read-back earned its place in the playbook: the FIRST publish reported success and changed nothing. New revision, new title, body unapplied — a malformed argument dropped silently. §5 demonstrating itself during its own publication. |
||
|
|
34e7a4780a |
Accept 0035 — a report is read from the system
The principle is obvious; the obligations it imposes are not, and those are what was actually being decided. The applier records what it did, including facts it never reads itself, purely so something else can read them back. It records them AFTER the thing works, so a half-finished run ends up with less metadata rather than optimistic metadata — rejecting the simpler alternative of tagging at creation with a status field, because a failed raise deliberately leaves wreckage standing and wreckage tagged at creation claims things that never happened. And it binds anything that reports on the mesh, not only the drawing. §5 carries the rule unqualified now. The constitution sync is owed for this and for 0034 and 0018, and has not been run. |
||
|
|
66a89cefbe |
Accept 0039 — a node owns no password, only an identity
Settled in the operator's own words: nodes should not own passwords, only an identity when communicating to the broker. Recorded that way at the top of the record, because it is the whole decision in one line and the rest is why. What this obliges, in order of newness: enrolment is the one mechanism that does not exist. Per-node broker users, virtual hosts and per-queue permissions are broker configuration. Mutual authority is certificates on a connection already open. And the boundary must fail legibly, which is the requirement the debugging objection earned. |
||
|
|
9ee0c63c7a |
0039: the link already exists, and most of the cost is already paid
Written first as though the link were a thing to build. It is not. ADR 0001 already has it — every node connects outbound to a single broker, nothing ever connects to a node, each node declaring an exchange named for itself and consuming from its own queue. Already outbound-only, already per-node addressed, already the one channel everything arrives through. So this record is not proposing a channel. It proposes that the channel carry per-node identity instead of one shared credential. The same as-is records the fault it fixes, for the broker rather than the database: the broker's credential is mesh-wide, rotating it is a mesh-wide operation, and doing it wrong has taken the broker down. That reframes the overhead objection, which was fair against what the record said and not against what it means. Of six properties, four are already true, one is broker configuration — users, vhosts and per-queue permissions the broker already implements — and exactly one is new machinery: enrolment. Meanwhile 0037 subtracts, since a node under it holds no database credential at all. Three hand-carried shared secrets become one identity that grants only identity. Adds the option that was actually being weighed and was missing: accept the exposure as the cost of simplicity. Rejected because the simplicity IS the unfixability — the credential cannot be rotated precisely because everything holds the same one. And adds a requirement from the debugging objection, which was the strongest part of it: it must fail legibly. A boundary that refuses a node without saying why is worse than the password it replaced, because a wrong password at least announces itself. That is §5 applied to a security mechanism. |
||
|
|
f5073b9b00 |
Accept 0034 and 0018; 0011 is superseded
0034: a test defends a decision. §5 carried it marked "proposed, pending review"; the marker is removed and the rule now stands unqualified. The lab was already built to it, which is the inversion §5 exists to catch — closed now rather than left standing. 0018: the mesh creates no symlinks. §3 said the installer owns the links today and the INTENT was that the mesh creates none. It is no longer an intent, so the wording says so, and ADR 0011 becomes superseded rather than edited — its reasoning is why the rule exists at all, and the incident behind it is the reason anyone believes either record. The links the installer still reconciles are a migration, not a permission. The constitution sync (§6 step 4, playbook 05) is NOT done. An unsynced rule is a rule the mesh does not enforce, whatever this document says — and publishing it changes what every design meeting is checked against, so it wants saying out loud rather than doing quietly. |
||
|
|
4866007c04 |
ADR 0038 accepted; ADR 0039 (proposed) — the link is the security boundary
0038 accepted: no two modes. One behaviour, two sources of declaration. 0039 fills the gap 0038 named. Proposed rather than measured — it designs a boundary that does not exist — but what it replaces IS measured, and that is the argument. Today adopt.sh asks the operator to paste in the postgres password and the object store password, the same ones on every node, and they are not discarded after adoption: wireguard and traefik open a pg connection on every reconcile. So every node permanently holds a credential to the control plane's database, and 00-as-is/06 records that nothing rotates it. Compromise of any node is compromise of the mesh's store, with no way back. Four properties make the link a boundary rather than a pipe: it is outbound and node-initiated, so a node has no listening control surface — which the topology already requires, since most nodes have no forwarded port. A node holds its own identity and nothing else, so compromise of a node is compromise of that node. Authority is mutual, because a host that applies whatever the link delivers must know the mesh from something impersonating it. And what may be pushed is bounded by FORM — declarations of known shape, never a command to run. That last property is stated with its limit rather than oversold: it bounds form, not impact. A compromised control plane can declare harmful state and the host will apply it faithfully. What it buys is a describable blast radius. Joining uses a one-time short-lived enrolment token, useless once used and useless after a while, in place of hand-carried shared secrets. Named rather than hidden: the first node's identity is self-issued and becomes the root of trust, which is the one place 0038's "no special first node" does not fully hold. Rotation becomes possible and is still not designed. And what may expire is constrained by 0036 — an identity needing refresh would make a laptop fail for being a laptop. |
||
|
|
72b22830f3 |
ADRs 0036, 0037, 0038 — what a node is, what the host does, how one joins
0036 (accepted): a node is a managed machine, and disconnection is a situation. The open question posed a class distinction — full nodes and lesser presences. There is none. Reachability is state, not kind, which promotes the host's local store from a component to a requirement: it is what makes disconnection ordinary rather than exceptional. The reduced contract the question reached for is real but it is capability, and that belongs in the profile. 0037 (accepted): the host applies, it does not decide. Measured rather than argued — the absorption is smaller than the machinery that already applies state, and eight of ten adapters carry no dependency to move. The two that do open a Postgres connection to the control plane, which inside tier 0 is the one thing the tier rule exists to forbid. So each concern splits: deciding needs every other node and stays in tier 2; applying needs root and locality and goes to tier 0. The host carries ONE concern, of which the six are instances. 0038 (proposed): a node joins by linking first. The operator's two-modes proposal, adopted as intent and corrected as structure. Two modes is two code paths where the first runs once per mesh and rots — and the mesh already has that fault in its worst form, as three hand-run shell scripts. Instead: one behaviour, two sources of declaration. The first node is not a different kind of node, it is a node whose mesh is not up yet, and its specialness is temporary and self-erasing. 0038 also shrinks the migration 0037 called expensive: a joining node never needs mesh-wide state, because the hard part of the overlay is only needed to compute the WHOLE mesh. It needs one peer. The rest arrives. Left open and said so: what may be pushed over the link and how a joining node proves it is entitled to join, and whether one host can raise the substrate alone. |
||
|
|
b225b07625 |
ADR 0035 (proposed): a picture is read from what runs
Drawing a scenario forced a choice that looks cosmetic and is not. A diagram built from the declaration and captioned "as raised" answers "is what is running what I asked for?" with the request, which always agrees with itself. So: a picture captioned as raised reads only the running system, and where the hypervisor does not hold a fact the picture needs, the raise records it on the resource. With the rule that makes the recording worth anything — a behavioural tag is written after the behaviour works, never at creation, because a failed raise leaves wreckage standing and a picture of wreckage must not badge what the wreckage was supposed to be. It earned itself on the first comparison: every VM showed no addresses, because a container's interface carries the device's name and a VM names its own. The two pictures disagreed, so a whole class of machine silently losing its addresses was visible in seconds. §5 of how-we-build gains the general form, marked proposed. The constitution sync is deliberately NOT done — a rule the mesh enforces before a second person agreed to it is what §6 exists to prevent. |
||
|
|
a2495e4d8e |
ADR 0034 (proposed): a test defends a decision
how-we-build 5 already says that if a document states a rule about the mesh, it says how the rule is verified — an unenforced rule being indistinguishable from a wrong one, and costing more because people believe it. That has never been applied to decisions, and a decision record states the same kind of claim. The gap was found by review: the lab reached 2,128 lines with 1,072 untested and no stated rule broken, because there is no testing posture in how-we-build at all. Every decision the lab embodies was verified by hand and none of it survives the terminal it ran in — which is 04-ISSUES/005 in miniature, coverage assumed rather than checked. Rejected a coverage percentage: it measures how much code a test touched, not whether anything important is defended, and would have been satisfied by testing the parser harder while the hypervisor integration stayed unasserted. Rejected test-driven development as a hard rule, and not because it is wrong in general. Half this implementation was discovery — that the hypervisor CLI reads a definition from stdin and hangs, that it assigns a MAC without recording it, that a stock image's networking flushes a static address. A test written first against undiscovered behaviour asserts a guess. So: structure and logic tested first, behaviour against a real system tested alongside, mocking the boundary forbidden, and a blocking gate as the definition of done. A test names the decision it defends, which is what makes the pairing checkable — a decision without one can be found rather than noticed. Stated as proposed rather than adopted: 6 requires review by someone who is not the proposer. Records 0001-0033 predate it and are not retroactively invalid, but each should acquire a test or an explicit note that it cannot have one, and until then the rule is aspirational for them — which is the state 5 warns about, recorded rather than hidden. |
||
|
|
eab4598494 |
ADR 0033: a router is scenery, not a node
ADR 0016 makes a lab node a virtual machine, and its reasoning is fidelity: a node boots a stock image and runs the real install, so it has to be a real machine or the thing under test is not the thing that ships. That reasoning does not reach a router. Nothing under test runs on one, it holds no identity, the mesh never installs anything on it, and no assertion is ever made about its internals. It exists so packets behave the way they behave in the world, which is the definition of scenery. So a router is a system container. What it must reproduce is kernel behaviour — translation, connection tracking, filtering, forwarding — and a container has the same kernel. Verified before deciding rather than assumed. In a plain unprivileged container: ip_forward and ipv6 forwarding both settable, nftables masquerade accepted and listed back, and the conntrack timeouts that mapping_ttl depends on both writable. No privileged mode, no nesting, no capability grants. Rejected letting the hypervisor provide NAT, on a stronger ground than speed: it makes the lab provide what the declaration is supposed to own, and it cannot express a mapping that expires, a gateway that refuses to forward, or policy between siblings. The model would shrink to fit the tool. The distinction is now load-bearing and has to stay legible: node means something under test, scenery means something that makes the test real. If the mesh ever installs anything on a router, it has become a node and this record no longer covers it. |
||
|
|
a253afe020 |
Scenario lifecycle, and how two scenarios coexist
ADR 0032: a scenario is a closed address space. Every segment materialises as its own isolated link belonging to one instance, so two scenarios raised from the same declaration hold the same addresses and never meet. The declaration keeps its literal addresses and they mean what they say — allocating from a pool would have made them a fiction, so a scenario reproducing a specific topology would stop reproducing it. The constraint that follows shapes everything: the lab never reaches into a scenario over IP. It talks to machines through the virtualisation layer's own channel. If it reached them by address, the workstation would need a route into each scenario, and two carrying the same prefix would give it two routes to one destination — failing not with an error but by one scenario's traffic arriving in another. That also makes reachability an honest question. Can this machine reach that one is asked from INSIDE, by executing on the first, rather than probed from a workstation that is not on the network and whose opinion would be a different question with a misleadingly similar answer. The lifecycle itself: six verbs, of which raise and destroy are enough to be useful and the rest are what make repetition cheap. Raising is convergent rather than incremental, because a lab behaving differently from the thing it tests teaches the wrong habit. A failed raise leaves the wreckage standing. Tearing down on failure destroys the only evidence, which is backwards — a scenario that failed to raise is more interesting than one that succeeded. Snapshots are whole-scenario. Per-machine would be cheaper and wrong: the mesh keeps state spanning nodes, so restoring one machine while its peers move on produces a mesh that has never existed, and faults found there would be artefacts of the lab. Closes the declaration's open question about running several scenarios at once. |
||
|
|
a72fea5342 |
ADR 0031 and the scenario declaration
The lab provides the underlay; the mesh builds the overlay. This is the boundary that decides whether the lab is worth having: a scenario that assigns overlay addresses, elects the hub and writes peer configuration certifies its own work — if the mesh's peering is broken, that scenario still comes up green. The most valuable thing the lab can test is exactly the part pre-building would replace. So a scenario declares what a hosting provider and a home router would provide: segments, which machine sits where at which address, what NAT is between them, which ports are forwarded, which machines are detached. It declares nothing about overlay addresses, hubs, peering, names or certificates, all of which become outcomes to observe. The declaration has four parts — segments, machines, place, snapshot — and the two scenario classes differ only in place. That is what makes one a strict subset of the other rather than a fork. Research 004's most important finding becomes a format constraint rather than a footnote: the routable segment must use RFC 5737 documentation space, because the mesh decides public versus private by matching the address, and a private range there makes the hub test as unreachable while the mesh silently never forms. A segment without behind: is routable, and a non-documentation address in it should be refused before anything is raised — ADR 0008 applied to a configuration file, since the failure it prevents has no error at all. Four things left open, including the one that matters most: a lab machine is always privileged, so the user and edge profiles have no scenario that exercises them. |
||
|
|
09489a298c |
ADR 0030: the repository structure, and the rule that names them
The tiers were settled and the product was named, but the repositories themselves existed only in a research sketch. That had already caused two problems. ADR 0029 makes the lab phase 0 of the migration and could not say where it lives, because no record named a repository. And the sketch contradicted an accepted record: it listed mesh-hq while ADR 0028 had decided novox/hq and explicitly rejected that name. A design resting on research is resting on something that can change without a decision. Corrected in the research too. The naming rule, which both earlier records implied and neither stated: a repository belonging to a product carries that product's prefix; a company-scoped one does not. That is why this repository is hq and the mesh's are mesh-*. Seven repositories recorded — host, substrate, control, surfaces, sdk, lab, and this one. The lab gets its own: it ships to nobody, outlives any single tier, and drives virtualisation on a workstation, which nothing else does. Inside the host it would couple development tooling to a shipped component; inside the control plane the bootstrap scenario would depend on a tier that does not exist when it is needed. Tier 4 is deliberately not decided. Whether the catalogue is one repository, one per domain or one per application stays open from ADR 0015 and is blocked on research 005 — how many repositories hold domains cannot be answered before knowing what the domains are. mesh-catalog appears in the sketch and is not decided by this record. The cost is stated rather than glossed: seven release cadences where there is one, and cross-repository changes that used to be one commit. |
||
|
|
b4904fec7e |
The lab comes first, and its first scenario has no pipeline
The lab was designed around a module under test, with a scenario being a complete mesh — forge, coordinator, cascade, verify. That is unusable for building the new mesh, because all four are tier 2 and do not exist yet. And research 009 had the sequence backwards. It placed the lab at phase B as verification of tiers already built, but tier 0 is the component that takes over a machine's packages, services and network. It cannot be developed against a machine anyone needs. The lab has to exist before the thing it will test. ADR 0029 splits scenarios into two classes. The bootstrap scenario is virtual machines, the host binary and a pinned bundle, with the verdict coming from what the host reports about the state it reconciled. The full scenario is the designed one. The first is a strict subset of the second — same virtualisation, same networking, same lifecycle, stopping before a control plane exists — so the second is reached by addition rather than rework. The consequence worth having: raising a node from nothing stops being the least-exercised path in the system and becomes the inner development loop. It also settles the runner's two jobs. Scenario lifecycle is needed immediately, because something must materialise and reset a mesh before anything can be written against it. Assertion execution waits for the full scenario. Corrects a stale claim in the design while amending it: it argued scenarios were affordable with system containers and would not be with virtual machines. ADR 0016 superseded that reasoning and the text had not followed. Issue 007: the lab's first requirement is installed and unusable. The virtualisation package is present and explicitly installed; both units are disabled, the operator is in no group, and the client reports the server unreachable. Not issue 001 again — that is an install failing while reporting success. This is an install succeeding when success was not the point. A package is files; a capability is a running service and an identity permitted to reach it, and the module model has no vocabulary for the second. |
||
|
|
c0ae8dec96 |
Remove two disclosures, and record Nox as the answer to 006
Found by a full scan before making the repository public, which is the moment the public rule stops being aspirational. A module name identified a specific laptop model — hardware inventory, which is operational detail about one installation rather than a lesson that travels. Generalised. ADR 0028 named a forge username in a repository path, which the public rule forbids, and the sentence had also gone stale: the repository it described was subsequently verified empty of anything unique and removed. Rewritten to state what happened without the username. Removing a disclosure from a record is the same class as fixing a path — the rule that permits it outranks the one that forbids editing. Issue 006 gains its proposed direction: Nox works from within this repository rather than these documents being synced into the knowledge base. Better on three counts — no copy, so no drift; no fourth knowledge system, which was the original objection; always current. But it changes the promise, and the issue says so. ADR 0019 promised these documents would surface BESIDE everything else in a symptom search. An agent that must be asked is reachable, not surfacing, and the two differ in precisely the case the operational memory exists for — someone debugging an error with no reason to suspect HQ knows anything about it. The question narrows to whether a symptom search finds this content without the searcher already suspecting it. |
||
|
|
93a1231e00 |
Retire the HAL name where it points forward
Skills take the hq- prefix: they are HQ process workflows, not mesh workflows, and HQ is company-scoped now. hq-new-research, hq-graduate, hq-new-issue, hq-diagnose, hq-amend-design, hq-handoff, hq-sync-constitution, hq-status. Forward-looking prose becomes Novox Mesh or simply the mesh — the root README, AGENTS.md, the 00-META README, the mission's module example, and one to-be document that addressed 'someone working on HAL'. Three categories deliberately keep HAL, per ADR 0027: The monorepo is still called hal on the forge. repos.md, every code: field and every located-in: field name a repository that exists under that name, and renaming them in prose would make them false. The as-is layer and the research that measured it describe the system that runs, and that system is called HAL. 124 modules, 9 daemons, a dead containerised node — those are observations, not intentions. Records 0001-0026 are immutable. A record says what was decided when it was decided, and no record is edited for a name. Also repoints ADR 0022's link at the renamed skill — a path fix, which the immutability rule permits, not a change of meaning. |
||
|
|
87f4f29cc6 |
Novox Mesh, Nox, and HQ becomes company-scoped
ADR 0027 — the product is Novox Mesh, shortened to mesh internally. HAL was never chosen: it arrived with the dotfiles repository this grew out of, it is borrowed, and it is borrowed from the canonical untrustworthy machine intelligence, which is an odd flag for infrastructure trusted with credentials. Timing is the substance of the decision, not an aside — the skeleton is not built, so renaming costs a search and replace now and a migration later. Nox is an identity of Novox, and specifically the agent of the MESH rather than of a node. Nodes keep their own identities. Nox addresses them, and a human mostly talks to Nox — which makes it the concrete form of the mission's vision: state an intent, and the mesh works out which node holds the thing. It holds no private channel. The gap this opens is recorded: ADR 0012 binds every agent to a home node, and a mesh-scoped agent has none, so the model needs extending. ADR 0028 — HQ is company-scoped, novox/hq, with the mesh as its first product. Checked rather than assumed: the company organisation already holds live projects that the mesh builds and deploys, so they are tenants rather than peers, and the mesh is the ground they stand on. There is also company work outside the mesh already, which strengthens the case and means the eventual split is closer than "some day" — so each document's scope is fixed now, in a table, making that split mechanical instead of archaeological. The folders are deliberately not restructured yet. The skeleton takes the new vocabulary: mesh-host, mesh-substrate, mesh-control, mesh-surfaces, mesh-catalog. Substrate drops to four services now that identity is a hosted workload rather than a dependency. Research 009 opens the migration, with the reframing that lowers its risk: replace the control plane, do not move the workloads. Their data never moves, so it is re-declared rather than adopted — which keeps adoption out of scope, as the lab design requires. Self-hosting is the last phase, or a failed cutover takes away the means to fix it. |
||
|
|
143f8ab2f1 |
research 005: which modules actually change together
The domain-grouping premise is testable, so it was tested before drawing a list. Co-change across the full history of the module catalogue, current modules only, platform namespace excluded. Nine commits in ten touch exactly one module, and 50 of 89 modules have never been edited alongside anything. Two clusters exist above that floor. Reachability holds up: proxy, resolver, firewall and VPN genuinely move together under one intent, three times in recent history. That is the shape ADR 0017 describes and the only place the measurement finds it. The provider cluster does not, and this is the finding worth having. Every multi-provider commit is a cross-cutting manifest change applied N times — feature detection, hook conventions, volume binds, network scoping. None is a change to what a database is. Merging them would not have prevented one of those commits, and the history already shows the fix that worked: verify by shape in the SDK rather than copying a script into every module. Move the concern into the machinery, do not merge the modules carrying it. ADR 0017 keeps its principle and gains a pointer to this narrowing. |
||
|
|
3f6d939930 |
Every decision is a record; the ledger is gone
papa-hq has no ledger. Its root is AGENTS.md, CLAUDE.md, README.md, every decision is a numbered record, and its graduation playbook has no path for an unrecorded decision. hal-hq now matches. The ledger's 41 entries classified as: 10 restating a record, 11 restating design docs, 15 describing how this repository works with the reasoning sitting in a README rather than anywhere citable, 3 small rules with no home, 2 superseded stubs. Mostly a copy — and a hand-maintained index, the exact pattern ADR 0022 had just rejected for the decision index on the grounds it drifted after one addition. Keeping one copy of that while removing another is not a position. It also collided by name with 02-DECISIONS/ in any directory listing. Nothing was dropped. Records 0019-0025 give the repository decisions the reasoning they never had: HQ is its own repository and is public, design has two layers, work moves through playbooks, status lives in frontmatter, issues have a front door, the numbering is the flow, HQ is the source of the constitution. 0026 records the ledger's own removal. The three orphan rules went to how-we-build, where a rule is enforced and keeps the incident that earned it — the package rule was genuinely unwritten anywhere. Two lab decisions stated only in the ledger went into the lab design. "Deliberately not decided" went to the research effort and design document each question actually belongs to. The chronological view the ledger provided is now generated from record frontmatter, which is what it was for. The cost, stated in 0026 rather than glossed: a record is more work than a table row, so the risk is a small decision going unrecorded because nobody wanted to write a document. how-we-build takes rules cheaply, which is the mitigation, not a solution. |
||
|
|
c0b35652d0 |
The numbering is the flow: decisions are 02, design is 03
papa-hq reads 01 research -> 03 decision -> 02 design. The order is a scar, not a choice: 02-DESIGN existed from its initial commit, and when adr/ was finally promoted on 2026-07-13 it took the next free number rather than its place in the sequence. By then design was too settled to renumber. hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and 02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks the process in the order it happens: research produces a decision, the decision authorises a design. 00-GENESIS becomes 00-META, matching papa's rename from the same restructure. Every path reference rewritten across documents, frontmatter, playbooks and skills. All links resolve; all 58 frontmatter blocks parse and their path fields still point at files that exist. |