6abfec74336718816703985417cb9ee1be676b71
534
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
eab4598494 |
ADR 0033: a router is scenery, not a node
ADR 0016 makes a lab node a virtual machine, and its reasoning is fidelity: a node boots a stock image and runs the real install, so it has to be a real machine or the thing under test is not the thing that ships. That reasoning does not reach a router. Nothing under test runs on one, it holds no identity, the mesh never installs anything on it, and no assertion is ever made about its internals. It exists so packets behave the way they behave in the world, which is the definition of scenery. So a router is a system container. What it must reproduce is kernel behaviour — translation, connection tracking, filtering, forwarding — and a container has the same kernel. Verified before deciding rather than assumed. In a plain unprivileged container: ip_forward and ipv6 forwarding both settable, nftables masquerade accepted and listed back, and the conntrack timeouts that mapping_ttl depends on both writable. No privileged mode, no nesting, no capability grants. Rejected letting the hypervisor provide NAT, on a stronger ground than speed: it makes the lab provide what the declaration is supposed to own, and it cannot express a mapping that expires, a gateway that refuses to forward, or policy between siblings. The model would shrink to fit the tool. The distinction is now load-bearing and has to stay legible: node means something under test, scenery means something that makes the test real. If the mesh ever installs anything on a router, it has become a node and this record no longer covers it. |
||
|
|
132f6a0626 | Merge pull request 'The scenario model, the lifecycle, and what the lab actually costs' (#6) from design/scenario-underlay-detail into main | ||
|
|
e88b448145 |
The fix is real: 76x, verified. And how the lab installs on a clean machine
Snapshot 9.9s -> 0.13s. Restore 10.4s -> 0.80s. Three snapshots sharing 1.36 GB instead of costing 4.8 GB. The projected four-machine reset cycle falls from ~90s, unbounded at worst, to ~15s dominated by a boot that cannot be avoided. ADR 0029's inner-loop argument holds with copy-on-write and did not without it. The consistency matters as much as the speed: three consecutive snapshots took 0.13, 0.12 and 0.13 seconds, against a dir second snapshot that never finished. One honest counter-observation recorded: launching onto the fresh copy-on-write pool was slower, 20.2s against 14.3s, because the image had to be unpacked into a pool that had never seen it. Paid once per pool, and dwarfed by what snapshotting saves, but it went the other way. Doing the measurement produced the answer to how the lab installs on a clean machine, because both failure modes appeared while doing it. Installed is not available: the daemon was present with units disabled and no group. Issue 007. Available is not adequate, and this is worse: with the storage tooling absent everything worked and snapshots were seventy-six times slower. Nothing failed, nothing warned. That is a variant the mesh has not catalogued — its usual failure is reported success and did nothing; this is reported success and did it seventy-six times slower, which no error surface catches because nothing is wrong. So the lab verifies CAPABILITY, never installation, and refuses to run degraded rather than warning — a warning about a slow inner loop is read once and ignored forever. Prerequisites may arrive from a mesh module or from the lab's own bootstrap, and the second path is required rather than convenient: a lab installable only by a mesh cannot host the development of the mesh that installs it. The lab is the second thing installed by hand, after the node host, and for the same reason: something has to be first, and pretending otherwise produces a circularity papered over by a script nobody exercises. |
||
|
|
98bcd5cc49 |
Measure the lab's inner loop — it is too slow, for a fixable reason
The lifecycle design closed on a question that was measurable rather than arguable, so it was measured. One virtual machine on a workstation with hardware virtualisation and NVMe. Raising: the launch call returns in 3.4s, the machine is actually usable after 14.3s. The gap is a design constraint — raise must wait for the second number, because reporting the first would be transport reported as effect, which is the mesh's own recurring failure. Snapshot: 9.9s and 1.6 GB for a 1.5 GB instance. A dir snapshot is a full copy; nothing is shared. Restore: 10.4s, usable again after 20.1s. The second snapshot exceeded two minutes and never completed. That is the more troubling number: snapshot cost here is not merely high, it is unpredictable, and a loop with a variable multi-minute step is one nobody trusts. Projected to a four-machine scenario, a reset-and-rerun cycle is about a minute and a half at best and unbounded at worst, before any of the mesh's own work begins. That is too slow for an inner loop, and ADR 0029's whole argument — that making the bootstrap path the inner loop turns the least-exercised code into the most-exercised — holds only while resetting is cheap. The cause is not virtual machines. Hardware virtualisation is present and machines boot in fourteen seconds. It is that the daemon offers exactly one storage driver, dir, which has no copy-on-write and therefore no cheap snapshot. The btrfs kernel module is available; btrfs-progs is simply not installed, which is the entire reason the driver is absent. The copy-on-write comparison was deliberately NOT run, because running it would mean installing a package by hand — which the rules forbid and which would have made the measurement unreproducible. So the honest statement is that the current configuration is too slow and the likely fix is known but unverified, rather than that btrfs fixes it. |
||
|
|
a253afe020 |
Scenario lifecycle, and how two scenarios coexist
ADR 0032: a scenario is a closed address space. Every segment materialises as its own isolated link belonging to one instance, so two scenarios raised from the same declaration hold the same addresses and never meet. The declaration keeps its literal addresses and they mean what they say — allocating from a pool would have made them a fiction, so a scenario reproducing a specific topology would stop reproducing it. The constraint that follows shapes everything: the lab never reaches into a scenario over IP. It talks to machines through the virtualisation layer's own channel. If it reached them by address, the workstation would need a route into each scenario, and two carrying the same prefix would give it two routes to one destination — failing not with an error but by one scenario's traffic arriving in another. That also makes reachability an honest question. Can this machine reach that one is asked from INSIDE, by executing on the first, rather than probed from a workstation that is not on the network and whose opinion would be a different question with a misleadingly similar answer. The lifecycle itself: six verbs, of which raise and destroy are enough to be useful and the rest are what make repetition cheap. Raising is convergent rather than incremental, because a lab behaving differently from the thing it tests teaches the wrong habit. A failed raise leaves the wreckage standing. Tearing down on failure destroys the only evidence, which is backwards — a scenario that failed to raise is more interesting than one that succeeded. Snapshots are whole-scenario. Per-machine would be cheaper and wrong: the mesh keeps state spanning nodes, so restoring one machine while its peers move on produces a mesh that has never existed, and faults found there would be artefacts of the lab. Closes the declaration's open question about running several scenarios at once. |
||
|
|
e88a6df924 |
Public networks are unrelated, and routed rather than bridged
Caught in review: every public address sat in one /24, which made the three of them look like one network. They are not. The internet is a very large number of unrelated networks routing to each other, and a machine in one is many hops from a machine in another with no shared broadcast domain between them. Putting them in one prefix would have quietly made four false things true in the lab: machines resolving each other by ARP and talking directly, TTL never decrementing, broadcast and multicast crossing between them, and any two being adjacent. The third is not hypothetical. The as-is layer records that mesh names are deliberately not multicast names, after a delay and a one-node-only failure mode. A lab where the internet is one broadcast domain would let a node discover a peer by multicast that it could never discover in production, and report success — the exact false green this effort exists to prevent. So a scenario has one public segment per public NETWORK, each with its own unrelated prefix, wired together through a router and never onto a shared bridge. That is a property of how the lab wires them rather than a field anyone sets, because no correct scenario has two public networks adjacent. Addresses now spread across all three RFC 5737 ranges plus RFC 3849 /48s, chosen to look nothing like each other, and a foreign private network uses someone else's RFC 1918 range rather than a documentation one. All three examples in the document rewritten, since two of them still showed a single flat internet segment and contradicted the new rule. |
||
|
|
a873088140 |
A worked example: the whole model applied to an ordinary mesh
The shape research 004 identified — one machine with a routable address, one publicly named but behind a household connection, one stationary on that network, one that roams — written out with every field the model has, in role names and documentation addresses. It shows the ISP modem doing nothing, because in bridge mode it is a media converter: it changes the physical medium and leaves the packets alone, so it creates no IP-level fact and appears nowhere. In router mode it would be a second gateway and publishing would need a rule on both, which is the one case the model still cannot express. It shows two segments sharing one gateway declaration, which means one gateway machine, and a policy rule between them that is asymmetric because useful ones almost always are. And it shows what is deliberately absent. Research 004 recorded overlay addresses, hub election and names for exactly this topology, and none of them appear: a scenario must not state what the mesh is responsible for. Given the declaration, whether a hub is elected, whether the NATed machine's endpoint is learned, and whether the roaming machine re-forms after moving are all observed rather than arranged. The absence is the point. The run at the end moves one identity through four positions — home, foreign network, asleep, home again — against a foreign gateway whose mapping expires in 30 seconds, which is why a number is there rather than a boolean. |
||
|
|
b1874f1d0f |
Segment policy, shared gateways, and what the model leaves out
Asked whether a real setup is coverable — router, modem, access points — the answer splits, and one part was a genuine gap. Most equipment is invisible and the omission is deliberate. The test: does the device change what an IP packet can do? A switch moves frames within a segment. An access point bridges wireless clients onto one — a machine on wifi and a machine on cable are the same machine to IP. A controller configures equipment and has no packets of its own. Modelling any of them adds a fixture with no fault to catch. Two entries in that list do matter. A modem in bridge mode is a media converter and invisible; in router mode it is a second gateway, which is double NAT — expressible as nested segments, but publishing through two gateways still is not, and that is now named as the one real absence. And VLANs are segments, which exposed the gap: inter-segment policy was inexpressible. inbound: is a HOST firewall, per machine. A segmented router enforcing rules between networks is a different thing and blocks traffic regardless of what the destination thinks — a node behind such a rule cannot be reached even by a peer that knows exactly where it is. policy: states it as a fact about a pair rather than a property of either, defaulting to allowed and asymmetric by design, because the useful configuration is almost always one-directional. Segments may also share a gateway: identical gateway declarations mean one gateway machine, not two, because that is what a VLAN-capable router is — and two routers sharing an address would not work anyway. |
||
|
|
274bd3b304 |
Close the missing axes — and address family changes the model
Address family was not a field. IPv6 usually has no NAT, so a machine behind a household gateway is typically unforwardable on v4 and DIRECTLY ATTACHED on v6, at the same moment. The three positions therefore apply per family, and reachability is a property of (machine, family) rather than of a machine. The consequence is bigger than the syntax: 'can these two nodes reach each other' stops being a yes/no question. It is asked once per family, and the asymmetric answers are the interesting ones. A mesh treating reachability as one fact per node reaches a peer over one family, fails over the other, and reports whichever it tried. That distinction did not exist in the model and would have been found by a failure rather than by reading. Two fields follow from it. inbound: allow|deny became necessary because with NAT unreachability was implied by topology, while a globally routable v6 address is reachable unless something refuses — so refusing has to be sayable or v6 addressing silently implies reachability. And nat: became a list of families rather than a boolean, because a real gateway translates v4 and routes v6 and a boolean cannot say that. mapping_ttl closes the keepalive gap: a mesh holding a connection through NAT without refreshing it works perfectly until the far side goes quiet for longer than the mapping lives. segments[].mtu closes the fragmentation gap: an overlay adds a header, so a tunnel over a reduced-MTU path establishes a connection and then silently drops large packets. at: takes a list, so a multi-homed machine is expressible — which the model already implicitly required, since a border machine sits on two segments. v6 uses RFC 3849 documentation space, the exact counterpart of the RFC 5737 rule and load-bearing for the same reason. Remaining: nested forwarding and an address changing in place, both extensible when needed. Path quality stays deliberately out — it changes performance, not correctness, and modelling it makes a network simulator rather than a fixture. |
||
|
|
b944904f1a |
Audit the scenario model for generality, and fix what it found
The question is not whether the model covers our mesh but whether it can express any mesh. Audited against the axes a deployment varies along, with the standard being every property that changes how the mesh BEHAVES rather than every property a network has — bandwidth does not change correctness, MTU does. One real bug, now fixed. A segment with no gateway was read as the internet, which made an isolated network inexpressible: a LAN with no route out would have been treated as public and forced onto documentation addresses. Segments now state kind: public or private, and a private segment with no gateway is an island. A mesh spanning a site with no internet is a real topology. One modelling error, now corrected. The three positions were framed by ownership — a gateway you control versus one you do not. The axis is forwardability. Carrier-grade NAT is your own connection and is still unforwardable, so it belongs with the café network. Gateways gain forwardable:, independent of nat:, and publishing through an unforwardable one is a declaration error because that is the constraint being reproduced. Three genuine gaps recorded in priority order. Address family: cidr is implicitly v4, and a v6-only node is not exotic — a mesh that assumes v4 fails there completely rather than partially, which makes this a second world rather than a refinement. Expiring NAT mappings: without them keepalive behaviour is hoped for rather than tested, and for a mesh mostly behind NAT that is the fault that shows up after an idle night. MTU: tunnels fragment, and a smaller-MTU path establishes a connection that then silently drops large packets — the exact shape this effort exists to stop shipping. Latency and loss are deliberately out: they change performance, not correctness, and modelling them makes a network simulator rather than a fixture. Also adds a NAT primer, because the three positions are consequences of it and the document should not assume the reader already knows why a mesh dials outward and never inward. |
||
|
|
e65e5809dc |
The scenario declaration gets a real network model
forwarded: [443] was the tell. It implied a destination-NAT rule while never saying from which address, and the address is the whole point: a household's public address is what a peer records as the endpoint when a machine there dials out, and what a public name for a published machine there resolves to. It was decoration in the old shape and is load-bearing in this one. The model now names three positions a machine can be in, because they are genuinely different and the mesh has to cope with all three. Directly attached, with its own routable address. Behind a gateway you control, reachable only through a forwarded port at the gateway's address. Behind a gateway you do not control, reachable not at all, with an apparent address belonging to someone else's router that changes when the machine moves. The third is the hard one and the one that breaks reachability assumptions first. A gateway now carries three facts instead of a boolean: the parent segment, the address the world sees the network as, and whether addresses are translated — so a routed range is expressible as well as ordinary household NAT. published names the gateway it forwards through, which is how a machine on a LAN that itself has a public address is stated, and publishing on a foreign gateway is a declaration error because that is exactly the constraint being reproduced. Moving a machine between positions becomes a lifecycle operation rather than a declaration: the same identity at home, then on a foreign network, then asleep, in one run. Whether the overlay survives that and notices the endpoint changed is observed, never arranged. All three RFC 5737 ranges are now allocated a job — the internet segment, a foreign network, and a spare — with private segments kept byte-identical to production because those addresses mean the same everywhere. New open question worth having: a real gateway forgets NAT mappings after a timeout, and whether a scenario can say so decides whether keepalive behaviour is testable or merely hoped for. |
||
|
|
d64b386d03 | Merge pull request 'Hand the lab design off to mesh-lab' (#5) from handoff/lab-to-mesh-lab into main | ||
|
|
a72fea5342 |
ADR 0031 and the scenario declaration
The lab provides the underlay; the mesh builds the overlay. This is the boundary that decides whether the lab is worth having: a scenario that assigns overlay addresses, elects the hub and writes peer configuration certifies its own work — if the mesh's peering is broken, that scenario still comes up green. The most valuable thing the lab can test is exactly the part pre-building would replace. So a scenario declares what a hosting provider and a home router would provide: segments, which machine sits where at which address, what NAT is between them, which ports are forwarded, which machines are detached. It declares nothing about overlay addresses, hubs, peering, names or certificates, all of which become outcomes to observe. The declaration has four parts — segments, machines, place, snapshot — and the two scenario classes differ only in place. That is what makes one a strict subset of the other rather than a fork. Research 004's most important finding becomes a format constraint rather than a footnote: the routable segment must use RFC 5737 documentation space, because the mesh decides public versus private by matching the address, and a private range there makes the hub test as unreachable while the mesh silently never forms. A segment without behind: is routable, and a non-documentation address in it should be refused before anything is raised — ADR 0008 applied to a configuration file, since the failure it prevents has no error at all. Four things left open, including the one that matters most: a lab machine is always privileged, so the user and edge profiles have no scenario that exercises them. |
||
|
|
4d387d2998 |
Hand the lab design off to mesh-lab
Playbook 04: the design names its owner and flips to in-progress. code moves from hal to mesh-lab, and decisions gains 0029 and 0030 — the design now rests on three records rather than one. repos.md marks mesh-lab as the one target repository that exists. The other six remain the target, not the present, and saying so is the point: a map that lists repositories which do not exist is a map that will be believed. |
||
|
|
14f1c035a9 | Merge pull request 'The lab comes first, and its first scenario has no pipeline' (#4) from design/lab-bootstrap-scenario into main | ||
|
|
4a628bf3fe |
Issue 007: the instance is fixed, the class is the issue
The incus hook landed and does the post-install work — group, subordinate id ranges, both units, storage pool, bridge, default profile. Verified independently here: group exists with the operator in it, both id files carry the range, service active, pool reports CREATED. The only thing that fails is a shell whose process tree predates the usermod, which is how group membership works and not a defect. So the mechanism was never missing. Hooks are the right place and they work. The gap is narrower and worse: the hook did six things, six checks were then performed by a human by hand, and nothing in the pipeline asserted any of them. A pipeline that dispatched a hook which silently never fired would have been green in the same 48 seconds — and a hook named for a feature its module does not carry is skipped without complaint, thirteen of which were found at once in the past. The six manual checks are, almost word for word, the module's own verification: outcomes rather than steps, which is exactly the shape the lab design asks for. They currently live in a chat message. In the module they would run on every delivery to every node. Status moves to diagnosing rather than resolved, and fixed-by records the instance explicitly as the instance only. |
||
|
|
09489a298c |
ADR 0030: the repository structure, and the rule that names them
The tiers were settled and the product was named, but the repositories themselves existed only in a research sketch. That had already caused two problems. ADR 0029 makes the lab phase 0 of the migration and could not say where it lives, because no record named a repository. And the sketch contradicted an accepted record: it listed mesh-hq while ADR 0028 had decided novox/hq and explicitly rejected that name. A design resting on research is resting on something that can change without a decision. Corrected in the research too. The naming rule, which both earlier records implied and neither stated: a repository belonging to a product carries that product's prefix; a company-scoped one does not. That is why this repository is hq and the mesh's are mesh-*. Seven repositories recorded — host, substrate, control, surfaces, sdk, lab, and this one. The lab gets its own: it ships to nobody, outlives any single tier, and drives virtualisation on a workstation, which nothing else does. Inside the host it would couple development tooling to a shipped component; inside the control plane the bootstrap scenario would depend on a tier that does not exist when it is needed. Tier 4 is deliberately not decided. Whether the catalogue is one repository, one per domain or one per application stays open from ADR 0015 and is blocked on research 005 — how many repositories hold domains cannot be answered before knowing what the domains are. mesh-catalog appears in the sketch and is not decided by this record. The cost is stated rather than glossed: seven release cadences where there is one, and cross-repository changes that used to be one commit. |
||
|
|
b4904fec7e |
The lab comes first, and its first scenario has no pipeline
The lab was designed around a module under test, with a scenario being a complete mesh — forge, coordinator, cascade, verify. That is unusable for building the new mesh, because all four are tier 2 and do not exist yet. And research 009 had the sequence backwards. It placed the lab at phase B as verification of tiers already built, but tier 0 is the component that takes over a machine's packages, services and network. It cannot be developed against a machine anyone needs. The lab has to exist before the thing it will test. ADR 0029 splits scenarios into two classes. The bootstrap scenario is virtual machines, the host binary and a pinned bundle, with the verdict coming from what the host reports about the state it reconciled. The full scenario is the designed one. The first is a strict subset of the second — same virtualisation, same networking, same lifecycle, stopping before a control plane exists — so the second is reached by addition rather than rework. The consequence worth having: raising a node from nothing stops being the least-exercised path in the system and becomes the inner development loop. It also settles the runner's two jobs. Scenario lifecycle is needed immediately, because something must materialise and reset a mesh before anything can be written against it. Assertion execution waits for the full scenario. Corrects a stale claim in the design while amending it: it argued scenarios were affordable with system containers and would not be with virtual machines. ADR 0016 superseded that reasoning and the text had not followed. Issue 007: the lab's first requirement is installed and unusable. The virtualisation package is present and explicitly installed; both units are disabled, the operator is in no group, and the client reports the server unreachable. Not issue 001 again — that is an install failing while reporting success. This is an install succeeding when success was not the point. A package is files; a capability is a running service and an identity permitted to reach it, and the module model has no vocabulary for the second. |
||
|
|
1570234ac0 | Merge pull request 'Publish-safe: remove two disclosures, and record Nox as the answer to 006' (#3) from chore/publish-safe into main | ||
|
|
c0ae8dec96 |
Remove two disclosures, and record Nox as the answer to 006
Found by a full scan before making the repository public, which is the moment the public rule stops being aspirational. A module name identified a specific laptop model — hardware inventory, which is operational detail about one installation rather than a lesson that travels. Generalised. ADR 0028 named a forge username in a repository path, which the public rule forbids, and the sentence had also gone stale: the repository it described was subsequently verified empty of anything unique and removed. Rewritten to state what happened without the username. Removing a disclosure from a record is the same class as fixing a path — the rule that permits it outranks the one that forbids editing. Issue 006 gains its proposed direction: Nox works from within this repository rather than these documents being synced into the knowledge base. Better on three counts — no copy, so no drift; no fourth knowledge system, which was the original objection; always current. But it changes the promise, and the issue says so. ADR 0019 promised these documents would surface BESIDE everything else in a symptom search. An agent that must be asked is reachable, not surfacing, and the two differ in precisely the case the operational memory exists for — someone debugging an error with no reason to suspect HQ knows anything about it. The question narrows to whether a symptom search finds this content without the searcher already suspecting it. |
||
|
|
d0ec3e9373 | Merge pull request 'Retire the HAL name where it points forward' (#2) from chore/retire-the-hal-name into main | ||
|
|
93a1231e00 |
Retire the HAL name where it points forward
Skills take the hq- prefix: they are HQ process workflows, not mesh workflows, and HQ is company-scoped now. hq-new-research, hq-graduate, hq-new-issue, hq-diagnose, hq-amend-design, hq-handoff, hq-sync-constitution, hq-status. Forward-looking prose becomes Novox Mesh or simply the mesh — the root README, AGENTS.md, the 00-META README, the mission's module example, and one to-be document that addressed 'someone working on HAL'. Three categories deliberately keep HAL, per ADR 0027: The monorepo is still called hal on the forge. repos.md, every code: field and every located-in: field name a repository that exists under that name, and renaming them in prose would make them false. The as-is layer and the research that measured it describe the system that runs, and that system is called HAL. 124 modules, 9 daemons, a dead containerised node — those are observations, not intentions. Records 0001-0026 are immutable. A record says what was decided when it was decided, and no record is edited for a name. Also repoints ADR 0022's link at the renamed skill — a path fix, which the immutability rule permits, not a change of meaning. |
||
|
|
f6867d88d1 | Merge pull request 'HQ: the as-is base layer, the process, and the names' (#1) from docs/as-is-base-layer-and-process into main | ||
|
|
87f4f29cc6 |
Novox Mesh, Nox, and HQ becomes company-scoped
ADR 0027 — the product is Novox Mesh, shortened to mesh internally. HAL was never chosen: it arrived with the dotfiles repository this grew out of, it is borrowed, and it is borrowed from the canonical untrustworthy machine intelligence, which is an odd flag for infrastructure trusted with credentials. Timing is the substance of the decision, not an aside — the skeleton is not built, so renaming costs a search and replace now and a migration later. Nox is an identity of Novox, and specifically the agent of the MESH rather than of a node. Nodes keep their own identities. Nox addresses them, and a human mostly talks to Nox — which makes it the concrete form of the mission's vision: state an intent, and the mesh works out which node holds the thing. It holds no private channel. The gap this opens is recorded: ADR 0012 binds every agent to a home node, and a mesh-scoped agent has none, so the model needs extending. ADR 0028 — HQ is company-scoped, novox/hq, with the mesh as its first product. Checked rather than assumed: the company organisation already holds live projects that the mesh builds and deploys, so they are tenants rather than peers, and the mesh is the ground they stand on. There is also company work outside the mesh already, which strengthens the case and means the eventual split is closer than "some day" — so each document's scope is fixed now, in a table, making that split mechanical instead of archaeological. The folders are deliberately not restructured yet. The skeleton takes the new vocabulary: mesh-host, mesh-substrate, mesh-control, mesh-surfaces, mesh-catalog. Substrate drops to four services now that identity is a hosted workload rather than a dependency. Research 009 opens the migration, with the reframing that lowers its risk: replace the control plane, do not move the workloads. Their data never moves, so it is re-declared rather than adopted — which keeps adoption out of scope, as the lab design requires. Self-hosting is the last phase, or a failed cutover takes away the means to fix it. |
||
|
|
daf3e17c32 |
self-hosting, provisioning and delivery efforts, and the dotfiles origin
The identity provider is settled as not-substrate: the mesh does not require one, tier 2 authenticates natively, and it is a hosted service like any other. Four substrate services, not five. The tier test's second step gains the verb that matters — can the control plane START without it, not function fully without it. That verb answers the forge and the registries. They are not substrate and they are not duplicated: the control plane starts and manages nodes without a forge, it just cannot change itself. One gitea module, tier 4, and the mesh's own instance is distinguished by what it is bound to rather than by being a different module — the same answer as postgres, from the same test. It also buys a property worth having: if the forge dies the mesh keeps running. Delivery needing them is not an upward dependency, resolved the way the constitution already says to: tier 2 declares requirements, tier 4 provides implementations, the binding is data. The mechanism is provisioning, and the new idea is that the control plane is itself a consumer. Self-hosting therefore becomes a state the mesh REACHES, not a precondition. A first node comes up from pinned external artifacts and re-binds to internal providers once they exist. Today's mesh assumes the second state from the first moment, which is why the first-node path needs a script that papers over an impossibility and is the least-exercised code in the system. Made explicit, the transition is also reversible. Research 007 and 008 opened for the two areas flagged as important and complex, scoped from the weaknesses the as-is layer already documents rather than started blank. And the origin: this began as a dotfiles repository. The first two days adopt dotfiles, add per-node overrides, and introduce service symlinking with an ignore file. The flat one-directory-per-tool catalogue, linking over copying, adoption of already-configured machines, per-node overrides and the desktop modules are all inherited rather than chosen for a mesh. That is the single most useful fact for anyone changing the catalogue, it strengthens ADR 0018 — the case for links was never made for a mesh — and it explains research 005's silent fifty: dotfiles-era entries for one tool never shared a domain because they never had one. |
||
|
|
7a20358113 |
research 006: the code skeleton, and where postgres lands
A tier test as a decision procedure — five ordered questions, first match wins — so placement is answerable rather than argued. Postgres was the test case and the naive answer is wrong. Not twice, once: the control plane cannot exist without a relational store, so it is tier 1 and lives in hal-substrate/store/postgres. What differs between the mesh's own database and a project's is not the module but how that instance is brought up — pinned bundle applied by the host, versus the ordinary delivery and provisioning path. Tier is a property of the module; the bundle is a property of the mesh's own instance. The naive answer would also have made substrate reach up into the catalogue, which the dependency rule forbids. Working the test across the catalogue surfaces a third fate that neither of research 005's options covers, and it is the most common one: absorbed into the host, ceasing to be a module at all. That explains 005's one positive measurement rather than confirming it — the reachability cluster is not four modules that should be one domain module, it is four facets of one thing the host should own, expressed as modules because a module was the only unit available. Under this skeleton the overlay and firewall modules stop existing. It also partly answers the silent fifty: several are host concerns, so silence was the right signal and grouping was the wrong inference. Flags rather than settles: the identity provider is a genuine boundary case (four substrate services or five), and absorbing six concerns into a binary whose argument is that it has no dependencies is the skeleton's biggest unproven claim. |
||
|
|
00b8398d07 |
skeleton: hal-agent -> hal-host, and the two gaps the questions found
Agent is a first-class concept here — a participant, some of whom are human, holding identity and memory. Using it for the tier-0 node binary put both meanings in one document: hal-agent at tier 0, agents at tier 2. That is the anatomy-naming failure again, an evocative domain word pointing at infrastructure, and how-we-build 4 exists to catch it. ADR 0015's own title supplies the fix: the mesh brokers, nodes host, agents think. So the control plane brokers (hal-mesh), the tier-0 binary hosts (hal-host), the participant thinks (agents, untouched). hal-node was rejected — Node is the inventory aggregate, the binary is what runs on it. Recorded in the document as a near-miss rather than quietly corrected. Tiers now explained before the tree, as a boot narrative, with the point they were carrying made explicit: dependencies point only downward, that is the whole bootstrap answer, and the current mesh violates it — the database is a module, modules come from the pipeline, the pipeline needs the database. Connectivity gains the part that was missing. Naming a context says who decides, not who runs it: overlay membership is tier 0 in the host, policy is tier 2, machinery is tier 4 modules. And the host's link to the control plane deliberately does not run over the overlay, or the overlay would have to exist before a node could be told how to join it. 'tier' is this document's coinage and did not land on first reading; that is now an open question rather than settled vocabulary. |
||
|
|
b4365d8aa4 |
research 006: the mesh designed from nothing
A skeleton laid out against the stated requirements rather than derived from the current shape: four tiers, repositories at the root, and a dependency rule that only points downward. Four moves the current shape does not have. The substrate is applied by the agent from a pinned bundle, not delivered by the pipeline. That is the bootstrap circularity removed rather than worked around — the first node is the ordinary path with no control plane on the other end, which also makes it the cheapest lab scenario instead of the one nobody exercises. One agent binary with a detected capability profile — managed, user, edge. A phone becomes a capability question rather than a platform question, so it needs no second implementation. Modules declare which profiles they can land on, and an impossible assignment fails at declaration. Connectivity becomes a context. ADR 0015 names nine and none owns the overlay, resolver, firewall or ingress, while research 005 measured reachability as the only cluster in the catalogue that genuinely changes together under one intent. Gap and evidence point the same way. That is an addition to an accepted record, so it needs its own record and is not written here. Feature splits into artifact (built once per version) and part (selected per node). The conflation of those two cardinalities under one word is what makes the delivery pipeline hard to reason about. Also makes explicit in how-we-build that the main-branch rule covers this repository too. The rule already said 'without exception'; nothing was amended, so nothing is recorded. |
||
|
|
143f8ab2f1 |
research 005: which modules actually change together
The domain-grouping premise is testable, so it was tested before drawing a list. Co-change across the full history of the module catalogue, current modules only, platform namespace excluded. Nine commits in ten touch exactly one module, and 50 of 89 modules have never been edited alongside anything. Two clusters exist above that floor. Reachability holds up: proxy, resolver, firewall and VPN genuinely move together under one intent, three times in recent history. That is the shape ADR 0017 describes and the only place the measurement finds it. The provider cluster does not, and this is the finding worth having. Every multi-provider commit is a cross-cutting manifest change applied N times — feature detection, hook conventions, volume binds, network scoping. None is a change to what a database is. Merging them would not have prevented one of those commits, and the history already shows the fix that worked: verify by shape in the SDK rather than copying a script into every module. Move the concern into the machinery, do not merge the modules carrying it. ADR 0017 keeps its principle and gains a pointer to this narrowing. |
||
|
|
3f6d939930 |
Every decision is a record; the ledger is gone
papa-hq has no ledger. Its root is AGENTS.md, CLAUDE.md, README.md, every decision is a numbered record, and its graduation playbook has no path for an unrecorded decision. hal-hq now matches. The ledger's 41 entries classified as: 10 restating a record, 11 restating design docs, 15 describing how this repository works with the reasoning sitting in a README rather than anywhere citable, 3 small rules with no home, 2 superseded stubs. Mostly a copy — and a hand-maintained index, the exact pattern ADR 0022 had just rejected for the decision index on the grounds it drifted after one addition. Keeping one copy of that while removing another is not a position. It also collided by name with 02-DECISIONS/ in any directory listing. Nothing was dropped. Records 0019-0025 give the repository decisions the reasoning they never had: HQ is its own repository and is public, design has two layers, work moves through playbooks, status lives in frontmatter, issues have a front door, the numbering is the flow, HQ is the source of the constitution. 0026 records the ledger's own removal. The three orphan rules went to how-we-build, where a rule is enforced and keeps the incident that earned it — the package rule was genuinely unwritten anywhere. Two lab decisions stated only in the ledger went into the lab design. "Deliberately not decided" went to the research effort and design document each question actually belongs to. The chronological view the ledger provided is now generated from record frontmatter, which is what it was for. The cost, stated in 0026 rather than glossed: a record is more work than a table row, so the risk is a small decision going unrecorded because nobody wanted to write a document. how-we-build takes rules cheaply, which is the mitigation, not a solution. |
||
|
|
c0b35652d0 |
The numbering is the flow: decisions are 02, design is 03
papa-hq reads 01 research -> 03 decision -> 02 design. The order is a scar, not a choice: 02-DESIGN existed from its initial commit, and when adr/ was finally promoted on 2026-07-13 it took the next free number rather than its place in the sequence. By then design was too settled to renumber. hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and 02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks the process in the order it happens: research produces a decision, the decision authorises a design. 00-GENESIS becomes 00-META, matching papa's rename from the same restructure. Every path reference rewritten across documents, frontmatter, playbooks and skills. All links resolve; all 58 frontmatter blocks parse and their path fields still point at files that exist. |
||
|
|
f05e4a0dce |
Follow papa-hq's research convention; the mesh links nothing
Research efforts move from status.md to 00-overview.md with active / graduated / abandoned, matching papa-hq so the two repositories read the same way. Playbooks, skills, README and the ledger follow. Reverses yesterday's withdrawal of the symlink note in GENESIS. The note was right and the withdrawal was wrong: the intent is that the mesh creates no symlinks at all, so a founding document listing "symlinks, not copies" as a design principle does point the opposite way from where this is going, and that is a contradiction rather than a stale detail. ADR 0018 records the position, proposed. ADR 0011 stays as it is — it is the historical decision and the incident behind it is why anyone believes either record — and is superseded in intent, not edited. Its one editorial line, which called the wider reading false, is corrected to state what is actually true: centralising who may link narrowed the incident class without closing it, because a link the installer makes resolves exactly like one made by hand. The argument that kept linking was staleness. ADR 0004 removed it: every managed file is already derived and reconciled, so a copy is the natural form and a pointer into source is the shape the mesh's own model forbids everywhere else. What is not settled, and is marked open, is how staleness gets detected — which is the decision that makes or breaks it. |
||
|
|
702efca6bb |
Base layer: the mesh as it is, under the mesh as it should be
HQ held only the to-be. Every reader had to already know the system the decisions were about, and an as-is claim had nowhere to live except inside an intention. Adds 02-DESIGN/00-as-is — eleven documents written from the implementation and the operational record, not from intent, including the parts nobody would choose again. The two existing designs move under 01-to-be. Layers are declared in frontmatter and never mix: a design that ships does not move, its as-is counterpart is written, and both stand. Back-fills adr/0001-0014 for decisions taken in implementation and never recorded — the broker, the module abstraction, the mesh database, managed files, provisioning, migrations, the workspace removal, failing loudly, the constitution, application placement, linking, the employee model, the artifact, the three silos. Each marked reconstructed, dated from the history, and citing the evidence it was recovered from. The two existing records renumber to 0015 and 0016 so the ledger runs oldest first; 0017 extends 0015 to modules outside the core, principle only — the domain list is deliberately not invented here. how-we-build.md becomes the source of the mesh constitution, with a sync playbook, so the enforced copy stops being the only one that is true. Process becomes explicit: five playbooks, eight thin skills that defer to them, a repository map, and AGENTS.md with CLAUDE.md as its include. The five Observations become 04-ISSUES 001-005 where they can be owned and closed. 006 is new and uncomfortable: HQ is not indexed into the knowledge base. That claim is what decision 27 rests on, it was never checked, and the README now says so instead of repeating it. Also corrects the ADR index into something generated, the "02-DESIGN is empty" claim, the VISION.md pointer that did not survive the repo split, and a note asserting the symlink rule was contradicted — it was a misreading; the rule forbids hand-made links, the installer links by design. |
||
|
|
cf9357e8e9 |
HQ — the mesh's own documentation
What the mesh is, what it is becoming, and why. Implementation lives in the code repositories; the reasoning lives here. 00-GENESIS mission, engineering context, effect, and the rules that hold 01-RESEARCH investigations, before they harden into design 02-DESIGN the authoritative specification adr numbered decisions — what was chosen, and what was rejected DECISIONS.md the ledger: every decision, in the order it was taken Written for a reader who is not its author and has no access to the mesh it describes. Addresses use the documentation ranges of RFC 5737 and RFC 1918; nodes are named by role. Single initial commit by intent. The prior history came from a private repository and carried operational detail — a routable address identified as a VPN hub, real domain names, a hosting provider — which sanitising a tip commit would not have removed from the log. |