94229d95ebcfbfcd5362c6458dcc0714b2714c8a
24
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
94229d95eb |
The installation, written out in full
Every step from a bare machine to a mesh that maintains itself, in three phases, with each step named as the installer prints it. The point of writing it out is the shape it exposes. The installer owns twelve steps and ends at a mesh that RUNS. Seven more turn that into a mesh that WORKS — the shared base, a store that is a provider rather than the control plane's own memory, the catalogue, the replay of what was built before the catalogue existed, the control plane rebuilt through the module path, the private network with the node actually placed on it, and the packet filter. None of those seven is the installer's. They are things somebody types, which is why a test had to be written to discover they were missing. Machines arrive last, in phase three, because a machine joining a mesh that cannot build anything proves enrolment works and nothing else. And five things that are not yet true are named rather than implied: phase two is manual, a second machine cannot pull what the mesh built, ADR 0014 assumes a private package registry that genesis has not installed when the first build needs it, the host agent does not survive a reboot, and nothing can contradict a claim that a machine was installed this way. Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx |
||
|
|
bc271d4de0 |
A worked guide: one module, four capabilities, four languages
What the SDK contains, answered by exclusion as much as by inclusion. It is the protocol and nothing else — no configuration loader, because configuration arrives as files the mesh wrote; no API clients, because a Plex client changes when Plex changes and that has nothing to do with any other module; no storage, HTTP or logging, because the language has those. The test for anything proposed is ADR 0039's: does editing it recompile unrelated modules, and does it change often. Both, and it stays out. Then the worked module: events in TypeScript, tools in Go, a provisioner in Rust, a scheduled job in Python. Four artifacts, four toolchains, four processes, one module — and each part is an ordinary project in its language depending on the mesh SDK the ordinary way, so a laptop resolves what a build resolves. And publishing a package as a module capability, which makes the SDK unspecial: it is simply the first module that published a library. A Plex client belongs to the Plex module because that is the only thing that knows when Plex changed. Three things left open rather than papered over: which registry (the catalogue holds verdaccio and a forge usually serves one too, and nothing says which is ours), who may publish (a credential that does not exist), and what a range means in a mesh where everything else is pinned by digest — a mesh that can rebuild a commit and get a different library is a real change, and should be decided rather than arrived at. Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx |
||
|
|
b637fe106f |
The module protocol, specified per capability
A floor every implementation needs and three capabilities independent of each other, so an SDK can implement the floor and events and be a real thing rather than an unfinished one. Written as a specification, which means it says what is required rather than how anything is arranged — and says plainly where it describes behaviour that is not yet true. Three places it does: x-causation-id and x-schema are specified and emitted by nothing; the Go side writes four headers and the TypeScript side declares six. A module may serve tools and may not call them, because a caller needs a reply queue its account may not declare. And the two implementations disagree about what a grant carries — in TypeScript consumer is the module, in Go it is the node and the module is From. One word, two meanings, in two halves of one mesh. Naming those in the specification rather than leaving them for conformance to discover, because a specification that only described what already works would have nothing to say about the things most likely to break. Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx |
||
|
|
5d1e6d0b09 |
Design: building a module, and why the recipe cannot always be a Dockerfile
12-a-module-repository says what a module may build and where it goes. Nothing said how a build is MODELLED, and the model is the problem: a recipe is implicit, singular and always a Dockerfile; a toolchain is not modelled at all, arriving as two build arguments the module hand-writes; a language is not a concept; and an archive is declared in the manifest and refused by the builder. The cost is measurable rather than theoretical. Adding a module with its own code means repeating an incantation - two ARG bases, a specific working directory so the SDK resolves upward, the compiler invoked by absolute path because the usual symlink is resolved away when the base is assembled, a second stage, an env var naming the entrypoints. Most of the catalogue is unconverted, and two conversions done in one session were each wrong twice with a working example open. So: recipe becomes explicit with three kinds, and toolchain becomes derived from a declared language rather than written by every author. A Dockerfile stays, and stops being compulsory - it is right for software needing a particular base and wrong for "compile my module's code", which is the same operation every time. The cost is stated before it is chosen: every language is permanent, and the contracts are already expressed twice - Go structs and TypeScript types kept in step by hand. A second language makes that drift. So language-neutral contracts come first, or the drift gets worse while hiding. Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx |
||
|
|
3d939b5c77 |
Describe how a mesh is raised, because only a test did
The one complete account of standing a mesh up was an integration test, and a fixture is free to invent what it needs — which is how a registry that exists in no production hid two faults for as long as the lab existed. Written from what the installer does, not what it should do: genesis and joining are separate moments, the lab runs the installer rather than describing installing, and three things that are not true yet are named rather than glossed, including one rule nothing checks. Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx |
||
|
|
36d342f176 |
What a module must be able to say, measured against 127 that exist
Every manifest in the system being replaced was read and every key counted, then set against what the new one can express. Three findings worth more than the table. **The most-used key was already covered and I expected a gap.** Depending on another module — 65 manifests, the commonest thing any of them says — is a requirement naming a module, which already means that module rather than anything providing the name. **The largest real gap is tool servers: 56 modules, over half.** A module can already run one; what is missing is anything saying it offers tools. That is plausibly a provision rather than new vocabulary, which would need nothing added — not yet decided, and recorded as undecided. **The gap most worth closing is health, at seven modules.** The mesh knows a container is running, which is not whether it answers, and this project has paid for that distinction twice. An action with a verify is exactly the right shape and may not arrive over the link, so a module cannot declare one. Two things are missing deliberately and say so: stage hooks, because the link may not carry an action and a module needing setup ships a program; and flavours, retired in favour of claims. Config merging is missing and should stay missing. A mechanism that understands TOML gets asked for YAML, then INI, which is how the thing being replaced became unholdable. Also records what the survey found that is not about coverage: manifests that had stopped matching what was actually brokered, one fact derived in two places giving two answers, and a live listing returning credentials in plaintext. |
||
|
|
cb1954e7b5 |
Accept 0024, and rewrite the work breakdown around what is actually being done
**0024 accepted.** Model access was decided, built, and proven in the lab, and two design documents rest on it; only the status had never moved. The gate is green again. **The work breakdown rewritten.** It planned a decomposition of the existing system in place — extract contexts, declared features, shrink the shared library. That is not the work. A replacement is being built beside it, and only the old Phase 0 survived contact with reality, so the one document meant to say what happens next was describing a system being retired. Now ordered by what "modules move across one at a time until the old registry is off" actually requires: - Phase 0 is marked done against the twenty-two lab assertions, **and carries its own limitation**: every module exercised was written to test the mechanism, so the vocabulary was shaped by its own fixtures. - Phase 1 is the vocabulary gaps found by asking what real modules need — an object-store provision, a session as a licence consumer, a network shape with ordering, public certificate issuance. - Phase 2 is one module, then a week of running it, because the point of going first is to find what Phase 1 missed. - Phase 3 picks modules that each prove something the first did not; the mail system is last because it is the one that may send work back into the declaration language. - Phase 4 is switching the registry off, named as a phase so it is not mistaken for the goal. Keeps the rules of engagement unchanged — they were about how work is done, not what it is — with one addition: stop and ask before anything that touches a machine outside the lab. Adds a section on keeping the list true, since the document it replaces was wrong for weeks and nothing said so. A claim here is counted, not reasoned, and a phase is done when the lab says so. |
||
|
|
3c6c16abdf |
The mesh has a session of its own, and it is the node session's mechanism
A session for the mesh itself, addressed as the mesh, differing from a node's in exactly three things: the context it starts in, its engram, and its licence binding. Not a new kind of agent — the same mechanism pointed at a different root. Two implementations of one mechanism drift, and the vocabulary collision 0001 exists to undo began exactly that way. It runs on the control-plane node, and the reasoning is easy to get backwards: not "the important agent on the important machine", but that this node is already the one place excepted from "compromise of a node is compromise of that node". Placed anywhere else it would create a second such place. It is an addition to per-node messaging and never a replacement. 0001 holds that losing the control plane costs change, not operation — and a mesh whose only conversational surface lived there would lose the ability to ask anything while every machine kept running perfectly. Writing it up exposed that the node session's setup was never designed at all. 0004 gives behaviour and stops: nothing said how a session starts, where its context lives, or how a broker message becomes a prompt. That gap was invisible until something had to be built *like* a node session. 15-the-agent-session.md covers both as one mechanism. It also makes "a consumer that is not a machine" undeferrable. The control-plane node now hosts two sessions that must hold different licences, and a per-machine binding cannot express that at all. Noted in 14-model-access.md against the gap it was already recorded as. Also completes the to-be index, which stopped at 10 and omitted four documents. Pre-existing broken ADR references in the older rows are left alone rather than guessed at. |
||
|
|
333356cff3 |
Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when things happened to be decided, which after consolidation is fictional anyway since record 5 alone folds decisions taken across a week. Concretely wrong before: the domain statement sat at 8, after five engineering rules; the constitution was scattered across 5, 12 and 17; the tiers landed at 15, 16, 21 and 22 with process records in between. Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what runs on them and how it gets there (9-10), how it is built (11-16), how it is checked (17-18), how we work (19-23). Two things made this safe rather than free. It is a permutation, not a compaction, so the renames go through temporary names -- otherwise two files want one slot and one is lost. And the reference rewrite is a single simultaneous pass, because almost every number moved into a slot another number was vacating; replacing one at a time would have cascaded and pointed things at the wrong record while still resolving. Verified: 284 [ADR NNNN](path) links across the repository, all with matching text and target. The ordering principle is now stated in 19 rather than left implicit -- the repository already said "the numbering is the flow" about its folders, and there was no reason for the records to be the exception. |
||
|
|
e1febe8e0f |
Renumber the records 1 to 23
The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18, 19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only the archaeology of what used to be there. Renumbered contiguously. Renames run in ascending order, so every target number is already free and no two files ever collide. The reference rewrite is one simultaneous pass rather than a sequence of replacements. Numbers moved into slots other numbers were vacating -- the node host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time would have cascaded and silently pointed things at the wrong record. Seven plain-text references survived the merges as prose rather than links, naming records that no longer existed: the enrolment token, the link boundary, what a declaration is, reachability, the repository structure. Each mapped to the consolidated record that now holds it. Verified rather than assumed: every [ADR NNNN](path) link now has matching text and target, checked across the whole repository, and the checker passes. Frontmatter `consolidates:` lists dropped -- they named records that are gone, and each consolidated record already says in prose what it absorbed. |
||
|
|
77f3a4cea7 |
Consolidate: 65 decision records to 23
Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.
the node host 8 -> 1 applies not decides, depends on nothing,
per operating system, root service, the
launcher, episodic, what a declaration is,
actions from the bundle only
a node and how it joins 4 -> 1 what a node is, joining, the link as
security boundary, the enrolment token
modules and the graph 7 -> 1 everything is a module, no domain modules,
three edges, provisioning, the core library
substrate and control 6 -> 1 the test, seven contexts, one control plane,
plane the authority is not a database, the named
products, the pinned bundle
connectivity 3 -> 1 a route is a grant, reachability declared,
filter rules
delivery 5 -> 1 reconciliation not a pipeline, artifacts,
the three silos, a failed step, the verdict
the lab 5 -> 1 (earlier)
how this repository 10 -> 1 (earlier)
works
Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.
The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.
The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
|
||
|
|
5e83ac2c22 |
Consolidate: 65 decision records to 52
Jochen: a normal application has 3-5 ADRs, maybe 10 for a large one, and we are at 65. Fair, and the cause is mine -- I recorded every FINDING as a decision rather than every fork in the road. Two merges, both cases where one decision had been split across many records because it was taken over several days rather than at once. 0019 absorbs ten records about how this repository works: what it is and that it is public, the folder flow, the two design layers, the issue front door, status in frontmatter, playbooks, the naming rule, the product name. Those were never ten decisions -- they were one, seen from ten angles as the repository took shape. 0016 absorbs the five about the lab: a node is a virtual machine, a router is scenery, a scenario declares the underlay, a scenario is a closed address space, and the two scenario classes. Same pattern -- one design, split by the order it was worked out in. The consolidated 0019 also raises the bar for what earns a record, since that is what produced 65: a record is warranted when there is a genuine fork -- a direction reversed, an alternative that will be proposed again, something contested. A finding is not a decision, and a bug is certainly not. Everything else belongs in the design document where the reasoning is actually read. The checker earned its place here. Deleting nine records left 13 dangling links across the repository and it named every one, including in AGENTS.md. Nothing was found by reading. Remaining clusters worth the same treatment: the host (8 records), delivery (5), modules (6), connectivity (4), substrate and control plane (4). That would be 52 down to roughly 30. |
||
|
|
10365f2eae |
Consolidate the design layer: one place per topic
Jochen: a jungle of specs that slightly contradict or patch each other, and what matters is a working state rather than history. Both are fair and both are mine. Measured rather than assumed. 05-the-node-host and 09-the-node-lifecycle both covered enrolment, the install commands, the unit file, the launcher and reconcile -- I wrote 09 without taking anything out of 05, so the same things were said twice and could drift apart. Split by what each document IS. 05 is the component: what the host is, its parts, the declaration vocabulary, the build order, how it is verified. 09 is what happens to it: install, enrol, run, upgrade, retire. The whole "The process" section left 05, and the unit file moved to 09 where installing is described. 05 goes from 338 lines to 245 and now points at 09 rather than restating it. 09 also carried a 105-line "Resolved" section -- six mechanisms framed as "these were open and here is the answer". The content is needed; the framing is history, and history is what makes a document read as a changelog rather than a description. Renamed to what it actually is and the was-open phrasing removed. Also added 10-delivery.md, which did not exist: four accepted decisions -- 0054, 0063, 0064, 0065 -- had no design document at all, which is the specific reason the delivery picture felt scattered. It is now one document covering modules, the three edges, the core library, and how a change becomes a running thing, with a table of what each property is designed against and what must exist before it can be built. |
||
|
|
2204b01909 |
Design the node lifecycle end to end
The host was described as a component and never as something that runs for years on a machine somebody else also uses. 09 covers every state a machine can be in and every transition between them. Four states: unmanaged, hosted, enrolled, disconnected. Only the last two are nodes, and they are the same node in two situations. `hosted` -- the host installed but never told which mesh it belongs to -- had no name before and is where a machine sits between the two adoption commands. Things that were unclear and now are not: The first node walks the same path in an unusual order: reconcile from the bundle, the control plane it just raised issues a token, enrol against it. Its specialness lasts two commands. A side effect worth having -- enrolment is exercised on node one, rather than being written and first used on node two. Enrolment reports profile and inventory BEFORE the control plane decides anything. The profile is the input to that decision, not a diagnostic; the control plane cannot decide what a machine should run without knowing what it can run. Rebooting mid-apply is safe by construction. The store records each resource after it worked, so a host that dies half way through comes back and applies the rest. The rule that stops the host lying about what it did also makes it crash-safe. Retiring splits in two. Graceful is a final empty declaration. A node that is gone will reconcile its last declaration forever -- the honest consequence of making disconnection ordinary. The answer is not to make the host expire but that the node holds nothing that outlives revocation: every grant is a per-node credential revoked at the provider. A lost node keeps running and stops being able to reach anything. Said plainly rather than implying the mesh can switch a machine off, which it cannot and should not. Losing the store is quiet and permanent, so it gets its own section. The host re-enrols and re-applies fine; what does not come back is removal, because resources it no longer has a record of become unowned and sit there indefinitely. Also corrects 0057, which said the mesh must not upgrade the host at all. That conflated two acts. Replacing the binary is safe -- Unix keeps the running inode. Stopping the unit is not. So the host may apply a package naming itself, and restarts by finishing its apply and exiting cleanly, letting the supervisor start it on the new binary. It never asks the service manager to restart it. That makes a fleet-wide host upgrade an ordinary declaration, which the first draft gave up on. 0057 remains proposed. |
||
|
|
e1f4c7d9e0 |
Approve 0054-0056, apply them, and fix the two smaller findings
0003 is now superseded by 0056. Nothing is left proposed. Applied: - 06 corrected from ten contexts to seven plus the api, each row now stating why it passes the more-than-one-node test. work, knowledge and stream are named as mesh-hosted rather than dropped; `ai` folds into config; `record` is deferred explicitly rather than listed. Its frontmatter now cites 0055. - how-we-build §4 amended per 0054, and the derived page republished by playbook 05. The sync found the drift the playbook exists to catch: the published §4 and the source did not say the same thing. The source said "four accidents, not four boundaries"; the published page said "one intent expressed four times", and only the published page carried the scope caveat. Same rule, two texts, already diverging. Verified the republish by reading back -- the new rule is present and the old section's body returns nothing -- rather than trusting the success message. The two smaller findings: - 0051 separated the transport identity from the declaring authority. It said the token carries "an address" and "the identity to expect" without saying what the node dials. It dials the broker, so pinning only that would make the control plane's authority transitive and let a compromised broker forge declarations -- which, since the host applies whatever the link delivers, is the whole machine. The token now carries four things, and declarations are signed and verified per declaration. Cost recorded: rotating the signing identity is fleet-wide. - 0026 no longer restates 0022's rule about generated views. 0022's own words are "prose does not restate status; one place, and two is one too many", which is what 0026 was doing to it. |
||
|
|
ef5dd0751b |
Approve 0049-0053; drop a to-be item superseded by ADR 0044
The 'domain grouping' item cited ADR 0017 as live guidance. 0044 superseded it -- there is no domain module to group into, so there is no domain list to settle. |
||
|
|
4e80820e2f |
Design connectivity in full: overlay, resolution, exposure, filtering, certificates
Written as one document because the five are one design. They share inputs, they must agree, and every one of them today is computed in a different place by a different module from a different copy of the same facts. The through-line is that none of the five can be answered by a machine alone, so all five are decided centrally and delivered as `file` resources. That costs no new host vocabulary and removes both remaining direct database connections from nodes -- wireguard and traefik are the only two, and both are connectivity. Three decisions fall out, all proposed: 0050 -- reachability is declared, not inferred from an address. The RFC1918 regex is wrong for carrier-grade NAT (100.64/10 tests as public, so an endpoint is written to an address nothing can reach), wrong for IPv6, and wrong for a routable address behind a closed firewall. The lab needing TEST-NET-3 to satisfy the regex is the same bug from the other side. Also kills hub election by address prefix, which fails silently and makes renumbering an outage. 0051 -- the enrolment token carries where the mesh is and how to recognise it. Closes two circles with one mechanism: verifying the mesh needed the CA, and obtaining the CA meant trusting whoever handed it over; and a node had to reach the mesh before it could resolve any mesh name. An address plus a fingerprint, carried out of band, resolves both -- and closes the CA question 0049 deferred. 0052 -- a filter rule names its source. `scope:` is declared in five manifests, is part of no rule type, and is referenced by no code, so those manifests appear to restrict ports and restrict nothing. Removed rather than implemented; the general fix is refusing unknown keys, which the host already does and manifests do not. Also corrects two claims in 0049 asserting wireguard was already handled. Research 006 says both modules still reach upward; neither is. |
||
|
|
60aea14935 |
Define the substrate, and answer 006's four-or-five conditionally
Same gap as the control plane: load-bearing and unpinned. The substrate is what the control plane CONSUMES AND CANNOT GRANT ITSELF. Every module needing a database asks provisioning for one; the control plane needs one too and cannot ask itself, because it is not running yet. That circularity is not an awkwardness to work around — it is the definition, and anything on the wrong side of it must be raised by the bundle the host carries. Which answers 006's open question in the honest form rather than with a number. The identity provider is substrate only if the control plane DELEGATES authentication — then it cannot serve anybody before the provider exists and cannot grant itself a client. If it authenticates natively, the provider is an ordinary hosted service. So the count follows from a decision not yet taken, and asserting four was asserting that decision. The test also rules out the tempting wrong answer: an identity provider, a mail server and an analytics service are all infrastructure by any ordinary reading, and none are substrate, because the control plane starts and runs without them. Important is not the test. Records why the bundle is pinned by hand — it is applied when no mesh exists, so nothing can resolve a version or ask a registry — and why it must be self-contained, which makes it an artifact built on a machine with a network for a machine that may have none. |
||
|
|
148395ca54 |
Define the control plane, which was used 79 times and defined nowhere
Nineteen files, seventy-nine mentions, no definition. That is how-we-build §5 failing on this repository's own vocabulary — ubiquitous language is checked, not assumed. The definition, and it is not arbitrary: the control plane is everything that needs to know about MORE THAN ONE NODE. It follows from ADR 0037, which has the host applying rather than deciding precisely because deciding needs knowledge the machine does not have. So the line falls exactly there — writing a file is the host's, choosing which nodes run the store is the control plane's, and anything a single machine could answer alone does not belong here at all. That last consequence is worth having: putting a single-machine concern in tier 2 is a mistake the tier rule will NOT catch, because the dependency direction stays correct. Also states what it is not — not the thing that changes machines, not a surface, not the substrate, and not privileged on a node beyond what the declaration vocabulary allows. And the property that makes tier 2 unlike the others: it is itself a consumer, with the same requirements as any module, which is the circularity the bundle exists to resolve rather than hide. Scoped deliberately: this defines the term and does not design the contexts inside it. Ten is the skeleton's claim rather than a settled list, and research 006 still asks whether the record belongs here or in the substrate. |
||
|
|
b9facf9375 |
Design the node host
Playbook 02 step 3, on four recorded decisions. Tier 0 has one job — apply declared state on this machine — and the six absorbed concerns are instances of it, not additions to it. Specifies the six parts and what each owns, and the two properties that make apply trustworthy rather than merely present: every applier reads back, because setting a value is not evidence the value took; and what was applied is recorded after it works, never before, because a failed apply leaves the machine wherever it reached and nothing must claim otherwise. Build order is staged so each stage is verifiable in the lab before the next exists. Stage 1 is profile and inventory — no control plane, no declarations, no network — and it is deliberately the smallest useful thing, because `place:` has nothing to place and the lab therefore raises empty machines. Stage 1 ends that, and every later stage is tested by a lab that already works. Stage 2 is the one that could invalidate the tier boundary: whether one host can raise the substrate alone is Move 1's assumption and has never been proved. Every decision the design rests on is given the test that asserts it, per 0034 — including the dependency-direction lint, which is what makes "the host never queries the mesh database" a rule rather than an intention. Six things left open and named, including the one that host-size.md could not measure: zero dependencies, but still six vocabularies. |
||
|
|
e88b448145 |
The fix is real: 76x, verified. And how the lab installs on a clean machine
Snapshot 9.9s -> 0.13s. Restore 10.4s -> 0.80s. Three snapshots sharing 1.36 GB instead of costing 4.8 GB. The projected four-machine reset cycle falls from ~90s, unbounded at worst, to ~15s dominated by a boot that cannot be avoided. ADR 0029's inner-loop argument holds with copy-on-write and did not without it. The consistency matters as much as the speed: three consecutive snapshots took 0.13, 0.12 and 0.13 seconds, against a dir second snapshot that never finished. One honest counter-observation recorded: launching onto the fresh copy-on-write pool was slower, 20.2s against 14.3s, because the image had to be unpacked into a pool that had never seen it. Paid once per pool, and dwarfed by what snapshotting saves, but it went the other way. Doing the measurement produced the answer to how the lab installs on a clean machine, because both failure modes appeared while doing it. Installed is not available: the daemon was present with units disabled and no group. Issue 007. Available is not adequate, and this is worse: with the storage tooling absent everything worked and snapshots were seventy-six times slower. Nothing failed, nothing warned. That is a variant the mesh has not catalogued — its usual failure is reported success and did nothing; this is reported success and did it seventy-six times slower, which no error surface catches because nothing is wrong. So the lab verifies CAPABILITY, never installation, and refuses to run degraded rather than warning — a warning about a slow inner loop is read once and ignored forever. Prerequisites may arrive from a mesh module or from the lab's own bootstrap, and the second path is required rather than convenient: a lab installable only by a mesh cannot host the development of the mesh that installs it. The lab is the second thing installed by hand, after the node host, and for the same reason: something has to be first, and pretending otherwise produces a circularity papered over by a script nobody exercises. |
||
|
|
a253afe020 |
Scenario lifecycle, and how two scenarios coexist
ADR 0032: a scenario is a closed address space. Every segment materialises as its own isolated link belonging to one instance, so two scenarios raised from the same declaration hold the same addresses and never meet. The declaration keeps its literal addresses and they mean what they say — allocating from a pool would have made them a fiction, so a scenario reproducing a specific topology would stop reproducing it. The constraint that follows shapes everything: the lab never reaches into a scenario over IP. It talks to machines through the virtualisation layer's own channel. If it reached them by address, the workstation would need a route into each scenario, and two carrying the same prefix would give it two routes to one destination — failing not with an error but by one scenario's traffic arriving in another. That also makes reachability an honest question. Can this machine reach that one is asked from INSIDE, by executing on the first, rather than probed from a workstation that is not on the network and whose opinion would be a different question with a misleadingly similar answer. The lifecycle itself: six verbs, of which raise and destroy are enough to be useful and the rest are what make repetition cheap. Raising is convergent rather than incremental, because a lab behaving differently from the thing it tests teaches the wrong habit. A failed raise leaves the wreckage standing. Tearing down on failure destroys the only evidence, which is backwards — a scenario that failed to raise is more interesting than one that succeeded. Snapshots are whole-scenario. Per-machine would be cheaper and wrong: the mesh keeps state spanning nodes, so restoring one machine while its peers move on produces a mesh that has never existed, and faults found there would be artefacts of the lab. Closes the declaration's open question about running several scenarios at once. |
||
|
|
a72fea5342 |
ADR 0031 and the scenario declaration
The lab provides the underlay; the mesh builds the overlay. This is the boundary that decides whether the lab is worth having: a scenario that assigns overlay addresses, elects the hub and writes peer configuration certifies its own work — if the mesh's peering is broken, that scenario still comes up green. The most valuable thing the lab can test is exactly the part pre-building would replace. So a scenario declares what a hosting provider and a home router would provide: segments, which machine sits where at which address, what NAT is between them, which ports are forwarded, which machines are detached. It declares nothing about overlay addresses, hubs, peering, names or certificates, all of which become outcomes to observe. The declaration has four parts — segments, machines, place, snapshot — and the two scenario classes differ only in place. That is what makes one a strict subset of the other rather than a fork. Research 004's most important finding becomes a format constraint rather than a footnote: the routable segment must use RFC 5737 documentation space, because the mesh decides public versus private by matching the address, and a private range there makes the hub test as unreachable while the mesh silently never forms. A segment without behind: is routable, and a non-documentation address in it should be refused before anything is raised — ADR 0008 applied to a configuration file, since the failure it prevents has no error at all. Four things left open, including the one that matters most: a lab machine is always privileged, so the user and edge profiles have no scenario that exercises them. |
||
|
|
c0b35652d0 |
The numbering is the flow: decisions are 02, design is 03
papa-hq reads 01 research -> 03 decision -> 02 design. The order is a scar, not a choice: 02-DESIGN existed from its initial commit, and when adr/ was finally promoted on 2026-07-13 it took the next free number rather than its place in the sequence. By then design was too settled to renumber. hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and 02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks the process in the order it happens: research produces a decision, the decision authorises a design. 00-GENESIS becomes 00-META, matching papa's rename from the same restructure. Every path reference rewritten across documents, frontmatter, playbooks and skills. All links resolve; all 58 frontmatter blocks parse and their path fields still point at files that exist. |